AI Agent Evaluation · San Francisco, CA
San Francisco ships AI products faster than anyone, and the ones that keep winning share a discipline: they measure. Eval sets built from real traffic, judges calibrated against human standards, regression gates that catch the Friday-night prompt tweak, and model swaps treated as routine experiments rather than acts of faith.
We build that machinery: first eval sets mined from your logs, LLM-as-judge calibration with provenance, tiered CI gates, trajectory testing for agents, and the eval-driven habits that make iteration safe at SF speed.
Tell us the product and what quality means for it.
First eval sets come from logs, not imagination: clustered production traffic sampled to mirror reality, failure cases from tickets, and 150 to 300 items labeled by the person whose taste is the product. Small, fresh, and maintained beats large, imagined, and stale every quarter of the year.
LLM judges scale measurement only after calibration: human grades first, judge tuned to demonstrated agreement, scores shipped with provenance and re-verified monthly. A judge score without calibration is a vibe wearing a decimal point, and investors are learning to ask.
CI gates earn their keep by being tiered and humane: smoke suites in minutes on every PR, full runs nightly with per-slice deltas, statistical thresholds that fail on signal rather than sampling noise, and an override path so the gate survives its first urgent Friday.
Model swaps become quarterly opportunities once the harness exists: per-slice comparison, cost and latency beside quality, canaried rollout, boring rollback. The eval set compounds into the one moat an AI product genuinely owns.
The standard build, tuned for fast-moving product teams.
Clustered, proportionally sampled, failure-enriched suites labeled with your quality owner, sized to maintain rather than admire.
LLM judges tuned to demonstrated human agreement, scores carrying their calibration numbers, monthly sampled re-verification.
Smoke on every PR, full nightly, adversarial weekly, with noise-aware thresholds, per-slice diffs, and a documented override.
Sandboxed replays scoring tool calls, recovery, efficiency, and stopping discipline for anything that acts rather than answers.
Per-slice candidate comparison, three-dimensional decision dashboards, canary plans, and rollback that stays boring.
Sampled scoring of live traffic, slice-level alerts, and the feedback loop that keeps the eval set growing with reality.
SF product teams change prompts daily, swap models quarterly, and compete against last month's batch; the only sustainable way to move that fast is measurement that moves faster. Evaluation here is not a compliance artifact, it is the product discipline that separates compounding teams from thrashing ones, and increasingly the diligence question that separates funded ones from passed-on ones.
We build the machinery and the habit: harnesses your team owns, runbooks that make swaps routine, and an eval culture that survives our departure, because rented discipline is not discipline.
We work with SF teams remotely, in Pacific hours, with first harnesses typically standing in two to three weeks.
Tell us the product, the traffic volume, and what quality means to you. We reply within one business day with a scope and a fixed price.