AI Agent Evaluation · New York, NY
New York deploys AI under examiners' eyes: model-risk teams fluent in SR 11-7, a local law governing automated hiring tools, and legal workflows where an unfaithful citation is a professional hazard. Evaluation here is not a quality nicety; it is the evidence layer that lets regulated institutions say yes.
We build evaluation systems for LLM applications and agents: representative and adversarial test sets from your real traffic, faithfulness harnesses with calibrated judges, trajectory testing for multi-step agents, and documentation written in your risk team's vocabulary.
Tell us the system and who has to sign off on it.
Model-risk teams engage when artifacts look challengeable: intended use, test design tied to harms, thresholds with sign-offs, drift monitoring, and change management. The generative translation (behavioral suites, adversarial sets, trajectory checks) is our job; the vocabulary stays theirs, and validation goes faster for it.
Faithfulness is measured by claim decomposition against your own documents, with LLM judges calibrated to attorney-graded samples before their scores count. Claim-support and fabrication rates by document type turn 'seems accurate' into a number a partner can gate releases on.
Hiring-adjacent tools get honest boundary treatment: the Local Law 144 bias audit belongs to independent auditors, while we build the instrumentation, continuous disparate-impact dashboards, and scope documentation that make audits tractable and legal exposure visible to counsel.
Agents earn production credentials through trajectory gates: tool-call correctness, permission discipline, recovery behavior, and stopping rules, replayed in sandboxes where misbehavior is free. The destination being right does not excuse the path being dangerous.
The standard build, tuned for New York's regulated deployments.
Representative, edge-case, and adversarial suites built from your documents and logs, labeled by the people whose judgment counts.
Claim-level verification against cited sources with judges calibrated to expert grades, reporting support and fabrication rates by class.
Sandboxed replays scoring tool calls, permissions, efficiency, recovery, and stopping discipline before any production credential.
Intended use, limitations, thresholds, monitoring, and change management written for model-risk review rather than translated after it.
Decision-relevant logging and continuous disparate-impact dashboards that make independent audits and counsel's scoping tractable.
Regression suites on every prompt or model change, production sampling with drift triggers, and ownership handed to your team.
New York institutions do not adopt AI on vibes: banks route it through model risk, firms route it through professional-responsibility instincts, and anything touching hiring routes through counsel. The evaluation layer is what converts a promising system into an approvable one, and it has to produce evidence those reviewers recognize as evidence.
The same machinery pays operational dividends: calibrated judges and CI gates catch regressions before clients do, and drift monitors catch the quiet degradations that erode trust without a headline.
We work with New York teams remotely, in Eastern hours, with initial harnesses typically standing in two to three weeks.
Tell us the application, the reviewers it must pass, and the harms that matter. We reply within one business day with a scope and a fixed price.