AI Agent Evaluation · Boston, MA
Boston deploys AI where the stakes are clinical: summaries a covering physician relies on, portal features a patient trusts, research tools a trial depends on. Evaluation here is the safety case: clinician-graded ground truth, escalation suites that gate at 100 percent, and hallucination rates broken out by the error classes that actually differ in danger.
We build evaluation systems for clinical and life-sciences AI: clinician-efficient grading workflows, refusal and escalation batteries, claim-level hallucination measurement, and the evidence layer that makes governance committees and counsel comfortable.
Tell us the feature and who it could hurt if wrong.
Clinician hours are the scarce resource, so the workflow engineers around them: drafted candidate sets from de-identified cases, mechanical pre-screening, focused grading blocks with rubrics in clinical language, and calibrated judges carrying volume after the standard is set. Two hundred graded items costs a clinician about a workday, and the set compounds from production corrections thereafter.
Refusal and escalation behavior gets its own battery because it shifts silently with every upstream change: hard refusals, mandatory escalations, and graceful boundaries, tested through adversarial phrasings and multi-turn pressure, gated near 100 percent where a miss ends programs.
Hallucination becomes a managed quantity through claim-level verification with a clinical error taxonomy: unsupported, contradicted, and omitted classes reported per note type, because fabricated medication lines and missing allergies are different emergencies than imprecise social history.
The SaMD boundary rewards early evidence either way: rigorous evaluation documents why an administrative feature stays administrative, and becomes the regulatory foundation if the roadmap crosses. Counsel rules; our evidence makes the ruling tractable.
The standard build, tuned for clinical and life-sciences stakes.
De-identified case sampling, mechanical pre-screening, and focused grading workflows that buy ground truth for hours, not months.
Red-flag, adversarial-variant, and multi-turn erosion suites with near-100-percent gates, re-run on every upstream change.
Atomic-claim verification against source records with a clinical error taxonomy, reported per class and per note type.
LLM judges tuned to clinician-graded standards with documented agreement, carrying volume under monthly re-verification.
Intended use, tested behaviors, thresholds, and monitoring cadence, written for review committees and counsel's boundary analysis.
Harnesses running inside BAA channels with de-identification discipline and audit logging matching your existing controls.
Boston's hospitals, digital-health vendors, and life-sciences teams answer to review committees that have read the case reports and counsel that has read the enforcement letters. AI features clear that bar with evidence, not enthusiasm: graded ground truth, gated safety behaviors, and hallucination rates with taxonomies attached. That is the layer we build.
The same rigor compounds operationally: calibrated judges catch regressions before clinicians do, and escalation suites turn the scariest failure mode into a tested property.
We work with Boston teams remotely, in Eastern hours, with first harnesses typically standing in two to three weeks.
Tell us the feature, the clinical context, and the committee it faces. We reply within one business day with a scope and a fixed price.