AI Agent Evaluation · Dallas, TX
Dallas produces regulated decisions and regulated language at portfolio scale: underwriting files summarized for pricing, servicing letters checked against FDCPA and fifty state variants, correspondence an examiner may sample a year from now. Evaluation here is the testing layer that makes all of it defensible.
We build evaluation systems for insurance and servicing AI: violation-corpus testing of rule layers, claim-level faithfulness verification for summaries, template regression suites, and the audit lineage that makes DOI examinations reconstructive rather than archaeological.
Tell us the pipeline and the examiner behind it.
Rule layers get attacked, not assumed: violation corpora built from compliance knowledge and systematic mutation, recall measured near 100 percent because leaks are regulatory events, false positives tracked with the asymmetry explicit, and every real-world near-miss converted into permanent test coverage.
Summary faithfulness is verified claim-by-claim against the file, with underwriting's own taxonomy: fabrications, contradictions, and the silent omissions that shape decisions worst. Calibrated judges carry volume; underwriter-graded samples anchor the standard; bind-relevant outputs hold the strictest bars.
Exam readiness is reconstructability: versioned harness runs, violation-corpus history, and per-letter lineage (template, rules passed, model) that lets a sampled letter from eleven months ago answer cleanly. The NAIC-shaped artifact set assembles itself from evaluation done properly.
Templates are versioned code with test suites: state-variant disclosures, merge-field edge cases, and full-diff reports on every change, so compliance reviews a diff instead of re-reading everything and 'small wording change' retires as an incident category.
The standard build, tuned for insurance and servicing operations.
Known-violating drafts plus systematic mutations, recall measured against a near-100-percent bar, and the corpus growing with every catch.
Claim-level verification with underwriting's error taxonomy, calibrated judges, and thresholds gated by decision stakes.
Per-version expected-behavior tests, edge-case record samples, and output diffs that make compliance sign-off a review, not a re-read.
Versioned runs, per-letter provenance, and remediation records assembled into the NAIC-shaped set examiners recognize.
Specialist-graded standards, measured judge agreement, and monthly anchor samples that keep the scores meaning something.
Production sampling with control bands, change-triggered re-runs, and alarms routed to the queue owners who act on them.
Dallas-Fort Worth operates correspondence and underwriting volumes where AI multiplies output and blast radius alike: a rule gap repeats across thousands of letters before lunch. The evaluation layer is what makes that multiplier governable, and the institutions here, fluent in audit culture already, adopt it naturally once it speaks their dialect: recall numbers, lineage, diffs, exhibits.
The same machinery compounds operationally: tested rule layers let drafting models right-size safely, and template CI converts the steadiest source of incidents into caught-before-send events.
We work with Dallas teams remotely, in Central hours, with first harnesses typically standing in two to three weeks.
Tell us the pipeline, the rules it answers to, and the exam horizon. We reply within one business day with a scope and a fixed price.