LLM Cost Optimization · New York, NY
New York's LLM bills grow the way its workloads do: filings analyzed twenty times over, legal documents processed at firm scale, media archives tagged in the millions, all on models sized for caution rather than measured need. Cost optimization here means auditing where the tokens actually go and fixing the spend without touching the compliance posture.
We audit a month of real usage, baseline it, and implement the ranked fixes: prompt trimming, caching architecture, eval-gated model routing, and batch migration, with every change measured against your own quality bar.
Tell us the workload and last month's bill.
Spend concentrates in four places, in rank order: prompt bloat that grew during development, frontier models doing commodity work, repeated context paying full price for lack of caching, and real-time pricing on workloads that run overnight. The audit instruments a month of traffic and ranks the fixes in dollars, because intuition about token costs is reliably wrong.
Document-heavy finance and legal workflows are caching's best case: stable prefixes around filings, agreements, and standing instructions turn the twentieth query against the same document into a 90-percent-discounted call. The work is prompt architecture, and it is measurable in the hit rate.
Routing in regulated shops is eval-gated or it is gambling: cheaper models take only the request classes where your own labeled set shows parity, every request logs its model and reason, and client-facing final drafts keep the strong model until the evidence says otherwise.
Compliance constraints shape the levers rather than removing them: retention rules and logged outputs coexist fine with caching and routing, and where a constraint genuinely blocks a fix (some zero-retention configurations limit caching), the report says so and prices the alternative honestly.
The standard sequence, tuned for New York's regulated document workloads.
A month of real traffic instrumented: tokens by feature, model, and prompt section, with the spend map your finance team has been guessing at.
System prompts and context trimmed against your eval set, often the cheapest 30 to 50 percent reduction available, with quality regression-tested.
Prompts restructured for stable prefixes, cache hit rates measured in production, and repeated-document workflows paying the discounted rate they deserve.
Cheaper models taking only the request classes where your labeled set proves parity, with model-and-reason logging for every request.
Overnight and non-interactive workloads moved to batch APIs at half price, with queue and retry plumbing included.
Dashboards per feature and per model, drift alerts, and the instrumentation to re-answer the cost question every quarter without us.
New York's LLM spend skews document-analytical: research over filings, contract and brief processing, compliance summarization, media tagging at archive scale. Those shapes respond unusually well to caching and routing, and unusually badly to one-size frontier-model defaults, which is why audits here routinely find large headroom.
The compliance overlay (retention, logging, version pinning) constrains nothing essential: every lever we use leaves the audit trail intact, and the deliverables include the documentation your reviewers will ask for.
We work with New York teams remotely, in Eastern hours, with the audit typically complete in two weeks and implementation in two to four more.
Tell us the workload, the stack, and roughly what last month cost. We reply within one business day with an audit scope and a fixed price.