Paid Audit
This is a 5-day paid audit that produces a written report showing which LLM calls are your top cost drivers, what you're paying per query, and which optimisations would have the biggest impact. It's a diagnostic (not an implementation engagement) and it answers the question that most LLM-heavy teams can't answer from their provider dashboard alone.
You get a ranked list of changes with projected savings and implementation effort, plus a 60-minute debrief call to walk through the findings. You keep the report regardless of whether you hire us to implement the fixes. Turnaround: 5 business days from receipt of your usage data.
Tell us what you're building.
Six areas of analysis, each producing specific findings, not a generic report with generic recommendations.
30–90 days of OpenAI, Anthropic, or Google API usage logs broken down by endpoint, model, and prompt template. We map spend to the features that generate it, not just aggregate totals by model.
Input vs. output token ratio, context window utilisation, and average tokens per call by endpoint. Output tokens cost 3–5× more than input tokens on most models. Teams are often surprised how much their output verbosity is costing them.
Are you using the right model for each task? Classification and summarisation tasks running on GPT-5 when GPT-5 mini would produce equivalent results are a common finding. This is usually the fastest win with no quality tradeoff.
What percentage of your queries are semantically similar to previous queries? Is semantic caching viable for your query distribution? We calculate the potential cache hit rate and estimate the monthly savings if caching is implemented.
The 10 most expensive query patterns in your system, with estimated monthly cost per pattern. Named, specific, ranked. Not "your chatbot is expensive" but "this prompt template running on GPT-5 accounts for $1,200 of your $3,400 monthly bill."
Every identified optimisation ranked by projected monthly savings divided by implementation effort in engineering days. You see exactly which changes pay back fastest so you or your team can act on the report immediately.
Three deliverables. No retainer. No lock-in.
10–15 pages covering all findings, charts showing cost distribution by endpoint and model, and the full prioritised optimisation list with effort estimates.
We walk through the findings, answer questions about any recommendation, and help you prioritise what to implement first based on your engineering capacity.
Take the written report to any engineer or agency to implement the fixes. You are not required to hire us for implementation. The audit cost is credited toward implementation if you proceed within 60 days.
The audit has a fixed scope and specific prerequisites. If any of these apply, it's not the right starting point.
Teams spending under $300/month on LLMs
At that spend level, the audit fee won't pay back within a reasonable timeframe. When your monthly LLM bill reaches $800–$1,000, the math changes and optimisation becomes worthwhile.
Companies without API usage logs or billing data
The audit requires your API usage export from your provider dashboard and billing data for the same period. Without usage logs, there is nothing to analyse.
Anyone who needs fixes this week
The audit is a diagnostic first. If you have a production incident caused by runaway LLM spend, you need emergency triage, not a 5-day structured audit. We can discuss emergency support separately.