AI Managed Services
The problem with AI in production is that it drifts. Models change, context windows expand, providers deprecate endpoints. A RAG system that scored 92% accuracy in January might score 78% in June if no one is watching. Most teams notice when a user complains, not before.
We run monthly evaluation cycles, apply model updates, monitor costs, and respond to incidents on production AI systems, ones we built or ones we inherited from another team. The retainer is designed to keep a working system working, not to add features.
Tell us what you're building.
All six items are included in every tier. The tiers differ in how much optimisation work is included and response priority.
We run your golden dataset through the system every month and report accuracy, faithfulness, and retrieval recall against the established baseline.
When providers update their models, we assess the impact on your system, pin or upgrade as appropriate, and document the decision.
Automated anomaly alerts when spend spikes unexpectedly. Monthly cost trend report so you can see patterns before they become problems.
Minor prompt optimisations based on eval results. Not a rewrite, targeted adjustments that move the needle on the scores that matter.
LangChain, OpenAI SDK, and related packages kept current. Security patches applied within 48 hours. Breaking changes handled with testing before rollout.
SLA: 4-hour response and 24-hour resolution target for P1 (system down or accuracy below 50%). 24-hour response and 72-hour resolution for P2 (degraded performance).
The retainer is for maintenance and incident response. Being clear about scope prevents billing surprises on both sides.
New feature development, billed separately as project work
Complete rewrites of existing systems
Incidents caused by client-side changes made without informing us
Traditional software degrades slowly, a bug is a bug, and it stays a bug until someone fixes it. AI systems degrade invisibly. A model provider releases a new default version, the output distribution shifts, and your application starts giving subtly worse answers without any error log to investigate.
RAG systems have a second failure mode: data freshness decay. The documents your retrieval layer was indexed against get updated, renamed, or deleted. The system keeps retrieving outdated chunks, and accuracy falls without any code change on your part.
Monthly eval runs catch both. A fixed golden dataset gives you a stable reference point. When scores drop, you have a number to point to, a timestamp, and a diff of what changed between the last passing run and the current one. That's the difference between a 20-minute investigation and a two-day debugging session.
A retainer is only as good as the baseline it watches against. Most of the value in month one is setup: we cannot tell you the system is drifting until we know what good looked like. Whether we built the system or inherited it, the first weeks follow the same path.
We learn the system. For an inherited system, week one is an audit: we map the architecture, the data flow, the prompts, the model versions, and the failure modes the previous team never documented. You get a written account of what you actually own, which for many teams is the first time that has existed on paper.
We build the golden dataset. We assemble a fixed evaluation set from real queries and known-good answers, reviewed by someone on your team who knows what correct looks like. This set is the stable reference every future run is measured against, so a drop in scores is a fact with a timestamp rather than a hunch from a support ticket.
We establish the baseline. We run the golden set, record accuracy, faithfulness, latency, and cost, and agree the thresholds that count as healthy. From then on, a monthly run either confirms the system is holding or flags exactly which metric moved and when, which is the difference between a short investigation and a long one.
We wire up the alerts. Cost-spike alerts, error-rate monitors, and the incident channel go live so a problem reaches us before it reaches your customers. By the end of month one you have visibility most teams never set up for themselves, and the retainer settles into the quieter rhythm of monthly runs and fast incident response.
Not every team needs a managed services retainer. Here's when it's not the right call.
Teams with in-house AI engineers
If you have engineers who can own model monitoring and eval, spend the budget on their time rather than ours. This service exists for teams without that internal capacity.
Genuinely stable projects
If your AI system has had no incidents, no model updates, and no data changes in six months, you may not need a retainer. Rare in practice, but worth asking.
Startups pre-product-market-fit
Before you have repeatable revenue and a stable product, spend the budget on shipping features instead. Come back when the system is in production and business-critical.
3-month minimum on all tiers. After that, month-to-month with 30 days notice to cancel.
Standard
Eval suite run against your golden dataset, cost monitoring with anomaly alerts, dependency updates, and incident response.
Standard + Optimisation
Everything in Standard, plus monthly prompt optimisation work based on eval results and a written improvement report.
Premium
Everything in Standard + Optimisation, plus weekly syncs, dedicated response queue, and priority scheduling for any additional project work.
Not sure which tier fits? Describe your system and we'll recommend one on the discovery call.
Tell us about your current AI system. What it does, what stack it runs on, and any issues you've noticed. We'll reply within one business day.