Paid Audit
AI companies are better at pitching than most investors are at evaluating. Benchmark numbers look impressive until you understand how the benchmark was constructed. A “proprietary model” may be a fine-tuned wrapper. A “data moat” may be a scraped dataset with unclear licensing. An accuracy claim may hold only on the demo inputs, not on representative production data.
We conduct independent technical due diligence on AI companies — evaluating model quality against claimed benchmarks, assessing architecture scalability, examining data defensibility, and identifying IP and compliance risks. We deliver a written technical memo with risk ratings, follow-up questions, and a clear summary of red flags and green flags.
Tell us about the company you’re evaluating.
Six dimensions of technical risk, each examined independently against the company's claims.
We independently evaluate the model or AI system against the benchmarks the company is using in their pitch. We design evaluation sets that test performance on representative production inputs, not cherry-picked demo scenarios. If the claimed accuracy figures hold up under independent testing, we say so. If they don't, we document the gap and the conditions under which claims hold versus fail.
We review the technical architecture against the company's claimed or implied scale trajectory. Can the current architecture handle 10× the current load? Where are the bottlenecks, the inference cost, the retrieval latency, the data pipeline throughput? Is the architecture typical of what an experienced team would build, or does it reflect shortcuts that will require expensive rewrites at scale?
AI companies frequently claim a data advantage. We assess whether the data asset is actually defensible, proprietary and hard to replicate, or replicable with money and time. Is the training data genuinely differentiated or is it a curated subset of publicly available data? Are the data collection mechanisms sustainable at scale? Is there a feedback loop that improves the model with usage?
Which components of the stack did the team build themselves versus wrapping third-party services? Building on top of OpenAI APIs is a legitimate engineering choice. It's also a different risk profile than training or fine-tuning your own models. We assess whether the build vs. buy decisions were deliberate and appropriate, or whether the company is claiming more proprietary technology than the architecture supports.
A 10-minute conversation with the technical founders or lead engineers tells you a lot about depth. We assess whether the team understands the failure modes of their own system, whether they have evaluation practices, how they approach model updates and reliability, and whether the codebase reflects engineering maturity or prototype-quality code extended past its intended scope.
Training data licensing, third-party model usage terms, output IP ownership questions, and data privacy compliance. AI companies frequently have IP exposure they haven't fully assessed, training on scraped data without clear licensing, using API terms that restrict commercial use, or processing personal data through third-party models in ways that create GDPR or CCPA exposure.
Three deliverables. The memo is yours to use in your investment process.
A written technical memo (10–20 pages) covering each assessment area with findings, evidence, and a risk rating on a four-point scale. Written for a technical reader, with an executive summary for non-technical partners.
A list of 10–20 specific follow-up questions for the company's technical founders or engineering team, based on our findings, designed to probe the areas where we identified risk or uncertainty.
A clear summary of the findings that should increase confidence versus the findings that warrant further investigation or price adjustment. Not a recommendation on whether to invest. That's your call, but a clear enumeration of the technical factors relevant to that decision.
Three types of organisations that commission AI technical due diligence.
Technical due diligence on AI companies is different from standard software company diligence. Most investment teams don't have in-house ML engineering capacity to evaluate benchmark claims, assess training data defensibility, or identify model architecture red flags. We provide that assessment as a standalone engagement.
The technical depth of an AI acquisition target matters significantly to integration cost, scalability, and IP risk. We assess whether the technology is what it appears to be, whether the architecture can be integrated into a larger system, and whether there are IP or compliance issues that affect deal terms.
Enterprise procurement teams evaluating AI vendor claims (particularly around accuracy, security, and data handling) benefit from independent technical assessment before signing multi-year contracts. We assess vendor claims against independent evaluation and flag contractual risk areas.
This is a third-party assessment for external evaluators. It has specific prerequisites.
Companies conducting self-assessment
This engagement is specifically for third parties assessing an AI company or vendor: investors, acquirers, or enterprise buyers. We don't conduct due diligence on your own company. If you want an independent technical review of your own system, our AI readiness assessment or AI security audit are the right engagements.
Assessments where the company won't provide access
The assessment requires access to the technical team for interviews, the codebase or architecture documentation, and the system for independent evaluation. If the company is unwilling to provide these under NDA, we can assess only the publicly available information, which is materially less useful for investment decisions.