AI startup due diligence comes down to seven checks: is the AI foundational, is the training data proprietary, are the models owned, can the team build them, do inference costs scale, is the company regulator-ready, and is any of it defensible.
Our repository of 13,500+ AI funding deals tracked since 2023 shows why the first check carries the most weight. A company’s claim to use artificial intelligence tells you nothing about its investability.
The spread inside that dataset is extreme. OpenAI closed a $40 billion round at a $300 billion valuation in March 2025, then raised $110 billion more at a $730 billion valuation eleven months later.
In the first quarter of 2024, Stability AI generated under $5 million in revenue while carrying close to $100 million in unpaid bills. Both companies were AI companies by any marketing definition.
That gap is why AI investment evaluation needs a different toolkit than traditional venture capital due diligence. This guide gives you the framework, the metrics, the red flags, and 25 questions to ask before you wire money.
On this page
- Why AI Startups Need Specialized Due Diligence
- The 7-Point AI Due Diligence Framework
- Technical Assessment: Evaluating AI Models, Data, and Architecture
- Financial Due Diligence: AI-Specific Metrics That Matter
- Red Flags: How to Spot AI Washing and Overblown Claims
- AI Governance, Model Risk, and Legal Diligence
- Due Diligence Checklist: 25 Questions to Ask Before Investing
- How the AI Due Diligence Process Runs
- FAQ: AI Startup Due Diligence
- Methodology
Why AI Startups Need Specialized Due Diligence
Standard venture capital due diligence evaluates team, market, product, and financials. For AI startups, that checklist misses the three variables that decide whether an advantage is real or manufactured: data quality, model defensibility, and compute economics.
Consider the difference between a company our AI startup classification framework puts in the AI Native tier, where the product cannot exist without AI, and a CRM tool that added a chatbot widget and rebranded. Anthropic sits in the first group. Plenty of Series A pitch decks sit in the second while using identical language.
Classification is the first filter precisely because investors rarely run it systematically. A company whose product collapses without its models is underwriting a different risk than one that shipped a chat interface, even when both raise on the same vocabulary.
Regulation adds a second kind of exposure. The EU AI Act’s prohibitions have applied since February 2025 and its general-purpose AI obligations since August 2025, with transparency duties arriving in August 2026.
The heaviest tier, the high-risk system obligations, was pushed back by the EU’s 2026 digital omnibus revision to December 2027 for Annex III systems and August 2028 for Annex I. Penalties reach 7% of total worldwide annual turnover or 35 million euros, whichever is higher, for prohibited practices, with lower tiers at 3% and 1%.
The SEC stood up a Cyber and Emerging Technologies Unit in February 2025 to prosecute AI washing. Investors who skip technical verification on AI companies now carry both financial and legal exposure.
Evaluating AI startups in 2026 means answering a question traditional diligence never asks: is the AI real, and does it matter?
The 7-Point AI Due Diligence Framework
The framework distills the evaluation into seven domains. Each one builds on the last, so skipping a step collapses everything downstream.
| Point | Domain | Core Question | Kill Criteria |
|---|---|---|---|
| 1 | AI Classification | Is AI foundational, augmentative, or cosmetic? | Company cannot articulate what AI removes from the workflow |
| 2 | Data Asset | Is the training data proprietary, licensed, or public? | Entire model built on publicly available datasets with no unique data flywheel |
| 3 | Model Architecture | Custom models vs. API wrappers vs. fine-tuned open-source? | Pure API wrapper with no proprietary layer |
| 4 | Technical Team | ML engineers vs. “AI strategists” on the org chart? | No in-house ML talent; outsourced model development |
| 5 | Unit Economics | Does inference cost scale favorably with revenue? | Gross margins below 50% (software) or 30% (AI-infra) with no path to improvement |
| 6 | Regulatory Readiness | Compliance with EU AI Act and SEC disclosure rules? | High-risk AI application with zero compliance documentation |
| 7 | Competitive Moat | Data moat, switching costs, or network effects? | Defensibility rests solely on “first mover advantage” |
The framework works as a sequential filter. Technical assessment layers on top of commercial diligence rather than running beside it as a separate workstream, because the commercial numbers mean something different once you know what is producing them.
Point 1 eliminates the largest category of wasted diligence time. A real share of companies that pitch themselves as AI startups classify as Non-AI the moment you examine what the product actually does. Running that check before any deep technical work saves weeks of analyst time.
Technical Assessment: Evaluating AI Models, Data, and Architecture
Technical due diligence for AI startups has three layers: data, model, and infrastructure. Each needs a different evaluation method, and each is where pitch-deck language stops being useful.
Data Quality Assessment
A data moat exists when a dataset improves with usage, cannot be replicated by competitors, and produces measurable performance advantages. Data quality on its own is not a moat in 2026, because open datasets commoditize it.
The moats that hold combine execution speed with a data loop that compounds every time a customer uses the product. Verify four things:
- Provenance: Where does training data originate? Licensed, scraped, user-generated, or synthetic?
- Volume and freshness: How often is the dataset refreshed? Stale data degrades model performance.
- Labeling quality: Who labels the data, domain experts, crowd workers, or automated systems?
- Compliance: Does collection satisfy GDPR, the EU AI Act’s training-data disclosure duties, and applicable copyright law?
Scale AI, valued at $29 billion after Meta’s 2025 investment, built its defensibility on proprietary data labeling infrastructure rather than on a model. That is the shape of a technical moat, and it is the kind of thing a data-quality review surfaces that a product demo does not.
Model Performance Evaluation
Model performance evaluation goes well past accuracy, whether the company is running classical machine learning, natural language processing over unstructured data, or generative AI. Ask for:
- Benchmark comparisons: How does the model perform against open-source alternatives on standardized benchmarks?
- Inference latency: Response time at production scale, not demo conditions.
- Hallucination rates: For generative systems, what share of outputs contains fabricated information?
- Degradation curves: How does performance move as input complexity rises?
A company calling a third-party model through an API with a custom prompt template carries a different risk profile than one running proprietary models on owned infrastructure. Both will describe themselves as AI-powered. The technical review is what tells you which one you are buying.
Financial Due Diligence: AI-Specific Metrics That Matter
AI startup valuation in 2026 does not behave like SaaS valuation. Revenue multiples for AI companies run at a large premium to comparable software businesses, and that premium is frequently applied to the label rather than to the underlying economics.
The metrics below are the primary defense against overpaying for it.
| Metric | What It Reveals | Red Flag Threshold |
|---|---|---|
| Gross Margin | Compute cost sustainability | Below 50% for software, below 30% for AI-infra |
| Revenue per Employee | Operational efficiency | Below $150K at Series B+ |
| Net Dollar Retention | Product stickiness | Below 110%; 120%+ signals strong retention |
| Inference Cost per Query | Scaling economics | Rising faster than revenue per query |
| Data Acquisition Cost | Moat sustainability | Increasing quarter-over-quarter with no quality gains |
| Customer Concentration | Revenue risk | Top 3 customers above 60% of ARR |
The most diagnostic line is gross margin trajectory. AI startups burning GPU compute often post gross margins of 30% to 40%, which is fine at seed and a structural problem at Series B if the trend line is flat.
Stability AI’s unpaid cloud bills and creditor debt in early 2024 are what that looks like from the inside. The company was recapitalized later that year, which is the outcome you are underwriting against, not the one you plan for.
Contrast that with Palantir, which reported an 84% adjusted gross margin in Q3 2025 on a software-led revenue mix. When you underwrite an AI margin, establish what share of revenue is software and what share is human delivery dressed as product.
Read the financial statements against the compute bill, not on their own. An AI company’s financial health lives in the relationship between cloud spend, headcount, and revenue, and a P&L that shows a deep-tech story with no meaningful infrastructure line is describing a different business than the one being pitched.
Venture capital due diligence for AI also has to assess revenue quality. A company showing $10 million ARR from three enterprise contracts is a different risk than one showing $10 million from 500 mid-market customers. The top line is identical. The concentration risk is not.
Red Flags: How to Spot AI Washing and Overblown Claims
AI washing, the practice of exaggerating or fabricating AI capabilities, is no longer only a reputational risk. It is a legal one.
In April 2025 the SEC and DOJ filed parallel actions against Albert Saniger, founder of Nate Inc. He allegedly raised $42 million on the claim that his shopping app was “fully automated based on AI.” The transactions were processed by hand, first by contract workers in the Philippines and later in Romania.
Nate is not an outlier, it is the case that drew an indictment. Agent washing is the current version of the same move: chatbots and scripted automation rebranded as autonomous agents, worth reading against the AI agent startups actually raising money.
The test is the one that applied to machine learning claims a decade ago. Ask what decision the system makes with no human in the loop, then ask to watch it make one.
Our guide on how to spot AI washing covers the behavioral patterns that precede these failures. During diligence, these are the signals that surface first.
Technical red flags
- No ML engineers on the team, only “AI strategists” or “prompt engineers”
- Inability to explain model architecture without pointing at a third-party provider
- Demo environments that cannot be reproduced in live production
- Reluctance to share model performance benchmarks or error rates
Commercial red flags
- Customer testimonials focused on potential rather than measured outcomes
- Revenue growth driven by one-time implementation fees instead of recurring usage
- AI features described in future tense across every piece of marketing
- Competitive positioning that rests entirely on “proprietary AI” with no specifics
Financial red flags
- R&D spending below 15% of revenue at a company claiming deep-tech differentiation
- No compute costs on the P&L at a company claiming to run its own models
- Valuation justified by an “AI premium” with no AI-driven unit economics behind it
AI Governance, Model Risk, and Legal Diligence
The three checks that most often get skipped sit between the technical review and the legal one. They are also the three that turn into write-downs after the round closes.
Algorithmic Bias and Model Risk
Ask how the company tests for algorithmic bias and what it does when a test fails. A team that has never measured performance across demographic or geographic slices has not measured the risk it is carrying, and in a regulated sector that is a product defect rather than an oversight.
Model risk also means drift. Any AI system in production degrades as the world moves away from its training data, so ask what monitoring exists, what threshold triggers a retrain, and who owns that decision.
AI Governance and Documentation
AI governance is the paperwork that proves the company can answer a regulator. At minimum, look for a model inventory, documented training data provenance, human-review procedures for consequential decisions, and an incident log.
Startups treat this as a Series C problem. Buyers and enterprise customers increasingly treat it as a procurement gate, which makes it a revenue question rather than a compliance one.
Data Privacy and Legal Exposure
Where a product touches personal or clinical data, confirm the specifics: GDPR lawful basis, HIPAA handling if health data is involved, and what the contracts say about customer data being used for training. That last clause is where a data moat turns into a lawsuit.
Legal due diligence for AI companies also has to cover training-data licensing and IP. A model fine-tuned on scraped copyrighted material carries an exposure that does not appear anywhere on the cap table.
Due Diligence Checklist: 25 Questions to Ask Before Investing
The checklist organizes 25 questions by domain. Each one targets a specific failure mode that repeats across the deals we track.
AI Classification (Questions 1-4)
- What specific tasks does AI perform in the product that a rules-based system cannot?
- If the AI component were removed, would the product still function? What would be lost?
- What percentage of the product’s value proposition depends on AI versus non-AI features?
- How does the company’s AI usage map to the AI startup classification framework (Native, Augmented, Adjacent, Platforms)?
Data and Model (Questions 5-11)
- What is the source, size, and refresh rate of the training dataset?
- Does the company own its training data, license it, or scrape it from public sources?
- What data labeling methodology is used, and what is the inter-annotator agreement rate (how often independent labelers assign the same label to the same data point)?
- How does model performance compare to open-source alternatives on standardized benchmarks?
- What is the model’s error rate in production, and how is it measured?
- Does the product use proprietary models, fine-tuned open-source models, or API calls to third-party providers?
- What is the model retraining cadence, and what triggers a retrain?
Team and Talent (Questions 12-15)
- How many team members hold ML or AI-specific roles versus general engineering?
- What is the founder’s technical background in AI and ML specifically?
- Has the team published peer-reviewed research or open-source contributions in AI?
- Does the founder have domain expertise in the field where the AI is applied? Founder-market fit correlates with better technical calls and faster iteration.
Unit Economics (Questions 16-19)
- What is the fully loaded cost per inference at current scale?
- How does inference cost change at 10x current volume?
- What is the current gross margin, and what is the path to 70%+?
- What percentage of revenue comes from AI-specific features versus non-AI features?
Regulatory and Legal (Questions 20-22)
- Has the company mapped its products against the EU AI Act risk tiers?
- What AI-specific IP protections (patents, trade secrets) does the company hold?
- Does the company meet AI regulatory compliance requirements in each target market?
Market and Competitive Position (Questions 23-25)
- What prevents a well-funded competitor from replicating this AI capability within 18 months?
- Does the company have a data moat that compounds with usage?
- What is the switching cost for customers once the AI is embedded in their workflows?
Score each answer on a 1 to 5 scale and flag any domain where the average falls below 3. Two or more domains below 3 is a pass signal for all but the most exceptional teams.
That scoring discipline is the whole method. It is also what separates a repeatable process from a partner’s instinct, which is worth checking against how the most active AI investors actually build their pipelines.
How the AI Due Diligence Process Runs
The due diligence process for an AI startup runs in three passes, and the order matters more than the depth of any single one.
Pass one, one to three days. Classification and desk research. Establish what the product does without AI, read the financial statements, and map the company against the seven points. Most passes end here, which is the point.
Pass two, one to two weeks. Technical and commercial verification. Live demo outside the sandbox, benchmark data, production error logs, customer references, and a walk through unit economics at 10x volume. Bring in an outside ML reviewer if nobody on the deal team can read a model card.
Pass three, one to two weeks. Legal, governance, and confirmatory work. Training-data licensing, IP, privacy posture, and the AI governance documentation above, run in parallel with standard corporate diligence.
Compressing this to a single sprint is where diligence outcomes go wrong. Technical verification generates the questions that commercial and legal diligence then have to answer, so running them at once produces three separate reports that never reconcile.
FAQ: AI Startup Due Diligence
What should investors look for when evaluating AI startups?
Three criteria are non-negotiable: whether the AI is genuinely foundational to the product, whether the company holds a proprietary data asset that improves with usage, and whether the unit economics support scaling as inference costs decline against revenue.
Companies that clear all three are rare. Most clear one and pitch as though they clear all three, which is exactly what the classification step is for.
How do you assess the quality of an AI startup’s data?
Evaluate four dimensions: provenance (is the data legally and ethically sourced?), uniqueness (can competitors reach the same data?), freshness (how often is it refreshed?), and labeling quality (expert-annotated, crowd-sourced, or synthetic).
Then request a data sample and compare model performance against the same model trained on publicly available data. The delta is your data moat, measured rather than asserted.
What are the red flags in AI startup due diligence?
The highest-signal ones: no ML engineers on staff, inability to run a live demo outside a controlled environment, R&D spending below 15% of revenue at a self-described deep-tech company, and customer references that describe potential rather than measured outcomes.
The SEC’s enforcement action against Nate Inc., where “fully automated AI” turned out to be manual labor, is the extreme version of what those signals point at.
How is AI due diligence different from traditional startup evaluation?
Traditional due diligence assumes the product works as described and concentrates on market size, team, and financials. AI due diligence adds a technical verification layer, because the gap between a company that markets AI and a company that runs it is invisible from a pitch deck.
You verify the AI exists, assess whether it creates defensibility, evaluate compute economics, and confirm regulatory compliance. None of those appear on a standard VC checklist.
What metrics matter most for AI startups?
Five metrics separate viable AI companies from hype. Gross margin should trend above 50% for software-layer AI and above 30% for AI infrastructure. Net dollar retention of 110% is the floor, and 120%+ signals strong retention. Inference cost per query has to decline with scale.
Two more complete the picture: data acquisition cost, which must not outpace data value, and revenue per AI feature, which isolates the AI contribution from legacy product revenue.
How do you detect AI washing during due diligence?
Three tests. Ask the CTO to whiteboard the model architecture without referencing third-party APIs, since inability to do so indicates a wrapper rather than a product. Request production error logs, because companies with real AI have real errors and real monitoring.
Then compare employee profiles against the claimed AI capabilities. If the team is 90% salespeople and 10% engineers, the “proprietary AI” is an integration. Our full breakdown of how to spot AI washing goes deeper on each test.
Methodology
This guide draws on Bot Memo’s repository of 13,500+ AI funding deals tracked from 2023 through 2026. Records are verified and deduplicated before they enter the repository, and each company is AI-classified.
Classification system: five tiers. AI Native (AI is foundational), AI Augmented (an existing product enhanced with AI), AI Adjacent (enables AI infrastructure), AI Platforms (trains foundation models or designs AI chips), and Non-AI (no meaningful AI component). Full detail sits in our data methodology.
Limitations: companies named here are illustrative examples, and inclusion is not an investment recommendation. Valuation and funding figures reflect publicly reported amounts as of August 2026. Valuation data is available for a subset of tracked deals.



