GROW TODAY · MEMO
Operational Readiness Is Now a Diligence Question
Why 2026's Best Rounds Get Scored on Ops Layer, Not Just Deck
By Zach Zelle, General Partner, Aureum Capital
PUBLISHED · AUG 10 · 2026
TL;DR
Diligence has moved from deck to ops layer. In 2026, GPs are scoring companies on five operational categories in parallel with financial diligence: AI orchestration, team leverage (revenue per employee), data infrastructure and provenance, cross-tool systems, and real-time visibility. The evidence is structural, not narrative. Menlo Ventures puts enterprise AI investment at $37B in 2025, up from $1.7B in 2023. ICONIQ tracks AI as a rising share of internal spend. Bessemer names AI-native operations as an underwriting frontier for 2026. Cursor, Lovable, and Ramp are the reference cases for what a legible ops layer looks like at scale. Builder.ai and Presto Automation are the reference cases for what happens when the ops layer is fabricated or absent.
The morning the ops layer showed up in every data room
In May 2025, Builder.ai filed for insolvency after a lender seized $37M following a fraud investigation. The company had spent years marketing an AI-native app-building platform. The internal probe, per Silicon.co.uk, found the platform had been delivered largely by more than 700 human developers across India and Ukraine. Reported revenue of $180M for 2023 and $220M for 2024 had actually been $45M and $55M, roughly a 300% inflation each year, allegedly enabled by a round-tripping invoicing scheme with VerSe Innovation.
I spent the following partner meetings watching diligence questions change shape in real time. GPs I have known for a decade began asking, in the first meeting, the questions they used to save for confirmatory diligence. What ran the model in production. Where the human hand-off lived. Who owned the eval loop. How booked revenue reconciled to product usage on a per-account basis. The deck stopped being the surface of the conversation. The ops layer became the surface.
At Aureum Capital we run Grow Today as our founder services platform for exactly this reason. Founders who could not answer those questions in April 2025 were losing rounds by August. Founders who could, without prep, were closing at term sheets their financial metrics alone would not have justified twelve months earlier.
What operational diligence means in 2026
The phrase has entered industry vernacular fast enough that its content needs pinning down. In 2026, operational diligence is a scored bundle of five categories that funds work in parallel with financial diligence, not after it.
AI orchestration
GPs increasingly separate AI adoption from AI system. VC Lab captured the distinction in a 2026 analysis: "A GP who responds with tooling names is describing adoption, while a GP who responds with calibration data, override rates, and an anomaly-routing design is describing a system." The same test is now applied to portfolio companies. Naming Claude, Cursor, or ChatGPT tells us very little. Override rates, cost per resolution, and a routing tree with escalation logic tell us what the company will look like at three times its current scale.
Team leverage and revenue per employee
Revenue per employee has become close to a first-order metric, tracked alongside ARR, NRR, and burn multiple. Nick Talwar's analysis of AI-native revenue density puts the top ten AI startups at an average of $3.48M revenue per employee, versus $610K for the leading traditional SaaS firms, a 5.7x spread. Cursor reached $1.67M per employee at sixty people and $3.3M per employee at three hundred. Lovable hit $2.7M ARR per employee at 146 people on $400M ARR. Lighter Capital's benchmarks put private SaaS median at roughly $130K per employee and public SaaS median at roughly $395K.
Data infrastructure and provenance
Capwave's 2026 enterprise-readiness analysis captured the shift in a single line: "The readiness bar moved from 'is your software stable' to 'can you prove the data path.'" GPs now ask for lineage: where training and inference data originates, what governance sits over it, what happens when a customer requests deletion, how the audit trail survives a subpoena. The last question sounds legal. In current diligence it is technical.
Cross-tool systems and internal automation
Ramp is the public reference case. Its AI Principles document builds explicitly toward autonomous finance operations, and Notion's case study on Ramp's internal AI operating system describes agents auto-flagging missing or incorrect data across the company's own tooling. The diligence-relevant point is not that Ramp uses AI. It is that Ramp's internal ops stack is legible to an outside investor as a system, with defined inputs, agents, escalation logic, and human review gates.
Real-time visibility
The financial model, historically the primary diligence artifact after the deck, has been joined by a live-metrics dashboard as a first-look asset. CapMaven's 2026 piece states it plainly: "Diligence is the new pitch. Your financial model has become a better storyteller than your 20-slide deck." Our view at Aureum is narrower. The dashboard behind the model has become a better storyteller than the model itself.
The evidence this is not a narrative shift
Narrative shifts in venture are cheap and structural ones are rare. This one is structural.
Menlo Ventures' 2025 State of Generative AI in the Enterprise report tracked enterprise AI investment at $37B in 2025, up from $11.5B in 2024 and $1.7B in 2023, now roughly 6% of the global SaaS market. ICONIQ's State of AI report, summarized by SaaStr, puts internal AI spend as a share of revenue at 11% in 2025, projected 16% in 2026 and 19% in 2027. High-growth companies in the same dataset allocate roughly 57% of R&D spend to AI, versus a 38% average. AI product gross margins moved from 41% in 2024 to a projected 52% in 2026.
Bessemer's 2026 AI Infrastructure Roadmap names AI-native operations, defined as infrastructure for grounding AI in operational context, as one of five underwriting frontiers for 2026, formally distinct from application-layer AI. The category is now on the underwriting map.
Sprinto's 2026 diligence roundup states that compliance standards like SOC 2 and AI bias audits are now treated as mandatory at the Seed stage, a bar that did not widely exist in 2023. A Bain-sourced finding cited in the same piece attributes the majority of deal disappointments to operational and commercial risks, with people-related issues a close third.
At the LP layer, which sits one level above the fund-to-startup conversation but signals the same rigor pressure, Blackbird reported that 79% of 150 LPs surveyed across North America, Europe, and Asia Pacific have significantly deepened operational scrutiny in the past year. LPs who hold their GPs to an operational standard produce GPs who hold their portfolio companies to one.
Revenue per employee, benchmark spread. Private SaaS median: roughly $130K. Public SaaS median: roughly $395K. AI-native leaders: $400K-plus for the top 55%, with a $3.48M average for the top ten. Cursor at three hundred people is running $3.3M each. Lovable at 146 people is running $2.7M each. This is the number GPs now request as a first-meeting follow-up, not a confirmatory diligence item.
The counter-signal, which I take seriously, is that despite $97B of AI investment in 2025 (roughly 40% of all VC dollars), Startups Magazine reports fewer than 15% of investing firms had any formal framework for assessing how portfolio companies actually source and handle data. The rigor is rising unevenly, but it is rising.
Case: Cursor, where the ops layer carried the round
Cursor's November 2025 Series D closed at a $29.3B valuation on $2.3B raised, led by Coatue, Nvidia, and Google. A subsequent 2026 round pushed the company to roughly $50B, co-led by a16z, Thrive, and Nvidia. The commentary from investors and analysts around both rounds treated a specific operational fact as central to the underwriting logic. The company had reached roughly $2B in ARR in three years on a headcount that Value Add VC characterized as the fastest B2B scaling on record.
The ops-layer artifact here was not a dashboard screenshot. It was the ratio itself. Revenue per employee is a compressed statement about hiring discipline, product leverage, and internal AI use in one number. When a GP reviews a company running at $3M-plus per employee, the follow-up questions write themselves: which functions have been automated, where the human decisions still live, what breaks first at four times current volume. The number invites the diligence rather than surviving it.
To be precise about causation: I could not find on-record GP attribution stating ops-layer discipline was the sole reason Cursor's rounds closed at those valuations. What is visible is that Cursor's team-leverage numbers were the publicly documented feature outside analysts consistently pointed to when explaining the underwriting. In a market where the reference case for a fabricated ops layer had just filed for insolvency, that documentation mattered. The ops layer was legible before the meeting. That legibility is the point.
Case: Builder.ai, where the ops layer voided the round after the fact
Builder.ai is not a story of failed diligence at the term sheet. It is a story of what happens when diligence at the term sheet does not include operational verification. The company had raised roughly $450M cumulative from investors including Microsoft, SoftBank's DeepCore, and Qatar Investment Authority. Its narrative was clean, its revenue growth on paper was strong, and its AI-native positioning matched the fundraising climate.
The SEC's February 2025 action against Presto Automation for failing to disclose that its voice-AI product was actually developed, owned, and operated by a third party, as covered by Global Investigations Review, is the same failure mode at smaller scale. IdeaProof's 2026 tracking analysis of more than 364 failed AI startups since 2023 identifies AI-washing exposure, GPU-burn without a data moat, and commoditization as recurring killers. The pattern rhymes across scale.
The legal frame, articulated in Gurpreet S. Bal's 2026 analysis of venture-backed startup fraud, has already reshaped GP behavior on the liability side: "A venture investor who sits on a board, receives materials claiming AI capabilities, approves fundraising rounds built on those claims, and fails to conduct any independent verification of the technical claims has a different exposure profile than one who documented diligence." That sentence, more than any market-timing argument, is why documented operational diligence has moved from optional to standard.
The steelman for pushing back
The honest question is whether any credible investor is defending pure narrative-based diligence in 2026. After a dedicated search, I could not find one. What exists is a methodology dispute, not a categorical one. SeedScope's 2026 analysis argues that "the market rewards discipline over speed" and that institutional-grade metrics now move companies through diligence faster than a strong narrative alone. That is the mainstream position, not the dissent.
The best steelman sits adjacent to the WRITER and Contrary Research findings on enterprise AI ROI. Roughly 60% of public-company CEOs report AI projects have not yet delivered positive ROI, and 79% of enterprises report adoption challenges despite continued investment. From that data, a fair argument runs: GPs may be over-indexing on visible ops-layer maturity (dashboards, agent stacks, calibration reports) as a proxy for quality, when the underlying enterprise data shows AI tooling adoption alone does not reliably predict outcomes. The signal risks becoming a checklist a well-resourced founder can produce without the substance underneath it.
That argument does not survive current LP pressure. When 79% of surveyed LPs are asking their GPs to deepen operational scrutiny, and when the legal exposure profile for undocumented diligence has changed, the response is not to abandon the ops layer as a scored category. The response is to sharpen how ops artifacts are read, so a Potemkin dashboard does not pass for the real thing. The correction runs toward more rigor, not less.
What operationally legible looks like before the meeting
For founders reading this, and for operators inside our own portfolio, the practical work is upstream of the pitch. Operational legibility is not built inside the data room. It is exposed there.
The layers we recommend building first, in order:
- A live metrics view with non-revenue signals. Product usage per account, AI cost per resolution, human override rate, time from lead to activation. These sit next to the ARR waterfall, not underneath it.
- A written eval loop for any AI in the product. What the model does, how quality is measured, how regressions are caught, who signs off on a shipped change. One page, with names attached.
- A team-leverage table. Revenue per employee, cost per closed customer, ratio of engineers to active users. Updated monthly. Three numbers, one page.
- Data lineage documentation. Where inputs originate, what governance sits above them, what happens on customer deletion. Technical, not legal, but readable by a lawyer.
- A cross-tool systems map. How internal tools talk to each other, where agents run without human review, where the escalation gates sit. Ramp's public documentation is a strong reference for the shape.
A strong data room in 2026 looks less like a folder of PDFs and more like a set of live dashboards with narrative annotations. The annotations exist because a GP will still read them. The dashboards exist because a GP will now check them. Two failure modes to avoid: operational polish without operational substance, which experienced diligence teams catch inside the first hour with a targeted product-usage reconciliation; and treating the ops layer as a deliverable rather than a running system, so the dashboards go stale between first partner meeting and confirmatory pass.
Through-line
Operational readiness is now a diligence question because the paper stage has caught up to the warm circuit. GPs are no longer diligencing the deck in confirmation of a story they already believe. They are diligencing the ops layer in construction of a story they can defend to their LPs and, if it comes to it, to a court. Founders who build operational legibility as a first-order asset will find diligence shorter and terms tighter. Founders who do not will spend the second partner meeting explaining why the slide numbers do not reconcile to the system underneath them.
Citations
- Silicon.co.uk, "Builder.ai Collapsed After Finding Sales 'Inflated By 300 Percent'"
- eDiscovery Today, "The Biggest AI Fraud in Startup History"
- VC Lab, "AI Tools Every Modern Venture Capital Fund Needs in 2026"
- Nick Talwar, "$4M Revenue Per Employee Is the New"
- Forbes, "AI-Native Firms Lead In Revenue Per Employee"
- Value Add VC, "Cursor AI Valuation: How a Code Editor Became a $9B Company"
- Capwave, "Is Your AI Startup Enterprise-Ready in 2026?"
- CapMaven, "Diligence is the New Pitch"
- Menlo Ventures, "2025: The State of Generative AI in the Enterprise"
- ICONIQ / SaaStr, "The Execution Era of AI"
- Bessemer, "AI Infrastructure Roadmap: Five Frontiers for 2026"
- Sprinto, "Deal Autopsy: How & Why Due Diligence Red Flags Quietly Kill Startup Transactions"
- Blackbird, "What Do LPs Really Look For in GP Due Diligence"
- Startups Magazine, "The new due diligence"
- Gurpreet S. Bal, "The Fraud Wave in Venture-Backed Startups"
- Notion, "How Ramp built an AI operating system for scalable work"
- Global Investigations Review, "US enforcement agencies intensify scrutiny of AI washing"
- IdeaProof, "364+ AI Startups That Failed (2023-2026)"
- CNBC, "Cursor AI Startup Funding Round Valuation"
- SeedScope, "What Investors Want in 2026"
- Lighter Capital, "Revenue per Employee Benchmarks"
- WRITER, "Enterprise AI adoption in 2026"
Related memos