A technical DD is not a conversation. It is a document request. Every claim a founder makes about the engineering has an artifact that either proves it or disproves it, and your job before the term sheet is to ask for the artifact instead of the assurance. This checklist pairs each item with what to request and the red flag that should slow the deal. Treat it as an underwriting problem, not a chemistry read on the founder.
What happens by default: you run DD as a call and a deck. “We’re cloud-native.” “We use AI.” “Security is a priority.” Each claim gets accepted without a diagram, an eval, or a pen test. The risks stay hidden until after the wire clears, when they change your returns instead of the price.
What good diligence does: it maps every claim to an artifact. Architecture to a load test. AI to a model, an eval harness, and a cost per call. Security to a current pen test and an SBOM. Team to a bus-factor read of the commit history. Risks surface before the term sheet, when they still change the price.
A finding that changes whether you invest, rather than only what you pay. Core “AI” that is a thin wrapper with no moat and the thesis depends on it. No IP assignment on contractor-built code. Copyleft linked into the shipped product. Undisclosed breach or open criticals with customer data exposed. One of these overrides a clean scorecard everywhere else.
Real debt that a credible team can pay down on a 100-day plan. Low test coverage. A missing SOC 2 that is two quarters away. Cloud waste with no owner. Weak DORA numbers from missing instrumentation rather than a broken culture. A stack of these is a negotiation, not a no. Price them into the valuation.
Not a threat today, but a trend that could become one. Cloud cost creeping faster than revenue. A bus factor of one where the person is still engaged and locked in. A dependency that would hurt if a vendor changed terms. Write these into the 100-day plan so they get an owner before they get expensive.
Send the data-room list before the first technical call. For each assertion, name the artifact that proves it: the diagram, the eval suite, the pen test, the commit log. If an artifact does not exist, that absence is itself a finding.
Get read access to the repo and the commit history. For any AI story, ask what is a trained model versus a prompt on a public API, then ask for the eval harness and the cost per request. A demo is not evidence. An eval and a unit-economics number are.
Score all nine sections, weight security and code quality heaviest, and total to a weighted 125. The band tells you the read: clean, priced risk, heavy, or stop. The number keeps the conversation about evidence instead of gut feel.
Carry every 4 and 5 into the triage. Sort each into deal-breaker, fixable, or monitor. Then act on the split: condition the close on the deal-breakers, price the fixables, and hand the monitors to the 100-day plan.
A strong thesis survives a hard technical read and comes out sharper, with the risks priced and a plan attached. A weak one falls apart the moment you ask for the artifact instead of the answer, which is exactly when you want it to fall apart. The founder’s confident AI story is not the signal. The eval harness, the pen test, and the commit history are. Ask for those before you wire, not after.
No. A strong technical founder expects it and has most of it ready. The reaction to the request is itself a signal. Defensiveness or “we’ll get to that after the round” is a yellow flag, not a personality quirk.
On the two findings that quietly change everything: unit economics that turn negative at scale, and a bus factor of one on the core system. Both are invisible in a deck and obvious in the artifacts. Catching either before the term sheet is worth the whole exercise.
Investors, corporate development teams, and acquiring CTOs writing a check on a software or AI company. Also useful for founders who want to pressure-test their own engineering before they raise.
Ask what they actually own. A fine-tuned or trained model with an eval suite and a unit-cost-per-request is real. A prompt on top of a public API, sold as proprietary AI with no evals, is the pattern behind Gartner’s projected 40%-plus cancellation rate. Make them show the harness.
The checklist structures the ask so anyone can send the document request and score the bands. Reading the repo, testing the AI claims, and judging the architecture takes engineers who ship to production. Run the structure yourself, bring in the read where it counts.