LPs should ask their GPs seven questions about AI diligence. What standard does it apply? Can you show me the reasoning? What happens when a claim fails? Where does the deal data go? Who is accountable? Is the verdict reproducible? Can you show me a real audit? None of them needs a technical background. Together they separate a diligence process an LP can examine from an expensive summary engine.
LP checklist
Seven questions to ask your GP about AI diligence
- 01 What standard does your diligence apply? Good answer: It names the tests a claim must pass before anyone relies on it.
- 02 Can you show me the reasoning behind a conclusion? Good answer: Every claim traces to a source document, and contradictions are named.
- 03 What happens when a claim fails the test? Good answer: The system flags the violation and stops. It does not smooth it over.
- 04 Where does the deal data go? Good answer: Named subprocessors, a stated retention window, no training, a written policy.
- 05 Who is accountable for the output? Good answer: A named partner signs off, with the reasoning trail attached.
- 06 Is the verdict reproducible? Good answer: Same recorded inputs, same verdict; a past verdict checked from its record.
- 07 Can you show me a real audit? Good answer: A redacted reasoning trail from the fund's own pipeline, not a vendor demo.
General partners have integrated AI into their diligence process at speed: summarization, extraction, first-pass screening, memo drafting. The tools are useful and the adoption is rational.
Oversight has not kept pace. The ILPA Due Diligence Questionnaire, the industry’s reference template for GP due diligence, was last updated in 2021, more than a year before ChatGPT’s public release.
So limited partners are being asked to trust a process they cannot see, run on tools they have not evaluated, producing conclusions they cannot independently verify. That is not a technology problem. It is a fiduciary problem layered above the one the GP is solving.
The gap is structural. GPs are accountable to LPs for their process. But when that process runs through a probabilistic system, the GP often cannot explain how a conclusion was reached, whether the reasoning was tested against anything, or what happened when a claim failed.
Most LPs have not asked, because the questions sounded technical. They are not. They are the questions a fiduciary has always been required to answer. Here are seven of them.
1. What standard does your diligence apply?
Not which AI tool you use. What standard a claim must meet before you rely on it.
Every fund has an implicit standard. The only question is whether it is explicit and testable. “We reviewed it carefully” is not a standard. “Unit economics must be internally consistent, revenue must have a commercial basis, and the cap table must not carry terminal governance risk” is.
A good answer names specific tests: revenue claims cross-checked against contracts, unit economics modeled at scale, cap table structure tested against governance thresholds. The GP can state what a claim must pass.
A weak answer describes the process without naming the standard. “We use AI to be more efficient.” “The team reviews every output.”
2. Can you show me the reasoning behind a conclusion?
Not the output. The reasoning.
If a GP tells you they passed on a deal because “the unit economics didn’t work,” ask what specifically didn’t work, and on what evidence. If the answer is a summary, you have an impression. If it is a claim tied to a source, traced to a document, with the contradiction preserved rather than smoothed over, you have a reasoning trail.
A good answer traces the conclusion to the underlying evidence. Every claim has a source. Every contradiction is named. The GP can point to the line in the spreadsheet that broke the thesis.
A weak answer is a summary with no traceability. The reasoning felt right. The model said so.
3. What happens when a claim fails the test?
This is the most important question on the list, and the one that separates two entirely different kinds of system.
A large language model, meeting a contradiction, produces a smoother answer. It is built to optimize for coherence: LLMs optimize for persuasion. A deterministic system does the opposite. It halts and flags the violation.
If your GP’s diligence tool is probabilistic, the failure mode is silent smoothing of exactly what you needed to see.
A good answer: “The system flags the violation and stops. We see the contradiction.” The GP can describe what a failure looks like.
A weak answer: “It summarizes the data room.” Summarization hides contradictions. It is the mechanism by which a fatal flaw becomes a plausible narrative.
4. Where does the deal data go?
Diligence runs on confidential material: cap tables, financial models, LP letters, legal agreements.
Ask where that data is processed, who can access it, how long it is retained, and whether any model is trained on it.
A good answer names the subprocessors, states the retention window, confirms the data is not used for training, and points you to a written policy. The best answers also say which of those statements are policy and which are enforced by the system, because they are not the same thing.
A weak answer is vague on any of those points. “It’s all secure.” “We use a leading provider.” Vagueness is not a security posture.
The LP has its own duty to the confidentiality of the underlying portfolio. If the GP cannot state where the data goes, the LP is carrying an unquantified risk.
5. Who is accountable for the output?
“A human reviews it” and “a human is accountable for it” are different statements. Only the second is a governance structure. The first is a process step. The second is a person who can defend the reasoning.
A good answer names the accountable party. “The deal partner signs off on every conclusion, and the reasoning trail is attached.” The GP can explain how the chain of accountability works.
A weak answer is generic. “The team reviews it.” “We have a process.” Nobody named, no signature required.
6. Is the verdict reproducible?
Run the same deal through the process twice. Do you get the same conclusion? And when you want to check a verdict from last year, what do you check it against?
A deterministic evaluation gives the same verdict for the same inputs under the same rules. That is what makes an audit defensible after the fact. But if an AI model reads the documents and extracts those inputs, the extraction can change when the model does. So the honest test of a past verdict is its record: the inputs that were scored, the version of the rules, and a fingerprint of the documents read.
A probabilistic system has no obligation to give the same answer twice. That is fine for creative work. It is not fine for diligence.
A good answer: “Yes. The evaluation is deterministic: the same recorded inputs under the same rule version return the same verdict. Each verdict keeps that record, so we can show you how any past verdict was reached months later. Extraction runs on a model, and we can tell you which one.”
A weak answer: “The AI helps, but a human always makes the final call.” That avoids the question. A human signing off on a probabilistic output does not make the reasoning underneath it reproducible.
7. Can you show me a real audit?
Not a demo. Not a sample from a vendor’s website. A real audit from the GP’s own pipeline.
Ask to see the reasoning trail for one deal you know something about, ideally one pass and one investment.
A good answer is a document: a reasoning trail with the sources attached.
A weak answer is an explanation of why they cannot share one. Confidentiality is a real constraint, and also a convenient one. If the process produces audit artifacts, they exist and can be redacted. If they do not exist, the process is producing summaries, not verdicts.
It is also worth checking that the AI is there at all. In March 2024 the SEC settled charges against two investment advisers for claiming AI capabilities they did not have, and named the practice “AI washing”. A GP who describes an AI-driven diligence process should be able to show you its output.
How do I put these questions into a DDQ?
The checklist above is selectable text; copy it as it stands, or use this wording, which names no vendor and fits alongside the ILPA template:
- Describe the standard a claim must meet before your diligence process relies on it, including the specific tests applied.
- For one recent investment and one recent pass, provide the reasoning trail that links each material conclusion to its source documents.
- Describe what your diligence process does when a claim fails a test, with an example.
- For any AI-assisted step, list the subprocessors that handle deal data, the retention period, and whether any data is used for model training. State which of these are written policy and which are enforced by the system.
- Name the person accountable for conclusions produced with AI assistance and describe the sign-off.
- State whether the same recorded inputs produce the same diligence verdict under the same rule version, how a past verdict is checked from its record, and which steps (for example AI extraction) can vary between runs.
- Provide a redacted audit record from your own diligence pipeline.
What do the answers reveal?
They do not tell you whether your GP is a good investor. They tell you whether your GP’s process is defensible.
Those are different questions, and only one of them is a fiduciary one.
A GP with a defensible process may still make bad calls; that is the nature of the business. But you can examine, question and improve the process. A GP with an indefensible process is running a system you cannot evaluate. Their judgment might be excellent. You have no way to know.
What does the input look like?
Here is the math on what arrives. In our benchmark of 2,488 pitch decks, scored by one engine version over four weeks, the median Clarity Score was 35 out of 100 and 70% drew a PASS verdict, the bottom band below 50. Business Model Physics was the weakest section by a wide margin: 67% of decks scored below half marks on it.
That is the input side, and it is not the GP’s fault. Catching it is the GP’s job. A tool that summarizes a weak business model will describe it fluently. It will not tell you that it does not close.
The standard LPs should set
If a GP cannot answer the seven questions, it is not running a defensible process. It may be running a good one. It may be running one that has produced exceptional returns. But it is not running one that an LP can independently evaluate.
That is the standard. Not that every fund should use the same tools, but that every fund should be able to explain its process in terms an LP can examine. The LP Governance Standard sets out how to write that into side-letter terms and quarterly reporting, and The LP’s Blind Spot makes the case for why decision architecture, not track record, is the leading indicator. If you want to test a manager’s reasoning yourself during selection, rather than ask about it, due diligence for allocators shows how LPs and family offices run that audit.
LPs are fiduciaries too. The capital you allocate carries its own duties, and one of them is oversight. GPs who adopt AI with rigor will be the ones worth backing. GPs who adopt it only for efficiency will be the ones whose process you cannot examine.
Ask the seven questions. The answers will tell you which is which.
Venture capital is the last unaudited asset class. askOdin provides the infrastructure to close the gap.
Methodology and defensibility
askOdin audits reasoning, not truth. We do not predict market outcomes or guarantee operational success. We compile investment narratives against structural constraints so that capital is not deployed on unexamined assumptions.
The compiler runs on four protocols, each the subject of a U.S. provisional patent application:
- RUNE Protocol™: domain-specific narrative compiler (U.S. Provisional Patent No. 63/948,559)
- RAVEN Protocol™: cross-document triangulation (U.S. Provisional Patent No. 63/994,876). The architectural mechanics of RAVEN’s triangulation engine are protected under U.S. Provisional Patent No. 63/994,876 and are not publicly disclosed.
- NORN Protocol™: temporal semantic drift detection (U.S. Provisional Patent No. 64/011,252)
- JUDGE Protocol™: runtime circuit breaker (U.S. Provisional Patent No. 64/017,488)
This essay builds on the working paper The Last Mile of AI: Judgment Infrastructure, Defensible Audit Logs, and the End of Information Retrieval (SSRN).