askOdin — AI Judgment Infrastructure for Capital Allocation

FIELD NOTE

The Fork in Data Retention: Two AI Giants, One Problem, Opposite Answers

Within 48 hours, OpenAI and Anthropic published opposing answers to the same question. Here is what it means for anyone running confidential deal documents through an AI vendor, including us.

By YekSoon Lok, Founder & CEO · Published · Updated · 9 min read

On 19 August 2026, OpenAI previewed something it calls Private Safety Processing: a system meant to catch misuse that only becomes visible across several related interactions, without keeping the content those interactions were made of.

The next day, Reuters reported that Anthropic was changing its enterprise retention policy. The company will still require 30 days of retention, but customers will be able to hold the data on their own cloud infrastructure.

Two announcements, 48 hours apart, are not a coincidence. They are two answers to the same question, and the question is genuinely hard: as models get powerful enough to be useful in a cyberattack, how does a vendor watch for that without holding its customers’ confidential material?

The two firms have answered it in opposite directions. If you run deal documents through anybody’s AI, the divergence is now your problem too.

Where this started

On 9 June 2026, Anthropic launched Claude Fable 5 and Mythos 5. Alongside the models came a policy: every prompt and every output from those models would be retained for 30 days, with content flagged by safety systems held for up to two years.

The part that made it a story was the absence of an exit. There was no opt-out, no platform exemption, no enterprise carve-out, and prior zero-retention agreements did not survive for Fable 5 traffic. Firms that had negotiated bespoke privacy terms found those terms overridden for the vendor’s most capable models. The available workaround was isolation, not escape: enable retention on one workspace and keep others clean.

The reaction was fast. Within days, Microsoft restricted its own employees from using Fable 5 while legal and compliance reviewed the terms. At the same time it was shipping that model to GitHub Copilot and Azure customers, where it arrived disabled by default and needed an administrator to switch it on. A company can hold both positions at once, but the signal is hard to miss. If Microsoft’s legal team wanted a closer look, a smaller firm’s compliance officer is entitled to the same instinct.

Two architectures, not two marketing positions

Both companies face an identical technical constraint. The most serious misuse patterns are invisible in any single interaction; you only see them across a sequence. Detection therefore requires memory of some kind. The disagreement is about what form that memory takes.

Anthropic’s path is retain, inspect, delete. Content sits in storage for a defined window where automated systems and, where necessary, people can examine it. The August change moves where it sits, onto customer-controlled cloud infrastructure, while keeping the 30-day requirement itself. Reuters reports the design has been in development for months, shaped with more than a hundred customers including Salesforce, with a new safety system expected later this year.

OpenAI’s path is detect, discard, signal. Content stays on customer infrastructure, or on OpenAI’s under customer-managed encryption keys that OpenAI does not hold. Automated systems look for patterns across related interactions and emit what OpenAI describes as narrowly scoped safety signals: the shape of the activity, not the material it was found in. Customers investigate alerts in their own systems and decide what, if anything, to share back. Broader rollout and a technical white paper are due in September.

That distinction is worth sitting with, because it determines what you are actually relying on. A retention policy is a commitment about behaviour: we will hold this, look at it under these conditions, and delete it on this schedule. You are trusting a process, and processes can be audited. A cryptographic guarantee is a claim about what is possible: we cannot read this, because we do not have the key. You are trusting an architecture, and architectures can be verified, but only once the white paper exists and someone competent has read it. September will tell.

Neither is obviously right. Retention gives investigators the material to reconstruct a real attack, which matters if you think the attacks are real. Non-retention removes a class of risk entirely, which matters if you think the likeliest breach is of the vendor holding your data rather than by the attacker they were watching for. Reasonable security teams land in different places. What is no longer available is the assumption that your vendor’s answer matches your policy by default.

What this means at askOdin

We process the most sensitive material our customers have: transaction documents, financial models, cap tables, the reasoning behind investment decisions. Publishing an article that tells you to interrogate your AI vendors, without answering the same questions ourselves, would be worth nothing. So here are our answers, including the one where we come off badly.

We call one model, for one job. We use the Google Gemini Developer API to extract structured data points from unstructured documents. That is the model’s entire role. We do not ask it to analyse, summarise, or form a view. Once variables are extracted, our deterministic engine takes over and does the actual work of compiling claims against business physics and producing a verdict. The architecture is model-agnostic behind an adapter layer, so the vendor can change without the logic changing, but today it is Gemini and we would rather name it than say “a leading model provider.”

Google does not train on it. We run on the paid tier, under which prompts and responses are not used for product improvement or model training. Google offers a Zero Data Retention option on that tier by request, which clears user content and identifiable metadata before logging. We have not yet obtained it. We should, and it is now on the list.

Raw documents are held for up to 30 days. Not destroyed the instant the audit finishes. Thirty days is a ceiling, and it is there so an audit can be reconstructed while it still matters.

And here is the part most vendors would leave out. That 30-day ceiling is a documented policy. It is not yet an automated control. The purge job is scheduled work that lands with our SOC 2 Type I process; today, deletion is something we carry out deliberately rather than a cron job you could point an auditor at. That gap matters most to the people already in the system. Crucible is free and live, founder submissions sit in it today, and the structural record derived from them carries their names. Those founders are data subjects. They have erasure rights, and they are owed exactly the control that an institution running a diligence file through us later would demand. The job is in build now, and until it lands we will not label that ceiling as anything other than what it is. When it ships we will say so on the security page, which is also where you will find which of our controls are implemented and which are still in progress.

We could have written “documents are automatically and permanently deleted after 30 days.” It reads better. It is also the exact sentence a SOC 2 auditor asks you to evidence, and we could not. A vendor who will not tell you the difference between a policy and a control is a vendor who will tell you other things that are not quite load-bearing either.

We do not read your documents. This is a procedural control, not an architectural one, and the distinction matters. We have not built a system that makes access technically impossible; we have a policy that staff do not access customer material. Access logging is built but not yet wired up, which I mention because I originally wrote that it was. If someone tells you their engineers cannot reach your data, ask them what enforces that. Sometimes the honest answer is a rule rather than a cryptographic boundary, and the rule is fine, as long as nobody dresses it up as mathematics.

And “who can access it” is not only a question about staff, which is the half most vendors answer and stop. On our shared processing tier, cache isolation between tenants is being hardened and that work is not finished. Dedicated instances are the answer today. If strict tenant isolation gates your firm, ask us where that stands before you upload anything. Ask every other vendor the same question in the same two halves, because a clean answer about employees tells you nothing about the customer in the next row.

What persists is derived, not documentary. After the retention window, what remains is the verdict, the judgment analysis, and structural data that calibrates the Judgment Graph, our benchmark corpus, built on public deal data. Your documents are not in it. The audit trail outlives the document, which is the point: an LP asking in three years how a decision was reached needs the reasoning, not the source PDF.

That structural record carries identifying details such as founder names, so we treat it as personal data and your erasure rights reach it. Stripping those details systematically is on the roadmap and is not yet built. An earlier version of this article described the corpus as carrying no identifying detail; an engine audit run days after publication established otherwise, and the sentence you are reading is the correction. That is the cost of the position I took above. If you publish which of your controls are real, you also have to publish it when you get one wrong.

The distinction almost everyone misses

Here is the thing the last week of coverage has largely got wrong.

OpenAI and Anthropic are arguing about infrastructure-layer security. Where the data sits. Who can read it. How long it stays. That argument matters and you should follow it.

But a retention policy tells you nothing about whether a conclusion is trustworthy.

A model can retain nothing at all and hand you an untraceable answer you cannot reproduce, cannot source, and cannot defend to an investment committee. Another can hold your data for 30 days and produce a verdict with every claim tied to a page and a line. Privacy at the infrastructure layer and provenance at the judgment layer are different properties. Buying one does not get you the other, and vendors are happy to let you assume otherwise.

This is why we split those layers deliberately. The model handles extraction and never forms a view. The deterministic engine forms the view, and because it is deterministic, the same inputs produce the same verdict, which is what makes it auditable at all. LLMs optimize for persuasion. We compile for physics. That property is unaffected by where anybody stores anything.

The questions to ask your vendors

If you are evaluating AI tools for investment decisions, this is the list.

QuestionWhy it matters
Where is my data stored?Determines jurisdiction, and therefore which regulator has a say
Who can access it?Personnel, process, and permission boundaries
How long is it retained?Does it fit your own governance policy, or quietly override it?
How is it deleted?The end of the lifecycle is where policies are usually vaguest
Is the processing scoped?Whole documents to a model, or only the fields required?
Are there separate abuse-monitoring logs?The model platform may keep logs your vendor does not control and cannot delete
Is the judgment auditable?Even with data deleted, is the conclusion reproducible and sourced?

Then one more, which cuts across all seven and is the one we would most want asked of us:

For each control you just named, is it implemented, or is it written down? A stated policy is not an implemented control, and an implemented control is not an evidenced one. Any vendor can produce a document. Ask for the artifact showing the control executed. The answer tells you less about their security than about whether they will be straight with you when something is not finished. Over the life of a relationship, that is worth considerably more.

Most AI vendors cannot answer the last three questions clearly. We have tried to answer all eight above, including where the answer is “not yet.”

Venture capital is the last unaudited asset class. askOdin provides the infrastructure to close the gap. An audit infrastructure that will not audit itself in public is not worth much.

A Dialogue on Institutional Judgment

The Judgment Gap is an existential threat to funds facing the mathematical crisis of scaling capital and deal flow. In the AI era, running on artisanal, unscalable judgment processes is no longer a viable strategy. We are building the infrastructure to solve this.

If you are a partner or principal at a growing venture capital fund and are committed to building a more scalable, defensible, and rigorous investment process, we invite you to a confidential discussion.

Share this Insight


Sources. OpenAI, Offering Zero Data Retention for frontier models, 19 August 2026. Reuters, Anthropic plans to change enterprise data retention policy, source says, 20 August 2026. Bloomberg, Anthropic Plans to Change Data Retention Policy for Advanced AI, 20 August 2026. Reporting on Anthropic’s 9 June 2026 Fable 5 / Mythos 5 retention policy and Microsoft’s internal restriction, June 2026. Google, Zero data retention in the Gemini Developer API and Data Logging and Sharing, ai.google.dev.