// AI GOVERNANCE / VENDOR REVIEW
AI Vendor Review: Where Your Data Actually Goes
Anthropic's own two newest models, Claude Fable 5.1 and Claude Mythos 5.1, are excluded from Anthropic's zero data retention option. Not eligible by default. Not eligible by self-serve application. Excluded, unless Anthropic grants a case-by-case exception. If you assumed a vendor's privacy page describes the model you are actually calling, that assumption breaks first on the models most companies reach for by default.
This closes the six-part AI governance series that opened with the case against treating governance as a policy document. The first five pieces covered access, review, logging, and audit scope, the surfaces a company controls internally. This one is the outward-facing surface: what you are actually buying when you sign with an AI vendor, and where your data goes after you hit send.
Zero data retention is not one setting
Every major AI vendor now advertises some version of zero data retention. Treated as a single feature you either have or do not, the phrase hides three different gates, and I checked each one directly against the vendor's own current documentation rather than a compliance blog's summary of it.
OpenAI's zero data retention is not self-serve. It requires an Enterprise Agreement, goes through sales for prior approval, and is granted endpoint by endpoint rather than account-wide, currently covering chat completions, the responses API, images, embeddings, audio, completions, moderations, and realtime. Without it, standard API logs are kept up to 30 days for abuse monitoring. On August 19, 2026, OpenAI said it would keep offering the option for its frontier models and previewed Private Safety Processing, a system built to catch risk patterns across related interactions without giving OpenAI staff a way to read the underlying content.
Anthropic's default runs the other direction: conversation content is not retained by default on the Claude API, no application required. Then the exception arrives. Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5 are designated Covered Models, and per Anthropic's own current retention documentation, they require 30-day retention and are not available under zero data retention unless Anthropic expressly authorizes it. A workspace running zero retention that sends a request to a Covered Model gets an error back, not a silent policy override. A customer who needs those specific models can turn on 30-day retention for one workspace and keep every other workspace at zero, which is a genuinely well-built control. It is also a control almost nobody checks for before they pick a model.
Google's paid Gemini API tier does not train on your prompts or responses. But actual zero retention takes more than picking the paid tier: it means turning off state storage on the Interactions API, avoiding session resumption on the Live API, skipping context caching, and staying off grounding tools entirely, since Search and Maps grounding store data for 30 days regardless of anything else you configure. For a guaranteed, contractual zero-retention commitment rather than a product default, Google points customers to Vertex AI and a Data Processing Addendum amendment instead.
Training and retention are two different promises
"We don't train on your data" answers one question. It does not answer whether your prompts are stored, for how long, or who can read them while they sit there. Vendors are not being dishonest when they lead with the training promise. It is the easier promise to keep, and the one buyers ask about most, because "does OpenAI train on my data" gets typed into a search bar a lot more often than "what is OpenAI's log retention window."
The gap shows up clearest at the tier boundary. On Google's Gemini API, an unbilled AI Studio key trains on your inputs and outputs. Attach that same key to a paid billing account and training stops, a distinction Google confirmed as current in June 2026. Anthropic's commercial terms state plainly that Anthropic may not train models on customer content from its paid services, while its separate consumer policy allows an optional extended retention window when a user opts into training. A team that piloted a tool on a personal or free account, liked it, and rolled it onto a business tier did not just upgrade seats. It changed which privacy policy actually governs the account, and almost nobody re-reads the fine print at that moment.
| Vendor | Training default (paid tier) | Retention default | Zero retention gate |
|---|---|---|---|
| OpenAI | No training since March 2023, unless opted in | Up to 30 days, abuse monitoring | Enterprise Agreement, sales approval, per endpoint |
| Anthropic | No training without express permission | Zero by default; 30 days for Covered Models | Automatic except Covered Models, which need an exception |
| Google (Gemini API) | No training once billing is attached | Short window, abuse monitoring only | Requires disabling several features yourself; Vertex AI for a contractual guarantee |
Where your data actually goes: subprocessors
Your contract is with one vendor. Your data's actual exposure runs through everyone on that vendor's subprocessor list: the cloud host, the vector database, the monitoring tool, the translation service quietly sitting behind a feature you didn't know needed one.
OpenAI publishes its subprocessor list and updated it June 2, 2026 with expanded hosting and processing-purpose detail, and it now offers email notifications when a new subprocessor is added, a real improvement over checking a static page on a schedule. Anthropic consolidated its own list into a public Trust Center that launched in January 2026, before which the subprocessor list was not published at all, and added TurboPuffer on May 6, 2026 to support web search across its products. Both changes are recent enough that a vendor review done a year ago is already reading stale information.
The relationship can run deeper than a single hop. OpenAI itself was added to the Microsoft Online Services Subprocessors List on June 23, 2026 and became available for use inside Microsoft's own products on July 9, 2026, meaning a company using Microsoft Copilot inherited an OpenAI dependency it may never have separately reviewed. That is the actual shape of the risk: not "did we vet our AI vendor," but "did we vet everyone our AI vendor depends on, and everyone that vendor's vendor depends on."
What to actually ask a vendor
Five questions, answered in writing, not inferred from a trust badge or a blog post summarizing one.
- What is the training default for the specific product and tier I am buying? Not the company's general policy. The product, the tier, the account type, since free and paid frequently diverge and consumer and commercial almost always do.
- What is the retention default, and is my chosen model excluded from any zero-retention option you advertise? A vendor's homepage claim about zero retention is a starting point. The model-specific exception list is the actual answer.
- Can I see your current subprocessor list, and how do I get notified when it changes? A list with no update mechanism is a snapshot, not a control.
- Will you put the training default and retention default in the DPA, not just the privacy page? A policy page can change without notice. Contract language cannot, without telling you.
- Are you a subprocessor to anything else I already use? The Microsoft and OpenAI relationship is not a special case. It is a preview of how most enterprise AI exposure actually arrives, one layer removed from the tool you thought you were reviewing.
None of this requires distrusting OpenAI, Anthropic, or Google. Every fact in this piece came from each vendor's own current documentation, checked directly rather than taken from a summary, which is exactly the discipline this series has argued for since the audit-scope piece: a vendor's policy page is a claim. Your own dated record of what you checked, and when, is the control.
That is the same test this whole series has run against a policy document, a review gate, a connector, a log, and now a vendor. A claim is not a control until someone can point to the mechanism behind it and the date it was last checked.
Frequently asked questions
Does OpenAI train on my data?
Not by default, since March 1, 2023, when OpenAI stopped using API data to train or improve its models unless a customer explicitly opts in. That is a training promise, not a retention promise: OpenAI still keeps standard API logs for up to 30 days for abuse monitoring. Removing your data from those logs entirely requires Zero Data Retention, which is not self-serve. It requires an Enterprise Agreement, goes through OpenAI's sales team for prior approval, and is granted endpoint by endpoint, not a single account-wide switch. On August 19, 2026, OpenAI said it would keep offering Zero Data Retention for frontier models and previewed Private Safety Processing, built to spot risk patterns across related interactions without giving OpenAI staff access to the content.
What is zero data retention and does every AI vendor offer it the same way?
Zero data retention means a vendor does not store your prompts or the model's responses after the request completes. No two of the three major vendors implement it the same way. OpenAI's requires an Enterprise Agreement and sales approval, applied per endpoint. Anthropic's is the API default, but Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5 are designated Covered Models requiring 30-day retention, not eligible for zero data retention unless Anthropic expressly authorizes an exception. Google's paid Gemini API tier does not train on your data, but true zero retention requires disabling several specific features yourself, since grounding tools store data for 30 days regardless. Google directs customers needing a guaranteed, contractual commitment to Vertex AI instead.
What should an AI vendor due diligence checklist include?
Five things, in writing. The training default for the specific product and tier you are buying, not the vendor's general statement. The retention default and whether your chosen model is excluded from any zero-retention option advertised. A current subprocessor list naming every third party that touches your data. A signed DPA stating the training and retention defaults as contract language, not a webpage that can change without notice. And whether the vendor is itself a subprocessor to something else you already use, since a model provider can sit two or three layers behind a tool you trust without anyone flagging it.
Is the free tier of an AI tool different from the paid tier for data privacy?
Usually, and the difference is not cosmetic. On Google's Gemini API, an unbilled AI Studio key trains on your inputs and outputs; attach that key to a paid billing account and training stops. Anthropic's commercial terms state Anthropic may not train models on customer content from paid services, while its separate consumer policy allows an optional extended retention window when training is enabled. A team that evaluates a tool on a free account and rolls it out on a business tier has evaluated a different privacy posture than the one it is now running in production.
What is a subprocessor and why does it matter for AI vendor review?
A subprocessor is any third party your AI vendor hands your data to in order to deliver the service: cloud hosting, a vector database, a monitoring tool. Your contract is with the vendor, but your data's exposure runs through everyone on that list. OpenAI publishes a subprocessor list updated June 2, 2026 and now offers change notifications. Anthropic consolidated its list into a public Trust Center that launched in January 2026 and added TurboPuffer on May 6, 2026 for web search. A vendor's own privacy promises only cover what the vendor itself does. Everyone on its subprocessor list is a risk you inherited without signing anything directly.
Sources
OpenAI API training and Zero Data Retention eligibility, endpoints, and abuse-monitoring retention window: OpenAI, "Data controls in the OpenAI platform" (checked September 2026). OpenAI's August 19, 2026 announcement continuing Zero Data Retention for frontier models and previewing Private Safety Processing: OpenAI, "Offering Zero Data Retention for frontier models", corroborated by Techstrong.ai and OpenAI's own August 19, 2026 announcement post on X. OpenAI subprocessor list and June 2, 2026 update: OpenAI, Sub-processor list. OpenAI's addition to Microsoft's subprocessor list: Microsoft Learn, "OpenAI as a subprocessor in Microsoft Online Services". Anthropic's retention defaults, Zero Data Retention scope, and the Covered Models list requiring 30-day retention: Anthropic, "API and data retention" (Claude Platform Docs, checked September 2026). Anthropic commercial no-training terms: Anthropic, Commercial Terms of Service. Anthropic Trust Center launch and subprocessor updates: Anthropic Trust Center. Google Gemini API training default by tier and Zero Data Retention configuration requirements: Google AI for Developers, "Zero data retention in the Gemini Developer API" (checked September 2026).
About the author
Jeff Brokaw is a sitting CMO and Certified Chief AI Officer who ships AI in production, not slideware. He has been building AI systems commercially since 2016. He built the commercial engine behind $185M in new-business revenue for a defense manufacturer, and authored the go-to-market behind a $114M institutional raise that came together in under 30 days.