Answer guide

Is our data used to train your AI models?

Of all the questions in the AI section, this is the one where a casual answer does the most damage — because it is a written statement about the buyer's own data. The honest, complete answer has two layers, and the second one lives in a contract that deserves a re-read before you answer: the data processing agreement with your model provider.

Why buyers ask this

The question is not really about machine learning. It is a confidentiality question wearing an ML costume: can our contracts, financials, customer records — whatever we put into your product — end up inside a model that serves other companies, outside our control and beyond deletion? That is why training and data-flow questions are standard in AI due-diligence instruments such as the AI-CAIQ and the SIG 2026 AI domain (see our AI-CAIQ guide) and in many custom AI sections, and why this one can be a hard gate: the answer feeds into the buyer's DPA review, not just the security scorecard.

The honest answer structure

Two layers, then evidence — the same structure as Step 2 of our method for AI sections:

  1. Your own practice. Do you train or fine-tune any model on customer data? If not, say so plainly. If yes, scope it precisely: which data, which purpose, which contractual basis, whether the customer can opt out.
  2. Your model provider's contractual terms. The standard API and enterprise terms of the major model providers generally exclude using customer API data for training unless the customer opts in — but do not assert that from memory. Terms differ between consumer and API tiers of the same provider, and between versions. Pull the data processing agreement you actually signed, confirm what it says about training, and cite it by name and date.
  3. The evidence around it. Retention of prompts and outputs (yours, and the provider's under the DPA), whether the provider appears on your public subprocessor list, and the processing region. If the provider processes in the US, expect the follow-up on international transfers — verify the transfer mechanism in the signed DPA before naming one.

Template — for orientation only. Replace every bracket with facts you have verified. The sentence about your provider is only as good as the agreement you actually signed — confirm it against the signed version before this leaves your outbox.

"Our practice: [your product] does not use customer data to train or fine-tune AI models. Provider terms: customer content submitted to [your product]'s AI features is processed via the [provider] API under [the provider's Data Processing Agreement, version/date], under which content submitted via the API is not used to train [provider]'s models. [Provider] appears on our subprocessor list at [URL]. Retention and location: we retain prompts and outputs for [N days] for [purpose]; retention by [provider] is governed by the same agreement; processing takes place in [region]."

The mistake that costs deals

Answering "no" from memory. The claim "our provider doesn't train on API data" is easy to write from a half-remembered blog post — and hard to stand behind, because its only real source is the agreement you signed. If the buyer's counsel pulls your provider's current terms and they do not match your answer (wrong tier, outdated version, a consumer product's policy quoted for an API product), you have not made a small error: you have made a written misstatement about their data, in a document you may be asked to warrant at contract stage.

The runner-up mistake is answering only about the provider. A reviewer who reads "our model provider does not train on API data" and nothing else will send the obvious follow-up: and do you? Cover both layers in one answer and the question closes; cover one and it becomes a thread.

Mini-FAQ

Do the major AI providers train on data sent through their APIs?

Their standard API and enterprise terms generally exclude training on customer API data unless the customer opts in — but tiers and versions differ, so verify the DPA you actually signed and cite it by name and date rather than answering from memory.

Is a one-word "No" a complete answer?

No. Buyers expect your own practice, the provider's terms, and the retention, subprocessor and region detail around them. A bare "no" reads as unexamined and invites a follow-up round.

What if we do fine-tune on customer data?

Say so precisely: which data, which purpose, which contractual basis, and any opt-out. A scoped "yes, limited to X under Y" survives due diligence; a "no" the buyer later disproves takes the whole questionnaire's credibility down with it.

Related: "Are you compliant with the EU AI Act?" — usually the next question on the same page — plus the first-24-hours playbook and the free Article 50 checker if your product ships AI features under its own brand.

This exact question is sitting in your questionnaire?

Send it — first 3 answers free within 24h, full delivery in 48h for $490 flat up to 60 questions, paid after delivery. Every legal claim cited to its source. Judge the work on the public sample first.

Send your questionnaire →