On this page
The short answer
Ask separate questions about processing, storage, training, human access, and deletion. A statement that data is not used for model training does not establish that nothing is retained. Map each provider and application component, then match every public claim to current settings, contractual scope, and observed behavior.
What to take away
- Training use and retention are different questions.
- Follow data beyond the model provider.
- State the scope and date of every verification.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Replace broad claims with specific questions
An AI application may send information to a model, a retrieval store, an observability tool, and a support platform. Each destination can have different purposes and retention behavior. Start with a data-flow map that names the input, destination, account, region where relevant, and responsible owner. Our recommendation is to phrase claims so they can be checked. A promise that an app does not use data is too broad if processing that data is the core function.
Separate training from other processing
OpenAI states that API data is not used to train or improve its models unless a customer explicitly opts in. That statement concerns a specific provider surface and does not describe every SaaS built on it. The application’s own logs, stored files, subprocessors, and support access require separate inspection. Check the actual account configuration and applicable endpoint behavior before repeating a provider claim in your product’s copy. Preserve the date and scope of the supporting evidence.
Evidence: Data controls in the OpenAI platform [1]
Measure verification coverage
A checklist marked complete is weak evidence if the system inventory is incomplete. Define the set of data destinations first, then track which ones have documented settings and tested behavior. Use synthetic records to trace deletion and access controls without exposing customer information. The measures below are proposed audit tracking metrics; they do not prove the absence of every possible data leak. Record unknowns explicitly so incomplete evidence cannot quietly become a passed control.
| Metric | How to calculate or check | Decision it supports |
|---|---|---|
| Destination coverage | Reviewed destinations / identified destinations | Are any data paths unexamined? |
| Claim coverage | Public data claims with current evidence / claims reviewed | Can the marketing language be supported? |
| Deletion verification | Synthetic records removed as expected / deletion test records | Does the documented process behave as stated? |
| Unauthorized access findings | Observed access-control failures in the agreed test scope | What must be fixed before a claim is made? |
Keep the conclusion bounded
An audit conclusion should identify the product version, configuration, tested environment, evidence reviewed, and unresolved gaps. Changes to providers, logging, or account settings can invalidate an earlier conclusion. If a badge is used, link it to a dated verification record and a clear description of what was assessed. Avoid converting a limited review into a claim of universal security. Give the product team an owner and a recheck trigger for each material data-handling commitment.
Limits of the evidence
This is technical evidence-planning guidance, not legal advice or a certification. Provider policies and application configurations change. A scoped review cannot prove that no data is ever retained or that every possible unauthorized access path is absent.
Common questions
Does no training mean no storage?
No. Verify retention, logs, stored application state, and deletion independently from the provider training-use setting.
What should a data-handling badge link to?
A dated record describing the assessed product, configuration, scope, evidence, limitations, and current verification status.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Data controls in the OpenAI platform OpenAI · Living reference · Maintained documentationReviewed: API training-use default and distinction between data controls. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.