# AI Vendor Data-Training Terms: An Executive Procurement Guide

> Use TRAIN-9 to turn “we do not train on your data” into a verifiable map of data, purposes, settings, retention, subprocessors and remedies.

- Canonical page: https://mtfinstitute.com/insights/ai-vendor-data-training-terms-executive-procurement-guide/
- Content type: Article
- Editorial category: Articles &amp; Analysis
- Publisher: MTF Institute of Management, Technology and Finance
- Author: MTF Institute Editorial Team- Published: 2026-09-26
- Updated: 2026-09-26
- Language: English
- Topics: AI governance, Third-party risk, Data Governance, AI Vendor Management, AI Procurement

## AI Vendor Data-Training Terms: An Executive Procurement Guide

When an AI vendor is asked whether customer data is used for model training, “no” is not a complete answer. A responsible buyer needs to know which data, which model or service component, for what purpose, under which default, with which subprocessors, for how long, and what happens to derived artefacts after deletion. The practical task is to convert a broad assurance into contractual scope, technical evidence and an operating control.

This guide provides the TRAIN-9 review. It covers nine decision areas: data inventory, purposes, model improvement, product telemetry, retention, subprocessors, security boundaries, rights and remedies, and change control. It is intended for executives, procurement leaders, legal teams, risk owners and product managers who need a shared set of questions. It is not legal advice; contractual conclusions require qualified review in the relevant jurisdictions.

## Why the phrase “we do not train on your data” can mislead

AI services process more than the text a user intentionally submits. Depending on the product, information can include prompts, uploaded files, retrieved documents, system instructions, outputs, user feedback, conversation history, account metadata, usage events, error logs, safety samples, support tickets and administrator configuration. Some elements may be excluded from foundation-model training but retained for abuse monitoring, service improvement or product analytics.

The word “training” can also be narrower than the buyer assumes. A provider may not update a general model with customer content but may use interactions to evaluate models, tune classifiers, improve retrieval, train a customer-specific component or create de-identified statistics. Those activities may be legitimate and well controlled. The procurement problem is ambiguity, not the existence of every secondary use.

The buyer should therefore avoid a binary questionnaire item. Replace it with a data-purpose-control map. The map forces the parties to describe each data class, permitted purpose, legal or contractual basis, storage location, retention period, access, model relationship and deletion behaviour.

## TRAIN-9 overview

Score each area from zero to two. Zero means the answer is absent or purely promotional. One means a documented answer exists but important scope or evidence is missing. Two means the contract, product setting and evidence align. A total of 15–18 supports conditional approval when no critical issue remains. Eleven to fourteen requires remediation or a tightly bounded pilot. Ten or below means the service is not ready for the proposed use. The score is a triage aid; a prohibited use or unacceptable contractual term overrides the total.

## 1. Build the data inventory

List what the service receives directly and indirectly. Ask the business owner to describe realistic use, not an idealised demonstration. Include prompts, attachments, database records, retrieved content, generated outputs, user identifiers, audit events and support material. Classify data by sensitivity and ownership. Mark personal data, confidential information, intellectual property, regulated records and information received under third-party obligations.

Ask the vendor for a product-specific data-flow diagram and data dictionary. A group-level privacy notice is insufficient when several products have different defaults. Verify which fields are stored, which remain transient and which are copied to observability, security or support systems.

Decision rule: do not approve a use case when the organisation cannot say which data will enter the service. An unknown inventory makes every later assurance unstable.

## 2. Separate purposes

Create a row for each purpose: provide the requested service, maintain security, prevent abuse, deliver support, bill the customer, monitor performance, improve the product, evaluate models, personalise the service or train a model. Require a specific data class and retention rule for every purpose.

Purpose language should not be endlessly expandable. Terms such as “improve our services” can cover many activities. Ask whether the purpose is necessary for delivery, optional, controlled by an administrator or dependent on user consent. Record whether opting out changes functionality, price, support or assurance.

Decision rule: a purpose without a bounded data class and control is not an operational answer.

## 3. Define model improvement precisely

Ask four separate questions. Is customer content used to train or fine-tune a shared model? Is it used to evaluate model performance? Is human review permitted for quality or safety? Is it used to improve customer-specific retrieval, ranking or configuration? The answers may differ.

Request the exact default for the contracted account type. Consumer, free, developer, enterprise and regulated offerings often have different terms. Confirm whether an administrator can enforce the setting for all users and integrations. Check whether an application programming interface, embedded assistant and support channel follow the same policy.

Decision rule: accept “no shared-model training” only when the covered data, accounts, endpoints and exceptions are named in binding documentation.

## 4. Examine telemetry and feedback

Operational telemetry can contain content fragments, identifiers and error context. Ask what is captured when a request fails, is flagged or receives negative feedback. Determine whether users can send a conversation to support and whether that action changes access or retention.

Telemetry should be proportionate to a defined purpose. Review masking, sampling, role access, regional storage, retention and deletion. Where the system records complete prompts for debugging, ask whether a content-minimising mode exists and whether it reduces diagnostic capability.

Decision rule: treat logs and feedback as data processing, not harmless metadata, until the vendor demonstrates otherwise.

## 5. Test retention and deletion

Record retention separately for active service data, history, backups, security logs, support copies, evaluation sets and derived artefacts. “Deleted within thirty days” is incomplete unless the trigger and scope are clear. Is deletion triggered by a user action, administrator action, contract termination or expiry? Are backups overwritten on a fixed schedule? Are legal holds possible?

Ask what can be exported before deletion and how completion is evidenced. If customer content entered a model-training pipeline despite the intended settings, determine whether removal is technically possible and what remedy applies. Do not assume that deleting an account removes an already trained model contribution.

Decision rule: approval requires a lifecycle from collection to verified deletion, including exceptions.

## 6. Map subprocessors and locations

Identify providers that host data, provide models, process safety signals, deliver support or operate analytics. Request the current subprocessor list, locations, functions and notification process. Distinguish a vendor’s own model from third-party model routes that may have separate terms.

For cross-border processing, involve privacy and legal specialists. Ask how routing is controlled and whether region selection applies to all data classes or only primary storage. A regional endpoint does not automatically mean that support, security or telemetry stays in the same region.

Decision rule: no unidentified model or processing route should sit behind a general “service provider” label for a material use case.

## 7. Verify security and isolation

Review identity controls, single sign-on, multifactor authentication, role design, encryption, tenant isolation, key management, secure development, vulnerability handling, incident notification, logging and independent assurance. Match depth to the proposed exposure. A public-content drafting tool and a system connected to client records require different evidence.

Ask how prompts and retrieved data are isolated from other tenants, how administrators can restrict connectors, and whether users can create unreviewed external actions. For agentic capabilities, examine permissions, confirmation gates, transaction limits, sandboxing and rollback.

Certificates and audit reports are evidence inputs, not automatic approval. Check scope, date, exceptions and whether the contracted service is included.

Decision rule: technical controls must support the contractual promise. A prohibition that cannot be configured, monitored or audited is weak protection.

## 8. Clarify rights, ownership and remedies

The contract should state rights in customer inputs and generated outputs, subject to applicable law and third-party rights. Review vendor licences to customer content, confidentiality, infringement allocation, warranties, limitations of liability and obligations after unauthorised use.

Ask what happens if data is processed outside the agreed purpose. Remedies may include investigation, notification, deletion, suspension of the affected feature, cooperation, audit evidence and contractual consequences. Ensure the organisation can stop data flow quickly while preserving records needed for incident response.

Decision rule: a promise without a practical remedy leaves the buyer carrying most of the downside.

## 9. Control changes

AI services change quickly. Models, subprocessors, product settings and terms can change during a contract. Define notice periods, material-change criteria, review rights and the ability to disable a new feature. Subscribe named owners to vendor notices and schedule periodic reassessment.

Create change triggers: new data category, new connector, new model provider, new autonomous action, new jurisdiction, material incident, revised retention or new use for improvement. Each trigger should reopen relevant parts of TRAIN-9.

Decision rule: approval applies to a defined configuration and use, not every future version of a vendor’s product.

## Copyable data-purpose-control table

Use these columns:

1. data class;
2. concrete example;
3. sensitivity;
4. source;
5. required service purpose;
6. optional or secondary purpose;
7. shared-model training status;
8. evaluation or human-review status;
9. administrator control;
10. storage location;
11. retention and trigger;
12. subprocessor;
13. access role;
14. deletion behaviour;
15. contractual source;
16. technical evidence;
17. unresolved issue;
18. owner and due date.

Complete the table jointly. Procurement owns commercial coordination, but it cannot infer architecture. Legal interprets terms, but it cannot define the business workflow. Security reviews control evidence, but it does not own the benefit case. The product or business owner remains accountable for the proposed use and user behaviour.

## Contract questions to send to a vendor

Ask the vendor to answer against the named product and account tier:

- Which customer-content and telemetry fields are processed?
- Which fields are used to provide the service, security, support, evaluation, product improvement or model training?
- Are prompts, files, retrieved content, outputs or feedback used to train any shared model?
- Which defaults apply at contract start, and can administrators enforce them?
- Do API, embedded, mobile and support interactions follow identical rules?
- When can humans access content and how is access approved and logged?
- What retention applies to each data class, backup and derived artefact?
- Which subprocessors and model providers receive which information?
- Which processing locations and transfer mechanisms apply?
- How is tenant isolation tested?
- What evidence demonstrates that opt-out controls operate as stated?
- How are incidents and unauthorised secondary uses handled?
- What changes trigger notice and reassessment?

Require answers to cite the governing contract, data-processing addendum, security document or product control. A sales email can clarify, but material commitments should appear in durable, authorised documentation.

## Worked example: proposal drafting assistant

A consulting company wants an assistant to draft proposals from a library of approved case studies. The team initially asks whether the vendor trains on data and receives a “no” response. TRAIN-9 reveals four separate flows: user prompts, retrieved case-study passages, generated drafts and diagnostic logs.

The enterprise account excludes prompts and retrieved passages from shared-model training. However, voluntary thumbs-down feedback can attach the conversation for human review, diagnostic logs retain limited content for thirty days, and a third-party model provider processes requests in a different region. The administrator can disable feedback submission and restrict connectors, but those settings are not on by default.

The team approves a bounded pilot after enabling the controls, using only approved case studies, prohibiting client-specific data, requiring human claim verification and adding a retention review. The outcome is not “vendor safe” in the abstract. It is “this configuration is conditionally approved for this workflow and data class until the review date.”

## Evidence hierarchy

Use the contract and data-processing addendum for binding commitments. Use product documentation and administrator screenshots for configuration. Use architecture diagrams and data dictionaries for flow. Use current assurance reports for tested control scope. Use penetration, isolation or evaluation summaries where proportionate. Use incident history and remediation evidence to understand operational maturity.

Marketing pages and verbal answers are lower in the hierarchy. They can help identify a claim, but they should not overrule specific contract language or observed configuration. Record document versions and capture dates because terms and interfaces change.

## Red flags

Pause when the vendor cannot distinguish enterprise and consumer policies; when “de-identified” is asserted without describing transformation or re-identification controls; when an opt-out depends on every user finding a personal setting; when the service can route data to unspecified models; when deletion excludes unexplained derived data; when subprocessor notice occurs only after a change; or when the vendor refuses to name the document containing a material assurance.

Another red flag is organisational: the buyer has no sanctioned alternative, so employees will use unapproved tools regardless of the policy. Procurement should pair restrictions with a workable approved route, training and support.

## Approval memo template

Write the decision in one page. State the use case, users, data classes, value hypothesis, configuration, prohibited uses, required review, vendor evidence, residual risks, accountable owners, monitoring, incident route, expiry date and decision. Attach the TRAIN-9 table and unresolved actions.

Use one of four outcomes: approved, conditionally approved, pilot only or not approved. Conditional approval must name the condition, owner and deadline. Pilot-only approval must limit population, data and duration. All approvals should expire or reopen when a change trigger occurs.

## What the reader can do next

Replace the binary training question in the next AI procurement review with the data-purpose-control table. Complete TRAIN-9 for the highest-exposure proposed workflow, verify the governing documents and issue a dated approval memo. Leaders building broader capability in procurement, governance, finance and transformation can also review the [Advanced Executive Management &amp; Business Administration programme](https://mtfinstitute.com/programs/advanced-executive-management-business-administration/#enroll) as one relevant learning route. The immediate procurement rule remains simple: approve a defined use, configuration and evidence set—not a brand-level promise.



## Citation

When citing or summarizing this material, link to the canonical HTML page: https://mtfinstitute.com/insights/ai-vendor-data-training-terms-executive-procurement-guide/
