Direct answer
Before approving an AI vendor, request evidence that connects the product's claims to accountable ownership, documented data practices, testing, security, human oversight, monitoring and incident response. A policy statement alone is not enough. Buyers need versioned documents, named owners, test results, operating records and contract terms that can be checked again at renewal.
This checklist is a procurement and governance framework, not legal advice and not a certification audit. Adapt it to the use case, affected people, sector, jurisdictions and risk level. High-impact decisions require deeper legal, security, privacy and technical review.
The 25-document evidence request
| Evidence item | What it should answer | Review signal |
|---|---|---|
| 1. AI system description | What the system does, for whom and in which workflow | Scope is specific enough to test |
| 2. Intended-use statement | Which decisions and users the product supports | Intended and prohibited uses are explicit |
| 3. System and data-flow diagram | Where data enters, moves, is stored and leaves | Subprocessors and boundaries are visible |
| 4. Ownership matrix | Who owns product, risk, privacy, security and incidents | Named accountable roles exist |
| 5. Model and dependency inventory | Which models, APIs, datasets and external services are used | Material dependencies are versioned |
| 6. Data-source register | Where training, tuning, retrieval and customer data originate | Rights and provenance can be explained |
| 7. Customer-data terms | Whether customer data trains models or is retained | Defaults, choices and deletion are clear |
| 8. Privacy assessment | Which personal-data risks were assessed | Use-case risks and mitigations are recorded |
| 9. Security architecture | How access, encryption, isolation and secrets are controlled | Controls match the architecture |
| 10. Secure-development evidence | How code, models and dependencies are reviewed | Testing is continuous, not one-off |
| 11. Threat model | Which abuse, prompt, model and integration threats were considered | Scenarios reflect the actual product |
| 12. Evaluation plan | What quality and risk characteristics are measured | Metrics map to intended use |
| 13. Evaluation results | How the current version performed | Results include limits and uncertainty |
| 14. Bias and impact assessment | Who may be affected differently and why | Populations and contexts are named |
| 15. Robustness and stress tests | How the system behaves outside normal inputs | Failure boundaries are documented |
| 16. Human-oversight design | Where a person reviews, overrides or stops the system | Authority and escalation are usable |
| 17. Transparency materials | What users are told about AI involvement and limitations | Notices match actual operation |
| 18. Logging specification | Which inputs, outputs, decisions and changes are recorded | Logs support investigation without excess collection |
| 19. Monitoring plan | Which production metrics and risks are watched | Thresholds and owners are defined |
| 20. Incident-response plan | How incidents are detected, contained and communicated | Timelines and responsibilities are explicit |
| 21. Change-management policy | What triggers retesting or customer notice | Model and policy changes are controlled |
| 22. Business-continuity plan | How service failure, supplier loss or model withdrawal is handled | Exit and fallback routes exist |
| 23. Independent assurance | Which audits, certifications or reviews apply | Scope and date are verifiable |
| 24. Contract control schedule | Which promises become enforceable obligations | Audit, notice, deletion and liability terms align |
| 25. Renewal evidence pack | What will be refreshed before renewal | Evidence has owners and expiry dates |
Why an evidence pack matters
AI governance is often presented as a policy exercise. Procurement needs an operating view. The central question is not whether a vendor has a responsible-AI page; it is whether the buyer can trace a material claim to current evidence and an accountable process.
NIST organizes AI risk work through four connected functions: Govern, Map, Measure and Manage. Governance establishes policy, roles and accountability. Mapping establishes context and affected parties. Measurement evaluates risk and performance. Management prioritizes and treats risk. The evidence request should therefore cover all four functions rather than focusing only on model accuracy.
ISO/IEC 42001 applies a management-system approach to establishing, implementing, maintaining and continually improving AI governance. A certificate may be useful evidence, but buyers still need to inspect its scope, issuing body, covered entity and relationship to the product being purchased.
The EU AI Act uses a risk-based framework. Specific obligations depend on the system and role. The European Commission's Article 50 transparency guidance, published in July 2026, addresses transparency obligations that apply from 2 August 2026 for certain providers and deployers. A buyer should determine applicability with qualified counsel rather than copying a generic checklist into a contract.
Step 1: define the use case before evaluating the vendor
A vendor cannot provide meaningful evidence if the buyer has not defined the decision context.
Document:
- business purpose;
- users and operators;
- people affected by outputs;
- input data;
- decisions supported or automated;
- consequences of error;
- required human review;
- jurisdictions;
- integration points;
- fallback process.
The same model can be low-risk in one workflow and unacceptable in another. A text assistant that drafts internal meeting summaries is not equivalent to a system that recommends employment, credit, health or public-service decisions.
Step 2: classify evidence by strength
Use four evidence levels:
- Claim: a statement on a website or sales document.
- Document: a policy, architecture note or test report.
- Operating record: logs, tickets, review minutes, monitoring results or change records showing the process is used.
- Independent evidence: a scoped audit, certification, assessment or test by a competent external party.
No level is automatically sufficient. Independent assurance can be too broad or outdated. Operating records can expose sensitive information. Claims can still be useful at early screening. Record which level is required for each risk.
Step 3: connect each claim to a control and an owner
Create a traceability table:
| Vendor claim | Evidence received | Control owner | Validation method | Expiry or review date |
|---|---|---|---|---|
| Customer data is not used for model training | Contract clause, product setting and data-flow diagram | Privacy lead | Configuration test and contract review | Annual or on material change |
| Human review is available | Workflow design and role matrix | Business process owner | Scenario test | Before launch and quarterly |
| Outputs are monitored | Monitoring plan and sample report | Product risk owner | Metric and alert review | Monthly |
This prevents a familiar procurement failure: approving a persuasive claim without knowing how it will remain true after a model, subprocessor or configuration changes.
Step 4: test the evidence against realistic failure scenarios
Ask the vendor to walk through scenarios such as:
- a user submits confidential or prohibited data;
- the model produces harmful or fabricated output;
- a third-party model changes without notice;
- a monitoring threshold is crossed;
- an affected person challenges an outcome;
- a regulator or customer requests an explanation;
- the service becomes unavailable;
- the contract ends and data must be returned or deleted.
The purpose is not to demand perfect prediction. It is to see whether responsibilities, records, controls and communication routes work together.
Step 5: make approval conditional and time-bound
Use an approval record containing:
- approved use case;
- prohibited uses;
- required controls;
- residual risks;
- named business owner;
- evidence gaps and deadlines;
- monitoring frequency;
- incident route;
- material-change triggers;
- renewal date.
Approval of one use case should not silently authorize every use of the product. Material changes in models, data, decisions, integrations or affected populations should trigger review.
A simple scoring model
Score each evidence item from 0 to 3:
- 0 - absent: no evidence;
- 1 - claim: assertion without supporting detail;
- 2 - documented: current and scoped document;
- 3 - operating evidence: current document plus records, testing or independent assurance.
Do not turn the total into an automatic approval. Weight critical controls and preserve stop conditions. A vendor with strong documentation but no acceptable data terms may still be unsuitable.
How this connects to MTF business services
MTF Institute's AI Governance and Responsible AI practice helps organizations define AI inventories, accountability, evidence and operating controls. The Vendor Trust and Compliance practice focuses on reusable evidence packs, customer due diligence and trust operations.
Teams building procurement capability can also explore the Strategic Procurement, Sourcing and Vendor Risk program and the AI, Digital Transformation and Platform Strategy program.
These services and programs support professional capability. They do not replace legal advice, regulatory decisions, certification audits or sector-specific assurance.
Frequently asked questions
Is an ISO/IEC 42001 certificate enough?
No single certificate answers every product and use-case question. Check the certificate's scope, validity, covered entity and issuing body, then connect it to product-specific evidence.
Should a buyer request a model card?
Often yes, but names vary. The useful document describes intended use, limitations, evaluation, data considerations and current version. It must still be checked against the deployed product and workflow.
What if the vendor cannot disclose sensitive technical details?
Agree on proportionate alternatives: redacted evidence, independent assurance, a controlled review, contract warranties or a higher-level test. Confidentiality does not remove the buyer's need for evidence.
How often should evidence be refreshed?
Set dates by risk and change frequency. Review at renewal and after material changes, incidents, regulatory developments or significant performance shifts.
Can this checklist be used for every AI system?
It is a starting framework. Adapt depth and ownership to the system, sector, role, impact and jurisdiction. Some uses require specialist legal, privacy, cybersecurity, safety or fundamental-rights assessment.
Procurement decision record
Before signing, make sure the file contains:
- a defined and approved use case;
- a completed evidence request with gaps;
- named business, privacy, security and risk owners;
- product-specific testing;
- data and subprocessor terms;
- human-oversight and escalation routes;
- monitoring and incident obligations;
- material-change notification;
- exit, deletion and continuity provisions;
- a dated renewal review.
Good AI governance does not mean collecting the most documents. It means requesting the evidence that can change a decision, assigning ownership and keeping that evidence current as the system and its context change.
Editorially reviewed for clarity and source currency on .