Direct answer

Before approving an AI vendor, request evidence that connects the product's claims to accountable ownership, documented data practices, testing, security, human oversight, monitoring and incident response. A policy statement alone is not enough. Buyers need versioned documents, named owners, test results, operating records and contract terms that can be checked again at renewal.

This checklist is a procurement and governance framework, not legal advice and not a certification audit. Adapt it to the use case, affected people, sector, jurisdictions and risk level. High-impact decisions require deeper legal, security, privacy and technical review.

The 25-document evidence request

Evidence item What it should answer Review signal
1. AI system description What the system does, for whom and in which workflow Scope is specific enough to test
2. Intended-use statement Which decisions and users the product supports Intended and prohibited uses are explicit
3. System and data-flow diagram Where data enters, moves, is stored and leaves Subprocessors and boundaries are visible
4. Ownership matrix Who owns product, risk, privacy, security and incidents Named accountable roles exist
5. Model and dependency inventory Which models, APIs, datasets and external services are used Material dependencies are versioned
6. Data-source register Where training, tuning, retrieval and customer data originate Rights and provenance can be explained
7. Customer-data terms Whether customer data trains models or is retained Defaults, choices and deletion are clear
8. Privacy assessment Which personal-data risks were assessed Use-case risks and mitigations are recorded
9. Security architecture How access, encryption, isolation and secrets are controlled Controls match the architecture
10. Secure-development evidence How code, models and dependencies are reviewed Testing is continuous, not one-off
11. Threat model Which abuse, prompt, model and integration threats were considered Scenarios reflect the actual product
12. Evaluation plan What quality and risk characteristics are measured Metrics map to intended use
13. Evaluation results How the current version performed Results include limits and uncertainty
14. Bias and impact assessment Who may be affected differently and why Populations and contexts are named
15. Robustness and stress tests How the system behaves outside normal inputs Failure boundaries are documented
16. Human-oversight design Where a person reviews, overrides or stops the system Authority and escalation are usable
17. Transparency materials What users are told about AI involvement and limitations Notices match actual operation
18. Logging specification Which inputs, outputs, decisions and changes are recorded Logs support investigation without excess collection
19. Monitoring plan Which production metrics and risks are watched Thresholds and owners are defined
20. Incident-response plan How incidents are detected, contained and communicated Timelines and responsibilities are explicit
21. Change-management policy What triggers retesting or customer notice Model and policy changes are controlled
22. Business-continuity plan How service failure, supplier loss or model withdrawal is handled Exit and fallback routes exist
23. Independent assurance Which audits, certifications or reviews apply Scope and date are verifiable
24. Contract control schedule Which promises become enforceable obligations Audit, notice, deletion and liability terms align
25. Renewal evidence pack What will be refreshed before renewal Evidence has owners and expiry dates

Why an evidence pack matters

AI governance is often presented as a policy exercise. Procurement needs an operating view. The central question is not whether a vendor has a responsible-AI page; it is whether the buyer can trace a material claim to current evidence and an accountable process.

NIST organizes AI risk work through four connected functions: Govern, Map, Measure and Manage. Governance establishes policy, roles and accountability. Mapping establishes context and affected parties. Measurement evaluates risk and performance. Management prioritizes and treats risk. The evidence request should therefore cover all four functions rather than focusing only on model accuracy.

ISO/IEC 42001 applies a management-system approach to establishing, implementing, maintaining and continually improving AI governance. A certificate may be useful evidence, but buyers still need to inspect its scope, issuing body, covered entity and relationship to the product being purchased.

The EU AI Act uses a risk-based framework. Specific obligations depend on the system and role. The European Commission's Article 50 transparency guidance, published in July 2026, addresses transparency obligations that apply from 2 August 2026 for certain providers and deployers. A buyer should determine applicability with qualified counsel rather than copying a generic checklist into a contract.

Step 1: define the use case before evaluating the vendor

A vendor cannot provide meaningful evidence if the buyer has not defined the decision context.

Document:

  • business purpose;
  • users and operators;
  • people affected by outputs;
  • input data;
  • decisions supported or automated;
  • consequences of error;
  • required human review;
  • jurisdictions;
  • integration points;
  • fallback process.

The same model can be low-risk in one workflow and unacceptable in another. A text assistant that drafts internal meeting summaries is not equivalent to a system that recommends employment, credit, health or public-service decisions.

Step 2: classify evidence by strength

Use four evidence levels:

  1. Claim: a statement on a website or sales document.
  2. Document: a policy, architecture note or test report.
  3. Operating record: logs, tickets, review minutes, monitoring results or change records showing the process is used.
  4. Independent evidence: a scoped audit, certification, assessment or test by a competent external party.

No level is automatically sufficient. Independent assurance can be too broad or outdated. Operating records can expose sensitive information. Claims can still be useful at early screening. Record which level is required for each risk.

Step 3: connect each claim to a control and an owner

Create a traceability table:

Vendor claim Evidence received Control owner Validation method Expiry or review date
Customer data is not used for model training Contract clause, product setting and data-flow diagram Privacy lead Configuration test and contract review Annual or on material change
Human review is available Workflow design and role matrix Business process owner Scenario test Before launch and quarterly
Outputs are monitored Monitoring plan and sample report Product risk owner Metric and alert review Monthly

This prevents a familiar procurement failure: approving a persuasive claim without knowing how it will remain true after a model, subprocessor or configuration changes.

Step 4: test the evidence against realistic failure scenarios

Ask the vendor to walk through scenarios such as:

  • a user submits confidential or prohibited data;
  • the model produces harmful or fabricated output;
  • a third-party model changes without notice;
  • a monitoring threshold is crossed;
  • an affected person challenges an outcome;
  • a regulator or customer requests an explanation;
  • the service becomes unavailable;
  • the contract ends and data must be returned or deleted.

The purpose is not to demand perfect prediction. It is to see whether responsibilities, records, controls and communication routes work together.

Step 5: make approval conditional and time-bound

Use an approval record containing:

  • approved use case;
  • prohibited uses;
  • required controls;
  • residual risks;
  • named business owner;
  • evidence gaps and deadlines;
  • monitoring frequency;
  • incident route;
  • material-change triggers;
  • renewal date.

Approval of one use case should not silently authorize every use of the product. Material changes in models, data, decisions, integrations or affected populations should trigger review.

A simple scoring model

Score each evidence item from 0 to 3:

  • 0 - absent: no evidence;
  • 1 - claim: assertion without supporting detail;
  • 2 - documented: current and scoped document;
  • 3 - operating evidence: current document plus records, testing or independent assurance.

Do not turn the total into an automatic approval. Weight critical controls and preserve stop conditions. A vendor with strong documentation but no acceptable data terms may still be unsuitable.

How this connects to MTF business services

MTF Institute's AI Governance and Responsible AI practice helps organizations define AI inventories, accountability, evidence and operating controls. The Vendor Trust and Compliance practice focuses on reusable evidence packs, customer due diligence and trust operations.

Teams building procurement capability can also explore the Strategic Procurement, Sourcing and Vendor Risk program and the AI, Digital Transformation and Platform Strategy program.

These services and programs support professional capability. They do not replace legal advice, regulatory decisions, certification audits or sector-specific assurance.

Frequently asked questions

Is an ISO/IEC 42001 certificate enough?

No single certificate answers every product and use-case question. Check the certificate's scope, validity, covered entity and issuing body, then connect it to product-specific evidence.

Should a buyer request a model card?

Often yes, but names vary. The useful document describes intended use, limitations, evaluation, data considerations and current version. It must still be checked against the deployed product and workflow.

What if the vendor cannot disclose sensitive technical details?

Agree on proportionate alternatives: redacted evidence, independent assurance, a controlled review, contract warranties or a higher-level test. Confidentiality does not remove the buyer's need for evidence.

How often should evidence be refreshed?

Set dates by risk and change frequency. Review at renewal and after material changes, incidents, regulatory developments or significant performance shifts.

Can this checklist be used for every AI system?

It is a starting framework. Adapt depth and ownership to the system, sector, role, impact and jurisdiction. Some uses require specialist legal, privacy, cybersecurity, safety or fundamental-rights assessment.

Procurement decision record

Before signing, make sure the file contains:

  • a defined and approved use case;
  • a completed evidence request with gaps;
  • named business, privacy, security and risk owners;
  • product-specific testing;
  • data and subprocessor terms;
  • human-oversight and escalation routes;
  • monitoring and incident obligations;
  • material-change notification;
  • exit, deletion and continuity provisions;
  • a dated renewal review.

Good AI governance does not mean collecting the most documents. It means requesting the evidence that can change a decision, assigning ownership and keeping that evidence current as the system and its context change.

Editorially reviewed for clarity and source currency on .