From Data Cleanup to AI Readiness: Eight Business Operating Practices for Trusted Data
AI initiatives often begin with a request that sounds technical: connect the assistant to the knowledge base, train a forecasting model, automate customer-service triage or give managers a natural-language interface to operational data. The first obstacle is frequently described as “dirty data”. That phrase is convenient, but it hides the real management problem.
An organization can clean thousands of records and still be unready for a defined AI-supported use. The team may not know which source is authoritative, who owns a disputed definition, whether a document is current, which transformations occurred, what uses are permitted, what quality threshold the business actually needs or who can accept a remaining limitation. Conversely, a dataset does not need to be perfect to support every decision. It needs to be sufficiently understood, controlled and fit for the intended purpose, with residual limitations made visible to the people who own the next decision.
That distinction is increasingly visible in current hiring. MTF Institute's accepted public-web corpus contains 105 live, deduplicated data-governance, stewardship, quality, metadata, master-data and AI-readiness vacancies from 94 employers. A reproducible keyword review of each title plus its retained short evidence excerpt produced non-exclusive lower-bound signals for data quality or fitness in 49 records, policy, standards or controls in 44, ownership or stewardship in 33, and metadata, cataloguing, lineage, provenance or discoverability in 28. Thirteen records were already classified in an explicit AI-readiness or data-and-AI role family. Because the excerpts contain only 8 to 20 words, these counts are minimum signals rather than estimates of market prevalence.
This practical guide is grounded in the MTF Institute Research Team's 105-vacancy study and its open Zenodo archive.
The operating lesson is clear: AI readiness is not a one-off cleaning project. It is a business capability that joins purpose, ownership, meaning, provenance, quality, permitted use, issue resolution and accountable handoff.
The following eight practices turn that capability into work managers can perform without purchasing a proprietary framework or pretending that one checklist fits every organization.
Practice 1: define the business use before cleaning the data
Start with a decision, service or workflow, not with an enterprise-wide demand to “fix the data”. A readiness team needs to know what the proposed AI-supported activity will do, who will use it, whose interests may be affected, what data it expects, what output it produces and which human remains accountable for the result.
The UK government's guidelines for making government datasets ready for AI begin from closely related questions: whether the right data exists, whether the capability meets the business need, whether it performs, whether it can be maintained and whether its use is appropriately governed. That document is public-sector guidance, not a universal obligation for private companies, but the sequence is broadly useful.
Imagine a fictional company, Alder & Row Services, that wants an internal assistant to answer account managers' questions about customer contracts, service entitlements and renewal dates. “Clean the customer data” is too broad. The actual readiness question is narrower:
Can the proposed assistant retrieve the currently effective contract terms for an identified customer and service, cite the controlled source, distinguish an executed amendment from a draft, and route uncertainty to the contract owner before an account manager acts?
That question exposes the necessary data: customer identity, contract identity, effective dates, amendment relationships, service definitions, document status, source location and ownership. It also establishes a boundary: the assistant may retrieve and summarize controlled facts, but it does not decide the legal meaning of a clause or authorize a commercial commitment.
A useful first artifact is a one-page Data Readiness Use Brief with these fields:
- business decision or workflow;
- intended users and affected parties;
- expected output and prohibited output;
- required data domains and time horizon;
- minimum response time and freshness;
- accountable business owner;
- downstream AI, product, privacy, security, legal and technical owners;
- evidence cut-off date;
- success measure; and
- conditions that require pause or escalation.
This brief prevents teams from spending months cleaning data that does not materially affect the proposed use.
Practice 2: assign owners, stewards and decision rights by data domain
Data governance becomes operational only when named people can make and challenge decisions. A committee cannot compensate for an undefined owner. A catalogue cannot resolve a semantic dispute if nobody has authority to decide what a field means.
The NIST Data Governance and Management Profile Concept Paper describes data governance as an organizing logic for authority and control over data-management activity. It also links accountability to documented rationale and understood impacts. Importantly, this is a concept paper that supported development discussions. It is not a final NIST profile, a standard or a certification scheme.
For each data domain, separate at least four responsibilities:
- Business owner — accepts the purpose, definition, priority, tolerance and business consequences.
- Data steward — maintains definitions, issue records, quality rules, metadata and coordination.
- Technical custodian — operates storage, pipelines, access mechanisms, monitoring and recovery.
- Independent specialist — reviews questions within privacy, security, legal, risk, compliance or assurance mandates when required.
The same person may hold more than one role in a small organization, but the decisions should remain explicit. Alder & Row's commercial director may own the contract domain, a revenue-operations analyst may steward customer and entitlement definitions, and the platform team may operate the document repository. Legal counsel retains interpretation of disputed contract language. The AI product owner receives the data-readiness evidence but does not silently inherit the commercial director's authority.
Create a Domain Decision Matrix rather than a generic RACI chart. For every recurring decision, record who proposes, who supplies evidence, who decides, who can challenge, who implements and who must be informed. Include decisions such as approving a business definition, changing a critical source, accepting a temporary quality exception, authorizing a new use and retiring a data product.
The completion test is practical: when two systems disagree about a renewal date, can the team identify within minutes who has authority to determine the source of truth and what evidence that person must review?
Practice 3: identify critical data and make business meaning testable
Not every field deserves the same governance effort. A business should identify the data elements whose failure could materially alter the selected decision, service, obligation or customer outcome. These are critical because of consequence and use, not because a platform vendor labels them important.
For the Alder & Row assistant, “contract status”, “effective date”, “customer legal entity”, “service entitlement”, “amendment sequence” and “document version” may be critical. A decorative marketing preference might not be. The list will differ for forecasting, recruitment, procurement, maintenance or customer-service use cases.
For each critical element, create a Critical Data and Definition Register containing:
- business term and plain-language definition;
- business example and counterexample;
- domain owner and steward;
- authoritative source and approved alternatives;
- calculation or derivation rule;
- allowed values and units;
- time basis and effective-date rule;
- population and exclusions;
- sensitivity or classification label supplied by the proper owner;
- downstream uses; and
- unresolved ambiguity.
Definitions need completion tests. “Active customer” is not complete if sales counts signed contracts, finance counts invoiced entities and support counts users who logged in during the past 90 days. The register should state which definition applies to which purpose and where conversions occur.
The NIST Research Data Framework Version 2.0 discusses catalogues, inventories, stewardship, ownership, access and lifecycle curation in a research-data context. Organizations can borrow the operating idea—make data findable and interpretable—without claiming NIST conformity or reproducing the framework's tables.
The readiness test is not “does a glossary exist?” It is “can a user, reviewer and system interpret the critical field consistently enough for the defined purpose, and can they see where interpretation still differs?”
Practice 4: preserve provenance, lineage and transformation evidence
Trusted data needs a traceable story. Where did it originate? Who or what changed it? Which version was used? Which records were excluded? Which transformation joined, aggregated, inferred or corrected it? Could the same result be reconstructed from retained evidence?
The NIST Research Data Framework describes provenance in terms of an attributed history of origin, acquisition, processing and alterations. The NIST DGM concept paper also highlights the difficulties created by uncertain provenance and lineage, especially when data comes from multiple sources or formats.
For business work, use a lightweight Source and Transformation Map. Each row should identify:
- source name and accountable owner;
- method and date of collection;
- original purpose where relevant;
- system and geographic or organizational location;
- upstream supplier or third party;
- version or effective period;
- transformation steps and responsible process;
- validation evidence;
- downstream data product or AI-supported use;
- known loss, inference or ambiguity; and
- retention or reconstruction reference.
In the fictional case, an executed contract may enter the repository through electronic signature, while an amendment is uploaded manually and a service code is enriched from a billing system. If the assistant retrieves a clause without knowing that a later amendment superseded it, document quality is not enough. The relationship among documents is the control.
Do not confuse lineage software with lineage evidence. A visualization can show that table B receives data from table A, yet omit the business rule that excluded dormant accounts or converted gross revenue to net revenue. The useful map connects technical movement to business meaning, decision ownership and change history.
For unstructured documents, add document type, approved status, effective date, supersession link, language, access class, extraction method and citation target. A retrieval system should be able to return the evidence supporting an answer, not merely a confident paragraph.
Practice 5: measure quality against purpose and manage it at the source
“Good quality” is not a measurable requirement. A team needs explicit rules, thresholds and consequences for the intended use.
The NIST DGM concept paper connects fitness for purpose with factors such as accuracy, timeliness, completeness, relevance and consistency, while emphasizing that context matters. The UK data-management guidance links its public-sector quality framework to intended purpose and a structured data-quality action plan. These sources support a general operating principle: quality should be assessed throughout the lifecycle, communicated to users and improved through owned action.
For each critical element, define a Quality Rule Card:
- rule statement;
- reason and affected decision;
- population and exclusions;
- calculation method;
- acceptable threshold and warning threshold;
- measurement frequency;
- data owner and operational resolver;
- source evidence;
- response when the threshold is breached;
- temporary exception authority and expiry; and
- change history.
For Alder & Row, one rule might be: every contract presented as effective must have an executed-status marker, a valid effective date, a customer legal-entity identifier and no later superseding amendment left unlinked. Another rule may require entitlement data to be no more than 24 hours behind the billing source for operational questions. Those are purpose-specific rules, not universal definitions of quality.
Measure defects where they originate, not only after they reach the AI interface. If account managers select the wrong customer entity during contract intake, a downstream matching model may hide the symptom while preserving the control weakness. The remediation plan should distinguish correction of existing records from prevention of future defects.
Report quality with context. A “98% complete” field may be unacceptable if the missing 2% contains the highest-value customers. A score should always state population, period, exclusions, method and decision consequence.
Practice 6: record access, purpose, retention and sharing conditions
Discoverable data is not automatically reusable data. A business needs to know who may access it, for what purpose, under which conditions, for how long and with which downstream restrictions.
Create a Permitted-Use and Sharing Handoff for every material data source. Record the requesting use, data categories, owner, source agreement, access group, retention expectation, onward-sharing condition, security and privacy review status, unresolved questions, decision date and expiry or reassessment trigger.
For personal data in scope of the EU General Data Protection Regulation, principles include purpose limitation, data minimisation, accuracy, storage limitation, integrity and confidentiality, and accountability. Those principles have defined material, territorial and role-specific scope. A course article cannot decide whether they apply to a particular organization or processing activity.
Other EU data laws also have specific boundaries. The European Commission explains that the Data Governance Act addresses particular public-sector reuse, data-intermediation and data-altruism mechanisms. Its Data Act explanation covers defined questions of access, use and interoperability, including connected-product data. Neither is a universal permission to collect, combine or reuse any available data.
The UK's ICO AI and data protection risk toolkit addresses risks to individuals from organizations' own AI systems. As of 20 August 2026, the ICO page states that the guidance is under review following legislative change. That status should travel with any reference.
In practice, the data steward documents the question and evidence. Qualified privacy, security, legal, procurement or regulatory owners make the decisions within their mandates. “AI ready” must never be used to bypass that handoff.
Practice 7: govern structured records and unstructured knowledge as products
AI-supported work increasingly combines structured tables with policies, contracts, tickets, emails, manuals and other documents. The accepted vacancy corpus includes roles that explicitly connect AI readiness with both structured and unstructured data. That does not mean every document should be indexed. It means readiness must cover the information forms the intended use actually depends on.
Treat each governed dataset or document collection as a Business Data Product Card. It should state:
- purpose and approved consumers;
- accountable owner and steward;
- content and exclusions;
- source and update process;
- master or reference identifiers;
- schema or document metadata;
- quality and freshness commitments;
- permitted-use and access conditions;
- known limitations;
- version and change notice;
- support and issue route; and
- retirement condition.
Master and reference data deserve special attention because they connect systems. Customer, product, supplier, legal-entity, location, currency and organizational-unit identifiers can determine whether records are joined correctly. A duplicate or ambiguous master record may cause an AI-supported process to retrieve the wrong evidence even when each source is locally accurate.
Unstructured collections need equivalent control. For Alder & Row, the contract knowledge product should exclude drafts from ordinary retrieval, link amendments to base agreements, expose document status and effective date, retain citations, and identify the owner who can resolve conflicting text. The assistant's retrieval index is a derived product with its own version and refresh evidence; it is not the authoritative archive.
This practice avoids two common failures: treating data as a one-time project output, and treating an AI index as if it were the source of truth.
Practice 8: close readiness with remediation, evidence and an accountable handoff
A readiness assessment should not end with a red, amber or green badge. It should produce decisions and work.
Build a Data Readiness Decision Pack containing:
- the approved use brief;
- domain decision matrix;
- critical-data and definition register;
- source and transformation map;
- quality rule cards and current results;
- permitted-use and sharing handoffs;
- business data product cards;
- open issues, exceptions and assumptions;
- prioritized remediation backlog;
- residual limitations;
- readiness decision for the data layer; and
- named downstream owners and next review date.
Prioritize remediation by decision consequence, affected population, likelihood, detectability, effort, dependency and reversibility. Do not prioritize only by the number of defective rows. A small ambiguity affecting contract authority may matter more than thousands of missing optional marketing fields.
Each backlog item should have an owner, due date, evidence of completion, retest method and closure authority. Temporary acceptance should state the compensating action, monitoring, expiry and escalation trigger. Repeated exceptions should lead to source-process improvement rather than permanent tolerance.
The final statement must be precise. For example:
The contract and entitlement data layer is sufficiently documented for a controlled pilot limited to internal retrieval of current standard-service terms. The assistant must cite the approved document, must not interpret disputed clauses, and must route missing amendment links to the commercial owner. Three remediation items remain open and are due before expansion to non-standard agreements.
That statement is more useful than “data is AI ready”. It defines purpose, scope, conditions, unresolved work and decision ownership. It also preserves the boundary to the next teams. Data governance prepares evidence; AI governance, product, engineering, privacy, security, legal and other accountable owners perform their own reviews.
A practical 30-day operating rhythm
Organizations do not need to complete enterprise-wide transformation before they can improve one use case. A bounded 30-day cycle can create durable evidence:
Week 1: purpose and ownership
- select one decision or workflow;
- approve the Data Readiness Use Brief;
- identify data domains and critical elements;
- assign owners, stewards and specialist handoffs.
Week 2: meaning and traceability
- complete business definitions;
- identify authoritative sources and alternatives;
- map provenance, transformations and document supersession;
- record master and reference identifiers.
Week 3: quality and permitted use
- define purpose-specific quality rules and freshness thresholds;
- measure the current population with stated exclusions;
- document access, purpose, retention and sharing questions;
- route specialist decisions without inventing answers.
Week 4: remediation and handoff
- prioritize defects and source-process changes;
- record accepted limitations and expiring exceptions;
- assemble the evidence pack;
- issue a bounded data-layer readiness decision and handoff.
The cycle can then repeat for another domain or use. Reusable definitions, ownership records and source maps reduce later effort, while each new use still receives a purpose-specific assessment.
What these practices do not prove
These eight practices do not certify data, an organization or an AI system. They do not establish legal compliance, model validity, fairness, safety, security or production approval. They do not replace testing, privacy assessment, security review, legal advice, model evaluation, product governance or independent assurance.
Article 10 of the EU AI Act contains data-governance and data-quality requirements in its defined context for providers of covered high-risk AI systems using model-training techniques and relevant datasets. It should not be generalized to every dataset, AI use or organization. Applicability, actor role and timeline require qualified review.
The NIST AI Risk Management Framework is voluntary, and NIST currently states that version 1.0 is being revised. The NIST DGM source used here is a concept paper. UK sources used here are public-sector or UK regulatory guidance. These materials provide evidence and operating ideas; they do not create a single global compliance checklist.
Conclusion
The move from cleanup to readiness is a move from records to decisions. A trusted data layer has a defined purpose, named owners, testable meaning, traceable sources, measurable quality, documented use conditions, maintained data products and an owned remediation process.
Current vacancies show that employers are assembling these responsibilities across governance, stewardship, quality, metadata, master-data and AI-readiness roles. The official sources reviewed here reinforce the same practical direction while preserving important boundaries: context matters, accountability must be explicit, and readiness evidence is not the same as approval.
For business professionals, the most valuable starting point is therefore not a platform purchase or an enterprise-wide cleaning campaign. It is one well-defined use, one accountable data domain and one evidence pack that lets the next decision-maker see what is known, what is controlled and what still requires work.
Sources and further reading
- NIST Joint Frameworks Data Governance and Management Profile Concept Paper — concept paper, not a final profile or standard
- NIST AI Resource Center and AI Risk Management Framework resources
- NIST Research Data Framework Version 2.0
- UK guidelines and best practices for making government datasets ready for AI
- UK National Data Library data-management guidance
- EU Artificial Intelligence Act
- EU General Data Protection Regulation
- European Commission explanation of the Data Governance Act
- European Commission explanation of the Data Act
- UK ICO AI and data protection risk toolkit — under review as stated by the ICO on the retrieval date
Evidence limitation
The vacancy corpus is a point-in-time public-web sample retrieved on 20 August 2026. Job pages may later change or close. It covers multiple regions but is not proportional to labour-market size. The keyword counts use only titles and short retained excerpts, so they are conservative topic signals rather than estimates of employer prevalence, hiring volume, candidate demand or course sales.