The Operating Shape of Data Governance and AI Readiness: Evidence from 105 Current Vacancies in 2026
MTF Institute Research Team
Research cut-off: 20 August 2026
Status: original website research draft; not published
Abstract
Data readiness is often discussed as if an organization could reach it by buying a platform or cleaning a dataset once. A bounded review of 105 current public vacancies shows a different operating picture. Employers describe work that connects data quality, ownership, stewardship, metadata, catalogues, lineage, rules, standards, measurement and remediation. Only 13 records in the sample use explicit AI- or data-readiness language in the title or preserved evidence excerpt. Much larger groups refer to the foundations on which readiness depends: 57 mention quality or fitness terminology, 41 mention rules, standards or policies, 33 mention ownership or stewardship, and 31 mention metadata, catalogues, glossaries or dictionaries.
This report analyses the operating shape visible in those vacancies. It does not estimate market size, hiring growth, salary or course demand. Its central finding is that AI readiness is better treated as a bounded business-data operating capability than as a label attached to every governance task. The practical work is to make data discoverable, owned, defined, traceable, usable for a stated purpose and accompanied by documented limitations and remediation decisions before accountable downstream teams proceed.
The complete research archive — searchable PDF, evidence workbook and accepted-vacancy dataset — is openly available at Zenodo DOI 10.5281/zenodo.22029878.
Key findings
- The corpus contains 105 unique current public vacancies from 94 employer labels, retrieved on 20 August 2026.
- 64 vacancies (61.0%) came from a primary employer or applicant-tracking page; 41 (39.0%) came from a public job board.
- The sample spans nine preserved regional labels, grouped into Europe, North America, Latin America, Asia-Pacific, Africa and the Middle East, plus one cross-region record for compact analysis.
- The largest normalized role families are leadership and enterprise governance (24 vacancies; 22.9%) and governance operations and policy (21; 20.0%).
- Hands-on families are also substantial: stewardship and ownership (16; 15.2%), metadata/catalogue/lineage (13; 12.4%) and data quality (12; 11.4%).
- Explicit AI- or data-readiness terminology appears in 13 records (12.4%), while quality or fitness terminology appears in 57 (54.3%).
- The sample includes 40 individual-contributor labels, 30 manager or lead labels, and 14 executive, director or head labels. The work therefore cannot be reduced either to executive policy or to technical administration.
Why examine vacancies?
Vacancies are imperfect but useful operating evidence. A course description can state aspirations; a vacancy normally has to describe work that an employer expects someone to perform. Across many employers, those descriptions reveal the verbs, objects and interfaces around a role: define, own, curate, monitor, document, trace, resolve, coordinate and improve.
Vacancies do not reveal everything. They may compress a complex role into a few paragraphs, mix strategic and operational responsibilities, or use organization- specific titles. Public pages may disappear after a position closes. The number of visible vacancies is not the number of people hired, and a purposive public-web sample is not a labour-force survey. These constraints make vacancies unsuitable for a growth forecast. They remain useful for a narrower question: what kinds of work are employers currently describing around data governance and readiness?
That question matters because “AI readiness” can otherwise become a vague slogan. If readiness is a professional capability, a learner should be able to produce evidence: which data a proposed business use needs, where that data comes from, who owns it, what it means, how its fitness will be assessed, which limitations remain and who must decide the next step.
Method
The MTF Institute Research Team assembled a dated corpus of 105 individually addressable public vacancies. Every accepted record exposed a current application surface on 20 August 2026 and had duties materially related to data governance, stewardship, data quality, metadata, master data or data readiness. The corpus excludes closed and expired pages, inaccessible or redirect-only records, search-results lists, duplicate or syndicated copies where a primary page was available, and roles too remote from the research question.
Each record contains a source identifier, employer label, title, location, regional label, source family, role-family code, seniority label, retrieval date, status, public URL and one short evidence excerpt. The final corpus contains 105 unique source identifiers, 105 unique URLs and no duplicate normalized employer- title-location key.
For analysis, mixed role labels were mapped one-to-one into eight normalized families. Regional labels were retained, then grouped only for compact crosstabs. Seniority labels were grouped with a fixed keyword precedence. Evidence-signal counts searched each title plus its short excerpt for disclosed English, Portuguese and Spanish terms. A vacancy could match several signals, so those counts are deliberately non-additive.
The complete coding rules and crosstabs are recorded in the accompanying vacancy analysis. This report uses aggregate counts and original synthesis. It does not reproduce individual job descriptions or infer facts that were not preserved in the corpus.
1. Readiness rests on quality and fitness for a stated purpose
Quality and fitness terminology is the most visible signal in the corpus. It appears in 57 of the 105 titles or short evidence excerpts. The matched language includes quality, integrity, accuracy, timeliness, reliability, trustworthiness and fitness for purpose.
This variety matters. “Good data” is not an actionable requirement. A commercial team deciding which customers should receive a service intervention may need complete contact preferences, an agreed customer identifier and sufficiently fresh activity data. A supply team planning replenishment may care more about product-master consistency, location codes and the time boundary of stock data. The same dataset can be fit for one decision and unsuitable for another.
The vacancy evidence therefore supports a purpose-first working question:
What decision or process will use this data, and what observable properties must the data have for that purpose?
Answering it requires more than a generic score. A usable fitness statement names the data element, population, period, rule, threshold, owner, evidence source and response when the rule is not met. For example, a business team might define the required population as active customer accounts, the relevant element as the account-country code, the test as conformance to an approved reference list, and the decision as whether regional service rules can be applied consistently. The example describes an operating method, not a universal threshold.
The important shift is from “clean the data” to “state the use, test the relevant properties and record the remaining limitation.” That shift makes quality work reviewable and connects it to a real business decision.
2. Rules and standards become useful through operating routines
Rules, standards or policy terminology appears in 41 records. This is the second-most frequent signal, but a written rule alone does not create reliable data. The operating question is how a rule enters daily work.
A workable routine needs at least five connections:
- a defined object, such as a critical field, dataset, report input or shared reference value;
- an accountable owner who can clarify meaning and priority;
- a control point where the rule can be observed;
- an exception path when the rule is not met; and
- a review cadence that decides whether the rule still serves its purpose.
This helps explain why the sample includes both leadership roles and operational roles. Leaders may define direction, scope, decision rights and priorities. Data owners, stewards, analysts and specialists translate those decisions into definitions, catalogues, checks, issue queues and working agreements. Neither layer replaces the other.
For professional learning, the implication is practical. A learner should not finish with a list of principles only. They should be able to convert one principle into a working record: the governed object, the rule, the accountable role, the evidence, the exception route and the next review date.
3. Ownership and stewardship are active work, not labels on an organization chart
Ownership or stewardship terminology appears in 33 records. The normalized stewardship-and-ownership family contains 16 vacancies, and similar responsibilities also appear in leadership, operations, quality and metadata roles.
An ownership label is weak if it does not connect to a decision. “Marketing owns customer data” leaves important questions unresolved. Who approves a business definition? Who decides whether a quality exception is tolerable for a particular use? Who coordinates a fix across a source system and a downstream report? Who records a limitation when remediation cannot be completed before a deadline?
The operating evidence suggests a clearer division of work:
- a business owner accepts priorities and business consequences;
- a steward maintains definitions, coordinates evidence and follows issues;
- a source or platform role explains technical origin and transformation;
- a data user states the decision context and fitness need; and
- an accountable downstream specialist makes the next decision that belongs to product, engineering, privacy, security, legal or AI governance.
These labels are an original teaching synthesis, not a universal organizational model. Their value is to expose handoffs. When a data issue remains unresolved, the team needs to know who can fix it, who can accept a bounded limitation, who must be informed and who does not have authority to proceed.
The corresponding learner artifact is not a large responsibility chart. It is a narrow ownership-and-decision map for a defined data domain, with concrete decisions, evidence responsibilities and escalation points.
4. Metadata and lineage turn discovery into evidence
Metadata, catalogue, glossary or dictionary terminology appears in 31 records. Lineage, traceability or provenance terminology appears in 17. These signals are related but not interchangeable.
A catalogue can help someone find a data asset. A glossary can help people agree what a business term means. Technical metadata can describe format, location or refresh. Lineage can show movement or transformation. Provenance can record where an input originated and under which conditions it was obtained. Readiness work needs a purposeful combination of these forms of evidence.
Consider a fictional organization planning an AI-supported service-summary tool. The team may locate customer cases, knowledge articles and interaction notes. That discovery is only the beginning. It still needs to distinguish current from retired articles, identify the source of each interaction record, understand whether important fields have a shared meaning, and record which transformations occur before the tool receives the data.
The practical chain is:
business term → data element → source → transformation → destination → use → known limitation
This chain is intentionally simple. It gives business professionals a way to ask better questions without pretending to replace technical lineage engineering. A learner can document the evidence available, mark unknown links and assign an owner to close each material gap.
The vacancy corpus supports this evidence-oriented approach. It does not establish that every employer uses the same catalogue, lineage platform or metadata model, and it does not support a vendor recommendation.
5. Explicit AI-readiness work is visible, but it is not the whole operating model
Thirteen records contain explicit AI- or data-readiness terminology in their title or preserved evidence excerpt. Twelve vacancies were also coded into the normalized AI-readiness and data/AI-interface role family. Seven of those twelve are in the North American portion of this purposive sample.
The concentration should be interpreted cautiously. It may reflect employer language, source access and sampling, not a regional difference in underlying practice. More importantly, the 13 explicit signal matches are far fewer than the 57 quality matches, 33 ownership/stewardship matches or 31 metadata/catalogue matches.
That pattern supports a bounded conclusion: AI readiness is an interface built on data operating work. It is not evidence that every data-governance role should own an AI system's approval, risk classification, technical validation or monitoring.
A business data-readiness assessment can answer questions such as:
- Is the required data discoverable and linked to a defined business purpose?
- Is there an owner for meaning, priority and remediation decisions?
- Are critical elements defined consistently enough for the intended use?
- Is the source and transformation path sufficiently documented?
- Are relevant fitness rules stated, tested and time-bounded?
- Are access, purpose and other use constraints recorded for handoff?
- Are known gaps visible, prioritized and assigned?
The assessment should then stop at an explicit handoff. Other accountable teams must decide questions outside the data layer. This boundary protects both the learner and the organization from false authority.
6. The sample shows an operating system distributed across levels
The seniority mix reinforces the need for an end-to-end operating view. Forty records carry individual-contributor labels, 30 carry manager or lead labels, and 14 carry executive, director or head labels. Eight are consultant labels, five are early-career labels, and eight use mixed or experience-only descriptions.
The role-by-seniority crosstab makes the distribution more concrete:
- the leadership and enterprise-governance family contains seven executive and 15 manager/lead records;
- governance operations and policy contains 12 individual contributors, three consultants, two executive records, two manager/lead records and two early- career records;
- metadata/catalogue/lineage contains nine individual contributors, two consultants, one early-career record and one mixed label; and
- data quality contains seven individual contributors, two managers/leads, two consultants and one early-career record.
This is not proof of a standard career ladder. It does show that the work crosses organizational levels. A director may establish priorities and decision rights; a manager may operate the governance cadence; an analyst or steward may maintain definitions and issue evidence; and a specialist team may implement technical controls. Effective learning should make these interfaces visible.
7. An original business data-readiness loop
The vacancy evidence can be synthesized into an original MTF learning loop. The loop is not an employer standard, a certification model or a claim that every organization should use the same sequence. It is a practical way to organize the work visible across the sample.
Step 1 — Define the business use
State the decision, process or service that needs data. Record the population, time horizon, intended user and consequence of a poor result. A readiness review without a defined use becomes an abstract data-cleanup exercise.
Step 2 — Identify the necessary data
List the domains and critical elements required for the use. Separate necessary inputs from desirable additions. This keeps the review bounded and prevents a team from attempting to fix the whole enterprise before delivering value.
Step 3 — Assign decisions and stewardship work
Name who decides meaning, priority and acceptable limitations; who maintains evidence; who can change the source; and who owns the downstream decision. Record gaps instead of inventing accountability.
Step 4 — Define meaning and locate sources
Connect business terms to data elements and locate the authoritative or currently used sources. Record competing definitions and unresolved source questions.
Step 5 — Map provenance and transformation
Describe where data originated, how it moved and which material transformations occurred. Use the level of detail necessary for the decision, and hand technical implementation questions to the appropriate specialists.
Step 6 — Set fitness rules
Translate “good enough” into observable rules for the stated purpose. Record the population, threshold, evidence date, exclusions and rule owner. Distinguish a measured fact from an estimate or assumption.
Step 7 — Record use constraints and interfaces
Document known access, purpose, contractual, retention or organizational constraints as facts to be reviewed by the accountable teams. Do not convert them into a legal or compliance conclusion.
Step 8 — Prioritize remediation
Create a backlog with the gap, business consequence, owner, dependency, evidence needed, next action and decision date. Prioritize by the ability of a gap to change the intended use, not by how easy it is to count.
Step 9 — Handoff with evidence
Package the definitions, ownership, provenance, fitness results, limitations and open actions for the next accountable team. The handoff says what is known, what is not known and which decision remains outside the data-governance team's authority.
The strength of the loop is traceability. A learner can move from a business need to a specific data element, from that element to an owner and source, from the source to a fitness test, and from a failed test to a remediation decision.
Implications for professional learning
The corpus supports a course built around reusable professional artifacts rather than knowledge checks alone. A coherent learning journey could ask a learner to produce:
- a business-use and data-dependency brief;
- a data-domain and ownership map;
- a critical-data register;
- a business glossary and definition-decision log;
- a source and provenance register;
- a business-level lineage map;
- a fitness-for-purpose rule set;
- a quality evidence scorecard;
- an issue and exception workflow;
- a remediation backlog;
- a master and reference-data accountability record;
- a third-party data evidence record;
- a readiness assessment for one defined use;
- a downstream handoff pack; and
- an operating cadence and 90-day improvement roadmap.
These are proposed learning tools, not artifacts observed as a single package in the vacancies. Their design should remain original, vendor-neutral and adaptable. Each should include a filled fictional case example, a blank learner template, context-rich AI prompts and a self-assessment rubric. The prompts should help a learner organize sanitized evidence, identify gaps and test internal consistency; they should not fabricate organizational facts or replace accountable review.
The sample also cautions against turning the course into an analytics or technical implementation program. The professional outcome is an evidence-controlled data readiness operating pack, not a dashboard, model, application, legal opinion or AI-system approval.
What this research does not show
This report does not show that vacancies grew in 2026, that one role title is more valuable than another, or that a particular qualification changes salary. It does not estimate the total number of jobs, the probability of employment or the commercial performance of a course.
It does not determine whether any employer satisfies a law, policy, security requirement or professional standard. Terms of that kind may appear in vacancy wording, but their applicability and interpretation are outside this analysis.
It also does not show that the ten evidence signals are complete job descriptions. Each record preserves only a short excerpt. A missing term means the term was not visible in the retained title and excerpt, not that the underlying role lacks the responsibility.
Finally, the report does not authorize copying an employer's job description or any proprietary framework. Counts, mappings, explanations and the readiness loop are original analytical expression by the MTF Institute Research Team.
Conclusion
Across 105 current vacancies, the operating shape of data governance is not one job title or one framework. It is a connected system of leadership, operating rules, ownership, stewardship, definitions, metadata, lineage, quality evidence, measurement and remediation. Explicit AI-readiness language is visible, but the larger signal is the foundational work that makes data understandable and usable for a defined purpose.
For business professionals, the practical opportunity is to learn how to assemble that evidence and manage the handoffs. A credible readiness assessment should say which data is needed, who owns it, what it means, where it comes from, how its fitness was evaluated, which limitations remain and who must decide what happens next. That is a bounded, observable professional outcome—and a stronger basis for AI-supported work than an unsupported claim that the data is simply “ready.”
Research basis and limitations
Primary research asset: MTF Institute Research Team, Data Governance and AI Readiness accepted vacancy corpus, 105 public vacancies retrieved 20 August 2026. The corpus records source URLs, provenance, retrieval status, short evidence anchors and deterministic deduplication fields.
Source composition: 64 primary employer or ATS pages and 41 public job-board
pages.
Geographic composition: 39 Europe, 20 Asia-Pacific, 20 Latin America, 18
North America, seven Africa and Middle East, and one cross-region record under the
report's grouping rule.
Core limitations: point-in-time status; purposive and access-influenced sample;
not weighted to labour-market size; short evidence excerpts; analytical role and
seniority normalization; no salary, growth, hiring-volume or sales inference.