Research Report MTF-RR-2026-09-21-02

Executive summary

Workplace AI is moving beyond single-turn question answering. Across employer implementations, current U.S. vacancy signals, occupational tasks, empirical studies, and platform or public-interest guidance, the recurring pattern is not “full autonomy.” It is bounded delegation: a worker defines an outcome, supplies approved context, selects tools and permissions, supervises progress, checks the result, and remains responsible for handoff or escalation.

This report analyzes 100 distinct evidence records drawn from 78 public source URLs. The records are distributed across five predeclared strata: 25 employer or workplace implementations, 20 U.S. labour-demand signals, 20 occupational task records, 20 empirical observations from five method-disclosed studies, and 15 platform, technical, safety, or public-interest records. The design is purposive, not representative. It identifies transferable skills and claim boundaries; it does not estimate national prevalence or prove productivity, income, or employment effects.

The central conclusion is practical: the durable capability is not familiarity with one chatbot or model name. It is the ability to turn a work objective into a controlled agent workflow and to know when the system should stop, ask, escalate, or hand work back to a person.

Research question and method

The study asks a deliberately narrow question: which practical capabilities recur when U.S. workers and organizations move from conversational AI use toward multi-step, tool-using work? It does not ask how many jobs will be automated, which vendor will lead the market, or what return an organization should expect from a particular system. Those questions need different sampling designs and, in many cases, longitudinal or causal evidence.

Before collection, the research geography was fixed as the United States and the five evidence strata and minimum counts were declared. A record was accepted only when it had a public URL, identifiable publisher, U.S. jurisdiction or a clearly bounded U.S. application, a distinct evidence proposition, a rights note, and a claim-use boundary. Global documentation from U.S. platform providers was allowed only to support dated capability or safety propositions. It was not allowed to establish U.S. adoption prevalence. Search snippets could help discover a source, but the record had to resolve to the underlying page or document.

The unit of analysis is one independently citable task, implementation, demand signal, empirical observation, capability, or safety proposition. Records were deduplicated using source identity, organization, title, evidence meaning, and the task/output pair. A single report could contribute more than one record only when the observations were analytically different. This is why the final corpus is described as 100 evidence records from 78 unique source URLs. It is not a set of 100 independent publications.

Each record was coded for work context, task, output, interaction mode, autonomy, tools or data, human checkpoint, verification method, risk, failure mode, prerequisite skill, evidence strength, and contrary evidence. The synthesis then mapped records to eight non-exclusive skill domains. Because a real workflow can require framing, permissions, supervision, and verification at the same time, domain totals are not intended to sum to 100.

Several balance controls were applied. No organization supplies more than two employer-implementation records. No empirical source group supplies more than a quarter of that stratum. No publisher supplies more than 15 percent of the full matrix, and at least 30 percent of records had to be non-vendor. The completed matrix exceeds those safeguards: its largest publisher contributes 11 percent, and 79 percent of records are classified as non-vendor under the conservative publisher rule.

What distinguishes agentic work from chatbot use

A chatbot interaction usually produces a response to a prompt. An agentic workflow may instead plan or decompose a task, call tools, retrieve or update records, act across several steps, retain state, and return a result that needs an acceptance decision. The distinction is functional, not promotional: a product label alone does not make a workflow agentic.

Employer evidence shows this progression in several contexts. AT&T describes agents that can update customer records or propose network fixes under oversight, while its earlier Ask AT&T work emphasized document synthesis, code drafting, and database queries. Walmart describes agent use across customer, associate, supply chain, and software workflows. Microsoft reports support workflows that connect agents to trusted sources and APIs. These examples establish that bounded action exists in practice; because they are first-party reports, they do not independently prove return on investment or organization-wide impact.

Sources: AT&T agentic AI, Walmart's agentic strategy, Microsoft modern support.

Results across the five evidence strata

The employer stratum contains 25 observations covering 19 organizations, 10 sectors, and 24 coded functions. Twenty-two records come from first-party employer sources. The examples include customer and employee support, research, software, network operations, retail, supply chain, human resources, revenue-cycle work, clinical decision support, marketing, and internal governance. The spread matters because it shows that agentic practice is not confined to software development. The records also show substantial variation in autonomy: some systems draft or recommend, while others take bounded action in connected systems. Across both forms, a person or accountable service owner remains part of the control design.

The labour-demand stratum contains 20 distinct U.S. or U.S.-eligible vacancies. Its strongest repeated signals are workflow design, tool integration, retrieval, evaluation, observability, guardrails, human approval, traceability, recovery, and safe deployment. These records indicate what sampled employers were seeking on the retrieval date. They do not reveal how frequently those skills appear across the whole labour market, and their technical-role concentration means they should not be projected onto every knowledge-work occupation.

The occupational stratum contains 20 records drawn from U.S. Bureau of Labor Statistics task descriptions and U.S. Department of Labor AI-literacy guidance. It anchors the analysis in work people already perform: analyzing information, coordinating projects, communicating findings, maintaining records, applying professional judgment, checking outputs, protecting data, and knowing when to escalate. These records make no claim that an agent can safely take over an occupation. Their role is to show where agentic tools meet existing human tasks and responsibilities.

The empirical stratum contributes 20 observations from five source groups, with no group contributing more than four. It captures adoption, reported use, organizational readiness, barriers, and changing patterns of cognitive work. The platform and safety stratum adds 15 dated records spanning OpenAI, Google, Anthropic, NIST, the U.S. Department of Labor, and the Federal Trade Commission. Taken together, these two strata add an essential corrective: capability is moving quickly, but readiness, trust, control, and actual use remain uneven.

Eight transferable skill domains

1. Select the right task and define the outcome

The first decision is whether a task should be delegated at all. A responsible workflow starts with a concrete work product, boundaries, acceptable sources, constraints, and a quality test. Current vacancy signals include workflow discovery, reference designs, acceptance criteria, and customer or business translation—not only model configuration. U.S. occupational sources similarly show that requirements gathering, decision support, communication, and quality control are core work activities.

2. Prepare context, sources, and data

Agents cannot infer which internal source is authoritative or which data may be used safely. Evidence across research synthesis, sales, support, and enterprise knowledge work repeatedly points to grounding, source selection, structured inputs, and controlled access. More data is not automatically better; irrelevant, stale, confidential, or unlicensed material can make the result worse or create a new risk.

3. Choose tools, environments, and permissions

The move from text generation to action creates an authority problem. A useful agent may need a browser, code environment, knowledge base, API, or enterprise system. Each connection expands what the system can see or change. The worker therefore needs to choose the minimum tool and permission set appropriate to the task, distinguish read from write access, and keep irreversible or sensitive actions behind approval.

4. Decompose and delegate

Agentic work often requires breaking a broad objective into steps, assigning tools or specialized agents, and preserving dependencies. The evidence does not support treating orchestration as automatic correctness. Good decomposition makes intermediate outputs inspectable and creates places where a person can intervene.

5. Supervise progress and handle exceptions

Human oversight is most useful when it is designed into the workflow. The evidence includes approval gates, escalation to specialists, review of clinical or financial outputs, and recovery paths for unexpected states. A human “in the loop” is not a safety guarantee unless that person has sufficient context, authority, time, and a clear decision rule.

6. Verify outputs and preserve traceability

Current U.S. vacancy signals repeatedly ask for evaluations, regression tests, observability, tracing, monitoring, benchmarks, and incident handling. These are not exclusively engineering ideas. A knowledge worker also needs to verify facts, sources, calculations, policy alignment, completeness, and whether the result meets the original acceptance test. The verification method should match the failure cost.

7. Integrate privacy, security, rights, and policy

Risk control belongs inside the workflow. NIST's AI Risk Management Framework profile, U.S. Department of Labor AI literacy guidance, Federal Trade Commission material, and vendor safety documentation all reinforce the need to consider data sensitivity, authority, provenance, security, and foreseeable harm before an agent acts. These sources inform operating practice; they are not legal advice and do not establish compliance in every jurisdiction.

Sources: NIST Generative AI Profile, U.S. Department of Labor AI Literacy Framework, Anthropic trustworthy agents.

8. Hand off, document, and operationalize

A useful workflow ends with an accountable handoff: what the system did, what sources and tools it used, what remains uncertain, who approved the result, and what should happen next. Documentation also makes a workflow teachable, repeatable, auditable, and easier to improve.

What the labour-demand evidence adds

The 20 U.S. or U.S.-eligible vacancy records are current demand signals rather than a labour-market census. Together they emphasize tool calling, retrieval, multi-agent design, guardrails, human review, evaluation, observability, traceability, rollback, policy controls, and safe deployment. The sample is deliberately broad across employers but is weighted toward technical, product, and AI-enabled roles. It therefore supports a conclusion about the kinds of capabilities employers are asking for, not how common those requirements are across all U.S. jobs.

Examples include Salesforce's Agentforce architecture role, DoorDash's agent evaluation role, and Docker's long-running agent workflow role. Vacancy pages may change or disappear; their status should be rechecked before external publication.

What empirical studies add

Five method-disclosed source groups provide 20 sample-bounded observations: Gallup, Epoch AI with Ipsos KnowledgePanel, Deloitte, Pew Research Center, and OpenAI Economic Research. Together they show that workplace AI adoption and use are uneven, that workers and organizations differ in readiness, and that observed use spans several kinds of cognitive work. The studies use different populations and methods—worker self-report, leader surveys, interviews, and product-use analysis—so their figures should not be pooled or treated as interchangeable.

Sources: Gallup workforce adoption study, Epoch AI and Ipsos polling, Deloitte agentic transformation study, Pew Research Center worker study, OpenAI work-use research.

Implications for business and work

The evidence supports four restrained conclusions.

First, agentic work is cross-functional. The employer stratum covers 19 organizations, 10 sectors, and 24 coded functions. Second, reliable use depends on workflow and control skills as much as on model interaction. Third, increasing autonomy raises the importance of permissions, monitoring, verification, and recovery. Fourth, the most transferable learning objective is accountable delegation, not dependence on a particular product interface.

The evidence does not show that an agent should replace an accountable worker, that one platform is universally superior, or that brief training guarantees productivity, career advancement, or income. High-stakes examples in health, finance, security, employment, or public services should be used to teach risk boundaries, not to encourage learners to experiment with real sensitive data.

How the skill stack works as one operating cycle

The eight domains are best understood as a connected cycle rather than a menu of independent techniques. A worker first decides whether the task is suitable and defines the intended output. The worker then assembles authorized context, selects the minimum necessary tools and permissions, and decomposes the work into inspectable steps. During execution, the worker watches for exceptions, pauses at approval points, and prevents the system from extending its authority merely because it can continue. The output is then checked against sources, calculations, policy, tests, or other acceptance criteria. Finally, the worker documents what happened and hands the result to the person or process that owns the next decision.

This cycle also explains why “human in the loop” is too vague to be a complete control. Oversight can occur before action, during execution, after output, or only when an exception is triggered. The right checkpoint depends on consequence and reversibility. A draft internal summary can usually tolerate a different review design from a customer-record update, financial decision, clinical alert, or security change. The corpus supports risk-proportionate controls, not one universal approval pattern.

Verification likewise changes with the task. Factual synthesis needs source and provenance checks. Code or configured workflows need tests and regression checks. Policy-sensitive work needs rule and authority review. Long-running or connected agents need monitoring, traceability, incident handling, and recovery. A fluent answer is therefore only the beginning of evidence, not proof that the work is correct or complete.

For organizations, this points to a practical division of responsibility. System owners define permitted tools, data boundaries, logging, and escalation paths. Managers define outcomes, tolerances, and accountability. Workers frame tasks, supervise execution, verify outputs, and make or escalate decisions. Risk, security, legal, and subject-matter specialists provide controls where the consequences require them. The evidence does not imply that every workflow needs all roles; it supports making ownership explicit before autonomy increases.

Method and evidence boundaries

The corpus contains exactly 100 non-duplicative analytical records across five predeclared strata. It resolves to 78 unique URLs because a single method-disclosed study, public framework, or technical document may support more than one distinct observation. Those repeated URLs are legitimate multi-observation sources; the corpus must not be described as 100 independent publications.

No publisher supplies more than 11 percent of the matrix. Twenty-two of 25 employer-implementation records are first-party employer sources. The empirical stratum contains five source groups, with no group contributing more than four of 20 observations. Under a conservative publisher classification, 79 percent of records are non-vendor evidence. These controls reduce, but do not eliminate, selection and self-report bias.

The full methodology, record-level claim boundaries, rights status, source ledger, geography-coherence matrix, claim-evidence matrix, and limitations are preserved in the frozen research package.

Trademark, product, and jurisdiction note

ChatGPT and Codex are OpenAI product names; Gemini is a Google product name; Claude is an Anthropic product name. Their use in comparative educational or research text should remain nominative and accurately attributed. This research does not imply endorsement, sponsorship, certification, or partnership by those companies. Do not use vendor logos or trade dress without separate permission. Model and feature names are volatile and require verification immediately before public release.

The evidence population is U.S.-focused. It does not establish compliance with Portuguese or EU law. Any Lisbon delivery needs a separate review of privacy, consumer information, accessibility, recording consent, certificate wording, and applicable AI-governance obligations.

Related professional learning

Professionals who want to turn this evidence into supervised practice can join MTF Institute's Professional Certificate in Agentic Systems and AI in Work and Business. The live two-hour workshop is delivered in Lisbon or online, uses each participant's own computer, and applies the report's operating cycle through instructor-guided work. It is professional education, not an academic degree; the cohort date, time and Lisbon venue are announced during recruitment.

Archive and citation

The complete public research package, including the searchable PDF, frozen evidence matrix, source ledger, claim-evidence matrix, geography-coherence matrix, methodology and limitations, is preserved in Zenodo record 10.5281/zenodo.22884074. The public PDF is the archival reading edition.

Author: MTF Institute Editorial Team.
Publication date: 21 September 2026.
Institution: MTF Institute.
Report number: MTF-RR-2026-09-21-02.

Readers who want a dated view of recent product and operating-model changes can also read From Answers to Actions: What Changed in Agentic AI for Work in the Last 90 Days.