# The Agentic Work Skill Stack: Evidence for U.S. Knowledge Work and Business Practice

> A 100-record U.S. evidence synthesis identifies eight transferable capabilities for framing, supervising, verifying and governing agentic work.

- Canonical page: https://mtfinstitute.com/insights/agentic-work-skill-stack-us-knowledge-work-business-practice/
- Content type: Article
- Editorial category: Research &amp; Reports
- Publisher: MTF Institute of Management, Technology and Finance
- Author: MTF Institute Editorial Team- Published: 2026-09-21
- Updated: 2026-09-22
- Language: English
- Topics: AI governance, Human Oversight, Agentic AI, Knowledge Work, Business Workflows

**Research Report MTF-RR-2026-09-21-02**

## Executive summary

Workplace AI is moving beyond single-turn question answering. Across employer
implementations, current U.S. vacancy signals, occupational tasks, empirical
studies, and platform or public-interest guidance, the recurring pattern is not
“full autonomy.” It is bounded delegation: a worker defines an outcome, supplies
approved context, selects tools and permissions, supervises progress, checks the
result, and remains responsible for handoff or escalation.

This report analyzes 100 distinct evidence records drawn from 78 public source
URLs. The records are distributed across five predeclared strata: 25 employer or
workplace implementations, 20 U.S. labour-demand signals, 20 occupational task
records, 20 empirical observations from five method-disclosed studies, and 15
platform, technical, safety, or public-interest records. The design is purposive,
not representative. It identifies transferable skills and claim boundaries; it
does not estimate national prevalence or prove productivity, income, or employment
effects.

The central conclusion is practical: the durable capability is not familiarity
with one chatbot or model name. It is the ability to turn a work objective into a
controlled agent workflow and to know when the system should stop, ask, escalate,
or hand work back to a person.

## Research question and method

The study asks a deliberately narrow question: which practical capabilities recur
when U.S. workers and organizations move from conversational AI use toward
multi-step, tool-using work? It does not ask how many jobs will be automated, which
vendor will lead the market, or what return an organization should expect from a
particular system. Those questions need different sampling designs and, in many
cases, longitudinal or causal evidence.

Before collection, the research geography was fixed as the United States and the
five evidence strata and minimum counts were declared. A record was accepted only
when it had a public URL, identifiable publisher, U.S. jurisdiction or a clearly
bounded U.S. application, a distinct evidence proposition, a rights note, and a
claim-use boundary. Global documentation from U.S. platform providers was allowed
only to support dated capability or safety propositions. It was not allowed to
establish U.S. adoption prevalence. Search snippets could help discover a source,
but the record had to resolve to the underlying page or document.

The unit of analysis is one independently citable task, implementation, demand
signal, empirical observation, capability, or safety proposition. Records were
deduplicated using source identity, organization, title, evidence meaning, and the
task/output pair. A single report could contribute more than one record only when
the observations were analytically different. This is why the final corpus is
described as **100 evidence records from 78 unique source URLs**. It is not a set
of 100 independent publications.

Each record was coded for work context, task, output, interaction mode, autonomy,
tools or data, human checkpoint, verification method, risk, failure mode,
prerequisite skill, evidence strength, and contrary evidence. The synthesis then
mapped records to eight non-exclusive skill domains. Because a real workflow can
require framing, permissions, supervision, and verification at the same time,
domain totals are not intended to sum to 100.

Several balance controls were applied. No organization supplies more than two
employer-implementation records. No empirical source group supplies more than a
quarter of that stratum. No publisher supplies more than 15 percent of the full
matrix, and at least 30 percent of records had to be non-vendor. The completed
matrix exceeds those safeguards: its largest publisher contributes 11 percent,
and 79 percent of records are classified as non-vendor under the conservative
publisher rule.

## What distinguishes agentic work from chatbot use

A chatbot interaction usually produces a response to a prompt. An agentic
workflow may instead plan or decompose a task, call tools, retrieve or update
records, act across several steps, retain state, and return a result that needs an
acceptance decision. The distinction is functional, not promotional: a product
label alone does not make a workflow agentic.

Employer evidence shows this progression in several contexts. AT&amp;T describes
agents that can update customer records or propose network fixes under oversight,
while its earlier Ask AT&amp;T work emphasized document synthesis, code drafting, and
database queries. Walmart describes agent use across customer, associate, supply
chain, and software workflows. Microsoft reports support workflows that connect
agents to trusted sources and APIs. These examples establish that bounded action
exists in practice; because they are first-party reports, they do not independently
prove return on investment or organization-wide impact.

Sources: [AT&amp;T agentic AI](https://about.att.com/blogs/2025/agentic-ai.html),
[Walmart&#039;s agentic strategy](https://corporate.walmart.com/news/2025/05/29/inside-walmarts-strategy-for-building-an-agentic-future),
[Microsoft modern support](https://www.microsoft.com/insidetrack/blog/enabling-modern-support-at-microsoft-with-ai/).

## Results across the five evidence strata

The employer stratum contains 25 observations covering 19 organizations, 10
sectors, and 24 coded functions. Twenty-two records come from first-party employer
sources. The examples include customer and employee support, research, software,
network operations, retail, supply chain, human resources, revenue-cycle work,
clinical decision support, marketing, and internal governance. The spread matters
because it shows that agentic practice is not confined to software development.
The records also show substantial variation in autonomy: some systems draft or
recommend, while others take bounded action in connected systems. Across both
forms, a person or accountable service owner remains part of the control design.

The labour-demand stratum contains 20 distinct U.S. or U.S.-eligible vacancies.
Its strongest repeated signals are workflow design, tool integration, retrieval,
evaluation, observability, guardrails, human approval, traceability, recovery, and
safe deployment. These records indicate what sampled employers were seeking on
the retrieval date. They do not reveal how frequently those skills appear across
the whole labour market, and their technical-role concentration means they should
not be projected onto every knowledge-work occupation.

The occupational stratum contains 20 records drawn from U.S. Bureau of Labor
Statistics task descriptions and U.S. Department of Labor AI-literacy guidance.
It anchors the analysis in work people already perform: analyzing information,
coordinating projects, communicating findings, maintaining records, applying
professional judgment, checking outputs, protecting data, and knowing when to
escalate. These records make no claim that an agent can safely take over an
occupation. Their role is to show where agentic tools meet existing human tasks
and responsibilities.

The empirical stratum contributes 20 observations from five source groups, with
no group contributing more than four. It captures adoption, reported use,
organizational readiness, barriers, and changing patterns of cognitive work. The
platform and safety stratum adds 15 dated records spanning OpenAI, Google,
Anthropic, NIST, the U.S. Department of Labor, and the Federal Trade Commission.
Taken together, these two strata add an essential corrective: capability is moving
quickly, but readiness, trust, control, and actual use remain uneven.

## Eight transferable skill domains

### 1. Select the right task and define the outcome

The first decision is whether a task should be delegated at all. A responsible
workflow starts with a concrete work product, boundaries, acceptable sources,
constraints, and a quality test. Current vacancy signals include workflow
discovery, reference designs, acceptance criteria, and customer or business
translation—not only model configuration. U.S. occupational sources similarly
show that requirements gathering, decision support, communication, and quality
control are core work activities.

### 2. Prepare context, sources, and data

Agents cannot infer which internal source is authoritative or which data may be
used safely. Evidence across research synthesis, sales, support, and enterprise
knowledge work repeatedly points to grounding, source selection, structured
inputs, and controlled access. More data is not automatically better; irrelevant,
stale, confidential, or unlicensed material can make the result worse or create a
new risk.

### 3. Choose tools, environments, and permissions

The move from text generation to action creates an authority problem. A useful
agent may need a browser, code environment, knowledge base, API, or enterprise
system. Each connection expands what the system can see or change. The worker
therefore needs to choose the minimum tool and permission set appropriate to the
task, distinguish read from write access, and keep irreversible or sensitive
actions behind approval.

### 4. Decompose and delegate

Agentic work often requires breaking a broad objective into steps, assigning
tools or specialized agents, and preserving dependencies. The evidence does not
support treating orchestration as automatic correctness. Good decomposition makes
intermediate outputs inspectable and creates places where a person can intervene.

### 5. Supervise progress and handle exceptions

Human oversight is most useful when it is designed into the workflow. The
evidence includes approval gates, escalation to specialists, review of clinical
or financial outputs, and recovery paths for unexpected states. A human “in the
loop” is not a safety guarantee unless that person has sufficient context,
authority, time, and a clear decision rule.

### 6. Verify outputs and preserve traceability

Current U.S. vacancy signals repeatedly ask for evaluations, regression tests,
observability, tracing, monitoring, benchmarks, and incident handling. These are
not exclusively engineering ideas. A knowledge worker also needs to verify facts,
sources, calculations, policy alignment, completeness, and whether the result
meets the original acceptance test. The verification method should match the
failure cost.

### 7. Integrate privacy, security, rights, and policy

Risk control belongs inside the workflow. NIST&#039;s AI Risk Management Framework
profile, U.S. Department of Labor AI literacy guidance, Federal Trade Commission
material, and vendor safety documentation all reinforce the need to consider
data sensitivity, authority, provenance, security, and foreseeable harm before an
agent acts. These sources inform operating practice; they are not legal advice and
do not establish compliance in every jurisdiction.

Sources: [NIST Generative AI Profile](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence),
[U.S. Department of Labor AI Literacy Framework](https://www.dol.gov/sites/dolgov/files/ETA/advisories/TEN/2025/TEN%2007-25/Attachment%20I%20%28Accessible%20PDF%29.pdf),
[Anthropic trustworthy agents](https://www.anthropic.com/research/trustworthy-agents).

### 8. Hand off, document, and operationalize

A useful workflow ends with an accountable handoff: what the system did, what
sources and tools it used, what remains uncertain, who approved the result, and
what should happen next. Documentation also makes a workflow teachable,
repeatable, auditable, and easier to improve.

## What the labour-demand evidence adds

The 20 U.S. or U.S.-eligible vacancy records are current demand signals rather
than a labour-market census. Together they emphasize tool calling, retrieval,
multi-agent design, guardrails, human review, evaluation, observability,
traceability, rollback, policy controls, and safe deployment. The sample is
deliberately broad across employers but is weighted toward technical, product,
and AI-enabled roles. It therefore supports a conclusion about the kinds of
capabilities employers are asking for, not how common those requirements are
across all U.S. jobs.

Examples include [Salesforce&#039;s Agentforce architecture role](https://careers.salesforce.com/en/jobs/jr343933/partner-technical-architect-agentforce/),
[DoorDash&#039;s agent evaluation role](https://job-boards.greenhouse.io/doordashusa/jobs/8013249),
and [Docker&#039;s long-running agent workflow role](https://jobs.ashbyhq.com/docker/348e2a4c-f794-4106-8c36-bb313ff15819).
Vacancy pages may change or disappear; their status should be rechecked before
external publication.

## What empirical studies add

Five method-disclosed source groups provide 20 sample-bounded observations:
Gallup, Epoch AI with Ipsos KnowledgePanel, Deloitte, Pew Research Center, and
OpenAI Economic Research. Together they show that workplace AI adoption and use
are uneven, that workers and organizations differ in readiness, and that observed
use spans several kinds of cognitive work. The studies use different populations
and methods—worker self-report, leader surveys, interviews, and product-use
analysis—so their figures should not be pooled or treated as interchangeable.

Sources: [Gallup workforce adoption study](https://www.gallup.com/workplace/704225/rising-adoption-spurs-workforce-changes.aspx),
[Epoch AI and Ipsos polling](https://epoch.ai/data/polling),
[Deloitte agentic transformation study](https://www.deloitte.com/us/en/insights/industry/technology/path-to-agentic-transformation.html),
[Pew Research Center worker study](https://www.pewresearch.org/social-trends/2025/02/25/workers-exposure-to-ai/),
[OpenAI work-use research](https://openai.com/index/how-ai-is-expanding-what-people-do-at-work/).

## Implications for business and work

The evidence supports four restrained conclusions.

First, agentic work is cross-functional. The employer stratum covers 19
organizations, 10 sectors, and 24 coded functions. Second, reliable use depends
on workflow and control skills as much as on model interaction. Third, increasing
autonomy raises the importance of permissions, monitoring, verification, and
recovery. Fourth, the most transferable learning objective is accountable
delegation, not dependence on a particular product interface.

The evidence does **not** show that an agent should replace an accountable worker,
that one platform is universally superior, or that brief training guarantees
productivity, career advancement, or income. High-stakes examples in health,
finance, security, employment, or public services should be used to teach risk
boundaries, not to encourage learners to experiment with real sensitive data.

## How the skill stack works as one operating cycle

The eight domains are best understood as a connected cycle rather than a menu of
independent techniques. A worker first decides whether the task is suitable and
defines the intended output. The worker then assembles authorized context,
selects the minimum necessary tools and permissions, and decomposes the work into
inspectable steps. During execution, the worker watches for exceptions, pauses at
approval points, and prevents the system from extending its authority merely
because it can continue. The output is then checked against sources, calculations,
policy, tests, or other acceptance criteria. Finally, the worker documents what
happened and hands the result to the person or process that owns the next decision.

This cycle also explains why “human in the loop” is too vague to be a complete
control. Oversight can occur before action, during execution, after output, or only
when an exception is triggered. The right checkpoint depends on consequence and
reversibility. A draft internal summary can usually tolerate a different review
design from a customer-record update, financial decision, clinical alert, or
security change. The corpus supports risk-proportionate controls, not one universal
approval pattern.

Verification likewise changes with the task. Factual synthesis needs source and
provenance checks. Code or configured workflows need tests and regression checks.
Policy-sensitive work needs rule and authority review. Long-running or connected
agents need monitoring, traceability, incident handling, and recovery. A fluent
answer is therefore only the beginning of evidence, not proof that the work is
correct or complete.

For organizations, this points to a practical division of responsibility. System
owners define permitted tools, data boundaries, logging, and escalation paths.
Managers define outcomes, tolerances, and accountability. Workers frame tasks,
supervise execution, verify outputs, and make or escalate decisions. Risk,
security, legal, and subject-matter specialists provide controls where the
consequences require them. The evidence does not imply that every workflow needs
all roles; it supports making ownership explicit before autonomy increases.

## Method and evidence boundaries

The corpus contains exactly 100 non-duplicative analytical records across five
predeclared strata. It resolves to 78 unique URLs because a single
method-disclosed study, public framework, or technical document may support more
than one distinct observation. Those repeated URLs are legitimate
multi-observation sources; the corpus must not be described as 100 independent
publications.

No publisher supplies more than 11 percent of the matrix. Twenty-two of 25
employer-implementation records are first-party employer sources. The empirical
stratum contains five source groups, with no group contributing more than four of
20 observations. Under a conservative publisher classification, 79 percent of
records are non-vendor evidence. These controls reduce, but do not eliminate,
selection and self-report bias.

The full methodology, record-level claim boundaries, rights status, source
ledger, geography-coherence matrix, claim-evidence matrix, and limitations are
preserved in the frozen research package.

## Trademark, product, and jurisdiction note

ChatGPT and Codex are OpenAI product names; Gemini is a Google product name; Claude
is an Anthropic product name. Their use in comparative educational or research
text should remain nominative and accurately attributed. This research does not
imply endorsement, sponsorship, certification, or partnership by those companies.
Do not use vendor logos or trade dress without separate permission. Model and
feature names are volatile and require verification immediately before public
release.

The evidence population is U.S.-focused. It does not establish compliance with
Portuguese or EU law. Any Lisbon delivery needs a separate review of privacy,
consumer information, accessibility, recording consent, certificate wording, and
applicable AI-governance obligations.


## Related professional learning

Professionals who want to turn this evidence into supervised practice can join MTF Institute&#039;s [Professional Certificate in Agentic Systems and AI in Work and Business](https://mtfinstitute.com/programs/professional-certificate-agentic-systems-ai-work-business/#enroll). The live two-hour workshop is delivered in Lisbon or online, uses each participant&#039;s own computer, and applies the report&#039;s operating cycle through instructor-guided work. It is professional education, not an academic degree; the cohort date, time and Lisbon venue are announced during recruitment.


## Archive and citation

The complete public research package, including the searchable PDF, frozen evidence matrix, source ledger, claim-evidence matrix, geography-coherence matrix, methodology and limitations, is preserved in [Zenodo record 10.5281/zenodo.22884074](https://doi.org/10.5281/zenodo.22884074). The [public PDF](https://zenodo.org/records/22884074/files/MTF-RR-2026-09-21-02.pdf?download=1) is the archival reading edition.

**Author:** MTF Institute Editorial Team.  
**Publication date:** 21 September 2026.  
**Institution:** MTF Institute.  
**Report number:** MTF-RR-2026-09-21-02.

Readers who want a dated view of recent product and operating-model changes can also read [From Answers to Actions: What Changed in Agentic AI for Work in the Last 90 Days](https://mtfinstitute.com/insights/from-answers-to-actions-agentic-ai-work/).



## Citation

When citing or summarizing this material, link to the canonical HTML page: https://mtfinstitute.com/insights/agentic-work-skill-stack-us-knowledge-work-business-practice/
