NIST presents its AI Risk Management Framework Playbook as a set of suggested actions, not a checklist. How large and operational is that action space? MTF Institute counted 592 visible suggested-action bullets across all 72 public Playbook subcategories on 7 September 2026 and analysed their distribution, action orientation and recurring themes.

Executive finding

AI risk management in the Playbook is strongly evidence-intensive. MEASURE contains 239 of 592 actions (40.4%), more than any other function. In transparent, non-exclusive keyword coding, measurement, testing and monitoring language matched 257 actions; risk, impact and tolerance matched 216; and human or stakeholder participation matched 155.

Function Subcategories Visible actions Share
GOVERN 20 152 25.7%
MAP 17 118 19.9%
MEASURE 22 239 40.4%
MANAGE 13 83 14.0%
Total 72 592 100%

The result does not mean MEASURE is more important. It shows that the public Playbook gives the densest menu of actions to evaluation and measurement. A policy-only governance model therefore omits much of the practical workload.

Research question

How is practical AI risk-management work distributed across the NIST AI RMF Playbook, which operational themes recur, and what minimum evidence packet can learners and managers build from the pattern?

Scope and method

The study covers the four public Playbook function pages - GOVERN, MAP, MEASURE and MANAGE - and all 72 subcategories visible on 7 September 2026. Every subcategory was expanded in the Codex in-app browser.

The unit of analysis was one visible bullet under “Suggested Actions”. Nested bullets were counted separately when the page presented them as distinct actionable units. The resulting census contained 592 observations. For every subcategory, the action count was recorded. Across the 72 records, the minimum was 1, the median 7, the mean 8.22 and the maximum 36.

The study normalized the first alphabetic token of each action to describe opening verbs. It also applied a predeclared set of case-insensitive keyword and phrase patterns. Theme flags were non-exclusive: one action could match more than one theme, so thematic totals must not be summed to 592.

The supporting workbook publishes the 72-row sampling frame, function totals, thematic aggregates, opening-verb counts, method and limitations. Official NIST action text remains at the primary source and is not redistributed in bulk.

Where the Playbook is most detailed

Subcategory Visible actions Operational emphasis
MEASURE 2.11 36 Fairness and bias
MEASURE 2.5 24 Validity and reliability
MAP 2.3 21 Scientific integrity and TEVV
GOVERN 1.4 20 Transparent process and documentation
GOVERN 2.1 18 Roles and communication
MAP 1.1 17 Context and intended use
GOVERN 1.2 15 Trustworthiness in policy and process
MEASURE 2.2 14 Human-subject representativeness
MANAGE 2.1 14 Resources and non-AI alternatives
MEASURE 2.8 13 Transparency and accountability

Action density is descriptive. More bullets do not prove greater importance, effort or maturity. It does show which subcategories a practitioner should expect to decompose into several distinct records or tests.

Recurring operational themes

Non-exclusive theme Matched actions Share of 592
Measurement, testing and monitoring 257 43.4%
Risk, impact and tolerance 216 36.5%
Human and stakeholder participation 155 26.2%
Documentation, evidence and traceability 115 19.4%
Governance, policy and accountability 113 19.1%
Data privacy and quality 73 12.3%
Incident, change and decommissioning 61 10.3%
Training, skills and resources 55 9.3%
Third party, supply chain and procurement 34 5.7%
Security and resilience 22 3.7%

Keyword matching is a transparent descriptive device, not a semantic model. An idea may appear without using the selected vocabulary, and a match does not prove depth. Still, the pattern is useful: test and monitoring work appears more often than policy/accountability vocabulary, and participation appears in more than one quarter of the action set.

Action orientation

The most common normalized opening verb was establish (102 actions), followed by identify (40), document (23), verify (21), evaluate (20), utilize (19), define (17) and assess (13). Opening verbs are imperfect because bullet grammar varies. They nevertheless reinforce an operational reading: practitioners are asked to create, locate, record, test and maintain evidence.

EVIDENCE-7: practical application

Step Manager action Minimum artefact
Mandate Define the use, owner and decision boundary Approved use-case brief
Context Map people, systems, data and affected groups Context and impact map
Test Specify performance, bias, privacy and security tests Test protocol and results
Decide Record thresholds, residual risk and authority Decision memo
Monitor Track drift, incidents, feedback and change Monitoring register
Respond Contain, correct, escalate or retire Incident or change record
Review Reassess evidence and improvement priorities Periodic review pack

For students, choose one low- or moderate-impact AI use case and build the seven records without confidential data. The quality test is traceability: a reviewer should be able to move from purpose and affected parties to tests, decision criteria, monitoring and an accountable response path.

What managers should do next

  1. Select one live or proposed AI use and name the accountable decision owner.
  2. Define affected groups, data, model/service, human role and non-AI alternative.
  3. Translate risk claims into testable criteria and evidence owners.
  4. Set release thresholds before seeing results.
  5. Create monitoring and incident triggers tied to actions and authority.
  6. Review the evidence packet after material changes, not only on a fixed calendar.

MTF Institute's AI Governance Manager: Lifecycle Controls, Evidence and Oversight programme connects inventories, lifecycle controls, decision evidence, monitoring, incidents and executive oversight. Use EVIDENCE-7 as a capstone structure. The programme is professional education, not an academic degree, legal advice or a guarantee of employment or compliance.

Limitations

The Playbook is a living public resource and may change after the capture date. Bullet granularity varies, so a count is not a measure of effort or importance. Nested bullets were treated as separate units when visibly presented that way. Keyword categories are non-exclusive and depend on the selected patterns. Opening verbs are syntactic signals, not full semantic classifications. The report studies guidance, not organizations, outcomes, adoption, legal duties or control effectiveness. NIST states that the Playbook offers suggested actions and is not a checklist or requirement to perform every action.

Reproducibility and archival record

The searchable PDF and supporting workbook are archived at Zenodo under DOI 10.5281/zenodo.22642165. The workbook contains the complete 72-subcategory sampling frame, aggregate coding tables and protocol.

Sources