# AI Evaluation Specialist Role SOP and Operating Playbook

A model AI evaluation operating playbook from authorized intake and evidence design through testing, scoring, escalation and decision handoff.

**Practise the AI evaluation operating cycle:** [Open the course and enrol](https://mtfinstitute.com/programs/ai-evaluation/#enroll)

**Resource type:** role sop operating playbook  
**Evidence geography:** United States  
**Evidence scope:** A structured purposive study of 107 directly verified current U.S. AI evaluation and adjacent vacancies from 51 employers, observed on 4 October 2026, plus a separate review of recent changes in evaluation work; the sample does not establish national prevalence.  
**Accepted source SHA-256:** `7c6bcd5b5444e69ac1874120c85a071151bb67811c4879b95e74ce763cdaf67b`

An AI Evaluation Specialist turns a bounded question about an AI system into testable criteria, reproducible observations, a clear finding, and a handoff to the people who own the system. This playbook is an evidence-derived model for adapting to a specific workplace. It does not assign the specialist universal access, permission to test a live target, or authority to release a product. The local system owner and approval process set those boundaries for each engagement.

## Purpose and operating boundary

Use this procedure when a team needs to know whether a model or AI application behaves acceptably for a stated use: before a change, after a reported failure, during a scheduled quality review, or at a release checkpoint. The result is a decision-ready record. It should show what was tested, what was not, how outputs were judged, what failed, and which owner must act next.

The evaluation object may be a base model, an agent with tools, a customer-facing feature, a safeguard, or an autonomous system. These settings need different cases, measures, environments and specialist reviewers. Do not transplant a language-model rubric into a vehicle or robotics release test. Use the relevant local safety process and domain owner for physical systems.

Before any run, record the employer-controlled boundaries: approved system and version, permitted environment and accounts, data and test-set rights, allowed test methods, evaluator access, logging and retention, stop conditions, escalation contacts, and the person who can accept or reject a release risk. A red-team or cyber test needs its own explicit authorization and controlled environment. A request to “try it and see” does not replace that approval.

## People, inputs and handoffs

One person may fill several roles in a small team, but name the decision owner in the record rather than assuming that the evaluator holds every right.

- **Requesting owner:** states the product or research question, intended use, affected users, change being assessed, and decision deadline. This person confirms that the question matters to the next decision.
- **Evaluation specialist:** drafts the test plan, prepares approved cases and scoring, runs or coordinates the assessment, checks reproducibility, and writes findings with limits. The specialist may recommend action within the approved assignment.
- **System or model owner:** supplies version identity, known constraints, telemetry and a route to reproduce and fix behavior. This owner accepts work on a defect and confirms the version to retest.
- **Domain reviewer or calibrated human rater:** judges cases that require subject expertise, applies the agreed rubric, records disagreement, and explains uncertainty. A single uncalibrated score should not be treated as settled truth.
- **Data, privacy, security or safety owner:** approves access and handling where the cases, environment or finding warrant it; defines special stop and reporting routes.
- **Release or risk decision owner:** receives the evidence and makes the launch, hold, monitoring or exception decision under local policy. The evaluation specialist does not infer this right from participation in the review.

The minimum inputs are a system/version identifier, use and risk question, approved environment and data, acceptance criteria or a named owner who will define them, decision date, relevant known failures, and the local contact and escalation route. Include a prior baseline if one exists. For a first version, record that no comparator exists and agree absolute criteria before the run. If a required input is missing, record the gap and seek a scoped answer before making a pass/fail claim.

## Trigger-to-close workflow

The steps below form one ordered evaluation cycle. A local team may add controls, but skipping a decision-critical step should be visible in the final record.

1. **Open the request.** Record the trigger, requester, intended use, affected population, system boundary, version, decision to be informed, deadline and owner. Decide whether the request is a routine quality check, a regression, a safety concern, a release review or an authorized adversarial exercise. Do not start a live test while the permitted system, environment or data rights are unclear.
2. **Confirm scope and approval.** Write a short plan that names the in-scope behaviors and explicit exclusions. Identify the local approver for access, special testing, data handling and any release decision. Record the stop conditions and the route for unexpected exposure or unsafe behavior. Ask a domain owner to review criteria that depend on local policy or specialist knowledge.
3. **Design the evidence.** Choose the case population and sampling logic. Include ordinary use, known failure modes, relevant edge cases and appropriate negative controls. Define the rubric or measurable outputs before viewing results. State how human raters, automated judges and disagreement will be calibrated. Protect holdouts and record case provenance. Reuse a benchmark only when its validity for this system and decision is defensible.
4. **Prepare a reproducible run.** Confirm the approved environment, model and application versions, prompts or input formats, tool configuration, data snapshot and evaluator permissions. Record the prior baseline if available; otherwise use the absolute criteria agreed before results are seen. Check that logging captures enough to reproduce a finding without collecting unnecessary private information. Run a small pilot to catch broken cases and ambiguous scoring; revise the plan and version it before the main run.
5. **Run and score within scope.** Execute the approved cases and capture outputs, traces or reviewer notes under the local data rules. Record failures, refusals, tool results, latency or cost only when relevant to the defined criteria. Keep raw observations separate from interpretation. If an automated grader is used, compare a suitable sample with human judgment and note where it is unreliable.
6. **Investigate and classify.** Reproduce significant failures where permitted. Separate a model-behavior issue from retrieval, tool, data, orchestration, test-case or scoring defects. Record severity using the local rubric, affected case class, frequency within this test set, uncertainty, and whether the issue is new or a regression. Do not project a small purposive case set onto all users.
7. **Hand off remediation and retest.** Route each issue to an accountable owner with a minimal reproduction record, expected behavior, observed behavior and proposed next check. Track the owner response and the version changed. Retest the affected cases and a relevant regression slice; record both improvements and newly introduced failures. Keep an unresolved issue open rather than converting a promised fix into a passed test.
8. **Prepare the decision record and close.** Summarize scope, coverage, methods, results, limitations, open risks and the options available to the authorized decision owner. Distinguish “meets this evaluation's criteria,” “does not meet them,” and “insufficient evidence.” Record the owner's decision or pending action, link the approved records, set the next review trigger, and close only when ownership of every material open item is explicit.

## Operating cadence and work rhythm

These are useful checkpoints, not a universal employer schedule. The local owner sets the actual frequency, service targets and meeting calendar. A role assigned only to a project may perform the same actions at its project milestones.

- **Daily or per active run:** check the approved run queue and version changes; confirm case and grader health; inspect new results or alerts; log material failures; route urgent findings through the local stop and escalation path; update owners and retest status.
- **Weekly or per evaluation cycle:** review coverage and error slices with the system and domain owners; calibrate human or automated scoring; identify brittle cases and false alarms; verify completed fixes against regressions; publish a short status readout with open decisions.
- **Monthly or per planned review:** revisit whether cases represent current use, whether holdouts remain protected, whether metrics still predict useful outcomes, and whether recurring failure classes need a new test population. Review access, retention and the list of unresolved accepted risks with the responsible owners.
- **Event-driven:** open or refresh the evaluation when a model, prompt, tool, retrieval source, policy, user population or deployment context changes; when an incident or customer report reveals a new failure mode; before a release gate; or when an authorized safety exercise is scheduled. Decide whether old results are still valid before reusing them.

## Decisions, escalation and exceptions

The specialist can normally choose and explain a method inside an approved evaluation plan, flag a case as ambiguous, request another reviewer, and recommend a fix or more evidence. Local policy decides whether the specialist can pause a run or change its scope. A release, production rollback, public disclosure, access exception, or acceptance of residual risk belongs to the named owner unless the employer explicitly assigns it to the specialist.

- **Stop and escalate immediately under the local route** if the test reaches a system or data source outside its approved scope, exposes personal or confidential information, or creates an unexpected safety or security concern. Preserve only the minimum evidence needed for the responsible team to investigate; do not keep probing to make the result more dramatic.
- **Pause the conclusion and repair the method** if a case is broken, a rubric conflicts with the requested behavior, an automated judge disagrees materially with calibrated humans, a holdout may have leaked, or the model version changed during the run. Label affected results invalid or provisional until rerun.
- **Escalate a consequential failure** with the observed behavior, reproducibility, affected use, uncertainty, owner and proposed next check. Do not assign severity from intuition alone where the employer has a severity scale or domain reviewer.
- **Mark insufficient evidence** when coverage is too narrow, observations are missing, a test environment differs materially from use, or the decision threshold was set only after results were seen. Ask the decision owner whether to expand the assessment, accept a bounded uncertainty, or defer a decision.
- **Document disagreement** between reviewers or teams rather than averaging away a substantive conflict. Record the differing interpretations, the rubric text, an adjudicator and the final rationale.

For authorized adversarial work, the local approval must state target, environment, timing, permitted methods, accounts, containment, observation and stop conditions. This playbook describes the approval and finding path; it is not a procedure for attacking an unapproved system.

## Records and quality measures

Keep the record small enough to be usable and complete enough that another authorized reviewer can check the conclusion. Link to employer-approved storage rather than copying sensitive payloads into a broad report.

- **Intake and approval:** requester, system/version, purpose, decision owner, permitted scope, environment, data rights and approval date.
- **Evaluation plan:** hypotheses, case population and exclusions, sampling, scoring rules, a prior baseline if available or pre-agreed absolute criteria, human/automated grader checks, stop conditions and version.
- **Case register and run log:** case IDs and provenance, protected-holdout status, run identity, relevant configuration, outputs or trace references, scorer and timestamp.
- **Finding and retest log:** expected and observed behavior, reproduction confidence, cause hypothesis versus confirmed cause, severity source, owner, mitigation and retest result.
- **Decision record:** findings by severity and coverage, unresolved risks, limits, recommendation, decision owner, decision and next review trigger.

Useful measures include coverage of agreed case classes, valid-case rate, rater agreement, automated-grader agreement with human review, reproducibility of material findings, time from finding to owner assignment, retest completion, and the count of unresolved material issues at a decision point. Define each numerator, denominator and target locally. A higher raw pass rate is not automatically better: the case set may have become easier, the judge may have drifted, or risky cases may have been excluded. Report those changes alongside the score.

## Reusable blank procedure model

Copy this blank form into the team's approved document system. Complete each row with local facts. If an owner has not made a decision, write “not yet decided” in that row and leave the evaluation conclusion open rather than guessing.

| SOP field | Local entry |
|---|---|
| System, application, version and evaluation name | |
| Requester, evaluation specialist and decision owner | |
| Trigger, intended use and decision to inform | |
| Approved environment, accounts, data and permitted methods | |
| Exclusions, approvals, stop conditions and escalation contact | |
| Case classes, sampling, protected holdouts and case owners | |
| Expected behavior, rubric, measures and human-grader check | |
| Prior baseline if available, or pre-agreed absolute criteria; run configuration, pilot and main-run window | |
| Observation and trace location; access and retention rules | |
| Validity checks, disagreement route and severity source | |
| Finding owners, reproduction detail and retest scope | |
| Decision options, authorized approver and open risks | |
| Closure owner, record location and next review trigger | |

The employer controls the actual field values, access list, severity scale, thresholds, retention period, release authority and reporting route. If a model or tool makes a draft, a human still checks it against the approved observations and local rules.

## Worked cycle

Fictional example for learning purposes.

A software team is preparing a new version of a customer-support assistant for a retail returns page. The assistant can retrieve a public policy excerpt and can ask a customer to contact support; it cannot issue refunds. The product owner asks whether the new version gives correct guidance when order status is unavailable. The evaluation specialist records the candidate version, staging environment, policy snapshot, decision owner and approved test accounts. The product owner approves twelve synthetic cases: eight ordinary questions, two ambiguous requests and two cases where the status tool returns unavailable. The criterion is to avoid claiming a refund or an order status the assistant cannot verify, and to give the approved support handoff when the tool cannot answer.

The specialist pilots two cases with a domain reviewer. One expected answer is unclear because the policy excerpt does not address exchanges. They mark that case invalid for the current run and ask the policy owner to clarify the expected behavior. The product owner and policy owner approve a substitute case and the versioned rubric before the main run. The main run then uses twelve valid cases. Ten meet the agreed criterion. In two unavailable-tool cases, the assistant states that a refund has already been issued. The run log links each case ID to the candidate version, policy snapshot, tool result and assistant response. A second reviewer confirms the same two failures under the rubric. The specialist reports **2 failures in this 12-case designed set**, without presenting that fraction as a customer-wide defect rate.

The system owner reproduces the issue and traces it to a fallback response that treats “tool unavailable” as “refund confirmed.” The specialist records that explanation as an owner-confirmed cause, asks for a fix, and recommends holding the affected release decision until retest. The owner changes the fallback and supplies a new version. The specialist reruns the two failing cases plus a small ordinary-use regression slice; the false refund claim does not recur in those cases, and the assistant now gives the approved support handoff. This limited retest does not establish that the new version passes the original release review. The product owner keeps the broader release on hold, assigns a rerun of the full twelve-case set and the agreed regression suite, and sets the next review date. The specialist records that hold decision, the case invalidated during the pilot, the two original failures, the limited retest result and remaining coverage limits. This evaluation cycle closes with a linked follow-up owner and date; the release gate stays open until the full rerun is reviewed.

## Adaptation and quality check

Before using this model in a workplace, confirm that the named people actually hold the stated responsibilities, the test environment and data are approved, the rubric fits the local product and users, and the stop/escalation route is known. Check that every finding can be traced to a system version, case, observation and scorer; that invalid cases and disagreement remain visible; that a fix was retested rather than merely promised; and that the final note separates observation, interpretation, recommendation and the authorized owner's decision. Remove any field the local team does not use, but retain the information needed to explain a consequential decision.

## Source and use notes

This model synthesizes a structured purposive review of 107 U.S. employer vacancies checked on 4 October 2026, not a single official occupation or a universal company policy. Eleven autonomous-system positions are a separate technical cohort; their physical-system evaluation methods are not ordinary assistant-evaluation steps. The [MTF Institute research report](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-107-us-vacancies-2026/) describes the sample and its limits; the separate [current-changes review](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-2026-recent-changes/) explains recent concerns about test validity, protected evidence and controlled high-risk evaluation. Direct employer examples include [Figma's AI evaluation research role](https://job-boards.greenhouse.io/figma/jobs/6112112004) for rubrics and decision readouts, [Abacus Insights' AI quality role](https://job-boards.greenhouse.io/abacusinsights/jobs/8853572002) for scenario suites and release criteria, [Innodata's agentic workflow reviewer](https://job-boards.greenhouse.io/innodatainc/jobs/4412191009) for rubric scoring and escalation, and [NVIDIA's LLM safety program role](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Senior-Technical-Program-Manager---LLM-Safety_JR2024325) for mitigation ownership and unresolved-risk escalation. Their duties support the model's alternatives; none sets the procedure for another employer.

## Connected role pathway

- [ats resume template](https://mtfinstitute.com/insights/ai-evaluation-specialist-ats-resume-template/)
- [model job description](https://mtfinstitute.com/insights/ai-evaluation-specialist-model-job-description/)
- [role sop operating playbook](https://mtfinstitute.com/insights/ai-evaluation-specialist-role-sop/)
- [Vacancy evidence](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-107-us-vacancies-2026/)
- [Current-practice analysis](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-2026-recent-changes/)

**Study the Professional Certificate in AI Evaluation:** [Open the course and enrol](https://mtfinstitute.com/programs/ai-evaluation/#enroll)

Canonical URL: https://mtfinstitute.com/insights/ai-evaluation-specialist-role-sop/
