Model job description

Model Job Description for an AI Evaluation Specialist

This model describes an AI Evaluation Specialist who plans valid tests, interprets bounded evidence and hands findings to the people who own fixes and release decisions. Adapt responsibilities, systems, qualifications and approvals to the actual employer; this is a template, not a live vacancy or universal policy.

Learn the AI evaluation workflow
Resource
Model job description
Evidence
United States
Reviewed
October 4, 2026
Format
Reusable professional guide

An evidence-derived model job description for AI evaluation work, covering responsibilities, outputs, methods, authority boundaries and operating cadence.

Evidence scope: A structured purposive study of 107 directly verified current U.S. AI evaluation and adjacent vacancies from 51 employers, observed on 4 October 2026, plus a separate review of recent changes in evaluation work; the sample does not establish national prevalence.

An AI Evaluation Specialist designs or applies tests of an AI system's behavior, judges results against a stated use case, and gives decision makers traceable evidence of quality, limitations and risk. This is an evidence-derived job-description model for local adaptation, not a live vacancy or a universal employer policy. It draws on a purposive set of 107 current U.S. employer vacancies reviewed on 4 October 2026; the set describes observed role requirements, not the prevalence of those requirements across the U.S. labor market. Eleven autonomous-system roles form a distinct technical cohort and should be adapted separately for simulation, physical-world and safety-critical work.

Role purpose and scope

The role helps a team answer a practical question: does a model or AI-enabled application perform acceptably for its intended users, tasks and operating conditions? The specialist makes the evaluation question, test population, scoring method, observed failures and limits of the evidence explicit. The relevant system may be a model, agent, retrieval workflow, voice assistant or other AI application. Local owners must specify the system and decision before adopting this model.

Evaluation can be performed by a human reviewer, an engineer, a researcher or a lead. The person may execute an established rubric, build automated tests, investigate failures or coordinate a program, depending on the approved role. The description below offers selectable duties; it does not imply that one person holds every duty or final release authority.

Responsibilities to select for the local role

  • Define the evaluation objective with the product, research or safety owner: intended use, users, known risks, comparison baseline, decision criteria and what the test cannot establish.
  • Create or apply representative cases, grading rubrics, benchmark tasks or controlled test datasets. Record case provenance, version and access restrictions so results can be reproduced without exposing protected material.
  • Run approved human or automated evaluations, including model or application tests, trace review, score aggregation and comparison with a baseline or prior version. Check that prompts, reference answers, graders and metrics measure the intended behavior.
  • Review errors and disagreement. Distinguish an AI-behavior failure from a test defect, ambiguous instruction, data problem or tool failure; document evidence and uncertainty before assigning a severity or recommending a fix.
  • Communicate results as an auditable finding or scorecard: method, sample, observed behavior, limitations, likely impact, proposed follow-up and responsible decision owner.
  • Re-test agreed fixes or regressions and maintain the link between the original finding, changed system version and later result. Escalate unresolved, unsafe or ambiguous findings through the local route.
  • Where explicitly authorized, conduct bounded safety or adversarial tests against an approved system and threat model, with agreed access, isolation, stop conditions and reporting. This is an optional specialization, not blanket permission to probe systems or users.

The 11 autonomous-system postings in the source set call for their own evaluation domain, such as scenario sampling, simulator evidence and physical-system interpretation. A general AI assistant example below is not a substitute for that domain's engineering and safety controls.

Expected work products

The local description should name the products the person actually owns. Common options in the reviewed vacancies are:

  • An evaluation brief specifying use case, risk, acceptance criteria, metric definitions, sample and known limits.
  • A versioned case set or benchmark with rubric, reference material, permitted data source and reviewer instructions.
  • A scorecard or dashboard with denominators, uncertainty or disagreement notes, meaningful slices and a baseline comparison.
  • A finding record with input or trace reference, observed behavior, expected behavior, severity rationale, reproduction steps and escalation status.
  • A regression record linking a finding to a fix, system version, re-test outcome and open decision.

These are alternative outputs across roles. A reviewer may chiefly produce scored cases and rationales; an engineer may own a harness and regression suite; a lead may own quality standards and decision-ready summaries.

Capabilities and working methods

Hard skills. Select capabilities that match the actual system: evaluation design; rubric and benchmark validity; sampling and statistical interpretation; test-case construction; structured error analysis; reproducible documentation; and privacy-aware handling of test data. Technical roles may also need Python, SQL, software testing, model or agent trace analysis, data pipelines or an evaluation harness. Human-review roles may place greater weight on consistent rubric use, calibration and written adjudication. A specialized safety role needs an approved threat model and controlled adversarial test method. Named languages or frameworks should appear as requirements only when the local work truly depends on them.

Observable working behaviors. The person explains why a case received a score, records uncertainty instead of guessing, compares disagreements against the rubric, invites relevant domain review, and changes a conclusion when stronger evidence appears. They communicate a finding in terms a product or safety owner can act on, preserve a clear decision trail and escalate beyond their authority. These behaviors can be assessed from work samples and review records; they are more useful than generic personality adjectives.

Tools and systems. Choose from the employer's approved evaluation environment, case or benchmark repository, annotation or review interface, model and application logs, trace viewer, version control, analysis notebook, dashboard and issue tracker. Use access-controlled test data and approved AI endpoints. The role model does not prescribe a universal vendor stack.

Levels, experience and entry route

The reviewed U.S. postings span reviewers, engineers, researchers and leaders. Title evidence includes 21 staff or principal, 16 senior, 11 manager or lead, four director or executive and two early-career entries; 53 of 107 do not support a reliable title-based level. These counts describe this selected sample only. No single degree, number of years or certification is a universal entry requirement for the model.

  • Developing or supervised evaluator: applies an approved rubric and case set, writes reproducible rationales, identifies unclear cases and escalates them; a named reviewer checks judgments before they support a consequential decision.
  • Independent specialist: helps design case coverage and measures, runs and interprets evaluations, diagnoses test versus system faults, and prepares evidence-backed recommendations within a defined product and access scope.
  • Senior specialist or lead: sets evaluation standards, challenges validity and coverage, calibrates reviewers, prioritizes risk, mentors others and presents trade-offs to decision owners. Hiring criteria should reflect the technical and domain depth of the actual vacancy.

An employer should publish its own formal minimum and preferred qualifications after checking the responsibilities, tooling, supervision and risk of the local role. Experience in software QA, data analysis, model evaluation, domain review or safety testing may be relevant, but the employer must decide which evidence is sufficient.

Interfaces, authority and escalation

The specialist may work with model researchers, product and engineering owners, data and annotation teams, domain experts, security and safety staff, privacy or policy reviewers, and customer-facing teams. Record who supplies the use case and test data, who approves the method, who owns remediation and who decides whether a system can ship or remain in use.

The specialist can normally document a finding, recommend additional testing, request adjudication and raise a concern. Any ability to block a release, change a production system, access sensitive data or conduct adversarial testing must be granted explicitly by local policy. For suspected harmful behavior, compromised test data, privacy exposure, an out-of-scope test or a contested high-severity result, preserve the relevant evidence and use the designated escalation route. Do not infer final deployment authority from responsibility for evaluation.

A workable cadence

These are planning prompts, not a universal schedule. Set frequency according to release rhythm, case volume and the organization's risk process.

  • Daily or each review session: check the approved version and rubric; run or review assigned cases; record scores, rationales and ambiguous results; raise urgent concerns promptly.
  • Weekly or each evaluation cycle: inspect score and failure patterns, calibrate disagreements, refine case coverage, review open findings with owners and confirm re-test priorities.
  • Monthly or each planning cycle: reassess whether the test set still reflects intended use and emerging failure modes; review metric drift, reviewer consistency, data access and documentation quality with the owning teams.
  • Event driven: before a material model, prompt, tool or policy change, agree the regression scope and decision threshold; after a serious incident or unexpected result, preserve evidence, contain testing as directed, investigate with the right owners and document the re-test and decision.

Reusable blank description model

Copy this form into a local draft. Before filling it, confirm the owning team and system, the evaluation decision, the duties and outputs actually assigned, and the people who approve methods, data access, remediation and release. Set formal minimums separately from optional qualifications; name only approved tools, work terms and escalation routes.

Field Local entry
Role title and level
Team, system and intended users
Purpose
Core duties
Work products
Minimum qualifications
Preferred qualifications
Methods and approved tools
Interfaces
Authority and escalation
Cadence and quality checks
Work arrangement

Complete the local-entry cells with approved facts and remove unused rows before sharing a filled description. Keep formal minimums distinct from preferred experience; do not turn an unlabeled applicant profile into a mandatory threshold.

Completed example

Fictional example for learning purposes.

Role title and level: AI Evaluation Specialist, independent contributor. Team and reporting line: Product Assurance; reports to the evaluation lead. System and users: A customer-support assistant that retrieves approved policy articles and drafts answers for human agents. Agents review every draft before sending it to customers.

Purpose: Determine whether assistant drafts answer the customer's question accurately, use the current policy source, avoid unsupported claims and hand off uncertain cases to a human. Provide the product owner with evidence for a change decision; the specialist does not approve production release.

Core duties: Agree test scope and acceptance measures with the product owner and support policy lead. Maintain a versioned set of routine, edge and high-impact support cases, with reference sources and a scoring rubric. Run approved tests against named assistant versions, check retrieved evidence and traces, review disagreements with a second reviewer and distinguish policy ambiguity from assistant failure. Document reproducible defects, likely user impact, uncertainty and recommended follow-up. Re-test approved fixes and report remaining limitations before each scheduled release review.

Work products: An evaluation brief; controlled case set and rubric; weekly scorecard with case counts and important slices; finding records linked to traces and policy versions; and a regression summary for the release review. Store customer-like examples only in the approved test environment and use sanitized data.

Minimum qualifications: Demonstrated experience in software or content QA, structured review or AI application evaluation. Ability to design cases, apply a rubric consistently, inspect logs or comparable records, analyze error patterns and write concise findings. Able to use the team's approved spreadsheet or analysis tool.

Preferred qualifications: Experience analyzing larger case samples with Python or SQL. Familiarity with retrieval-based assistants and support policy review.

Working behaviors: Explains scoring with cited case evidence; flags uncertain policy interpretations; seeks second review when the rubric does not resolve a case; and tells the product owner what the evidence can and cannot support.

Interfaces and authority: Product owns the intended use and release decision; support policy owns the reference answers; engineering owns fixes; privacy approves test-data handling; the evaluation lead approves method changes. The specialist may recommend a hold and escalate a high-impact or privacy-relevant finding to the evaluation lead immediately, but cannot change production settings or authorize launch.

Cadence: Record assigned case outcomes each review day. Calibrate disputed cases and share an open-finding summary weekly. Revisit case coverage, policy versions and reviewer agreement monthly. Run the agreed regression suite whenever the model, retrieval source, prompt or policy materially changes; for an urgent incident, preserve traces and follow the incident route.

Local adaptation and quality check

Local adaptation and quality check

  • Identify the actual system, users, geography and evaluated decisions; avoid copying responsibilities that belong to a different AI product or an autonomous-system team.
  • Mark each duty as owned, shared, supervised or out of scope. Remove optional red-team work unless there is an approved threat model, access scope, stop condition and reporting owner.
  • Match every minimum qualification to an essential duty; place genuinely optional experience in preferred qualifications. Check local employment, accessibility and legal wording with the appropriate owner.
  • Name the real tools and permitted data; define benchmark or rubric custody, result storage, retention and trace access under local policy.
  • Check that each promised work product has a user and decision, and that measurements disclose their sample, baseline and limitations.
  • Name method approval, remediation and final release owners. State the escalation route for unsafe, privacy-relevant, ambiguous or contested findings.
  • Read the filled description as a candidate and as a decision owner: can each party tell what the person will do, what evidence they will deliver and what they may decide?

The underlying U.S. vacancy study explains the sample and coding limits. The separate current-changes analysis adds dated context on evaluation planning, test validity, safe test environments and finding escalation; it does not turn a single organization's practice into a universal job requirement. Representative employer descriptions include Anthropic's Cyber Evaluations Engineer, Innodata's Agentic Workflow Reviewer, Apple's evaluation-framework engineer and Waymo's Statistical Evaluation and Sampling engineer. The Waymo page illustrates the separate autonomous-system context, not a general requirement for AI assistant evaluation.

Quick reference

Use the resource in five moves

  1. Read the role purpose and expected outputs.
  2. Compare the model with the local role and authority boundaries.
  3. Select only statements supported by real evidence.
  4. Adapt the reusable fields without inventing experience or approvals.
  5. Review the result with the accountable person before operational use.