AI Evaluation and Red Teaming Work: Evidence from 107 Current U.S. Vacancies
MTF Institute Research Report
Report no. MTF-CF-RR-2026-10-04-AIQE-01 | 4 October 2026
Author: MTF Institute Research Team
Executive summary
AI evaluation work appears under many titles. In a structured purposive snapshot of 107 distinct U.S. employer vacancies across 51 normalized employers, people were asked to design evaluations, build or curate test material, measure model and application behavior, diagnose failures, maintain evaluation infrastructure, or make quality evidence usable in product decisions. Some roles also conduct authorized adversarial testing. These are duties observed in a selected group of employer postings on 4 October 2026; the study does not estimate their frequency across the U.S. labor market.
The fixed role taxonomy separates 11 autonomous-system evaluation positions from 96 other AI evaluation positions. This matters: testing vehicle perception or embodied world models has different data, metrics, tooling, and release context from evaluating a language-model assistant or a cybersecurity safeguard. Within the 107 selected postings, the coded pages mention test data or cases in 73/107, failure triage or remediation in 57/107, release or lifecycle assurance in 45/107, and authorized adversarial testing in 19/107. These counts describe statements in the included vacancies. Near-universal evaluation and measurement counts are partly a result of the inclusion rule itself and must not be read as market prevalence.
The practical picture is a chain of evidence: define the system and risk being tested; assemble appropriate cases and human or automated scoring; compare outputs against explicit criteria; reproduce and classify failures; document uncertainty; and hand findings to people who can improve, approve, hold, or monitor the system. Employer pages differ sharply in the level of authority they assign. Professional preparation can practise these work products and decision handoffs while respecting each employer's scope and approval rights.
Question, scope, and method
The research question is: What work do U.S. employer vacancies ask practitioners to perform when they evaluate the quality, behavior, safety, or security of AI systems? The unit of analysis is one current employer vacancy with a substantial, recurring AI-system evaluation duty. It is not an official single occupation, and its title need not contain “evaluator” or “red team.” The corpus includes research, engineering, application quality, human review, program leadership, and autonomous-system positions where the visible duty passed the same inclusion test.
Candidate postings were checked against public employer-controlled detail pages on 4 October 2026. To be retained, a page had to show a U.S. location or U.S.-remote eligibility, substantive recurring evaluation, quality, or authorized adversarial-testing work, and a visible application route. Selected employer pages are linked throughout this report. For each retained posting, the study recorded its retrieval date, location and application evidence, and the basis for inclusion and deduplication. The verified archival record contains the PDF and a dated source register listing every retained posting.
The sample is structured and purposive. Searches targeted relevant employers, titles, and duties, followed by checks of the employer's current detail page. There was no complete sampling frame of every U.S. vacancy, no probability selection, and no weighting by employer size, hiring volume, industry, or job board. “Current” means that the checked employer page offered an application route at retrieval; the page may later close.
Deduplication used employer job identity and normalized role. Location mirrors and repeated representations of one opening count once; distinct levels or functions count separately only where the job identifiers and pages show separate openings. The source screen excluded closed or redirected pages, generic software QA, generic use of AI, one-off paid studies, marketplace talent pools, and listings whose direct-employer vacancy identity could not be established.
The Role Requirements Matrix codes nine families and also captures responsibilities, outputs, hard methods and skills, observable workplace behaviors, named tools, education and experience, required versus preferred language, level, cadence, interfaces, authority, and escalation. not_disclosed means the employer page did not state the item; it is missing evidence, not evidence of absence. The table below counts a family when a source-supported code was recorded. Each cell uses its own fixed cohort denominator, showing both the observed mention and the size of the reviewed cohort.
This report paraphrases duties and uses only short necessary role descriptions. It does not reproduce job advertisements, proprietary test sets, benchmark items, or vendor training material. The separate MTF article on recent AI evaluation changes is independent context and supplies no vacancy count here.
The role cluster is wider than its titles
The fixed taxonomy assigns each accepted vacancy to one main subgroup. These are analytic labels for this corpus, not standard occupational categories.
| Fixed subgroup | Accepted postings |
|---|---|
| Evaluation engineering and platforms | 32/107 |
| Evaluation research and benchmarks | 26/107 |
| AI application quality | 14/107 |
| Evaluation program leadership | 14/107 |
| Autonomous-system evaluation | 11/107 |
| AI safety and authorized red teaming | 6/107 |
| Human AI evaluation | 4/107 |
| Total | 107/107 |
The 11 autonomous-system positions include work such as Waymo's statistical release evaluation and perception evaluation, and Atoms' autonomy evaluation lead, which links scored log replays, rare-event coverage, safety metrics, and release criteria. The 96 other positions span, for example, Figma's AI evaluation research leadership, ServiceNow's evaluation platform leadership, and Reddit's AI security model evaluation. Similar family labels across these settings do not imply interchangeable technical preparation or identical authority.
Titles are an imperfect guide. 56/107 selected titles explicitly name evaluation, evals, or benchmarks; 10/107 name red team or safety; 6/107 name quality or QA; 1/107 is human-review named. The other 34/107 fall into broader engineering, research, product, leadership, or other title families while their duties still met the inclusion rule. This is one reason title-only searches would miss parts of the cluster. In the fixed level coding, 16/107 are senior, 21/107 staff or principal, 11/107 manager or lead, and 4/107 director or executive. 53/107 did not disclose a level clearly enough for classification, and 2/107 are in a combined early-career or internship bucket. The level buckets describe these postings; they do not establish an entry-level hiring rate. Hippocratic AI's HCP Evaluator is a part-time contractor position; it should be read separately from full-time professional roles.
What the accepted postings state
The following counts are coded mentions within the included vacancies. They are not the prevalence of a skill in U.S. employment, and a not_disclosed result cannot be read as “the employer does not need it.” The first four rows are especially inclusion-conditioned: every retained vacancy had to show substantial AI-system evaluation work.
| Role Requirements Matrix family | All reviewed U.S. postings | Autonomous-system positions | Other AI evaluation positions |
|---|---|---|---|
| Evaluation scope and design | 107/107 | 11/11 | 96/96 |
| Test data and cases | 73/107 | 7/11 | 66/96 |
| Measurement and analysis | 106/107 | 11/11 | 95/96 |
| AI application or model-system quality | 103/107 | 11/11 | 92/96 |
| Authorized adversarial testing | 19/107 | 1/11 | 18/96 |
| Failure triage and remediation | 57/107 | 3/11 | 54/96 |
| Release and lifecycle assurance | 45/107 | 4/11 | 41/96 |
| Tools and environment | 100/107 | 11/11 | 89/96 |
| Observable behavioral skills | 106/107 | 11/11 | 95/96 |
The more discriminating rows show that the selected cluster extends beyond red teaming: authorized adversarial testing is stated in 19/107 postings, while ordinary failure triage and release/lifecycle work appear in larger, different sets of pages. 14/19 postings with a positive adversarial-testing code also carry a positive failure-triage/remediation code. That co-occurrence describes these selected postings, not the share of red-team jobs nationally.
Evaluation design, cases, and measurement
Employers describe evaluation as a design task before it becomes a score. Figma's Director, Research – AI Evals is asked to define quality dimensions, rubrics, golden datasets, and decision-ready readouts for AI features. Scale AI's frontier risk research role asks for risk measures, harnesses, and datasets, including clearly scoped high-risk capability tests. Innodata's speech and audio research role illustrates modality-specific judgment: coverage across languages, accents, acoustic conditions, and human versus automated scoring cannot be replaced by a generic LLM score. Turing's research scientist role links task generation, grading, contamination control, difficulty calibration, and validation to the usefulness of a benchmark.
The common work product is a defensible answer to “what did this test measure?” In some postings that is a rubric and human-calibration record; in others it is a reproducible pipeline or statistical comparison. Amazon's frontier assessment scientist is tasked with sampling and error analysis to make quality findings defensible. Google DeepMind's agent-quality research engineer is asked to calibrate automatic raters against human evaluation and diagnose agent trajectories. These examples show why a metric, its test population, its calibration, and its uncertainty need to travel together when results are handed to another team.
Application quality, failure analysis, and release evidence
For customer-facing agents, quality work includes the surrounding system, not just the base model. Cresta's senior ML role describes hallucinations, retrieval failures, tool misuse, context drift, and multi-step reasoning breakdowns, paired with offline benchmarking, online experiments, and measurable task/latency/cost criteria. Glean's Assistant Quality engineer connects evaluation and monitoring loops to real enterprise workflows. Abacus Insights' AI Systems Quality Engineer is asked to create scenario suites, automated quality gates, and measurable go/no-go criteria for agentic healthcare data systems.
The release connection varies by role. Atoms' autonomy lead explicitly owns safety/behavior metrics and go/no-go criteria for vehicle releases. Inflection AI's principal research engineer links model evaluation, regression detection, release criteria, readiness reviews, and post-release monitoring. DoorDash's GenAI infrastructure engineer builds shared traces, scores, judges, and simulations for product teams; its posting describes platform enablement rather than final product sign-off. Those distinctions should be preserved: “builds evidence for a decision” and “owns the decision” are separate responsibilities.
Authorized adversarial and safety testing
The adversarial subset is specific, bounded work. OpenAI's Red Team Specialist – Cyber describes evaluating model cyber capability and safeguards through repeatable hands-on testing, then turning results into risk assessments for security, research, product, policy, and engineering partners. Reddit's AI Security ML engineer describes adversarial robustness and feedback from red-team findings in the model lifecycle. NVIDIA's LLM safety program manager connects agentic safety evaluations to mitigation owners, residual-risk reporting, and release-readiness records. These postings do not imply that every evaluator is authorized to probe live systems, choose an attack target, or approve residual risk. Scope and permissions come from the employer's process.
Who is hiring in this selected corpus
The 107 postings belong to 51 normalized employers, but the corpus is concentrated: Apple contributes 11/107, Innodata 10/107, and Waymo 8/107; together these three account for 29/107 pages. Anthropic and Amazon each contribute 5/107, while OpenAI contributes 4/107. Multiple current pages from one employer are informative about that employer's work, but they do not give that employer proportional weight in the U.S. job market. The employer name xAI is used for the SpaceXAI posting board in this count.
The retrieval channels are concentrated too. 43/107 accepted URLs are on employer-controlled Greenhouse boards and 32/107 on employer-controlled Ashby boards. Apple Careers accounts for 11/107, Waymo Careers 8/107, Amazon Jobs 5/107, and other employer/ATS domains the remaining 8/107. This mix reflects where the study found and could verify relevant public pages. It is a source-access characteristic of the study, not a claim about the share of U.S. employers using any hiring system.
Required, preferred, and unstated criteria
Job-ad language distinguishes a responsibility from a hiring requirement. Figma's AI-evals leadership page presents automated evaluation pipeline/tooling experience as an added plus, while describing extensive research and management experience in its candidate profile. Innodata's speech research page expressly requires a technical bachelor's degree and prefers an advanced degree. Apple's ML/GenAI Evaluation engineer describes several minimum qualifications while marking a master's degree as strongly preferred and considering a bachelor's degree with longer experience. These are employer- and role-specific statements, not one qualification rule for the whole cluster.
The matrix stores explicit required and preferred phrases separately. Some pages use “minimum,” “basic,” or “required”; others use “ideally,” “you might thrive,” or “nice to have.” A family coded as stated in a duty section does not become a required candidate qualification. Conversely, a preferred tool mentioned in a posting should not be presented as a universal prerequisite. The source set is mixed in role level, with many unclassified titles, so it supports no single degree or experience threshold for the cluster.
Implications for professional preparation
The observed work suggests that preparation should produce inspectable evidence of method and judgment, rather than a claim of familiarity with a particular model or tool. Depending on the target role, useful practice artifacts include:
- An evaluation scope note stating intended use, system boundary, quality or risk hypotheses, test population, and decision that results will inform.
- A case set and rubric with sampling logic, coverage and exclusions, versioning, human-review instructions, and a record of how graders were calibrated.
- A measurement record showing baseline, metrics, uncertainty or disagreement, error slices, repeatability checks, and the limits of automated judging.
- A failure and regression log that reproduces a model or agent behavior, separates model, retrieval, tool, data, and orchestration causes, and tracks a retest after a proposed change.
- A decision handoff that makes unresolved risks, tradeoffs, owner, escalation route, and evidence needed for launch, hold, or monitoring explicit—without claiming approval authority the role does not hold.
These artifacts follow the work visible across the Figma evaluation, Abacus application-quality, Waymo release-evaluation, and OpenAI cyber red-team postings. They are examples of work products to practise across different specialisms. An employer's approved environment, data rights, safety policy, and decision process remain decisive.
Selected primary posting examples
The links below illustrate the different work settings in the sample; the report's counts use all 107 accepted pages.
| Work setting | Employer posting |
|---|---|
| AI product quality research | Figma, Director, Research – AI Evals |
| Frontier risk evaluation | Scale AI, Research Scientist – Frontier Risk Evaluations |
| Speech and audio assessment | Innodata, Research Scientist – Speech & Audio |
| Evaluation infrastructure | ServiceNow, Senior Director – AI Evaluation Platform |
| AI application release criteria | Abacus Insights, Senior AI Systems Quality Engineer |
| Authorized cyber red teaming | OpenAI, Red Team Specialist – Cyber |
| Autonomous-system release evaluation | Atoms, Technical Lead – Autonomy Evaluation |
| Model assessment methodology | Amazon, Senior Applied Scientist – Frontier AI Assets Assessments |
Limits and next checks
This corpus cannot estimate U.S.-wide demand, hiring growth, wages, vacancy-to-hire conversion, or the chance that a novice will qualify. The sampling frame was not exhaustive or random. AI product and model categories changed quickly; pages can close or change. Employer and ATS concentration, seniority skew, title ambiguity, the combined early-career/intern classification, and cross-domain differences limit generalization. A missing field in a posting remains unknown, not zero. The 11 autonomous-system roles must not be pooled into claims about general LLM application practice without showing the split.
The separate trend article uses a different evidence stream. It is not a source of vacancy counts in this study.
Continue learning
Apply the evidence from this report through MTF Institute's Professional Certificate in AI Evaluation. The programme turns the identified capabilities into structured theory, guided AI practice and reusable workplace artifacts.