Online professional certificate
Professional Certificate in AI Evaluation
Practical AI Quality Assurance and Authorized Red Teaming. Learn to test AI applications with valid cases, human and automated scoring, controlled synthetic red teaming, and evidence that product owners can act on.
- Format
- Online, self-paced
- Study time
- Up to 1 month
- Curriculum
- 4 modules · 20 lessons
- Language
- English
Practical capability
Build the evidence behind an AI quality decision
Move from a vague concern to a bounded test, then turn the result into a clear next action for the people who own the system.
Name the AI application, intended use, authority boundary and question the evidence must answer.
Select a defensible case population, record provenance and protect unseen tests.
Write observable rubrics, calibrate reviewers and check automated graders against human judgment.
Track versions, tool states, permissions and case outputs so another evaluator can inspect a result.
Separate an observed failure from a weak test item, a tool problem or an unconfirmed cause.
Explain uncertainty, verify fixes and hand a bounded release or hold recommendation to the decision owner.
Who this course is for
For professionals who need dependable AI evidence
The course starts with practical workplace documentation and builds the AI-specific methods step by step.
The operating cycle
From question to owner action
The sequence follows the Model Role SOP: open and authorize the work, build valid evidence, run and score it, investigate material findings, then hand a decision record to the right owner.
Curriculum
Four modules. Twenty applied lessons.
Define the Evaluation
A useful evaluation begins before the first test. You will clarify the system, intended users, decision, and people authorized to approve methods, data and release. You will turn a vague quality concern into a focused request, set observable criteria, choose methods and select a case population. The resulting plan shows what the test can answer and where it needs more evidence.
01 Who Decides What in an AI Evaluation
Identify who performs, approves and decides each part of a bounded AI evaluation.
Five practical steps
- Identify the evaluated system and decision
- List owners and needed approvals
- Separate own actions from owner decisions
- Mark escalation triggers
- Review the map with a named owner
Primary deliverable: role responsibility and decision-rights map.
02 Turn a Quality Concern into an Evaluation Request
Convert an ambiguous concern into a scoped request with the inputs needed before testing.
Five practical steps
- State the trigger and affected workflow
- Name the system and candidate version
- Write the decision question
- Ask for missing permissions and evidence
- Record exclusions and owner response
Primary deliverable: evaluation intake record.
03 Set Quality Criteria Before Testing
Define observable acceptance criteria and evidence limits before reviewing results.
Five practical steps
- Choose user-relevant behaviors
- Describe expected and unacceptable outcomes
- Select measures and denominators
- Agree decision thresholds before results
- Record owner approval and evidence limits
Primary deliverable: quality-criteria and decision-threshold register.
04 Choose Methods for the AI System
Choose and justify methods that fit the application and decision.
Five practical steps
- Describe the system and decision
- Match each question to an evidence method
- State inputs and permissions
- Explain each method’s blind spots
- Select a proportionate combination
Primary deliverable: assessment-method selection brief.
05 Sample the Cases That Matter
Select a valid case population and explain what the resulting sample can support.
Five practical steps
- Define the population
- Name important strata and failure modes
- Choose a defensible selection rule
- Record exclusions and sample size
- State what the sample cannot estimate
Primary deliverable: case-population and sampling plan.
Build Valid Test Evidence
Good scores depend on good test material. You will build versioned cases and rubrics, inspect weak items, protect holdouts and prepare a reproducible run. Each step leaves evidence another authorized reviewer can inspect. You will keep case provenance, scoring rules, tool settings and data rights attached to the results.
06 Build a Versioned Case Set
Create traceable cases that can be rerun and audited.
Five practical steps
- Draft cases from approved sources
- Set expected behavior and case class
- Record provenance and permissions
- Assign stable version and change log
- Review coverage with an owner
Primary deliverable: versioned evaluation case register.
07 Write Rubrics for Consistent Judgment
Apply a rubric so another reviewer can inspect and reproduce a score.
Five practical steps
- Define one dimension per judgment
- Write observable score anchors
- Pilot ambiguous cases
- Adjudicate disagreement
- Version the guidance and rescore affected cases
Primary deliverable: rubric and adjudication guide.
08 Check Whether Evaluation Items Are Valid
Decide whether a test item is valid enough to support a result.
Five practical steps
- Inspect the task and expected answer
- Check hidden or automated tests
- Reproduce the alleged item defect
- Ask a domain reviewer to adjudicate
- Record include, revise or exclude impact
Primary deliverable: benchmark-item validity audit log.
09 Protect Holdouts and Test Data
Protect test material and preserve the credibility of held-out results.
Five practical steps
- Identify each data owner and permitted use
- Classify holdout exposure
- Set access and storage rules
- Plan contamination checks
- Define stop and escalation for exposure
Primary deliverable: test-data custody and holdout register.
10 Run Reproducible Tests in Approved Environments
Run and record an authorized test another worker can reproduce.
Five practical steps
- Confirm scope and accounts
- Record version and configuration
- Pilot cases and grader health
- Run approved cases with minimum necessary logging
- Store trace references and stop events
Primary deliverable: reproducible evaluation run manifest.
Measure and Diagnose Behavior
Once results arrive, the evaluator must decide what the evidence means. You will calibrate human review, check automated graders, calculate measures and investigate surprising behavior. You will distinguish system defects from test or tool faults, then present quality and uncertainty in a scorecard a product owner can use.
11 Calibrate Human Reviewers
Resolve scoring disagreement without hiding uncertainty.
Five practical steps
- Collect independent scores
- Compare disagreements by case class
- Request evidence-based explanations
- Adjudicate or revise rubric
- Record rescore and remaining uncertainty
Primary deliverable: reviewer-calibration and disagreement record.
12 Validate Automated Graders
Decide where automated scores are reliable enough to support review.
Five practical steps
- Define where a grader may help
- Select human-reviewed comparison cases
- Calculate agreement by slice
- Inspect disagreement and failure modes
- Record permitted use and limits
Primary deliverable: automated-grader validation report.
13 Analyze Results and Uncertainty
Calculate and explain results without overstating what the sample proves.
Five practical steps
- Validate case inclusion and denominator
- Calculate overall and slice results
- Compare with approved criteria or prior baseline
- Examine uncertainty and disagreement
- State a bounded interpretation
Primary deliverable: evaluation analysis and uncertainty worksheet.
14 Diagnose AI System Failures
Classify a failure and hand it to the right owner with reproducible evidence.
Five practical steps
- Reproduce under approved conditions
- Separate observation from cause hypothesis
- Inspect trace and reference source
- Classify fault and local severity
- Route evidence to the accountable owner
Primary deliverable: AI-system failure triage record.
15 Build a Decision-Ready Quality Scorecard
Present quality evidence a decision owner can understand and challenge.
Five practical steps
- Select decision-relevant measures
- Show counts and meaningful slices
- Link major findings and grader limits
- Explain uncertainty and open risks
- Request an owner action
Primary deliverable: AI-application quality scorecard.
Turn Findings into Responsible Action
Findings matter when they lead to accountable action. You will route defects, retest changes, prepare a recommendation and keep escalation moving at the cadence your team needs. You will also practice an approved synthetic safety test and identify cases that require physical-system specialists. These boundaries help you contribute evidence while the right owners retain access and release decisions.
16 Verify Fixes and Regression Results
Verify observed improvement rather than closing a finding on a promise.
Five practical steps
- Confirm owner and proposed fix
- Record new system version
- Retest original failure
- Check relevant regression cases
- Keep unresolved issues and limits visible
Primary deliverable: remediation and regression evidence record.
17 Recommend a Release Decision with Evidence
Give the decision owner a defensible recommendation without taking their approval right.
Five practical steps
- Restate scope and pre-agreed criteria
- Summarize observed evidence
- Surface open issues and uncertainty
- Recommend an option with rationale
- Route to the authorized decision owner
Primary deliverable: release-evidence recommendation memo.
18 Manage Work Rhythm and Escalation
Coordinate evaluation work and escalate material findings at the rhythm the employer requires.
Five practical steps
- Read the local release rhythm and case volume
- Assign check and calibration triggers
- Set periodic method-review interval
- Define urgent stop and handoff
- Write an actionable status message
Primary deliverable: evaluation operations and escalation handoff plan.
19 Conduct an Authorized Synthetic Red-Team Evaluation
Run a bounded approved synthetic test and document what happened and when to stop.
Five practical steps
- Verify written target and environment scope
- Design benign synthetic cases inside allowed bounds
- Obtain named-owner approval of case plan
- Run only approved cases in the mock sandbox
- Record observations, stop response and handoff
Primary deliverable: authorized synthetic adversarial test plan, run and findings record.
20 Recognize Autonomous-System Evaluation Boundaries
Recognize when autonomous-system evaluation needs domain-specific methods and owners.
Five practical steps
- Describe the new system and real-world stakes
- Compare evidence objects and environments
- Identify methods that do not transfer
- Name specialist owners and approvals
- Write a limited handoff recommendation
Primary deliverable: autonomous-system applicability and specialist-handoff memo.
Applied capstone
Make an evidence-based AI pilot decision.
Use the relevant course methods to interpret an approved synthetic packet and give the product owner one clear recommendation.
The situation
An internal expense assistant gives approval language after its status tool times out during a supervised staging check.
Your task
Use the supplied policy, pre-agreed criteria, synthetic outputs and trace excerpts to advise the product owner whether a controlled pilot is ready.
The people behind MTF
Meet MTF faculty and the learner community.
Explore the professional backgrounds of MTF faculty and learn more about the international community studying with the Institute.
Enrollment
Enroll in Professional Certificate in AI Evaluation
One-time course price: €10, including applicable taxes. Payment is processed securely by Stripe. No card details are stored on the MTF Institute website.
You will receive an email with access to the course. If you have any difficulties, please write to welcome@gtf.pt.
Questions and details
Frequently asked questions
Open the sections that matter to you, including delivery format, AI-supported practice and the evidence used to design the curriculum.
Who is this AI evaluation course for?
It is designed for software QA, data and ML, product, safety and technical-risk professionals who need a practical way to evaluate AI applications. It begins with scope, evidence and authority boundaries before asking learners to interpret results or recommend action.
How does the course work?
The course is online and self-paced, with four modules, 20 applied lessons and one capstone. Each lesson explains a method, works through a connected case and gives you a reusable professional artifact with a completed example. Study time is up to one month at a pace that fits your practice.
How is AI used in the exercises?
Each lesson includes an AI-supported draft and a separate critic step. You use only approved case facts and tools, check the complete output against the taught method, and keep scoring, access and release decisions with the named human owners.
What evidence supports the curriculum?
The curriculum follows an MTF Institute study of 107 directly verified current U.S. AI evaluation and adjacent vacancies, plus a separate review of recent changes in evaluation work. The vacancy sample is purposive, so its counts describe the selected postings rather than national prevalence. The research report is also archived with a Zenodo DOI.
What practical work will I complete?
You will practise evaluation requests, quality criteria, sampling plans, versioned case sets, rubrics, data and holdout controls, run manifests, human and automated score checks, failure triage, scorecards, remediation evidence and bounded recommendations.
What does the applied capstone involve?
You will assess a supplied synthetic expense-assistant test packet against pre-agreed criteria after a tool-unavailable response raises a quality concern. The principal deliverable is one evaluation decision brief that reports results and limits, recommends a next step and names the pilot decision owner.
What certificate and access will I receive?
Enrollment gives you access to the MTF learning platform and the MTF Institute certificate activity for Professional Certificate in AI Evaluation.