# Professional Certificate in AI Evaluation

Canonical URL: https://mtfinstitute.com/programs/ai-evaluation/
Official publisher: MTF Institute of Management, Technology and Finance
Language: English
Topics: AI Evaluation, AI Quality Assurance, Model Evaluation, AI Application Testing, Benchmark Validity, Human and Automated Scoring, Authorized Red Teaming, Evaluation Evidence, Release Decisions

> Practical AI Quality Assurance and Authorized Red Teaming. Learn to design valid AI tests, judge results, investigate failures and give product owners evidence they can use.

## Program facts

- Format: Online, self-paced
- Recommended duration: Up to 1 month
- Study time: Flexible self-paced study
- Tuition: €10
- Credential: Certificate of completion: Professional Certificate in AI Evaluation
- Enrollment: https://edu.gtf.pt/course/view.php?id=109


## Professional Certificate in AI Evaluation

Practical AI Quality Assurance and Authorized Red Teaming. Learn to test AI applications with valid cases, human and automated scoring, controlled synthetic red teaming, and evidence that product owners can act on.

## Who this course is for

The course starts with practical workplace documentation and builds the AI-specific methods step by step.
- Software QA professionals: Extend familiar test practice to variable AI outputs, valid cases and owner-ready findings.
- Data and ML practitioners: Connect metrics and benchmark results to application behavior, uncertainty and reproducible records.
- Product and AI operations teams: Use evaluation evidence to manage changes, priorities, handoffs and release questions.
- Trust and safety specialists: Plan bounded synthetic checks and escalate findings within approved access and decision routes.

## What you will be able to do

Move from a vague concern to a bounded test, then turn the result into a clear next action for the people who own the system.
- **Scope the decision:** Name the AI application, intended use, authority boundary and question the evidence must answer.
- **Design valid cases:** Select a defensible case population, record provenance and protect unseen tests.
- **Set reliable criteria:** Write observable rubrics, calibrate reviewers and check automated graders against human judgment.
- **Run reproducible checks:** Track versions, tool states, permissions and case outputs so another evaluator can inspect a result.
- **Diagnose behavior:** Separate an observed failure from a weak test item, a tool problem or an unconfirmed cause.
- **Recommend responsible action:** Explain uncertainty, verify fixes and hand a bounded release or hold recommendation to the decision owner.

## Curriculum

Four connected modules contain 20 applied lessons and a separate decision-brief capstone.

### Module 1: Define the Evaluation

A useful evaluation begins before the first test. You will clarify the system, intended users, decision, and people authorized to approve methods, data and release. You will turn a vague quality concern into a focused request, set observable criteria, choose methods and select a case population. The resulting plan shows what the test can answer and where it needs more evidence.

1. **Who Decides What in an AI Evaluation** — Identify who performs, approves and decides each part of a bounded AI evaluation. You will practise: Identify the evaluated system and decision; List owners and needed approvals; Separate own actions from owner decisions; Mark escalation triggers; Review the map with a named owner. Primary deliverable: role responsibility and decision-rights map.
2. **Turn a Quality Concern into an Evaluation Request** — Convert an ambiguous concern into a scoped request with the inputs needed before testing. You will practise: State the trigger and affected workflow; Name the system and candidate version; Write the decision question; Ask for missing permissions and evidence; Record exclusions and owner response. Primary deliverable: evaluation intake record.
3. **Set Quality Criteria Before Testing** — Define observable acceptance criteria and evidence limits before reviewing results. You will practise: Choose user-relevant behaviors; Describe expected and unacceptable outcomes; Select measures and denominators; Agree decision thresholds before results; Record owner approval and evidence limits. Primary deliverable: quality-criteria and decision-threshold register.
4. **Choose Methods for the AI System** — Choose and justify methods that fit the application and decision. You will practise: Describe the system and decision; Match each question to an evidence method; State inputs and permissions; Explain each method’s blind spots; Select a proportionate combination. Primary deliverable: assessment-method selection brief.
5. **Sample the Cases That Matter** — Select a valid case population and explain what the resulting sample can support. You will practise: Define the population; Name important strata and failure modes; Choose a defensible selection rule; Record exclusions and sample size; State what the sample cannot estimate. Primary deliverable: case-population and sampling plan.

### Module 2: Build Valid Test Evidence

Good scores depend on good test material. You will build versioned cases and rubrics, inspect weak items, protect holdouts and prepare a reproducible run. Each step leaves evidence another authorized reviewer can inspect. You will keep case provenance, scoring rules, tool settings and data rights attached to the results.

6. **Build a Versioned Case Set** — Create traceable cases that can be rerun and audited. You will practise: Draft cases from approved sources; Set expected behavior and case class; Record provenance and permissions; Assign stable version and change log; Review coverage with an owner. Primary deliverable: versioned evaluation case register.
7. **Write Rubrics for Consistent Judgment** — Apply a rubric so another reviewer can inspect and reproduce a score. You will practise: Define one dimension per judgment; Write observable score anchors; Pilot ambiguous cases; Adjudicate disagreement; Version the guidance and rescore affected cases. Primary deliverable: rubric and adjudication guide.
8. **Check Whether Evaluation Items Are Valid** — Decide whether a test item is valid enough to support a result. You will practise: Inspect the task and expected answer; Check hidden or automated tests; Reproduce the alleged item defect; Ask a domain reviewer to adjudicate; Record include, revise or exclude impact. Primary deliverable: benchmark-item validity audit log.
9. **Protect Holdouts and Test Data** — Protect test material and preserve the credibility of held-out results. You will practise: Identify each data owner and permitted use; Classify holdout exposure; Set access and storage rules; Plan contamination checks; Define stop and escalation for exposure. Primary deliverable: test-data custody and holdout register.
10. **Run Reproducible Tests in Approved Environments** — Run and record an authorized test another worker can reproduce. You will practise: Confirm scope and accounts; Record version and configuration; Pilot cases and grader health; Run approved cases with minimum necessary logging; Store trace references and stop events. Primary deliverable: reproducible evaluation run manifest.

### Module 3: Measure and Diagnose Behavior

Once results arrive, the evaluator must decide what the evidence means. You will calibrate human review, check automated graders, calculate measures and investigate surprising behavior. You will distinguish system defects from test or tool faults, then present quality and uncertainty in a scorecard a product owner can use.

11. **Calibrate Human Reviewers** — Resolve scoring disagreement without hiding uncertainty. You will practise: Collect independent scores; Compare disagreements by case class; Request evidence-based explanations; Adjudicate or revise rubric; Record rescore and remaining uncertainty. Primary deliverable: reviewer-calibration and disagreement record.
12. **Validate Automated Graders** — Decide where automated scores are reliable enough to support review. You will practise: Define where a grader may help; Select human-reviewed comparison cases; Calculate agreement by slice; Inspect disagreement and failure modes; Record permitted use and limits. Primary deliverable: automated-grader validation report.
13. **Analyze Results and Uncertainty** — Calculate and explain results without overstating what the sample proves. You will practise: Validate case inclusion and denominator; Calculate overall and slice results; Compare with approved criteria or prior baseline; Examine uncertainty and disagreement; State a bounded interpretation. Primary deliverable: evaluation analysis and uncertainty worksheet.
14. **Diagnose AI System Failures** — Classify a failure and hand it to the right owner with reproducible evidence. You will practise: Reproduce under approved conditions; Separate observation from cause hypothesis; Inspect trace and reference source; Classify fault and local severity; Route evidence to the accountable owner. Primary deliverable: AI-system failure triage record.
15. **Build a Decision-Ready Quality Scorecard** — Present quality evidence a decision owner can understand and challenge. You will practise: Select decision-relevant measures; Show counts and meaningful slices; Link major findings and grader limits; Explain uncertainty and open risks; Request an owner action. Primary deliverable: AI-application quality scorecard.

### Module 4: Turn Findings into Responsible Action

Findings matter when they lead to accountable action. You will route defects, retest changes, prepare a recommendation and keep escalation moving at the cadence your team needs. You will also practice an approved synthetic safety test and identify cases that require physical-system specialists. These boundaries help you contribute evidence while the right owners retain access and release decisions.

16. **Verify Fixes and Regression Results** — Verify observed improvement rather than closing a finding on a promise. You will practise: Confirm owner and proposed fix; Record new system version; Retest original failure; Check relevant regression cases; Keep unresolved issues and limits visible. Primary deliverable: remediation and regression evidence record.
17. **Recommend a Release Decision with Evidence** — Give the decision owner a defensible recommendation without taking their approval right. You will practise: Restate scope and pre-agreed criteria; Summarize observed evidence; Surface open issues and uncertainty; Recommend an option with rationale; Route to the authorized decision owner. Primary deliverable: release-evidence recommendation memo.
18. **Manage Work Rhythm and Escalation** — Coordinate evaluation work and escalate material findings at the rhythm the employer requires. You will practise: Read the local release rhythm and case volume; Assign check and calibration triggers; Set periodic method-review interval; Define urgent stop and handoff; Write an actionable status message. Primary deliverable: evaluation operations and escalation handoff plan.
19. **Conduct an Authorized Synthetic Red-Team Evaluation** — Run a bounded approved synthetic test and document what happened and when to stop. You will practise: Verify written target and environment scope; Design benign synthetic cases inside allowed bounds; Obtain named-owner approval of case plan; Run only approved cases in the mock sandbox; Record observations, stop response and handoff. Primary deliverable: authorized synthetic adversarial test plan, run and findings record.
20. **Recognize Autonomous-System Evaluation Boundaries** — Recognize when autonomous-system evaluation needs domain-specific methods and owners. You will practise: Describe the new system and real-world stakes; Compare evidence objects and environments; Identify methods that do not transfer; Name specialist owners and approvals; Write a limited handoff recommendation. Primary deliverable: autonomous-system applicability and specialist-handoff memo.

## How the course works

Study online at your own pace over up to one month. Each lesson connects a method to a worked case, a blank professional template, a completed example and two guided AI prompts. The first prompt supports a bounded draft; the second asks you to critique the complete output against the taught method and the approved facts. You retain the decision and verify the result.

## Certificate

Completing the required learning activities gives access to the MTF Institute certificate activity for Professional Certificate in AI Evaluation.

## Evidence behind the course

The curriculum draws on 107 directly verified current U.S. AI evaluation and adjacent vacancies and a separate review of recent changes in the field. The vacancy sample is purposive and describes the selected postings, rather than national prevalence. Read the [vacancy research report](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-107-us-vacancies-2026/) and [current-changes article](https://mtfinstitute.com/insights/ai-evaluation-red-teaming-2026-recent-changes/). The research report is archived at [Zenodo](https://zenodo.org/records/23145132).

## Start the course

[ENROLL NOW](https://edu.gtf.pt/course/view.php?id=109)

## Frequently asked questions

### Who is this AI evaluation course for?

It is designed for software QA, data and ML, product, safety and technical-risk professionals who need a practical way to evaluate AI applications. It begins with scope, evidence and authority boundaries before asking learners to interpret results or recommend action.

### How does the course work?

The course is online and self-paced, with four modules, 20 applied lessons and one capstone. Each lesson explains a method, works through a connected case and gives you a reusable professional artifact with a completed example. Study time is up to one month at a pace that fits your practice.

### How is AI used in the exercises?

Each lesson includes an AI-supported draft and a separate critic step. You use only approved case facts and tools, check the complete output against the taught method, and keep scoring, access and release decisions with the named human owners.

### What evidence supports the curriculum?

The curriculum follows an MTF Institute study of 107 directly verified current U.S. AI evaluation and adjacent vacancies, plus a separate review of recent changes in evaluation work. The vacancy sample is purposive, so its counts describe the selected postings rather than national prevalence. The research report is also archived with a Zenodo DOI.

### What practical work will I complete?

You will practise evaluation requests, quality criteria, sampling plans, versioned case sets, rubrics, data and holdout controls, run manifests, human and automated score checks, failure triage, scorecards, remediation evidence and bounded recommendations.

### What does the applied capstone involve?

You will assess a supplied synthetic expense-assistant test packet against pre-agreed criteria after a tool-unavailable response raises a quality concern. The principal deliverable is one evaluation decision brief that reports results and limits, recommends a next step and names the pilot decision owner.

### What certificate and access will I receive?

Enrollment gives you access to the MTF learning platform and the MTF Institute certificate activity for Professional Certificate in AI Evaluation.

## Professional education notice

Professional courses and certificates are taught under the terms of paragraph 3 of article 3 of Decree-Law No. 474/2010, published on July 8th by the Portuguese Ministry of Labour and Social Solidarity. The professional programs are related to professional / business education and are provided without official recognition (certificates are provided at a professional level and not academic degrees or diplomas and do not confer academic credits).

## Citation guidance

When quoting or summarizing this program, cite the canonical HTML page: https://mtfinstitute.com/programs/ai-evaluation/
