Evaluation of Agentic Systems: testing and quality assessment of AI agents
Information about training in this course.
Course objective:
to provide the practical skills required to evaluate, test and monitor AI agents — from tracing and observability,
through building evaluation suites (code-based checks, LLM-as-a-judge, human review) and designing verifiable tasks
and test environments, to public benchmarks, agent red teaming and production monitoring.
Training takes place in the centre of Tallinn (Tartu mnt. 18, Tallinn) and/or online.
All educational materials are included in the course price.
A laptop is provided for the duration of the training if needed.
Target group:
This course is for you if you:
- are a QA / test automation engineer and want to add the fastest-growing specialisation in quality assurance — evaluation of AI agents;
- are a software engineer building (or planning to build) LLM and agent features and need to know how to measure their quality;
- are a data / ML specialist moving towards applied AI evaluation and LLMOps;
- are an IT professional in Estonia or the EU looking for remote-friendly work — AI evaluation platforms hire worldwide and value the skills this course teaches;
- are a product manager or analyst of an AI product and need a systematic way to judge whether your agents actually work;
- already work on AI evaluation platforms (annotation, response rating) and want to move up to engineering-level agent evaluation.
Key skills you will gain on this course:
- Understand agent architectures: planning loops, tools, memory, MCP
- Use coding agents (Claude Code, Cursor) as everyday working instruments
- Instrument agents with tracing (OpenTelemetry, Arize Phoenix, LangSmith)
- Build golden datasets and run structured experiments
- Design and validate LLM-as-a-judge evaluators
- Evaluate agent trajectories, tool choice and convergence
- Design verifiable tasks and test environments for agents
- Run agents on public benchmarks (GAIA, SWE-bench) and interpret results
- Red-team agents: unsafe shortcuts, tool-use abuse, prompt injection
- Monitor agents in production: online evals, regressions, cost and latency
Requirements for students:
- confident PC user
- English sufficient to follow technical training (approximately B1/B2)
- basic Python skills are strongly recommended (no formal prerequisite: practical labs come with ready-made templates and can be completed with the help of AI coding assistants)
- It is desirable to have your own laptop (Windows / Mac / Linux, 8 GB RAM+); a laptop will be provided for the duration of the training if needed.
Learning outcome:
Those who complete this course:
- explain the architecture and typical failure modes of agentic AI systems
- instrument and trace agents, read and analyse agent trajectories
- build evaluation suites combining code-based checks, LLM-as-a-judge and human review
- design verifiable tasks and test environments with acceptance tests
- run and interpret benchmark evaluations and maintain regression suites
- conduct basic agent red teaming and produce quality and vulnerability reports
Training methods:
The total course volume is 80 academic hours, of which 40 academic hours are classroom contact hours with the instructor (lectures, practical work, seminars) and 40 academic hours are independent work.
Evaluation criteria for learning outcomes:
Learning outcomes are assessed based on independently completed practical work and a final project (a complete evaluation harness for an AI agent).
Evaluation methods:
Upon successful completion, practical and homework assignments receive a "pass" grade.
Course completion conditions:
To successfully complete the course and receive a certificate, it is necessary to achieve a "pass" grade on 75% of the homework assignments and defend the final project.
Additional information:
Training programme group: 0613 - Software and applications development and analysis (0613 - Tarkvara ja rakenduste arendus ning analüüs)
Basic rules for training organisation (in Estonian)
Basic rules for ensuring the quality of the educational process (in Estonian)
Course program
| Module | Main topics | Volume |
| 1. Agentic systems and the evaluator's toolkit |
|
4 ac/h |
| 2. Observability and tracing |
|
4 ac/h |
| 3. Evaluators I: datasets and judges |
|
4 ac/h |
| 4. Evaluators II: trajectories and failure analysis |
|
4 ac/h |
| 5. Environments and task design |
|
8 ac/h |
| 6. Benchmarks |
|
4 ac/h |
| 7. Agent safety and red teaming |
|
4 ac/h |
| 8. Production monitoring |
|
4 ac/h |
| 9. Final project |
|
4 ac/h |
Course information
Time of conduct:
29.09.2026 - 24.11.2026
27.10.2026 - 22.12.2026
24.11.2026 - 19.01.2027
Timetable:
Tue, Thu 17:45–21:00 (evening group).
Apply → We'll reply within 1 business day
Course length:
8 weeks
Format and place of conduct:
Address: Tartu mnt. 18, Tallinn / Online.

The course is conducted in a classroom format (up to 12 people in class) and/or online (Zoom / Microsoft Teams); total group size up to 18 people.
Training language: English
Price: 2500 EUR (VAT 24% included)
Total course volume: 80 ac/h (40 classroom + 40 independent)
Format: lectures + practical work + independent work.
Tutors
Nikolay Zubrilov
Qualification: Python developer and Data Scientist; AI/LLM developer at Mentastic; owner of Dataocean Analytics OÜ; commercial experience in backend development, data analysis and AI/LLM.Specialisation: Python, AI/LLM development, data analysis and data science (Pandas, NumPy, Scikit-Learn), SQL (PostgreSQL / MySQL), REST API (FastAPI / Flask), automation.
Teaching experience: instructor on Python, data analysis and LLM testing courses at Gamma Intelligence Training Centre.
Education: Master's degree — computer and systems engineering.