Evaluation of Agentic Systems: testing and quality assessment of AI agents


Information about training in this course.

Course objective: to provide the practical skills required to evaluate, test and monitor AI agents — from tracing and observability, through building evaluation suites (code-based checks, LLM-as-a-judge, human review) and designing verifiable tasks and test environments, to public benchmarks, agent red teaming and production monitoring.

Training takes place in the centre of Tallinn (Tartu mnt. 18, Tallinn) and/or online. All educational materials are included in the course price. A laptop is provided for the duration of the training if needed.


Target group:

This course is for you if you:

  • are a QA / test automation engineer and want to add the fastest-growing specialisation in quality assurance — evaluation of AI agents;
  • are a software engineer building (or planning to build) LLM and agent features and need to know how to measure their quality;
  • are a data / ML specialist moving towards applied AI evaluation and LLMOps;
  • are an IT professional in Estonia or the EU looking for remote-friendly work — AI evaluation platforms hire worldwide and value the skills this course teaches;
  • are a product manager or analyst of an AI product and need a systematic way to judge whether your agents actually work;
  • already work on AI evaluation platforms (annotation, response rating) and want to move up to engineering-level agent evaluation.

Key skills you will gain on this course:

  • Understand agent architectures: planning loops, tools, memory, MCP
  • Use coding agents (Claude Code, Cursor) as everyday working instruments
  • Instrument agents with tracing (OpenTelemetry, Arize Phoenix, LangSmith)
  • Build golden datasets and run structured experiments
  • Design and validate LLM-as-a-judge evaluators
  • Evaluate agent trajectories, tool choice and convergence
  • Design verifiable tasks and test environments for agents
  • Run agents on public benchmarks (GAIA, SWE-bench) and interpret results
  • Red-team agents: unsafe shortcuts, tool-use abuse, prompt injection
  • Monitor agents in production: online evals, regressions, cost and latency

Requirements for students:

  • confident PC user
  • English sufficient to follow technical training (approximately B1/B2)
  • basic Python skills are strongly recommended (no formal prerequisite: practical labs come with ready-made templates and can be completed with the help of AI coding assistants)
  • It is desirable to have your own laptop (Windows / Mac / Linux, 8 GB RAM+); a laptop will be provided for the duration of the training if needed.

Learning outcome:

Those who complete this course:

  • explain the architecture and typical failure modes of agentic AI systems
  • instrument and trace agents, read and analyse agent trajectories
  • build evaluation suites combining code-based checks, LLM-as-a-judge and human review
  • design verifiable tasks and test environments with acceptance tests
  • run and interpret benchmark evaluations and maintain regression suites
  • conduct basic agent red teaming and produce quality and vulnerability reports

Training methods:

The total course volume is 80 academic hours, of which 40 academic hours are classroom contact hours with the instructor (lectures, practical work, seminars) and 40 academic hours are independent work.

Evaluation criteria for learning outcomes:

Learning outcomes are assessed based on independently completed practical work and a final project (a complete evaluation harness for an AI agent).

Evaluation methods:

Upon successful completion, practical and homework assignments receive a "pass" grade.

Course completion conditions:

To successfully complete the course and receive a certificate, it is necessary to achieve a "pass" grade on 75% of the homework assignments and defend the final project.

Additional information:

Training programme group: 0613 - Software and applications development and analysis (0613 - Tarkvara ja rakenduste arendus ning analüüs)
Basic rules for training organisation (in Estonian)
Basic rules for ensuring the quality of the educational process (in Estonian)

Course program

Module Main topics Volume
1. Agentic systems and the evaluator's toolkit
  • Agent anatomy: planning loop (ReAct), router, tools, memory; MCP
  • Chatbots vs RAG vs agents; where agents fail: hallucinated actions, loops, unsafe shortcuts
  • Why agent evaluation differs from software testing and from LLM evaluation
  • Coding agents (Claude Code, Cursor) as everyday instruments of the evaluator
  • 4 ac/h
    2. Observability and tracing
  • Traces and spans; OpenTelemetry instrumentation
  • Arize Phoenix and LangSmith / Langfuse — reading agent runs
  • Token, cost and latency monitoring
  • Practice: instrumenting and inspecting a working agent
  • 4 ac/h
    3. Evaluators I: datasets and judges
  • Three evaluator types: code-based, LLM-as-a-judge, human review
  • Golden datasets and structured experiments
  • Designing and validating judge prompts
  • Practice: building an evaluation suite for a demo agent
  • 4 ac/h
    4. Evaluators II: trajectories and failure analysis
  • Router and tool-choice evaluations
  • Trajectory evaluation and convergence scoring
  • Failure-mode taxonomy and error analysis
  • Practice: from raw traces to a prioritised defect report
  • 4 ac/h
    5. Environments and task design
  • Realistic test environments: codebase, infrastructure, context ("virtual company")
  • Designing verifiable tasks from intermediate environment states
  • Acceptance tests: accept every valid solution, reject invalid ones
  • Terminal conditions, reproducibility and fairness of evaluation protocols
  • Practice: building and reviewing an evaluation task end-to-end
  • 8 ac/h
    6. Benchmarks
  • Public agent benchmarks: GAIA, SWE-bench Verified, tau-bench
  • Reliability metrics: pass@k and pass^k; non-determinism
  • Custom, domain-specific benchmarks — when and how to build them
  • Practice: running an agent on a GAIA subset and interpreting the score
  • 4 ac/h
    7. Agent safety and red teaming
  • Unsafe shortcuts: benign goal, unsafe path; scope creep
  • Tool-use abuse and prompt injection through tools
  • Guardrails and supervisor models; EU AI Act perspective
  • Practice: red-teaming a demo agent and reporting vulnerabilities
  • 4 ac/h
    8. Production monitoring
  • Online evaluations and user-feedback loops
  • Regression suites and CI integration
  • Cost / latency budgets; quality dashboards
  • Practice: wiring evals into a CI pipeline
  • 4 ac/h
    9. Final project
  • A complete evaluation harness for an AI agent
  • Rubric + automated evals + red-team pass + report
  • Defence of the final project
  • 4 ac/h

    Course information

    Time of conduct:

    29.09.2026 - 24.11.2026
    27.10.2026 - 22.12.2026
    24.11.2026 - 19.01.2027


    Timetable:

    Tue, Thu 17:45–21:00 (evening group).


    Apply → We'll reply within 1 business day

    Course length:

    8 weeks



    Format and place of conduct:

    Address: Tartu mnt. 18, Tallinn / Online.
    Gamma Intelligence Training Centre
    The course is conducted in a classroom format (up to 12 people in class) and/or online (Zoom / Microsoft Teams); total group size up to 18 people.

    Training language: English

    Price: 2500 EUR (VAT 24% included)

    Total course volume: 80 ac/h (40 classroom + 40 independent)
    Format: lectures + practical work + independent work.


    Tutors

    Nikolay Zubrilov

    Qualification: Python developer and Data Scientist; AI/LLM developer at Mentastic; owner of Dataocean Analytics OÜ; commercial experience in backend development, data analysis and AI/LLM.

    Specialisation: Python, AI/LLM development, data analysis and data science (Pandas, NumPy, Scikit-Learn), SQL (PostgreSQL / MySQL), REST API (FastAPI / Flask), automation.

    Teaching experience: instructor on Python, data analysis and LLM testing courses at Gamma Intelligence Training Centre.

    Education: Master's degree — computer and systems engineering.

    Review the CV