LLM Tester: quality assessment of AI systems (with Python basics)


Information about training in this course.

Course objective: to provide the basic theoretical knowledge and practical skills required to test and assess the quality of systems built on large language models (LLMs) — from manual, rubric-based response evaluation to automated eval pipelines, testing of RAG systems and vulnerability discovery (red teaming). The course includes a Python foundations block — the main working tool of this profession.

Training takes place in the centre of Tallinn (Tartu mnt. 18) and/or online. All educational materials are included in the course price. A laptop is provided for the duration of the training if needed.


Target group:

This course is for you if you:

  • are a beginner or practising software tester and want to add the most in-demand specialisation to your profile — testing of AI systems;
  • are a specialist from another field and are considering entering the AI industry through quality evaluation of models — a direction with a low entry barrier;
  • are a linguist, editor, translator or analyst and want to apply attention to detail and strong language skills to AI work (response evaluation, annotation, rubrics);
  • already work on AI evaluation platforms (annotation, response comparison) and want to move to more complex, better-paid tasks with automation;
  • are a developer or engineer and need to know how to verify the quality and safety of LLM features in your products;
  • are an international professional in Estonia or the EU — AI evaluation platforms hire worldwide, and English-speaking QA roles exist in local product companies.

Key skills you will gain on this course:

  • Write Python programs: data types, loops, functions, working with files
  • Understand how LLM systems work: tokens, context, non-deterministic outputs
  • Distinguish AI defect classes: hallucinations, bias, prompt injections
  • Craft prompts and design response evaluation rubrics
  • Run manual evaluation: side-by-side comparison, ranking, recording failure modes
  • Apply quality metrics and the LLM-as-a-judge method
  • Build automated eval pipelines (promptfoo, DeepEval, pytest)
  • Test RAG systems and AI agents (tool use, tracing)
  • Perform basic red teaming: prompt injection, jailbreak tests
  • Write defect and vulnerability reports for AI systems

Requirements for students:

  • confident PC user; prior programming experience is NOT required (the course includes a Python foundations block)
  • English sufficient to follow technical training (approximately B1/B2); international AI evaluation platforms typically expect B2 or higher
  • It is desirable to have your own laptop (Windows / Mac, 8 GB RAM+); a laptop will be provided for the duration of the training if needed.

Learning outcome:

Those who complete this course:

  • write Python scripts for data processing and automated checks
  • understand the architecture of LLM systems (chatbots, RAG, agents) and their typical defect classes
  • evaluate model responses against rubrics and document the results
  • build automated eval pipelines with metrics and regression runs
  • test RAG systems and agents, apply basic red teaming techniques
  • produce quality and vulnerability reports for AI systems

Training methods:

Lectures, practical work, seminars, demonstration, independent work.

The total course volume is 120 academic hours, of which 64 academic hours are classroom contact hours (including practical work).

Evaluation criteria for learning outcomes:

Learning outcomes are assessed based on independently completed practical work.

Evaluation methods:

Upon successful completion, practical and homework assignments receive a "pass" grade.

Course completion conditions:

To successfully complete the course and receive a certificate, it is necessary to achieve a "pass" grade on 75% of the homework assignments.

Additional information:

Training programme group: 0613 - Software and applications development and analysis (0613 - Tarkvara ja rakenduste arendus ning analüüs)
Basic rules for training organisation (in Estonian)
Basic rules for ensuring the quality of the educational process (in Estonian)

Course program

Module Main topics Volume
Block 1. Python foundations
  • Held jointly with the group of the «Fundamentals of the Python Programming Language» course (lessons 1–8).
  • Setup and environments (IDLE, Jupyter Notebook), variables and data types
  • Conditions, loops, lists, strings, dictionaries and sets
  • Functions and modular code
  • Files (CSV), installing packages (pip), exception handling
  • 32 ac/h
    1. Python for LLM testing
  • Virtual environments (venv), project organisation
  • HTTP requests, calling LLMs via API (OpenAI, Anthropic)
  • JSON: test-case datasets, parsing model responses
  • pytest basics; first result summaries (pandas)
  • 4 ac/h
    2. LLM and GenAI systems: what we test
  • How LLMs work: tokens, context, temperature, non-determinism
  • System types: chatbots, RAG, agents; LLM application lifecycle
  • Defect classes: hallucinations, bias, injections, degradation
  • How LLM testing differs from classic QA; EU AI Act, ISO/IEC 25059
  • 4 ac/h
    3. Prompt engineering and manual evaluation
  • Prompting techniques for the tester
  • Rubric design; side-by-side comparison and response ranking
  • Human-in-the-loop: how AI evaluation platforms work
  • Recording failure modes: written defect reports
  • 4 ac/h
    4. Quality metrics and LLM-as-judge
  • Metrics: accuracy, completeness, consistency, faithfulness, relevance
  • Reference-based and reference-free evaluation
  • LLM-as-judge: designing and validating judge prompts
  • Error analysis: failure taxonomy, prioritisation
  • 4 ac/h
    5. Automation: eval pipelines
  • The pipeline: dataset → metrics → run → regression
  • Tools: promptfoo, DeepEval + pytest
  • Golden datasets and synthetic test cases
  • CI integration; quality dashboards
  • 4 ac/h
    6. Testing RAG and AI agents
  • RAG architecture and its failure points: retrieval, grounding
  • RAG metrics (the Ragas approach): faithfulness, context relevance
  • Agents: tool-use correctness
  • Tracing and observability (LangSmith)
  • 4 ac/h
    7. Red teaming and AI safety
  • Prompt injection and jailbreaks: techniques and defences
  • Adversarial testing; bias/fairness auditing
  • Guardrails and content moderation
  • AI system vulnerability report
  • 4 ac/h
    8. Final practical project
  • End-to-end evaluation of a training LLM application
  • Rubric + automated eval run + red-team pass
  • Writing the quality report
  • Defence of the final project
  • 4 ac/h

    Course information

    Time of conduct:

    06.10.2026 - 19.11.2026
    27.10.2026 - 11.12.2026


    Timetable:

    Block 1 (Python): Tue, Thu, Fri 17:45–21:00 (3 weeks).
    Block 2 (LLM): Tue, Thu 17:45–21:00 (4 weeks).


    Apply → We'll reply within 1 business day

    Course length:

    7 weeks



    Format and place of conduct:

    Address: Tartu mnt. 18, Tallinn / Online.
    Gamma Intelligence Training Centre
    The course is conducted in a classroom format (up to 12 people in class) and/or online (Zoom / Microsoft Teams); total group size up to 18 people. The Python foundations block is held jointly with the group of the Fundamentals of the Python Programming Language course.

    Training language: English

    Price: 1700 EUR (VAT 24% included)

    Total course volume: 120 ac/h (64 classroom + 56 independent)
    Format: lectures + practical work + independent work.


    Tutors

    Roman Kutselepa (Python foundations block)

    Roman Kutselepa Qualification: over 5 years of Python development and over 3 years of JavaScript development; participated in software-integration projects.

    Specialisation: Python development, web development, software solution integration.

    Teaching experience: over 5 years of teaching and staff-training experience.

    Education: Anglia Ruskin University, higher education (2010).

    Review the CV


    Nikolay Zubrilov (LLM block)

    Qualification: Python developer and Data Scientist; AI/LLM developer at Mentastic; owner of Dataocean Analytics OÜ.

    Specialisation: Python, AI/LLM development, data analysis and data science (Pandas, NumPy, Scikit-Learn), SQL (PostgreSQL / MySQL), REST API (FastAPI / Flask), automation.

    Teaching experience: instructor on data analysis, Python and LLM testing courses at Gamma Intelligence Training Centre.

    Education: Master's degree — computer and systems engineering.

    Review the CV