Back to courses
IT & ENGINEERING Advanced

LLM Evaluation and Testing: Shipping Reliable AI

Read the first lesson free — in full No account, no card · plus the interactive platform demo and the AI Professor Start now

A premium, complete and advanced course on evaluating and testing LLM-powered systems, updated for 2026. You will learn why evals are the hardest and most important discipline in AI engineering, the full evaluation taxonomy (offline vs online, reference-based vs reference-free), why classic metrics such as BLEU, ROUGE and exact match fail on open-ended text, and how embedding-based metrics help and mislead. You will master LLM-as-a-judge (rubric design, position/verbosity/self-preference bias, calibration against humans), building golden datasets with edge cases and slices, and the real 2026 toolchain: promptfoo, LangSmith, Langfuse, Braintrust, DeepEval, Ragas and pytest. The course goes deep on evaluating RAG (retrieval and generation), evaluating agents (task completion and trajectory), safety and guardrail evals, red-teaming, regression testing prompts in CI/CD, online A/B testing, quality-drift monitoring, human annotation and the real cost of evaluation. Includes runnable code, production patterns, a strong focus on data privacy and GDPR for eval data, and a comprehensive final assessment. Educational content, not legal advice.

10 modules
30 lessons
~25h duration
v1.0 version
AI professor An AI agent built into every lesson — ask questions and get instant answers based on the course content
Hands-on exercises Real scenarios and practical exercises directly on the platform, with instant feedback
Progress & analytics A personal dashboard with statistics, streaks, scores and structured learning paths
Interactive AI quizzes Questions generated by AI and adapted to your level, with detailed explanations
Individual course access
€49
+ VAT / month
Get started
All lessons AI quizzes AI professor included Cancel anytime
Or read the first lesson free
or
Recommended
IT Pro bundle
€399
+ VAT / month
See the IT Pro bundle
  • Every IT Pro courseA full library, not just this course
  • AI professor in every lessonAnswers when you need them, included in your subscription
  • Quizzes, progress, streaks & statistics
  • Content updated regularly
Cancel anytime
Secure payment
Updated regularly
Content in English
Built-in AI agent Exclusive Ask anything about the lesson and get an instant answer — the agent knows the course content
Interactive AI chat Automatic summaries Personalized quizzes

What you will learn

Practical skills you gain by completing this course

Foundations: Why LLM Evaluation Is Hard and Essential
Classic Metrics and Why They Fail on Open-Ended Text
LLM-as-a-Judge: Design, Bias and Calibration
Building Evaluation Datasets
The Evaluation Tooling Landscape
Evaluating RAG Systems
Evaluating AI Agents
Safety, Guardrails and Red-Teaming Evaluation
Evaluation in Production: CI/CD, A/B Testing, Monitoring and Cost
Final Quiz — LLM Evaluation and Testing

Who it is for

Developers Software engineers Solution architects CTOs / Tech Leads Data Scientists ML Engineers DevOps Engineers

Recommended level

Advanced

Assumes hands-on experience with AI and complex scenarios.

Updates

Regular

Last update: Aug 8, 2026. Content kept up to date.

Category

IT & Engineering

A technical course for IT professionals — available with individual course access or the IT Pro / All Access bundle.

Advanced level

Hands-on experience required

Assumes practical experience with AI. Covers complex scenarios and advanced strategies.

Always up to date

Last update: Aug 8, 2026

The course is updated regularly with the latest information, tools and practices from the industry.

Practical and applied

30 lessons with real examples

Each lesson includes practical scenarios, actionable checklists and quizzes to check your understanding.

Curriculum

10 modules, 30 lessons — structured to learn step by step.

10 modules
30 lessons
~25h of content
Interactive quizzes
Free preview available Why LLM Evaluation Is Hard — and Non-Negotiable
Read the preview
1 Free preview lesson Why LLM Evaluation Is Hard — and Non-Negotiable
Read the preview
2 The Evaluation Taxonomy: Offline vs Online, Reference-Based vs Reference-Free
50 min
3 Building Your First Eval Harness with pytest
50 min
4 Reading Eval Results: Variance, Significance and Confidence
50 min
1 Exact Match, BLEU, ROUGE and String Metrics
50 min
2 Embedding-Based Metrics: BERTScore and Semantic Similarity
50 min
3 When Classic Metrics Mislead You
50 min
1 Designing an LLM Judge
50 min
2 Judge Biases and How to Mitigate Them
50 min
3 Calibrating and Validating Your Judge Against Humans
50 min
1 Golden Sets and Test Case Design
50 min
2 Edge Cases, Adversarial Inputs and Slices
50 min
3 Synthetic Data Generation for Evals
50 min
1 promptfoo: Config-Driven Evaluation and Red-Teaming
50 min
2 LangSmith, Langfuse and Braintrust: Tracing and Eval Platforms
50 min
3 DeepEval: Pytest-Native LLM Testing
50 min
1 Retrieval Metrics: Recall, Precision, MRR and NDCG
50 min
2 Generation Metrics: Faithfulness, Answer Relevancy and Ragas
50 min
3 End-to-End RAG Evaluation and Failure Attribution
50 min
1 Task Completion and Success Metrics
50 min
2 Trajectory and Tool-Call Evaluation
50 min
3 Multi-Turn and Simulation-Based Agent Evaluation
50 min
1 Safety Evals and Guardrail Testing
50 min
2 Adversarial Testing and Red-Teaming
50 min
3 Eval Data Privacy, Governance and GDPR
50 min
1 Regression Testing Prompts in CI/CD
50 min
2 Online Evaluation and A/B Testing in Production
50 min
3 Monitoring Quality Drift and Human Annotation
50 min
4 The Cost of Evaluation
50 min
1 Final Assessment — Shipping Reliable AI
45 min
Access this course from €49 / month

Ready to start learning?

Create an account and choose how you want to learn — just this course, or the full IT Pro bundle.

30 hands-on lessons Content updated regularly AI professor included in your subscription