Cursuri AI.ro International — AI e-learning for professionals

Try the platform before you pay

Explore the interactive account demo and read the first lesson of LLM Evaluation and Testing: Shipping Reliable AI for free, in full — no account, no card.

AI Professor in every lesson Quizzes with feedback Cancel anytime
cursuri-ai.ro/app/courses/llm-evaluation-and-testing Interactive demo · sample account
Cursuri AI.ro International — AI e-learning for professionals

Welcome! 👋

Welcome to Cursuri AI International — AI e-learning for professionals, now in English. Activate your subscription and start learning.

AI virtual professor Exclusive to Cursuri AI

A mentor who knows every lesson in detail and guides you personally.

Included in your subscription

Everything you get with a single subscription

Full access to the entire platform, with no hidden limits or extra costs.

50 premium courses Practical content created by specialists, constantly updated with the latest in AI
24/7 AI virtual professor A personal mentor built into every lesson — ask any question and get an instant answer
1459 hands-on lessons Short, structured lessons with practical exercises you can apply at work right away
Smart AI quizzes Automatically generated questions, tailored to your weak spots, with detailed feedback
Save and organize Bookmarks, personal notes and a spaced-repetition system for maximum retention
Progress and analytics Streaks, weekly goals and personalized insights into your growth

Start learning AI today

Pick the right plan and get instant access to premium courses, an AI virtual professor, smart quizzes and hands-on exercises.

No long-term commitment Cancel anytime, no penalties Instant access after payment
This is a preview with sample data. Activate your subscription to save your conversations with the AI Professor.

AI Conversations

The complete history of your dialogues with the AI Virtual Professor

4 Conversations
36 Messages
2 Courses
AI Fundamentals 2 conversations
You

Can you explain how an artificial neuron works?...

AI Professor
AI Professor

An artificial neuron simulates how a biological neuron works: 1. Inputs — it receives multiple numeric values 2. Weights — each input has a weight that determ...

You

What is the ReLU activation function?...

AI Professor
AI Professor

ReLU (Rectified Linear Unit) is the most popular activation function: - Formula: f(x) = max(0, x) - If the input is positive → it passes through unchanged - If the input is negative → it turns it into ...

Prompt Engineering Masterclass 2 conversations
This is sample data Activate a subscription to save and export your real conversations with the AI Professor.
See plans
This is a preview with sample data. Activate a subscription to see your real scores.

Scores

Your performance across 5 assessed courses

79% Overall average
Very good
5
Courses assessed
4 / 5
Passed (≥70%)
22
Quizzes completed
168 / 220
Total points
Score distribution
4 courses ≥70% 1 course 40-69%
Best result Prompt Engineering Masterclass — 92%

Performance by course

92
Prompt Engineering Masterclass Passed
46 / 50 points 5 quizzes 92%
85
Introduction to AI Engineering Passed
34 / 40 points 4 quizzes 85%
76
AI for Digital Marketing Passed
38 / 50 points 5 quizzes 76%
71
RAG — Retrieval Augmented Generation Passed
21 / 30 points 3 quizzes 71%
58
AI Agents & Automation Not passed
29 / 50 points 5 quizzes 58%
This is sample data Activate a subscription to see your real quiz scores.
See plans
Structured learning paths

From zero to expert,
step by step

Career paths and specializations with a clear structure, from fundamentals to advanced level. Each course prepares you for the next — no guessing what to learn next, the platform guides you.

50 Premium courses
1459 Hands-on lessons
24/7 AI Professor

How a learning path works

Four simple steps from your first lesson to expert level

01

Choose the right path

Two complete career paths plus focused specializations — for engineers, managers, marketers, or professionals in specialized fields.

02

Complete the stages in order

Each path is divided into clear stages. Each course prepares you for the next — no skipped steps, no lost context.

03

The AI Professor guides you

Ask anything in every lesson. You get explanations, practical examples, and personalized quizzes.

04

Track your progress

Progress per stage and per path, quiz scores, and automatic recommendations for your next step.

AI Virtual Professor Exclusive

Built into every lesson of every path

Throughout your entire path, you have access to an AI professor that knows the content of every lesson. Ask anything — you get detailed explanations, industry examples, and quizzes adapted to your level.

AI chat included in your subscription Answers tailored to the lesson you're studying
Automatic summaries Key points intelligently extracted
Personalized quizzes Unique questions, calibrated to you

Start your first learning path today

Choose the right plan and get instant access to the AI Professor, smart quizzes, and all premium courses.

Instant access Cancel anytime AI Professor included
Individual courses

Choose the courses you're interested in

Subscribe to each course individually. €49 + VAT / month per course, cancel anytime.

IT PRO Advanced

LLM Evaluation and Testing: Shipping Reliable AI

The dedicated, advanced masterclass on evaluating and testing LLM systems — judges, datasets, RAG and agent evals, CI/CD and production monitoring, updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Coaches and Course Creators

Build and grow your coaching or course business with AI — niche, curriculum, client delivery, honest marketing, lead magnets, community and support. Ready-to-use prompts. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Research and Academia

Use AI ethically and effectively across the research lifecycle — literature, writing, grants, references, and dissemination — with academic integrity as the spine. Ready-to-use prompts, updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Operations and Supply Chain

Apply AI across the end-to-end supply chain the practical way — forecasting, inventory, procurement, logistics, production, maintenance and risk, with human oversight. Updated for 2026.

30 lessons ~25h
BUSINESS Beginner Popular

AI for Your Job Search and Career Growth

Run a smarter, honest job search and grow your career with AI — resume, LinkedIn, interviews, salary negotiation, networking and upskilling. Ready-to-use prompts. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Legal Professionals

Use AI across legal research, contracts, due diligence and writing the responsible way — verify every citation, protect privilege, keep the lawyer in control. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Consultants and Freelancers: Build a Solo AI-Powered Business

Run a lean, high-margin solo consulting or freelance business with AI as your unfair advantage — win clients, deliver faster, automate the boring parts. Guardrails included. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for E-commerce and Retail

Apply AI across your online store the right way — listings, imagery, personalization, search, pricing, inventory, reviews, chatbots and marketplaces, with EU compliance built in. Updated for 2026.

30 lessons ~25h
BUSINESS Beginner

AI for Educators and Teachers

Save hours every week and teach better with AI — lesson planning, materials, differentiation, feedback and quizzes — the responsible, privacy-safe way. Beginner-friendly, updated for 2026.

31 lessons ~25h
BUSINESS Intermediate

AI for Project and Product Management

Use AI across the whole project and product lifecycle the smart way — discovery, PRDs, roadmaps, backlog, planning, analytics and experiments — with humans in control. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate Popular

AI Video Creation and Editing: From Script to Publish

Create professional video with AI end to end — script, generate, voice, edit and publish — using Runway, Veo, Kling, HeyGen, ElevenLabs, Descript and CapCut, the legal and ethical way. Updated for 2026.

31 lessons ~26h
IT PRO Advanced

Time Series Forecasting with Machine Learning

The complete, hands-on time series forecasting course — from stationarity, ARIMA and exponential smoothing to gradient boosting, global models, N-HiTS, Transformers, 2026 foundation models and conformal intervals, with honest evaluation and production deployment.

30 lessons ~25h
IT PRO Advanced

Edge AI and On-Device Intelligence

Ship real Edge AI: compress models, master LiteRT/ONNX/Core ML/ExecuTorch, run SLMs on phones and microcontrollers, and deploy privately with OTA updates and monitoring. Updated for 2026.

30 lessons ~25h
IT PRO Advanced

Recommender Systems with AI: From Collaborative Filtering to Deep Learning

The complete, hands-on recommender systems course — content-based and collaborative filtering, matrix factorization, deep learning, two-tower retrieval, sequential transformers, evaluation, LLMs and GDPR, updated for 2026.

30 lessons ~25h
IT PRO Advanced

AI for Cybersecurity: Threat Detection and SOC Operations

A hands-on blue-team course on using AI and machine learning for threat detection, SOC operations, UEBA, alert triage, SOAR automation and threat intelligence, updated for 2026.

30 lessons ~25h
IT PRO Advanced

Reinforcement Learning and RLHF: Training and Aligning AI Models

The complete, university-level RL and RLHF course — from MDPs, Q-learning, DQN and PPO to reward models, DPO, GRPO, Constitutional AI and verifiable rewards — with real Gymnasium, Stable-Baselines3 and TRL code, updated for 2026.

30 lessons ~25h
IT PRO Advanced

Data Engineering for AI: Pipelines, Vector Stores and Data Quality

The complete, advanced course on the data foundation for AI: pipelines, lakehouse, data quality, feature stores, embeddings and vector stores - updated for 2026.

31 lessons ~26h
IT PRO Advanced

Computer Vision with AI: From Detection to Multimodal Understanding

The complete, advanced masterclass on modern computer vision - from OpenCV and CNNs to detection, segmentation, Vision Transformers, multimodal VLMs and production MLOps, updated for 2026.

31 lessons ~25h
IT PRO Advanced

MLOps: The Machine Learning Lifecycle in Production

The complete, advanced masterclass on operating machine learning models in production — tracking, pipelines, serving, monitoring, drift, retraining and governance — updated for 2026.

30 lessons ~25h
IT PRO Intermediate

Deep Learning and Neural Networks with PyTorch

The complete, hands-on deep learning course with PyTorch — from tensors and autograd to CNNs, Transformers, GPU training and production deployment, updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI Data Privacy and EU AI Act Compliance for Professionals

Master AI data privacy and EU AI Act compliance the right way — GDPR, risk tiers, provider vs deployer duties and internal AI governance. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Finance and Accounting Professionals

Use AI to speed up finance and accounting work — reporting, analysis, forecasting, reconciliation, AP/AR, audit support and document processing — with every number verified by a human. Updated for 2026.

30 lessons ~25h
BUSINESS Intermediate

AI for Customer Support and Service

Deploy AI across customer support the right way — deflection, chatbots, RAG, agent copilot, triage and voice — with human escalation, EU AI Act and GDPR built in. Updated for 2026.

30 lessons ~25h
BUSINESS Beginner Popular

The AI Productivity Masterclass for Knowledge Workers

Work faster and think better with ChatGPT, Claude and Gemini — a practical, tool-agnostic AI productivity course for every knowledge worker.

30 lessons ~25h
BUSINESS Intermediate

AI for HR: Recruiting, Onboarding and L&D — Done Compliantly

Use AI across recruiting, onboarding and L&D the compliant way — human-in-the-loop, EU AI Act and GDPR ready. Updated for 2026.

30 lessons ~25h
IT PRO Advanced

AI Security: Defending LLM Applications (OWASP LLM Top 10, Guardrails, Red-Teaming)

The complete, advanced masterclass on defending LLM applications: OWASP LLM Top 10, guardrails, and authorized red-teaming — updated for 2026.

30 lessons ~25h
IT PRO Advanced

AI for DevOps and SRE: AIOps in Practice

The complete, advanced masterclass on AIOps for DevOps and SRE — intelligent observability, incident response, RCA, CI/CD, IaC and safe automation, updated for 2026.

30 lessons ~25h
IT PRO Advanced

Voice AI and Realtime Multimodal Agents

The complete, advanced masterclass on building realtime voice and multimodal AI agents — STT, TTS, the Realtime API, telephony, and production deployment for 2026.

30 lessons ~25h
IT PRO Advanced

Build and Ship a Production AI SaaS: From Idea to Paying Users

The complete, advanced masterclass on building and shipping a production AI SaaS in 2026 — idea, stack, LLM integration, RAG, auth, Stripe billing, observability, deployment and go-to-market.

30 lessons ~25h
IT PRO Advanced

Fine-Tuning and Customizing Open-Source LLMs: LoRA, QLoRA and Self-Hosting

The complete, advanced masterclass on fine-tuning open-source LLMs with LoRA, QLoRA and self-hosting — updated for 2026.

30 lessons ~25h
BUSINESS Beginner

Microsoft 365 Copilot for Office Work: Role-Based Productivity

Turn Word, Excel, Outlook, Teams and PowerPoint into AI assistants — real productivity for your role, not generic tricks.

32 lessons ~24h
BUSINESS Intermediate Popular

SEO and AEO/GEO in the AI Era: Optimizing for Google, AI Overviews and Generative Engines

Modern SEO plus end-to-end AEO/GEO: how to appear in Google and AI Overviews and get cited by ChatGPT, Perplexity, Gemini

30 lessons ~26h
BUSINESS Intermediate

Manager in the AI Era: Leading Your Team Through the AI Transformation

Lead your team through the AI transformation with proven frameworks: ADKAR, Kotter and the psychology of adoption

30 lessons ~24h
BUSINESS Intermediate Popular

AI Image Generation: The Complete Guide from Prompt to Publishing

Premium course: visual prompt engineering, AI platforms, production at scale, legal, ROI and hands-on tutorials — updated for 2026

26 lessons ~27h
BUSINESS Beginner

No-Code Data Analysis with AI: ChatGPT, Excel and SQL for Non-Programmers

Analyze data with conversational AI, no code: ChatGPT, Excel and SQL for non-programmers, with verification and GDPR.

27 lessons ~24h
BUSINESS Intermediate Popular

AI for Entrepreneurs and Startups: The Complete Guide

Premium course: idea validation, no-code MVP, go-to-market, finance, scaling and the startup ecosystem with AI — updated for 2026

30 lessons ~28h
BUSINESS Intermediate Popular

AI for Sales and CRM

The complete AI masterclass for sales and intelligent CRM — updated for 2026

28 lessons ~25h
BUSINESS Intermediate Popular

AI for Content Creation and Copywriting

Premium course: copywriting, editorial strategy, social media, video, audio, brand voice, SEO, KPIs and scaling with AI — updated for 2025-2026

29 lessons ~29h
BUSINESS Beginner Popular

AI for Digital Marketing

The complete AI masterclass for digital marketers — updated for 2026

31 lessons ~27h
IT PRO Advanced Popular

Context Engineering and Memory for AI Agents: Beyond Prompting

Engineer AI agents' context and memory for reliability at scale: context anatomy, memory types, retrieval, compaction, cost.

25 lessons ~24h
BUSINESS Intermediate Popular

AI for Business Leaders

The complete AI masterclass for leaders and management — updated for 2026

24 lessons ~22h
IT PRO Advanced Popular

Claude Code Mastery: Agentic Coding from the Terminal (multi-file, git, CI, MCP)

Terminal-first agentic coding with Claude Code: multi-file, git, headless, subagents, hooks, CI/CD, MCP, security and routing across the 2026 Anthropic lineup (including Claude Fable 5).

28 lessons ~25h
IT PRO Beginner

Vibe Coding: From Prompt to Application with Lovable, v0, Bolt and Replit Agent

Build full-stack applications from prompts and ship them safely, legally and responsibly — without writing code.

22 lessons ~24h
IT PRO Expert Popular

Cursor as a Pro: AI-Native IDE, Composer and Multi-Agent 2026 (Enterprise Edition)

Cursor 3 Pro: from Composer to Background Agents and Enterprise — April 2026

30 lessons ~25h
IT PRO Advanced Popular

MCP (Model Context Protocol) — Building Servers and Integrations (Enterprise Edition)

The complete MCP guide: from architecture to production servers — Python, TypeScript, security, deployment

22 lessons ~25h
IT PRO Advanced Popular

AI Agents: Architecting and Automating Autonomous Systems

The complete AI Agents and Automation masterclass for engineers — updated for 2026

30 lessons ~26h
IT PRO Advanced Popular

Advanced LLM Integration in Production Applications

The complete masterclass on LLM integration in production — updated for 2026

24 lessons ~26h
IT PRO Advanced Popular

RAG: Retrieval-Augmented Generation in Practice

The complete RAG masterclass for AI Engineers — updated for 2026

27 lessons ~24h
IT PRO Intermediate Popular

The Complete Prompt Engineering Masterclass

Advanced prompting techniques for professionals — updated for 2026

32 lessons ~27h
IT PRO Beginner Popular

Introduction to AI Engineering

The complete AI Engineering masterclass for developers — updated for 2026

28 lessons ~25h
This is a preview with sample data. Activate a subscription to save your lessons and notes.

Saved & Notes

Your saved lessons and personal notes, organized by course.

5 saved lessons 3 notes
Prompt Engineering Masterclass 3
Chain-of-Thought Prompting — The Complete Guide Mar 18
Few-Shot vs Zero-Shot — When to Use Each Mar 16
Advanced System Prompts for GPT-5.5 and Claude Mar 14
AI for Digital Marketing 2
Generating Ad Copy with AI — Meta & Google Ads Mar 12
Automating Campaigns with AI Agents Mar 10
This is sample data Activate a subscription to save your own lessons and personal notes.
View plans
This is a preview with sample data. Activate a subscription to generate personalized review flashcards.

Review

Key concepts from your lessons, generated by the AI Virtual Professor

3 Lessons
10 Cards
2 Courses
AI Fundamentals 7 cards
Prompt Engineering Masterclass 3 cards
This is sample data Activate a subscription to generate personalized review flashcards for every lesson.
View plans
This is a preview with demo data. Activate a subscription to see your real statistics.

Learning Insights

Your learning statistics

15h 30m total time 7 achievements
12/7 Lessons 175/180 Minutes 14/21 Streak
12
Lessons completed
175
Minutes invested
59%
5
Quizzes taken
14
Day streak

Personal records

21 days Longest streak
175m Most active week
45 days Total active days
22m/day Daily average
This is demo data Activate a subscription to see your real statistics and achievements.
See plans

Saved articles

The blog articles you saved for later, all in one place.

0 saved articles

No saved articles yet

When you read an article on the blog, hit the “Save” button to find it again quickly right here.

My account

Welcome

The premium AI education platform. Manage your account, subscription and progress.

Demo data. Activate a subscription to see your real progress.

Your progress

See courses
3
Courses in progress
5
Courses completed
47
Lessons completed
14
Day streak

You don't have an active subscription.

Choose a plan
Premium feature

Available with an account

This section becomes active once you create an account.

See available plans
50 premium courses 1459 hands-on lessons 24/7 AI Professor Secure payments with Stripe Cancel anytime
Full lesson, free Module 1 · Lesson 1

Why LLM Evaluation Is Hard — and Non-Negotiable

From the course LLM Evaluation and Testing: Shipping Reliable AI

50 min read Advanced Quiz included + 29 lessons with a subscription
Tip: select any passage in the lesson and hit “Ask the AI Professor” — that's exactly how it works inside the platform.

Built-in AI Professor Exclusive

Ask anything about the lesson and get an instant answer. The AI Professor knows the course content and helps you learn more effectively.

Select a passage in the lesson → it explains it on the spot Interactive chat — answers any question Automatic summaries with key points Personalized AI-generated quizzes

Most engineering teams can ship a web service they can reason about: the same input yields the same output, a failing assertion is unambiguous, and a green test suite means the behavior is locked in. LLM-powered systems break every one of those assumptions. The model is stochastic, the space of acceptable outputs is enormous and fuzzy, and a change that looks harmless — swapping a model version, editing one line of the system prompt, reordering few-shot examples — can silently shift quality across thousands of requests. Evaluation is the discipline that turns "it feels better" into "it is measurably better on the cases we care about, with no regression on safety." Without it you are flying blind, and in production that blindness is expensive.

This lesson frames the whole course. It explains why the familiar testing playbook collapses when the unit under test is a language model, what a real evaluation harness looks like at its core, and the maturity ladder that separates a demo from a dependable product. Everything that follows — metrics, judges, datasets, RAG and agent evals, CI and monitoring — is an elaboration of the ideas introduced here.

Why the usual testing playbook fails

Traditional software testing rests on two properties: determinism and exact correctness. A function that adds two numbers has exactly one right answer, and a unit test pins it forever. LLM outputs violate both properties at once, and they violate them in ways that are not fixable with a cleverer assertion.

  • Non-determinism. With any temperature above zero the same prompt produces different completions. Even at temperature zero you cannot rely on byte-identical output: providers batch requests, floating-point reductions on GPUs are not associative, and the vendor may silently update the served checkpoint. "Reproducible" for an LLM means "statistically stable," not "identical."
  • No single correct answer. For "summarize this support ticket" there are thousands of good summaries and thousands of bad ones. Correctness is a fuzzy region in output space, not a point. An exact-match assertion against one reference rejects perfectly good outputs and accepts nothing useful.
  • Semantic, not lexical, equivalence. "The payment failed because the card expired" and "Charge declined: expired card" mean the same thing and share almost no tokens. Any metric that compares surface strings scores them as wildly different, which is exactly backwards.
  • Emergent, high-dimensional failure modes. LLMs fail in ways ordinary functions cannot: hallucinated facts, subtle instruction-following drift, format violations, tone problems, prompt-injected content, refusal on benign requests, and sycophancy. A single pass/fail bit cannot represent this surface.

The practical consequence is blunt: you cannot assert output == expected. You need graded, multi-dimensional evaluation that tolerates paraphrase while still catching real regressions. The rest of the discipline is about building measurements that behave well under exactly these conditions.

The cost of shipping without evals

Teams that skip systematic evaluation do not avoid evaluation — they outsource it to production and to their users. The symptoms are predictable. A prompt tweak that improves one demo quietly degrades ten other intents. A model upgrade the vendor promised was "strictly better" changes the output format just enough to break a downstream JSON parser. A retrieval change lifts recall but floods the context with distractors and lowers answer quality. Because nobody measured, the regression surfaces days later through complaints, churn, or a support escalation, and debugging starts from zero because there is no baseline to compare against.

Consider a concrete scenario. A fintech support assistant is upgraded from one model generation to the next because the newer model is cheaper and faster. In a five-minute manual check it looks great. Two weeks later, refund-related tickets spike. Investigation reveals the new model is subtly more cautious and now appends a disclaimer that trips a regex the downstream ticketing system uses to auto-categorize messages. No error was thrown anywhere; the pipeline "worked." The only missing ingredient was an eval that scored format conformance and category routing on a representative set before the change shipped. The fix took an afternoon; finding the problem took two weeks and eroded trust.

"Vibes-based development" — eyeballing a handful of outputs after each change — feels fast and is genuinely useful in the first days of a project. But it does not scale, it misses rare-but-severe failures, and it cannot run automatically on every commit. The entire purpose of an eval suite is to convert that manual squinting into a repeatable, quantitative gate that a machine can enforce while you sleep.

Evaluation is a first-class product surface

The maturity ladder for LLM applications is worth memorizing, because most teams can locate themselves on it honestly:

Level Practice What it catches What it misses
1 Manual spot checks in a playground Obvious breakage Everything rare or subtle
2 A saved set of hard prompts re-run by hand Known failure cases New regressions; runs only when remembered
3 Automated offline eval suite with metrics + thresholds Quantified regressions on a dataset Live-traffic drift
4 That suite wired into CI as a merge gate Regressions before they merge Production-only behaviors
5 Online evaluation + monitoring on real traffic Drift, novel inputs, real user harm — (this is the goal)

The gap between a hobby project and a reliable product is almost entirely the climb from level 1 to levels 3–5. The point is not to reach level 5 on day one; it is to keep climbing and to never let a serious change — new model, new prompt, new retrieval strategy, new tool — ship without being judged against the same yardstick as the last one.

A minimal harness makes the problem concrete

Here is the trap that motivates everything in this course. The naive test uses exact string match, and it fails on a semantically perfect answer:

# naive_test.py - why exact match is the wrong default
reference = "The payment failed because the card had expired."
model_output = "Charge declined: the customer's card was expired."

def exact_match(pred: str, ref: str) -> bool:
    return pred.strip() == ref.strip()

assert exact_match(model_output, reference)  # FAILS - yet the answer is correct

Because exact_match returns False, this "test" would block a good output. The fix is not a better string comparison; it is a different kind of check. A robust harness scores along the axes that matter and asserts on thresholds, not identity:

# harness.py - the shape every eval framework generalizes
from dataclasses import dataclass

@dataclass
class Case:
    prompt: str
    expected_facts: list[str]   # things that MUST appear (semantically)
    forbidden: list[str]        # things that must NOT appear (PII, competitors)

def evaluate(case: Case, output: str, judge) -> dict:
    covers = judge.entailment(output, case.expected_facts)   # 0..1 semantic coverage
    leaks = any(judge.contains(output, f) for f in case.forbidden)
    return {
        "coverage": covers,
        "safe": not leaks,
        "passed": covers >= 0.8 and not leaks,
    }

Three ideas recur across every tool we will study. First, a dataset of cases that represents the traffic you care about, including the rare and the adversarial. Second, a scorer that understands meaning rather than surface form — an embedding metric, an entailment check, or an LLM judge. Third, an explicit pass threshold that encodes your quality bar so the result is a decision, not a vibe. LangSmith, promptfoo, Braintrust, DeepEval and Ragas are all production-grade elaborations of exactly this skeleton. When a tool feels overwhelming, ask which of these three pieces each feature belongs to and it snaps into place.

The five properties of a healthy eval practice

A mature evaluation practice has five properties. Use them as a self-audit checklist for any suite you inherit or build.

  • Representative. The dataset mirrors real traffic — the common cases and the long tail of rare, hard, and adversarial inputs. An eval that only contains easy cases lies to you by reporting numbers that production will not reproduce.
  • Multi-dimensional. Correctness, faithfulness, format conformance, safety, tone, cost and latency are separate signals. Collapsing them into one number hides which dimension moved and why.
  • Comparative. You always score a change against a baseline, so you read deltas ("+3% coverage, no safety regression"), not uninterpretable absolutes ("0.81").
  • Automated. It runs on every change, not when someone remembers. Automation is what turns evaluation from a pre-launch ritual into an always-on gate.
  • Honest. It is designed to surface regressions, not to produce a flattering number. An eval you can game by prompt-tuning to the test set is worse than none, because it manufactures false confidence.

Common pitfalls before you write a single metric

Even experienced teams stumble on the same rocks. Testing on the same handful of demo prompts that were used to build the feature — the model looks perfect because you tuned to those exact cases. Measuring only average score and missing that the average hides a catastrophic 2% of inputs where the system does real harm. Treating one green run as proof despite the underlying non-determinism, when you should run enough samples to separate signal from noise (a topic this module returns to). Letting the eval set leak into prompts or fine-tuning, which quietly turns your held-out measurement into training data and inflates every number. And optimizing the metric instead of the product — Goodhart's law applies with force to LLMs, because they are extraordinarily good at satisfying the letter of a rubric while missing its spirit.

From a vibe to a number: a worked example

Suppose your team keeps saying the assistant "feels worse at refunds lately." That sentence is unactionable. Turn it into a measurement in four moves. First, collect the cases: pull thirty real refund conversations from the last month, redacting personal data. Second, define the rubric for each: the good answer confirms the order, states the refund policy accurately, and never promises a timeline the business cannot keep. Third, score each output on three binary sub-checks — order confirmed, policy accurate, no over-promise — and compute the pass rate per check. Fourth, compare to a baseline: run the previous prompt version over the same thirty cases.

Now "feels worse" becomes "policy accuracy dropped from 93% to 74% after last week's prompt edit, concentrated in partial-refund cases." That statement is debuggable: you know the dimension (policy accuracy), the magnitude, and the slice (partial refunds). You can form a hypothesis, change one thing, and re-run the exact same thirty cases to confirm the fix. Notice how little machinery this required — thirty cases, three checks, one baseline — and how much clarity it bought. Every advanced technique in this course is a way to make this basic loop more representative, more automatic, and more trustworthy, but the loop itself is always the same: cases, rubric, score, compare.

In practice: where to start

If you are staring at a system with no evals, do not try to leap to level 5. Collect twenty to fifty real inputs, write down for each what a good and a bad answer looks like, and build the smallest harness that scores coverage and safety and prints a pass rate. That single afternoon converts "I think it's fine" into a number you can move deliberately. Then grow the set as production surprises you, wire the suite into CI, and only then reach for online monitoring. Evaluation is not a chore bolted on before launch; it is the instrument panel that lets you fly the system at all.

This material is educational and does not constitute legal advice. Where later lessons touch privacy, data protection and compliance, treat the guidance as a starting point and verify it against your own obligations and jurisdiction.

Real quiz · from Lesson 1

**[Easy]** Why can you not reliably use `assert output == expected` to test an LLM feature?

First lesson — read in full, for free

Enjoyed it? All 30 lessons look like this.

You just read a complete lesson, exactly as it appears in the platform. Create your account in under a minute and pick the option that fits you best:

Just this course €49+ VAT / month All 30 lessons + AI Professor (fair-use limit)
Recommended IT Pro bundle €399+ VAT / month Every IT Pro course + smart quizzes + full AI Professor
Instant access to Lesson 2 All 30 lessons in the course Interactive quizzes with feedback AI Professor built into every lesson Notes & progress saved automatically Access from any device
Instant access Cancel anytime Secure payment via Stripe

Up next in the course

  • 2 The Evaluation Taxonomy: Offline vs Online, Reference-Based vs Reference-Free 50 min
  • 3 Building Your First Eval Harness with pytest 50 min
  • 4 Reading Eval Results: Variance, Significance and Confidence 50 min
Unlock all 30 lessons
Course curriculum

Everything you'll learn in this course

10 modules
30 lessons
~25h of content
Advanced level
1 Foundations: Why LLM Evaluation Is Hard and Essential 4 lessons
  • Why LLM Evaluation Is Hard — and Non-Negotiable Reading now 50 min
  • The Evaluation Taxonomy: Offline vs Online, Reference-Based vs Reference-Free 50 min
  • Building Your First Eval Harness with pytest 50 min
  • Reading Eval Results: Variance, Significance and Confidence 50 min
2 Classic Metrics and Why They Fail on Open-Ended Text 3 lessons
  • Exact Match, BLEU, ROUGE and String Metrics 50 min
  • Embedding-Based Metrics: BERTScore and Semantic Similarity 50 min
  • When Classic Metrics Mislead You 50 min
3 LLM-as-a-Judge: Design, Bias and Calibration 3 lessons
  • Designing an LLM Judge 50 min
  • Judge Biases and How to Mitigate Them 50 min
  • Calibrating and Validating Your Judge Against Humans 50 min
4 Building Evaluation Datasets 3 lessons
  • Golden Sets and Test Case Design 50 min
  • Edge Cases, Adversarial Inputs and Slices 50 min
  • Synthetic Data Generation for Evals 50 min
5 The Evaluation Tooling Landscape 3 lessons
  • promptfoo: Config-Driven Evaluation and Red-Teaming 50 min
  • LangSmith, Langfuse and Braintrust: Tracing and Eval Platforms 50 min
  • DeepEval: Pytest-Native LLM Testing 50 min
6 Evaluating RAG Systems 3 lessons
  • Retrieval Metrics: Recall, Precision, MRR and NDCG 50 min
  • Generation Metrics: Faithfulness, Answer Relevancy and Ragas 50 min
  • End-to-End RAG Evaluation and Failure Attribution 50 min
7 Evaluating AI Agents 3 lessons
  • Task Completion and Success Metrics 50 min
  • Trajectory and Tool-Call Evaluation 50 min
  • Multi-Turn and Simulation-Based Agent Evaluation 50 min
8 Safety, Guardrails and Red-Teaming Evaluation 3 lessons
  • Safety Evals and Guardrail Testing 50 min
  • Adversarial Testing and Red-Teaming 50 min
  • Eval Data Privacy, Governance and GDPR 50 min
9 Evaluation in Production: CI/CD, A/B Testing, Monitoring and Cost 4 lessons
  • Regression Testing Prompts in CI/CD 50 min
  • Online Evaluation and A/B Testing in Production 50 min
  • Monitoring Quality Drift and Human Annotation 50 min
  • The Cost of Evaluation 50 min
10 Final Quiz — LLM Evaluation and Testing 1 lessons
  • Final Assessment — Shipping Reliable AI 45 min
What you get on the platform

Everything you need to learn effectively

Interactive quizzes

Check your knowledge at the end of every lesson with scored quizzes and feedback.

Personal notes

Save notes on every lesson, accessible anytime from your dashboard.

Scheduled reviews

Revisit lessons exactly when it matters, at the right intervals — so you remember for the long term.

Progress & Achievements

Track your progress, unlock achievements, and visualize what you've learned.

Bookmarks

Save the lessons that matter and find them instantly when you need them.

Questions & Answers

Ask questions right on the lesson and get answers from our team.

Frequently asked questions

Good to know before you start

How do I get access to the course?

You can read the first lesson in full for free, right on this page — no account needed. For the rest of the course you create an account, pick the subscription that fits — a single course or a bundle — and get access immediately after your payment is confirmed. Everything happens 100% online.

Can I cancel my subscription anytime?

Yes. Cancel anytime, straight from your account, in just a few clicks. Your access stays active until the end of the period you have already paid for.

What does the subscription for this course include?

All 30 lessons in the course, interactive quizzes, the AI professor built into every lesson (select any passage and it explains it on the spot), personal notes, automatically saved progress, and content updates included.

Is there a fixed learning schedule?

No. You learn at your own pace, on any device. Lessons are structured step by step, and the platform saves your progress automatically, so you can pick up right where you left off — anytime.

Ready to unlock all the content?

Just this course — €49 + VAT / month — or every IT Pro course, with smart quizzes and the full AI Professor, in the bundle at €399 + VAT / month.

Instant access Cancel anytime Secure payment via Stripe
From €49 + VAT / month · Cancel anytime
Read the free lesson