AI curriculum

Course outline for LLM and Agent Evaluation

Agent completion claims need independent evidence. Design evaluation suites that distinguish answer quality, task execution, retrieval performance, and human judgment.

About LLM and Agent Evaluation

Agent completion claims need independent evidence. Design evaluation suites that distinguish answer quality, task execution, retrieval performance, and human judgment.

LLM and Agent Evaluation Course Objectives

  • Define quality and task success metrics.
  • Build a representative regression dataset.
  • Compare human and automated judgments with explicit limitations.

Pre-requisites

  • Familiarity with LLM or agent application behavior.
  • Basic testing and data analysis skills.

Lab Setup

  • Computer with a test runner and spreadsheet or analysis notebook.
  • Use supplied interaction logs and model outputs; online behavior can be simulated without live users or paid judges.

Detailed Course Outline

Proposed modules

Module 1: Evaluation settings

  • Offline evaluation
  • Online evaluation
  • Quality metrics

Practical outcome: Choose measures and evaluation settings for a bounded task.

Module 2: Task and retrieval checks

  • Agent success metrics
  • RAG metrics
  • Regression suites

Practical outcome: Define tests for tool effects and retrieved evidence.

Module 3: Evaluation evidence

  • Evaluation datasets
  • Human evaluation
  • Automated judges

Practical outcome: Compare judgments and document disagreement.

Practical exercise

  • Evaluate two agent configurations on a small suite and explain their failure categories.

How we train

Contact us for full course details, including duration, delivery options and lab requirements.

Ask about exercises, instructor feedback and the prior knowledge you need. Tool-specific courses marked provisional may become modules in a broader course.

Discuss your learning goals