innovator · learns

Evaluation & Red-Teaming

Master AI safety by learning to benchmark, test, and ethically challenge language models through adversarial techniques. Build real-world red-teaming skills to make AI systems more robust.

20 modules·Difficulty: ★★★★★· 5 Free
Start module 1

Modules

  • 1
    Free8 min
    What Is AI Evaluation?
    Discover why testing AI is different from testing traditional software and explore the evaluation landscape.
  • 2
    Free10 min
    Benchmark Types and Metrics
    Learn about accuracy, perplexity, BLEU scores, and other metrics that measure language model quality.
  • 3
    Free9 min
    Safety vs. Capability Tradeoffs
    Understand how making AI safer can sometimes limit what it can do and why balance matters.
  • 4
    Free11 min
    Introduction to Red-Teaming
    Meet the ethical hackers of AI who test systems by thinking like adversaries to find flaws.
  • 5
    Free10 min
    The Evaluation Pipeline
    Build your first automated test suite to check if a model gives consistent answers over time 🔧
  • 6
    Paid12 min
    Crafting Adversarial Prompts
    Design tricky questions and edge cases that expose how models handle ambiguous or conflicting input.
  • 7
    Paid11 min
    Bias Detection in Outputs
    Test whether a model treats different groups fairly by analyzing its responses to demographic prompts.
  • 8
    Paid10 min
    Prompt Injection Basics
    Explore how attackers slip hidden instructions into prompts to override system guidelines.
  • 9
    Paid12 min
    Hallucination and Factuality Tests
    Create benchmark questions with known answers to measure how often a model invents false information.
  • 10
    Paid11 min
    Multi-Turn Attack Sequences
    Learn how adversaries build trust over several messages before attempting a policy violation.
  • 11
    Paid10 min
    Jailbreak Taxonomy
    Classify different jailbreak strategies from role-play tricks to encoding schemes and linguistic exploits.
  • 12
    Paid12 min
    Ethical Boundaries of Testing
    Understand when red-teaming crosses into harmful use and how to test responsibly within legal limits.
  • 13
    Paid11 min
    Automating Adversarial Search
    Write scripts that generate thousands of prompt variations to find vulnerable input patterns quickly.
  • 14
    Paid10 min
    Guardrail Mechanisms
    Examine content filters, output validators, and other defenses that prevent harmful model responses.
  • 15
    Paid12 min
    Documenting Vulnerabilities
    Practice writing clear, reproducible reports that developers can use to patch security issues.
  • 16
    Paid11 min
    Capstone: Red-Team a Chatbot (Part 1)
    Apply benchmark design and adversarial prompting techniques to audit a real conversational agent.
  • 17
    Paid12 min
    Capstone: Red-Team a Chatbot (Part 2)
    Expand your attack surface by testing multi-turn sequences and hallucination triggers from earlier modules.
  • 18
    Paid10 min
    Capstone: Red-Team a Chatbot (Part 3)
    Evaluate guardrails and automated defenses, then measure your success rate against each barrier.
  • 19
    Paid11 min
    Capstone: Red-Team a Chatbot (Part 4)
    Compile your findings into a professional vulnerability report that balances detail with ethical responsibility.
  • 20
    Final exam9 min
    Final Exam: Evaluation & Red-Teaming Mastery
    Demonstrate your understanding of benchmarks, adversarial techniques, and ethical testing practices.