AI evaluation, without the jargon
Test one chatbot. Learn every eval platform.
Use the same five customer questions, the same pass rules, and the same before-and-after comparison in every guide.
AI trainingChoose your platform
One lesson, six implementations.
Each course starts with Northstar Shoes, five support questions, and clear yes-or-no checks.
Arize Phoenix: Evaluate a Simple Chatbot
Use Phoenix to save five support questions, run a chatbot against them, and compare one change without losing the failed examples.
Open guideLangfuse: Evaluate a Simple Chatbot
Use Langfuse to save five support questions, score each answer, and compare a new chatbot version with the current one.
Open guideLangSmith: Evaluate a Simple Chatbot
Follow LangSmith’s chatbot workflow: create a small set of expected answers, run two versions, and compare the failed questions.
Open guideBraintrust: Evaluate a Simple Chatbot
Use Braintrust’s data, task, and scorer pattern to test a support chatbot on five understandable cases.
Open guideHelicone: Evaluate a Simple Chatbot
Use Helicone to save useful chatbot requests and keep simple quality scores beside cost and latency.
Open guideOpenAI Evals: Evaluate a Simple Chatbot
Use OpenAI Evals to define five chatbot cases once, run them against a model, and inspect each grader result before release.
Open guideMore training
Longer courses for when the five-question workflow feels familiar.
Tool Use with Claude: Connect AI to External Tools and APIs
Learn how to connect Claude to external tools and APIs. Understand client vs server tools, the agentic loop, tool definitions, parallel tool calls, strict mode, and production patterns.
AI Evals for Beginners: How to Measure LLM Quality
Learn the fundamentals of AI evaluation — what evals are, why they matter, and how to build your first eval pipeline from scratch with practical code examples.
Advanced AI Evals: Production Pipelines, Custom Judges, and CI/CD
Take your eval game to production: custom LLM judges, multi-dimensional scoring rubrics, RAG evaluation, CI/CD integration, and real-world eval architecture patterns.