Your First Langfuse Evaluation
24-second evaluation loop
Turn chatbot failures into permanent tests
A production chatbot failure becomes a versioned test case, receives deterministic and semantic scores, and enters a repeatable release gate.
Read the visual transcript
- 01A support chatbot incorrectly promises that a customs fee is always refunded.
- 02Capture that production failure as one versioned dataset case with the expected policy behavior.
- 03Score hard rules with code and semantic quality with a calibrated judge or human review.
- 04Release on evidence, then make every meaningful production failure a permanent regression test.
Begin with one small support bot
Imagine a shoe-store chatbot with four rules: unworn items can be returned within 30 days, original shipping is not refundable, damaged items go to a specialist, and unanswered policy questions must not be guessed.
An evaluation asks the bot the same five questions before and after a change. Each answer gets a clear pass or fail. That is all you need to understand before using Langfuse.
Langfuse connects each saved question to the chatbot run and its score, which makes it easier to inspect exactly why a case failed.
What Langfuse calls the parts
In Langfuse, a dataset holds your test cases. A task calls your chatbot for each case. An evaluator checks the answer, and the result is stored as a score in an experiment run.
The workflow
Create a dataset → add question-and-answer expectations → run the chatbot → attach simple scores → compare the new run with the current run.
Interactive test / Langfuse
Try three support-chatbot cases in Langfuse
Save each customer question with the answer you expect.
Answer Northstar Shoes customers using only the return policy.
Can I return unworn shoes after 20 days?
Say yes and mention the 30-day window
Yes. Unworn shoes can be returned within 30 days of delivery.
- -Start with five understandable questions
- -Write the pass rule before running the chatbot
- -Keep the questions the same when comparing versions