How to Evaluate a Simple AI Chatbot
A beginner-friendly guide to testing a support chatbot with five questions, simple pass-or-fail checks, and a repeatable spreadsheet workflow.
An evaluation is a saved test for your chatbot. You ask the same questions, check the answers using the same rules, and see whether a change made the bot better or worse.
You do not need a large dataset, a special platform, or another AI model to begin. Five good questions and a spreadsheet are enough.
This guide uses a fictional shop called Northstar Shoes. Its chatbot answers questions about this return policy:
- Unworn items can be returned within 30 days of delivery.
- Original shipping fees are not refundable.
- Damaged items must be sent to a support specialist.
- The policy says nothing about customs fees.
Step 1: Write five test questions
Start with questions real customers ask. Each row needs a question and a plain-English description of what a good answer must do.
| Customer question | A passing answer must | Why this case matters |
|---|---|---|
| Can I return unworn shoes after 20 days? | Say yes and mention the 30-day window | Common question |
| Can I return shoes after 45 days? | Say the standard window has passed | Boundary case |
| Is my original shipping fee refunded? | Say no | Exact policy fact |
| My shoes arrived damaged. What should I do? | Send the customer to a specialist | Required handoff |
| Will you refund my customs fee? | Say the policy does not say; do not guess | Missing information |
This table is your first test set. Evaluation tools may call it a dataset, but it is simply a saved list of examples.
Step 2: Use three simple checks
Read each answer and mark yes or no:
- Correct: Does it match the policy?
- No invention: Did it avoid promises or facts that are not in the policy?
- Useful next step: Did it tell the customer what to do next when needed?
Do not start with a vague score such as “helpful: 82%.” A yes-or-no result is easier to understand and easier to fix.
Example: a clear pass
Question: Can I return unworn shoes after 20 days?
Answer: “Yes. Unworn shoes can be returned within 30 days of delivery.”
- Correct: yes
- No invention: yes
- Useful next step: yes
- Result: pass
Example: a confident failure
Question: Will you refund my customs fee?
Answer: “Yes, customs fees are always refunded.”
- Correct: no—the policy does not answer this
- No invention: no—the chatbot made up a promise
- Useful next step: no—it should offer a handoff
- Result: fail
This is why fluency is not the same as quality. The failing answer sounds polished but could cost the business money.
Step 3: Run the test by hand
Create a spreadsheet with these columns:
| Question | Expected behavior | Chatbot answer | Correct? | No invention? | Useful next step? | Notes |
|---|---|---|---|---|---|---|
| Will you refund my customs fee? | Admit the policy is silent | “Yes, always.” | No | No | No | Invented refund promise |
Now run all five questions. A case passes only when every required check is yes.
Your first result might be 3 of 5 cases passed. That number is not the main lesson. The two failed rows tell you what to improve.
Step 4: Change one thing and run the same test again
Suppose the chatbot guessed about customs fees. Add one instruction:
If the policy does not answer the question, say you do not know and offer to connect the customer with support.
Run the same five questions again. Do not rewrite the questions or change the model at the same time. If the score moves from 3/5 to 5/5, you have useful evidence that the instruction helped.
That is the core evaluation loop:
- Save representative questions.
- Describe what a pass requires.
- Run the chatbot.
- Review the failed rows.
- Change one thing.
- Run the same questions again.
Step 5: Grow the test set from real mistakes
Five cases are enough to learn the process, not enough to prove the chatbot is ready for everyone. Add cases gradually:
- different ways of asking the same question
- follow-up questions such as “What if it was a gift?”
- unclear questions that require clarification
- requests the chatbot must hand to a person
- real failures from production, after removing private information
A practical first target is 15–25 carefully chosen cases. Quality matters more than volume.
When should you automate scoring?
Automate only after the manual rules feel stable.
- Use code for exact checks, such as whether an order number is present or a forbidden promise appears.
- Keep a human reviewer for nuanced cases at first.
- Use an AI judge only when the answer cannot be checked reliably with simple rules, and compare the judge with human decisions before trusting it.
The interactive test below lets you inspect three example cases. Treat each row as a conversation to understand, not just a score to maximize.
How the evaluation platforms fit
The six platforms covered on this site package the same basic loop in different ways:
- Phoenix runs a task and evaluators over a saved dataset, then compares experiments.
- Langfuse connects dataset items, experiment runs, traces, and scores.
- LangSmith provides a chatbot tutorial built around a golden dataset, metrics, and comparisons.
- Braintrust structures an evaluation as data, a task, and scorers.
- Helicone stores evaluation scores beside request, cost, and latency data.
- OpenAI Evals separates reusable test criteria from the model runs you compare.
Choose a platform when the spreadsheet becomes hard to manage. The platform does not decide what “good” means for your chatbot—you still have to write the questions and pass rules.
A simple release rule
For this support bot, use a rule anyone can explain:
- all return-policy facts must be correct
- the bot must never invent a refund or policy
- every damaged-item case must reach a specialist
- the new version must pass at least as many ordinary cases as the current version
If a new version breaks one of those rules, do not release it. Open the failed row, fix the cause, and test again.
Once this workflow feels natural, continue with tool-calling evaluations or RAG chatbot evaluations. The platform courses are there when you want to move the same test set into a dedicated tool.
Interactive test
Try three support-chatbot test cases
Answer Northstar Shoes customers using only the return policy.
Can I return unworn shoes after 20 days?
Say yes and mention the 30-day window
Yes. Unworn shoes can be returned within 30 days of delivery.
24-second evaluation loop
Turn chatbot failures into permanent tests
A production chatbot failure becomes a versioned test case, receives deterministic and semantic scores, and enters a repeatable release gate.
Read the visual transcript
- 01A support chatbot incorrectly promises that a customs fee is always refunded.
- 02Capture that production failure as one versioned dataset case with the expected policy behavior.
- 03Score hard rules with code and semantic quality with a calibrated judge or human review.
- 04Release on evidence, then make every meaningful production failure a permanent regression test.
Frequently asked questions
What is a chatbot evaluation?
A chatbot evaluation is a repeatable test: give the bot a saved set of questions, check each answer against clear pass rules, and compare the results before and after a change.
How many test questions do I need to start?
Start with five to ten useful questions. Include common requests, missing-information cases, and at least one mistake the chatbot has made before.
Do I need an evaluation platform or an LLM judge?
No. A spreadsheet and human review are enough for a first evaluation. Add automation only after your questions and pass rules are clear.
Bhaulik Patel
Forward deployed AI engineer and creator of Deployed Engineer.