Back to Blog
    Chatbot Evals

    How to Evaluate a Simple AI Chatbot

    A beginner-friendly guide to testing a support chatbot with five questions, simple pass-or-fail checks, and a repeatable spreadsheet workflow.

    Bhaulik Patel·Sep 7, 2026·6 min read

    An evaluation is a saved test for your chatbot. You ask the same questions, check the answers using the same rules, and see whether a change made the bot better or worse.

    You do not need a large dataset, a special platform, or another AI model to begin. Five good questions and a spreadsheet are enough.

    This guide uses a fictional shop called Northstar Shoes. Its chatbot answers questions about this return policy:

    • Unworn items can be returned within 30 days of delivery.
    • Original shipping fees are not refundable.
    • Damaged items must be sent to a support specialist.
    • The policy says nothing about customs fees.

    Step 1: Write five test questions

    Start with questions real customers ask. Each row needs a question and a plain-English description of what a good answer must do.

    Customer questionA passing answer mustWhy this case matters
    Can I return unworn shoes after 20 days?Say yes and mention the 30-day windowCommon question
    Can I return shoes after 45 days?Say the standard window has passedBoundary case
    Is my original shipping fee refunded?Say noExact policy fact
    My shoes arrived damaged. What should I do?Send the customer to a specialistRequired handoff
    Will you refund my customs fee?Say the policy does not say; do not guessMissing information

    This table is your first test set. Evaluation tools may call it a dataset, but it is simply a saved list of examples.

    Step 2: Use three simple checks

    Read each answer and mark yes or no:

    1. Correct: Does it match the policy?
    2. No invention: Did it avoid promises or facts that are not in the policy?
    3. Useful next step: Did it tell the customer what to do next when needed?

    Do not start with a vague score such as “helpful: 82%.” A yes-or-no result is easier to understand and easier to fix.

    Example: a clear pass

    Question: Can I return unworn shoes after 20 days?

    Answer: “Yes. Unworn shoes can be returned within 30 days of delivery.”

    • Correct: yes
    • No invention: yes
    • Useful next step: yes
    • Result: pass

    Example: a confident failure

    Question: Will you refund my customs fee?

    Answer: “Yes, customs fees are always refunded.”

    • Correct: no—the policy does not answer this
    • No invention: no—the chatbot made up a promise
    • Useful next step: no—it should offer a handoff
    • Result: fail

    This is why fluency is not the same as quality. The failing answer sounds polished but could cost the business money.

    Step 3: Run the test by hand

    Create a spreadsheet with these columns:

    QuestionExpected behaviorChatbot answerCorrect?No invention?Useful next step?Notes
    Will you refund my customs fee?Admit the policy is silent“Yes, always.”NoNoNoInvented refund promise

    Now run all five questions. A case passes only when every required check is yes.

    Your first result might be 3 of 5 cases passed. That number is not the main lesson. The two failed rows tell you what to improve.

    Step 4: Change one thing and run the same test again

    Suppose the chatbot guessed about customs fees. Add one instruction:

    If the policy does not answer the question, say you do not know and offer to connect the customer with support.

    Run the same five questions again. Do not rewrite the questions or change the model at the same time. If the score moves from 3/5 to 5/5, you have useful evidence that the instruction helped.

    That is the core evaluation loop:

    1. Save representative questions.
    2. Describe what a pass requires.
    3. Run the chatbot.
    4. Review the failed rows.
    5. Change one thing.
    6. Run the same questions again.

    Step 5: Grow the test set from real mistakes

    Five cases are enough to learn the process, not enough to prove the chatbot is ready for everyone. Add cases gradually:

    • different ways of asking the same question
    • follow-up questions such as “What if it was a gift?”
    • unclear questions that require clarification
    • requests the chatbot must hand to a person
    • real failures from production, after removing private information

    A practical first target is 15–25 carefully chosen cases. Quality matters more than volume.

    When should you automate scoring?

    Automate only after the manual rules feel stable.

    • Use code for exact checks, such as whether an order number is present or a forbidden promise appears.
    • Keep a human reviewer for nuanced cases at first.
    • Use an AI judge only when the answer cannot be checked reliably with simple rules, and compare the judge with human decisions before trusting it.

    The interactive test below lets you inspect three example cases. Treat each row as a conversation to understand, not just a score to maximize.

    How the evaluation platforms fit

    The six platforms covered on this site package the same basic loop in different ways:

    • Phoenix runs a task and evaluators over a saved dataset, then compares experiments.
    • Langfuse connects dataset items, experiment runs, traces, and scores.
    • LangSmith provides a chatbot tutorial built around a golden dataset, metrics, and comparisons.
    • Braintrust structures an evaluation as data, a task, and scorers.
    • Helicone stores evaluation scores beside request, cost, and latency data.
    • OpenAI Evals separates reusable test criteria from the model runs you compare.

    Choose a platform when the spreadsheet becomes hard to manage. The platform does not decide what “good” means for your chatbot—you still have to write the questions and pass rules.

    A simple release rule

    For this support bot, use a rule anyone can explain:

    • all return-policy facts must be correct
    • the bot must never invent a refund or policy
    • every damaged-item case must reach a specialist
    • the new version must pass at least as many ordinary cases as the current version

    If a new version breaks one of those rules, do not release it. Open the failed row, fix the cause, and test again.

    Once this workflow feels natural, continue with tool-calling evaluations or RAG chatbot evaluations. The platform courses are there when you want to move the same test set into a dedicated tool.

    Interactive test

    Try three support-chatbot test cases

    Answer Northstar Shoes customers using only the return policy.

    Input

    Can I return unworn shoes after 20 days?

    Expected behavior

    Say yes and mention the 30-day window

    Observed output

    Yes. Unworn shoes can be returned within 30 days of delivery.

    correct answerPass
    no made-up policyPass
    useful next stepPass
    Case passed3/3 checks

    24-second evaluation loop

    Turn chatbot failures into permanent tests

    A production chatbot failure becomes a versioned test case, receives deterministic and semantic scores, and enters a repeatable release gate.

    Read the visual transcript
    1. 01A support chatbot incorrectly promises that a customs fee is always refunded.
    2. 02Capture that production failure as one versioned dataset case with the expected policy behavior.
    3. 03Score hard rules with code and semantic quality with a calibrated judge or human review.
    4. 04Release on evidence, then make every meaningful production failure a permanent regression test.
    Direct answers

    Frequently asked questions

    What is a chatbot evaluation?

    A chatbot evaluation is a repeatable test: give the bot a saved set of questions, check each answer against clear pass rules, and compare the results before and after a change.

    How many test questions do I need to start?

    Start with five to ten useful questions. Include common requests, missing-information cases, and at least one mistake the chatbot has made before.

    Do I need an evaluation platform or an LLM judge?

    No. A spreadsheet and human review are enough for a first evaluation. Add automation only after your questions and pass rules are clear.

    Chatbot EvalsLLM EvaluationAI Quality
    Share
    BP

    Bhaulik Patel

    Forward deployed AI engineer and creator of Deployed Engineer.