How to Evaluate a Simple RAG Chatbot
Test whether a RAG chatbot found the right passage, answered from that passage, cited it, and admitted when the answer was missing.
Start with the simple chatbot evaluation guide if you have not created a small test set yet.
A RAG chatbot searches your documents before it answers. That gives you two separate things to test:
- Did search find the right passage?
- Did the chatbot use that passage correctly?
Keeping those questions separate is the difference between “the bot was wrong” and knowing what to fix.
We will keep using the Northstar Shoes support bot. Its knowledge base contains three short documents:
- Returns policy: unworn items, 30-day window, original shipping not refundable
- Damaged items guide: send damaged-item cases to a specialist
- International orders FAQ: duties and customs fees vary; contact support for order-specific help
Step 1: Save the evidence each question needs
Start with five cases:
| Question | Passage that must be found | A passing answer must |
|---|---|---|
| Can I return unworn shoes after 20 days? | Returns policy | Say yes, mention 30 days, cite the policy |
| Is original shipping refunded? | Returns policy | Say no, cite the policy |
| My shoes arrived damaged | Damaged items guide | Offer a specialist handoff |
| Are customs fees refunded? | International orders FAQ | Say it varies and offer support |
| Do you repair shoes? | None | Say the documents do not answer this |
The expected answer does not need to be a perfect paragraph. Save the facts that must appear, the claims that must not appear, and the source the chatbot should cite.
Step 2: Check retrieval first
For each question, record the passages returned by search.
Question: Is original shipping refunded?
Expected passage: Returns policy
Retrieved passages: Returns policy, Size guide, Store locations
Retrieval passes because the required passage is present. If the returns policy is missing, fix search or document indexing first. Rewriting the answer prompt cannot help the chatbot use evidence it never received.
For a beginner test, use one yes-or-no retrieval check:
Did the search results include the passage needed to answer the question?
Step 3: Check the answer second
When the right passage was retrieved, mark four checks:
- Correct: Does the answer match the passage?
- Supported: Can every important claim be found in the retrieved text?
- Cited: Does the answer point to the right document?
- Honest when missing: Does it avoid guessing when no passage answers the question?
Passing answer
“No. The original shipping fee is not refundable. [Returns policy]”
Failing answer
“No, but the store will give you a credit for the shipping fee.”
The second answer adds a store-credit promise that does not appear in the evidence. It fails even though its first word is correct.
Step 4: Use the failure to find the right fix
| What happened | Likely cause | First thing to fix |
|---|---|---|
| Right passage was not retrieved | Search or document problem | Indexing, query, or ranking |
| Right passage was retrieved, but answer was wrong | Answer-generation problem | Prompt or model behavior |
| Answer was right, but citation was wrong | Citation mapping problem | Source IDs or formatting |
| No passage answered, but bot guessed | Missing-answer behavior | Add an explicit “do not guess” rule |
This small table is the main reason to separate retrieval from answering. One overall quality score cannot tell you which part broke.
Step 5: Run one missing-evidence test
Remove the returns-policy passage for the shipping-fee question and run the chatbot again.
Pass: “I cannot find that answer in the available documents. I can connect you with support.”
Fail: The chatbot repeats a policy from memory or invents one.
This test proves the bot depends on your documents instead of merely producing a plausible answer.
Step 6: Compare one change on the same cases
Run the five cases with your current search settings. Then change only one thing, such as the number of passages returned, and run the same cases again.
| Version | Retrieval passed | Answers passed | Unsupported claims |
|---|---|---|---|
| Top 3 passages | 4/5 | 3/5 | 1 |
| Top 5 passages | 5/5 | 5/5 | 0 |
More passages are not always better. Extra irrelevant text can confuse the model, so inspect the rows instead of assuming a larger number wins.
Step 7: Add a multi-step agent only when needed
If the chatbot can also look up orders or change account data, keep the RAG checks above and add the tool checks from the tool-calling guide.
For a request such as “Find the return policy, check whether order 1042 qualifies, and start a return,” test each step:
- Did it retrieve the right policy?
- Did it look up the correct order?
- Did it apply the policy correctly?
- Did it ask before creating the return?
- Did the final reply match the actual result?
That is an agent evaluation in plain language. You are checking the steps and the outcome, not assigning one mysterious score to the whole run.
A simple release rule
For this RAG chatbot, release only when:
- every critical question retrieves the required passage
- every policy claim is supported by retrieved text
- every answer cites the right document
- every missing-answer case admits that the evidence is missing
- the new version does not break a case the current version passes
Use the interactive example below to see retrieval and answer failures separately. When the test set outgrows a spreadsheet, the Phoenix, Langfuse, and LangSmith courses show how to run the same loop in a platform.
Interactive test
Try three RAG chatbot test cases
Find the right Northstar Shoes document and answer only from that evidence.
Is original shipping refunded?
Find Returns policy; answer no; cite it
No. Original shipping fees are not refundable. [Returns policy]
Frequently asked questions
What should I evaluate in a RAG chatbot?
Check retrieval and answering separately: did search find the needed passage, and did the chatbot answer only from that evidence with the right citation?
What should a RAG chatbot do when the answer is not in its documents?
It should say the available documents do not answer the question and offer a useful next step instead of guessing.
Bhaulik Patel
Forward deployed AI engineer and creator of Deployed Engineer.