Tutorials
    Chapter 1 of 3

    Your First Helicone Evaluation

    24-second evaluation loop

    Turn chatbot failures into permanent tests

    A production chatbot failure becomes a versioned test case, receives deterministic and semantic scores, and enters a repeatable release gate.

    Read the visual transcript
    1. 01A support chatbot incorrectly promises that a customs fee is always refunded.
    2. 02Capture that production failure as one versioned dataset case with the expected policy behavior.
    3. 03Score hard rules with code and semantic quality with a calibrated judge or human review.
    4. 04Release on evidence, then make every meaningful production failure a permanent regression test.

    Begin with one small support bot

    Imagine a shoe-store chatbot with four rules: unworn items can be returned within 30 days, original shipping is not refundable, damaged items go to a specialist, and unanswered policy questions must not be guessed.

    An evaluation asks the bot the same five questions before and after a change. Each answer gets a clear pass or fail. That is all you need to understand before using Helicone.

    Helicone is most useful after requests are already being logged. It lets you attach your own pass-or-fail results to those requests and review trends.

    What Helicone calls the parts

    Helicone stores evaluation results, but your team still decides and computes what passes. Start with boolean scores such as policy_correct and safe_handoff, then attach them to each request.

    The workflow

    Log chatbot requests → choose five useful cases → score them manually or in your own code → report the scores → compare quality with cost and latency.

    The five-step evaluation loop
    1
    Save
    Write five real customer questions and what a passing answer must say.
    2
    Run
    Ask the current chatbot all five questions and save its answers.
    3
    Check
    Mark each answer correct, invented, or missing a required handoff.
    4
    Change
    Change one prompt, model, or retrieval setting—not several at once.
    5
    Compare
    Run the same five questions and inspect every case that became worse.

    Interactive test / Helicone

    Try three support-chatbot cases in Helicone

    Choose a chatbot request you want to review.

    Answer Northstar Shoes customers using only the return policy.

    Input

    Can I return unworn shoes after 20 days?

    Expected behavior

    Say yes and mention the 30-day window

    Observed output

    Yes. Unworn shoes can be returned within 30 days of delivery.

    correct answerPass
    no made-up policyPass
    useful next stepPass
    Case passed3/3 checks
    Key Takeaways
    • -Start with five understandable questions
    • -Write the pass rule before running the chatbot
    • -Keep the questions the same when comparing versions
    Knowledge Check

    How should you compare the current chatbot with a new version?