All book labs

    THE COST OF USEFUL AI / LAB 06

    What does the score hide?

    Fictional teaching examples. Calculations stay in your browser. No provider calls or input uploads.

    Fictional examples. Human reference labels are resolved judgements under the rubric, not infallible truth. Intervals assume independent representative Bernoulli observations from a stable population; they do not correct sampling bias, clustered cases or grading errors. The zero-event bound applies only when zero failures were observed.

    Inspect the grader

    Enter the four counts from a reference-labelled audit. Accept is the positive decision.

    Keep reviewer disagreement

    Agreement can look high while the disputed cases contain the most useful rubric feedback.

    Open the average by slice

    The sample has more routine work than risk-sensitive work. Production weights let you ask whether that mix matches the real queue.

    Relative whole-number weight

    Relative whole-number weight

    Put a range beside a rate

    This is a separate sample of task acceptance, not a confidence interval for the grader matrix above.

    When no failures appeared

    A third, separate sample with exactly zero failures of the specified kind.

    YOUR GRADER AUDIT

    Overall agreement
    85%
    All 200 audited outputs
    False-accept rate
    40%
    Among 50 reference failures
    Release precision
    87.5%
    Among 160 outputs the grader accepts
    False-reject rate
    6.67%
    Among 150 reference passes

    Reviewer calibration

    84% agreement

    Keep all 8 disagreements with both labels and the resolution reason.

    Slice view

    Routine

    98%

    Risk-sensitive

    70%

    Raw sample average: 92.4%

    Production-weighted average: 95.2%

    Weights are normalized automatically. They currently total 100. A weighted average still cannot excuse a failed mandatory slice.

    Separate task sample

    90% accepted

    Two-sided 95% Wilson interval: 82.56% to 94.48%

    The 95% level describes this method's long-run coverage under its assumptions. It is not a guarantee that 95% of future tasks will pass.

    Separate zero-failure sample

    0.99%

    Exact one-sided 95% upper bound on the failure probability. Zero observed failures does not establish zero risk.

    Try the printed exercise

    The default grader agrees on 85% of cases while accepting 40% of reference failures. The slice view shows why 92.4% across the sample does not describe the 70% risk-sensitive result.

    Interval methods: NIST/SEMATECH e-Handbook. Neither this lab nor a single score supplies a universal release threshold.