Benchmarks · 30 September 2026 · proposal, not yet filed

Detector bias-by-bias plan

Rebuild the detector one bias at a time: switch every bias off, add one back, tune it until it clears its bars, then add the next. The conversations are generated once, up front; the labels are made one bias at a time, each only once.

Words used here

Moment
One message of the person in a conversation; the opening is one too. A conversation of 4 to 9 turns gives 4 to 9 moments.
Gold
What two raters agree a moment shows, for one bias: fits, doesn't fit, or can't tell. It is about the message, never about the detector, so tuning the detector never needs new gold. Only rewriting a bias's own definition does, for that bias alone.
Sensitivity
Of the moments where the bias fits, the share where the detector fires.
Specificity
Of the moments where the bias does not fit, the share where the detector stays silent.
Precision
Of the detector’s fires, the share where raters agree the bias fits. It fires on 50 moments, raters agree on 30: precision 60%. It is what a person feels in use. (Not in the glossary yet: the spec and the code use it.)
Tuning and checking slices
The dataset split in two, by conversation, once and for good. We tune on the tuning slice as often as we like. The checking slice is used only to confirm a bias, and every look is counted.
Calibration moments
About 7% of each version's new moments, drawn at random at each Promote, the same share in each slice (about 140 in all), labelled for every bias. They show how the detector behaves on ordinary moments.
Validated conversation
A generated conversation an admin has read and accepted into the draft. Reading is not labelling.

What was decided

The dataset

PartConversationsMoments
With a bias: 12 biases × 48 conversations, about 2 biases per scenario≈ 274≈ 1,640
No bias (the pilot's 40 included)≈ 60≈ 360
Validated in all≈ 330≈ 1,980

About 370 scenarios requested to end with about 330 validated (some refused at review; about 385 conversations played, some replayed after rejection). Model cost ≈ $145: about $0.06 a scenario, $0.32 a conversation.

Step by step

  1. Generate

    Pilot: 40 scenarios with no bias

    Ask for 40 with No bias only, French a third. Review, generate, read and promote them. They stay in the dataset and count toward the ~60 with no bias. We learn the share refused and rejected, the real turns and cost, and whether the scripted steps read naturally; rated for the first biases, they show how often a bias appears on its own.

    ≈ $15 · about 1½ hours of reading

  2. Generate

    The full set

    Request the rest in batches of up to 100, with Fill the gaps, until the balance says every bias has its 48 conversations. Read the first 20 before asking for the rest: they show whether scenarios keep to 1 to 4 biases. Review each scenario, read each conversation, validate or reject, add the validated ones to the draft.

    ≈ $130 · about 15 hours of reading, shared between admins

  3. Generate

    Promote, with the split

    Promote shows, before confirming, each bias's count in each slice and each language, and the calibration moments: the same share of each slice's new moments. After Promote the split never changes.

    $0 · minutes

  4. Bias 1

    Label bias 1

    A benchmark scored by people, bias 1 only, on bias 1's 48 conversations plus the calibration moments: about 430 moments, each answered by two raters.

    $0 · about 6½ hours of rater time in all

  5. Bias 1

    Check the labels

    Per slice: at least 40 moments where bias 1 fits in the checking slice, and raters agreeing well enough (the trust bar's 0.60). Too few fits: Fill the gaps for bias 1 and label only the new moments. Low agreement: the definition is unclear; it is rewritten in the code with an agent, and bias 1 is labelled again.

    $0

  6. Bias 1

    Tune bias 1

    The detector with only bias 1 on. The baseline run twice once, to know the noise. Then each round: a new detector prompt is written in the code with an agent and shipped, then run against gold on the tuning slice. Repeat until it clears both bars.

    ≈ $0.30 and 8 minutes a round; the real pace is the change to the prompt

  7. Bias 1

    Confirm bias 1

    One run on the checking slice, the only time the detector sees it. If it passes, a look at a few dozen real moments from Audit as a sanity check, then bias 1 can go live. If it fails, back to tuning; the screen counts looks at the checking slice, and after three the honest fix is fresh checking conversations.

    ≈ $0.35

  8. Bias 2

    Label bias 2

    A benchmark scored by people for biases 1 and 2, on bias 2's conversations plus the calibration moments. Moments that already have bias 1's gold are skipped for bias 1, so raters answer bias 1 only on bias 2's conversations, where bias 1's false alarms matter most.

    $0 · about 8 hours of rater time in all

  9. Bias 2

    Check, tune and confirm bias 2

    As steps 5 to 7, with biases 1 and 2 both on. Bias 2 must clear its bars and bias 1 must stay above its own; a round that helps bias 2 but drops bias 1 is rejected. Confirming bias 2 also checks bias 1 again on its checking slice, and counts as a look for both.

    ≈ $0.60 a round

  10. Next biases

    Steps 8 and 9 again, for each bias

    No new generation: the conversations exist. A bias short of fits gets a top-up before its turn.

    about 8 hours of rater time and ≈ $10 of detector runs per bias

Where it happens: the Plan tab

A Plan tab in Admin › Benchmarks shows one row per bias: its step, where it is, its sensitivity and specificity, and one next action. Each action's tooltip gives its count and cost before you click. The Scenarios board stays as it is, with the balance table folded away.

ActionWhen
Fill the gapsbefore Promote, a bias has fewer than 48 conversations
Promoteevery bias has its 48
Labelthe bias has no gold yet
Top upfewer than 40 moments where it fits in the checking slice
Tunelabels checked; runs on the tuning slice
Confirmthe tuning slice clears the bars; runs on the checking slice
Real checkconfirmed; a few dozen real moments from Audit
Resultonce a result exists

Done outside the screen: a definition and the detector's prompt are rewritten in the code, with an agent; switching a bias on stays with the technique flags. Datasets keeps Browse items for a draft or a version, with a new Slice filter.

The bars

Goal: a precision of at least 75% in real use, so at most one fire in four is wrong.

BarValueMeasured on
Sensitivity85% or moreat least 40 moments where the bias fits
Specificityset per bias from how common it is, so precision reaches 75%every labelled moment where the bias doesn't fit
Raters agreetrust bar's 0.60the bias's labelled moments

Why specificity moves with the bias: most messages don't show a given bias, so a small share of wrong fires outnumbers the right ones.

Bias shows in…Specificity needed for 75% precision
15% of messages95%
5%98.5%
2%99.4%

How common each bias really is comes from real moments (Audit), not from this dataset, where we chose how often each bias appears. A rare bias whose specificity stays too low is better left off. Precision as measured on this dataset flatters the detector and is shown as "not comparable with real use".

Aiming higher: what each tier costs

A bar is only as good as the number of moments behind it. Higher bars need more labelled moments, and the cost grows fast. Two separate dials:

Specificity: more moments where the bias doesn't fit

The conversations already exist: the checking slice holds about 1,200 moments. Going up a tier means raters label the bias on more of them. It costs rater time, no generation.

Specificity to proveMoments where it doesn't fitExtra rater time per biasFits
95% (common bias)≈ 200none, in the planthe base plan
98.5% (bias in 5% of messages)≈ 500+ 4 hlabel the bias on part of the other conversations
99.4% (bias in 2%)≈ 1,200+ 14 hlabel the bias on the whole checking slice
99.8% (90% precision on a rare bias)≈ 3,000+ 40 h, and ≈ 300 more conversations (≈ $100, 10 h reading)out of reach for now

Rule of thumb: to show that wrong fires stay under 1 in N, you need about 3 × N moments where the bias doesn't fit, and more if any wrong fire occurs.

Sensitivity: more moments where the bias fits

These come only from conversations written with the bias, so going up a tier means generating, reading and labelling more of them. Two biases per scenario halves the generation.

How sureFits in the checking sliceConversations per biasFor all 12 biases
Trust bar: 85% measured on 40, true value within about ±11 points4028the base plan
Within about ±7 points10067+ 230 conversations: ≈ $80, 8 h reading, + 3½ h labelling per bias
Within about ±5 points200135+ 640 conversations: ≈ $225, 21 h reading, + 9 h labelling per bias

In short: the base plan proves the trust bar and a 75% precision for common biases. A bias in about 5% of messages needs the next specificity tier (+4 hours of labelling). Beyond that, labelling time grows about three times with each step, and generated data stops being the limit: real moments are.

Raters

First two biasesIn allPer person (÷3)
Reviewing scenarios and reading conversations≈ 17 h≈ 5½ h
Labelling bias 1≈ 6½ h≈ 2¼ h
Labelling bias 2≈ 8 h≈ 2¾ h
Total≈ 31 h≈ 10½ h

Still to decide

Biggest unknown: the plan assumes a bias shows in about 1.5 moments of its conversation. Step 5 measures it on bias 1. If it is closer to 1, every bias needs about a third more conversations; each costs about $0.35 and two minutes of reading.