Bias Benchmark Schema

How we will benchmark the decision engine

One page for review. The engine stays the only thing that writes to the page. Everything we measure comes from two readers that never write: one reads what the person says, one reads what the person sees. The vocabulary is fixed: a bias in what the person says; a debias shown when the question that answers it is on the page; then successful, ignored or dismissed. Nothing here is built yet. It is meant to be woven into the architecture clean-up, and this is the moment to say what is missing.

Glossary

The words used on this page and in every ticket from now on. Where an older ticket says "heuristic", read bias or debias; that word is retired.

WordMeaning
BiasA reasoning error in what the person says: a fact offered as proof that proves nothing, a figure read one way, an option weighed alone, evidence read for one side only, advisers who already lean, the same considerations circling, an extreme stretch taken as the norm.
DebiasThe question on the page that answers a bias. It exists for the person, not for us.
Debias shownThe debias is visible on the page, in any wording, within the window. The only event that counts as delivered. Reply prose never counts.
Successful / ignored / dismissedWhat the person did with a shown debias. Successful: they answered it and the line settled on their answer. Ignored: left open while they talked past it. Dismissed: pushed away. Counted separately from shown.
MisfireA debias shown where no bias was present: on a story, a figure of speech, a restatement, or a line already answered.
EngineThe one model call that writes the reply and edits the page. The only writer.
Debias instructionA per-bias text that can ride along in the engine's prompt. Optional; kept only for a bias the engine alone does not serve, and only after the paired replay shows it helps.
The pageWhat the person sees: the open questions, the settled lines, the "at stake" line. The only surface the benchmark reads.
DetectorA reader that lists the biases present in the person's message, given the page as it stood. It writes nothing. Its list is the expectation the scorecard checks, and the trigger for any debias instruction that is switched on.
Page-line readerA reader that looks only at the page after a turn and says, per bias, whether a debias is shown.
Outcome readerA reader that looks at the later turns and records successful, ignored or dismissed.
Labelled setAbout sixty moments labelled blind by a person: bias present or not, debias shown or not. Frozen and versioned. Every reader is scored against it before its numbers count.
GoldThe label the raters agreed on for a sample and a bias, formed by majority, tied to the bias's written definition. Readers and variants are scored against it.
Trust barA reader counts only when its agreement with the founders, the lower of its true-positive and true-negative rates on a frozen split, is at least 0.85 on at least 40 items per class. Below that it is uncalibrated and gives no verdict.
Noise floorThe same variant run twice against itself; how much two identical runs disagree. No verdict is given without it, and a Better must clear it.
Split / contestedA sample where raters disagree is a split; settled by a third rater or an adjudication with a reason, or it stays contested and out of gold.
Paired replayThe same frozen turn run twice, with and without one change, three draws each, scored by the same readers. The one experiment for every change.
WindowHow long after a bias is identified a debias may appear and still count as shown. Not limited for now: until the conversation concludes.
Prompt stampThe recorded identity of the prompt and model that ran a turn, so every number is tied to an engine version.

1. What happens on every turn

The person writes. The engine writes the reply and edits the page. The detector reads the message together with the page as it stood, and lists the biases it finds. That list does two jobs from one reading: it is the expectation the scorecard checks the page against, and it is the trigger for a debias instruction, for any bias where a fallback is switched on. Today that trigger is the detector's reading of the question, the page and the new message (since 2026-09-27 it no longer reads earlier turns); every bias that fits rides, with no per-turn cap and no "once per page" rule, by decision, and clogging is watched rather than prevented. Trigger accuracy and detection accuracy are therefore the same number. After the turn, the page-line reader reads only the page and says, for each bias, whether a debias is now visible. On later turns, the outcome reader says what the person did with it. The reply prose is never scored.

Person's message this turn Page as it stood the growing evidence Engine the only writer one prompt, one model Reply prose never scored The page after the turn what the person sees Debias instruction fires only when the detector finds its bias Detector lists the biases present Page-line reader debias shown? per bias Outcome reader successful / ignored / dismissed Scorecard per bias, per turn identified / shown / outcome writes reads reply page edits rides only where shown-rate is low reads the message and the page before reads the page only reads the next turns shown? identified fires it, where a fallback is on
Figure 1. Per turn. One detector reading serves as both the expectation and the trigger; the readers are the whole measurement; the engine is not asked to report what it did; the dashed instruction is the only optional part of the system.
the page, the only surface that countsreaders, rounded: they never writedashed: optional, added per bias only when measured as needed

2. What counts as a debias shown

CountsDoes not count
A new, revised or settled line on the page that carries the debias, in any wording, any time after the bias was identified and before the conversation concludes. Settled lines count (answered Q1). The "at stake" line is taken to count for the same reason; say if not.The same point made in the reply prose and absent from the page.
The line shown once and then answered: it must not be asked again.A debias re-asked after the person answered it, or after the line already stood on the page. That is a misfire.
Toward the shown rate: a debias that appears where the labels say a bias was present.Toward the misfire rate instead: a debias that appears where no bias was present, on a story, a figure of speech, an emotional spill or a restatement. Same event, different column, decided by whether the bias was there.
Answered. A settled line counts (Dominique). The window is not limited for now: a debias counts whenever it appears before the conversation concludes; a re-ask after the answer is still a misfire.

3. The biases in scope, and the debias each one calls for

Bias in what the person saysDebias the page should show
A fact offered as proof that would be just as true under either answer (diagnosticity)What would tell the two answers apart
A figure read one way only (framing)The same figure read the other way
One option weighed with no alternative on the table (single-option evaluation)What they would do if this option vanished
A fact read as evidence for one side only (one-sided evaluation)Would they read it the same way from the other side
Advisers who already lean, or one source carrying the whole read (biased sample)Who could give a straight read before knowing their lean
The same considerations circling with nothing settling (deferral, status quo)What someone taking over tomorrow would do
An extreme recent result taken as the new normal (regression neglect)How much of it a typical stretch would have produced

The list is right for now and will change substantially with the wiki (Dominique). Base rate, the devil's advocate and Franklin are out of scope for this benchmark for now, by decision. Each bias's written definition is the contract the labels encode; a reader that infers extra conditions is measuring its own judgement.

4. Who checks the checkers

Every reader is a model, and a model reader counts only after it agrees with a human-labelled set. The set is small, blind, frozen and versioned, and it grows slowly with real conversations. When a reader's model changes, it is scored against the set again before its numbers count. This is the step every earlier benchmark skipped.

Labelled set about 60 moments, labelled blind by Dominique, frozen, versioned bias present? debias shown? Rated real conversations admin import, person removed Detector scored: precision, recall Page-line reader scored: agreement per bias Quality judge rubric from the gold and no-go rulings A reader's model changes or its prompt, or its criteria its numbers stop counting until re-scored on the set validates validates seeds grows the set slowly re-validate
Figure 2. Ground truth flows one way: from the human-labelled set to the readers. Nothing a reader produces feeds back into the set.
human ground truthreaders, roundeddashed: the re-validation trigger
Answered. Labellers: Dominique, Wilfried and Stef for now. Agreement follows the rules already built into the app's benchmark (admin, /admin/benchmarks; raters at /benchmarks), specified in docs/spec/SPEC-benchmarks.md and grounded in Wilfried's research note Benchmark flow research: how the smart teams evaluate judges and raters. The rules that carry over to this schema:

The readers are judges, and the research note already sets their trust bar. This part is specified, not built: the spec calls it "the judge path" and leaves it for after the picker. What it says, and what applies to the page-line reader and the quality judge:

For the three readers of this schema: the detector is the picker, already the benchmark's built subject, scored by precision with recall as its guard. The page-line reader and the quality judge are the judge path above. The outcome reader may need no model at all: whether a line settled on the person's answer is in the page versions.

5. The one experiment we run for every change

To know whether a debias instruction, a page rule, a prompt edit or a model swap changes anything, we replay the same frozen turns twice, with and without the change, three draws each, and let the readers score both. Criteria are written before the run. The difference per bias is the result. The same shape answers "should we retire this instruction" and "did this prompt edit help".

Frozen turn page before + message, rebuilt from the record prompt stamp, model named Arm A: engine alone 3 draws Arm B: engine + the change 3 draws the one thing that differs Page after, A read by the same readers Page after, B read by the same readers Difference shown-rate per bias misfires per bias criteria set before same turnsame turn
Figure 3. The paired replay. Same turns, same readers, three draws per arm; the only difference between the arms is the change under test, so the difference in the scorecard belongs to it.

Decision rule per bias. If arm A already shows the debias at the rate we want, the instruction is retired and the bias is served by the engine alone. If not, we add the cheapest thing that could raise it, a page rule first, a fallback instruction second, and keep it only if arm B beats arm A at three draws. The rate we want is an open question below.

How a verdict is read. One primary metric fixed before the run. Per sample, whether B beat A, A beat B, or both were the same; an exact sign test on the samples that differ. Better when the sign test favours B at p below 0.05 (8 of 8, 9 of 10, 12 of 15 or 15 of 20 differing samples). Worse is the mirror. Not detectably worse is the verdict for a cost-saving swap, and applies only when the stated non-inferiority margin is larger than the swing the benchmark can detect. Otherwise no detectable difference, with the swing size said. A candidate is ready to ship after two independent Better verdicts, or one Better whose gap is at least twice the A-versus-A noise floor. Adding samples until a result turns significant is refused; a larger suite is a new benchmark.

What the benchmark can see. With 36 paired samples and 30% of them differing, the smallest net difference it can detect at 80% power is about 25 points. At 56 samples about 20; at 100 about 15. Two identical picker runs today agree on 100 of 155 cases, so each variant runs each sample 3 times and takes the majority; that is the "three draws" of this schema.

What this means for retiring an instruction. "Retire" is a cost-saving swap: engine alone in place of engine plus instruction. Its verdict is therefore "not detectably worse", which needs a non-inferiority margin stated before the run, and a suite large enough that the margin exceeds the detectable swing. With today's 36 openings that margin cannot be tighter than about 25 points. So Q4 below is not a shown-rate threshold; it is the margin we accept, and the suite size that margin requires.

The statistics already exist in the app's benchmark and carry over unchanged. A variant's answer on a sample is the majority of its k runs; self-agreement across the runs is reported. No verdict is given while a variant has more than 5% of its calls in error, while no A-versus-A noise floor exists on the same dataset version, or while a split waits for a person. The comparison is a paired sign test per opening with Wilson intervals, per-bias rows flagged only after Holm correction; time and cost sit beside the verdict, never inside it. "Compared with last time" re-scores both benchmarks against today's gold with no model call. Every number is held against a hand-worked toy benchmark and a committed golden benchmark inside npm test. What this schema adds is only the subject: today the benchmark's one subject is the picker; the page-line reader is its second subject, and the spec already reserves the place for it under the name "the judge path".

6. What the scorecard holds, per bias, per engine version, per model

NumberWhat it saysWhere it comes from
DetectionPrecision and recall of "bias present" on the person's messageDetector against the labelled set
Debias shownOf the moments where a bias was present, how often the page carried the debias within the windowPage-line reader
MisfireDebiases shown on a story, a restatement or an already-answered linePage-line reader + detector
QualityPlain question, names the other answer, asks for an observationQuality judge, rubric from the rulings
PersistenceThe line stayed until answered, not silently settled awayPage versions
OutcomeSuccessful, ignored or dismissed, counted separately from shownOutcome reader on later turns
Cost, latencyPer turn, per modelThe call log
Later: conclusionsConclusion quality with the debias shown versus not; where priority between biases will come fromRated conversations, once there are enough

7. The same reader in production

The page-line reader can run on production pages and report counts only. No text leaves. That gives a quality signal for a beta where nobody can read the conversations, and it is the first time production would tell us what people actually saw rather than what rode in a prompt.

Production page, each turn stays where it is Page-line reader runs inside the boundary Quality dashboard shown, misfire, outcome, per bias readscounts only, no textprivacy boundary
Figure 4. Production. Only counts cross the boundary.

8. Where this touches the app today

These are the pieces the architecture clean-up will meet. Each becomes part of the benchmark or is deleted on purpose; none should disappear by accident.

Piece in the app nowRole in the benchmarkWhat the clean-up must keep
The picker and its evidence spansBecomes the detector; its output is an expectationThe per-bias definitions and the evidence quote
The firings table and its outcome columnBecomes the outcome record: shown, successful, ignored, dismissedThe link from a line to the person's later turns
Page versions with the sweep turn idThe page-line reader and the replay both rebuild "page before" from itOne version per user turn, written before the edits
The "asked / declined" report from the engineNone; attribution stops mattering once the reader existsCan go
The eval-only prompt arms and leversOne paired-replay path, plus the plain-tail lever for models that refuse the thinking clauseOne path, not five
The heuristic column on page lines, the registry, the iconsThe labeller stamps the bias a line embodies; the icon is what the person seesMust survive even if every instruction is retired
The prompt stamp on turns; the rated-conversation importPer-model, per-version runs; corpus growth with real peopleBoth, unchanged
The benchmark app: Admin › Benchmarks (/admin/benchmarks), Admin › Datasets (/admin/datasets), raters promoted in Admin › Users, the raters' own screen at /benchmarks; seats, gold, adjudication, verdict statisticsThe labelling and verdict machinery for every reader in this schema; the "judge path" the spec reserves is where the page-line reader landsAll of it; extend the subject from the picker to page lines, do not build a second benchmark

What we will not do, and why

Practices the research note reviewed and rejected for a three-person panel. Listed so nobody re-proposes them.

9. Risks we already know

  1. Ground truth is the bottleneck, and it is one person. Synthetic conversations with a simulated user drift from real people. The set stays small and grows only through rated real conversations. The risk is tuning the engine to 36 openings.
  2. An unvalidated reader is worse than none. The last table we produced scored the formal wording, missed the engine's own phrasing, and inverted the result. Page lines are short and well-formed, which is why reading them is more tractable than reading reply prose.
  3. One draw is noise. The same prompt gave 27 and then 33 settled lines on one bench. Three draws, paired, criteria written first, every time.
  4. Production is blind without the reader. The firings table says what rode, not what the person saw.
  5. Windows and memory, not single turns. A debias shown one turn late is a success; one re-asked after the answer is a failure.
  6. Model swaps are their own axis. Fable refused one clause of the production prompt on every turn after the first. The benchmark should find the next one on purpose, per model, with the prompt stamp recorded.

10. Open questions for this review

Q1, answered. A settled line counts. The "at stake" line is assumed to count as well.
Q2, answered. No limit for now.
Q3, answered. Dominique, Wilfried and Stef label; agreement per Wilfried's existing benchmark rules.
Q4, restated. Retiring an instruction is a "not detectably worse" verdict, so the number to set before the run is the non-inferiority margin per bias, and the suite size it needs. At 36 openings the benchmark cannot see a swing under about 25 points; a tighter margin means a bigger suite. Which margin, and which suite?
Q5, answered. The seven are right for now; the list will evolve with the wiki. Base rate stays out for now.
Q6. Are counts-only from production acceptable under our privacy stance, and who sees the dashboard?
Q7, first answer. The trigger was missing: what fires a debias instruction. Now drawn: the detector fires it, from the same reading that sets the expectation. Still open to more.
Q8, for Wilfried. The research note compares versions with a sign test on the (sample, bias) pairs that changed; the spec (INV-bench-15) pairs per opening, never per (opening, bias), with per-bias rows flagged only after Holm. For debias shown, is the unit the opening or the (moment, bias)? The schema's scorecard is per bias, so the answer decides how a per-bias verdict is ever reached.

Status: schema for review. The detector, the page-line reader, the record and the audit view are built (GOD-1650, on main 2026-09-27); the labelled set, the trust bar and the paired replay are next (GOD-1671).2026-09-28