One page for review. The engine stays the only thing that writes to the page. Everything we measure comes from two readers that never write: one reads what the person says, one reads what the person sees. The vocabulary is fixed: a bias in what the person says; a debias shown when the question that answers it is on the page; then successful, ignored or dismissed. Nothing here is built yet. It is meant to be woven into the architecture clean-up, and this is the moment to say what is missing.
The words used on this page and in every ticket from now on. Where an older ticket says "heuristic", read bias or debias; that word is retired.
| Word | Meaning |
|---|---|
| Bias | A reasoning error in what the person says: a fact offered as proof that proves nothing, a figure read one way, an option weighed alone, evidence read for one side only, advisers who already lean, the same considerations circling, an extreme stretch taken as the norm. |
| Debias | The question on the page that answers a bias. It exists for the person, not for us. |
| Debias shown | The debias is visible on the page, in any wording, within the window. The only event that counts as delivered. Reply prose never counts. |
| Successful / ignored / dismissed | What the person did with a shown debias. Successful: they answered it and the line settled on their answer. Ignored: left open while they talked past it. Dismissed: pushed away. Counted separately from shown. |
| Misfire | A debias shown where no bias was present: on a story, a figure of speech, a restatement, or a line already answered. |
| Engine | The one model call that writes the reply and edits the page. The only writer. |
| Debias instruction | A per-bias text that can ride along in the engine's prompt. Optional; kept only for a bias the engine alone does not serve, and only after the paired replay shows it helps. |
| The page | What the person sees: the open questions, the settled lines, the "at stake" line. The only surface the benchmark reads. |
| Detector | A reader that lists the biases present in the person's message, given the page as it stood. It writes nothing. Its list is the expectation the scorecard checks, and the trigger for any debias instruction that is switched on. |
| Page-line reader | A reader that looks only at the page after a turn and says, per bias, whether a debias is shown. |
| Outcome reader | A reader that looks at the later turns and records successful, ignored or dismissed. |
| Labelled set | About sixty moments labelled blind by a person: bias present or not, debias shown or not. Frozen and versioned. Every reader is scored against it before its numbers count. |
| Gold | The label the raters agreed on for a sample and a bias, formed by majority, tied to the bias's written definition. Readers and variants are scored against it. |
| Trust bar | A reader counts only when its agreement with the founders, the lower of its true-positive and true-negative rates on a frozen split, is at least 0.85 on at least 40 items per class. Below that it is uncalibrated and gives no verdict. |
| Noise floor | The same variant run twice against itself; how much two identical runs disagree. No verdict is given without it, and a Better must clear it. |
| Split / contested | A sample where raters disagree is a split; settled by a third rater or an adjudication with a reason, or it stays contested and out of gold. |
| Paired replay | The same frozen turn run twice, with and without one change, three draws each, scored by the same readers. The one experiment for every change. |
| Window | How long after a bias is identified a debias may appear and still count as shown. Not limited for now: until the conversation concludes. |
| Prompt stamp | The recorded identity of the prompt and model that ran a turn, so every number is tied to an engine version. |
The person writes. The engine writes the reply and edits the page. The detector reads the message together with the page as it stood, and lists the biases it finds. That list does two jobs from one reading: it is the expectation the scorecard checks the page against, and it is the trigger for a debias instruction, for any bias where a fallback is switched on. Today that trigger is the detector's reading of the question, the page and the new message (since 2026-09-27 it no longer reads earlier turns); every bias that fits rides, with no per-turn cap and no "once per page" rule, by decision, and clogging is watched rather than prevented. Trigger accuracy and detection accuracy are therefore the same number. After the turn, the page-line reader reads only the page and says, for each bias, whether a debias is now visible. On later turns, the outcome reader says what the person did with it. The reply prose is never scored.
| Counts | Does not count |
|---|---|
| A new, revised or settled line on the page that carries the debias, in any wording, any time after the bias was identified and before the conversation concludes. Settled lines count (answered Q1). The "at stake" line is taken to count for the same reason; say if not. | The same point made in the reply prose and absent from the page. |
| The line shown once and then answered: it must not be asked again. | A debias re-asked after the person answered it, or after the line already stood on the page. That is a misfire. |
| Toward the shown rate: a debias that appears where the labels say a bias was present. | Toward the misfire rate instead: a debias that appears where no bias was present, on a story, a figure of speech, an emotional spill or a restatement. Same event, different column, decided by whether the bias was there. |
| Bias in what the person says | Debias the page should show |
|---|---|
| A fact offered as proof that would be just as true under either answer (diagnosticity) | What would tell the two answers apart |
| A figure read one way only (framing) | The same figure read the other way |
| One option weighed with no alternative on the table (single-option evaluation) | What they would do if this option vanished |
| A fact read as evidence for one side only (one-sided evaluation) | Would they read it the same way from the other side |
| Advisers who already lean, or one source carrying the whole read (biased sample) | Who could give a straight read before knowing their lean |
| The same considerations circling with nothing settling (deferral, status quo) | What someone taking over tomorrow would do |
| An extreme recent result taken as the new normal (regression neglect) | How much of it a typical stretch would have produced |
The list is right for now and will change substantially with the wiki (Dominique). Base rate, the devil's advocate and Franklin are out of scope for this benchmark for now, by decision. Each bias's written definition is the contract the labels encode; a reader that infers extra conditions is measuring its own judgement.
Every reader is a model, and a model reader counts only after it agrees with a human-labelled set. The set is small, blind, frozen and versioned, and it grows slowly with real conversations. When a reader's model changes, it is scored against the set again before its numbers count. This is the step every earlier benchmark skipped.
/admin/benchmarks; raters at /benchmarks), specified in docs/spec/SPEC-benchmarks.md and grounded in Wilfried's research note Benchmark flow research: how the smart teams evaluate judges and raters. The rules that carry over to this schema:The readers are judges, and the research note already sets their trust bar. This part is specified, not built: the spec calls it "the judge path" and leaves it for after the picker. What it says, and what applies to the page-line reader and the quality judge:
For the three readers of this schema: the detector is the picker, already the benchmark's built subject, scored by precision with recall as its guard. The page-line reader and the quality judge are the judge path above. The outcome reader may need no model at all: whether a line settled on the person's answer is in the page versions.
To know whether a debias instruction, a page rule, a prompt edit or a model swap changes anything, we replay the same frozen turns twice, with and without the change, three draws each, and let the readers score both. Criteria are written before the run. The difference per bias is the result. The same shape answers "should we retire this instruction" and "did this prompt edit help".
Decision rule per bias. If arm A already shows the debias at the rate we want, the instruction is retired and the bias is served by the engine alone. If not, we add the cheapest thing that could raise it, a page rule first, a fallback instruction second, and keep it only if arm B beats arm A at three draws. The rate we want is an open question below.
How a verdict is read. One primary metric fixed before the run. Per sample, whether B beat A, A beat B, or both were the same; an exact sign test on the samples that differ. Better when the sign test favours B at p below 0.05 (8 of 8, 9 of 10, 12 of 15 or 15 of 20 differing samples). Worse is the mirror. Not detectably worse is the verdict for a cost-saving swap, and applies only when the stated non-inferiority margin is larger than the swing the benchmark can detect. Otherwise no detectable difference, with the swing size said. A candidate is ready to ship after two independent Better verdicts, or one Better whose gap is at least twice the A-versus-A noise floor. Adding samples until a result turns significant is refused; a larger suite is a new benchmark.
What the benchmark can see. With 36 paired samples and 30% of them differing, the smallest net difference it can detect at 80% power is about 25 points. At 56 samples about 20; at 100 about 15. Two identical picker runs today agree on 100 of 155 cases, so each variant runs each sample 3 times and takes the majority; that is the "three draws" of this schema.
What this means for retiring an instruction. "Retire" is a cost-saving swap: engine alone in place of engine plus instruction. Its verdict is therefore "not detectably worse", which needs a non-inferiority margin stated before the run, and a suite large enough that the margin exceeds the detectable swing. With today's 36 openings that margin cannot be tighter than about 25 points. So Q4 below is not a shown-rate threshold; it is the margin we accept, and the suite size that margin requires.
The statistics already exist in the app's benchmark and carry over unchanged. A variant's answer on a sample is the majority of its k runs; self-agreement across the runs is reported. No verdict is given while a variant has more than 5% of its calls in error, while no A-versus-A noise floor exists on the same dataset version, or while a split waits for a person. The comparison is a paired sign test per opening with Wilson intervals, per-bias rows flagged only after Holm correction; time and cost sit beside the verdict, never inside it. "Compared with last time" re-scores both benchmarks against today's gold with no model call. Every number is held against a hand-worked toy benchmark and a committed golden benchmark inside npm test. What this schema adds is only the subject: today the benchmark's one subject is the picker; the page-line reader is its second subject, and the spec already reserves the place for it under the name "the judge path".
| Number | What it says | Where it comes from |
|---|---|---|
| Detection | Precision and recall of "bias present" on the person's message | Detector against the labelled set |
| Debias shown | Of the moments where a bias was present, how often the page carried the debias within the window | Page-line reader |
| Misfire | Debiases shown on a story, a restatement or an already-answered line | Page-line reader + detector |
| Quality | Plain question, names the other answer, asks for an observation | Quality judge, rubric from the rulings |
| Persistence | The line stayed until answered, not silently settled away | Page versions |
| Outcome | Successful, ignored or dismissed, counted separately from shown | Outcome reader on later turns |
| Cost, latency | Per turn, per model | The call log |
| Later: conclusions | Conclusion quality with the debias shown versus not; where priority between biases will come from | Rated conversations, once there are enough |
The page-line reader can run on production pages and report counts only. No text leaves. That gives a quality signal for a beta where nobody can read the conversations, and it is the first time production would tell us what people actually saw rather than what rode in a prompt.
These are the pieces the architecture clean-up will meet. Each becomes part of the benchmark or is deleted on purpose; none should disappear by accident.
| Piece in the app now | Role in the benchmark | What the clean-up must keep |
|---|---|---|
| The picker and its evidence spans | Becomes the detector; its output is an expectation | The per-bias definitions and the evidence quote |
| The firings table and its outcome column | Becomes the outcome record: shown, successful, ignored, dismissed | The link from a line to the person's later turns |
| Page versions with the sweep turn id | The page-line reader and the replay both rebuild "page before" from it | One version per user turn, written before the edits |
| The "asked / declined" report from the engine | None; attribution stops mattering once the reader exists | Can go |
| The eval-only prompt arms and levers | One paired-replay path, plus the plain-tail lever for models that refuse the thinking clause | One path, not five |
| The heuristic column on page lines, the registry, the icons | The labeller stamps the bias a line embodies; the icon is what the person sees | Must survive even if every instruction is retired |
| The prompt stamp on turns; the rated-conversation import | Per-model, per-version runs; corpus growth with real people | Both, unchanged |
The benchmark app: Admin › Benchmarks (/admin/benchmarks), Admin › Datasets (/admin/datasets), raters promoted in Admin › Users, the raters' own screen at /benchmarks; seats, gold, adjudication, verdict statistics | The labelling and verdict machinery for every reader in this schema; the "judge path" the spec reserves is where the page-line reader lands | All of it; extend the subject from the picker to page lines, do not build a second benchmark |
Practices the research note reviewed and rejected for a three-person panel. Listed so nobody re-proposes them.
Status: schema for review. The detector, the page-line reader, the record and the audit view are built (GOD-1650, on main 2026-09-27); the labelled set, the trust bar and the paired replay are next (GOD-1671).2026-09-28