Logbook: the detector, the reader and the instructions
As of 27 September 2026. For Wilfried, Stef and anyone joining: how things worked before, why we changed them, how, and what we decided. Nothing below is deployed yet.
Terms used in this document
Each term means one thing, and nothing else is used for it.
- The engine: the main model call. Each time the person sends a message, it writes a reply and updates the page.
- The reply: the text the engine writes back. The person never sees it as such; what they see is the page.
- The page: what the person sees: their question, what is settled, and the open questions still to answer.
- A line: one open question on the page, with a short closing condition under it (what the person must say to close it).
- A bias: a thinking trap in what the person says (for example, sunk cost: "we've already spent so much").
- An instruction: a paragraph, one per bias, added to the engine's prompt when that bias is detected. It tells the engine which question to ask and how to put it on the page as a line.
- A switch: a per-bias setting that allows or blocks its instruction.
- The icon: the small tag on a line that names the bias it addresses.
- A debias: a line on the page that addresses a bias. After it, the person's reaction is recorded as successful, ignored or dismissed.
- The detector: a small, fast model call that reads the person's message and says which biases are in it, quoting their exact words.
- The reader: a second small model call that reads the lines that changed on the page and says which bias each line addresses, if any.
The rules we follow
- A bias is addressed only when a line about it lands on the page. What the engine says in its reply does not count.
- The truth is what the person sees. We judge the words on the page, not what the engine claims it did.
- We test a change by comparing the same moments with and without it, and we read the results example by example, not only as totals.
- One change at a time, each tested before the next.
- Nothing is deployed before the detector and the reader are scored against labels we make by hand.
1. One instruction per bias (25โ26 September)
Before. Until 27 August the engine made two calls per turn: one wrote the reply, a second updated the page. So each bias had two instructions: one for the first call ("ask this question in your reply") and one for the second ("if the reply asked it, put it on the page as a line"). When the engine moved to a single call, both instructions kept being sent to it, still describing a hand-off to a second call that no longer existed. The engine was also asked to report on itself: "did I ask the question, and on which words?"
Why change. The two instructions had drifted apart as each was patched separately, and the self-report was unreliable. On 55 real turns, instructions were sent 65 times; the engine reported "asked" 7 times, 19 of its reports could not be read, and only 3 lines actually reached the page.
What we did. Merged each pair into one instruction: every sentence that existed in only one of the two was kept (that is the feedback gathered over time), the later wording kept where both said the same thing. The engine's self-report was removed. Unused switches and old prompt versions were deleted.
Result. In the final check, each of the eight biases landed its line 3 times out of 3, and every case where the bias was absent stayed silent.
See: The Turn, Before and After ยท The Fold, Side by Side.
2. The detector reads every bias, and a reader was added (26โ27 September)
Before. The detector only checked the biases whose switch was on and whose timing allowed them. After step 1, the engine put a hidden mark on any line it landed because of an instruction, and that mark decided which bias icon the person saw on the line.
Why change. Two problems.
- Most debiasing questions are asked without any instruction. When we reread the archived conversations, we found the engine often asks a question that addresses a bias on its own. Those lines got no mark, so they were invisible in our records.
- The engine was still grading its own work, in a smaller form: its mark could not be checked against anything independent.
Dominique's rule for this work: every turn must be plainly auditable (which bias was detected, which question was shown, and why), in words a 10-year-old can follow.
What we did.
- The detector now checks every bias on every turn. When an instruction is not sent (switch off, wrong timing), the detection is still recorded with the reason.
- We added the reader. After the page changes, it reads only the lines that changed and says which bias each addresses, with a one-line reason. It does not care whether an instruction was sent: it judges the words the person sees.
- The reader's label now decides the icon. The engine's hidden mark was deleted.
- An audit screen in Admin shows a conversation turn by turn: the biases detected on the person's words, the lines shown with their bias and reason, and what the person did with each.
Result. When an instruction was sent, the reader agreed with the detector every time. On lines the engine wrote on its own, the reader over-labels: on Dominique's test page, it gave 3 of 4 such lines a bias they do not address.
Decision. We do not tune the reader on one conversation. It is scored on our hand-labelled set first, and deployment waits for that score. The tiebreaker bias was also retired as a trigger: the detector kept seeing it in stories and emotional outbursts instead of in a fact offered as a reason.
3. New biases from the research (26โ27 September)
Before. The app used its own original set of eight biases (Plan B, the successor, the tiebreaker, the mirror, both frames, a typical stretch, outside read, base rate).
Why change. The research catalogue lists 42 biases, each with an established way to counter it. We want the app's biases to come from that research, not the reverse.
What we did. Six new biases were wired for testing, each with its instruction: three numbers, one reason against, objectives first, the master list, the next spend (sunk cost), the expectation. The words bias and debias, as defined at the top, replaced the older word "heuristic".
See: Godspeed bias (a Claude Doc, to be kept here once exported; its content is in the Linear document Technique surfacing and its companion), the catalogue: what each question does, what triggers it, what the engine asks, and the line it lands on the page.
4. The detector reads less (27 September)
Before. On each turn, the detector read the person's question, the page, the new message, and the last four exchanges word for word. An exchange there meant the person's message and the engine's reply.
Why change. The person never sees the reply. So the detector was reasoning over four turns of text the person never read, and could detect a bias in words the engine wrote rather than words the person wrote. That breaks rule 2. Dominique's ruling: the detector reads only what the person has written or can see on screen, nothing more.
What we did. The engine's reply was removed from what the detector reads outright, with no test: it breaks rule 2. That left one open question: keep the person's earlier messages or not? We tested two versions on the same 169 moments, 3 tries each:
- Version A: the question, the page, the new message, plus each earlier exchange (the line the person answered and what they wrote).
- Version B: the question, the page, the new message. Nothing older.
Result.
- A and B disagreed on 111 pairs of moment and bias.
- On 72 moments where A and B carried exactly the same information and differed only by a few framing words, they still disagreed on 29 pairs. So the detector is sensitive: a few words change its answer about as much as the whole history does.
- Detections that actually needed an earlier message: only 10, and 3 of those were wrong by construction.
- Repeated detections of the same opening words were the same under both (31 and 28 over 36 conversations).
Decision: version B. The earlier messages added no clear value, and B reads exactly what the person has on screen. It is also shorter and cheaper, and it is fixed word for word, which the hand-labelling needs. Merged, not deployed.
See: What the Detector Reads (an artifact not shared with the team; to be kept here once its owner saves it).
5. We keep the instructions (27 September)
Question. Does the engine ask these questions on its own? If it does, the instructions are useless and we drop them.
What we did. 30 real past moments where one of the six new biases was present. For each, the same prompt was run 3 times with the instruction and 3 times without. Cost $2.01.
Result.
| Instruction | Without it | With it |
|---|---|---|
| The next spend (sunk cost) | 0 of 18 | 12 of 12 where the situation was really there |
| The expectation | 0 of 15 | 9 of 15 |
| One reason against | 0 of 15 | 7 of 15 |
| Objectives first | 0 of 15 | 7 of 15, but it copied the instruction's own wording every time |
| The master list | 0 of 12 | 4 of 12, always two wants instead of three or four |
| Three numbers | 0 of 15 | 0 of 15 |
Without an instruction, the engine asked none of the six questions (0 of 90). At most it named "sunk cost" in its reply, which the person never sees as a line.
Decision: keep an instruction per bias. Keep three as written (the next spend, the expectation, one reason against). Reword the other three (section 6). Objectives first, version 2, is accepted: the copied wording is gone (5 clean lines out of 6 landed).
We also decided that any bias may be addressed from the very first message. Nothing is held back for later turns. We will watch whether the first page gets too crowded.
6. We are fine-tuning each instruction, one at a time (in progress)
Why. Section 5 shows the instruction does the work, so its wording decides what the person sees. Wording the engine does not understand gets copied into the line, and a vague count gets ignored.
How. One instruction at a time: reword it, re-run only the "with" side on the same moments as before, compare line by line, Dominique accepts or asks for another pass, then the next. Done so far: objectives first (version 2 accepted). In progress: the master list (the example now shows three wants, "three, never fewer"). Then three numbers, dropped if it still never lands. After the rewords, the instructions we keep are shortened, with the same check.
What comes next
- Finish fine-tuning the instructions (section 6): the master list, three numbers, then shortening (~$1 per re-run).
- Labelling session with Wilfried and Stef: we label real moments by hand, against the fixed detector (version B). The reader plugs into Wilfried's benchmark app, not a new tool.
- Score the detector and the reader against those labels, decide bias by bias, then deploy.
- Parked: whether the same opening words should be detected only once per page.