How Godspeed Finds Biases

How Godspeed finds biases

Two products use the same machinery to spot a bias in what a person writes. The app then acts on it during the conversation; the interview gathers what it found into a report at the end. State as of 28 September 2026, main at a1b1abf1.

Shared by both

The two pieces both products use

The definitions. One short paragraph per bias, saying what the bias looks like in a person's own words, and ending with Not this: the look-alikes that must not count. Three definitions are the very same text in both products: confirmation bias, sunk cost and advice from the choir.

The detector. One model call that reads a person's message against the definitions and answers with each bias it finds and the exact words that show it (2 to 30 words, quoted). A check then confirms the quote really is in what the person wrote; a quote that isn't there is thrown out. In the code it is still called the picker (heuristic-picker.ts). It is one component: if it fails, it fails in both products, and it is fixed in one place.

What the detector reads. Only what the person can see: their new message, and in the app the opening and the page's lines. Never an earlier message and never Godspeed's own reply (ruled on 27 September, GOD-1668).

Part 1 · The app

Detecting, triggering and reading, every turn

In the app a person is working through a decision that is still ahead of them. Each time they send a message, a bias found in it can make Godspeed ask one targeted question on their page.

  1. The person sends a message.
  2. The detector reads itshared against the 14 app definitions, together with the opening and the page's lines. It returns, for instance, mirror, because of "That tells me everything".
  3. The arbiter decides which of those may rideapp. Three conditions. The person's flag for that bias is on (all these flags are off in production today, on in development). The bias is not retired (the tiebreaker and a typical stretch are). Its window is open (only the master list and base rate still have one). Everything that passes rides; there is no cap. A bias found but not allowed to ride is still recorded, as suppressed, with the reason: off, window or retired.
  4. Triggering: the instruction ridesapp. Each riding bias adds its instruction paragraph to the engine's prompt for this one turn. The paragraph tells the engine to add one Open line to the page asking the technique's question in the person's own words, with a closing condition, and never to name the technique. Base rate works differently: it launches a web search, and the figure it finds rides a later turn.
  5. The engine replies and edits the pageapp, as it would anyway, plus the question.
  6. The reader reads the pageapp. After the turn it reads every line the turn added, rewrote or moved on the page, never the chat reply: a move counts only when it reaches the page. For each line it says which of the 14 biases the line works against, or none. It is blind by design: it does not know what the detector chose, so it is an independent check. Its answer is the icon on the line (the latest reading wins).
  7. The loop closes in the recordapp. The two readings are compared, one entry per turn and bias (see below), with the quote and the line. Admin › Audit shows a conversation turn by turn.

Example (a recorded test conversation):

My cofounder skipped the investor call again and then asked me to send him the notes. That tells me everythingMirror

The detector finds the mirror; the arbiter lets it ride. The engine adds an Open line, and the reader labels it:

Your cofounder asked for the notes after skipping: what else could that mean besides not caring?Mirror

Closing condition: Name what you would need to see to be sure. Detector and reader agree, so the record says the question landed.

Closing the loop: two ways a debias reaches the page

The reader does not care who asked the question. It only checks whether a line on the page now works against one of the 14 biases. So the same count, debias shown, fills from two sources, and the record keeps them apart. That is how we measure both what our 14 techniques add and what the engine does by itself.

Triggered by our techniques
  1. The detector finds the bias
  2. Its instruction rides
  3. The engine writes the question
  4. The reader sees it

Your cofounder asked for the notes after skipping: what else could that mean besides not caring?Mirror

Recorded as: debias shown, triggered.

Done by the engine alone
  1. No instruction rode (not detected, its flag off, or retired)
  2. The engine asks the question anyway
  3. The reader sees it

Does the second site's rent work at your normal 18k, not your best month?Typical stretch

Recorded as: debias shown, by the engine alone. This is how a typical stretch was caught doing its move without us, and retired.

One more case stays visible: the instruction rode but the reader finds no such line. Recorded as rode, debias not shown. The instruction failed to reach the page.

The two outcomes of a debias shown

Once the question is on the page, what the person does with it ends in one of two states:

SuccessfulThey answered it, and the line settled on their answer.
IgnoredThe page concluded with the line still open.

There is no third outcome. "Dismissed" was a legacy of Set aside, from when a question could be set aside; since Set aside stopped being a drop zone, nothing in the app dismisses a line, and the word is retired (2026-09-28).

Successful means the question was taken up, not that the bias is gone; that is measured later, on the conclusion. The record has a place for the outcome, but how the app fills it has not been checked yet.

The app's biases today

StateBiases
Ride (12)Plan B, the successor, the mirror, both frames, base rate, outside read, three numbers, one reason against, objectives first, the master list, the next spend, the expectation
Retired (2)The tiebreaker, a typical stretch: still detected and recorded, never ride
TestedThe wording of the instructions: each rewritten and replayed on real moments, with and without it (GOD-1669, GOD-1674).
Not measuredWhether the detector finds the right biases, and whether the reader labels lines correctly. That is the labelling session (GOD-1671).
DeployedNo. Held until the reader is scored.
Part 2 · The interview

Listening during the chat, detecting for the report

In the interview a person tells two or three decisions they have already made. Nothing is asked of them to fix a decision; the goal is a report on how they decide. So there is no page, no instruction and no reader.

  1. The person tells a decision in a chat, one message at a time.
  2. The interviewer listensinterview. It is a model with its own prompt, carrying each bias's cue and check question from the catalogue (for confirmation bias: "What was the best argument against it, and what did you do with it?"). When it hears a cue it may ask the check question. It never names a bias. The detector plays no part in this.
  3. The interview closesinterview with every turn recorded.
  4. The detector runsshared, once per message the person wrote inside a decision, reading only that message (no page, no earlier turns), against the 42 definitions. The twelve headline biases use the definitions approved on 28 September; the other thirty still use the catalogue's own description.
  5. The findings are counted per decisioninterview: a bias is seen in 2 of 3 decisions, never counted twice in one. The strongest become the headline.
  6. The report is writteninterview: each bias named, the person's own words quoted back, the research behind it, a book, and one thing to work on. The same content comes as a bias.md file.

Example (illustrative, not a recorded interview):

We'd looked at flats for four months. The landlord needed an answer by Friday, so I signedDecision by deadline. I was 100% sure it was the oneOverconfidence.

The outside deadline is given as the reason for deciding then, and nothing says what would have been enough beforehand. If the same shape shows up in a second decision, the report can headline it as seen in 2 of 3 decisions and quote both sentences.

The interview's biases today

GroupDefinition the detector reads
12 headlineApproved on 28 September. Three are the app's own text (confirmation bias, sunk cost, advice from the choir). Three are interview-only for now (decision by deadline, information as reassurance, deciding hot).
30 othersThe catalogue's description, until the twelve pass labelling.
On mainThe detector for the interview and the twelve definitions (#721).
On a branchThe report generation and the report page, with the design session (GOD-1659).
Not measuredHow well the detector finds these biases. The first probe, with the older text, flagged 20 to 24 of 42 in every interview. Labelling is joint with the app (GOD-1671).
Side by side

What differs

AppInterview
The decisionStill aheadAlready made
When the detector runsEvery message, before the replyOnce, after the close
What it readsThe message, the opening and the page's linesThe message alone
Biases14 (12 ride, 2 retired)42
What a finding doesAdds one question to the pageGoes into the report
InstructionYes, one per techniqueNone
ReaderYes, labels each changed lineNone (no page)
Questions askedThe technique's question, on the pageThe interviewer's check questions, in the chat