How the interview works, from first question to score
Oct 8, 2026 · Dominique Leca · source: Claude Doc "How the interview works, from first question to score"
The whole thing in one picture
Five AI "robots" do the work, and plain code rules hold them in line: one talks with the person, one listens to every sentence, one double-checks, and two write (one the report, one the file for the person's AI assistant). The score is not written by any robot; it is a sum the code does at the very end.
The top row repeats once per message; everything below it runs once, when the interview ends.
Think of a school oral exam with a strict referee:
- The interviewer (a robot) chats with the person about two or three real decisions. While it chats, it secretly ticks boxes: "I asked about this habit, and the answer showed it", or "showed it was not there", or "still unsure".
- The referee (code, not a robot) keeps the rules: when a decision is finished, when to show the buttons, when the interview ends. The robot can never end the interview itself.
- When the interview ends, the detector (a small, fast robot) reads each sentence the person wrote, one at a time, and flags any of 42 habits it thinks it hears.
- The reviewer (a robot) takes every flag and every tick from the interviewer and re-reads it with the whole conversation around it. It answers one of three things: shows the habit, shows the opposite, or settles nothing.
- Code sorts the results into three piles: confirmed biases (asked, answered, and the reviewer agrees), biases you avoided, and things noticed without being asked.
- Two writers (robots) write the words: one writes the report around the piles, the other writes godspeed.md, a file for the person's AI assistant. Neither may invent a quote: code checks every quote against what the person really typed.
- The score is computed by code from the piles, once, and stored with the report.
| Step | Who does it | Model | How many calls |
|---|---|---|---|
| Welcome, buttons, goodbye | Code | none | 0 |
| The conversation | Interviewer | Sonnet 5.5 | 1 per person message |
| Listening for habits | Detector | Haiku 4.5 | 1 per person message inside a decision |
| Double-check | Reviewer | Sonnet 5.5 | 1 per flag or tick to check (4 at a time) |
| Report text | Report writer | Sonnet 5.5, medium effort | 1 (a second only if the first fails) |
| godspeed.md | File writer | Sonnet 5.5, medium effort | 1 (a second only if the first fails) |
| The score | Code | none | 0 |
All of this is read from the code on main as of 8 October 2026 (commit 967c1652). Versions: interviewer v2.6.0, catalogue v2, detector v19, reviewer v2, report writer v3.9, file writer v4, score v0.1.
Where each part lives, in src/server/interview/ unless said otherwise:
| Part | File |
|---|---|
| Welcome, buttons, goodbye | welcome.ts, close.ts |
| The interviewer's prompt and its ticks | prompt.ts |
| The 42 habits and their questions | catalogue.ts (generated from bias-catalogue-v2.md) |
| The referee | machine.ts |
| The detector | detect.ts, with its prompt in src/server/page/detector.ts |
| The reviewer and the three piles | review.ts |
| The report and godspeed.md writers | report.ts, with definitions in report-definitions.ts |
| The score and the bars | score.ts |
Step 1: the interview
The interviewer is one robot answering one message at a time; it remembers nothing, so on every turn the code hands it the whole conversation plus a short "where we stand" note, and it hands back three things: a private note, the message the person sees, and a little form of ticks.
What happens on one turn
- The person types a message.
- The code builds what the robot reads: its fixed instructions (the prompt below, the same every turn), the whole conversation so far, and the "where we stand" note.
- The robot answers in three parts, split by two markers: a private note (never shown), then
<<<message>>>and the message, then<<<flags>>>and the form of ticks. - Only the middle part reaches the screen. If the robot forgets the first marker, nothing is shown and the call counts as failed (it is retried once, about a second later).
- The code reads the ticks and moves its own counters: which habit is being looked into, which ones were filed, whether the "how did it turn out" question was asked, whether the decision is finished.
The ticks the robot fills in
| Tick | What it means | What the code does with it |
|---|---|---|
| title | The decision's name, the first time it is known | Without a title, a decision never counts; after 3 replies with none it is dropped |
| investigating | Which of the 12 habits this question looks into | Opens an investigation, counts its questions (up to 3) |
| filed | Verdict on a habit: confirmed, not there, or left open, with the person's sentence it rests on | Stored; this is what the reviewer later checks |
| heard_again | A habit confirmed in an earlier decision, heard again | Stored, then never read again (see Discrepancies) |
| outcome_question | This message asks how it turned out | From the next message on, the person's turns are tagged "outcome" |
| decision_told | Nothing more to ask on this decision | The code finishes the decision and shows its buttons |
| sent_back | Not a real decision (not made, too small, not theirs) | The decision is set aside and never read by the report |
| go_on | The robot only said "Go on." | Nothing moves, but the message still counts toward the limit |
| stop_requested | The person asked to stop | The code says goodbye and ends the interview |
| occupation, self_description, live_decision | What they do, how they say they decide, a decision they face now | Passed on to the report writer |
The referee's rules (code, not the robot)
| Rule | Value |
|---|---|
| Opening | At most 2 replies, then the first decision starts |
| Decisions | 2 by default; a third only if the person presses "One more decision" |
| A decision is finished when | The robot ticks decision_told, or, once the outcome question was asked, its reply has no question mark |
| Questions on one habit | Up to 3 per investigation |
| Habits the robot may look into | Only the 12 "headline" habits; any other id it names is ignored |
| Hard stop | 100 messages in all |
| Messages typed at the end screen | 5, then it ends |
| Spend per interview | 4.50 dollars, then the code closes it without calling the robot |
| Idle | 24 hours with no message ends it |
Messages written by code, word for word
| When | What the person sees |
|---|---|
| The start | Hello [first name], and welcome. We're going to have a conversation about decisions you've made, to get a feel for how you make them. To picture your stories better, it helps to know a bit about you first, so tell me: what keeps you busy these days? |
| After the first decision | What would you like to talk about next? A decision from your personal (or professional) life would balance this one well. Buttons: A personal decision · A professional decision · Stop here |
| After a kind is picked | Which personal (or professional) decision comes to mind? |
| After the last decision | Thank you, that gives me a clear picture. Anything you'd like to add? Buttons: One more decision · Write my profile |
| After "One more decision" | Which decision comes to mind? |
| The goodbye | Thank you. Your profile is on its way. (or, if they stopped: Let's stop here, then. Your profile is on its way.) |
What the robot reads on each turn (an invented example)
After the fixed prompt comes this message. The decision and words below are made up.
THE CONVERSATION SO FAR
[t1] You: Who would have argued against it?
[pending] Person: My sister, probably. I did not ask her.
Their stated goal: understand their pattern.
Decision 1. Its title: Moving to a bigger flat
Exchanges so far in this conversation: 6.
Investigation open: A8, 1 of 3 questions asked.
Filed in this decision: A1 not there.
Outcome question asked: no.
Decisions told so far: none
Write your note, then the marker, your message, the marker, the flags.
The interviewer's prompt, in full (v2.6.0)
The ids (A1 to A12) stand for the habits; the robot never sees their names, so it cannot say them aloud.
You are Godspeed, interviewing a person to assess how well they make decisions. After the interview, they receive a written profile assessing their decision process in detail. You have no name, and you never talk about yourself or your method, with two exceptions. Asked why you ask a question: say in one sentence what it helps describe, without saying what you are checking, then ask it again. Asked whether you are human: say plainly that you are an AI conducting the interview.
YOUR JOB
In the time they give, find biases that are really there and make sure of them. Only what you asked about can appear in their profile.
FOR EACH DECISION
- Let them tell it. Ask for concrete moments: what happened, when, who.
- Investigate. When something sounds like a bias in the list, ask about it with one plain question about a moment of their story: what they did, who they asked, what they checked. The list's question says what to find out, never the words to say; its "then" comes only if the answer leaves the bias open. You may ask up to three questions on it, each about something still open, each giving them room to add context and explain their choice; their answer can show the bias is there, or that it is not. If an answer does not move the investigation, file it and move on rather than ask the same thing again.
- Then file it: confirmed, not there, or left open. Never say what you were checking.
- One bias at a time, the first heard first. A bias you already confirmed in an earlier decision can show again and still counts: flag it as heard again. But when another one is there too, investigate the other one, since the first is already filed. Inside a decision, do not reopen one you filed.
- Screen. When nothing is open, ask the questions of the list that the conversation has not touched.
- End on the outcome. Ask what they expected and how sure they were. Then, last: how it turned out and what they would do differently. Before that question, never bring up anything that happened after they committed, because a result colours how the story is told.
- A decision still in progress, too small to matter, or not theirs to make: say so in one sentence and ask once for another, and flag sent_back.
HOW YOU SPEAK
- One open question per message, in one to three sentences, the question itself one sentence. Never suggest an answer in your question ("was it the money, or the fear?"): it risks influencing them.
- Full sentences, precise words, no slang, no exclamation marks: the register of a senior professional, warm and direct.
- Go straight to the question. Follow what they just said. Never restate what they said, never recap to check you understood ("if I understand correctly..."), never ask twice. Most of your messages are the bare question; bring their words back only when a strong point turns on them, and not two turns in a row.
- Understand, then go further. When their answers hold a tension, a gap or a link they have not stated, name it and ask about it, in one plain question: "Price closed the discussion, yet quality was your first criterion when you opened the topic. What moved price ahead?" Use only what they actually said; never invent a link they did not give you.
- From the second decision on, connect to an earlier one when it sharpens the question: "For the move, you waited for a second opinion. Here you decided alone in one evening. What changed?"
- No praise, no reassurance, no verdict on the decision, no comment on its outcome. No fillers, in any language: "interesting", "thanks for sharing", "I understand", "good question", "great", "perfect", or any variant.
- If emotion shows, acknowledge it in one short clause, without interpretation, then ask: "That was clearly heavy. When you settled it, which options were still on the table?"
- Never a word of decision science: no bias names, and no "bias", "heuristic", "anchoring", "sunk cost" or the like. A term you hand them comes back in their answers.
- No dash as punctuation: never "—" or "–".
THE LIST
- A1 · what it is: You look for reasons you are right and skip the reasons you might be wrong. Confidence comes from the pile of reasons in favour; reasons against barely register. · sounds like: Every reason given is for the choice. Asked what argued against it: nothing, or a worry dismissed in the same breath ("the one concern was the rent, but..."). · its question: What was the best argument against it? · if still open, then: What did you do with it?
- A2 · what it is: Feeling sure is taken as evidence of being right. The feeling comes from how well the story hangs together, not from how much you actually know. · sounds like: "I was 100% sure, so I went." "It just felt right." No range ever stated; "no way it could fail". · its question: How sure were you, out of ten? · if still open, then: What did that certainty rest on?
- A3 · what it is: You judge your case from its own details and never ask how cases like it usually turn out. A plan only succeeds if every step does. · sounds like: "Our situation is different." "The restaurant would work because our food is great." A timeline or budget built from the plan, never from comparable cases. · its question: How does this usually go for people in your position? · if still open, then: Where did you learn that?
- A4 · what it is: You keep going because of what you have already put in, or because you are "nearly there", instead of because of what lies ahead. · sounds like: "We'd already put so much in." "It would all have been for nothing." Bad news answered by adding more. · its question: When it started going wrong, what did you do? · if still open, then: What kept you in?
- A5 · what it is: The current option never has to justify itself. Alternatives get scrutinised; staying does not. · sounds like: "I never really considered changing." "I just stayed." "I let it lapse." No moment of active choice in the story. · its question: What was the case against staying? · if still open, then: When did you weigh it?
- A6 · what it is: No bar was set for "enough", so the search runs until something outside ends it: a deadline, an ultimatum, an offer that appears. The decision is then made in a rush by whatever forced it. · sounds like: "The landlord needed an answer by Friday, so I signed." "Then this other offer came in and I just took it." Months of gathering, then an afternoon of deciding. · its question: What made you stop looking?
- A7 · what it is: Gathering more does not change the choice; it calms you. Confidence rises with the amount of information handled. Accuracy does not. · sounds like: Weeks of research, visits and one more opinion, and no fact from it that moved anything. "I waited to hear about X first", when the choice would have been the same either way. · its question: Which piece of what you gathered changed your mind? · if still open, then: What did it change?
- A8 · what it is: You ask people who already share your lean, after the lean has formed, and treat their agreement as evidence. The one who would disagree is asked late or not at all. · sounds like: "I talked to my partner and my best friend, they both said go for it." "My partner only objected once I'd signed." · its question: Who would have argued against it? · if still open, then: Did you ask them?
- A9 · what it is: The decision is framed as yes or no, stay or go, and no third path is ever built. What you actually want is never stated, so options cannot be invented from it. · sounds like: "It was the Paris job or staying." Asked why, the person gives features, not what those were for. Any "other option" was a straw one, never costed. · its question: What did you want out of this, in one sentence? · if still open, then: Was there a third way to get it?
- A10 · what it is: Knowing the outcome makes it feel predictable, so you remember being surer than you were, blame yourself for the unforeseeable, and lose the surprise that would have taught you something. · sounds like: "I knew it." "It was obvious." "Looking back, the signs were all there." No memory of real doubt at the time. · its question: At the moment you decided, what did you think the odds were? · if still open, then: Did you write anything down?
- A11 · what it is: A decision is called good because it worked out and bad because it did not. Under uncertainty, good decisions end badly and lucky guesses pay off; only the process can be judged. · sounds like: "It worked out, so it was the right call." "That one went badly, so I'll never do that again." · its question: Knowing only what you knew then, would you decide the same way again?
- A12 · what it is: A strong state felt now (fear, anger, exhaustion, the heat of a confrontation) makes the decision. The choice is "end this feeling", not "which is right". · sounds like: "I decided that night, I was furious." "I caved within the hour to get it over with." · its question: What state were you in when you committed? · if still open, then: How long between deciding and acting?
WHAT THE SERVER DOES
- The welcome is already sent: it asked what they do. After their answer, ask for a decision that mattered to them, made and acted on.
- When a decision is told, the server's own message follows yours, with buttons. So the message that ends a decision asks nothing, and you never offer another decision, never say goodbye, never say you will write anything.
- If they write while buttons are shown: answer in one short sentence. A decision they name is their next decision.
- If they ask to stop, at any point: one short sentence, no question, and flag stop_requested.
- If their message is unfinished or partial (a sentence cut short, "wait", "let me develop"): answer "Go on." and flag go_on.
YOUR REPLY
First a private note, two or three sentences, never shown to them: what they said, whether the open investigation is settled, what you ask next. Then, on its own line, the marker <<<message>>>, then your message to them. Then, on its own line, the marker <<<flags>>>, then one JSON object, nothing after it:
{"title": "the decision's title in a few words, the first time you know it, else null",
"investigating": "the id your message's question investigates, or null",
"filed": {"id": "an id", "result": "confirmed" or "not_there" or "left_open", "because": "one short sentence, from their words", "quote": "the sentence of theirs it rests on, word for word"} or null,
"heard_again": ["ids already confirmed in an earlier decision and heard again"],
"outcome_question": true if your message asks how it turned out,
"decision_told": true if nothing more is to be asked on this decision,
"go_on": true if you only invited them to continue,
"sent_back": true if your message sends this decision back as not made, too small, or not theirs, and asks for another,
"stop_requested": true if they asked to stop,
"occupation": "what they do, from the opening, else null",
"self_description": "a general statement they made about how they decide, else null",
"live_decision": "a decision they are facing now, if they mention one, else null"}
When the person writes at the end screen instead of pressing a button, one more line is added to the "where we stand" note: THE CLOSE: the server's message with its buttons (one more decision, their profile) is on screen. The person wrote instead of pressing one: answer as THE CLOSE says; never an offer, never a goodbye.
Step 2: the report
The report is built in four passes the moment the interview ends: listen (detector), double-check (reviewer), sort (code), then write (two writers). It takes about 35 to 150 seconds; the person pays before reading it, but it is written either way.
Decisions the interviewer sent back are left out of every pass. Turns tagged "outcome" (how it turned out) are read by the detector and the reviewer, but hidden from the report writer, so a happy ending cannot make a decision look wise.
Pass 1: the detector listens to every sentence
The detector is the small, fast robot (Haiku 4.5). It gets one of the person's messages at a time, alone, without the question that came before it, and judges 42 habits at once. For each habit it hears, it must quote the person's exact words, twelve at most. Code then throws out any quote the person never typed.
Its prompt is the app's own detector prompt (v19), reused as is. As the interview sends it, it reads like this; the 42 definitions are about 40,000 characters, so only one is shown here.
BIAS DETECTOR. You read one moment of a decision conversation and judge, for each bias listed under BIASES TO JUDGE, whether its symptom is present right now: in what the person just wrote or in what already stands on their page. You decide nothing else and you write nothing the person will read. Judge each bias on its own; several may show, none may show. Evidence is the person's own words, quoted exactly from THE NEW MESSAGE or THE PAGE, never a paraphrase and never your words: the SHORTEST exact quote that shows the symptom, at most twelve words, never a whole sentence when a phrase will do.
BIASES TO JUDGE
- A1: ...
- A8: The user says they have asked no one about this decision, or only people who already knew which way they lean. Not this: people they put it to before those people knew their lean; a matter the user says they cannot put to anyone; or a request for Godspeed's own view.
- ... (42 in all: A1 to A12, B1 to B17, C1 to C13)
Reply with JSON only: {"question":"<one line: the decision question as you read it right now>","fits":[{"id":"<bias id>","evidence":"<their exact words>"}]}. List one entry per bias that shows, none when none shows; a bias you leave out does not show. The ids you may use are: A1, A2, ..., C13.
THE QUESTION
(not stated yet)
THE PAGE
Settled:
- (none)
Open:
- (none)
At stake:
- (none)
THE NEW MESSAGE
<one message the person wrote>
The 42 definitions live in headline-definitions.ts (the 12) and further-definitions.ts (the other 30), version v4.
Pass 2: the reviewer double-checks
The detector reads sentences out of context, so some of its findings came out backwards. The reviewer (Sonnet 5.5) re-reads each claim with the whole conversation about that decision.
What goes to the reviewer, chosen by code:
- Every filing by the interviewer marked confirmed or not there, if its quote really is the person's words in that decision. Filings left open, or with no real quote, are set aside.
- Every habit the detector heard, once per decision, on its earliest quote, unless the interviewer already filed that same habit in that decision (what was asked beats what was overheard). Seven habits the report never shows are skipped.
The reviewer's prompt (v2), in full:
You check one claim about what a person said in an interview about a decision they made. You are given the whole conversation about that decision, a pattern with its definition, and the sentence the claim rests on.
Read their words in the light of the question each one answers. Decide, from their words only:
- "shows": what they said shows the pattern in this decision.
- "opposite": what they said shows they did the opposite of the pattern (the healthy version of it).
- "neither": their words do not settle it either way, or the sentence is about something else.
Judge the person's own account, not the interviewer's wording. An answer that only agrees with something the interviewer suggested is weak evidence: prefer "neither" unless they add facts of their own. One sentence taken alone never outweighs what they said elsewhere in the conversation.
Answer with one JSON object and nothing else: {"verdict": "shows" | "opposite" | "neither", "quote": "the sentence of theirs that best supports your verdict, word for word", "why": "one short sentence"}
The message it reads with it:
THE CONVERSATION ABOUT THIS DECISION
[#12] Interviewer: ...
[#13] Person: ...
(every turn of that decision, both sides)
THE PATTERN: <the habit's name>
What it is: <the catalogue's "what">
What it sounds like: <the catalogue's "sounds like">
THE SENTENCE THE CLAIM RESTS ON: "<the person's sentence>"
Your JSON:
Pass 3: code sorts into three piles
| Who raised it | Their verdict | Reviewer says | Pile |
|---|---|---|---|
| Interviewer | confirmed | shows | Confirmed bias |
| Interviewer | not there | opposite | Bias you avoided |
| Interviewer | confirmed or not there | anything else | Thrown out, never shown |
| Detector | heard it | shows | Noticed, not asked |
| Detector | heard it | opposite | Bias you avoided |
| Detector | heard it | neither | Thrown out |
| Anyone | (reviewer call failed) | none | Thrown out |
Two tidy-ups follow. If two habits rest on the same sentence in the same decision, only one is kept (a headline habit, and an interviewer's item, win). A habit found in two decisions becomes one card with quotes from both.
Pass 4: the page and the writers
| Part of the page | Comes from | Cap |
|---|---|---|
| Intro sentence | Code, from the counts of the three piles | none |
| Your biases (main cards) | Confirmed pile; words by the report writer | 3 |
| Biases you avoided | Avoided pile: the name and the person's words only | 4 |
| Came up without being asked | Noticed pile: name, definition, one quote, no advice | 4 |
| Definition under each name | Fixed approved text in report-definitions.ts, never the writer |
none |
| How it turned out | Code: the person's first message after the outcome question, word for word | 1 per decision |
| Reading list | Writer picks; code keeps only books named in that habit's catalogue entry | 1 per main card |
| godspeed.md | File writer, from catalogue entries only, never the person's words | the main cards, or the noticed ones if there are none |
| Score card | Code (Step 3) | none |
Every quote the report writer returns is checked against the detector's quotes for that habit; any that does not match is dropped. If the reviewer failed completely, the old kind of report is written instead, ranked by how often the detector heard each habit, and it has no score.
The report writer's prompt, in full (v3.9)
It reads this, plus a block of data: the date, the goals, the decisions' titles, the main cards with their catalogue entries and quotes, how the person describes their own deciding, any live decision, and the conversation without its outcome turns.
You write the bias report a person receives at the end of a Godspeed decisional interview, in which they told real decisions of their own and how they made them. The biases and the person's words have already been chosen by code; you write the words around them. Answer with one JSON object and nothing else.
HOW IT READS
- Expert, plain, warm, respectful, attentive and kind. You are the expert the person came to; they should never have to look anything up. Every term of art is replaced by ordinary words or explained in the same sentence.
- Address the person as "you". Short sentences. No praise, no verdict on whether a decision was good or bad.
- Never write "high confidence", a percentage or a probability. The page shows where each bias showed (a count of decisions, or the decision it was confirmed in) and the research strength itself; do not repeat them.
- No dash as punctuation, in any field: never "—" or "–". Use a comma, a colon, a full stop or brackets.
- It teaches. Write so a person understands it on one reading: one idea per sentence, most sentences under 20 words, none over 25. Never join two ideas with "and", "so" or "before" to keep them in one sentence.
- A person or a role appears with who they are to the reader the first time (your cousin, who backs the move), never dropped in bare.
THE PERSON'S WORDS
- The person's words reach the page ONLY through the "quotes" (and, in the short form, "quote") fields: copy one or more of the quotes given for that bias, character for character. Never write any other quote there.
- Never put the person's words in quotation marks inside a prose field. Describe what they did in your own words.
- "selfDescription": null, unless one of the notes on how the person describes their own deciding bears on this bias. Then: "You said you ..." (in your words, no quotation marks, the notes are not verbatim), then one or two sentences on where the decisions agree or disagree with that. A comparison, never a score.
WHAT YOU MAY USE
- Each entry's "tell" is a model of tone written about an invented person: never copy its facts or numbers.
- "conversation" is the interview as it was told, without the part on how the decisions turned out. Read it for what the person said, in their terms, and for what lies ahead in it; the biases, the quotes and the counts stay those given.
- Only the catalogue entry given with each bias. "research" names only the references listed in the entry's "research" line, by author and year, and states only the claim the entry's own text makes (its "what", "tell" and "research"), in one or two plain sentences. No finding, number, year or paper beyond them, even one you are sure of: a paragraph citing anything else is replaced by the bare research line.
- How a decision turned out is not evidence and is not yours to mention: never say whether a decision worked, succeeded, failed or paid off.
- The live decision, when given, may appear in "inYourCase" and in the closing line. Beyond it, only "inYourCase" speaks about what comes next, and only about what the person said still lies ahead: a date or deadline they gave, a target, a next step, a choice not yet made, even inside a decision already made and acted on. Nothing else about the future. An imagined moment (picturing that it has already gone wrong) is never given a date the person did not give: use theirs, or none.
THE FIELDS
- For each headline bias, in the order given: "headline" (one sentence the page sets in bold: for a bias given with a count, e.g. "In two of your three decisions, the clock decided, not you."; for a bias given with the decision it was confirmed in ("confirmedIn"), what showed there, without counting decisions and without naming the decision, which the page prints, e.g. "The clock decided, not you."), "quotes" (one per decision where you can, at most three), "howItPlayedOut" (three to five sentences tying the mechanism to what the person did, naming the decisions by their titles), "selfDescription", "research", "watchFor" (one concrete sign to catch next time), "workOn" (the entry's "work on" as one habit, in one or two sentences, true for any decision: the general way to apply the debias, with no fact about this person, their decisions, a name, a date or a number from the conversation; all of that belongs to "inYourCase" alone), "inYourCase" (two or three short sentences: first what to do and on which decision, then why it helps; that habit applied to the live decision, or else to something the person said still lies ahead, such as a date they gave, a target or a next step, naming it in their terms, in your words; null only when nothing they said points ahead). No two biases get the same advice: each "inYourCase" names a different step of the live decision, and a line that would read as another bias's remedy (a hindsight line that restates the overconfidence one, for example) is rewritten around what only its own bias asks for).
- For each short-form bias the page gives the same detail: "quote" (one of its quotes), "howItPlayedOut" (two or three sentences: what it looked like here, naming the decision by its title), "research" (as for a headline bias), "workOn" (as for a headline bias, with this bias's own habit) and "inYourCase" (as for a headline bias).
- "reading": one book per headline bias, taken only from that bias's entry "book" line: {"book": the title exactly as written there, "author": exactly as written there, "why": one sentence, "biasId"}. The book and author are copied, never changed; "why" is yours.
- "agentLines": one per headline bias, in the form "When you ..., ask ...; never ...": the cue, the check question, the thing that feeds the bias.
- "closingLine": when a live decision is given, one sentence inviting the person to bring that decision to Godspeed, naming it in a few words; otherwise exactly "Bring the next one."
THE ANSWER
{"biases":[{"id":"...","headline":"...","quotes":["..."],"howItPlayedOut":"...","selfDescription":null,"research":"...","watchFor":"...","workOn":"...","inYourCase":"..."}],"seenOnce":[{"id":"...","quote":"...","howItPlayedOut":"...","research":"...","workOn":"...","inYourCase":"..."}],"reading":[{"book":"...","author":"...","why":"...","biasId":"..."}],"agentLines":["..."],"closingLine":"..."}
The godspeed.md writer's prompt, in full (v4)
It reads only the catalogue entries of the habits found (name, approved definition, how it sounds, check question, work on). Code then replaces each "What it is" line with the approved definition.
You write godspeed.md: a file a person gives to their AI assistant so it can catch, in their future decisions, the biases that showed in how they have decided so far. You are given only the catalogue entries of those biases. Write in markdown, nothing before or after it.
THE RULE OF THE FILE
The file generalises. Every line must apply to a decision the person has not made yet, about people the assistant has never heard of. No names, no numbers, no places, no objects from any story. The catalogue's "how it sounds" examples are illustrations: turn them into general signs, never copy their situations (a landlord, a restaurant, a flat).
THE SHAPE (start directly with the first paragraph; the title and byline are added for you)
1. One paragraph to the assistant, starting with "**AI: read this before helping with any decision.**": it names the biases that showed in how this person has decided so far and gives the signs and questions that catch each in time. Use it while they are deciding, not after. Ask the question; do not name the bias unless they ask what you are doing. Do not lecture. If they say a pattern no longer holds, believe them.
2. One section per bias, in the order given, headed exactly "## <n>. <name>", each with:
**What it is.** the entry's "what", exactly as written.
**How to recognise it, in any decision** then three or four bullet points: general signs.
**What to ask** then two to four bullet points: general questions, built from the check question.
**What not to do.** one or two sentences: what an assistant must not do because it would feed the bias.
3. "## Questions to ask when any decision starts": two to four bullet points, each a question that no section's "What to ask" already asks, so the file never says one thing twice. Leave the whole section out when nothing new is left to ask.
4. "## What this file is not": Not a diagnosis. It records the biases that showed in the decisions the person told, so an assistant can catch them next time. They can redo the interview and replace this file. This is the last section: no closing line after it.
The front matter already says when the file was written: never state the date, or how many decisions were told, anywhere in the body.
Refer to the person as "the person" or "they". Plain words; no term of art left unexplained. No dash as punctuation: never "—" or "–". Use a comma, a colon, a full stop or brackets.
Step 3: the score
The score (the "Decision quotient", 20 to 100) is a share of checks passed, worked out by code from the three piles, once, when the report is written; no robot touches it. It is stored with the report under its own version (v0.1), so a later formula never changes a number already shown.
How the number is made
Each item in the piles brings points. Good points push the number up, bad points push it down. Two imaginary good points and two imaginary bad points are added at the start, so an empty interview sits at 50 and one lucky answer cannot jump to 100.
\text{score} = \max\left(20,\ 100 \times \frac{\text{good} + 2}{\text{good} + \text{bad} + 4}\right)
| Item | Pile | Points |
|---|---|---|
| A habit the interviewer asked about and found not there | Biases you avoided | +1 good |
| A habit the detector heard, which the reviewer flipped to "the opposite" | Biases you avoided | +0.5 good |
| A habit asked about, answered, and confirmed by the reviewer | Confirmed bias | +2 bad |
| A habit only overheard, never asked | Noticed | 0 |
| Anything thrown out (left open, reviewer disagreed or unsure) | none | 0 |
Items are counted per habit per decision, and every item counts, not just the ones shown on the page.
Worked examples
| What the piles hold | Good | Bad | Score |
|---|---|---|---|
| Nothing at all | 0 | 0 | 50 |
| One habit avoided (overheard only) | 0.5 | 0 | 56 |
| One habit avoided (asked) | 1 | 0 | 60 |
| One confirmed bias | 0 | 2 | 33 |
| One avoided (asked) and one confirmed | 1 | 2 | 43 |
| Three avoided (asked), one habit confirmed in two decisions | 3 | 4 | 45 |
| Five confirmed, nothing avoided | 0 | 10 | 20 (14 before the floor) |
The five bars under the number
The bars group the 12 habits the interviewer asks about. Each habit takes its worst result across the decisions (confirmed beats unsettled beats clear). A bar's length is the share of its settled habits found clear. Only counts leave the server, never a habit's name, because the card is made to be shared.
| Bar | Habits in it |
|---|---|
| Exploring your options | Two options only, Status quo bias, Decision by deadline |
| Testing your belief | Confirmation bias, Advice from the choir, Information as reassurance |
| Calibrating your confidence | Overconfidence, The inside view |
| Learning from outcomes | Hindsight, Judging by the outcome, Sunk cost |
| Keeping your cool | Deciding hot |
| Word by the bar | When |
|---|---|
| Strong | Every settled habit in the bar is clear |
| Good | More than half are clear |
| Watch | Half or fewer are clear, but at least one |
| Needs work | None clear, at least one confirmed |
| Not settled | Asked about, but nothing was settled |
| Not tested | Never asked about |
Discrepancies
Ten places where what one part says does not match what another part does, most serious first. Each was checked against the code on main; none was reproduced on a real interview, so each says what the code would do, not what a person has seen.
| # | What is said | What the code does | Invented example | Confidence |
|---|---|---|---|---|
| 1 | The score's own notes: an overheard weakness never moves the number, because the person never had the chance to explain. That reason applies just as much to an overheard strength | An overheard strength does move it (+0.5). Overhearing can only push the score up | The detector hears five habits; the reviewer flips two to "the opposite". The person gains 1 point for things nobody asked, and loses nothing for the three it kept. A noisier detector gives higher scores | High on the code; unknown whether intended |
| 2 | The interviewer's prompt: "Only what you asked about can appear in their profile" | The report shows a "came up without being asked" section, and godspeed.md is built from those unasked habits when nothing was confirmed | Nothing confirmed, two habits overheard: the AI-assistant file is written entirely about habits nobody asked about | High |
| 3 | One habit, one card on the page | The score counts each habit once per decision; the intro counts only the cards shown (capped at 3 and 4) | Confirmation bias confirmed in two decisions: one card, but 4 bad points. Five biases confirmed: the intro says "Three patterns", the score counts five | High |
| 4 | Prompt and spec: a habit heard again in a later decision "still counts" | The heard-again tick is stored and never read by the reviewer, the report or the score | Overconfidence confirmed in the flat move, heard again in the job change: it counts once, and the job change shows nothing for it | High |
| 5 | One bias, one meaning | Each habit has five or six wordings, one per reader: the interviewer's, the detector's, the reviewer's, the page's, the writers' entry, the landing page's name | Advice from the choir: the page names it so, its definition on that same card starts "An echo chamber forms", the landing page calls it "Echo chamber", and the detector's text says "The user" and mentions "Godspeed's own view" | High |
| 6 | The detector prompt describes a page, a question and "the user" | The interview has none of these: three of the five sections are blank on every call, and each message is read alone, without the question it answers | "No, I didn't." is read with no idea what was asked. This is why the reviewer had to be added | High |
| 7 | A bias the interviewer asked about shows on the bars as asked | If the interviewer's quote differs from the person's words by even one word (capitals and punctuation are forgiven), the filing is thrown out as "no quote", and its bar reads "Not tested" | The robot fixes a typo or drops a word in the quote; the habit it spent three questions on reads "Not tested" | Moderate |
| 8 | A decision sent back is read by nothing (spec, invariant 24) | The report skips its turns, but the interviewer is still told "confirmed in earlier decisions" for habits filed in it, and the reviewer's list still takes its filings (then drops them) | The person starts on a decision they have not made yet; the robot confirms a habit, sends the decision back, and later skips that habit as "already confirmed" | Moderate |
| 9 | Your rulings of 28 September: vague odds hidden when a decision gives a number; anchoring never a main finding; two habits never on top on one finding | Written down in report-grouping.ts, applied nowhere |
A decision with "I gave it 70%" can still show vague odds as noticed | High |
| 10 | The spec and the code comments describe the current flow | Several lines are stale: "the next six" (the cap is 4, and the new report has no short form at all); open questions about close steps the code now owns; machine.ts cites invariant 21 for sent-back decisions, the spec numbers it 24; the report writer's prompt still explains the short form and decision counts, which the current report never sends |
A reader of the spec expects six short findings and finds none | High |
Two more facts that are not contradictions but surprise a reader:
- If every reviewer call fails, the person silently gets the older kind of report, ranked by how often the detector heard each habit, and no score.
- The detector judges all 42 habits on every message, but 7 of them can never reach the page, and the 30 that are not headline habits can only ever appear as "noticed".
Recommendations
My main recommendation: take the detector out of the report, and let the report and the score rest only on what the interviewer asked (confidence: moderate). Today two judges look at the same conversation, and a third robot is paid to settle their disagreements. Since 8 October your own rule is that the report judges only what was asked; the detector's output now feeds only the "noticed" section and half-points on the score, and it caused discrepancies 1, 2, 6 and 9.
The strongest case against: the interviewer asks about one habit at a time and misses things, and the detector is the only net for the other 30 habits. If you value the "noticed" section as a teaser that makes people want to come back, keep the detector, but on that section only and out of the score.
Walking it through, in order
- Question the requirements. Who asked for 42 habits when the interviewer can only investigate 12? Who asked for overheard strengths to score? Both predate the 8 October rule; neither has an owner today.
- Delete. The detector pass (or at least its part in the score). The old report shape kept as a fallback. The unapplied rulings in
report-grouping.ts. The short-form instructions in the writer's prompt that the current report never uses. The heard-again tick, unless it starts to count. - Simplify what survives. One wording per habit, read by every robot and shown on the page, instead of five or six. One counting unit everywhere (per habit, not per habit per decision), so the intro, the cards and the score always agree.
- Speed up the checks. A made-up interview, replayed through the whole pipeline on the fake model in seconds, that prints the intro's counts, the cards and the score side by side.
- Automate last. Turn that replay into a test that fails whenever the three counts disagree.
What I would do, ranked by how much of your time it needs
| Proposal | Why it matters | Rough cost | Your review |
|---|---|---|---|
| Make the score count only what was asked (overheard strength from +0.5 to 0) | Removes the upward tilt (discrepancy 1); one number in score.ts, a version bump to v0.2 |
~20 min, no paid calls | One yes or no |
| Count per habit, not per decision, in the score; intro counts everything | Card, intro and number agree (discrepancy 3) | ~45 min, no paid calls | Read one before/after table |
| Make heard-again count, or delete the tick and its prompt line | The prompt stops promising something the code ignores (discrepancy 4) | ~30 min; a prompt change, so a bench run (about 1 to 2 dollars) | Read the prompt diff |
| Fix sent-back leaks and the "no quote" bar (discrepancies 7, 8) | Small correctness fixes, no wording changes | ~45 min, no paid calls | None beyond the merge |
| Refresh the spec and the writer's prompt (discrepancy 10) | The documents say what the code does | ~30 min for the spec; the writer prompt needs a bench | Read the diff |
| Drop the detector from the report, or keep it for "noticed" only | Simpler pipeline, fewer robots, roughly one cheap call less per message | ~2 to 3 hours, plus a paid replay on invented interviews (about 3 to 5 dollars) | A decision, then one report to look at |
| One wording per habit across all readers | Ends discrepancy 5 at the root | Half a day; every prompt changes, so benches | Approve 12 to 42 texts: the largest ask |
One more thought on the score itself: with nothing asked it reads 50, which a person will read as "average". A card that says "not enough checks" below, say, three settled habits would be more honest than a number.