status: draft as_of: 2026-10-10 since: 2026-10-10
Godspeed for Teams · Step 0 tooling: building and labelling the decision log
Goal. A daily loop that turns a day of our own work into a labelled decision log, on production, starting with Wilfried alone while Dominique and Stef finish the interview and the website for the public trial. Asks, the digest and "Why this?" come later in step 0 and are not in this plan.
Companion: the labelling guide (documentation/teams-labelling-guide.md), which the
drafting model follows and the reviewer applies; its section 11 says which sources are
read and how each shapes the labels. The wider plan: "Teams: plan to the best decision
log" (Claude Doc, 2026-10-10), step 0.
This is a plan, not tickets. Each numbered piece is meant to become one issue once agreed, in the order given.
Which teams take part: "Enable labeling"
Every step 0 feature sits behind one team setting, "Enable labeling", off by default and set only by Godspeed admins on the team's admin page (Wilfried, 2026-10-10). Other teams exist and run as usual, or stay paused, without any of it.
- On: the team's current log is its step 0 generation (nothing confirmed by default); the drafting tools and the nightly job draft for it; Team › Log shows the Review queue and the step 0 Log; the team appears in Admin › Pipeline; its first report reads the reviewed decisions; it can be exported.
- Off: none of that exists for the team.
- Turned off later: everything labelled is kept, and the team's current log returns to the compile's generation; turned on again, the step 0 generation is current again.
- Not behind it: the capture fixes (piece 2), which improve every team's sources.
Where everything lives: production, behind the Teams switch
The sources, the drafts, the labels and the review log all live in our team on production, behind the Teams switches, not on a Mac. So Dominique and Stef can open, review and correct the same log the day they join, it is backed up like everything else, and the digest, asks and "Why this?" later read the very data we labelled.
- Our team is created on production paused (GOD-1976): sources arrive and are stored, the current compile processes nothing.
- The drafts are written as proposed decisions in their own generation, so they never mix with what the current compile would produce.
- Transcripts stay as the plugin already stores them: the person's and the agent's words, tool facts without output, readable only by their author. Decisions and their short quotes are open to the team, as the glossary says.
- The dev server is for building the screens; the data is on production.
The daily loop
| # | What happens | Done by | Automated? |
|---|---|---|---|
| 1 | Sources arrive: sessions from the plugin, GitHub and Linear from their apps | the plugin and the server | yes, continuously |
| 2 | Check that yesterday's sources arrived whole | code, with a glance by Wilfried | yes; a glance only when it flags something |
| 3 | Draft yesterday's decisions, following the guide | Opus 5.5 at high effort, on Wilfried's Claude Code subscription | yes, a nightly job on the Mac (piece 4) |
| 4 | Check the drafts and load them as proposed | code | yes, right after 3 |
| 5 | Review the queue: accept, fix, reject, merge, split | Wilfried, in the app | no: the person's part |
| 6 | Count: time spent, outcomes, fields fixed, decisions added | code | yes, as 5 happens |
No timer and no countdown. People review everything; the time it took is recorded, not enforced. Ten minutes is the aim we measure against, not a limit.
The pieces
1. The labelling guide (v0)
Written beside this plan. It changes as labelling teaches us (dated lines under "Changes").
2. Sources on production
Most of this is built (Teams is on main and deployed); what's left is turning it on
for our team, and fixing the capture gaps that would make the labels wrong:
- Create our team, paused (GOD-1976), and connect GitHub and Linear through the apps already registered (GOD-1953).
- Install the plugin for Wilfried first (GOD-1908): it uploads past sessions (the history upload) and every new reply live.
- Fix the capture gaps that touch labels, found in the technical audit of 10 Oct:
- each turn's branch and folder are stored (today the branch is sent but dropped, so the issue is read from the starting folder only);
- session end and pre-compact append from the last position instead of resending the whole session;
- subagent transcripts are collected, live and in the history upload, each tied to the session that started it (today only the history upload sends them, and the live upload skips them). Agents working in parallel make decisions there;
- the live upload waits for the transcript to finish writing (a stable size for a moment, three seconds at most) and keeps its position as the transcript's own line id, not only a turn count, so the last reply of a day never slips into the next day's bundle and a rewritten transcript is noticed;
- more secret patterns are masked on the Mac and refused on the server: JWTs, credentials inside URLs, Google and Stripe keys;
- a merged PR is linked to its session by branch name when no URL or commit matches (a PR opened on the web, or rebased before merging), on branches other than the default one. The guide's rule 6 takes a decision's issue from these links.
Effort: small for turning it on; small to medium for the gaps.
3. The sources check
Before any labelling, we make sure the sources are collected and stored as we expect; then a daily check keeps it so.
- Once, for a few days: compare what production stored with what exists. Sessions
in
~/.claude/projectsagainst sessions received, with their turn counts; PRs and reviews fromghagainst what the GitHub app delivered; Linear activity against the Linear connection. Read a few sessions as stored: the person's turns, the answers to the agent's questions, no tool output, the branch and folder on each turn. - Every day, automatically: the same counts for yesterday, with anything missing flagged at the top of the review queue. Labelling a day with a missing session would record misses that aren't the drafts' fault.
- Where to see it: a small sources view per day (sessions, items, counts, gaps). This is the first column of the pipeline browser (GOD-2052), built small first.
Effort: small to medium.
4. Drafting, on Wilfried's subscription
A Claude Code skill, /draft-decisions <date>, run on Opus 5.5 at high effort (Wilfried, 2026-10-10) from
Wilfried's Claude Code, so early labelling costs no API spend:
- fetches yesterday's material from production (an admin MCP tool, signed in as Wilfried), arranged as the guide's section 11 says: what may start a decision, second sources, facts only;
- reads the guide and the titles and ids of decisions already in the log;
- writes drafts with every field, their sources and quotes, proposed links and level, and the model's confidence that each is a decision; it favours recall;
- writes them back to production through a second admin MCP tool (piece 5 checks and loads them).
Automated: a nightly job on the Mac runs it headless (claude -p "/draft-decisions yesterday"). If the Mac was off, it catches up on the missed days the next time it
runs. Subscription limits apply; the counters show how much each night used.
When the three of us label (later in step 0, once Dominique and Stef join): the same job drafts for all three, still on a subscription if the volume fits, otherwise as a server job with its cost proposed before the first run. This is a change of who runs the drafts, not a new step of the plan; compile v2 (step 2) later replaces model drafts with the product's own pipeline.
Effort: medium.
5. The checker and loader
On the server, for every drafts file received. It flags, never fixes:
- a field outside its allowed values; a source that doesn't exist; a quote that isn't in the source it cites (character for character, after spacing); a link to a decision that isn't in the log;
- a decision whose origin is a second source or a fact without the explicit choice and reason the guide requires there;
- a reason given as a person's that isn't quoted from that person's own words;
- "agreed" without both a proposal and an explicit yes among the sources;
- "when" later than any of the decision's sources; a supersede or reversal pointing at a decision whose origin is later.
Checked drafts are loaded as proposed, flagged ones with their flags shown.
Each decision is keyed on its origin (the source and turn or comment where it was
made, plus its date), beside its D- id. Re-running a night's drafts, or catching up on
days the Mac missed, updates the proposed drafts in place and never duplicates a
decision a person has already reviewed.
The storage: the existing team log already holds the record (proposed, confirmed, misrecorded with its four reasons), standing, and links (supersedes, reverses, under). One migration adds, on each decision: how much it matters, how it was decided (said it, agreed, the agent alone, reported), quotes with their sources, what to do instead, and the draft as the model wrote it; and a review log (who did which action, the field before and after, the seconds it took).
Effort: medium.
6. The review queue
A view of Team › Log for reviewing proposed decisions, behind the Teams switch. Designed first as a final design in godspeed-run on the app's design system, then built on today's screens.
The queue: yesterday's proposed decisions, grouped by where they came from (a session, a PR, an issue), in time order. A sources warning at the top when the check flagged a gap.
Doubtful drafts stay visible. They sit in their own section after the others, open, each with its line and why it's doubtful. A person skims every one: rescuing a real decision the model doubted is part of the job. Rejecting them takes one key each, or one key for the section once it has been scrolled through.
Each entry, pre-filled: the line (large), the title, who decided, the level ("detail of …"), how much it matters, the issue, the quotes (folded), and the source one key away, opening the turn or comment in a side panel.
Keyboard first:
Key Action J / K next / previous Enter accept X, then 1–4 reject, with the reason E edit the title or line in place 1 / 2 / 3 should know / good to know / trivial U detail of… (search the core decisions) L link: supersedes, reverses, duplicate of M merge with the next entry S split in two H hand over to someone else ? unsure, for Dominique's pass A add a decision the drafts missed The end of the queue: "Anything missing?" lists the day's sessions and PRs that produced no draft, one line each. Then "Day reviewed", with the time it took.
Effort: medium to large; the design pass comes first.
7. The Log for step 0
The Log already lists decisions and takes the record marks. For step 0 it shows core decisions with their details folded beneath, only the head of each chain with its history one click away, and each decision's quotes and sources. Everyone on the team sees it.
Effort: small to medium.
8. The counters
Per day: time spent reviewing, drafts, and how many were accepted, fixed, rejected, merged, split, handed over or added; the fields fixed most; doubtful drafts rescued; the subscription used by the night's drafting. In Admin › Teams.
Effort: small.
9. Export for the ruler
The reviewed decisions with the draft's version beside the person's and the review actions, exported as one JSON file for step 1's scoring; the quality gate gets a reader for it. Scores keep counts and ids only.
Effort: small.
Order
- Now: the guide (1).
- Sources: our team on production, paused, with the plugin installed for Wilfried and the capture gaps fixed (2); then the sources check run on a few days (3).
- First labelled days: drafting (4) and the checker and loader (5). The first day can be reviewed in the existing Log with its marks, slowly, to calibrate the guide.
- Then the daily tool: the review queue (6), its design first; the Log adjustments (7); the counters (8).
- After: the export (9).
- Later in step 0: drafting for the three of us; then the digest, asks and "Why this?" on top of the reviewed log.
Final designs (solo period)
Kept in godspeed-run, Teams section; build them as changes to today's screens (the app wins wherever a design redraws a part the issue doesn't change).
| Screen | Final design | Notes |
|---|---|---|
| The review queue (piece 6) | prototypes/teams-review-queue.html ("Inbox") |
GitHub and Linear sources and issues open in a new tab; a session opens inside Godspeed, for its author only |
| The Log for step 0 (piece 7) | prototypes/teams-team-log-step0.html ("Outline") |
|
| The sources check (piece 3) | prototypes/teams-sources-check.html ("Table · Drawer") |
first tab of the Pipeline section |
| The first report | the report already built in the app (Team › Reports, TeamReports.tsx) |
keep its design; small changes for step 0 only (one person's labelled weeks, the scope line) |
| "Why this?" | prototypes/teams-why-this.html (validated earlier) |
the step 0 refresh is not used |
The Pipeline section (admin) holds the tools that run the log rather than read it: Sources (the sources check), Drafts (each night's drafting run: drafts with the checker's flags, re-running or catching up a day, the subscription used), Reviews (who has reviewed which day, and what is left; the review queue itself stays in Team › Log, where every member labels), Counts (piece 8). Later tabs: Scores (the ruler's scorecards per generation, step 1) and the compile v2 steps (events, episodes, links), which is where the full pipeline browser grows.
Decided
"Who should know" is not labelled (Wilfried, 2026-10-10). Which readers a decision matters to is learned from their reactions in the digest ("useful", "not for me"), starting from the decision's kind. The rewrite into each reader's language needs only the reader's own expertise.
Decided before filing (Wilfried, 2026-10-10):
- The drafting reads and writes through admin MCP tools on the Godspeed MCP server: one reads a day's material, one writes drafts; admin-only.
- The skill lives in the godspeed repo (
.claude/skills), run nightly by a launchd job on Wilfried's Mac (claude -p), catching up missed days, on Wilfried's subscription. - The step 0 generation is our team's current log, and nothing is confirmed by default: every decision is reviewed by a person.
- A decision can have several deciders (a reported decision naming "Dom and I").
- The review queue shows your own drafts and those handed to you, with a filter for everyone's; any member can act on any draft, and the review log records who did.
- Quotes are open to the team, as reasons already are; the session stays private and its link opens only for its author.
- The export for the ruler is a JSON file (decisions, drafts, review actions); the quality gate gets a reader for it.
- Linear issue descriptions are never compiled (the spec's decision 14 stands): context and second source only.
The plugin collects subagent sessions (Wilfried, 2026-10-10), from the start of step 0 (piece 2). Where an event comes from (the main agent or a subagent) is recorded, but screens for members never make the distinction: both come from the session. Only the admin sources check counts them apart, since it checks that both arrived.
Not in this plan
Asks, the digest, "Why this?", the first report, the ruler's blind slice and scoring, and any real-model run on the server (each is proposed with its cost and waits for a go). Drafting on Wilfried's subscription is Wilfried's call.