Cartograph thesis validation — final merged report (H1–H7)
Synthesized 2026-07-30. Two workflow runs (P1: H1–H5, P2: H6–H7), 16 agents total: one researcher + one adversarial auditor per hypothesis, plus synthesis. Every load-bearing claim below survived (or was corrected by) independent verification against primary sources; where an audit weakened a claim, this report states the post-audit position. Per-hypothesis detail with full audits: h1-*.md … h7-*.md in this directory. Confidence tags: high / moderate / low.
Digest — the short version
The thesis survives, but two of its three named pillars don't. "Its identity is closure" holds — though on different scientific footing than the notes assumed. "Its moat is the outcome-reviewed corpus" is refuted: OpenAI's Dreaming memory (June 2026) already does automatic outcome-updating for free, and to the extent personal context locks anyone in, it locks users to the labs, who hold more of it. What the labs will not build — because it adds deliberate user friction — is the ratification-and-closure workflow itself. The moat and the identity turn out to be the same thing: the ritual, not the data. Which means the corpus should be aggressively exportable, readable by any assistant — fighting portability is both losing and off-brand.
The three concrete product changes the evidence demands, in order of urgency:
- Record accept provenance from day one (silent-default vs. actively touched). Silent accepts are governed by inertia, not judgment — so without this field, the corpus silently fills with dispositions nobody judged, and the one real advantage over lab memory (human-ratified labels) evaporates. Cheap now, impossible to reconstruct later.
- Capture "what do you expect will happen?" at conclusion. "Reviewing decisions makes you a better decider" has no evidential support for life decisions — drop that claim entirely. What is supported: hindsight-bias protection from an ex-ante record, and expectation-vs-outcome comparison at review.
- Wire closure into recall-at-new-goal-creation as one mechanism, not two features. Four literatures converge on the same death shape — a review ritual that ends things without feeding the next one (the corporate post-mortem that gets filed and never opened; the Army's version works only because the last mission's review briefs the next mission). And the ceremony literature itself collapsed under audit (Zeigarnik's recall effect is 0.99 — nothing — in the 2025 meta-analysis; the Norton/Gino ritual papers are retracted or unreproducible in the fraud investigation's wake). What survives is better for the product anyway: the Ovsiankina effect — unfinished things nag to be finished, the way a half-written email pulls you back (not remembered better, which is the dead claim); attentional-residue relief — closing one thing releases the attention it was silently occupying, which is what lets the next decision get your whole head; and the goal-gradient pull — countable progress toward a visible end accelerates effort, the café loyalty-card effect.
Every effect and study named in this digest is unpacked — mechanism in plain words, an everyday example, and the Cartograph translation — in ACTION-PLAN-cartograph.md, the action-oriented companion to this report.
Method note: every hypothesis was researched and then adversarially audited by a second agent that fetched the cited sources and tried to break the claims. The audits caught real errors — a miscited paper in H7, an overstated portability claim in H6, and the invented 70%/85% acceptance thresholds in the original notes — so the surviving citations are ones a skeptic already tried and failed to break.
Executive summary
The thesis survives contact with the evidence, but three of its pillars need replacing:
- The moat is not the corpus (H6, refuted). The labs already do automatic outcome-updating of memory at zero user cost, and to the extent accumulated personal context locks anyone in, it locks users to the labs, who hold more of it. Defensibility must move from the data to the workflow.
- The scientific backbone of closure is partly rotten (H7, mixed). Zeigarnik-as-memory is dead (recall ratio 0.99 in the 2025 meta-analysis) and the ritual/ceremony literature is retracted or unreproducible (Gino scandal). What survives is better suited to the product anyway: the Ovsiankina resumption pull (open loops demand disposition, 67% resumption), attentional-residue relief, and the goal-gradient effect — countable, visibly-closable loops.
- The opt-out default poisons the moat it was meant to feed (H2 + H6). Silent accepts are governed by inertia or decayed vigilance, not judgment — and the corpus's one residual advantage over lab memory (user-ratified labels are higher quality than inferred updates) only exists if accepts were actually judged.
The convergent picture across all seven hypotheses: the defensible core of Cartograph is the functional closure→recall cycle — conclude with an expectation on record, get it surfaced at the next relevant decision — not ceremony, and not accumulated data. That cycle is the one shape with replicated evidence behind it (debriefs d≈0.79), the one thing the labs' friction-free memory won't replicate, and the documented survival mechanism for every system in the graveyard.
The findings most likely to change the product direction (re-ranked across H1–H7)
1. The moat must move from the corpus to the ritual (H6). The corpus-as-moat claim is refuted from both directions: labs do outcome-tracking automatically and for free (OpenAI "Dreaming V3", June 2026), and the lock-in narrative favors whoever holds the most context — the labs. Concrete change: reposition defensibility as the workflow the labs won't build because it adds deliberate user friction (ratification discipline, closure, trigger-based recall), and make the corpus aggressively exportable and readable by any assistant — fighting portability is both losing and off-brand.
2. The moat has a silent-corruption problem, and provenance is the fix (H2 + H1 + H6). Feedback-free default accepts are governed by inertia; when proposal quality is high, vigilance decays exactly then. Either way, a corpus of silent accepts is a corpus nobody judged — which also destroys the only residual advantage over lab memory (ratified labels). Concrete change: record accept provenance (silent-default vs actively touched) from day one, and spend the friction budget only on closure, goal edits, and outcome marking. Cheap now, impossible to reconstruct later.
3. Closure must feed the next decision or it dies as ceremony — and ceremony now has zero citable evidence of its own (H5 + H4 + H7 + H1). Four literatures converge: the AAR's documented death mode is a terminal wrap-up with no forward link; a saved record's only evidenced path to reuse is trigger-based surfacing at the next task; "useful capture," not capture, killed the IBIS family; and the ritual literature that would have justified standalone ceremony is retracted or null. Concrete change: wire the closure ceremony into recall-at-new-goal-creation as one mechanism, not two features — and ground it in Ovsiankina/attentional-residue/goal-gradient, not Zeigarnik or Norton/Gino.
Runners-up: (4) Drop the judgment-improvement claim; capture "what do you expect will happen?" at conclusion and shape held/revised/reversed review as expectation-vs-outcome — the only review structure with real evidence, and the hindsight-bias protection is the honest pitch (H4). (5) Budget proposals shown, delivered at breakpoints or batched at closure; keep rejection at ghost-text cost; treat acceptance rate as a within-system trend dial only — the 70%/85% thresholds are invented (H3).
Part I — P1 hypotheses (H1–H5)
H1 — The IBIS graveyard: does Cartograph escape the documented causes of failure?
Verdict: mixed (high confidence on the failure causes; moderate on whether Cartograph escapes them).
The graveyard is real and well-documented, and Cartograph's mechanic targets the right cause — but only half of it. The systems died of two things: (1) live capture cost (four measured cognitive subtasks: unbundling, classification, naming, structuring), which "AI classifies, human ratifies" genuinely removes, and (2) recorded rationale that nobody ever usefully consumes — which automation does nothing about. The survivors (NCR's 2,300+ captured decisions; commercial Dialogue Mapping) all had one shape: the person paying the capture cost got immediate value, usually via a trained human facilitator. Cartograph's bet, in the field's own terms, is replacing that human facilitator with an AI one. The two rigorous LLM-era tests (EchoMind and MeetMap, both CSCW 2025) say the AI absorbs the mechanical cost (94% extraction coverage, ~4× fewer user edits) but fails at judging what the conversation is about right now (28% precision when automating focus-switching) — and MeetMap adds that users who structure material themselves report better sense-making, so full automation may cost comprehension, not just vigilance.
Strongest sources:
- Buckingham Shum et al., KMI retrospective on argumentation-based rationale — https://kmi.open.ac.uk/publications/pdf/KMI-05-18.pdf — the field's own post-mortem: the four capture subtasks, Grudin's inequality, and "it is not merely a 'capture' problem, but 'useful capture'".
- Shipman & Marshall, "Formality Considered Harmful" (1999) — https://people.engr.tamu.edu/shipman/formality-paper/harmful.html — the direct intellectual ancestor: four failure causes, and its incremental-formalization remedy never demonstrated adoption either.
- Chen et al., EchoMind (CSCW 2025) — https://dl.acm.org/doi/10.1145/3757587 — first rigorous test of LLM live dialogue-structuring: extraction solved, direction/judgment not.
- Chen et al., MeetMap (CSCW 2025) — https://dl.acm.org/doi/10.1145/3711030 — AI-Map vs Human-Map: low effort vs hands-on sense-making, the exact tradeoff Cartograph is making.
- Conklin & Yakemovic, NCR field study (HCI 1991) — https://dl.acm.org/doi/10.1207/s15327051hci0603%264_6 — the one long-running success, achieved by near-zero capture disruption.
Strongest disconfirming finding: capture cost was only half of what killed these systems, and automation addresses only that half. A perfectly-filed archive nobody re-reads is still a write-only archive — so H1's escape depends entirely on H4 (re-reference actually happening), plus the AI's weakest measured skill is exactly Cartograph's "At stake" read (knowing what matters now).
If negative, do differently: treat re-reference (H4) as a survival requirement of H1, not a nice-to-have, and keep the human's judging role focused on what's at stake (where LLMs measurably misread intent) rather than ratifying classifications (where LLMs are ~94% right and ratification will decay into rubber-stamping).
H2 — Do opt-out defaults transfer to repeated in-flow ratification?
Verdict: mixed (high confidence on components; moderate on synthesis). The transfer the thesis assumes does not hold as stated — Johnson & Goldstein is the wrong citation.
Three findings, post-audit: (1) In repeated choices where the user experiences the outcome of each accept, the default effect disappears — acceptance tracks experienced proposal quality (88–89% when the default is usually right, 33–36% when its benefit is rare, even though it was mathematically the better choice). But — audit qualification — most Cartograph accepts produce no immediate experienced payoff, and in feedback-free streams defaults DO stick (classic 401(k)-style inertia). That's worse, not better: silent accepts are then governed by inertia, not judgment. (2) Even the flagship organ-donation result is shaky on its home turf: a 2025 24-country analysis found opt-out raised registrations but not actual deceased donations (+7%, non-significant) and reduced living donations (−29%). (3) Automation bias is robust, occurs in experts, and cannot be trained away — though the audit notes some rubber-stamping is rational reliance on a reliable system, so the design target is cheap recovery from the first failure, not elimination of rubber-stamping. Clinical alert fatigue (49–96% overrides; 2024 pooled estimate ~90%) is the worst-case picture.
Strongest sources:
- Roth, Waldman & Erev 2024, Judgment and Decision Making — https://www.cambridge.org/core/journals/judgment-and-decision-making/article/impact-of-experience-on-the-tendency-to-accept-recommended-defaults/48A391D40E7A5A54347BFDC757240F46 — repeated-choice defaults are governed by experienced payoffs, not the default itself.
- Güntürkün et al. 2025, PNAS Nexus — https://academic.oup.com/pnasnexus/article/4/10/pgaf311/8303887 — the Johnson & Goldstein citation overstates real-world effects even for one-shot defaults.
- Parasuraman & Manzey 2010, Human Factors — https://journals.sagepub.com/doi/10.1177/0018720810376055 — automation bias/complacency is structural and not fixable by training.
- Felisberto et al. 2024, Health Informatics J — https://journals.sagepub.com/doi/10.1177/14604582241263242 — ~90% alert override rate, persistent despite fixes: the ceiling of what a high-volume proposal stream becomes.
- JAMIA 2019 systematic review — https://academic.oup.com/jamia/article/26/10/1141/5519579 — interrupting modals were accepted less (38.7% vs 61.6%); tailoring to the right recipient was the only clear win.
Strongest disconfirming finding: the corpus problem is a pincer. Where accepts carry no experienced feedback, defaults stick by inertia (non-judgment); where proposal quality is high, vigilance decays precisely then (also non-judgment). Either way, high acceptance is indistinguishable from nobody actually judging — which is exactly the moat poison.
Do differently: drop the Johnson & Goldstein framing; record accept provenance (silent-default vs actively touched) so the outcome-reviewed corpus can distinguish ratified from rubber-stamped, and put deliberate friction only on the few high-stakes moments (closure, goal edits, outcome marking) while letting low-stakes accepts flow.
H3 — Is the 70%/85% acceptance threshold right, or is rejection cost the real variable?
Verdict: mixed, high confidence — the 70% threshold is refuted outright; the rejection-cost/timing replacement hypothesis is supported (moderate-to-high), with one self-correcting nuance.
No 70%/85% threshold exists anywhere in the literature — treat the figure as invented. GitHub Copilot runs at 27–30% acceptance with ~90% developer satisfaction; inline text suggestion systems run at 10–15% and nobody calls Smart Compose a notification tax. What actually separates tolerated from resented systems is interruptiveness and rejection cost: clinical alerts sit in the same low-acceptance range as Copilot with opposite valence, and the clinical literature's own dividing line is interruptive/modal vs passive/review-at-will (with the audit caveat that clinical stakes and liability also differ — not a controlled isolation). Every shown suggestion costs attention even when rejected (Quinn & Zhai), so the budget should cap proposals shown. The nuance: within a single system, acceptance rate remains the best available predictor of perceived value (ρ=0.24 — the best of a weak field), so a falling trend is still the right early-warning dial.
Strongest sources:
- Ziegler et al., "Productivity Assessment of Neural Code Completion" — https://arxiv.org/abs/2205.06537 — 27% acceptance, 2,631 developers; acceptance rate the best (but weak) within-system predictor of perceived productivity.
- GitHub/Accenture enterprise study — https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-in-the-enterprise-with-accenture/ — ~30% acceptance, ~90% fulfillment: high value at 70% rejection.
- Quinn & Zhai, CHI 2016 — https://dl.acm.org/doi/10.1145/2858036.2858305 — every suggestion shown costs attention even when rejected; preference and objective cost diverge.
- Iqbal & Bailey, "Oasis," TOCHI 2010 — https://dl.acm.org/doi/10.1145/1879831.1879833 — deferring interruptions to task breakpoints reduces frustration and cost.
- Fitz et al. 2019, Computers in Human Behavior — https://www.sciencedirect.com/science/article/abs/pii/S0747563219302596 — batched notifications beat streaming; fully off increases anxiety (single unreplicated field experiment — directional).
Strongest disconfirming finding: against the brief's own replacement hypothesis — acceptance rate is not a red herring; Ziegler found it's the single best within-system predictor of perceived value, so it stays on the dashboard even though no cross-system floor exists.
Do differently: budget proposals shown (not an acceptance target), deliver at conversational breakpoints or batched at closure rather than streamed mid-flow, keep rejection at true one-tap ghost-text cost, and monitor acceptance as a trend signal only.
H4 — Do people re-reference concluded records, and does outcome review improve judgment?
Verdict: mixed, moderate confidence — and the audit pushed it further negative: the one positive evidence base for "outcome review improves judgment" is now contested.
Three sub-verdicts. (1) Base rates are as bad as feared — high confidence: only 16% of bookmarked targets were retrieved via bookmarks (mostly the always-visible browser bar), bookmarked sites re-found no better than non-bookmarked; NASA managers "rarely consult" the lessons-learned corpus they're required to consult. Deliberate saving predicts almost nothing about re-use. (2) "Outcome review improves judgment" is now weak/contested (downgraded by audit): outcome feedback alone doesn't teach in noisy, delayed-feedback ("wicked") environments — and life decisions are canonical wicked environments. The one well-evidenced loop (Tetlock/Mellers forecasting training, ~10% improvement) required hundreds of fast, unambiguously-scored predictions — and a 2025 Psychological Science reanalysis of the same data found the training/teaming effects substantially eliminated or reversed. Decision-journal claims (the "19%" figure) trace only to product marketing; no primary study exists. (3) Trigger-based recall is directionally supported (moderate): just-in-time retrieval users consumed ~3× more stored documents than search-only users, because the system did the remembering — but the studies are small and from 2000, so directional only.
Strongest sources:
- Bergman, Whittaker & Schooler 2020/21 — https://journals.sagepub.com/doi/abs/10.1177/0961000620949652 — bookmarks: deliberate keeping structures don't get used and don't even help re-finding.
- NASA OIG via Nextgov 2012 — https://www.nextgov.com/people/2012/03/nasa-knowledge-management-database-used-rarely/205923/ — even mandated institutional decision corpora go unconsulted.
- Hogarth, Lejarraga & Soyer 2015 — https://journals.sagepub.com/doi/abs/10.1177/0963721415591878 — learning from outcomes requires kind environments; wicked ones produce the feeling of improvement without the reality.
- Hauenstein et al. 2025, Psychological Science — https://pubmed.ncbi.nlm.nih.gov/39630638/ — reanalysis eliminating/reversing the forecasting-training effect: the best judgment-improvement evidence is now contested.
- Rhodes, JITIR PhD thesis (MIT 2000) — https://www.bradleyrhodes.com/Papers/rhodes-phd-JITIR.pdf — proactive contextual surfacing makes stored information actually get used (n=12/13; directional).
Strongest disconfirming finding: life decisions are a wicked learning environment where outcome feedback doesn't improve judgment; the only evidenced improvement loop needed a volume and clarity of scored outcomes a personal life-decision corpus cannot generate — and even that loop is now contested by a published reanalysis.
Do differently: never promise "reviewing past decisions makes you a better decision-maker." The defensible payoffs are narrower and should be built explicitly: (1) trigger-based recall at new-goal creation — the system surfaces, the user never searches; (2) the ex-ante record as protection against hindsight/outcome bias (robust classics: Fischhoff 1975, Baron & Hershey 1988) — which argues for recording "what did you expect would happen?" at conclusion and comparing at review.
H5 — Which non-document forms already produce voluntary repeated decision review?
Verdict: mixed, moderate confidence — the audit slightly strengthened the positive claim.
The core loop exists in the wild in exactly two evidenced forms. (1) Structured debriefs/AARs: ~20–25% performance improvement (46-sample meta-analysis, d=.67), independently replicated and strengthened by a larger 2021 meta-analysis (61 studies, d=0.79) — whose moderators (alignment + objective performance-review media) also independently confirm the central limit below. But corporate AAR transplants characteristically degrade into pro-forma ceremony (Senge: "people reduce the living practice of AARs to a sterile technique") — authoritative assertion, not adoption statistics; the failure is transplant fidelity, not technique efficacy. (2) Auto-generated closure artifacts in game cultures (roguelike morgue files, engine game review): genuine zero-user-cost closure documents produced and shared at scale — existence proof, not experimental evidence. The crucial correction to the brief's premise: millions consume machine-generated recaps; effortful deliberate review stays an elite minority behavior even in chess, where every enabling condition is maximally favorable (free perfect oracle, objective outcome, bounded state — and serious solitary study is still what separates grandmasters, ~5,000 hours, ~5× intermediates). The gamification-backfire literature does NOT indict game form (bounded state, endings, replays, progress tracks — no documented backfire); the poison is in reward loops (points/streaks) and mandatoriness.
Strongest sources:
- Tannenbaum & Cerasoli 2013, Human Factors — https://journals.sagepub.com/doi/abs/10.1177/0018720812448394 — debriefs improve performance ~20–25%: the strongest experimental evidence anywhere in H5 for a repeated outcome-review ritual.
- Keiser & Arthur 2021, J. Applied Psychology — https://doi.org/10.1037/apl0000821 — larger replication (d=0.79); its "objective review media" moderator confirms the objective-outcome-signal limit.
- Darling, Parry & Moore, "Learning in the Thick of It," HBR 2005 — https://hbr.org/2005/07/learning-in-the-thick-of-it — why AAR transplants die: reduction to a terminal ceremony with no forward link.
- Charness et al. 2005, Applied Cognitive Psychology — https://onlinelibrary.wiley.com/doi/10.1002/acp.1106 — even under ideal conditions, effortful review is the scarce elite behavior, not a mass one.
- Levitt, NBER w22487 — https://www.nber.org/papers/w22487 — coin-flip study: people nudged into change on major decisions were happier at 6 months (status-quo bias in life decisions is real; caveated sample).
Strongest disconfirming finding: the one condition that makes game review self-correcting — an objective, fast, individually-attributable outcome signal — is unavailable in principle to a life-decision tool; life outcomes are noisy, slow, small-N, no counterfactual. And the only transplant with adoption evidence (AAR) works through structure + facilitation + link-to-next-action, and dies precisely when reduced to ceremony.
Do differently: build closure as a cycle, not an event — the system writes the entire closure artifact for free at the moment of ending (morgue-file model, already Cartograph's architecture), and outcome review is triggered by and linked to the NEXT decision (the AAR stick-factor), which independently validates recall-at-new-goal-creation over any standalone "review your journal" surface. Do not build the closure ceremony as a terminal ritual — that is the documented death shape.
Part II — P2 hypotheses (H6–H7)
H6 — Is an outcome-reviewed personal decision corpus a real moat against the labs?
Verdict: refuted — confidence: moderate. The corpus is at best a modest switching cost, not a moat. Two independent mechanisms both point the same way. First, the labs already do outcome-tracking for free: OpenAI's "Dreaming V3" memory (announced June 4, 2026) automatically revises memories as time passes — "you're going to Singapore" becomes "you went to Singapore" — at zero user effort, with free-tier rollout announced. Second — the audit's sharper reframe — the prevailing mid-2026 analyst view is that personal AI context does create real lock-in, but that lock-in accrues to whoever holds the most context, which is the labs, not a startup with a smaller slice. Either way the moat claim fails: if context doesn't bind, Cartograph's corpus doesn't either; if it binds, the labs' bigger context binds harder. The audit weakened one supporting leg (Anthropic's "memory import" is a copy-paste prompt hack, not proof of an industry portability norm) but the verdict survives without it.
Strongest sources:
- OpenAI, "Dreaming: Better memory for a more helpful ChatGPT" (June 4, 2026) — openai.com/index/chatgpt-memory-dreaming/ (page 403s to fetchers; corroborated by 5+ independent write-ups, e.g. https://letsdatascience.com/news/openai-upgrades-chatgpt-memory-architecture-for-fresher-pers-b26b51d5) — labs now do automatic outcome-updating with zero user work, collapsing "labs won't do outcome review."
- Casado & Lauten, "The Empty Promise of Data Moats," a16z, May 2019 — https://a16z.com/the-empty-promise-of-data-moats/ — data accumulation rarely defends on its own; defensibility lives in workflow and switching costs (single-user corpus has no network effect at all — an acknowledged extrapolation from the essay's multi-user scope).
- Cloverpop pivot post-mortem — https://medium.com/@Lonsequitur/cloverpop-s-pivot-helps-businesses-make-smarter-decisions-d6097421a9a — the one venture-scale personal decision tool abandoned consumers in 2015 because demand for personal decision help didn't grow (n=1, pre-LLM, but the only data point that exists).
- 9to5Mac on Anthropic memory import/export (March 2, 2026) — https://9to5mac.com/2026/03/02/free-claude-users-can-now-use-memory-and-import-context-from-rivals/ — facts verified, but audit downgraded it: the import is a copy-paste prompt, and the dominant contrary narrative (memory as the labs' moat, e.g. https://bdtechtalks.substack.com/p/openais-moat) refutes H6 harder via lock-in-favors-the-labs.
- Casey Newton, "Why note-taking apps don't make us smarter," Platformer 2023 — https://www.platformer.news/why-note-taking-apps-dont-make-us/ — in the closest observed market (tools for thought), large personal corpora didn't even function as switching costs; users churned anyway (moderate: pattern evidence, no public retention numbers).
Strongest disconfirming finding: Dreaming V3 plus the memory-lock-in narrative attack both halves of the claim at once — the labs get most of the outcome-tracking value automatically and for free, and to the extent accumulated personal context binds users at all, it binds them to the labs, who hold more of it. The residual advantage (user-ratified held/revised/reversed labels are higher quality than inferred updates) is real but is a quality difference, not a defensibility mechanism, and no evidence was found that users pay or stay for it.
Do differently: Move the defensibility story from the data to the workflow — the closure ceremony, ratification discipline, and trigger-based recall are "the ritual the labs won't build because it adds user friction" — and make the corpus aggressively exportable and readable by any assistant, since fighting portability is both losing and off-brand.
H7 — Does the Zeigarnik effect survive replication, and does ceremony have evidence?
Verdict: mixed — confidence: high. The split, verified at primary sources by the audit:
- Zeigarnik as a memory effect: refuted. Ghibellini & Meier's 2025 meta-analysis finds a recall ratio of 0.99 for interrupted vs. completed tasks — no effect; the authors say it "lacks universal validity."
- Open loops as a resumption pull (Ovsiankina effect): supported. Same meta-analysis: 67% resumption rate across 21 publications, robust and consistent. This is what Cartograph actually needs — open loops demand to be closed, not remembered.
- Ritual/ceremony: currently uncitable. The flagship literature (Norton/Gino/Brooks, HBS) is wrecked by the Gino fraud investigation: Brooks et al. 2016 ("rituals decrease anxiety") is formally retracted, Tian et al. 2018 is retracted, and Norton reports he cannot reproduce the grief-rituals results for any of its four studies. The one independent pre-registered test (Karl & Fischer 2018) found no direct anxiety-suppression from ritual.
- Masicampo & Baumeister 2011 (plan-making defuses intrusive thoughts) exists as described but has no located direct replication — suggestive, not established.
- One audit downgrade: the "one-shot rituals do nothing" point from Hobson et al. 2017 measured intergroup bias, not personal benefit — so "ceremonial weight comes from repetition" is an analogy, not evidence.
Strongest sources:
- Ghibellini & Meier (2025), Humanities and Social Sciences Communications — https://www.nature.com/articles/s41599-025-05000-w — meta-analysis of 59 publications: Zeigarnik recall ratio 0.99 (dead), Ovsiankina resumption 67% (robust); fetched and verified, no published rebuttal found as of 2026-07-30.
- Brooks et al. 2016 retraction, OBHDP — https://www.sciencedirect.com/science/article/pii/S074959781630437X — the standard "rituals improve performance by decreasing anxiety" citation is formally retracted.
- Many Co-Authors registry, Gino entry #82 — https://manycoauthors.org/gino/82 — Norton answers "No" on ability to reproduce the grief-rituals results for all four studies.
- Karl & Fischer (2018), Human Nature — https://link.springer.com/article/10.1007/s12110-018-9325-3 — the only independent pre-registered test (N=180): no direct support for ritual suppressing anxiety.
- Adjacent effects that can carry the design instead: Leroy 2009 on attentional residue (completion aids disengagement, OBHDP 109:168–181) and Kivetz, Urminsky & Zheng 2006 on the goal-gradient effect (effort accelerates near a visible endpoint, JMR) — both moderate-high confidence, both supporting countable, visibly-closable loops.
Strongest disconfirming finding: Both named pillars fail scrutiny simultaneously — Zeigarnik's recall advantage is 0.99 (nothing) in the only modern meta-analysis, and the entire ceremony literature is retracted, unreproducible, or fraud-adjacent, with the sole independent pre-registered test coming up null. As of mid-2026 there is no citable evidence that a brief closure ceremony produces psychological benefit.
Do differently: Keep the closure feature but swap its foundations — drop Zeigarnik and every Norton/Gino ritual citation, and ground closure in the supported chain: Ovsiankina resumption pull (open loops demand disposition), attentional-residue relief from marking things finished, goal-gradient motivation from countable progress, and plain interaction design (recognition over recall, explicit endings), which needs no psychology citation at all. Treat "ceremonial weight builds through repeating the same closing move" as a design hypothesis, not an evidenced claim.