sufficient contextualization: what an ideal conclusion looks like

Honcho’s memory is not a transcript archive. It is a body of conclusions: products of reasoning—statements in plain natural language, each carrying one claim, derived from what a person says. Every conclusion is also a potential premise—an input to later reasoning that produces new conclusions the person never stated. The whole structure is a tree of logic: original messages at the base, conclusions built on them, further conclusions built on those. No statement in the tree is terminal; any of them may become load-bearing. Retrieval quality, reasoning quality, the ability to correct the record, and what a customer sees when they inspect their own representation all inherit from the quality of these statements.

A prior piece defined the extraction step: decompose what someone says into statements small enough to have a single truth value. That covers one half of the quality. This piece defines the other half—the one that is now the priority: sufficient contextualization. It states the ideal pure, as the aspiration and the roadmap. What the system does today, and how to measure the distance, is time-bound work—it lives alongside the work in Linear (see the end of this page).

The test

A conclusion is sufficiently contextualized when a future reasoning process can use it as a premise without ever seeing the conversation it came from.

That single sentence is the entire standard. Everything else in this piece is derived from it.

It matters that the standard is a test and not a taxonomy. A taxonomy of failure types—“unresolved pronoun,” “vague referent,” “relative date”—is useful for diagnosis, and this piece uses those labels freely. But a taxonomy makes a poor definition. Enumerate the failures and you have told the extractor (and the judge) exactly what to look for, which is also exactly what to miss: the next failure type is always missing from the list, and the list becomes a maintenance burden that grows forever. Define the standard as a test and both the model that writes conclusions and the model that judges them can regenerate every failure type— including ones nobody has named yet—from one principle. The prescription is exactly as tight as it needs to be and no tighter.¹

The test has a concrete form: the stranger test. Hand the statement, alone, to a reader with no access to the conversation—a new employee, or a model with an empty context window. Can they answer: who is this about? what exactly is being claimed? when does it hold? If they need the conversation to answer, the conclusion failed. It is not memory; it’s slop.

Why this one property carries so much

The cost of failing the test is paid three times.

Once in trust. People read their own representations. A representation studded with statements like “responded affirmatively” and “said hi” reads as noise, and users say so. The statement isn’t wrong—it’s inert, and inert entries erode confidence in every entry around them.

Once in reasoning. Conclusions exist to be combined—with each other, and with premises from entirely different conversations, months apart. A statement that cannot stand alone cannot combine: it has no hooks. Worse, it still occupies space. Every inert premise pollutes the context of every argument it enters, and because conclusions feed later conclusions, a defect at the bottom compounds upward through everything scaffolded on top of it.

Once in correction. Two statements can only confirm, supersede, or contradict each other if each says, on its own, what it is about. An under-contextualized statement is invisible to correction—a later, better statement about the same subject slides past it without collision, and the record never heals.²

This is why solving contextualization is the high-leverage move: retrieval misses, reasoning junk, uncorrectable errors, and user-visible noise are, to a large degree, four downstream symptoms of this one upstream defect. Solve this fully first, then see what remains.

The needle

The ideal sits between two failure directions. Naming both is what makes the standard usable.

Too little, and the statement is worthless. From a real production representation:³

courtland responded with ‘yes’ on July 2, 2026 at 2:13:45 AM

Note what this statement has: a named subject, perfect attribution, a timestamp precise to the second. Note what it lacks: the one thing that mattered—what the “yes” answered. Precision is not context. No future argument can use this premise, because it does not say what was agreed to.

Too much, and the statement stops being one statement. Pack the missing context in indiscriminately and you bundle several claims into one line. A bundled statement has no single truth value; and every extra detail is an additional claim that can independently be wrong, poisoning the whole unit—you can no longer correct one part without losing the rest.⁴

The resolution: specificity within atomicity. Added context that pins down what the claim means does not make a statement less atomic—it makes it more atomic, because it narrows the conditions under which the statement is true. “Alice is excited” is one vague claim; “Alice is excited about her February 2027 wedding” is one sharper claim. Context that fixes interpretation is required. Context that asserts something extra is bundling.

The check between the two is aboutness: does the addition sharpen what the statement was already about, or does it change what the statement is about? “Alice is excited about her February 2027 wedding” is still about exactly one thing—Alice’s excitement—now sharp enough to use. “Alice, who owns a dog, is excited about her wedding” is about two things: a second claim has been smuggled in as a subordinate clause, where it displays no truth value of its own and can never be corrected independently. Specificity sharpens the claim. Addition hides another claim inside it.

When context is needed to pin down who or what, prefer enduring descriptors— stable attributes that will still identify the referent months later—over incidental ones. “Bob, Alice’s boss at ACME” endures; “Bob, who was mentioned after the standup on Tuesday” does not.

Worked example. On October 14, 2026, Alice says: “my boss Bob at ACME is retiring next month.”

  • Too little: Bob is retiring. —Which Bob? Retiring from what? When? A stranger can do nothing with this.
  • Too much: Bob, Alice’s boss at ACME whom she has worked under for three years and gets along well with, is retiring next month. —Three claims wearing one truth value; and “next month” still drifts with the calendar.
  • Ideal: Bob, Alice’s boss at ACME, is retiring in November 2026. —One claim, uniquely resolved, anchored in absolute time.

The payoff shows up later, in a different conversation entirely. In January, Alice says she is thinking about going for an open management role. Combined with the ideal version, reasoning can form a real hypothesis—the opening may be Bob’s seat. Combined with “Bob is retiring,” it can form nothing: which Bob, whose boss, still retiring or long gone? Composability is not a bonus property of good conclusions. It is what conclusions are for.

What the test generates

Run the stranger test against typical statements and the same few questions fall out again and again. These are not a checklist to maintain—they are what the test produces, and a reader who forgets the list can regenerate it in one minute from the test itself.

About whom, exactly? Named and disambiguated—no orphaned pronouns, no “someone.” Production data shows this is the most common failure: in one scored corpus of 1,926 conclusions, roughly 8% were too vague to use—“is best friends with someone,” “asked someone if they got something,” and “is going to talk to a woman named Roy” (Roy is a man; the vagueness and the error travel together). The passing version costs one clause: “courtland and Vince Trost pivoted from web3 to AI after ChatGPT launched in October/November 2022”—a stranger knows exactly who, what, and when.

When a referent could be confused with another, the ideal resolves the confusion proactively. This production specimen has the right instinct—and fails the standard twice, which makes it worth dissecting:

courtland’s project ‘LexDOGE’ is a civic governance experiment in Lexington, KY; the name is a portmanteau of Lexington and DOGE and is not named after his dog Lex

First failure: it is a bundle. Count the truth values—what LexDOGE is, what the name derives from, what it is not named after: three claims sharing one line. Second failure: one of the bundled claims is false. Courtland has no dog named Lex; somewhere upstream an under-contextualized statement invented one, and the error has returned as context—a false rider with no displayed truth value of its own, uncorrectable without touching the whole statement.⁴ The molecular repair keeps the instinct and fixes both:

courtland’s project LexDOGE is a civic-governance experiment in Lexington, KY

the name ‘LexDOGE’ is a portmanteau of ‘Lexington’ and ‘DOGE’

Two statements, each with one truth value, each passing the stranger test—and the dog claim simply does not exist, because nothing said ever licensed it. Contextualization failures don’t stay where they happen; they come back as the context of future statements.

Claiming what—including the object. The act and what the act was about. A reply needs what it replied to; an emotion needs its object; a plan needs its content. “plans to do that tomorrow” fails twice in five words. The same requirement extends to scope: a preference recorded without the conditions under which it holds silently overgeneralizes. “Courtland prefers high-level, distilled summaries (TLDRs or ELI5s) when reviewing complex technical updates or PRs” is safe to reuse anywhere precisely because it says where it applies.

Anchored when? Absolute dates, always. “Today,” “tomorrow,” “next month” are different claims depending on the day they were written, and once stored, the writing day is not reliably recoverable—a statement anchored in relative time decays into ambiguity within a week.⁵ In the same scored corpus, about a third of all time references were left relative. The pair from that corpus: “michael is doing I.D. photos today” is meaningless within a day; “michael is hosting Casino Night on May 7, 2006” is usable forever. The fix costs a few tokens at derivation time and is impossible later.

With what force? Said, believes, might, did—preserved, not flattened. “michael said that he could murder Dwight” keeps “said” and “could” visible, so hyperbole reads as hyperbole; strip them and the record asserts something monstrous. The line between what a person said and what they did is force too: “said he prefers Python” and “chose Python unprompted in four consecutive projects” license different downstream bets, and the statement should preserve which one it is. Force also includes register—the same corpus stores “michael believes homeless people are very much alive”: every word faithfully kept, the joke entirely lost, a comedic beat archived as a belief. Hedges, attribution, and register are content.

And one generative consequence the lists miss: sometimes nothing passes the test. “hi” yields no statement a stranger could use. An assistant’s own suggestion reflected back is not the user’s claim at all. In some cases the ideal output is no conclusion. Note what is not being said here: there is no list of forbidden statement types, no rules about which utterances in which circumstances may not become conclusions—that would be the taxonomy again, by the back door. The test does the filtering by itself: if nothing mintable from an utterance passes, mint nothing. And when such an utterance does carry signal, the signal is what passes—a bare “hi” yields silence, while an uncharacteristically effusive greeting yields “X greeted Y with marked enthusiasm on 2026-07-10,” which stands alone because the claim is the enthusiasm, not the greeting. The ideal includes silence.⁶

Where context comes from

A conclusion is contextualized by resolving references, not decorating them—and resolution needs a source. The sources, in order of proximity: the surrounding conversation, prior conclusions about the same people, and the existing representation as a whole. This has a structural implication: derivation time is the only moment when the full resolving context is cheaply available. The conversation is present; the representation can be consulted. A week later, reconstructing what “that” meant requires archaeology. So the ideal system spends its contextualization effort at write time, once, rather than at every read forever.

Naive fixes trip on the switch point: the moment a conversation changes topic or referent. Mechanically inheriting the surrounding context there attaches the wrong context with full confidence—measurably worse than attaching none.¹ Contextualization is resolution against what the statement actually meant, not proximity-based decoration.


the distance

Measuring the current system against this ideal—the benches that run today, the DeriverBench proposal, the first molecular-bench baseline, and six conceptual directions—is deliberately not on this page: that material changes weekly, and this page should not. It lives alongside the work in Linear: Sufficient Contextualization: What an Ideal Conclusion Looks Like in the Deriver Conclusion Investigation project (ML-331).


All verbatim, from the two sources in footnote 3. The right-hand labels are what a judge applying the stranger test generates—illustrations of the test, not a taxonomy.

Pass:

SpecimenWhy it stands alone
courtland and Vince Trost pivoted from web3 to AI after ChatGPT launched in October/November 2022named parties, absolute anchor, one claim
Bob, Alice’s boss at ACME, is retiring in November 2026referent uniquely resolved with enduring descriptors; relative time made absolute
Courtland uses Tailscale (Tailscale IP 100.77.127.72) to remotely access the M1 MBP gateway while travelinga derived (never directly stated) statement that still passes: concrete, scoped, usable cold
Courtland prefers high-level, distilled summaries (TLDRs or ELI5s) when reviewing complex technical updates or PRsthe preference carries its own scope
michael is hosting Casino Night on May 7, 2006absolute date; usable years later
michael said that he could murder Dwightforce preserved—reads as hyperbole, not threat

Fail:

SpecimenWhat the test catches
courtland responded with ‘yes’ on July 2, 2026 at 2:13:45 AMthe object of the reply is missing; precision without context
courtland plans to do that tomorrowunresolved referent and unresolved time in five words
michael asked someone if they got somethingno recoverable referents at all
michael is going to talk to a woman named Royvague referent and wrong detail traveling together
michael is doing I.D. photos todayrelative time; false or meaningless within a day
michael believes homeless people are very much aliveregister lost; a joke archived as a belief
courtland is the CEO/founder of Plastic Extrusions (Honcho)fabricated entity; re-derived after correction because nothing could collide with it
courtland finds that clog Cove is doing really well with the Dialecticgarbage-in (speech-to-text garble memorized as an entity)—input quality, an axis of its own upstream of contextualization
Plastic Labs is EEO compliant— derived 4× in six minutesflat, context-free phrasing piling up as near-duplicates

Footnotes

  1. The formal grounding is Gunjal & Durrett, Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification (2024). Terminology map: what this piece calls sufficient contextualization, the paper calls decontextuality—interpretability of a statement removed from its context; molecular = atomic and decontextualized and minimal. Their minimality criterion is functional, not cosmetic: among all rewrites that uniquely fix a claim’s interpretation, prefer the one supportable by the most evidence—which is why enduring descriptors (“Bob, Alice’s boss at ACME”) beat incidental ones (Bob’s birthday). The switch-point warning is theirs too: in their experiments, naive context-adding scored worse than raw atomic facts immediately after a passage changed referents, because it confidently inherited the wrong entity’s context.
  2. Collision is mechanical, not magical: a contradiction check can only fire when two statements are recognizably about the same subject. In the production representation, the correction “…There is no parent entity named ‘Plastic Extrusions’…” and the error “courtland is the CEO/founder of Plastic Extrusions (Honcho)” coexisted without ever colliding—related content, no shared anchor. Part of contextualization’s payoff is rendering the same subject the same way, so that later statements land on earlier ones instead of sliding past them. A statement no future statement can land on is permanent, whether or not it is true.
  3. Two sources, both quoted verbatim throughout. (a) Courtland’s own production representation—Honcho workspace cotillion, 88,044 conclusions—sampled read-only via the platform API on July 10, 2026; the curated specimen set with full metadata is the internal vault note conclusion specimens 2026-07-10. (b) Abigail’s scored corpus of 1,926 conclusions derived from The Office test dataset—the Linear document Results from Office data, attached to DEV-1988 Honcho Conclusion Quality, which also carries the six-criteria rubric those numbers score against.
  4. This is the error-localization argument: if a bundled statement contains one wrong detail, the whole unit is wrong, and the true parts go down with it. The paper measures naive context-adding approaches producing ~10% bundled statements—a tax paid on every downstream use. The LexDOGE specimen above is the production version: one false rider (“his dog Lex”) inside an otherwise correct statement.
  5. Storage timestamps do not rescue relative time: conclusions are written in batches, so a statement’s stored time is ingestion time, not utterance time. “Tomorrow” plus a batch timestamp is still a guess.
  6. Silence is compatible with maximalism, not a retreat from it. The maximalist commitment is to extract all latent information—and an information-free utterance has none to extract. A statement minted from one adds noise, not coverage, and noise premises degrade every argument they enter.