method
how this was researched — parallel agent lanes, the evidence model, and how to correct it.
the lanes
the corpus was produced by four parallel agent research lanes — company history (69 events, 66 sources), founder writing and early posting, coverage and sentiment, and the unfinished idea (85 cataloged sources) — each writing a lane report against one shared scratch source catalog before synthesis. the pattern comes from the stripe-history formula: bounded collections, typed sources, and a review ledger. roam’s version drops the live-updating machinery (the company is quiet) and deepens the qualitative lanes: what the founder said, what the discourse made of it, and what the idea still owes.
what the formula proved this run: the source catalog is the real shared artifact — lane disagreements resolve as catalog entries, not arguments. where it needed help: a shared id convention upfront (lanes minted overlapping names that had to be reconciled), an explicit per-claim confidence field from the start, and a declared owner for deduplication so two lanes don’t archive the same thread under different ids.
the evidence rules
- every event, quote, and claim carries at least one source id.
- dates keep their precision — a year, a month, or a day — and approximate dates are labeled rather than sharpened.
- confidence stays honest: confirmed, reported, and inferred are distinct claims, not stages of the same claim.
- sources are typed by role — primary, interview, reporting, archive, community — so a claim’s foundation is visible before it is read.
- sentiment labels describe the coverage, not the truth of it; the same source can be evidence for both a hype wave and a backlash.
known gaps
- roam’s founding date is reported across 2017–2019 depending on whether you count the prototype, the company, or the public beta — this dossier keeps the spread rather than picking one.
- the white paper’s original roam-hosted url link-rots; the readable copy is a community mirror, and both are cited.
- no public current usage or revenue numbers exist; arr figures are community-reported and labeled as such.
- the twitter corpus comes through threadreader archives and the atlas.fm creator index — complete threads survive, but standalone tweets and engagement counts are partial.
- the reddit corpus is thin by design: the subreddit’s history is contested and largely unarchived, so moderation claims lean on conor’s own threads and contemporaneous coverage.
- trial length and believer-plan pricing drifted over time; the catalog keeps the prices each source actually recorded.
- several sources are future-dated relative to the events they describe — karpathy’s 2026 gist, the roam-tools changelog — and are treated as retrieved primary material, not normalized backward.
corrections
corrections and additional primary sources are welcome — cite the claim, the source, and what it changes. the catalog is designed to be audited, not just read.
AI-drafted at Ben Guo's direct request and credited to Hraness; every claim links to its cataloged source.