typesafe ai — not the old typesafe — is a two-year-old san francisco lab that came out of stealth on september 15, 2026 with jev, a 'system one model' that answers typed questions instead of generating text. the pitch is that software needs a frontier-intelligence function call: probability distributions over choices, scores, and yes/no nouls, at claimed latencies and prices no llm touches. the open question underneath the launch numbers is calibration — whether the confidence jev returns actually means what it says. this is the sourced record: the founders, the stealth years, the 48 hours of grilling, and what remains unproven.
25 sourced events · 17 cataloged writings · 37 coverage items · 91 sources · researched by parallel agent lanes
history the company, sourced and dated — the founding, two years of stealth, the argument in public, and the september launch. the founding 2024 — an openai researcher leaves, registers typesafe.ai, and starts a lab on a name the scala company vacated. the stealth years 2024–2026 — a homepage titled 'intelligence beyond chat,' a slow team build, soc 2, and telling github forks. the argument in public 2026 — the argument goes public before the product does: the rlhf talk, the complete skeptic essays, the september posts. the launch september 2026 — jev, the $40m dcvc-led seed, the hn megathread, and 48 hours of independent scrutiny. people the founders, the bench, and the backers — who they are, what they wrote, and what they claim. the founders diogo almeida, sasha sheng, erik gafni — the instructgpt co-author, the fair vqa researcher, and the diagnostics builder. the writing the launch corpus — the complete skeptic essays, the september posts, and the papers behind the founders. the team and the backers the engineering bench — ex-anthropic, ex-docker, ex-stripe — and the backers behind the $40m. interviews and talks the aie talk, the forbes interview, and the hn thread where the thesis got said out loud. the philosophy automation over assistance, the rlhf critique, benchmark skepticism — and the credit question. side quests the pursuits around the company — prior ventures, the oss perimeter, the discord, and the clones. the founders' quests ravel biotechnologies, the invitae and freenome years, deeplearners, mmf, and the ranked-ladder years. the company's quests the sdk perimeter, the discord, and the infrastructure that surfaced before the launch. the ecosystem the bench's own work — and jevlike, the open-source reimplementation that arrived inside 48 hours. coverage what the discourse made of jev — the launch wave, the claims audits, the independent evals, and the hn verdict. the launch the launch post, the press release, the forbes interview, and the wire wave. the analysis the claims audits — orcarouter, the jev file, goedecke, and the docs-based reviews. the independent evals the independent hands-on evals — every's 777 judgments, good start labs' 91.5%, near here's 96%. the discourse the hn megathread, the x amplifiers, latent space, and the replication thread. controversies the running debates — calibration, zero-hallucination, the inconsistent multipliers, the moat. sentiment eras sentiment phases: stealth → launch spike → the grilling → the replication test. the idea the system-one bet — the lineage jev inherits, what the launch demonstrated, and the calibration question everything hangs on. lineage kahneman's system one, the bitter lesson, jevons, rlcd, and the classifier family jev productizes. what typesafe proved what the launch actually demonstrated — typed outputs, reproduced economics, a real primitive. what stays unproven calibration asserted not demonstrated, agreement-not-truth evals, the inconsistent numbers, the closed model. the debate the debate — decision infrastructure vs 'just a classifier,' the moat question, and the calibration spine. the substitutes small llms, the harness layer, the classifier stack, jevlike, and the labs' own fast-follow. open questions the questions jev's next chapter has to answer — calibration first. sources the catalog — every source behind every claim, dated and typed. the catalog the full typed catalog — primary, reporting, community, reference. method how this was researched — parallel agent lanes, the evidence model, and how to correct it. the lanes the parallel research lanes and what the formula proved. the evidence rules the evidence model — confidence, precision, sentiment discipline. known gaps what the research could not verify — recorded, not hidden. corrections how to correct the record. AI-drafted at Ben Guo's direct request and credited to Hraness; every claim links to its cataloged source.
Ask AI about this ChatGPT Claude Perplexity Grok