hraness

the idea

the system-one bet — the lineage jev inherits, what the launch demonstrated, and the calibration question everything hangs on.

lineage

  1. 2011 → 2026system onekahneman's system 1 / system 2 split, branded onto inference: jev is sold as the fast, pre-attentive half and llms as the slow deliberate half. goedecke noted the psychology is 'partially discredited' but 'the branding works'; simonw reached for 'pre-attentive processing'; near here's eval framed jev as 'a system one for llm system twos.'seangoedecke.com, 2026-09-16latent space, 2026-09-16nearhere.events, 2026-09-16typesafe.ai, 2026-09-15
  2. 2019–2026the bitterest lessonsutton's bitter lesson pushed one level up: 'doing the right task > data > compute > algorithms.' the instructgpt anecdote is the evidence — a >100x smaller model trained on the right task (follow instructions) beat gpt-3 trained on the wrong one (predict the next token). rlcd is typesafe's bet that 'calibrated decisions' is the right task for machine consumers.typesafe.ai, 2026-09-10typesafe.ai, 2026-09-15completeskeptic.com, 2026
  3. 1865 → 2026the jevons betjevons paradox as naming thesis: when a resource gets cheaper, total consumption rises. jev is named for the economist — if a structured judgment costs 1/400th of an llm call, software asks orders of magnitude more questions. the faq answers 'who is jevons' before 'what is jev.'typesafe.ai, 2026-09-15typesafe.ai, 2026-09-16
  4. 2017–2026reinforcement learning for calibrated decisionsrlhf → rlvr → rlcd: the claimed third training regime. reinforcement learning for calibrated decisions optimizes the probability distribution itself rather than preference or verifiable answers — the mechanism behind 'we sell confidence.' the ceo ran post-training at openai; the method's details are undisclosed.typesafe.ai, 2026-09-15hacker news, 2026-09-15forbes, 2026-09-15
  5. 2018–2026zero-shot classification, productizedthe classifier family the community mapped jev onto within a day — deberta/gliner/gliclass-style zero-shot classification, bert-era encoders, 'productionized conformal prediction.' the ceo accepted the reframe in-thread ('exactly right!') while holding the architecture close. jev is the first of these sold as a frontier-priced api product.hacker news, 2026-09-15sesamedisk.com, 2026-09dailytexasnews.com, 2026-09-16
  6. 2023–2026structured outputs as afterthoughtthe harness layer jev bypasses — instructor, dspy, baml, grammar-constrained decoding: bolting typed outputs onto a generative model after the fact. jxnlco (instructor's author) amplified the launch; the contrast is that jev emits the typed distribution natively instead of constraining a sampler into producing one.latent space, 2026-09-16techmeme.com, 2026-09-16hacker news, 2026-09-15
  7. 2011–2016the other typesafethe name's prior life — typesafe inc., the scala/akka company founded 2011 by odersky, bonér, and phillips, renamed lightbend in 2016 and now does business as akka. the new company is unrelated; the name's meaning is the product claim (type-safe model outputs) and the search confusion is a live, self-inflicted cost.wikipedia, 2026globenewswire.com, 2016-02-23hacker news, 2026-09-15
  8. 2026the anti-benchmark positionthe anti-benchmark stance as method — 'lies, damned lies, and benchmarks' argued benchmarks are 'a shortcut to get around the hard problem of trust… a loan,' so typesafe shipped workflow evals measuring agreement with an llm panel instead of standard tables. legible as conviction; also conveniently the eval design jev looks best under.typesafe.ai, 2026completeskeptic.com, 2026-07agenccy.ai, 2026-09-16

what typesafe proved

  1. typed output is structural, not cosmeticthe type-safety guarantee is real and every serious reader grants it — jev never emits a malformed output, and the 0%-vs-8% chart (however gamed) points at a genuine property: the output contract is enforced by the model, not the parser. 'type safety is not factual correctness' is the caveat, not a rebuttal of the mechanism.typesafe.ai, 2026-09-15kingy.ai, 2026-09-15the-decoder.com, 2026-09-16hacker news, 2026-09-15
  2. the economics reproducethe speed and cost reproduce — three independent hands-on evals on launch day: mike taylor ran 777 real moderation judgments in under 0.7s for ~¼¢; near here measured ~0.59s average vs mistral small 4's 2.9s at ~8.6x cheaper; good start labs priced a million judgments at $160 vs ~$33k. the multipliers differ from marketing's, the direction doesn't.every, 2026-09-15nearhere.events, 2026-09-16goodstartlabs.com, 2026-09-15
  3. the decision primitive is reala machine-native decision primitive is a real slot — the 'zero-shot classifier as a service' consensus formed within a day and the ceo accepted it; goedecke called it 'a genuinely new computational primitive'; the per-turn economics argument ('technically possible and economically absurd' to verify every agent step with an llm) held up across independent reads.hacker news, 2026-09-15seangoedecke.com, 2026-09-16orcarouter.ai, 2026-09-16every, 2026-09-15
  4. the caveats were in the artifactlaunch discipline existed — the post self-flagged the subsidized-pricing caveat, the 'dumber than a frontier llm' limits, the agreement-based evals, and the asterisked comparisons; the ceo spent launch day conceding 'confidently wrong' in the hn thread instead of fighting it. whatever the marketing overshot, the caveats were in the artifact.typesafe.ai, 2026-09-15hacker news, 2026-09-15every, 2026-09-15

what stays unproven

  1. calibration is asserted, not demonstratedcalibration — the claim the whole pitch rests on — is publicly unverified. 'confidence metrics actually mean something' and 'a frontier-intelligence function call' require stated confidence to match outcome frequency; at launch there were no published calibration curves, no ece, no reliability diagram. hn's thduabmd asked whether per-answer calibration survives composition into workflows; mentlo asked whether it transfers across domains. the company agreed to share an internal calibration result; nothing public yet.hacker news, 2026-09-15typesafe.ai, 2026-09-16pasqualepillitteri.it, 2026-09-16progressiverobot.com, 2026-09-16typesafe.ai, 2026-09-16
  2. the evals measure agreementagreement is not ground truth — typesafe's workflow evals measure agreement with an llm panel, and on its own dashboard jev scores 67.8% vs gpt-5.1's 74.1%; zmmmmm read the first chart as jev below sonnet-5 accuracy. the independent evals are agreement-with-llm too: they show jev lands near the llm distribution, not that the distribution is right.agenccy.ai, 2026-09-16hacker news, 2026-09-15the-decoder.com, 2026-09-16orcarouter.ai, 2026-09-16
  3. which multiplierthe headline numbers don't agree with each other — 193.6x/444.6x on the homepage, 40–200x/40–400x in the launch post, 20–200x in the x thread, ~75x implied by the doom demo pair. dailytexas cataloged the spread; nobody has reconciled it.dailytexasnews.com, 2026-09-16typesafe.ai, 2026-09-16typesafe.ai, 2026-09-15x, 2026-09-15
  4. the closed modelthe model is closed — 'architecture is close to the chest for now'; no weights, no paper, ~32k context, text/json only, waitlist-gated. dailytexas's tarball forensics (jax/flax, embedding-layer only, ~800m-class) are the best public guess; every architectural claim in the discourse is inference, not disclosure.hacker news, 2026-09-15dailytexasnews.com, 2026-09-16actionbox.cloud, 2026-09-15
  5. a notch below frontier, except when it isn'taccuracy sits a notch below frontier on the vendor's own dashboard — orcarouter's 'mid-tier judgment,' 67.8% vs 74.1% aggregate. the wild card cuts the other way: near here's listing-moderation test had jev beat mistral and gemini outright (96% vs 84%/86%). the honest ledger: jev is mid-tier where the eval needs reasoning, competitive where it's pattern judgment — and 'just use a small llm' stays a live substitute either way.orcarouter.ai, 2026-09-16nearhere.events, 2026-09-16goodstartlabs.com, 2026-09-15agenccy.ai, 2026-09-16top5apps.ai, 2026-09-15
  6. the replication testthe moat is untested — goedecke argued the shape is 'one prefill forward pass plus a single token,' replicable by any lab; jevlike shipped an option-attention clone in 48 hours and a satirical two-hour qwen-rlcd video circulated. the counter — data, rlcd training, serving economics — is asserted, not proven.seangoedecke.com, 2026-09-16github, 2026-09-16hacker news, 2026-09-16hacker news, 2026-09-15

the debate

  1. forthe next category after chatbots and agents is decision infrastructure — the dcvc gp said it in the press release, the manifesto argues it as 'databases before sql,' and the per-turn economics make it concrete: software can't call a $0.01, multi-second model inside every loop, and at $42/billion it can.businesswire.com, 2026-09-15typesafe.ai, 2026-09-16orcarouter.ai, 2026-09-16every, 2026-09-15
  2. fora usable confidence signal is the missing primitive — if stated confidence tracks outcome frequency, software can branch on it (auto-accept above threshold, escalate below) and the last blocker to machine-consumed intelligence falls. this is the load-bearing version of the pitch.typesafe.ai, 2026-09-16typesafe.ai, 2026-09-16hacker news, 2026-09-15
  3. forjevons: at 1/400th the cost, demand isn't displaced, it's created — verification per agent turn, moderation per post, a judgment call per state transition. the naming faq is the thesis.typesafe.ai, 2026-09-15seangoedecke.com, 2026-09-16latent space, 2026-09-16
  4. againsttype safety is not factual correctness — a typed, fast, cheap wrong answer is still wrong, and jev can't abstain; 'yes, but not now' is a valid typed output. the guarantee covers the schema, not the world.kingy.ai, 2026-09-15hacker news, 2026-09-15the-decoder.com, 2026-09-16sesamedisk.com, 2026-09
  5. againstthere may be no moat — if jev is an ~800m encoder trained on distilled decisions, any lab can ship the shape in a quarter; jevlike already did the interface. the durable asset would have to be the calibration data flywheel, which is exactly the part nobody can inspect.seangoedecke.com, 2026-09-16github, 2026-09-16dailytexasnews.com, 2026-09-16hacker news, 2026-09-15
  6. againstit's mid-tier judgment at discount prices — 67.8% vs 74.1% on the vendor's own dashboard, misses that need reasoning, and the cheap substitute (a small llm with json mode) is 'good enough' for most of the workload jev targets.agenccy.ai, 2026-09-16orcarouter.ai, 2026-09-16nearhere.events, 2026-09-16pasqualepillitteri.it, 2026-09-16
  7. nuancecalibration is the question everything hangs on and it's open — the strongest version of the product requires per-answer and composed-decision calibration, neither of which has a public curve. read jev as 'a classifier with a confidence head pending validation,' not as proven decision infrastructure.hacker news, 2026-09-15pasqualepillitteri.it, 2026-09-16progressiverobot.com, 2026-09-16typesafe.ai, 2026-09-16
  8. nuancethe complement framing survives even if the infrastructure framing doesn't — good start labs' 'another reader, not the last word,' near here's 'system one for llm system twos': jev as the cheap first-pass filter in front of a model that can reason is useful at almost any accuracy.goodstartlabs.com, 2026-09-15nearhere.events, 2026-09-16seangoedecke.com, 2026-09-16

the substitutes

open questions

  1. does jev's stated confidence match outcome frequency on ground truth — will typesafe (or anyone) publish calibration curves, ece, or a reliability diagram, and does the promised internal calibration result ever go public?
  2. does per-answer calibration survive composition — when hundreds of jev calls chain into a workflow, do the confidences compound correctly or does the pipeline inherit the worst node's miscalibration (the thduabmd question)?
  3. does calibration transfer across domains — is a model calibrated on typesafe's eval suite still calibrated on your moderation queue, your triage rules, your corpus (the mentlo question)?
  4. the workflow evals measure agreement with an llm panel, not correctness — what does jev score against actual ground truth, and does anyone independent run that eval?
  5. what is jev — dailytexas's tarball forensics say jax/flax, embedding-layer only, ~800m-class, ~1.5gb q8_0; does the promised paper confirm it, and how much of 'frontier' survives disclosure?
  6. is the price real — the launch post itself flags 'we can't prove it isn't subsidized'; what happens to the 40–400x claim at scale, and does the training-time inference optimization scale the way the cost story needs (orcarouter's caveat)?
  7. is there a moat — jevlike cloned the interface in 48 hours; is the durable asset the rlcd training data and serving economics, or is the shape a quarter away for every lab?
  8. does a ~$200m seed valuation hold on 'another reader, not the last word' economics — is the market for machine-consumed decisions large enough, fast enough, to carry a frontier-lab valuation?
  9. what are the real limits — ~32k context, text/json only, no multimodality, waitlist access, can't abstain: how much of the 'decision infrastructure' market fits inside them?
  10. who else is on the cap table and the board — dcvc led, everyone else is undisclosed; at $200m post-money the governance story matters for a company selling trust.

AI-drafted at Ben Guo's direct request and credited to Hraness; every claim links to its cataloged source.