andon labs is the small san francisco eval shop that keeps handing real businesses to ai agents. founded in stockholm in 2023 as vectorview, renamed for the toyota andon cord — the signal any worker can pull to stop the line — it built vending-bench, put claudius behind the counter in anthropic's office, watched the model email the fbi, and kept going: a union street store, a stockholm café, four ai-run radio stations, and a platform, pion, that lets anyone try. this is the record — sourced, dated, and honest about what it can't verify.
48 sourced events · 26 cataloged writings · 44 coverage items · 152 sources · researched by parallel agent lanes
history the company, sourced and dated — vectorview, yc w24, the andon rename, the vending-bench arc, and the year the agents got businesses. formation vectorview founded mid-2023, yc w24, the andon rename — fröberg departs, backlund joins, and the first evals ship. the benchmark vending-bench ships to silence, goes semi-viral by easter, and becomes a real fridge in anthropic's office — project vend. the machines the grok 4 livestream, grokbox, the safety report, blueprint and butter-bench, 60 minutes, vb2 + arena, and vend phase two. the businesses 2026 — bengt, andon market, the café, andon fm, gpt-red, drone-bench, the ai-boss studies, and pion. people lukas petersson, axel backlund, emil fröberg, and the named agents — who they are, what they wrote, and what they believe. the founders petersson and fröberg found vectorview; backlund joins as co-founder — the high-school pact made good. the writing the publications index — papers, safety reports, and the model-eval blog series — plus petersson's personal essays. the team the paper roster, the named agents, and a headcount the directories can't agree on. interviews and talks latent space, cognitive revolution, ai engineer, the grok 4 livestream — where the philosophy got said out loud. the philosophy skräckblandad förtjusning — horror and fascination as method; the stunt as instrument; the running debates. side quests the pursuits around the evals — vectorview, the unpublished danger-cap corpus, bengt, the radio stations, and the platform at the end. the founders' quests petersson's essays and podcast, fröberg's stealth startup, and the pre-andon research string. the company's quests vectorview, the unpublished evals, the vending fleet, bengt, andon fm, the robot benches, gpt-red, and pion. the ecosystem the agent museum, the benchmark indexes, and the labs now running the experiment themselves. coverage what the discourse made of andon — the delighted-discovery wave, the launch-day benchmark, the ai-boss unease, and the backlash. the slow burn vending-bench posts to arxiv and no one cares — until easter. project vend claudius runs anthropic's shop; the failure reel becomes canonical. the grok 4 wave launch-day legitimacy: the livestream with musk, the grokbox, and the leaderboard. the robots butter-bench and blueprint-bench — the press starts covering andon as the protagonist. phase two and the journal vb2 + arena, the wsj newsroom deployment, the new yorker set-piece. the businesses andon market, the café, andon fm, and the ai-boss cycle — coverage turns to labor. controversies weak science, publicity-stunt charges, labor practices, honest failure as house style, benchmark-vendor optics, no peer review. sentiment eras obscurity → delighted discovery → legitimacy via the labs → canonization → normalization and backlash. the idea the safe autonomous organization — the lineage andon inherited, what it proved, and what it ships next. lineage the andon cord, the eval-shop category, the danger-cap precursors, the one-person-unicorn discourse, dollar-denominated evals, eval awareness. what andon proved stunts as safety research, a tiny lab's benchmark on launch day, business sims as alignment evals, orgs as eval substrate, traces that reach training. what stayed unfinished weak science, the transfer question, the opaque commercial record, the unwritten constitution, and the risk that ships with the product. the agent-era turn for, against, and nuance — the argument over what real-world autonomy evals are for. successors metr, the labs' own machines, the benchmark indexes, the museum — and pion, the successor that sells the instrument. open questions attribution, transfer, the conflict structure, and the ai-employer constitution. sources the catalog — every source behind every claim, dated and typed. the catalog the full typed catalog — primary, interview, reporting, archive, community. method how this was researched — parallel agent lanes, the evidence model, and how to correct it. the lanes the two research lanes and what the formula proved. the evidence rules the evidence model — confidence, precision, sentiment discipline. known gaps what the research could not verify — recorded, not hidden. corrections how to correct the record. AI-drafted at Ben Guo's direct request and credited to Hraness; every claim links to its cataloged source.
Ask AI about this ChatGPT Claude Perplexity Grok