saved
Everything hackable will get hacked
gist
Malte Ubl argues that AI cyber offense is already here: unsafeguarded open-weight models such as Kimi K3 can map kernel attack surfaces and hunt microVM escapes, even if they failed to break Vercel Sandbox. Defenders still have a temporary edge because frontier models (Sol 5.6, not Fable 5) will do security review today, so waiting for Mythos-class access is paralysis. He wants teams to run deepsec-style codebase reviews now, then keep repeating them as open-weight offense catches up.
ideas
- The offensive gap is already open. Unsafeguarded near-frontier open-weight models such as Kimi K3 already do cyber research at Opus-class strength.
- A failed sandbox escape is still a research loop. Kimi did not leave Vercel Sandbox, but it mapped guest-kernel surface, privilege-escalation paths, and wrote a fuzzer; the right bug would finish the job.
- Waiting for Mythos is paralysis. Frontier models except Fable 5 already do defensive review; source-code access is the working signal of defensive intent.
- Spend the temporary advantage. Sol 5.6 still beats Kimi on defense; run full deepsec reviews now, especially for IDOR, XSS, and SSRF, and keep rerunning as open weights catch up.
- Treat AI offense as a budgeted loop. Vercel spends tens of thousands on quarterly full reviews plus PR checks, opened Hobby egress firewall, and is redirecting offensive models into a sandbox HackerOne.
quotes
“Kimi K3 is an Opus 4.X-class model with no relevant cybersecurity safeguards.”
“While none of this produced an escape from Vercel Sandbox, it did show Kimi conducting an investigation on its own”
“As defender, you have the better tool at your disposal, but you have to use it.”
“Defenders can already use stronger models than those broadly available for offensive work.”