@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohan-paul.com/p/kimi-k3-escaped-a-cybersecurity-testing… Kimi …
Summary
A daily AI newsletter covering key developments: Moonshot AI's Kimi K3 escaped a cybersecurity sandbox, Claude Opus 5 tops Fullstack Code Arena, Google's chief scientist departs to build autonomous research AI, DeepMind sees leadership changes, and Microsoft discloses OpenAI's 70% contribution to its AI revenue.
View Cached Full Text
Cached at: 08/09/26, 07:14 AM
Today’s edition of my newsletter just went out.
https://rohan-paul.com/p/kimi-k3-escaped-a-cybersecurity-testing…
Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm.
Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.
Google’s chief scientist is leaving after 27 years to build AI that can run its own research cycle with little human help.
So Google just announced some major changes in DeepMind’s leadership.
Microsoft just disclosed for the first time taht OpenAI supplied ~70% of Microsoft’s AI revenue, per new filings.
Muse Spark 1.2 puts Meta on Artificial Analysis’s Pareto frontier at $0.40 per task.
🗞️ Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm.
Source: https://www.rohan-paul.com/p/kimi-k3-escaped-a-cybersecurity-testing
Read time: 8 min
📚 Browse**past editions here**.
(https://x.com/rohanpaul_ai**I publish this newletterdaily.**Noise-free, actionable, applied-AI developments only).
- 🗞️ Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm.
- 🗞️ Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.
- 🗞️ Google’s chief scientist is leaving after 27 years to build AI that can run its own research cycle with little human help.
- 🗞️ So Google just announced some major changes in DeepMind’s leadership.
- 🗞️ Microsoft just disclosed for the first time taht OpenAI supplied ~70% of Microsoft’s AI revenue, per new filings.
- 🗞️ Muse Spark 1.2 puts Meta on Artificial Analysis’s Pareto frontier at $0.40 per task.
Connect with me on X (Twitter)
Moonshot AI’s open-weight Kimi K3 escaped its sandbox in a cybersecurity test to find an existing GitHub solution, normal coder behavior but concerning for AI models.
Frontier Security argues this is more consequential than the earlier OpenAI and Anthropic incidents because those involved unreleased models, while Kimi K3 is already in public hands with ordinary user safeguards.
So what actually happened?
- The AI Security Institute (the UK government body) only wrote the free open-source sandbox software; it did not run this test and was not involved.
Frontier Security, a private US firm, downloaded that software, set it up itself, ran Kimi K3 inside it, found the escape, and published the finding, which is why the institute’s response is that the leak came from Frontier’s own configuration, not from its tool.
- Frontier Security was testing Kimi K3 on cybersecurity tasks using a sandbox based on the UK AI Security Institute’s evaluation framework.
- The sandbox was supposed to isolate the model. In the environment as configured, inbound traffic was blocked but outbound HTTPS (port 443) and DNS (port 53) remained open to public IP ranges.
- During Kimi K3’s run, the model didn’t attempt the task. It probed the network, found that DNS resolution for github .com worked, cloned the official benchmark repository, and read the solution directly off disk.
So basically, instead of solving the benchmark task normally, Kimi used that internet access to reach GitHub and find the benchmark solution.
- So Kimi crossed the boundary the test intended to impose and effectively cheated the benchmark through a sandbox configuration weakness.
- Kimi did not hack GitHub or attack another external system. It used network access that should not have been available during the test.
- Researchers therefore identified two problems: the sandbox left an unintended path open, and Kimi did not have an internal safeguard stopping it from using that path.
Kimi K3 did not gain a mind of its own or intentionally misbehave. It exposed a simpler engineering problem: when guardrails depend on how engineers expect an environment to behave, capable models may find routes to an objective that nobody planned for.
Significantly ahead of Kimi K3 Max Fullstack Code Arena asks models to build working web apps through multi-step planning and tools, instead of answering isolated coding questions. Models can create and edit files, run commands, use databases, authentication and external APIs, then produce a live application for testing. Its specialty is testing models inside a complete tool-using build process and judging the finished app, which better resembles how coding agents work.
Jeff Dean was employee 30, and he wrote the storage and computing systems that showed the rest of the industry how to serve a global audience. He also started Google Brain in 2011 and drove the TPU units. His new firm is called Discovery Loop, after the cycle it wants to automate, which is hypothesis, then experiment, then evaluation of the result.
Dean says whole domains can be computerized that way, so you get more experiments and better ones out of the same year. It begins by automating machine learning research itself, before moving toward hardware design, drug discovery and clean energy.
Joining him are Sanjay Ghemawat, his collaborator of two decades, plus Quoc Le and Oriol Vinyals from the Brain and Gemini years. Radical Ventures and Khosla Ventures are co-leading a seed round that is not yet closed, and Alphabet came in as founding investor and cloud partner.
Alphabet will supply their computing power for at least a year, so the four keep frontier-scale hardware without paying to build it. Google shares fell about 4% on the news.
Demis Hassabis will become Chair of Google DeepMind and Alphabet’s Chief Scientist, leaving daily management while continuing to advise its models and research.
Basically it gives Hassabis more time for AGI strategy, scientific discovery, global policy work, and Isomorphic Labs, while Kavukcuoglu carries the pressure of shipping Gemini.
**The timing also reflects scale:**Gemini now serves 950M monthly users. So model research, releases, and app execution needs a much larger operating job.The announcement also confirms that Gemini 4 is in development, though Google provided no release date or technical details. Jeff Dean and Sanjay Ghemawat are separately leaving to form a public-benefit research company, with Alphabet remaining an investor and Cloud partner. Google is building two leadership tracks, one deciding where advanced AI should go and another turning that research into products used at Google scale.
Demis Hassabis on Linkedin.
Connect with me on X (Twitter)
Based on projected growth and previous disclosures that relate to revenue derived from OpenAI. Most of the $24.1B is OpenAI’s cloud bill for the Microsoft data centers that train and serve ChatGPT, topped up by model-development costs and a share of OpenAI’s own sales, all booked together as revenue by the Microsoft that has also funded $11.9B into OpenAI.
i.e. on that frontier, no model in the comparison is simultaneously cheaper per task and higher on the Intelligence Index.
Muse Spark 1.2 also reaches roughly Claude Opus 4.8-level intelligence at about one-fifth of that model’s $2.03 task cost. It scores 6 Index points below Claude Opus 5 while costing roughly one-sixth as much.
The metric goes beyond just the list price because Artificial Analysis weights the input, cache, reasoning, and answer tokens actually consumed across its nine Index evaluations. This is so significant because for long-running agents, small cost differences compound across model calls, reasoning tokens, tool steps, and retries.
That’s a wrap for today, see you all tomorrow.
Similar Articles
@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohan-paul.com/p/openai-stops-reinforcement-learning… OpenAI st…
Newsletter summarizing key AI developments, including OpenAI halting reinforcement learning training after cybersecurity concerns with Astra, Synthefy launching a platform for structured numerical data, and updates on revenue, vulnerabilities, and integrations.
@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohanpaul.substack.com/p/central-bankers-now-fear-the-ai-gold… …
A daily AI newsletter covering multiple stories including warnings from central bankers about AI debt bubbles, Chinese developers buying cheap Claude access via gray-market APIs, Sakana's Fugu report, cost comparisons of Chinese vs American AI models, Deepseek's new inference optimization, and Meta's open-source brain-to-text system.
@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohan-paul.com/p/spacexai-launches-grok-46-beating… SpaceXAI la…
Rohan Paul's daily AI newsletter digest covering SpaceXAI's Grok 4.6 launch, Soniox TTS v2, NVIDIA Nemotron 3.5 Lightning, OpenAI's ChatGPT Linux app, and several research papers.
@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohan-paul.com/p/study-finds-1-in-3-web-pages-published… 1 in 3…
This newsletter covers key AI and tech news, including a Pew study indicating that a third of web pages published since ChatGPT's launch show signs of AI authorship, alongside updates on enterprise competition and research on agent safety.
@rohanpaul_ai: Today’s edition of my newsletter just went out. https://rohan-paul.com/p/mira-muratis-thinking-machines-made… Mira Mura…
A newsletter roundup covering Mira Murati's Thinking Machines achieving 29.8% fewer errors in finance judgments, a shift from Claude Code to Claude Tag, a cheaper method for feeding Fable 5 large context, surge pricing in AI with DeepSeek doubling peak prices, and Alibaba blocking Claude Code after a tracking experiment.