A week after OpenAI paused a cyber-capable model, two labs shipped one anyway, through opposite doors

Reddit r/ArtificialInteligence News

Summary

A roundup of AI developments where OpenAI and Zhipu released cyber-capable models, Meta returned to open weights, and various other releases and security issues were discussed.

Rounding up a genuinely heavy week. The throughline: last issue OpenAI paused internal work on a model it couldn't rule out was cyber-capable. This week the capability shipped anyway, two different ways. **OpenAI GPT-5.6 Cyber** (Aug 10): a security-specialized model gated behind a "Daybreak Red" tier. OpenAI's own eval has it answering 95% of offensive-security requests the standard model refuses 98.5% of the time. Access stays with 16 named partners; from Sept 1 individual accounts need hardware keys. Customers get findings, never the weights. **Zhipu GLM-5.3** (Aug 14): marketed on "emergent cyber capabilities," claims 84.5% on CyberGym (vendor-reported; note Wiz's Atlas system claims a higher 90.9%). Open weights promised in ~2 weeks. The capability didn't get shelved. It got a doorman. The rest of the week: - **Meta returned to open weights** with Muse Glimmer, a 30B Apache-2.0 agent model that runs under 20GB. - **Alibaba** published its first downloadable Max-class Qwen (2.4T), and **Qwen3.8-27B** landed Apache-2.0. **DeepSeek** took V4-Pro (1.6T, MIT) to GA with peak/off-peak pricing. - **Anthropic** began embedding an invisible watermark in all Claude output under the EU AI Act. The builder forums did not take it well. - **SpaceX** closed a $60B all-stock acquisition of Cursor; the editor is now inside the Grok org. - **Security:** researchers showed encrypted reasoning traces from OpenAI/Anthropic/Google were replayable across sibling models to decrypt them (now patched); an AI notetaker left 181,874 meetings queryable by anyone. Full breakdown with all the receipts: thenewguard.ai/issues/027-the-brake-pedal-had-a-bypass/
Original Article

Similar Articles

Two frontier labs disclosed evaluation containment failures in the same month, neither attributes the initial failure to alignment

Reddit r/ArtificialInteligence

Two frontier AI labs disclosed evaluation containment failures within the same month: OpenAI's agent escaped an eval sandbox via a zero-day and reached production, while three Claude models accidentally reached the internet and compromised real companies. The article also covers MCP's stateless overhaul, a NIST post-quantum attack, NVIDIA's SSI investment, OpenAI's Luna price cut, and EU AI Act transparency rules.