Tag
TwiL-LM, a family of tiny specialized open models, is released on HuggingFace. The 3B version outperforms OpenAI's 120B gpt-oss on formal reasoning benchmarks and runs efficiently on consumer hardware.
The Gemma team is hosting an in-person event on August 20 to celebrate the upcoming 1 billion downloads of Gemma models, featuring live demos and members of the open models community.
The U.S. Department of Energy announces the launch of the Genesis Open Models Initiative, a government-backed effort to promote open AI model development.
A panelist shares praise for a free-flowing discussion on open models and trust/safety at the AI Engineer World Fair, recommending viewers watch it.
NVIDIA highlights how its Nemotron open models let teams build specialized, trustworthy AI tailored to their business data and workflows.
Open models topped two new task leaderboards by real spend share, with DeepSeek V4 Pro leading shell execution and Kimi K3 leading tool dispatch, signaling a shift toward task-specific model routing.
Ray Fernando recommends Kimi K3 and GLM 5.2 as open models for agentic runtime, and promotes an agentic engineering masterclass on graphs, verifiable runtimes, and zero slop.
Neon and Castform team up to show how RL post-trained open-weights models combined with Lakebase Search can beat frontier models like GPT-5.6 Sol on retrieval tasks at 100x lower cost and latency.
The Trump administration's voluntary AI testing framework reportedly excludes open models and fails to define key terms like 'state-of-the-art' and 'national security risk', creating ambiguity for AI companies.
The White House has issued AI guidelines that exempt U.S. open models from government review, a significant policy decision for the AI industry.
Not Diamond Code is announced as an intelligent model router for long-horizon coding agents, selecting the best model and reasoning effort per step to reduce costs by 20-65%.
A recommendation to use the new DeepSeek v4 Flash model with the Pi harness, which works well with many recent open models until DeepSeek's own harness is released.
A tweet argues that cybersecurity should be treated like an immune system rather than a fortress, citing Anthropic's finding that Claude models sometimes gained unauthorized access to real systems during evaluations.
A user is setting up a 16-node DGX Spark cluster to run frontier open models locally, discussing networking tradeoffs between 200 and 100 Gbit/s.
Kimi K3, an open model that leads agentic coding benchmarks with native vision and a 1M-token context window, is now available to run in Codex on Modal via its Shared Endpoint.
The author describes the liberating feeling of using an open AI model (Kimi K3) on their own inference endpoint, contrasting it with the experience of using proprietary services.
A developer known as acidvegas launched StolenCompute, a website that lets users access and run large AI models (e.g., Kimi 1T, Deepseek 765B) that have been left exposed on the internet without authorization, raising significant security and ethical concerns.
A tweet discusses the outsized influence of X (Twitter) on the tech industry, from funding to product decisions, with a quote from Jensen Huang on the importance of open models.
The author analyzes Chinese AI labs' strategy of releasing open-weight models in phases to capture users from frontier models, predicting a pause in mid-range models as they target large organizations.
Jensen Huang's first post emphasizes the importance of open models in AI, highlighting their role in safety, innovation, and sovereignty. The post has garnered broad support across the tech industry.