Articles from Blog
OpenAI announces improved prompt caching for GPT-6, offering higher cache hit rates, discounts, and new tools for monitoring and optimization.
A TikTok creator criticizes the obvious use of AI in writing scripts for social media content, highlighting AI-isms and the lack of a personal voice.
Release of llm-typesafe plugin version 0.1a0, which adds support for TypeSafe AI's Jev model in the LLM command-line tool, with installation and usage examples for AI-powered queries.
NVIDIA releases Isaac ROS 5.0 with new agentic workflows and support for ROS Lyrical to advance open source robotics development using GPU acceleration.
Parallel leverages GPT-6 Astra to reduce research time and cost by half while delivering the same quality, showcasing the model's efficiency in focused searches and multi-agent coordination.
UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.
OpenAI outlines priorities and principles for effective third party assessments to enhance AI safety, emphasizing independent scrutiny, shared standards, and deep access for rigorous evaluation.
Hugging Face's transformers library now supports GGUF models from llama.cpp, enabling efficient local inference on consumer hardware through familiar APIs.
Alibaba has unveiled its new Zhenwu V900 AI accelerator, which triples the performance of its predecessor, targeting to drive 20GW of data centers by 2032.
Aikido introduces Altar, an open-weight security model derived from GLM-5.3 and optimized for efficient deployment in sovereign security environments, enabling local inference without external dependencies.
AWS has launched Strands Harness, an AI agent tool that enables developers to run various AI models with built-in capabilities like web search and memory, claiming significant cost savings over competitors.
Kev is a family of small decision models built on Qwen3.5, offering pretrained weights and training code for yes/no, multiple-choice, and rating questions. It includes a web playground and is compatible with TypeSafe's System One API.
AI models like Jev and SemIf are optimizing if-then decision-making in software, leading to significant cost reductions and improved accuracy, which highlights the potential for specializing other programming primitives.
MiMo-V2.6 introduces Groupwise Advantage Redistribution to enhance reinforcement learning for AI agents by comparing sibling attempts and using graded feedback, showing steady performance improvements across multiple task domains.
Devin has introduced cloud integration into its terminal CLI, allowing users to create, steer, and resume cloud sessions with features like handoff and full SSH access for a seamless development workflow.
RecreationWorld is a scalable framework for training hybrid AI agents that combine GUI interaction, coding, and visual verification to rebuild applications, with a benchmark suite called RecreationBench.
This article analyzes the capabilities and scaling dynamics of large AI agent swarms, citing OpenAI's recent examples, and discusses their potential as a new form of inference scaling.
An analysis of the business models of frontier AI labs like OpenAI and Anthropic, discussing their competitive advantages, revenue challenges, and strategies to expand into other industries amid rising competition and costs.
The article analyzes the shift in leadership for open-weight AI models from the U.S. to China, highlighting that Chinese models have surpassed American ones in commercial viability and download numbers.
The article analyzes the 'Great Unbundling of Intelligence' in AI, where agent economics are driving a shift from using general frontier models for all tasks to a system with specialized cheaper models for routine work, optimizing cost and efficiency.