Tag
Jev, a structured text classification model, demonstrated the ability to process 724 live ads from 37 brands in 40 seconds for just $0.09 in tokens, achieving 216ms median latency per record through parallel processing and integration with Gemini.
OpenAI expands its GPT-6 model family in GitHub Copilot with two new models: GPT-6 Sol for balanced agentic coding and GPT-6 Luna for cost-efficient tasks.
Introducing Claude Opus 5.5, the first model in the Claude 5.5 family, which performs at the level of Claude Fable 5.1 but costs 40% less to run with increased rate limits.
Jev is a system that can play nine classic games simultaneously using a single API call, costing $1.80 per hour, showcasing efficient AI performance in real-time gaming.
Jev is a new AI model focused on rapid decision-making, offering 20-200x faster and 40-400x cheaper performance than frontier LLMs, as released by Diogo Almeida.
The co-inventor of ChatGPT announces the release of a new AI model named Jev, trained with RLCD, claiming it is 20-200x faster, 40-400x cheaper, and optimized for composable intelligence as a path to AGI.
Occamy-1.0 is a cost-efficient open-source AI model for co-work agents, achieving strong performance on complex multi-step tasks and being competitive with larger frontier systems while maintaining broad agentic capabilities.
Trained a small transformer model from scratch to achieve 44% accuracy on the ARC-AGI-1 benchmark for only 67 cents, demonstrating improvements in speed, accuracy, and cost-effectiveness over previous methods.
The article discusses benchmark comparisons for the GLM-5.3-Flash model, highlighting its frontier intelligence and cost efficiency from a release blog post.
Mixedbread introduces Toast 1, a specialized search agent that matches frontier model quality while being up to 10x cheaper and 12x faster. It automates agentic search loops and achieves state-of-the-art results on benchmarks like OfficeQA Pro V2 and legal knowledge tasks.
Raindrop launches Signals 2.0 powered by rd-signal-2, a new model pipeline for building task-specific binary classifiers from production traces. It claims near GPT-5.6 Sol xhigh accuracy at 1600x lower cost, and also introduces Signal Builder for custom classifiers with zero data retention.
Hark Handoff reportedly outperforms GPT-5.5 and Opus 4.8 on several benchmarks at 90% lower cost, using SFT and asynchronous RL with GRPO on an undisclosed base model. The author expresses skepticism about latency in computer-use agents but is bullish on the demo.
Hark has unveiled Handoff, an AI model that reportedly achieved a record score on the OM2W benchmark, outperforming top models while running at a significantly lower cost.
browser-use announces a new agent powered by GPT-5.6 Luna, which can read top Hacker News posts and generate a full report for only 3 cents, highlighting the low cost of AI-driven web browsing.
Alibaba announces Qwen3.8-Max, a 2.4T-parameter MoE frontier model with open weights coming next week, claiming autonomous operation for 16 days and significantly lower cost than GPT-5.6 Sol and Claude Fable 5.
RAG-HAR+ is a retrieval-first, cost-optimized extension of RAG-HAR for human activity recognition from wearable sensors. It uses a retrieval designer agent and majority voting to reduce LLM usage while maintaining accuracy, and demonstrates feasibility for edge deployment.
This handbook teaches developers how to build a production-grade RAG system using Cloudflare Workers, Vectorize, and Workers AI, focusing on cost efficiency and reliability.
Google launches Gemini 3.5 Flash Cyber, a cost-efficient AI security model for vulnerability detection, alongside Gemini 3.6 Flash and 3.5 Flash-Lite, positioning it as a cheaper alternative to Anthropic's Mythos.
This paper proposes a generative AI-assisted summarization framework using GPT-5 model variants to address input-length limitations in automated essay scoring, demonstrating trade-offs between model capacity, summary fidelity, and computational cost on the ASAP 2.0 dataset.
Cognition launches SWE-1.7, a highly capable AI model for agentic software engineering that achieves frontier-level performance at reduced cost, with improvements in RL training, multi-cluster infrastructure, data curation, and self-compaction for long tasks.