Tag
The title references 'opus 5.5', likely referring to Anthropic's AI model, suggesting a version update or release.
Hi3D launched Hi3D Agent, an AI agent that combines a large language model with 3D tools to streamline workflows from model generation to final production in game, animation, and design.
Yandex has released AliceAI Foundation 80B A3B Base, a custom AI model competing with Qwen 35B and DeepSeek V4 Flash, featuring its own architecture without post-training.
This paper presents Jarvis, an open-source voice assistant framework for autonomous racecars that uses a fine-tuned Mistral 7B model for offline intent recognition, achieving 97.63% accuracy with low latency and outperforming larger online models.
DAPO is an open-source reinforcement learning system for large language models, developed by ByteDance Seed and Tsinghua AIR, which achieves state-of-the-art performance on the AIME 2024 benchmark with the Qwen2.5-32B model.
StepFun introduces Step 5 Preview, a new flagship AI model for agentic work with frontier-level performance in software engineering and finance, featuring a 600B parameter scale.
Sherpa is an advanced AI fiction writing tool by Pocket FM that uses specialized narrative models to enhance serialized content creation, boosting production and enabling significant earnings for creators.
This paper empirically investigates the susceptibility of LLM-based GUI agents to digital nudges, finding that reasoning configuration redirects rather than reduces nudge effects, positioning interface design as a governance concern for autonomous AI.
The article presents K2-Horizon-7B-Uno, a diffusion-augmented LLM that combines autoregressive and diffusion pathways to achieve 5200 tokens per second throughput without quality loss, with benchmarks showing competitive performance across various tasks.
The STaR paper introduces a technique for large language models to bootstrap their own reasoning by iteratively generating rationales, filtering for correctness, and fine-tuning on successful examples.
Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.
A Reddit post highlights the strong performance of the Qwen 3.8 27B model, with mention of 1.3B tokens used recently.
The article refutes the idea that AI development is slowing down, highlighting that Grok4.8, a 2.5T parameter model, has finished training and that major tech companies are still actively training AI models.
Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.
This paper proposes an LLM-enhanced multi-agent reinforcement learning framework to simultaneously optimize electric vehicle charging scheduling, station profitability, and grid stability, using LLMs for feature selection and adaptive weighting, outperforming state-of-the-art methods with reduced training time.
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms that uses large language models to enable natural language device control, focusing on architecture design and key technologies like KV Cache optimization and progressive prompt disclosure to reduce latency in sterile environments.
OntologyAligner presents a three-stage framework combining ontology-aligned retrieval and LLM reranking for biomedical ontology normalization, achieving state-of-the-art performance on HPO tasks and introducing the PhenoNormBench benchmark.
The paper proposes a Dynamic Fusion Large Language Model (DF-LLM) for traffic flow prediction, integrating spatiotemporal embedding, fusion modules, and an LLM backbone with adaptation strategies to enhance performance over traditional methods in intelligent transportation systems.
REVA is a framework that mines historical attention traces to create reusable evidence views for context-efficient RAG serving, improving generation quality while reducing compression overhead and latency.
Cognition's SWE-2 model, post-trained from Kimi K3, achieves a score of 92.8 on Terminal-Bench 2.1, offering competitive performance with lower cost compared to frontier models like Fable 5.1 and GPT-6 Astra.