Tag
DKV is an open-source framework for compressing KV-cache during local LLM inference, providing a CLI and a technical report.
The technical report of PrismML's Bonsai-27B model is trending on Papers with Code, with evaluations and Hugging Face models available.
Li Auto's Mach-Mind-4-Flash is a 35B MoE model with only 3B active parameters that, through post-training optimization using a unified RL/on-policy distillation loss, achieves performance matching or surpassing 100B-parameter models across multiple benchmarks.
Gemma 4 introduces a new generation of open-weight, natively multimodal language models with dense and Mixture-of-Experts architectures, featuring thinking mode for advanced reasoning, improved efficiency, and long-context capabilities.
This technical report presents iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision, language, and action for embodied agents, using a brain-cerebellum collaboration architecture and a four-stage training strategy.
Kyutai Labs introduces MIRA, a new multiplayer world model developed with Gen Intuition and Epic Games, releasing a technical report, dataset, and online demo.
A tweet sharing a chapter from a UC Berkeley EECS technical report from 2016, expressing excitement about its content.
The Gemma 4 Technical Report introduces a new generation of open-weight, natively multimodal language models with diverse architectures, enhanced reasoning capabilities, and improved performance across tasks. The models range from 2.3B to 31B parameters and feature a thinking mode for generating reasoning traces.
The Sakana Fugu technical report introduces a trained orchestrator that dynamically selects and coordinates specialist models for tasks, with a faster version (Fugu) and a slower workflow version (Fugu-Ultra) that can design custom teamwork patterns per request.
This technical report from TUM details the path a packet takes through the Linux kernel, covering networking internals.
A detailed blog post explaining the Sakana Fugu technical report, which introduces orchestrator AI models that route tasks to specialized models, achieving collective intelligence.
This technical report introduces Ling and Ring 2.6, a family of large language models at the trillion-parameter scale designed for efficient and instant agentic intelligence.
Qwen-RobotWorld is a language-conditioned video world model that predicts future visual trajectories across multiple robotic domains using a double-stream diffusion transformer and an 8.6M video-text corpus. It unifies embodied world modeling for robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer, achieving top benchmarks on EWMBench and DreamGen Bench.
This technical report presents Qwen-Image-2.0, a new image generation model from Alibaba's Qwen team, detailing its architecture and capabilities.
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
This technical report introduces X-OmniClaw, a unified mobile agent system designed for multimodal understanding and interaction on Android devices. It details the architecture for perception, memory management, and action execution using on-device AI capabilities.
OpenAI releases gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, open-weight reasoning models designed for policy-based content classification with full chain-of-thought reasoning. The technical report provides baseline safety evaluations and demonstrates the models' capabilities for content labeling tasks under the Apache 2.0 license.