Tag
SR-Fraud is an outcome-supervised reflective LLM agent framework for non-stationary payment fraud detection, improving detection metrics over traditional methods on a production benchmark.
A user found an agent offering up to 20 hours of free daily access to GLM 5.3 Flash and DeepSeek 4.1 Flash AI models, with bonuses for GitHub sign-ups and streaks, requiring no payment information.
The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.
Mark Zuckerberg announces the launch of Muse connectors, enabling developers to integrate their services with an AI agent and browser for seamless user interactions.
The paper proposes a method for proactively detecting implicit conflicts in user-side human-LLM dialogue, introducing a benchmark and synthesis approach to enhance lightweight LLMs' performance.
Claude Cowork and chat features are merging into a single Claude service, making it a more integrated general agent. This update is rolling out to Pro and Max plans first.
This paper introduces Agent as Policy (AGP), a framework that enables general-purpose agents to directly control physical robots for manipulation tasks without task-specific training, achieving high success rates across various real-world scenarios.
A user showcases a native personal trip planner agent built with Swift and Hono worker using Cloudflare services like Project Think, Browser Run, Durable Objects, and R2, demonstrating the versatility of Cloudflare beyond web apps.
T1 is a 122B Mixture-of-Experts model trained with reinforcement learning for long-horizon terminal tasks, achieving state-of-the-art results on benchmarks like Terminal-Bench 2.1 and surpassing models such as GPT-5.4 and GLM-5.1.
This article reveals how to use Topview and Agent tools to create an AI video with tens of millions of views across the web, with the entire process taking only about an hour and a half.
Qwen 3.8 27B, a 27-billion parameter AI model, outperforms Opus 4.8 in personal benchmarks when running locally on an RTX 5090 GPU, achieving up to 200 tokens per second without internet or API access.
Karpathy's method for LLM wiki automation emphasizes iterative source ingestion, filing answers back into the wiki, and regular linting for consistency, promoting user involvement in the process.
A developer seeks community feedback on releasing a self-hosted, open-source AI tool for research, document creation, and agent capabilities, inspired by Manus and Perplexity.
Blender Agent Bridge is an open-source MCP bridge designed to integrate AI workflows into Blender, enhancing 3D modeling capabilities with AI tools.
DeepSeek announced the GA release of DeepSeek-V4-Pro with major agent upgrades, flexible reasoning effort levels, native OpenAI Responses API support, and updated API pricing introducing peak and off-peak rates with off-peak 50% lower.
This paper proposes SDAM, a memory-based framework for complex Text-to-SQL that uses structure-difference aware reasoning, contradiction-aware reflection, and schema-grounded memory evolution to improve SQL generation. Experiments show modest gains on BIRD-dev and Spider-test benchmarks.
dots-studio releases dots3-note preview, the first open-weight model in the dots3 family: a 280B-parameter multimodal MoE with 16B activated parameters, 512K context, and support for text, image, video, and audio understanding.
The tweet introduces Deepseek Harness's trajectory mode, which allows viewing the Agent's thinking steps, tool calls, and system prompts, accompanied by a Chinese translation for easy learning.
DeepSeek Harness has been released, featuring a trace mode that visualizes Agent's thinking steps, tool calls, and return results, and uses colors to distinguish different types of content. It can be installed via npm and is suitable for developers to observe Agent work progress.
OpenAI introduces GPT-5.6, a new model family that delivers frontier-level agent performance at a fraction of the cost, with new API primitives for reasoning persistence, multi-agent orchestration, and programmatic tool calling.