agent

Tag

Cards List
#agent

SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection

arXiv cs.LG ↗ · yesterday Cached

SR-Fraud is an outcome-supervised reflective LLM agent framework for non-stationary payment fraud detection, improving detection metrics over traditional methods on a production benchmark.

0 favorites 0 likes
#agent

I found this agent that gives you up to 20 hours of GLM 5.3 Flash or DeepSeek 4.1 Flash for free every day, no payment info required!

Reddit r/AI_Agents ↗ · 3d ago

A user found an agent offering up to 20 hours of free daily access to GLM 5.3 Flash and DeepSeek 4.1 Flash AI models, with bonuses for GitHub sign-ups and streaks, requiring no payment information.

0 favorites 0 likes
#agent

VideoGen-Agent: Reinforcing Video Generation Agents

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.

0 favorites 0 likes
#agent

@yoheinakajima: (M)ark’s (C)connector (P)latform

X AI KOLs Timeline ↗ · 5d ago Cached

Mark Zuckerberg announces the launch of Muse connectors, enabling developers to integrate their services with an AI agent and browser for seamless user interactions.

0 favorites 0 likes
#agent

Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

arXiv cs.CL ↗ · 2026-09-18 Cached

The paper proposes a method for proactively detecting implicit conflicts in user-side human-LLM dialogue, introducing a benchmark and synthesis approach to enhance lightweight LLMs' performance.

0 favorites 0 likes
#agent

Claude Cowork and chat are now one Claude

Simon Willison's Blog ↗ · 2026-09-16 Cached

Claude Cowork and chat features are merging into a single Claude service, making it a more integrated general agent. This update is rolling out to Pro and Max plans first.

0 favorites 0 likes
#agent

Agent as Policy for Robotic Manipulation

arXiv cs.CL ↗ · 2026-09-14 Cached

This paper introduces Agent as Policy (AGP), a framework that enables general-purpose agents to directly control physical robots for manipulation tasks without task-specific training, achieving high success rates across various real-world scenarios.

0 favorites 0 likes
#agent

@fayazara: Fyi, you are not just limited to building web apps with @cloudflaredev This is my native personal trip planner agent I …

X AI KOLs Following ↗ · 2026-09-12 Cached

A user showcases a native personal trip planner agent built with Swift and Hono worker using Cloudflare services like Project Think, Browser Run, Durable Objects, and R2, demonstrating the versatility of Cloudflare beyond web apps.

0 favorites 0 likes
#agent

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Hugging Face Daily Papers ↗ · 2026-09-10 Cached

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning for long-horizon terminal tasks, achieving state-of-the-art results on benchmarks like Terminal-Bench 2.1 and surpassing models such as GPT-5.4 and GLM-5.1.

0 favorites 0 likes
#agent

@AI_Jasonyu: Unveiling How My AI Video Went Viral with Tens of Millions of Views! Full Recording of the Production Process – Watch Patiently. This video has conservatively reached tens of millions of views across the web, and even reposts on X have nearly 2 million views. It's my first time creating an AI short drama, and this 5-minute video took only about one hour to produce.

X AI KOLs Timeline ↗ · 2026-08-28 Cached

This article reveals how to use Topview and Agent tools to create an AI video with tens of millions of views across the web, with the entire process taking only about an hour and a half.

0 favorites 0 likes
#agent

@RoundtableSpace: Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens per second with n…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

Qwen 3.8 27B, a 27-billion parameter AI model, outperforms Opus 4.8 in personal benchmarks when running locally on an RTX 5090 GPU, achieving up to 200 tokens per second without internet or API access.

0 favorites 0 likes
#agent

@undefinedKi: Karpathy never meant for you to automate the LLM wiki. Most people set it up backwards. Three things are in his file an…

X AI KOLs Timeline ↗ · 2026-08-22 Cached

Karpathy's method for LLM wiki automation emphasizes iterative source ingestion, filing answers back into the wiki, and regular linting for consistency, promoting user involvement in the process.

0 favorites 0 likes
#agent

Is there any interest in a self hosted open source version of manus/perplexity?

Reddit r/LocalLLaMA ↗ · 2026-08-21

A developer seeks community feedback on releasing a self-hosted, open-source AI tool for research, document creation, and agent capabilities, inspired by Manus and Perplexity.

0 favorites 0 likes
#agent

Blender Agent Bridge

Product Hunt ↗ · 2026-08-16

Blender Agent Bridge is an open-source MCP bridge designed to integrate AI workflows into Blender, enhancing 3D modeling capabilities with AI tools.

0 favorites 0 likes
#agent

DeepSeek peak/off-peak pricing update

Hacker News Top ↗ · 2026-08-14 Cached

DeepSeek announced the GA release of DeepSeek-V4-Pro with major agent upgrades, flexible reasoning effort levels, native OpenAI Responses API support, and updated API pricing introducing peak and off-peak rates with off-peak 50% lower.

0 favorites 0 likes
#agent

SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper proposes SDAM, a memory-based framework for complex Text-to-SQL that uses structure-difference aware reasoning, contradiction-aware reflection, and schema-grounded memory evolution to improve SQL generation. Experiments show modest gains on BIRD-dev and Spider-test benchmarks.

0 favorites 0 likes
#agent

dots-studio/dots3-note-prev · Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-08-13 Cached

dots-studio releases dots3-note preview, the first open-weight model in the dots3 family: a 280B-parameter multimodal MoE with 16B activated parameters, 512K context, and support for text, image, video, and audio understanding.

0 favorites 0 likes
#agent

@Saccc_c: In Deepseek Harness's trajectory mode, you can also view its system prompt (see Figure 1). I translated it into Chinese; interested friends can take a look (see Figures 2 and 3)

X AI KOLs Following ↗ · 2026-08-13 Cached

The tweet introduces Deepseek Harness's trajectory mode, which allows viewing the Agent's thinking steps, tool calls, and system prompts, accompanied by a Chinese translation for easy learning.

0 favorites 0 likes
#agent

@Saccc_c: Deepseek harness's trace mode is so interesting, now you can completely know the progress of Agent's work, and your Agent employees will never dare to slack off again. Its specific thinking steps, tool calls, and tool return results are all clearly visible. The most interesting thing is that it uses different colors to distinguish when AI executes tasks…

X AI KOLs Following ↗ · 2026-08-13 Cached

DeepSeek Harness has been released, featuring a trace mode that visualizes Agent's thinking steps, tool calls, and return results, and uses colors to distinguish different types of content. It can be installed via npm and is suitable for developers to observe Agent work progress.

0 favorites 0 likes
#agent

The builder’s guide to GPT‑5.6

OpenAI Blog ↗ · 2026-08-13 Cached

OpenAI introduces GPT-5.6, a new model family that delivers frontier-level agent performance at a fraction of the cost, with new API primitives for reasoning persistence, multi-agent orchestration, and programmatic tool calling.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback