self-improving

Tag

Cards List
#self-improving

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

arXiv cs.AI ↗ · 2d ago Cached

The paper introduces JAZ, a minimalist agent framework that uses a single invoke primitive to enable LLM agents to handle long-horizon tasks and self-improvement without external harnesses, outperforming existing systems like MemGPT and ACE in evaluations.

0 favorites 0 likes
#self-improving

@rohanpaul_ai: New Harvard + MIT + other labs paper shows a cleaner path to self-improving financial AI: let the agent learn from SEC …

X AI KOLs Following ↗ · 4d ago Cached

A new paper from Harvard, MIT, and other labs introduces FINSKILLOPS, a method that enables financial AI agents to continuously learn from SEC filing errors while using regression tests to maintain correctness and safely update skills.

0 favorites 0 likes
#self-improving

Show HN: AutoBot – live voice control for long-running AI work

Hacker News Top ↗ · 2026-09-17 Cached

AutoBot is an open-source agentic harness that enhances AI task completion by 18.5% over baseline on benchmarks like OSWorld 2.0 and AssistantBench, featuring self-improvement, hierarchical memory, and live voice control for complex workflows.

0 favorites 0 likes
#self-improving

AI agent hit a bug it had predicted 20 minutes earlier-in its own self-written code

Reddit r/AI_Agents ↗ · 2026-09-17

A custom AI agent autonomously built a tool, audited its own code, predicted a bug, encountered it later, and recovered by adapting its approach, showcasing advanced self-improvement capabilities.

0 favorites 0 likes
#self-improving

@kentcdodds: I've got a package that automatically handles feedback I get from people using @kodykoala and even ships fixes and impr…

X AI KOLs Following ↗ · 2026-09-15 Cached

Kent C. Dodds shares a package that automates feedback handling and improvements, integrating with tools like Kodykoala, Discord, GitHub, CodeRabbit AI, and Cursor AI for self-improving software.

0 favorites 0 likes
#self-improving

@MaxBrodeurUrbas: Introducing Gumball Model agnostic, proactive, self improving. All in your company’s private cloud

X AI KOLs Following ↗ · 2026-09-10 Cached

Gumball is introduced as a model-agnostic, proactive, and self-improving system available in private clouds.

0 favorites 0 likes
#self-improving

@ycombinator: Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

Y Combinator discusses the importance of harnesses in AI, highlighting their role in improving model performance, self-improving agents, and real-world applications such as personal AI and work automation.

0 favorites 0 likes
#self-improving

@tobi: Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursiv…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

A tweet from @tobi discusses how training tiny models for specialized use cases with a self-improving flywheel is highly effective, noting that Shopify ML team's finetuned 0.8b model outperforms GPT 5.6-sol xhigh in a specific task.

0 favorites 0 likes
#self-improving

Prime Agent: A Self-Improving RLM Harness

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

Prime Agent is an open-source harness that uses recursive subagents and persistent computation to extend language models' long-horizon capabilities across coding and reasoning tasks, significantly improving performance on benchmarks like ARC-AGI-3.

0 favorites 0 likes
#self-improving

@no_stp_on_snek: Woot woot, latest contribution to Hermes for bots. Found when developing my ios/android app to access my Hermes session…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

A contribution to the Hermes Agent, a self-improving AI agent by Nous Research, featuring memory, skill creation, and multi-platform support for developers.

0 favorites 0 likes
#self-improving

@with_gene2626: wtf? 397b nvfp4 is a 3 or 4 spark setup... showing better benchmarks then all of the current open weight models? anyone…

X AI KOLs Following ↗ · 2026-08-19 Cached

Ornith-1.5, a family of open-source LLMs, is introduced with variants up to 397B MoE, achieving state-of-the-art performance among comparable models and rivaling Claude Opus in benchmarks.

0 favorites 0 likes
#self-improving

@sunmer575399: Everyone doing AI should check out hermes-agent. Not because of the 204.0k stars. But because it solves the core problems cleanly. It's known as the 'agent that grows' — it doesn't stick to a fixed workflow, but automatically adjusts strategies based on your usage habits... I've been running it for two weeks, and it feels like it understands you more the more you use it, much more flexible than those rigid auto...

X AI KOLs Timeline ↗ · 2026-08-15 Cached

Hermes Agent is a self-improving AI agent framework developed by Nous Research, supporting multi-environment deployment and model switching, aimed at improving development efficiency and user experience.

0 favorites 0 likes
#self-improving

@tom_doerr: Aeon is an autonomous agent framework that ships features to your repos, finds real vulnerabilities, and deploys live a…

X AI KOLs Timeline ↗ · 2026-08-07 Cached

Aeon is an autonomous agent framework that ships features to repos, finds real vulnerabilities, and deploys live apps without approval loops, running unattended on GitHub Actions. It supports 60+ skills across six harnesses and can write new skills for itself.

0 favorites 0 likes
#self-improving

@LinusEkenstam: Just before bed time. Let me sleep. plz 95.5% on ARC-AGI-3 (’huge if true”)

X AI KOLs Timeline ↗ · 2026-08-05 Cached

Linus Ekenstam highlights Prime Intellect's release of Prime Agent, a self-improving harness for coding and long-running autonomous tasks, reportedly scoring 95.5% on ARC-AGI-3, above the human baseline.

0 favorites 0 likes
#self-improving

@latkins: Yo

X AI KOLs Following ↗ · 2026-08-05 Cached

Prime Intellect introduced Prime Agent, a self-improving RLM harness for coding and long-running autonomous tasks, featuring programmatic tool calling, context as a variable, multi-agent messaging, and self-modifiable harness state.

0 favorites 0 likes
#self-improving

@rvaniaaaa: someone built an AI agent that learns new skills on its own. the only human step is the final approval. 8 agents. one l…

X AI KOLs Timeline ↗ · 2026-08-04 Cached

A new AI agent system uses eight specialized agents to autonomously discover, validate, and integrate new skills from GitHub, requiring only final human approval before merging.

0 favorites 0 likes
#self-improving

@RohOnChain: Don't waste 2 years learning to build self-improving AI agents. Stanford just dropped a 3 hour course on building AI ag…

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Tweet highlighting a Stanford 3-hour course on building self-improving AI agents from scratch, covering basics, multi-step reasoning, and learning from feedback, with a mention of high salaries for AI engineers.

0 favorites 0 likes
#self-improving

@0xMovez: Don't waste 2 years figuring out how AI agents actually work. ex. Anthropic & Google engineers at Stanford just dropped…

X AI KOLs Timeline ↗ · 2026-08-03 Cached

A tweet highlights a Stanford lecture by Anthropic and Google engineers covering self-improving AI agents, agent loop patterns, and the generator-verifier gap.

0 favorites 0 likes
#self-improving

A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification

arXiv cs.CL ↗ · 2026-07-22 Cached

SIFT introduces a self-improving document classifier that uses a cheap SPLADE-LightGBM pipeline and an LLM judge to continuously teach itself, while a frozen-gate safety mechanism prevents silent regression, enabling autonomous retraining without human labeling overhead.

0 favorites 0 likes
#self-improving

The $110/month self-improving pipeline (5 minute read)

TLDR AI ↗ · 2026-07-16 Cached

A developer shares their $110/month automated pipeline that uses Claude AI to triage, decompose, implement, and test GitHub issues, resulting in 27 merges over 2 weeks with minimal failures.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback