self-correction

Tag

Cards List
#self-correction

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

arXiv cs.CL · yesterday Cached

This paper introduces Hybrid Search, a method to enhance automatic speech recognition in large audio language models by leveraging hidden-state interactions between the ASR-LLM and base LLM for targeted token correction, improving performance beyond global LLM-correction strategies.

0 favorites 0 likes
#self-correction

Why Self-Correction Loops Can Degrade Reliability in LLM Pipelines (85% Down to 62%)

Reddit r/artificial · 2026-08-22

Adding a self-correction loop to an LLM pipeline for structured data extraction reduced consistency from 85% to 62%, due to compounding noise and regeneration drift. The article discusses potential solutions like granular diff mechanisms or deterministic gates.

0 favorites 0 likes
#self-correction

When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

arXiv cs.AI · 2026-08-18 Cached

The paper empirically studies self-correction in code generation using uncertainty estimation methods, finding that uncertainty-based approaches fail to improve Pass@1 accuracy, while verification-based methods yield significant gains.

0 favorites 0 likes
#self-correction

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

arXiv cs.CL · 2026-08-13 Cached

This paper introduces DARC, a diagnosis-guided recovery harness that makes agent self-correction selective by profiling failure modes and pruning mismatched interventions before test-time correction, improving performance on ALFWorld, AppWorld, and XBRL Finance.

0 favorites 0 likes
#self-correction

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

arXiv cs.CL · 2026-08-13 Cached

This paper introduces SFS-DPO, a reinforcement learning two-stage framework for step-level self-verification and self-correction in LLMs, with a teacher-assisted variant SFS-DPO-R. It demonstrates improvements in self-correction effectiveness across multiple LLMs with less training data than prior approaches.

0 favorites 0 likes
#self-correction

@marfinxx: Microsoft researchers published a landmark study defining the shift from Harness Engineering to Loop Engineering in cod…

X AI KOLs Timeline · 2026-08-09 Cached

Microsoft researchers published a landmark study introducing LoopsBench, a framework for evaluating self-correcting coding agent loops, and outlining the shift from harness engineering to loop engineering for reliable autonomous development.

0 favorites 0 likes
#self-correction

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv cs.AI · 2026-08-07 Cached

A new verifier-free breadth-depth refinement framework improves LLM reasoning at test time by sampling multiple rollouts, iteratively refining each via self-critique, and aggregating with majority voting. It consistently outperforms greedy decoding, majority voting, and verifier-based selection across several math benchmarks and open-weight models.

0 favorites 0 likes
#self-correction

@jakevin7: Sharing a god-tier review prompt methodology. The LLM self-correction survey 'When Can LLMs Actually Correct Their Own Mistakes?' concludes that without reliable external feedback such as test results or tool outputs, a model relying only on self-reflection often cannot steadily correct errors...

X AI KOLs Following · 2026-08-06 Cached

Shared a prompt methodology based on LLM self-correction research, emphasizing that self-checking is limited without external feedback, and recommending progressively enhanced prompting strategies such as adversarial review.

0 favorites 0 likes
#self-correction

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

arXiv cs.CL · 2026-08-06 Cached

This paper shows that apparent LLM self-correction gains often stem from format repair rather than improved reasoning. Across multiple model scales, format effects dominate content effects, with content margins near zero on capable models, suggesting the field has misattributed a minority of measured self-correction to actual content improvement.

0 favorites 0 likes
#self-correction

An agent just coded for 10 days with nobody watching. Qwen 3.8 max

Reddit r/AI_Agents · 2026-08-04

Alibaba's Qwen agent autonomously coded for over 10 days in an empty repo, filing issues, writing code, running tests, fixing failures, and merging. It still required some feedback, but demonstrates a self-correcting autonomous loop.

0 favorites 0 likes
#self-correction

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

arXiv cs.LG · 2026-08-03 Cached

This paper introduces the Human–LLM Reflection Framework (HRF) to compare human and LLM revision behavior, finding that LLM reflection often yields zero or negative information gain and behaves more like conditioned re-generation than genuine error-driven revision.

0 favorites 0 likes
#self-correction

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

arXiv cs.AI · 2026-08-03 Cached

Presents AMTFV, an agentic framework that decouples mathematical verification modeling from execution via a Mathematical Tool Flow interface, improving LLM answer verification and revision on five challenging math datasets.

0 favorites 0 likes
#self-correction

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

arXiv cs.AI · 2026-08-03 Cached

ViSAGE is a multimodal agentic memory framework for long-form video understanding that builds self-correcting, entity-centric memories via cross-modal binding, bidirectional memory refinement, and multi-agent cross-verification, achieving 5.9% higher accuracy than baselines.

0 favorites 0 likes
#self-correction

qwen agentworld can self-correct in reasoning traces

Reddit r/LocalLLaMA · 2026-07-27

A user experiments with Qwen AgentWorld and finds a system prompt that enables self-correction in reasoning traces, as demonstrated by the classic car wash test.

0 favorites 0 likes
#self-correction

@no_stp_on_snek: Ran this on Laguna S 2.1 in Poolside's own agent (pool), pointed at a local instance on a DGX Spark, using the prompt l…

X AI KOLs Timeline · 2026-07-24 Cached

A comparison of two AI coding agents building a Mario game: Laguna S 2.1 in Poolside's agent took 62 minutes with self-correction and passed tests, while a previous Qwen model took hours and needed human help; highlights oracle discipline and native harness advantages.

0 favorites 0 likes
#self-correction

Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation

arXiv cs.CL · 2026-07-21 Cached

Introduces Scientific Feasibility Control (SFC), a conformal prediction framework that provides statistical guarantees for scientific reasoning validity in LLMs, achieving 50.1% on PhyX physics reasoning, outperforming DeepSeek-R1 and GPT-4 while reducing scientific violations by 73%.

0 favorites 0 likes
#self-correction

Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making

arXiv cs.AI · 2026-07-21 Cached

This paper proposes a reward-driven LLM agent workflow that integrates POMDP routing and self-correcting reward models, achieving a 24.5% improvement in task success rate on benchmarks like ALFWorld and WebShop.

0 favorites 0 likes
#self-correction

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

arXiv cs.LG · 2026-07-16 Cached

Introduces Self-Correcting Coupled Markov Jump Processes (SC-CMJP) and a training-free sampler CO2Jump for concurrent image understanding and generation, achieving state-of-the-art joint performance on editing, maze, and nonogram tasks.

0 favorites 0 likes
#self-correction

Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

arXiv cs.AI · 2026-07-14 Cached

This paper explores using a compact Small Language Model (Qwen2.5-1.5B) retrained with GRPO and combined with a validator-guided correction loop for autonomous industrial control. The framework achieves high alignment accuracy and low latency, demonstrating practical viability for edge deployment.

0 favorites 0 likes
#self-correction

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation

arXiv cs.LG · 2026-07-14 Cached

This paper introduces SPARC, a spectral-algebraic theory explaining the self-correction blind spot in autoregressive language models, where models fail to correct their own errors but can fix identical external errors. The theory proves the blind spot arises when the spectral radius of an error-propagation operator is at least one, derives a threshold for correction markers, and provides convergence guarantees for RL-based self-correction training.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback