verifier

Tag

Cards List
#verifier

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Hugging Face Daily Papers · 2026-08-17 Cached

This paper finds that prior audit and repair episodes in context reduce false alarms in LLM verifiers by shifting decision thresholds, with repair content and audit verdict complementarily affecting different model families.

0 favorites 0 likes
#verifier

Verifier-Induced Support Reshaping in On-Policy Optimization

arXiv cs.LG · 2026-08-04 Cached

The paper introduces verifier-induced support reshaping, showing that on-policy RL with verifiable rewards can improve the current objective while making successful behaviors for later objectives too rare to sample. Experiments across math reasoning and instruction following demonstrate that endpoint improvements do not guarantee future trainability.

0 favorites 0 likes
#verifier

@feltanimalworld: Master, your post kept me pondering for two whole days! I originally wanted to write a long piece, but there were too many threads and I didn't know where to start. To sum it up, this is why I've been on Twitter lately—I felt something crucial was missing in my development understanding; it's also why, besides my hardware repair work, I've had to look into text-to-CAD recently. I...

X AI KOLs Timeline · 2026-07-12 Cached

The author discusses the lack of intermediate representation (IR) and verifier in the deployment of AI in serious industries, using text-to-CAD as an example to illustrate the key role of unified IR and verification in the feasibility of AI solutions.

0 favorites 0 likes
#verifier

New sampler + verifier *drastically* improves tiny 0.5b model coding performance

Reddit r/LocalLLaMA · 2026-06-25 Cached

The paper introduces VGB, a process-guided sampling algorithm with probabilistic backtracking, which significantly improves coding performance on tiny 0.5B models by being robust to verifier errors.

0 favorites 0 likes
#verifier

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

Hugging Face Daily Papers · 2026-06-08 Cached

Sign-Gated On-Policy Distillation (SG-OPD) enhances standard on-policy distillation by using a binary verifier as a trust signal for teacher supervision, improving performance on competition-level math reasoning benchmarks.

0 favorites 0 likes
#verifier

Agentic Systems as Boosting Weak Reasoning Models

arXiv cs.AI · 2026-05-15 Cached

This paper studies verifier-backed committee search as inference-time boosting for reasoning language models, showing that a committee of weak reasoning models can match the performance of much stronger models on code repair tasks like SWE-bench Verified.

0 favorites 0 likes
#verifier

Logic-Regularized Verifier Elicits Reasoning from LLMs

arXiv cs.CL · 2026-05-08 Cached

Introduces LoVer, an unsupervised verifier that uses logical rules (negation consistency, intra-group and inter-group consistency) to improve LLM reasoning without labeled data, achieving performance close to supervised verifiers on reasoning benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback