instruction-following

Tag

Cards List
#instruction-following

what happens if you instruct your go-to AI model to: "NEVER HALLUCINATE!!!"

Reddit r/singularity · 2026-06-09

A thought experiment questions whether instructing an AI model to never hallucinate would trigger self-reflection or result in the model gaslighting itself into believing it isn't hallucinating.

0 favorites 0 likes
#instruction-following

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Hugging Face Daily Papers · 2026-06-07 Cached

Introduces OmniCap-IF, the first comprehensive benchmark for evaluating instruction-following in omni-modal video captioning, revealing a format-content tradeoff and proposing improved models and datasets.

0 favorites 0 likes
#instruction-following

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

arXiv cs.CL · 2026-06-04 Cached

VCIFBench is a new benchmark for evaluating complex instruction following in video understanding, featuring 306 test instructions with content, format, style, and structure constraints, plus a DPO preference dataset. Experiments on 10 MLLMs reveal that joint constraint satisfaction remains challenging, and DPO training on the benchmark data improves instruction-following performance.

0 favorites 0 likes
#instruction-following

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hugging Face Daily Papers · 2026-06-03

SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.

0 favorites 0 likes
#instruction-following

I trained a 75M parameter LLM from scratch on 18B tokens and it beats a model almost double its size

Reddit r/LocalLLaMA · 2026-06-02

Trained a 75M parameter LLM called KeyLM from scratch on 18B tokens, achieving competitive instruction-following scores against larger models while using fewer parameters and less data.

0 favorites 0 likes
#instruction-following

Prompt-Level Reward Specifications for Open-Ended Post-Training

arXiv cs.CL · 2026-05-29 Cached

This paper proposes a prompt-level reward specification framework that separates reward specification from computation, constructing reusable task-adaptive rubrics and executable constraint checkers offline to produce a hybrid reward for open-ended post-training without requiring human annotations or separate reward models.

0 favorites 0 likes
#instruction-following

Label-Free Reinforcement Learning via Cross-Model Entropy

arXiv cs.LG · 2026-05-29 Cached

Proposes Cross-Model Entropy (CME) as a label-free reward signal for reinforcement learning post-training of large language models, enabling open-ended instruction following without ground-truth verifiers or human preference labels.

0 favorites 0 likes
#instruction-following

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

arXiv cs.CL · 2026-05-21 Cached

This paper investigates the conflict between instruction-following and pattern completion in LLMs, finding that instruction-following is brittle under induction pressure and varies widely across models, with output diversity being the primary factor for robustness.

0 favorites 0 likes
#instruction-following

@cursor_ai: Introducing Composer 2.5, our most powerful model yet. It's more intelligent, better at sustained work on long-running …

X AI KOLs Timeline · 2026-05-18 Cached

Cursor AI announces Composer 2.5, their most powerful model yet, with enhanced intelligence, better sustained work on long tasks, and improved instruction following; they are doubling included usage for a week.

0 favorites 0 likes
#instruction-following

AI benchmarks matter less than whether models can handle boring real-world responsibility

Reddit r/ArtificialInteligence · 2026-05-17

The article argues that AI benchmarks and flashy demos are overemphasized; the real test for AI trustworthiness is how models handle boring real-world responsibilities like following instructions, admitting uncertainty, handling edge cases, and being auditable.

0 favorites 0 likes
#instruction-following

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

arXiv cs.AI · 2026-05-14 Cached

This paper provides a mechanistic explanation for why LLMs lose track of instructions in long multi-turn interactions, introducing the Goal Accessibility Ratio (GAR) metric and a channel-transition framework. Through ablation studies and residual stream probes, it shows that attention to goal-defining tokens closes over turns while goal information persists in residual representations, with architecture-specific failure modes.

0 favorites 0 likes
#instruction-following

Macro-Action Based Multi-Agent Instruction Following through Value Cancellation

arXiv cs.AI · 2026-05-14 Cached

Proposes MAVIC, a method for multi-agent reinforcement learning that corrects value estimates at instruction boundaries to enable compliance with external natural language instructions while preserving base task performance.

0 favorites 0 likes
#instruction-following

I taught my 1B to follow instructions. It got worse at following instructions...

Reddit r/LocalLLaMA · 2026-05-14

The author trained 1B, 2B, and 3B models with the same SFT recipe and observed that instruction-following (IFEval) regressed for the 1B and 2B models but improved for the 3B, possibly due to different learning rates or model capacity.

0 favorites 0 likes
#instruction-following

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

Hugging Face Daily Papers · 2026-05-14 Cached

This paper introduces AbstractEdit, a benchmark for abstract image editing, and Entity-Rubrics, a framework for entity-level assessment, revealing challenges in balancing intent and preservation for abstract instructions and highlighting the need for LLM integration.

0 favorites 0 likes
#instruction-following

Instructions shape Production of Language, not Processing

arXiv cs.CL · 2026-05-13 Cached

This research paper investigates how instructions influence large language models, finding that they primarily affect the language production stage rather than the processing stage. The study uses attention-based interventions and probing to demonstrate this asymmetry across various model families and tasks.

0 favorites 0 likes
#instruction-following

SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

arXiv cs.CL · 2026-05-08 Cached

This paper introduces Sequor, a new benchmark for evaluating how well AI models follow constraints in long, multi-turn conversations. It highlights that current models struggle significantly with maintaining instruction adherence over extended interactions.

0 favorites 0 likes
#instruction-following

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

Hugging Face Daily Papers · 2026-05-08 Cached

This paper introduces SEIF, a self-evolving reinforcement learning framework that enhances LLM instruction-following capabilities through iterative difficulty adaptation and co-training of instructor and follower components.

0 favorites 0 likes
#instruction-following

Dual-View Training for Instruction-Following Information Retrieval

Hugging Face Daily Papers · 2026-04-20 Cached

A dual-view data synthesis method using polarity reversal boosts instruction-following retrieval performance by 45% on the FollowIR benchmark.

0 favorites 0 likes
#instruction-following

Introducing Gemma 3 270M: The compact model for hyper-efficient AI

Google DeepMind Blog · 2025-10-23 Cached

Google introduces Gemma 3 270M, a compact 270-million parameter model designed for efficient on-device AI with strong instruction-following capabilities and extreme energy efficiency (0.75% battery for 25 conversations on Pixel 9 Pro).

0 favorites 0 likes
#instruction-following

Introducing GPT-4.1 in the API

OpenAI Blog · 2025-04-14 Cached

OpenAI launches GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano models via API with major improvements in coding (54.6% on SWE-bench), instruction following, and 1M token context windows at lower costs. GPT-4.5 Preview will be deprecated on July 14, 2025.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback