Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Summary
This paper introduces PatternEval, a diagnostic benchmark for evaluating response-pattern failures in hybrid-thinking multimodal large language models, and proposes PatternRL for aligning these patterns through reinforcement learning with specific penalties.
View Cached Full Text
Cached at: 08/24/26, 04:40 AM
Paper page - Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Source: https://huggingface.co/papers/2608.12781 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Hybrid-thinking multimodal language models suffer from response-pattern misalignment between thinking and non-thinking modes, which is addressed by a diagnostic benchmark and pattern-specific reinforcement learning penalties.
Hybrid-thinkingmultimodal large language models(MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap throughresponse-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introducePatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning.PatternEvaltests four recurrent failures:chain-of-thought leakage, response repetition,logical contradiction, andperformative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we developPatternRM, a response-level reward model, andPatternRL, which introduces pattern-specific penalties duringreinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties intoreinforcement learningcan mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together,PatternEvalandPatternRLprovide an evaluation-and-training framework for aligning user-visible response patterns acrosshybrid-thinkinginterfaces.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2608\.12781
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.12781 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.12781 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.12781 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
This paper introduces ReasonMatch-Bench, a benchmark for wide-baseline matching in multimodal LLMs, and proposes Dynamic Correspondence Reinforcement Learning (DCRL) to improve spatial reasoning. Experiments show significant gains on the benchmark while maintaining general performance.
Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages
The article introduces VSysBench, a benchmark for evaluating multimodal large language models on constraint compliance and answer correctness under system messages, finding that system messages reduce task accuracy and compliance varies between open-weight and proprietary models.
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
HumanLLM presents a framework for benchmarking and improving LLM anthropomorphism by modeling psychological patterns as interacting causal forces, constructing 244 patterns from academic literature and 11,359 multi-pattern scenarios. The approach demonstrates that authentic human alignment requires cognitive modeling rather than shallow behavioral mimicry, with HumanLLM-8B outperforming larger models like Qwen3-32B on multi-pattern dynamics.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
The paper introduces a behavioral alignment framework for personalized LLM judges in recommendation evaluation, addressing bidirectional rationalization where off-the-shelf LLMs argue both for and against user engagement on the same item. Their fine-tuned and preference-optimized approach achieves a 32.19% Macro-F1 lift over zero-shot and matches production feature-engineered baselines.