Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

arXiv cs.AI Papers

Summary

The paper introduces UXBench, a multimodal benchmark for evaluating MLLMs on mobile UX reasoning tasks, and presents UI-UX, a fine-tuned MLLM based on Qwen3-VL-4B-Thinking that achieves state-of-the-art performance on this benchmark.

arXiv:2606.13192v1 Announce Type: new Abstract: User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large language models (MLLMs) in the field of user interfaces is evolving rapidly, such as visual element grounding, graphical user interface (GUI) agents, and design-to-code generation. However, research efforts on evaluating UX based on UI screenshots are still immature. To address this, we propose UXBench, a novel multimodal benchmark consisting of 2,000 VQA data samples designed to assess MLLMs' ability to perform UI-based reasoning. UXBench includes 8 tasks based on real-world UI screenshots that require fine-grained diagnosis of UX issues across layout relationships, visual hierarchy, and content consistency. Our extensive evaluation of mainstream MLLMs shows that they remain fundamentally limited in their capacity for UI-based reasoning. The results underscore the need for further advancements in this area. To bridge this gap, we propose UI-UX, an MLLM based on Qwen3-VL-4B-Thinking foundation model and enhanced via reinforcement learning with two key innovations: a reward routing mechanism that dynamically balances perceptual understanding and logical reasoning during inference, and an asymmetric transition reward that suppresses redundant or insufficient reasoning steps. Experiments demonstrate that UI-UX achieves state-of-the-art (SOTA) performance on UXBench, attaining an accuracy of 0.7963 -- surpassing Claude-4.5-Sonnet's 0.6550 -- while exhibiting strong generalization across diverse UI tasks and maintaining low inference latency.
Original Article
View Cached Full Text

Cached at: 06/12/26, 08:55 AM

# Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Source: [https://arxiv.org/html/2606.13192](https://arxiv.org/html/2606.13192)
Ruichao Mao, Zhou Fang, Teng Guo, Hao Yang, Yaping Li, Shaohua Peng, Maji Huang, Xiaoyu Lin, Shuoyang Liu, Xuepeng Li, Yuyu Zhang, Hai Rao Ant Group \{maoruichao\.mrc, yingnan\.fz, gt447200, charles\.yh, pinky\.lyp, shaohua\.psh\}@antgroup\.com \{huangmaji\.hmj, axiao\.lxy, liushuoyang\.lsy, healy\.lxp, yuyu\.zyy, raohai\.rh\}@antgroup\.com

###### Abstract

User experience \(UX\) centered on usability, perceived consistency, and functional clarity is fundamental to real\-world user interfaces \(UI\)\. The application of multimodal large language models \(MLLMs\) in the field of user interfaces is evolving rapidly, such as visual element grounding, graphical user interface \(GUI\) agents, and design\-to\-code generation\. However, research efforts on evaluating UX based on UI screenshots are still immature\. To address this, we propose UXBench, a novel multimodal benchmark consisting of 2,000 VQA data samples designed to assess MLLMs’ ability to perform UI\-based reasoning\. UXBench includes 8 tasks based on real\-world UI screenshots that require fine\-grained diagnosis of UX issues across layout relationships, visual hierarchy, and content consistency\. Our extensive evaluation of mainstream MLLMs shows that they remain fundamentally limited in their capacity for UI\-based reasoning\. The results underscore the need for further advancements in this area\. To bridge this gap, we propose UI\-UX, an MLLM based on Qwen3\-VL\-4B\-Thinking foundation model and enhanced via reinforcement learning with two key innovations: a reward routing mechanism that dynamically balances perceptual understanding and logical reasoning during inference, and an asymmetric transition reward that suppresses redundant or insufficient reasoning steps\. Experiments demonstrate that UI\-UX achieves state\-of\-the\-art \(SOTA\) performance on UXBench, attaining an accuracy of 0\.7963—surpassing Claude\-4\.5\-Sonnet’s 0\.6550—while exhibiting strong generalization across diverse UI tasks and maintaining low inference latency\.

## 1Introduction

In modern human\-computer interaction systems, the determinant of success has shifted from mere functional implementation to the quality ofUser Experience \(UX\)\[[1](https://arxiv.org/html/2606.13192#bib.bib47)\]\. While UI design focuses on visual and interactive components such as layout, typography, and screen elements, UX design covers the entire user journey—emotional, cognitive, and behavioral responses before, during, and after interaction\[[7](https://arxiv.org/html/2606.13192#bib.bib48)\]\.

The rise of large language models \(LLMs\) and multimodal large language models \(MLLMs\) has greatly advanced UX in human\-computer interaction, enabling UI understanding\[[41](https://arxiv.org/html/2606.13192#bib.bib52),[21](https://arxiv.org/html/2606.13192#bib.bib53),[15](https://arxiv.org/html/2606.13192#bib.bib54)\], UI generation\[[39](https://arxiv.org/html/2606.13192#bib.bib39),[22](https://arxiv.org/html/2606.13192#bib.bib49),[13](https://arxiv.org/html/2606.13192#bib.bib50),[25](https://arxiv.org/html/2606.13192#bib.bib51)\], and user affect recognition\[[43](https://arxiv.org/html/2606.13192#bib.bib55),[30](https://arxiv.org/html/2606.13192#bib.bib56),[14](https://arxiv.org/html/2606.13192#bib.bib57)\]from raw visual or multimodal inputs\. These models provide capabilities such as pixel\-level interface parsing, intent\-driven UI synthesis, and emotion\-aware interaction modeling, positioning foundation models as scalable infrastructure for data\-driven, user\-centric interface systems\.

However, UX issues often arise from misalignments between design conventions and user mental models, rather than visible layout defects\. For example, a modal dialog occluding navigation bars or inconsistencies between advertised and actual content may appear visually correct yet result in frustration, eroded trust, and user churn\. This shift from “perceiving interfaces” to “inferring experiences” poses fundamental challenges for multimodal models aiming at automated UX evaluation\.

Existing UI\-based benchmarks such as Screen2Words\[[32](https://arxiv.org/html/2606.13192#bib.bib43)\], Mobile\-bench\[[12](https://arxiv.org/html/2606.13192#bib.bib58)\], and VisualWebBench\[[19](https://arxiv.org/html/2606.13192#bib.bib44)\]predominantly evaluate visual perception tasks \(e\.g\., caption generation, element detection, layout parsing\) rather thanUX reasoning\. These benchmarks lack objectives grounded in behavioral outcomes, cognitive psychology, and trust dynamics, making them insufficient for detecting design patterns that trigger user errors or negative emotions\. Consequently, current MLLMs remain limited in applications such as UI design assessment, automated testing, and intelligent design assistance\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x1.png)Figure 1:UXBench samples spanning Efficiency, Trustworthiness, and Usability dimensions\.Each case requires visual\-semantic reasoning to detect UX issues \(e\.g\., overlapping modals, deceptive content, missing controls\)\. Red bounding boxes indicate ground\-truth defect regions for quantitative assessment\. Tasks involve inferring experiential consequences beyond pixel\-level perception\.Recent work has explored some UI\-based tasks such as design\-to\-code generation\[[39](https://arxiv.org/html/2606.13192#bib.bib39)\], automated GUI operation, and component recognition\. Yet these approaches typically rely on manually labeled instruction–action pairs or oversimplify UX evaluation into binary “good/bad” classification, ignoring the multi\-dimensional, context\-dependent nature of UX issues\. On complex cases\(e\.g\., nested pop\-ups making operations irreversible, mismatched service names and functionalities\) current models often produce vague, erroneous, or contradictory reasoning\. This reveals deficiencies in semantic association, counterfactual reasoning, and intent modeling; aligning vision and text alone is insufficient for genuine UX understanding\.

To address these limitations, we introduceUXBench, the first multimodal benchmark for evaluating MLLMs’ UI\-UX reasoning capabilities\. UXBench contains over 2,000 real UI screenshots with user feedback, categorizing UX issues into three dimensions: usability, efficiency, and trustworthiness, and further into 8 fine\-grained diagnostic tasks \(See Fig\.[1](https://arxiv.org/html/2606.13192#S1.F1)\)\. Each task is posed as a two\- or three\-choice question requiring causal reasoning and mapping to design principles rather than keyword matching\. Large\-scale MLLM\-assisted annotation, combined with two rounds of expert validation by four senior UX specialists, ensures high label consistency\.

We further proposeUI\-UX, a reinforcement learning\-based MLLM enhancement framework built upon Qwen3\-VL\-4B\-Thinking\[[5](https://arxiv.org/html/2606.13192#bib.bib11),[33](https://arxiv.org/html/2606.13192#bib.bib12),[6](https://arxiv.org/html/2606.13192#bib.bib13)\]\. UI\-UX adopts task\-adaptive reward routing with accuracy rewards for UX diagnostic tasks, ROUGE\-L scores for semantic alignment in general UI understanding, and hit rewards for grounding tasks\. Additionally, we introduce an Asymmetric Transition Reward mechanism to suppress redundant reasoning and reduce inference latency\. Without manual preference annotation, the model is optimized end\-to\-end via the GRPO\[[27](https://arxiv.org/html/2606.13192#bib.bib45)\]algorithm, achieving 73\.8% accuracy on UXBench and demonstrating strong cross\-domain generalization\.

Our contributions are as follows:

- •First UX reasoning benchmark: We propose UXBench, defining 8 fine\-grained, evaluable, user\-centered UI diagnostic tasks, filling a critical gap in the field\.
- •Reward routing: We introduce hard negative sampling and semantic preservation augmentation strategies to address positive sample scarcity and extreme class imbalance in real scenarios\.
- •Efficient reasoning: The proposed UI\-UX model combines reward routing and overthinking penalty, achieving MLLM UX reasoning performance surpassing human experts, with low latency and high robustness\.

## 2Related Work

### 2\.1UI\-based Benchmarks

Recent benchmarks have focused on evaluating models’ visual understanding of user interfaces \(UI\)\. Screen2Words\[[32](https://arxiv.org/html/2606.13192#bib.bib43)\]and GUI\-Text\[[10](https://arxiv.org/html/2606.13192#bib.bib72)\]train models to generate UI descriptions from paired screenshot\-text data, but are limited to static element recognition and naming, without modeling interaction logic or user intent\. RICO\[[11](https://arxiv.org/html/2606.13192#bib.bib75)\]and VisualWebBench\[[19](https://arxiv.org/html/2606.13192#bib.bib44)\]further introduce layout parsing and element localization tasks, remaining within the “perception” layer—detecting buttons, reading text, and extracting structure—rather than judging whether a design may confuse users\. None of these benchmarks defines evaluation dimensions related to User Experience, and thus cannot measure reasoning such as whether a pop\-up obscures critical operations or a badge misleads users\. Our work fills this gap by constructing the first benchmark centered on fine\-grained UX diagnostics, enabling models to move from understanding interfaces to understanding users\.

### 2\.2UX Issue and Detection Methods

User Interfaces \(UIs\) often suffer from issues such as text overlap, missing images, and layout distortion arising from device diversity and design complexity, undermining usability and accessibility\. Vision\-based tools like OwlEyes\-online\[[29](https://arxiv.org/html/2606.13192#bib.bib19)\]and Nighthawk\[[20](https://arxiv.org/html/2606.13192#bib.bib20)\]can localize visual bugs, while Metamorphosis\[[28](https://arxiv.org/html/2606.13192#bib.bib22)\]detects scaling defects via metamorphic testing\. However, most taxonomy\-based approaches address only predefined issue types in limited or synthetic datasets, overlooking higher\-level design smells and cross\-component inconsistencies\. UISGPT\[[38](https://arxiv.org/html/2606.13192#bib.bib26)\]applies large language models to identify guideline violations with explanations, but struggles with complex, context\-sensitive reasoning\. Our work differs by constructing a rich, real\-world dataset that captures both visual and structural UX issues, and by employing MLLMs with causal reasoning to detect and explain complex design flaws\.

### 2\.3Reasoning in MLLMs

Reasoning\-capable MLLMs have advanced complex visual reasoning by introducing explicit*Chain\-of\-Thought*\(CoT\) generation and self\-consistency verification\. Examples include LLaVA\-1\.5\[[16](https://arxiv.org/html/2606.13192#bib.bib79)\]for multi\-step image reasoning, MiniGPT\-4\-v2\[[44](https://arxiv.org/html/2606.13192#bib.bib46)\]and InternVL\[[9](https://arxiv.org/html/2606.13192#bib.bib38),[45](https://arxiv.org/html/2606.13192#bib.bib33),[34](https://arxiv.org/html/2606.13192#bib.bib32)\]for instruction\-tuned logical scene analysis, and CogVLM and Visual\-CoT\[[26](https://arxiv.org/html/2606.13192#bib.bib60)\]for structured output reasoning in tasks like MMMU and ChartQA\. Recent works such as GRPO\-λ\\lambda\[[23](https://arxiv.org/html/2606.13192#bib.bib61)\], Step Pruner\[[36](https://arxiv.org/html/2606.13192#bib.bib65)\], and CoRE\-Eval\[[42](https://arxiv.org/html/2606.13192#bib.bib70)\]address reasoning efficiency by pruning redundant steps and assessing step\-level importance, providing insights for reducing latency while preserving accuracy\. Despite these advances, most models still emphasize fact\-based reasoning over*experience\-based reasoning*, leaving them unable to identify UX issues such as inappropriate button placement, disruptive pop\-ups, or misleading descriptions\. Our work is the first to embed such reasoning into the human–computer interaction context, enabling MLLMs to perform user\-centric causal inference in UI scenarios\.

## 3UXBench

Existing vision–language benchmarks mainly address general scene understanding and image generation, with limited evaluation of models’ reasoning and comprehension capabilities in real\-world UI\. To address this gap, we presentUXBench—the first vision–language benchmark for UX defect diagnosis—designed to systematically assess multimodal large language models \(MLLMs\) on fine\-grained, higher\-order reasoning within realistic UI scenarios\.

### 3\.1Task Definition

Traditional UI recognition focuses on visual perception, detecting and classifying visible interface components from screenshots and reconstructing their layout\. These methods rely on explicit feature extraction, producing outputs derivable directly from pixel information—without modeling user behavior or intent\. In contrast, UX diagnosis requires identifying components and evaluating whether their arrangement, interaction logic, and semantics pose potential experience risks\. Such issues often stem from mismatches between design conventions and human usage habits\. For example, assessing whether “a popup occludes a button” involves detection, spatial analysis, and causal inference about operational impact—reflecting reasoning beyond pixel\-level perception\.

For fine\-grained evaluation, we organize UX diagnosis into three dimensions operationalized from rigorous HCI frameworks:

1. 1\.Usability: Clarity of operation and feedback\. This dimension aligns with ”Operability”\[[35](https://arxiv.org/html/2606.13192#bib.bib76)\]in HCI frameworks, focusing on the visibility and accessibility of screen objects\. 1. \(a\)BubbleOcclT: Textual overlay occluding page text\. 2. \(b\)BubbleOcclBtn: Textual overlay blocking clickable elements\.
2. 2\.Efficiency: Minimizing operational and cognitive cost\. Empirical data demonstrates that efficiency is the highest\-rated usability factor \(mean: 4\.07/5\)\[[35](https://arxiv.org/html/2606.13192#bib.bib76)\]\. 1. \(a\)PopupNoClose: Popup without explicit close control\. 2. \(b\)PopupBlockClose: Popup affecting native close button clickability\. 3. \(c\)PopupStack: Multiple modal popups present simultaneously\.
3. 3\.Trustworthiness: Maintaining consistency and credibility\. This dimension operationalizes ”Persuasiveness” and ”Security”\[[8](https://arxiv.org/html/2606.13192#bib.bib77)\]from HCI frameworks\. 1. \(a\)MismatchBadge: Badge content inconsistent with landing page\. 2. \(b\)MismatchContent: Service name inconsistent with page text\. 3. \(c\)MismatchFunc: Description inconsistent with provided functionality\.

### 3\.2Data Pipeline

UXBench is built through a multi\-stage pipeline combining large\-scale real user feedback, MLLM\-assisted annotation, and expert quality control\.

Raw data collection: screenshots and textual descriptions from in\-app feedback across diverse mobile application scenarios\.

Relevance filtering: Gemini\-2\.5\-Pro classifies relevant feedback; a fine\-tuned Qwen3\-VL\-2B replicates decisions at scale to retain high\-confidence UX samples\.

Core categorization: relevant samples are labeled by Gemini into dimensions and subtasks using few\-shot prompting, with multi\-round voting for ambiguous cases\.

Human verification: Four senior UX researchers conduct two\-stage manual validation, beginning with independent annotation by all four researchers followed by cross\-validation with disagreement resolution\.

Final dataset: balanced sampling from positive and negative pools yields 2,000 high\-quality image–question pairs covering typical issues and normal UI scenarios\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x2.png)Figure 2:Data distribution in UXBench\.\(A\) Distribution across different subtasks in the benchmark dataset \(2000 samples total\)\. \(B\) Distribution of user interaction options
### 3\.3Data Distribution

UXBench emphasizes diversity across tasks and option distribution\. \(See Fig[2](https://arxiv.org/html/2606.13192#S3.F2)\)

Tasks: three dimensions with 2–3 subtasks each, averaging 200–300 instances; positive and negative cases are balanced\.

Options: 40% with 2 choices, 60% with 3 choices\.

Coverage: includes iOS/Android platforms, light/dark themes, and multiple orientations for better generalization\.

## 4UI\-UX

To bridge the significant gap between existing Multimodal Large Language Models \(MLLMs\) and human experts in user interface \(UI\) reasoning, we introduce UI\-UX, a reinforcement learning\-enhanced model specifically optimized for UX diagnostic tasks\. Built upon the Qwen\-VL\-4B\-Thinking architecture, UI\-UX employs task\-aware reinforcement learning \(RL\) to provide end\-to\-end guidance for the reasoning generation process, achieving substantial improvements in reasoning efficiency while maintaining diagnostic accuracy\.

### 4\.1Training Data Composition

Dataset Collection and Annotation\.We collect a large\-scale dataset of real\-world UI screenshots via automated testing scripts executed on eight physical devices spanning Android, HarmonyOS, and iOS\. By systematically navigating core user flows in 1,200\+ popular applications and webpages, we gather 832,432 raw screenshots\(See Fig[3](https://arxiv.org/html/2606.13192#S4.F3)\)\. To mitigate visual redundancy caused by deterministic execution, we apply perceptual hashing \(pHash\) and retain only one sample per group of images with Hamming distance≤5\\leq 5, where the Hamming distance between two pHash \(nnbit\) vectors𝐡i,𝐡j∈\{0,1\}n\\mathbf\{h\}\_\{i\},\\mathbf\{h\}\_\{j\}\\in\\\{0,1\\\}^\{n\}is defined as:

dH​\(𝐡i,𝐡j\)=∑k=1n𝕀​\[hi​\[k\]≠hj​\[k\]\]d\_\{H\}\(\\mathbf\{h\}\_\{i\},\\mathbf\{h\}\_\{j\}\)=\\sum\_\{k=1\}^\{n\}\\mathbb\{I\}\\big\[h\_\{i\}\[k\]\\neq h\_\{j\}\[k\]\\big\]\(1\)
and𝕀​\[⋅\]\\mathbb\{I\}\[\\cdot\]denotes the indicator function\. This yields 68,138 unique UI screenshots\. For scalable, high\-fidelity labeling, we propose a two\-stage annotation pipeline\. First, we leverage multiple MLLMs \(GPT5, Gemini2\.5\-pro, Claude4\.5\) to predict UX issues across three dimensions \(Usability, Efficiency, Content Consistency\) and eight sub\-tasks, using carefully engineered prompts\. Predictions are aggregated via majority voting to generate robust pseudo\-labels\. Second, all pseudo\-labels are manually verified and corrected by five experienced UX researchers, ensuring label accuracy\.

Positive–Negative Sample Balancing\.UX issue detection suffers from severe class imbalance where positive samples \(defective UIs\) are heavily outnumbered\. This causes models to bias toward the majority class and makes GRPO training degenerate when all samples in a group share the same label\.

We apply hard negative mining by sampling each negative image 8 times with Qwen3\-VL\-Thinking\-4B \(temperature = 1\.0\)\. Only samples with inconsistent predictions \(≤5/8\\leq 5/8votes\) are retained as hard negatives, filtering out easy cases and ensuring challenging training examples\.

Mixed\-Task Regularization\.Training exclusively on the UX issue dataset triggers catastrophic forgetting, eroding the model’s general UI understanding and degrading performance\. To mitigate this, we incorporate 4,919 randomly sampled examples from the MultiUI dataset\[[18](https://arxiv.org/html/2606.13192#bib.bib74)\]—a large\-scale corpus of 7\.3M UI samples\.

Final Training Data Composition\.Our final training dataset contains 26,680 samples strategically distributed across UX\-specific and multi\-domain data\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x3.png)Figure 3:UI\-UX training pipeline overview\.\(1\) Raw Data Collection:6M\+ screenshots from 1,200\+ apps/websites, deduplicated via pHash\.\(2\) Label Generation:MLLM\-based pseudo\-labeling with 5× positive augmentation and 8× hard negative mining\.\(3\) Final Datasets:21,761 UX samples across eight tasks \+ 4,919 MultiUI samples for regularization\.\(4\) Training:RL optimization with asymmetric transition reward and reward router on Qwen3\-VL\-4B\-Thinking\.
### 4\.2Reward Design

Our reward function is formulated as:

ℛ=ℛrouting\+ℛtransition\\mathcal\{R\}=\\mathcal\{R\}\_\{\\text\{routing\}\}\+\\mathcal\{R\}\_\{\\text\{transition\}\}\(2\)
whereℛrouting\\mathcal\{R\}\_\{\\text\{routing\}\}is a task\-adaptive reward that dynamically selects an accuracy\-oriented metric according to the downstream task \(e\.g\., UX reasoning accuracy, OCR edit distance, or grounding IoU\), andℛtransition\\mathcal\{R\}\_\{\\text\{transition\}\}encourages concise reasoning by penalizing redundant or superfluous inference steps\. Empirical results demonstrate that incorporatingℛtransition\\mathcal\{R\}\_\{\\text\{transition\}\}significantly improves both training efficiency and inference speed\(FLOPs and latency\), without compromising task performance, thereby promoting succinct yet accurate multimodal reasoning\.

Reward Routing\.We propose a task\-aware reward routing mechanism that dynamically selects the most appropriate reward function based on the origin and format of each training sample, effectively decoupling optimization objectives across heterogeneous multimodal tasks\. For UX issue detect samples, typically presented as multiple\-choice questions requiring precise mathematical or logical reasoning\. we adopt answer accuracy \(MathAcc\) as the reward, computed via robust LaTeX parsing and semantic equivalence verification to determine whether the model’s final prediction matches the ground\-truth solution\. For MultiUI general understanding samples, where the model is required to generate natural\-language descriptions of UI elements or scenes, we employ ROUGE\-L to quantify textual fidelity, defined as:

ℛROUGE\-L=\(1\+β2\)⋅PLCS⋅RLCSβ2⋅PLCS\+RLCS\\mathcal\{R\}\_\{\\text\{ROUGE\-L\}\}=\\frac\{\(1\+\\beta^\{2\}\)\\cdot P\_\{\\text\{LCS\}\}\\cdot R\_\{\\text\{LCS\}\}\}\{\\beta^\{2\}\\cdot P\_\{\\text\{LCS\}\}\+R\_\{\\text\{LCS\}\}\}\(3\)
wherePLCSP\_\{\\text\{LCS\}\}andRLCSR\_\{\\text\{LCS\}\}denote precision and recall based on LCS\.

For visual grounding tasks:

ℛhit=𝕀​\(\(xcpred,ycpred\)∈\[x1gt,x2gt\]×\[y1gt,y2gt\]\)\\mathcal\{R\}\_\{\\text\{hit\}\}=\\mathbb\{I\}\\left\(\(x\_\{c\}^\{\\text\{pred\}\},y\_\{c\}^\{\\text\{pred\}\}\)\\in\[x\_\{1\}^\{\\text\{gt\}\},x\_\{2\}^\{\\text\{gt\}\}\]\\times\[y\_\{1\}^\{\\text\{gt\}\},y\_\{2\}^\{\\text\{gt\}\}\]\\right\)\(4\)
Asymmetric Transition Reward\.Transition markers indicate critical reasoning transitions within chains\. Motivated by the correlation between transition markers and reasoning quality, we design an asymmetric reward function to balance reasoning sufficiency and conciseness\. Given a model predictionppwith transition marker countT​\(p\)T\(p\)and ground truth labelyy, we define the correctness indicator𝟏p=y\\mathbf\{1\}\_\{p=y\}\. The complete reward function is formulated as: The reward function is:

R​\(p,y\)\\displaystyle R\(p,y\)=1p=y⋅Rcorrect​\(T​\(p\)\)\\displaystyle=1\_\{p=y\}\\cdot R\_\{\\text\{correct\}\}\\big\(T\(p\)\\big\)\(5\)\+\(1−1p=y\)⋅Rincorrect​\(T​\(p\)\)\\displaystyle\\quad\+\\big\(1\-1\_\{p=y\}\\big\)\\cdot R\_\{\\text\{incorrect\}\}\\big\(T\(p\)\\big\)
Where the two branch functions are defined as:

Rcorrect​\(T\)\\displaystyle R\_\{\\text\{correct\}\}\(T\)=max⁡\(rbase−α⋅T,rmin\)\\displaystyle=\\max\\big\(r\_\{\\text\{base\}\}\-\\alpha\\cdot T,\\;r\_\{\\min\}\\big\)\(6\)Rincorrect​\(T\)\\displaystyle R\_\{\\text\{incorrect\}\}\(T\)=min⁡\(α⋅T,rmax\)\\displaystyle=\\min\\big\(\\alpha\\cdot T,\\;r\_\{\\max\}\\big\)Hererbase=1\.0r\_\{\\text\{base\}\}=1\.0denotes the base reward,α\\alphais the penalty coefficient, andrminr\_\{\\min\},rmaxr\_\{\\max\}represent the reward boundaries for correct and incorrect predictions, respectively\. This design embodies three core mechanisms: \(1\) Penalty for correct predictions: reward decreases linearly with transition markers, discouraging over\-deliberation after reaching correct answers, while the lower boundrminr\_\{\\min\}ensures that even verbose correct answers outperform incorrect ones; \(2\) Exploration incentive for incorrect predictions: reward increases linearly with transition markers, encouraging more thorough reasoning exploration when uncertain, while the upper boundrmaxr\_\{\\max\}prevents exploiting the reward by simply adding redundancy; \(3\) Correctness\-first guarantee: the constraintrmin\>rmaxr\_\{\\min\}\>r\_\{\\max\}mathematically ensures that any correct answer strictly yields higher reward than any incorrect answer\.

The key innovation of this design lies in establishing an insurmountable reward gap𝒢\\mathcal\{G\}that mathematically enforces the correctness\-first principle:

𝒢=infT≥0Rcorrect​\(T\)−supT≥0Rincorrect​\(T\)=rmin−rmax\>0\\mathcal\{G\}=\\inf\_\{T\\geq 0\}R\_\{\\text\{correct\}\}\(T\)\-\\sup\_\{T\\geq 0\}R\_\{\\text\{incorrect\}\}\(T\)=r\_\{\\min\}\-r\_\{\\max\}\>0\(7\)This strictly positive reward gap guarantees that for any transition marker countsT1,T2≥0T\_\{1\},T\_\{2\}\\geq 0, we haveRcorrect​\(T1\)\>Rincorrect​\(T2\)R\_\{\\text\{correct\}\}\(T\_\{1\}\)\>R\_\{\\text\{incorrect\}\}\(T\_\{2\}\)\. During policy optimization, letπθ\\pi\_\{\\theta\}denote the parameterized policy\. The expected reward naturally decomposes as:

maxθ⁡𝔼πθ​\[R\]\\displaystyle\\max\_\{\\theta\}\\ \\mathbb\{E\}\_\{\\pi\_\{\\theta\}\}\[R\]=maxθ\(𝔼πθ\[1p=y\]⋅𝔼\[Rcorrect∣p=y\]\\displaystyle=\\max\_\{\\theta\}\\Big\(\\mathbb\{E\}\_\{\\pi\_\{\\theta\}\}\[1\_\{p=y\}\]\\cdot\\mathbb\{E\}\[R\_\{\\text\{correct\}\}\\mid p=y\]\(8\)\+\(1−𝔼πθ\[1p=y\]\)⋅𝔼\[Rincorrect∣p≠y\]\)\\displaystyle\\quad\+\\big\(1\-\\mathbb\{E\}\_\{\\pi\_\{\\theta\}\}\[1\_\{p=y\}\]\\big\)\\cdot\\mathbb\{E\}\[R\_\{\\text\{incorrect\}\}\\mid p\\neq y\]\\Big\)
Since𝒢\>0\\mathcal\{G\}\>0, the gain from improving accuracy𝔼πθ​\[𝟏p=y\]\\mathbb\{E\}\_\{\\pi\_\{\\theta\}\}\\big\[\\mathbf\{1\}\_\{p=y\}\\big\]always dominates conciseness optimization, theoretically preventing the degenerate solution of ”sacrificing correctness for conciseness\.”

For transition marker counting, we employ structured matching: let𝒱t\\mathcal\{V\}\_\{t\}denote the transition marker vocabulary and𝒫\\mathcal\{P\}the position prefix set \(e\.g\., ”\\n”, ”–”\)\. For textp=\{w1,…,wn\}p=\\\{w\_\{1\},\\dots,w\_\{n\}\\\}, we define

T​\(p\)=∑i=2n𝟏wi∈𝒱t⋅𝟏∃k∈𝒫:wi−1=kT\(p\)=\\sum\_\{i=2\}^\{n\}\\mathbf\{1\}\_\{w\_\{i\}\\in\\mathcal\{V\}\_\{t\}\}\\cdot\\mathbf\{1\}\_\{\\exists k\\in\\mathcal\{P\}:w\_\{i\-1\}=k\}\(9\)The position constraint filters out casual in–sentence transitions, capturing only explicit logical transitions in the reasoning path\.

### 4\.3Training Details

We train all models using GRPO, a sample\-efficient on\-policy RL algorithm that stabilizes policy updates via per\-token reward weighting and group\-wise normalization\. Training is conducted over 130 hours on a 16\-PPU cluster using Qwen3\-VL\-4B\-Thinking as the base model\. We apply LoRA with rankr=8r=8and scalingα=32\\alpha=32to all linear layers, while freezing the vision encoder \(ViT\) and vision–language aligner to preserve pretrained representations\. Training uses DeepSpeed ZeRO\-2 for memory efficiency, and vLLM with tensor parallelism \(4 GPUs per node\) and 50% GPU memory utilization to enable high\-throughput rollout generation\. Maximum sequence length is 16K tokens, with a completion length cap of 8K to support extended reasoning chains\. Each batch contains 2 samples per GPU \(effective batch size=32=32\), and 16 responses are generated per prompt during rollouts usingtemperature=1\.0\\mathrm\{temperature=1\.0\},top−\\mathrm\{top\-\}k=80k=80, andtop−\\mathrm\{top\-\}k=1\.0k=1\.0\. We optimize with a learning rate of2×10−62\\times 10^\{\-6\}, linear warmup over 5% of steps, followed by cosine decay over one epoch\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x4.png)Figure 4:Positive\-negative distribution before and after balanced sampling\.\(a\) Original data shows severe class imbalance with positive samples \(red\) heavily outnumbered by negative samples \(blue\)\. \(b\) After applying hard negative mining and positive augmentation, the dataset achieves improved balance across all eight tasks\.Table 1:Performance comparison of different models on UXBench\. Reasoning models generally outperform instruct models across tasks\.∗indicates format parsing failures or excessively long reasoning outputs during evaluation\. Bold values represent the best performance across all models for each metric\.ModelParamsBubble OccTBubble OccBtnPopup No ClosePopup Block ClosePopup StackMismatch BadgeMismatch ContentMismatch FuncAVG\.Instruct ModelLlava3\-Next\[[17](https://arxiv.org/html/2606.13192#bib.bib27)\]8B0∗0\.10\.220\.220\.260\.190∗0∗0\.1200Qwen2\.5\-VL\[[6](https://arxiv.org/html/2606.13192#bib.bib13)\]72B0\.09∗0\.510\.700\.550\.560\.690\.680\.710\.5482Qwen3\-VL\[[24](https://arxiv.org/html/2606.13192#bib.bib17)\]235B0\.300\.570\.790\.400\.520\.620\.640\.700\.5600InternVL\-3\[[34](https://arxiv.org/html/2606.13192#bib.bib32)\]2B0\.30\.340\.610\.00\.340\.520\.110\.080\.2875InternVL\-3\.5\[[34](https://arxiv.org/html/2606.13192#bib.bib32)\]2B0\.420\.340\.430\.260\.370\.610\.720\.760\.4888MiniCPM\-V\-4\.5\[[40](https://arxiv.org/html/2606.13192#bib.bib41)\]8B0\.30\.390\.560\.330\.550\.530\.650\.760\.5086Reasoning ModelGLM\-4\.1\-Thinking\[[31](https://arxiv.org/html/2606.13192#bib.bib73)\]9B0∗0\.60\.620\.340\.320\.320\.04∗0\.01∗0\.2813Qwen3\-VL\-Thinking\[[24](https://arxiv.org/html/2606.13192#bib.bib17)\]4B0\.470\.490\.720\.390\.500\.520\.630\.650\.5254Qwen3\-VL\-Thinking\[[24](https://arxiv.org/html/2606.13192#bib.bib17)\]235B0\.520\.610\.820\.470\.560\.540\.640\.700\.5854InternVL\-3\.5\[[34](https://arxiv.org/html/2606.13192#bib.bib32)\]4B0∗0\.390\.650\.380\.500\.590\.680\.720\.4888MimoVL\-0528\[[37](https://arxiv.org/html/2606.13192#bib.bib40)\]7B0\.580\.560\.740\.410\.520\.540\.400\.60\.5438Claude\-3\.7\-Sonnet\[[2](https://arxiv.org/html/2606.13192#bib.bib14)\]–0\.640\.660\.780\.560\.530\.610\.650\.760\.6488Claude\-4\-Sonnet\[[4](https://arxiv.org/html/2606.13192#bib.bib15)\]–0\.640\.520\.770\.460\.520\.680\.660\.740\.6238Claude\-4\.5\-Sonnet\[[3](https://arxiv.org/html/2606.13192#bib.bib16)\]–0\.650\.550\.770\.530\.660\.70\.660\.720\.6550UI\-UX\(ours\)4B0\.790\.880\.790\.770\.940\.720\.710\.760\.7963

## 5Experiment

### 5\.1Main Results on UXBench

Evaluation Setup\.We evaluate all models under consistent conditions with temperature set to 0 and maximum generation length of 8,192 tokens to ensure fair comparison across different architectures\.

Overall Performance\.Table[1](https://arxiv.org/html/2606.13192#S4.T1)presents the comprehensive evaluation results across eight UXBench tasks spanning usability, efficiency, and trustworthiness dimensions\. Our UI\-UX model, with only 4B parameters, achieves the best overall performance with an average score of 0\.7963, substantially outperforming both instruct and reasoning models\. Specifically, UI\-UX \(ours\) surpasses the best reasoning model Claude\-4\.5\-Sonnet \(0\.6550\) by 21\.6% and the top instruct model Qwen3\-VL 235B \(0\.56\) by 42\.2%, demonstrating the effectiveness of our domain\-specific design on UXBench despite being significantly smaller in scale\. Additionally, reasoning models generally demonstrate superior performance compared to instruct models on UXBench\. For instance, Qwen3\-VL\-Thinking 235B \(0\.5854\) outperforms its instruct counterpart Qwen3\-VL 235B \(0\.56\) by 4\.5%, highlighting the importance of chain\-of\-thought reasoning capabilities for these challenging benchmarks even when controlling for model scale\.

Overthinking Problem\.Despite their superior accuracy, reasoning models face a significant challenge on UXBench: overthinking that causes generation to exceed the 8,192 token limit, resulting in parsing failures \(marked with∗\)\. This issue disproportionately affects smaller models\. For instance, GLM\-4\.1\-Thinking \(9B\) suffers from multiple parsing failures \(Bubble OccT: 0∗, Mismatch Badge: 0\.04∗, Mismatch Content: 0\.01∗\), achieving only 0\.2813 average, while InternVL\-3\.5 \(4B\) fails on Bubble OccT \(0∗\) with 0\.4888 overall\. In contrast, larger reasoning models like Qwen3\-VL\-thinking \(235B\) and Claude\-4\.5\-Sonnet successfully avoid these failures\. This suggests that model scale is critical for controlling reasoning verbosity while preserving reasoning quality on UXBench’s complex visual reasoning tasks\.

### 5\.2Ablation Study on Training Strategies

To investigate the impact of different data sampling and augmentation strategies on model performance, we conduct comprehensive ablation studies on UXBench, as shown in Table[4](https://arxiv.org/html/2606.13192#S5.T4)\. All experiments are based on Qwen3\-VL\-4B\-Thinking\.

Baseline Performance\.The baseline model achieves an accuracy of 52\.54% on UXBench\. However, directly training with the full dataset before balancing suffers from severe reward hacking—the model exploits spurious correlations in training data, leading to inflated training metrics that fail to generalize to the evaluation set\. This makes the unbalanced training configuration impractical for real\-world deployment\.

Hard Negative Mining \(HNM\)\.Incorporating hard negative mining effectively addresses the reward hacking issue by prioritizing challenging samples that expose model weaknesses\. By focusing on frequently misclassified cases \(e\.g\., subtle modal overlaps, deceptive UI patterns\), HNM guides the model to learn more discriminative decision boundaries\. This strategy alone yields a substantial gain of \+19\.41% \(71\.95%\), demonstrating its critical role in mitigating dataset bias and establishing a robust training foundation\.

Positive Sample Strategies\.Building upon HNM, we compare two approaches for enriching positive samples\. Positive resampling increases the sampling frequency of existing defect instances, achieving 75\.79% \(\+23\.25%\)\. In contrast, positive augmentation generates diverse variations through layout transformations and content perturbations, reaching 77\.71% \(\+25\.17%\)\. The augmentation approach outperforms resampling by \+1\.92%, as it introduces greater sample diversity, exposing the model to varied visual contexts while preserving semantic defect patterns, thereby reducing overfitting to specific UI layouts\.

MultiUI Integration\.Finally, combining HNM and positive augmentation with MultiUI training achieves the best performance at 79\.63% \(\+27\.09%\)\. The MultiUI approach exposes the model to diverse interface designs, interaction paradigms, and visual styles beyond the target domain\. This cross\-domain training enhances generalization capability, enabling robust UX assessment across heterogeneous applications\. Notably, MultiUI contributes an additional \+1\.92% improvement over positive augmentation alone, validating its effectiveness in building cross\-domain robustness\.

The ablation study reveals three critical insights: \(1\) Unbalanced training causes reward hacking: under extreme imbalance, the model predicts all negatives, yielding 40% accuracy but 0% meaningful accuracy, while hard negative mining provides a \+19\.41% gain by addressing this issue\. \(2\) Augmentation\-based diversity \(\+25\.17%\) is more effective than resampling \(\+23\.25%\) for positive samples, with a \+1\.92% advantage\. \(3\) Multi\-domain exposure through MultiUI is essential for achieving state\-of\-the\-art performance \(79\.63%\), demonstrating that generalizable UX reasoning requires both within\-domain hard case coverage and cross\-domain knowledge transfer\.

Table 2:Ablation study on training strategies for UXBench\.
### 5\.3Analysis of Asymmetric Transition Reward

To validate the design rationale of the Asymmetric Transition Reward proposed in Section 4\.2, we conduct a systematic quantitative analysis focused on the core reward mechanism\. Note that this analysis uses augmented data from 8 visual understanding tasks without MultiUI integration, allowing us to isolate and evaluate the effectiveness of the transition reward mechanism independent of cross\-domain training effects\. We collect 2,000 reasoning samples and categorize them into “correct” \(1,092 samples, 54\.6%\) and “incorrect” \(908 samples, 45\.4%\) groups based on the correctness of final answers\. For each sample, we count the number of transition markersTT\(e\.g\., “but”, “however”\)\.

Significant Distributional Differences in Transition Markers\.Table[3](https://arxiv.org/html/2606.13192#S5.T3)presents key statistics on transition marker distribution\. Incorrect samples exhibit an average transition marker count \(14\.38\) nearly 6 times higher than correct samples \(2\.48\)\. Most critically, 31\.5% of incorrect samples haveT\>3T\>3, compared to only 11\.7% of correct samples—a 2\.7\-fold difference that provides empirical support for selectingT=3T=3as our penalty threshold\.

Table 3:Transition Marker Distribution Statistics for Correct vs\. Incorrect SamplesMetricCorrectIncorrectConclusionTotal Samples1,092 \(55%\)908 \(45%\)\-MeanTT2\.4814\.38Incorrect samples exhibitsignificantly higher“overthinking”MedianTT00Both distributionsskewed towardlow valuesT=0T=0Ratio69\.0%50\.7%Correct samples tendto be “confidentand correct”T\>3T\>3Ratio11\.7%31\.5%Key finding:T\>3T\>3isa strong indicator ofhigh error ratesRelationship Between Transition Markers and Accuracy\.Figure[5](https://arxiv.org/html/2606.13192#S5.F5)demonstrates that accuracy monotonically decreases from 61\.0% to 14\.8% as transition markers increase—a 4\.1\-fold reduction\. The reward metric exhibits a parallel declining trend, validating that our transition marker\-based mechanism effectively captures reasoning quality\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x5.png)Figure 5:Accuracy and Reward vs\. Transition Marker Intervals\. \(a\) Accuracy decreases monotonically as transition marker count increases; \(b\) Average reward exhibits a similar declining trend, validating the effectiveness of the transition marker\-based reward mechanism\.Empirical Validation of the Reward Function\.Figure[6](https://arxiv.org/html/2606.13192#S5.F6)validates our asymmetric reward design\. For correct samples, transition markers exhibit a strong negative correlation with reward \(r=−0\.728r=\-0\.728,p<0\.001p<0\.001\), matching our theoretical curveRcorrect=max⁡\(1\.0−0\.03​T,0\.5\)R\_\{\\text\{correct\}\}=\\max\(1\.0\-0\.03T,0\.5\)\. For incorrect samples, the correlation is strongly positive \(r=\+0\.774r=\+0\.774,p<0\.001p<0\.001\), followingRincorrect=min⁡\(0\.03​T,0\.4\)R\_\{\\text\{incorrect\}\}=\\min\(0\.03T,0\.4\)\. These correlations confirm that our reward functions accurately reflect the intended design\.

![Refer to caption](https://arxiv.org/html/2606.13192v1/x6.png)Figure 6:Reward vs\. Transition Markers for Correct and Incorrect Samples\. \(a\) Correct samples: negative correlation \(r=−0\.728r=\-0\.728\) between transition markers and reward; \(b\) Incorrect samples: positive correlation \(r=\+0\.774r=\+0\.774\), with upper bound at 0\.4 preventing over\-rewarding of verbosity\.Effectiveness Validation\.We compare our method against a baseline that provides rewards solely based on answer correctness without transition marker penalties\. Our method jointly optimizes MathAccuracy and Asymmetric Transition Reward \(weights 2\.0 and 1\.0\) with a protection period for the first 40% of training steps\.

Table 4:Effectiveness validation results\. \(Acc: Accuracy, TR: Transition Reward, Len: Mean output length in tokens\.\)Our method achieves comparable accuracy \(77\.71% vs\. 76\.75%\) while reducing generation length by 81\.1% \(from 1770 to 334 tokens\)\. The transition reward of 0\.926 approaches the ideal value of 1\.0, demonstrating effective verbosity control\. This validates that transition markers successfully identify and penalize inefficient reasoning without sacrificing accuracy\.

## 6Conclusion

We presentedUXBench, a multimodal benchmark for fine\-grained UI\-based reasoning, andUI\-UX, a reinforcement learning framework that improves MLLM performance via task\-adaptive rewards and overthinking penalties\. Our approach achieves strong accuracy and generalization without manual preference data\.

Future work will expand UXBench with richer dimensions from cognitive psychology and real interaction logs, and integrate UI\-UX into practical design assistants and automated testing, advancing AI\-driven, human\-centric interface evaluation\.

## References

- \[1\]\(2025\)The role of large language models in ui/ux design: a systematic literature review\.External Links:2507\.04469,[Link](https://arxiv.org/abs/2507.04469)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p1.1.1)\.
- \[2\]Anthropic\(2025\)Claude 3\.7 sonnet and claude code\.Note:[https://www\.anthropic\.com/news/claude\-3\-7\-sonnet](https://www.anthropic.com/news/claude-3-7-sonnet)Accessed: 2025\-02\-25Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.19.11.1.1.1)\.
- \[3\]Anthropic\(2025\)Introducing claude 4\.5\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-4\-5](https://www.anthropic.com/news/claude-sonnet-4-5)Accessed: 2025\-09\-30Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.21.13.1.1.1)\.
- \[4\]Anthropic\(2025\)Introducing claude 4\.Note:[https://www\.anthropic\.com/news/claude\-4](https://www.anthropic.com/news/claude-4)Accessed: 2025\-05\-23Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.20.12.1.1.1)\.
- \[5\]J\. Bai, S\. Bai, S\. Yang, S\. Wang, S\. Tan, P\. Wang, J\. Lin, C\. Zhou, and J\. Zhou\(2023\)Qwen\-vl: a versatile vision\-language model for understanding, localization, text reading, and beyond\.arXiv preprint arXiv:2308\.12966\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p7.1)\.
- \[6\]S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. Lin\(2025\)Qwen2\.5\-vl technical report\.arXiv preprint arXiv:2502\.13923\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p7.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.6.4.2.1.1)\.
- \[7\]P\. Bharath, D\. Damodhar,et al\.\(2023\)From leader to laggard: an analysis of blackberry’s ui/ux missteps and the decline of a tech giant\.Milestone Transactions on Futuristic Engineering1\(1\),pp\. 1–12\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p1.1)\.
- \[8\]E\. Brangier, J\. G\. Urrutia, V\. Senderowicz, and L\. Cessat\(2018\)Beyond ”usability and user experience” , towards an integrative heuristic inspection: from accessibility to persuasiveness in the ux evaluation a case study on an insurance prospecting tablet application\.External Links:1806\.11291,[Link](https://arxiv.org/abs/1806.11291)Cited by:[item 3](https://arxiv.org/html/2606.13192#S3.I1.i3.p1.1)\.
- \[9\]Z\. Chen, J\. Wu, W\. Wang, W\. Su, G\. Chen, S\. Xing, M\. Zhong, Q\. Zhang, X\. Zhu, L\. Lu,et al\.\(2024\)Internvl: scaling up vision foundation models and aligning for generic visual\-linguistic tasks\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 24185–24198\.Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[10\]C\. Cui, T\. Li, J\. Wang, C\. Chen, D\. Towey, and R\. Huang\(2024\)Large language models for mobile gui text input generation: an empirical study\.arXiv preprint arXiv:2404\.08948\.Cited by:[§2\.1](https://arxiv.org/html/2606.13192#S2.SS1.p1.1)\.
- \[11\]B\. Deka, Z\. Huang, C\. Franzen, J\. Hibschman, D\. Afergan, Y\. Li, J\. Nichols, and R\. Kumar\(2017\)Rico: a mobile app dataset for building data\-driven design applications\.InProceedings of the 30th annual ACM symposium on user interface software and technology,pp\. 845–854\.Cited by:[§2\.1](https://arxiv.org/html/2606.13192#S2.SS1.p1.1)\.
- \[12\]S\. Deng, W\. Xu, H\. Sun, W\. Liu, T\. Tan, L\. Liujianfeng, A\. Li, J\. Luan, B\. Wang, R\. Yan,et al\.\(2024\)Mobile\-bench: an evaluation benchmark for llm\-based mobile agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8813–8831\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p4.1)\.
- \[13\]P\. Duan, J\. Warner, and B\. Hartmann\(2023\)Towards generating ui design feedback with llms\.InAdjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,pp\. 1–3\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[14\]Z\. Lian, H\. Chen, L\. Chen, H\. Sun, L\. Sun, Y\. Ren, Z\. Cheng, B\. Liu, R\. Liu, X\. Peng,et al\.\(2025\)Affectgpt: a new dataset, model, and benchmark for emotion understanding with multimodal large language models\.arXiv preprint arXiv:2501\.16566\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[15\]K\. Q\. Lin, L\. Li, D\. Gao, Z\. Yang, S\. Wu, Z\. Bai, S\. W\. Lei, L\. Wang, and M\. Z\. Shou\(2025\)Showui: one vision\-language\-action model for gui visual agent\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 19498–19508\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[16\]H\. Liu, C\. Li, Y\. Li, and Y\. J\. Lee\(2024\)Improved baselines with visual instruction tuning\.External Links:2310\.03744,[Link](https://arxiv.org/abs/2310.03744)Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[17\]H\. Liu, C\. Li, Y\. Li, B\. Li, Y\. Zhang, S\. Shen, and Y\. J\. Lee\(2024\-01\)LLaVA\-next: improved reasoning, ocr, and world knowledge\.External Links:[Link](https://llava-vl.github.io/blog/2024-01-30-llava-next/)Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.5.3.4.1.1)\.
- \[18\]J\. Liu, T\. Ou, Y\. Song, Y\. Qu, W\. Lam, C\. Xiong, W\. Chen, G\. Neubig, and X\. Yue\(2024\)Harnessing webpage uis for text\-rich visual understanding\.arXiv preprint arXiv:2410\.13824\.Cited by:[§4\.1](https://arxiv.org/html/2606.13192#S4.SS1.p6.1)\.
- \[19\]J\. Liu, Y\. Song, B\. Y\. Lin, W\. Lam, G\. Neubig, Y\. Li, and X\. Yue\(2024\)VisualWebBench: how far have multimodal llms evolved in web page understanding and grounding?\.External Links:2404\.05955,[Link](https://arxiv.org/abs/2404.05955)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p4.1),[§2\.1](https://arxiv.org/html/2606.13192#S2.SS1.p1.1)\.
- \[20\]Z\. Liu, C\. Chen, J\. Wang, Y\. Huang, J\. Hu, and Q\. Wang\(2022\)Nighthawk: fully automated localizing ui display issues via visual understanding\.IEEE Transactions on Software Engineering49\(1\),pp\. 403–418\.Cited by:[§2\.2](https://arxiv.org/html/2606.13192#S2.SS2.p1.1)\.
- \[21\]Q\. Lu, W\. Shao, Z\. Liu, L\. Du, F\. Meng, B\. Li, B\. Chen, S\. Huang, K\. Zhang, and P\. Luo\(2025\)GUIOdyssey: a comprehensive dataset for cross\-app gui navigation on mobile devices\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 22404–22414\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[22\]Y\. Lu, Z\. Tong, Q\. Zhao, C\. Zhang, and T\. J\. Li\(2023\)UI layout generation with llms guided by ui grammar\.External Links:2310\.15455,[Link](https://arxiv.org/abs/2310.15455)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[23\]P\. Parthasarathi, M\. Reymond, B\. Chen, Y\. Cui, and S\. Chandar\(2025\)GRPO\-λ\\lambda: credit assignment improves llm reasoning\.External Links:2510\.00194,[Link](https://arxiv.org/abs/2510.00194)Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[24\]Qwen\(2025\)Qwen3\-vl: sharper vision, deeper thought, broader action\.Note:[https://qwen\.ai/blog?id=99f0335c4ad9ff6153e517418d48535ab6d8afef&from=research\.latest\-advancements\-list](https://qwen.ai/blog?id=99f0335c4ad9ff6153e517418d48535ab6d8afef&from=research.latest-advancements-list)Accessed: 2025\-09\-23Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.11.3.1.1.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.16.8.1.1.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.17.9.1.1.1)\.
- \[25\]D\. Ran, H\. Wang, Z\. Song, M\. Wu, Y\. Cao, Y\. Zhang, W\. Yang, and T\. Xie\(2024\)Guardian: a runtime framework for llm\-based ui exploration\.InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis,pp\. 958–970\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[26\]H\. Shao, S\. Qian, H\. Xiao, G\. Song, Z\. Zong, L\. Wang, Y\. Liu, and H\. Li\(2024\)Visual cot: unleashing chain\-of\-thought reasoning in multi\-modal language models\.CoRR\.Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[27\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo\(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p7.1)\.
- \[28\]Y\. Su, C\. Chen, J\. Wang, Z\. Liu, D\. Wang, S\. Li, and Q\. Wang\(2022\)The metamorphosis: automatic detection of scaling issues for mobile apps\.InProceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering,pp\. 1–12\.Cited by:[§2\.2](https://arxiv.org/html/2606.13192#S2.SS2.p1.1)\.
- \[29\]Y\. Su, Z\. Liu, C\. Chen, J\. Wang, and Q\. Wang\(2021\)OwlEyes\-online: a fully automated platform for detecting and localizing ui display issues\.InProceedings of the 29th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering,pp\. 1500–1504\.Cited by:[§2\.2](https://arxiv.org/html/2606.13192#S2.SS2.p1.1)\.
- \[30\]Y\. Sun and T\. Zhou\(2025\)DialogueMLLM: transforming multimodal emotion recognition in conversation through instruction\-tuned mllm\.IEEE Access\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[31\]V\. Team, W\. Hong, W\. Yu, X\. Gu, G\. Wang, G\. Gan, H\. Tang, J\. Cheng, J\. Qi, J\. Ji, L\. Pan, S\. Duan, W\. Wang, Y\. Wang, Y\. Cheng, Z\. He, Z\. Su, Z\. Yang, Z\. Pan, A\. Zeng, B\. Wang, B\. Chen, B\. Shi, C\. Pang, C\. Zhang, D\. Yin, F\. Yang, G\. Chen, J\. Xu, J\. Zhu, J\. Chen, J\. Chen, J\. Chen, J\. Lin, J\. Wang, J\. Chen, L\. Lei, L\. Gong, L\. Pan, M\. Liu, M\. Xu, M\. Zhang, Q\. Zheng, S\. Yang, S\. Zhong, S\. Huang, S\. Zhao, S\. Xue, S\. Tu, S\. Meng, T\. Zhang, T\. Luo, T\. Hao, T\. Tong, W\. Li, W\. Jia, X\. Liu, X\. Zhang, X\. Lyu, X\. Fan, X\. Huang, Y\. Wang, Y\. Xue, Y\. Wang, Y\. Wang, Y\. An, Y\. Du, Y\. Shi, Y\. Huang, Y\. Niu, Y\. Wang, Y\. Yue, Y\. Li, Y\. Zhang, Y\. Wang, Y\. Wang, Y\. Zhang, Z\. Xue, Z\. Hou, Z\. Du, Z\. Wang, P\. Zhang, D\. Liu, B\. Xu, J\. Li, M\. Huang, Y\. Dong, and J\. Tang\(2025\)GLM\-4\.5v and glm\-4\.1v\-thinking: towards versatile multimodal reasoning with scalable reinforcement learning\.External Links:2507\.01006,[Link](https://arxiv.org/abs/2507.01006)Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.9.7.4.1.1)\.
- \[32\]B\. Wang, G\. Li, X\. Zhou, Z\. Chen, T\. Grossman, and Y\. Li\(2021\)Screen2Words: automatic mobile ui summarization with multimodal learning\.External Links:2108\.03353,[Link](https://arxiv.org/abs/2108.03353)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p4.1),[§2\.1](https://arxiv.org/html/2606.13192#S2.SS1.p1.1)\.
- \[33\]P\. Wang, S\. Bai, S\. Tan, S\. Wang, Z\. Fan, J\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, Y\. Fan, K\. Dang, M\. Du, X\. Ren, R\. Men, D\. Liu, C\. Zhou, J\. Zhou, and J\. Lin\(2024\)Qwen2\-vl: enhancing vision\-language model’s perception of the world at any resolution\.arXiv preprint arXiv:2409\.12191\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p7.1)\.
- \[34\]W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.\(2025\)InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.12.4.1.1.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.13.5.1.1.1),[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.8.2.1.1)\.
- \[35\]P\. Weichbroth\(2025\)Factors influencing the perceived usability of mobile applications\.External Links:2502\.11069,[Link](https://arxiv.org/abs/2502.11069)Cited by:[item 1](https://arxiv.org/html/2606.13192#S3.I1.i1.p1.1),[item 2](https://arxiv.org/html/2606.13192#S3.I1.i2.p1.1)\.
- \[36\]C\. Wu, Q\. Cao, C\. Li, Z\. Wang, C\. Xue, Y\. Fan, W\. Xi, and X\. He\(2025\)Beyond token length: step pruner for efficient and accurate reasoning in large language models\.External Links:2510\.03805,[Link](https://arxiv.org/abs/2510.03805)Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[37\]L\. Xiaomi\(2025\)MiMo\-vl technical report\.External Links:2506\.03569,[Link](https://arxiv.org/abs/2506.03569)Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.18.10.1.1.1)\.
- \[38\]B\. Yang and S\. Li\(2024\)Uisgpt: automated mobile ui design smell detection with large language models\.Electronics13\(16\),pp\. 3127\.Cited by:[§2\.2](https://arxiv.org/html/2606.13192#S2.SS2.p1.1)\.
- \[39\]H\. Yang, W\. Qiu, R\. Zhang, Z\. Fang, R\. Mao, X\. Lin, M\. Huang, Z\. Huang, T\. Guo, S\. Liu, and H\. Rao\(2025\)UI\-ug: a unified mllm for ui understanding and generation\.External Links:2509\.24361,[Link](https://arxiv.org/abs/2509.24361)Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1),[§1](https://arxiv.org/html/2606.13192#S1.p5.1)\.
- \[40\]Y\. Yao, T\. Yu, A\. Zhang, C\. Wang, J\. Cui, H\. Zhu, T\. Cai, H\. Li, W\. Zhao, Z\. He,et al\.\(2024\)MiniCPM\-v: a gpt\-4v level mllm on your phone\.arXiv preprint arXiv:2408\.01800\.Cited by:[Table 1](https://arxiv.org/html/2606.13192#S4.T1.10.14.6.1.1.1)\.
- \[41\]K\. You, H\. Zhang, E\. Schoop, F\. Weers, A\. Swearngin, J\. Nichols, Y\. Yang, and Z\. Gan\(2024\)Ferret\-ui: grounded mobile ui understanding with multimodal llms\.InEuropean Conference on Computer Vision,pp\. 240–255\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[42\]J\. Zhao, B\. Wang, G\. Tu, Y\. Zhang, Q\. Wang, B\. Liang, J\. Li, and R\. Xu\(2025\)CoreEval: automatically building contamination\-resilient datasets with real\-world knowledge toward reliable llm evaluation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 22284–22306\.External Links:[Link](http://dx.doi.org/10.18653/v1/2025.acl-long.1085),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1085)Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[43\]S\. Zhao, M\. Hong, Y\. Liu, D\. Hazarika, and K\. Lin\(2025\)Do llms recognize your preferences? evaluating personalized preference following in llms\.arXiv preprint arXiv:2502\.09597\.Cited by:[§1](https://arxiv.org/html/2606.13192#S1.p2.1)\.
- \[44\]D\. Zhu, J\. Chen, X\. Shen, X\. Li, and M\. Elhoseiny\(2023\)MiniGPT\-4: enhancing vision\-language understanding with advanced large language models\.External Links:2304\.10592,[Link](https://arxiv.org/abs/2304.10592)Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.
- \[45\]J\. Zhu, W\. Wang, Z\. Chen, Z\. Liu, S\. Ye, L\. Gu, H\. Tian, Y\. Duan, W\. Su, J\. Shao,et al\.\(2025\)Internvl3: exploring advanced training and test\-time recipes for open\-source multimodal models\.arXiv preprint arXiv:2504\.10479\.Cited by:[§2\.3](https://arxiv.org/html/2606.13192#S2.SS3.p1.1)\.

Similar Articles

What We are Missing in Multimodal LLM Evaluation?

arXiv cs.AI

This paper reviews current multimodal LLM evaluation benchmarks and identifies key gaps such as temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention, arguing that existing isolated-task benchmarks fail to measure true cross-modal integration.

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

arXiv cs.AI

MobileForge is a benchmark for project-level multi-screen mobile app generation, evaluating multimodal LLMs on build success, cross-page navigation, visual fidelity, maintainability, and efficiency. Experiments on six frontier multimodal LLMs show current models can compile and reach pages but still struggle with interactive navigation and visual quality.