STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

arXiv cs.CL 论文

摘要

This paper introduces STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets for psycholinguistic plausibility judgments. Experiments show that adding a global reasoning scratchpad and evaluator-guided refinement substantially improves generation quality, though near-boundary events remain challenging.

arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.
查看原文
查看缓存全文

缓存时间: 2026/08/06 07:49

# Probing Reasoning Limits in Graded Plausibility Generation and Evaluation
Source: [https://arxiv.org/html/2608.04567](https://arxiv.org/html/2608.04567)
Bhiman Kumar Baghel1Anna Chrabaszcz1Tessa Warren1 Michael Walsh Dickey1Haley C\. Dresang2Xiang Lorraine Li1 1University of Pittsburgh, Pittsburgh, PA, USA 2University of Wisconsin–Madison, Madison, WI, USA Correspondence:[bkb45@pitt\.edu](https://arxiv.org/html/2608.04567v1/mailto:[email protected]),[xiangli@pitt\.edu](https://arxiv.org/html/2608.04567v1/mailto:[email protected])

###### Abstract

Event knowledge concerns who does what to whom\. Psycholinguists use event\-plausibility judgments to examine how this knowledge supports human language processing\. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed\. Constructing such sets manually is labor\-intensive\. We therefore introduce STRIVE, an LLM\-based framework for jointly generating and evaluating controlled event sets crossing plausibility class \(plausible vs\. implausible\) with intended classification difficulty \(easy vs\. hard\)\. Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed\. In experiments with six models across 60 verbs, GPT\-5\.1 produced high\-quality sets only 16\.7% of the time using the baseline generation prompt\. Adding a global reasoning scratchpad and evaluator\-guided refinement raised this rate to 75\.0%\. Greater reasoning effort also improved evaluator–human agreement\. Nevertheless, events near the plausibility boundary remain most difficult\. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible\-hard condition, indicating a need for human input\. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event\-set generation and evaluation for psycholinguistic studies\.111Under Review

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

Bhiman Kumar Baghel1Anna Chrabaszcz1Tessa Warren1Michael Walsh Dickey1Haley C\. Dresang2Xiang Lorraine Li11University of Pittsburgh, Pittsburgh, PA, USA2University of Wisconsin–Madison, Madison, WI, USACorrespondence:[bkb45@pitt\.edu](https://arxiv.org/html/2608.04567v1/mailto:[email protected]),[xiangli@pitt\.edu](https://arxiv.org/html/2608.04567v1/mailto:[email protected])

![Refer to caption](https://arxiv.org/html/2608.04567v1/x1.png)Figure 1:STRIVE’s four plausibility conditions for the root verbcatch\. The verb form, patient, instrument, and location remain fixed while the agent varies\. Easy and Hard indicate intended distance from the plausibility boundary; Clearly and Somewhat are annotator\-facing labels\. The gradient shows intended ordering, not calibrated probabilities or fixed semantic boundaries\.## 1Introduction

Understanding and producing language relies on knowledge about real\-world events, i\.e\., expectations about who does what to whom, with what, and where\(Elman and McRae,[2019](https://arxiv.org/html/2608.04567#bib.bib11)\)\. Psycholinguistic studies probe event knowledge in humans by asking them to judge the plausibility of events described in sentences or depicted in images\(Ivanova et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib16); Dresang et al\.,[2019](https://arxiv.org/html/2608.04567#bib.bib9)\)\. Here, event plausibility refers to the degree to which a complete described situation accords with event knowledge and ordinary world knowledge\(Wang et al\.,[2018](https://arxiv.org/html/2608.04567#bib.bib37); Porada et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib30)\)\. To attribute differences in these judgments to plausibility, such studies require matched verb in which only part of the event properties are modified\. These sentences are refereed as stimuli\. To capture finer distinctions in plausibility judgments, we construct gradient sentence stimuli using a two\-by\-two design crossing plausibility \(plausible vs\. implausible\) with intended classification difficulty \(easy vs\. hard\) \(Figure[1](https://arxiv.org/html/2608.04567#S0.F1)\)\.222These conditions are operational stimulus\-design targets rather than natural semantic categories or calibrated probability intervals\.

To generate stimuli spanning four plausibility conditions for a single verb, we swap values of one of the event frame’s slots\.333We vary the agent in the main experiments and evaluate patient variation in §[6](https://arxiv.org/html/2608.04567#S6)\.This design localizes plausibility differences within a set to that slot and can support psycholinguistic research on lexical prediction and event integration\(McRae et al\.,[2005](https://arxiv.org/html/2608.04567#bib.bib25); Khalkhali et al\.,[2012](https://arxiv.org/html/2608.04567#bib.bib20)\), sentence comprehension\(Bicknell et al\.,[2010](https://arxiv.org/html/2608.04567#bib.bib4); Warren et al\.,[2015](https://arxiv.org/html/2608.04567#bib.bib38)\), and language and cognitive impairments\(Dresang et al\.,[2019](https://arxiv.org/html/2608.04567#bib.bib9)\)\. Figure[1](https://arxiv.org/html/2608.04567#S0.F1)shows the root verbcatchexpanded into a shared soccer scene, with a different agent selected for each condition\. Constructing such controlled, multilevel sets manually is labor\-intensive, motivating NLP\-based generation\.

To achieve this goal, we developedSTRIVE, a framework for generating and evaluating controlled event\-plausibility stimulus sets\. In a valid set, each sentence must match its intended plausibility condition, while all non\-target slots remain fixed across sentences\. STRIVE evaluates both requirements and uses feedback to revise sets that fail either\.

To determine whether generation benefits from considering multiple agents, evaluator\-guided refinement, or their combination, we compare four generation strategies\. Reasoning Base \(RB\) generates a stimulus set in one structured pass\. Reason\-to\-Verbalize \(R2V\) uses verbalized sampling\(Zhang et al\.,[2026](https://arxiv.org/html/2608.04567#bib.bib42)\)to generate several fillers for the target slot before selecting one for each condition\. Reason\-to\-Refine \(R2R\) revises an RB output using feedback from a separate evaluator\(Madaan et al\.,[2023](https://arxiv.org/html/2608.04567#bib.bib23); Wang et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib36)\), while R2V\-initialized R2R \(R2VR\) applies the same refinement to an R2V output\. All four strategies use an expert\-designed prompt that requires a global reasoning scratchpad and per\-stimulus rationales\(Wei et al\.,[2022](https://arxiv.org/html/2608.04567#bib.bib39); Nye et al\.,[2022](https://arxiv.org/html/2608.04567#bib.bib27)\)\. For evaluation, we vary model family and reasoning effort and use the Alternative Annotator Test \(AAT\)\(Calderon et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib5)\)to assess whether model\-human agreement is non\-inferior to human\-human agreement\.

Our experiments yield three findings\. \(F1\) Generation: GPT\-5\.1 produces 28\.3% high\-quality stimulus sets under RB and 75\.0% under R2R, a 2\.6×\\timesimprovement\. Removing the global reasoning scratchpad reduces RB performance to 16\.7% and R2R performance to 40\.0%, while the full methods require 2\.7–3\.5×\\timesmore output tokens \(Appendix[J](https://arxiv.org/html/2608.04567#A10)\)\. \(F2\) Evaluation: Across six evaluator models, reasoning generally improves model\-human agreement, but only GPT\-5\.1 and Sonnet 4\.6 pass the primary AAT criterion, with bootstrap 5th percentiles above the−0\.05\-0\.05threshold\. This indicates that reasoning effort alone is insufficient and that underlying model capabilities also matter\. \(F3\) Boundary difficulty: cases nearest the plausibility boundary elicit the greatest human disagreement and remain the hardest for LLM evaluators, with the best evaluator reaching only 57% accuracy on the implausible\-hard condition \(§[5\.2](https://arxiv.org/html/2608.04567#S5.SS2)\)\.

To our knowledge, STRIVE is the first framework to jointly generate and evaluate matched, slot\-controlled stimulus sets across four event\-plausibility conditions\. More broadly, STRIVE’s root\-verb\-to\-frame design provides a scalable approach to controlled stimulus construction across psycholinguistic studies of event knowledge\.444The present experiments are restricted to English; cross\-linguistic extension requires language\-specific validation\.

## 2Related Work

This section focuses on work most directly related to plausibility event generation and evaluation\. Appendix[A](https://arxiv.org/html/2608.04567#A1)discusses broader psycholinguistic and plausibility research\.

##### Plausibility generation\.

ADEPT\(Emami et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib12)\)pairs a sentence with a version formed by adding an adjective to a noun, then assigns the pair a five\-way label indicating how the adjective affects plausibility\. These labels encode changes within sentence pairs, not five matched conditions in a shared event frame\.Eichel and Schulte im Walde \([2023](https://arxiv.org/html/2608.04567#bib.bib10)\)create pseudo\-implausible events by replacing two constituents of corpus\-derived event triples and collect absolute ratings\. PRobELM\(Yuan et al\.,[2024](https://arxiv.org/html/2608.04567#bib.bib41)\)ranks alternatives constructed from Wikidata, whileTang et al\. \([2023](https://arxiv.org/html/2608.04567#bib.bib35)\)target one generation band of relevant but less\-likely hypotheses\. Unlike these methods, STRIVE generates all slot values from a root verb and constructs a matched four\-condition set while holding non\-target slots fixed and varying only the target role\.

##### Plausibility evaluation and refinement\.

Existing methods estimate the plausibility of individual items using sentence probabilities\(Kauf et al\.,[2024](https://arxiv.org/html/2608.04567#bib.bib18)\), direct LM judgments\(Amouyal et al\.,[2024](https://arxiv.org/html/2608.04567#bib.bib2)\), or a dedicated estimator such as VERA\(Liu et al\.,[2023](https://arxiv.org/html/2608.04567#bib.bib22)\)\. Self\-Refine\(Madaan et al\.,[2023](https://arxiv.org/html/2608.04567#bib.bib23)\)improves generated outputs using self\-feedback, while Cross\-Refine\(Wang et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib36)\)uses a separate critic\. STRIVE connects these lines by using feedback on both condition assignment and set\-level role control to revise matched four\-condition sets, while validating the evaluator against human judgments\.

## 3STRIVE Framework

STRIVE has two parts: ageneratorthat produces graded plausibility stimuli from a root verb, and anevaluatorthat serves two purposes: quality assessment of generated stimuli, and structured feedback that drives iterative refinement \(§[3\.1\.3](https://arxiv.org/html/2608.04567#S3.SS1.SSS3)\), as illustrated in Figure[2](https://arxiv.org/html/2608.04567#S3.F2)\. Each is a standalone structured prompt555The prompts usethematic fitas a heuristic for reasoning about agent substitutions within a fixed event frame\. This heuristic supports candidate construction, but it is not the evaluated construct: all reported labels and human judgments assess complete\-event plausibility\.designed through iterative human\-AI collaboration between psycholinguists and AI researchers, refined based on failure patterns observed on three random verbs \(catch,break,erase\)\. Although we instantiate the variable slot as theagentthroughout this paper, the framework generalises to other slots; we discuss patient\-varying generation as a natural extension in §[6](https://arxiv.org/html/2608.04567#S6)\.

![Refer to caption](https://arxiv.org/html/2608.04567v1/x2.png)Figure 2:STRIVE framework\. Given a root verb, the generator produces a sentence set𝒮=\{sP​E,sP​H,sI​H,sI​E\}\\mathcal\{S\}=\\\{s\_\{PE\},s\_\{PH\},s\_\{IH\},s\_\{IE\}\\\}varying only the agent across the four plausibility conditions\. Four generator variants share the pipeline \(scene design, agent selection, visual distinctiveness checking, composition\):RBis the single\-shot reasoning base;R2Vadds verbalized plausibility scores over disjoint ranges per condition and samples the best agent;R2Radds iterative refinement that routes agent\-level and scene\-level failures separately;R2VRinitialises R2R with R2V\. The evaluator scores plausibility \(adversarial check, IH overlap audit\), visual evaluation, and scene assessment\.### 3\.1Generator

Given a root verbvv, the generator produces a set𝒮=\{sPE,sPH,sIH,sIE\}\\mathcal\{S\}=\\\{s\_\{\\text\{PE\}\},\\,s\_\{\\text\{PH\}\},\\,s\_\{\\text\{IH\}\},\\,s\_\{\\text\{IE\}\}\\\}of four sentences that share a fixed five\-slot template:

s=⟨Agent,v′,Patient,Instrument,Location⟩s=\\langle\\,\\textit\{Agent\},\\ v^\{\\prime\},\\ \\textit\{Patient\},\\ \\textit\{Instrument\},\\ \\textit\{Location\}\\,\\rangle\(1\)wherev′v^\{\\prime\}is a chosen inflected form ofvv\. Only theAgentslot varies across the four sentences; the other slots are held constant\. This single\-variable design localizes within\-set plausibility differences to the agent substitution, eliminating lexical and syntactic confounds that would otherwise complicate experimental interpretation\. Figure[1](https://arxiv.org/html/2608.04567#S0.F1)defines the four conditions\.

#### 3\.1\.1RB \(Reasoning Base\)

RB instantiates the generator as a single forward pass through a structured prompt that decomposes stimulus generation into five sequential subtasks, each depending on the output of the previous\. We adopted this decomposition after preliminary experiments showed that directly prompting for the full stimulus set repeatedly produced outputs violating multiple task requirements simultaneously, with no clear leverage for targeted correction\.

The sub\-tasks are: \(1\)Scene design: fix the static part of the stimulus \(thescene\) by choosing the verb form, patient, instrument, and location together, conditioned on whether the resulting frame can support the full plausibility gradient; forcatchin Figure[1](https://arxiv.org/html/2608.04567#S0.F1), the verb form becomesis catching, yielding the scene“\[Agent\] is catching a soccer ball with their hands on the soccer field\.”\(2\)Agent selection: select one agent per plausibility condition, conditioned on the scene; in the running example,a goalie,a referee,a hockey player, anda ballerinaare picked for PE, PH, IH, and IE respectively\. \(3\)Visual Distinctiveness: Some psycholinguistic event\-knowledge studies use image\-based stimuli\(Ivanova et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib16); Dresang et al\.,[2019](https://arxiv.org/html/2608.04567#bib.bib9)\), so agents intended for downstream image generation should be visually distinguishable\. STRIVE therefore applies a text\-based pre\-filter based on visible cues such as occupational attire \(Figure[1](https://arxiv.org/html/2608.04567#S0.F1)\)\. This heuristic does not validate actual images; image generation and image\-grounded validation are outside the scope of this work\. \(4\)Composition: insert each selected agent into the shared scene frame to compose the four sentences\. \(5\)Validation: answer a 16\-item YES/NO checklist covering every constraint from the previous steps; anynoanswer must be rectified before output \(Figure[9](https://arxiv.org/html/2608.04567#A2.F9)\)\.

The output begins with a free\-form scratchpad where the model reasons through these sub\-tasks before committing to the structured outputs that follow\. The structured output contains the scene, the four agents \(each with a per\-stimulus rationale justifying its gradient placement\), and the four composed sentences\. The full prompt and output schema are in Appendix[B](https://arxiv.org/html/2608.04567#A2)\.

#### 3\.1\.2R2V \(Reason\-to\-Verbalize\)

R2V replaces RB’s direct agent selection \(Step 2\) with verbalized samplingZhang et al\. \([2026](https://arxiv.org/html/2608.04567#bib.bib42)\), because we hypothesize that explicit per\-condition score ranges keep candidates within their intended plausibility grade, preventing the cross\-grade contamination that arises when grade boundaries are left implicit\.666We use these verbalized values only as heuristic plausibility scores for organizing candidate generation\. They are neither token\-level model probabilities nor calibrated estimates of real\-world event likelihood or human responses, and we do not evaluate their numerical accuracy\.We partition the\[0,1\]\[0,1\]scoring scale under three principles: i\) clear gradient ordering with no overlap, \(ii\) anchored endpoints at PE and IE that fix the gradient to the extremes of\[0,1\]\[0,1\], and \(iii\) non\-zero buffers between adjacent conditions\. The exact range boundaries are a design choice; any partition satisfying these principles would induce the same gradient structure\. We adopt PE: 0\.80–1\.00, PH: 0\.40–0\.65, IH: 0\.15–0\.35, IE: 0\.00–0\.05, with buffer widths reflecting conceptual proximity between adjacent conditions: PE–PH \(0\.15\) and IH–IE \(0\.10\) separate “clearly” from “somewhat” judgments \(annotator\-facing labels, Figure[1](https://arxiv.org/html/2608.04567#S0.F1)\), while the narrower PH–IH buffer \(0\.05\) sits between two adjacent “somewhat” conditions across the plausibility cut\. The model first generates a pool of candidate agents spanning the full plausibility range, then selects one candidate per plausibility condition from within the corresponding region\. The full prompt and output schema are in Appendix[C](https://arxiv.org/html/2608.04567#A3)\.

#### 3\.1\.3Iterative Refinement \(R2R, R2VR\)

R2R and R2VR augment RB and R2V with iterative refinement using evaluator feedback\. The refinement prompt already contains the same task definitions and validation rubric as RB\. Its feedback field does not repeat this rubric or include the full evaluator output; instead, it contains a compact, instance\-specific diagnosis with per\-condition verdicts, evaluator\-inferred conditions, brief issue descriptions, relevant overlap findings, and gradient and scene verdicts\. Refinement therefore adds error localization rather than new task rules\. These diagnoses drive two refiner behaviours\. Agent\-level failures concern a specific agent: it belongs in a different condition than assigned \(e\.g\., a hockey player placed in PH instead of IH\), or it is not visually identifiable, or both\. The refiner replaces only the failing agents, preserving the scene and the passing agents\. Scene\-level failures concern the entire scene: the scene does not support the full gradient \(e\.g\., a domain so narrow that no agent can plausibly fill IH\)\. The refiner discards the output and selects a new scene\. R2R and R2VR begin with RB and R2V outputs, respectively, then use a common refinement prompt with an instruction to rectify rather than regenerate and a feedback\-history placeholder that accumulates prior compact diagnoses to prevent regression \(Appendix[D](https://arxiv.org/html/2608.04567#A4)\)\.

### 3\.2Evaluator

The evaluator scores a generated stimulus set against task requirements and produces structured feedback\. Given a verb and four candidate stimuli, it returns three verdicts, evaluating correctness of three major generation components: the agents, their visual distinctiveness, and the scene holding them on the gradient\.

\(1\)Per condition adversarial verdict\(per condition\): the evaluator articulates the strongest argument that the agent belongs one level higher and one level lower on the gradient; the verdict is CORRECT only if both counter\-arguments are clearly weaker than the assigned placement, otherwise INCORRECT\. Specifically for IH, the evaluator audits the five overlap dimensions against PE \(occupational domain, setting, tools, patient type, physical actions\) and applies the negative test \(removing the shared verb, does the association persist?\); an IH agent with fewer than two overlap dimensions is reclassified as IE\. We selected this two\-of\-five cutoff by inspecting seed examples during prompt development; it is a tunable operational choice rather than a universal semantic threshold\.

\(2\)Visual distinctiveness verdict: each agent is rated IDENTIFIABLE, AMBIGUOUS, or UNIDENTIFIABLE based on whether its profession is recognisable from a photograph against a plain white background\. Visual distinctness is then checked across all pairs of agents\. The visual verdict is ALL\_DISTINCT \(all agents IDENTIFIABLE and all pairs distinct\), PROBLEMATIC \(any agent UNIDENTIFIABLE, or three or more pairs indistinct\), or MOSTLY\_DISTINCT otherwise\.

\(3\)Scene assessment: after per\-condition evaluation, the evaluator assesses the scene on instrument correctness, prototype clarity, imageability, and gradient supportability; if a condition is INCORRECT but a viable replacement agent exists, the verdict is AGENT\_FIXABLE, otherwise NEEDS\_REDESIGN\. The full prompt and output schema are in Appendix[E](https://arxiv.org/html/2608.04567#A5)\.

### 3\.3Why Graded Plausibility Is Hard

Three challenges compound across the framework\. The IH constraint sits at a semantic knife’s edge: too much overlap with PE pushes the agent into PH; too little drops it to IE\. Semantic distinctness does not imply visual distinctness: two professions that read as different on paper may still look the same in a photograph, making them unusable as paired stimuli\. And all four conditions must be jointly valid within one shared scene: a domain that is too narrow collapses IH because every adjacent professional becomes plausible, invalidating the entire set\. The evaluator must detect all three failure types and distinguish between those requiring agent replacement and those requiring full scene redesign\.

## 4Experiments

### 4\.1Generator

We run RB, R2V, R2R, and R2VR across six models on 60 picturable verbs drawn from three English verb\-naming assessments\(Cho\-Reyes and Thompson,[2012](https://arxiv.org/html/2608.04567#bib.bib6); Masterson,[2000](https://arxiv.org/html/2608.04567#bib.bib24); Swinburn et al\.,[2004](https://arxiv.org/html/2608.04567#bib.bib34)\)and filtered to retain verbs that can instantiate all four event roles in our template \(agent, patient, instrument, location\)\. The assessments provide only the root verbs; STRIVE generates the verb form and all event\-role values\. The six models span scale and openness: GPT\-5\.1 and Claude Sonnet 4\.6 \(closed\-source\), Qwen3\-Next\-80B, Qwen3\-30B, Ministral\-3\-14B, and Qwen3\-4B \(open\-source\)\. Iterative methods \(R2R, R2VR\) run up tok=3k=3iterations and exit early when the evaluator returns a GOOD scene verdict and all per\-condition verdicts are CORRECT\. We set temperature to 0\.3 to balance two requirements: enough creativity to support diverse candidate selection in the middle of the gradient \(PH and IH, where conceptual distinctions are subtler\), while remaining deterministic enough for reliable results\. For the same reason, native reasoning and thinking modes are disabled: these either constrain temperature to a fixed value, or, for closed\-source models, return summarised traces rather than the verbatim reasoning\. The scratchpad approach in §[3\.1](https://arxiv.org/html/2608.04567#S3.SS1)preserves full reasoning traces at our chosen temperature\. Complete implementation details are in Appendix[F](https://arxiv.org/html/2608.04567#A6)\.

### 4\.2Evaluation Experiments

Evaluator outputs comprise per\-condition, scene, and visual\-distinctiveness verdicts, which do not yield a single comparable label for generation methods\. We therefore map them to GOLD, SILVER, BRONZE, or FAIL using the hierarchical rule in Table[1](https://arxiv.org/html/2608.04567#S4.T1), with plausibility as the prerequisite gate\. Within STRIVE, GOLD denotes sets that pass all three evaluator checks\. For downstream tasks, GOLD serves only as STRIVE’s text\-based filter; image\-grounded and domain\-expert validation remain required\. We compare evaluator models and reasoning settings against human annotations\.

### 4\.3Human Annotation Study

We collected human judgments mimicking the evaluator \(§[3\.2](https://arxiv.org/html/2608.04567#S3.SS2)\) on a subset of 30 stimulus sets \(120 sentences\) sampled across the GOLD, SILVER, and FAIL tiers from our main generation experiments\. Eight undergraduate researchers in psycholinguistics777All annotators participated as volunteers; no compensation was provided\.each rated all 120 sentences for plausibility, yielding 960 sentence\-level ratings, and assessed visual distinctiveness for each four\-agent set\. Both are categorical tasks, so we adopt unweighted Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2608.04567#bib.bib7)\)as the primary metric for the AAT\(Calderon et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib5)\); this matches the tier\-assignment success criterion \(Table[1](https://arxiv.org/html/2608.04567#S4.T1)\), which requires exact category match\. Visual distinctiveness is decomposed into two sub\-tasks\. AgentID asks whether each agent’s profession is identifiable from a photograph; the primary metric is 3\-point unweightedκ\\kappa\. PairDist asks which agent pairs would look indistinguishable; the primary metric is binaryκ\\kappa, with Gwet’s AC1\(Gwet,[2008](https://arxiv.org/html/2608.04567#bib.bib14)\)reported to address prevalence skew\. Full survey details in Appendix[M](https://arxiv.org/html/2608.04567#A13)\.

We use two evaluators to support both AAT\-validated primary results and a fully open\-source pipeline\. Two configurations pass non\-inferiority on the unweighted Cohen’sκ\\kappaAAT against human annotators: GPT\-5\.1 \(high reasoning effort\) and Claude Sonnet 4\.6 \(non\-thinking mode\)\. We use GPT\-5\.1 as the primary evaluator for its lower per\-token cost, applying its verdicts to all primary generation results in §[5](https://arxiv.org/html/2608.04567#S5)\. To validate that our generation findings do not depend on closed\-source evaluation, we additionally evaluate every generator output with Qwen3\-30B, the open\-source evaluator whose unweighted Cohen’sκ\\kappawith human annotators comes closest to the human\-humanκ\\kappa\(Appendix[G](https://arxiv.org/html/2608.04567#A7)\)\. Although this margin does not pass the AAT non\-inferiority threshold, Qwen3\-30B achieves a GOLD\-tier concordance of 81\.5% with GPT\-5\.1 \(Wilson 95% CI excludes 50%; details in Appendix[K](https://arxiv.org/html/2608.04567#A11)\)\. This cross\-evaluator agreement at the gold\-stimulus decision level means the open\-source pipeline recovers the same downstream design conclusions as the closed\-source pipeline\.

Table 1:Hierarchical quality\-tier assignment\. Plausibility is the first gate; scene and visual verdicts distinguish tiers only after all four conditions are correct\.Table 2:GOLD \(%\) per \(model, method\) cell, reported as*GPT judge / Qwen judge*\. “—” indicates the GPT\-judge evaluation is unavailable due to budget constraints\.\*GPT\-judge aggregate covers only the two closed\-source generators \(n=115n\{=\}115for R2R,n=120n\{=\}120for R2VR\)DimensionMetricH\-HL\-HΔ\\DeltaDecisionPlausibility \(sentence\-level,n=120n\{=\}120\)4\-pt \(primary\)κ\\kappa0\.5290\.530\+\+0\.001PASSδ≤0\.05\\delta\{\\leq\}0\.054\-pt \(support\)κw\\kappa\_\{w\}0\.8060\.821\+\+0\.015PASSδ≤0\.01\\delta\{\\leq\}0\.01Visual distinctivenessAgentID \(n=120n\{=\}120\)κ\\kappa0\.2770\.210−\-0\.067PASSδ≤0\.15\\delta\{\\leq\}0\.15PairDist \(n=180n\{=\}180\)κ\\kappa0\.4070\.113−\-0\.293FAILPairDist \(n=180n\{=\}180\)AC10\.7470\.774\+\+0\.026PASSδ≤0\.01\\delta\{\\leq\}0\.01

Table 3:AAT results\. H\-H = mean H\-H pairwise agreement \(28 pairs\); L\-H = mean LLM–human agreement \(8 pairs\);Δ=κ¯L\-H−κ¯H\-H\\Delta=\\bar\{\\kappa\}\_\{\\text\{L\-H\}\}\-\\bar\{\\kappa\}\_\{\\text\{H\-H\}\}; Decision = smallestδ\\deltawhere bootstrap 5th percentile ofΔ\\Deltaexceeds−δ\-\\delta\(B=10,000B\{=\}10\{,\}000\)\.![Refer to caption](https://arxiv.org/html/2608.04567v1/x3.png)Figure 3:Per\-condition distribution of human plausibility ratings\. Diamonds mark median ratings; the dashed line traces the perfect gradient\.

## 5Results

### 5\.1Generation Quality

Table[2](https://arxiv.org/html/2608.04567#S4.T2)report GOLD rate of different models/methods judged by both GPT and Qwen and Table[15](https://arxiv.org/html/2608.04567#A14.T15)some samples\.Iteration is the central lever \(F1\)\.Both judges agree that iterative refinement substantially improves GOLD rate over the standalone methods\. Under the GPT judge, RB and R2V reach 11–13% GOLD while R2R and R2VR reach 74–77% on the two closed\-source generators\. Under the Qwen judge, standalone methods achieve 36–37% GOLD across all six generators while iterative methods reach 65–68%\. The qualitative ranking RB≈\\approxR2V≪\\llR2R≈\\approxR2VR holds across both evaluators\.

Refinement improves the plausibility gate\.We measure the fraction of sets with all four condition verdicts correct before applying scene and visual checks\. Under the GPT judge, this rate rises from 51\.7% under RB to 86\.7% under R2R for GPT\-5\.1, and from 43\.3% to 89\.1% for Sonnet 4\.6\. Thus, refinement gains cannot be attributed only to the later checks\. Appendix[L](https://arxiv.org/html/2608.04567#A12)reports all methods\.

Judge differences and caveats\.The two judges disagree on the magnitude of iteration’s effect for the two generators where a clean cross\-judge comparison exists\. The Qwen judge is more generous than the GPT judge on standalone outputs from closed\-source generators \(51\.7% vs 28\.3% for RB\) but less generous on their iterative outputs \(48\.7% vs 74\.8% for R2R\)\. For open\-source generators on R2R/R2VR, GPT\-judge evaluation was not run due to budget constraints\. The iteration feedback in those cases also came from the Qwen judge, so the Qwen\-judge column for those cells is a self\-evaluation rather than an independent check; we therefore do not directly compare these GOLD rates against the closed\-source rates\. The Qwen judge serves as cross\-evaluator trend confirmation on the closed\-source pipeline, not as a substitute for the GPT judge\.

### 5\.2Evaluator Validation

Humans distinguish the intended conditions, but IH remains most variable\.Before assessing evaluator–human agreement, we first ask whether human ratings differ overall across the four intended conditions\. As expected, ratings become progressively less plausible from PE to IE, as shown by the observed means \(PE 1\.21, PH 1\.89, IH 2\.58, IE 3\.86; Figure[3](https://arxiv.org/html/2608.04567#S4.F3)\)\. A Friedman test\(Friedman,[1937](https://arxiv.org/html/2608.04567#bib.bib13)\)confirms that this overall difference is statistically reliable \(χ2​\(3\)=76\.5\\chi^\{2\}\(3\)=76\.5,p<10−15p<10^\{\-15\}\), showing that humans do not treat all four conditions alike\. Because this overall effect could be driven only by the extreme conditions, we next ask whether humans distinguish each adjacent pair\. Holm\-corrected Wilcoxon signed\-rank tests\(Wilcoxon,[1945](https://arxiv.org/html/2608.04567#bib.bib40); Holm,[1979](https://arxiv.org/html/2608.04567#bib.bib15)\)find reliable differences at every adjacent boundary \(all adjustedp<0\.001p<0\.001\), including PH–IH\. Thus, with 960 sentence\-level ratings across 30 matched sets, the four conditions are both ordered and separable in human judgment, although IH remains the most variable \(SD=1\.03=1\.03\)\. Appendix[M](https://arxiv.org/html/2608.04567#A13)provides the test rationale and complete pairwise results\.

Plausibility classification: LLM and humans agree as much as humans agree with each other\.For the GPT\-5\.1 evaluator with high reasoning effort, L\-Hκ\\kappamatches the H\-H baseline on the primary 4\-point classification metric \(0\.530 vs 0\.529,Δ=\+0\.001\\Delta=\+0\.001; Table[3](https://arxiv.org/html/2608.04567#S4.T3)\)\. This passes the AAT non\-inferiority test atδ≤0\.05\\delta\\leq 0\.05: with 95% confidence, the LLM is statistically no worse than a human annotator by more than 0\.05 kappa\. The supporting weightedκ\\kappametric\(Cohen,[1968](https://arxiv.org/html/2608.04567#bib.bib8)\)passes at the tighterδ≤0\.01\\delta\\leq 0\.01, confirming the LLM and humans agree on the gradient ordering: if humans rate a stimulus as PH, the LLM rates it PH or at worst PE, not IH or IE\. The LLM evaluator can substitute for a human annotator at the primary classification task\. Complete results is Appendix[N](https://arxiv.org/html/2608.04567#A14)\.

Visual distinctiveness validation is unreliable\.Human\-human agreement on visual distinctiveness is itself low \(κ=0\.277\\kappa=0\.277for AgentID,0\.4070\.407for PairDist\), so humans do not form a reliable benchmark\. On AgentID, the LLM tracks the human baseline closely but the baseline itself is weak\. On PairDist,κ\\kappashows a large LLM\-human gap while Gwet’s AC1 \(which accounts for the high prevalence of distinguishable pairs\) passes\. We treat visual distinctiveness as unvalidated: the low human\-human agreement reflects that annotators judged each profession’s appearance from text alone, with each annotator imagining the visual details differently\. Validation requires image\-based annotation, presenting actual stimulus photographs to constrain this variability\.

## 6Verb Difficulty Is Slot\-Dependent

Aggregating across all models and methods in agent\-varying generation, per\-verb GOLD rates span42\.4%42\.4\\%\(throw\) to3\.4%3\.4\\%\(tickle,pay\), Table[11](https://arxiv.org/html/2608.04567#A9.T11)\. The distribution divides into three bands: Easy \(≥20%\\geq 20\\%GOLD,n=11n=11\), Moderate \(10–19%,n=34n=34\), and Hard \(<10%<10\\%,n=15n=15\)\. The Hard band concentrates two kinds of verbs: \(a\) actions tied to no specific occupation \(eat,drink,tickle,pay\), which leave no prototypical PE agent, and \(b\) actions with very narrow occupational scope \(tow,iron,erase\), which leave no room for a distinct IH agent\. Both failure modes are properties of the agent slot, not of the action itself\.

A natural question arises from the per\-verb GOLD distribution: does verb difficulty reflect an inherent property of the verb, or an artefact of fixing the agent slot? We therefore re\-ran the pipeline with the patient slot varied\. The framework is otherwise identical: only the five overlap dimensions and the visual identifiability criterion were recast from role\-centered to object\-centered properties\. Across the same 60 verbs and five models \(Sonnet 4\.6 and the four open\-source models\), patient\-varying R2V reaches16\.4%16\.4\\%GOLD versus6\.7%6\.7\\%for RB \(Table[5](https://arxiv.org/html/2608.04567#S7.T5)\)\. Crucially, only3%3\\%of model\-verb pairs achieve GOLD in both slot conditions under R2V \(Figure[25](https://arxiv.org/html/2608.04567#A9.F25)\): the agent\-GOLD and patient\-GOLD verb sets are largely disjoint\. Verb difficulty is substantially slot\-dependent, not verb\-inherent\. Agent\-only generation has a structural ceiling set by which verbs admit a valid agent gradient; suggesting comprehensive stimulus coverage requires multi\-slot variation\.

Table 4:GOLD% by model×\\timesstrategy on 60 verbs\.ResSP= reasoning scratchpad;Viz= visual identifiability constraint\.Bold\- each column’s best configuration\.![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/glazier.png)

![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/contractor.png)

![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/karate_instructor.png)

![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/personal_trainer.png)

Figure 4:Visual distinctiveness failure and recovery for the verbbreak, visualized for illustration \(STRIVE outputs sentences, not images\)\.Left\(without visual constraints\): a glazier and a home renovation contractor, satisfying their intended plausibility conditions but visually indistinguishable in generic workwear\.Right\(with visual constraints\): a karate instructor and a personal trainer, satisfying the same plausibility conditions but visually distinct through attire\.
## 7Ablations

We ablate two components of the generation framework using the GPT\-judge evaluation: the visual identifiability constraint \(Viz\) and the reasoning scratchpad \(ResSP\)\. Results are in Table[4](https://arxiv.org/html/2608.04567#S6.T4)\.

The visual pre\-filter strongly affects GOLD eligibility\.Under the text\-based evaluator, removing the visual constraint from RB reduces GOLD% to≤1\.7%\\leq 1\.7\\%across all models\. Without it, the model produces plausibility\-appropriate agents predicted from text to look similar, e\.g\., a glazier and a home renovation contractor in generic workwear \(Figure[4](https://arxiv.org/html/2608.04567#S6.F4), left\)\. Adding the constraint steers the model toward agents with more distinctive occupational cues \(Figure[4](https://arxiv.org/html/2608.04567#S6.F4), right\): for GPT\-5\.1 on RB, GOLD% rises from0\.0%0\.0\\%\(no viz\) to16\.7%16\.7\\%\(viz, no scratchpad\)\. This ablation measures compliance with the text\-based pre\-filter, not image\-level distinctiveness, which requires image\-grounded validation\.

Reasoning scratchpad consistently improves performance\.The scratchpad lifts GOLD% across models and methods\. On single\-shot methods, gains range from\+1\.6\+1\.6to\+11\.6\+11\.6across the open\-source lineup \(e\.g\., Qwen3\-30B\-A3B:\+3\.4\+3\.4on RB,\+4\.6\+4\.6on R2V; Ministral\-14B:\+5\.1\+5\.1on RB\) and GPT\-5\.1 \(\+11\.6\+11\.6on RB,\+5\.0\+5\.0on R2V\)\. On iterative methods, GPT\-5\.1 gains\+35\.0\+35\.0on R2R \(40\.0→75\.040\.0\\to 75\.0\) and\+41\.7\+41\.7on R2VR \(30\.0→71\.730\.0\\to 71\.7\)\.888Open\-source iterative cells are not in the ablation because those runs used Qwen\-judge for iteration feedback \(§[4](https://arxiv.org/html/2608.04567#S4)\); we restrict ablation reads to GPT\-judge for consistency\.The exception is Qwen3\-Next\-80B\-A3B, which regresses by roughly88points on both single\-shot methods\.

Table 5:Aggregate GOLD% by slot condition and method, across the five models with both\-slot coverage\.n=60n=60verbs per cell\.
## 8Conclusion

We introduced STRIVE, a framework for jointly generating and evaluating matched four\-condition event\-plausibility stimulus sets under shared\-frame and single\-slot constraints\. Our results show that structured evaluator feedback and explicit reasoning improve generation, while reasoning effort improves alignment between LLM and human plausibility judgments\. These gains are not uniform: evaluation remains most difficult near the plausibility boundary, and generation difficulty depends on which event slot is varied\. In particular, verbs that support successful agent\-varying sets do not necessarily support patient\-varying sets, showing that difficulty is not inherent to the verb\. These findings position controlled graded\-plausibility generation as a joint problem of event frame construction, slot choice, and fine\-grained evaluation\. STRIVE provides a scalable foundation for psycholinguistic stimulus design and motivates future methods that choose whether to vary the agent, patient, or another event slot for each verb while accounting for uncertainty near the plausibility boundary\.

## 9Limitations

##### Text\-based visual evaluation\.

The LLM evaluator and human annotators both judged visual distinctiveness from text descriptions of agent attire rather than from actual images\. This is a deliberate design choice: STRIVE produces sentence\-level stimuli meant to be consumed by any downstream image\-generation model, so the visual constraint serves only as a text\-based pre\-filter for downstream image generation, not as validation of actual visual distinctiveness\. The low human–human agreement on the visual sub\-tasks \(AgentIDκ=0\.277\\kappa=0\.277, PairDistκ=0\.407\\kappa=0\.407, §[5\.2](https://arxiv.org/html/2608.04567#S5.SS2)\) suggests that text descriptions leave room for annotator imagination to diverge; image\-grounded validation on the downstream stimulus images is a natural complement\.

##### Occupational stereotyping\.

The visual\-identifiability constraint steers agent selection toward professions with standardised attire, which encodes demographic and gender stereotypes \(§[10](https://arxiv.org/html/2608.04567#S10)\)\. Diversifying the agent pool beyond such defaults remains an open methodological question that interacts directly with the visual constraint\.

##### English\-only scope\.

All experiments use English verbs and sentences because the source verb\-naming assessments are English\-language instruments\. Extending STRIVE to other languages requires language\-specific validation of its event schemas and plausibility conditions\.

##### Clinical validation\.

STRIVE produces sentence\-level stimuli intended for downstream image generation and patient\-facing assessment\. Image generation, behavioural validation with persons with aphasia, and clinical reliability studies remain necessary before clinical deployment\.

## 10Ethical Considerations

##### Clinical scope\.

STRIVE is designed to support psycholinguistic research and the development of assessment materials for aphasia; it is not intended for clinical diagnosis or autonomous clinical decision\-making\. Any deployment involving patient populations must be conducted under appropriate institutional ethical oversight and clinician supervision\. Generated stimuli should be reviewed by a qualified clinician before use in patient\-facing settings\.

##### Occupational stereotyping\.

Agent selection in STRIVE relies on occupational roles with distinctive visual presentations\. This design choice encodes stereotyped associations between professions and appearance — for example, defaulting to gender\-specific visual representations of certain roles\. These defaults may not reflect the diversity of actual practitioners and could skew stimuli in ways that affect responses from patients with different cultural or demographic backgrounds\. We flag this as a limitation requiring human review before clinical deployment, and encourage future work to diversify the agent pool beyond role stereotypes\.

##### Annotator welfare\.

The human annotation study involved rating event plausibility from text descriptions only; no sensitive, distressing, or harmful content was included in the stimuli\. Annotators participated voluntarily, and were not exposed to patient data or clinical material\. No personally identifiable information was collected\.

##### Model and data transparency\.

All models used in this work are commercially available or publicly released systems; no proprietary patient data were used at any stage of the pipeline\. The generated stimuli, evaluation outputs, and annotation data will be released for research reproducibility, subject to review to ensure no harmful content is included in the release\.

## 11Use of AI Assistants

We used Claude \(Anthropic\) as a research, coding, and writing assistant during this project\. Claude assisted with brainstorming and iterative design of the generator and evaluator prompts \(§[3](https://arxiv.org/html/2608.04567#S3)\), implementation of the STRIVE framework and analysis code, and drafting and polishing prose in this manuscript\. All design decisions, scientific claims, and analytical conclusions are the authors’ own\. The language models studied as part of STRIVE appear in the paper as research subjects, not as authorial assistants\.

## References

- Alshemali et al\. \(2026\)Safeyah Khaled Alshemali, Daniel Bauer, and Yuval Marton\. 2026\.[Uncovering autoregressive LLM knowledge of thematic fit in event representation](https://doi.org/10.18653/v1/2026.conll-main.11)\.In*Proceedings of the 30th Conference on Computational Natural Language Learning*, pages 165–177, San Diego, California, USA\. Association for Computational Linguistics\.
- Amouyal et al\. \(2024\)Samuel Amouyal, Aya Meltzer\-Asscher, and Jonathan Berant\. 2024\.[Large language models for psycholinguistic plausibility pretesting](https://doi.org/10.18653/v1/2024.findings-eacl.12)\.In*Findings of the Association for Computational Linguistics: EACL 2024*, pages 166–181, St\. Julian’s, Malta\. Association for Computational Linguistics\.
- Anthonio et al\. \(2022\)Talita Anthonio, Anna Sauer, and Michael Roth\. 2022\.[Clarifying implicit and underspecified phrases in instructional text](https://aclanthology.org/2022.lrec-1.354/)\.In*Proceedings of the Thirteenth Language Resources and Evaluation Conference*, pages 3319–3330\. European Language Resources Association\.
- Bicknell et al\. \(2010\)Klinton Bicknell, Jeffrey L\. Elman, Mary Hare, Ken McRae, and Marta Kutas\. 2010\.[Effects of event knowledge in processing verbal arguments](https://doi.org/10.1016/j.jml.2010.08.004)\.*Journal of Memory and Language*, 63\(4\):489–505\.
- Calderon et al\. \(2025\)Nitay Calderon, Roi Reichart, and Rotem Dror\. 2025\.[The alternative annotator test for LLM\-as\-a\-judge: How to statistically justify replacing human annotators with LLMs](https://doi.org/10.18653/v1/2025.acl-long.782)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 16051–16081, Vienna, Austria\. Association for Computational Linguistics\.
- Cho\-Reyes and Thompson \(2012\)Soojin Cho\-Reyes and Cynthia K\. Thompson\. 2012\.[Verb and sentence production and comprehension in aphasia: Northwestern assessment of verbs and sentences \(navs\)](https://doi.org/10.1080/02687038.2012.693584)\.*Aphasiology*, 26\(10\):1250–1277\.
- Cohen \(1960\)Jacob Cohen\. 1960\.[A coefficient of agreement for nominal scales](https://api.semanticscholar.org/CorpusID:15926286)\.*Educational and Psychological Measurement*, 20:37 – 46\.
- Cohen \(1968\)Jacob Cohen\. 1968\.[Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit](https://doi.org/10.1037/h0026256)\.*Psychological Bulletin*, 70:213–220\.
- Dresang et al\. \(2019\)Haley C\. Dresang, Michael Walsh Dickey, and Tessa C\. Warren\. 2019\.[Semantic memory for objects, actions, and events: A novel test of event\-related conceptual semantic knowledge](https://doi.org/10.1080/02643294.2019.1656604)\.*Cognitive Neuropsychology*, 36\(7\-8\):313–335\.
- Eichel and Schulte im Walde \(2023\)Annerose Eichel and Sabine Schulte im Walde\. 2023\.[A dataset for physical and abstract plausibility and sources of human disagreement](https://doi.org/10.18653/v1/2023.law-1.4)\.In*Proceedings of the 17th Linguistic Annotation Workshop \(LAW\-XVII\)*, pages 31–45\. Association for Computational Linguistics\.
- Elman and McRae \(2019\)Jeffrey L\. Elman and Ken McRae\. 2019\.[A model of event knowledge](https://doi.org/10.1037/rev0000133)\.*Psychological Review*, 126\(2\):252–291\.
- Emami et al\. \(2021\)Ali Emami, Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung\. 2021\.[ADEPT: An adjective\-dependent plausibility task](https://doi.org/10.18653/v1/2021.acl-long.553)\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 7117–7128\. Association for Computational Linguistics\.
- Friedman \(1937\)Milton Friedman\. 1937\.[The use of ranks to avoid the assumption of normality implicit in the analysis of variance](http://www.jstor.org/stable/2279372)\.*Journal of the American Statistical Association*, 32\(200\):675–701\.
- Gwet \(2008\)Kilem Li Gwet\. 2008\.[Computing inter\-rater reliability and its variance in the presence of high agreement](https://doi.org/10.1348/000711006X126600)\.*British Journal of Mathematical and Statistical Psychology*, 61\(1\):29–48\.
- Holm \(1979\)Sture Holm\. 1979\.[A simple sequentially rejective multiple test procedure](http://www.jstor.org/stable/4615733)\.*Scandinavian Journal of Statistics*, 6\(2\):65–70\.
- Ivanova et al\. \(2021\)Anna A\. Ivanova, Zachary Mineroff, Vitor Zimmerer, Nancy Kanwisher, Rosemary Varley, and Evelina Fedorenko\. 2021\.[The language network is recruited but not required for nonverbal event semantics](https://doi.org/10.1162/nol_a_00030)\.*Neurobiology of Language*, 2\(2\):176–201\.
- Ivanova et al\. \(2025\)Anna A\. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U\. Kumar, Setayesh Radkani, Thomas H\. Clark, Carina Kauf, Jennifer Hu, R\. T\. Pramod, Gabriel Grand, Vivian C\. Paulun, Maria Ryskina, Ekin Akyürek, Ethan G\. Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Joshua Tenenbaum, and Jacob Andreas\. 2025\.[Elements of world knowledge \( EWoK \): A cognition\-inspired framework for evaluating basic world knowledge in language models](https://doi.org/10.1162/tacl.a.38)\.*Transactions of the Association for Computational Linguistics*, 13:1245–1270\.
- Kauf et al\. \(2024\)Carina Kauf, Emmanuele Chersoni, Alessandro Lenci, Evelina Fedorenko, and Anna A Ivanova\. 2024\.[Log probabilities are a reliable estimate of semantic plausibility in base and instruction\-tuned language models](https://doi.org/10.18653/v1/2024.blackboxnlp-1.18)\.In*Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP*, pages 263–277, Miami, Florida, US\. Association for Computational Linguistics\.
- Kauf et al\. \(2023\)Carina Kauf, Anna A\. Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci\. 2023\.[Event knowledge in large language models: The gap between the impossible and the unlikely](https://doi.org/10.1111/cogs.13386)\.*Cognitive Science*, 47\(11\):e13386\.
- Khalkhali et al\. \(2012\)S\. Khalkhali, Jeffrey D\. Wammes, and Ken McRae\. 2012\.[Integrating words that refer to typical sequences of events](https://doi.org/10.1037/a0027369)\.*Canadian Journal of Experimental Psychology*, 66\(2\):106–114\.
- Li et al\. \(2024\)Huihan Li, Yuting Ning, Zeyi Liao, Siyuan Wang, Xiang Lorraine Li, Ximing Lu, Wenting Zhao, Faeze Brahman, Yejin Choi, and Xiang Ren\. 2024\.[In search of the long\-tail: Systematic generation of long\-tail inferential knowledge via logical rule guided search](https://doi.org/10.18653/v1/2024.emnlp-main.140)\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 2348–2370, Miami, Florida, USA\. Association for Computational Linguistics\.
- Liu et al\. \(2023\)Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah Smith, Yejin Choi, and Hannaneh Hajishirzi\. 2023\.[Vera: A general\-purpose plausibility estimation model for commonsense statements](https://doi.org/10.18653/v1/2023.emnlp-main.81)\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 1264–1287, Singapore\. Association for Computational Linguistics\.
- Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark\. 2023\.[Self\-refine: Iterative refinement with self\-feedback](https://openreview.net/forum?id=S37hOerQLB)\.In*Thirty\-seventh Conference on Neural Information Processing Systems*\.
- Masterson \(2000\)Jackie Masterson\. 2000\.[An object and action naming battery](https://doi.org/10.1037/t41579-000)\.*Journal of Neurolinguistics*\.
- McRae et al\. \(2005\)Ken McRae, Mary Hare, Jeffrey L\. Elman, and Todd Ferretti\. 2005\.[A basis for generating expectancies for verbs from nouns](https://doi.org/10.3758/bf03193221)\.*Memory & Cognition*, 33\(7\):1174–1184\.
- McRae et al\. \(1998\)Ken McRae, Michael J\. Spivey\-Knowlton, and Michael K\. Tanenhaus\. 1998\.[Modeling the influence of thematic fit \(and other constraints\) in on\-line sentence comprehension](https://doi.org/10.1006/jmla.1997.2543)\.*Journal of Memory and Language*, 38\(3\):283–312\.
- Nye et al\. \(2022\)Maxwell Nye, Anders Johan Andreassen, Guy Gur\-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena\. 2022\.[Show your work: Scratchpads for intermediate computation with language models](https://openreview.net/forum?id=HBlx2idbkbq)\.In*Deep Learning for Code Workshop*\.
- Padó et al\. \(2006\)Ulrike Padó, Matthew Crocker, and Frank Keller\. 2006\.[Modelling semantic role pausibility in human sentence processing](https://aclanthology.org/E06-1044/)\.In*11th Conference of the European Chapter of the Association for Computational Linguistics*, pages 345–352, Trento, Italy\. Association for Computational Linguistics\.
- Porada et al\. \(2019\)Ian Porada, Kaheer Suleman, and Jackie Chi Kit Cheung\. 2019\.[Can a gorilla ride a camel? learning semantic plausibility from text](https://doi.org/10.18653/v1/D19-6015)\.In*Proceedings of the First Workshop on Commonsense Inference in Natural Language Processing*, pages 123–129\. Association for Computational Linguistics\.
- Porada et al\. \(2021\)Ian Porada, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung\. 2021\.[Modeling event plausibility with consistent conceptual abstraction](https://doi.org/10.18653/v1/2021.naacl-main.138)\.In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 1732–1743\. Association for Computational Linguistics\.
- Pyatkin et al\. \(2021\)Valentina Pyatkin, Shoval Sadde, Aynat Rubinstein, Paul Portner, and Reut Tsarfaty\. 2021\.[The possible, the plausible, and the desirable: Event\-based modality detection for language processing](https://doi.org/10.18653/v1/2021.acl-long.77)\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*\. Association for Computational Linguistics\.
- Qian et al\. \(2025\)Ming Qian, Terry Patten, Spencer Lynn, Aaron Winder, and Maxwell Pickering\. 2025\.[Generating neurolinguistic stimuli using llm prompting](https://doi.org/10.1007/978-3-032-13184-3_25)\.In*HCI International 2025 – Late Breaking Papers: 27th International Conference on Human\-Computer Interaction, HCII 2025, Gothenburg, Sweden, June 22–27, 2025, Proceedings, Part XV*, page 404–419, Berlin, Heidelberg\. Springer\-Verlag\.
- Resnik \(1996\)Philip Resnik\. 1996\.[Selectional constraints: an information\-theoretic model and its computational realization](https://doi.org/10.1016/S0010-0277(96)00722-6)\.*Cognition*, 61\(1\):127–159\.Compositional Language Acquisition\.
- Swinburn et al\. \(2004\)K\. Swinburn, G\. Porter, and D\. Howard\. 2004\.[*Comprehensive Aphasia Test*](https://books.google.com/books?id=MMVdPwAACAAJ)\.Psychology Press\.
- Tang et al\. \(2023\)Liyan Tang, Yifan Peng, Yanshan Wang, Ying Ding, Greg Durrett, and Justin Rousseau\. 2023\.[Less likely brainstorming: Using language models to generate alternative hypotheses](https://doi.org/10.18653/v1/2023.findings-acl.794)\.In*Findings of the Association for Computational Linguistics: ACL 2023*, pages 12532–12555, Toronto, Canada\. Association for Computational Linguistics\.
- Wang et al\. \(2025\)Qianli Wang, Tatiana Anikina, Nils Feldhus, Simon Ostermann, Sebastian Möller, and Vera Schmitt\. 2025\.[Cross\-refine: Improving natural language explanation generation by learning in tandem](https://aclanthology.org/2025.coling-main.77/)\.In*Proceedings of the 31st International Conference on Computational Linguistics*, pages 1150–1167, Abu Dhabi, UAE\. Association for Computational Linguistics\.
- Wang et al\. \(2018\)Su Wang, Greg Durrett, and Katrin Erk\. 2018\.[Modeling semantic plausibility by injecting world knowledge](https://doi.org/10.18653/v1/N18-2049)\.In*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\)*, pages 303–308\. Association for Computational Linguistics\.
- Warren et al\. \(2015\)Tessa Warren, Evelyn Milburn, Nikole D\. Patson, and Michael Walsh Dickey\. 2015\.[Comprehending the impossible: What role do selectional restriction violations play?](https://doi.org/10.1080/23273798.2015.1047458)*Language, Cognition and Neuroscience*, 30\(8\):932–939\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H\. Chi, Quoc V Le, and Denny Zhou\. 2022\.[Chain of thought prompting elicits reasoning in large language models](https://openreview.net/forum?id=_VjQlMeSB_J)\.In*Advances in Neural Information Processing Systems*\.
- Wilcoxon \(1945\)Frank Wilcoxon\. 1945\.[Individual comparisons by ranking methods](http://www.jstor.org/stable/3001968)\.*Biometrics Bulletin*, 1\(6\):80–83\.
- Yuan et al\. \(2024\)Moy Yuan, Eric Chamoun, Rami Aly, Chenxi Whitehouse, and Andreas Vlachos\. 2024\.[PRobELM: Plausibility ranking evaluation for language models](https://openreview.net/forum?id=k8KS9Ps71d)\.In*First Conference on Language Modeling*\.
- Zhang et al\. \(2026\)Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael Tomz, Christopher D Manning, and Weiyan Shi\. 2026\.[Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity](https://openreview.net/forum?id=zloIrd77G5)\.In*Forty\-third International Conference on Machine Learning*\.
- Zhao et al\. \(2024\)Wenting Zhao, Justin T\. Chiu, Jena Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Lorraine Li, and Alane Suhr\. 2024\.[UNcommonsense reasoning: Abductive reasoning about uncommon situations](https://doi.org/10.18653/v1/2024.naacl-long.469)\.In*Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 8487–8505, Mexico City, Mexico\. Association for Computational Linguistics\.

## Appendix ADetailed Related Work

##### Psycholinguistic uses of controlled event stimuli\.

Controlled event stimuli support several psycholinguistic paradigms\.McRae et al\. \([2005](https://arxiv.org/html/2608.04567#bib.bib25)\)show that nouns generate expectancies for verbs, andKhalkhali et al\. \([2012](https://arxiv.org/html/2608.04567#bib.bib20)\)examine the integration of words referring to typical event sequences\.Bicknell et al\. \([2010](https://arxiv.org/html/2608.04567#bib.bib4)\)study how event knowledge affects the processing of verbal arguments, whileWarren et al\. \([2015](https://arxiv.org/html/2608.04567#bib.bib38)\)examine the comprehension of impossible events\. Across modalities,Ivanova et al\. \([2021](https://arxiv.org/html/2608.04567#bib.bib16)\)compare plausibility judgments for sentences and line drawings, andDresang et al\. \([2019](https://arxiv.org/html/2608.04567#bib.bib9)\)develop an assessment of event\-related conceptual semantic knowledge\. These studies demonstrate the breadth of psycholinguistic uses for controlled event stimuli, but do not address their automated construction as matched, multilevel sets\.

##### Relation to selectional preference and thematic fit\.

Selectional preferencedescribes a predicate’s probabilistic expectations about the semantic properties of fillers in its argument slots\(Resnik,[1996](https://arxiv.org/html/2608.04567#bib.bib33)\), whereasthematic fitdescribes the graded compatibility of a particular filler with an event slot in context\(McRae et al\.,[1998](https://arxiv.org/html/2608.04567#bib.bib26); Padó et al\.,[2006](https://arxiv.org/html/2608.04567#bib.bib28); Alshemali et al\.,[2026](https://arxiv.org/html/2608.04567#bib.bib1)\)\. Both constructs are broadly related to STRIVE’s prompt design\. Requiring an agent\-slot filler to be a living human identified by a profession constrains the semantic class of possible fillers\. In the workedcatchexample, the model considers how well a goalie, referee, hockey player, or ballerina fills the agent slot in the fixed soccer\-ball scene; this resembles thematic\-fit reasoning\. However, these choices are made internally by the model\. STRIVE does not independently manipulate or score either construct, nor does it test how accurately the model represents them\. Our evaluation concerns only the resulting complete\-event plausibility, defined as the degree to which the complete situation accords with event knowledge and ordinary world knowledge\. Selectional preference and thematic fit therefore describe possible mechanisms informing generation, not controlled variables or measured outcomes in this study\.

##### Plausibility modeling and measurement\.

Wang et al\. \([2018](https://arxiv.org/html/2608.04567#bib.bib37)\)construct a 3,062\-item dataset of subject, verb, and object triples from crowdsourced subject\-verb and verb\-object pairs and collect binary physical\-plausibility labels\.Porada et al\. \([2019](https://arxiv.org/html/2608.04567#bib.bib29)\)learn plausibility from naturally occurring corpus events using self\-supervision\.Porada et al\. \([2021](https://arxiv.org/html/2608.04567#bib.bib30)\)train on Wikipedia events paired with random argument perturbations and improve consistency across conceptual abstractions\. These methods estimate the plausibility of individual events rather than construct graded matched sets\. Graded resources operationalize different quantities\. ADEPT\(Emami et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib12)\)pairs a sentence with a modified version formed by adding an adjective to a noun, then assigns the pair a five\-way label indicating how the adjective affects the event’s plausibility\. These labels encode changes within sentence pairs, not five matched conditions in a shared event frame\.Eichel and Schulte im Walde \([2023](https://arxiv.org/html/2608.04567#bib.bib10)\)collect absolute slider ratings for corpus\-derived original and pseudo\-implausible event triples; because the midpoint cannot be submitted, the scale provides four response values \(1, 2, 4, and 5\)\. PRobELM\(Yuan et al\.,[2024](https://arxiv.org/html/2608.04567#bib.bib41)\)constructs one most\-plausible scenario and ten less\-plausible alternatives from Wikidata, then evaluates LM rankings of those scenarios\. Related tasks assess other units, including candidate clarifications in CLAIRE\(Anthonio et al\.,[2022](https://arxiv.org/html/2608.04567#bib.bib3)\)and event\-based modality categories\(Pyatkin et al\.,[2021](https://arxiv.org/html/2608.04567#bib.bib31)\)\. These resources provide binary labels, pairwise changes, absolute ratings, rankings, or modality categories rather than slot\-controlled matched sets\.

##### Controlled and atypical content generation\.

Zhao et al\. \([2024](https://arxiv.org/html/2608.04567#bib.bib43)\)generate abductive explanations that make unexpected outcomes more likely in context\.Li et al\. \([2024](https://arxiv.org/html/2608.04567#bib.bib21)\)use logical\-rule\-guided search to generate factually correct but low\-confidence inferential statements\.Qian et al\. \([2025](https://arxiv.org/html/2608.04567#bib.bib32)\)use GPT\-4o to generate structurally controlled neurolinguistic vignettes with expected and unexpected conditions\. More directly,Eichel and Schulte im Walde \([2023](https://arxiv.org/html/2608.04567#bib.bib10)\)replace two constituents of corpus\-derived event triples to produce pseudo\-implausible examples, whileTang et al\. \([2023](https://arxiv.org/html/2608.04567#bib.bib35)\)train models to generate relevant but less\-likely alternative hypotheses\. These methods provide vignette\-level structural control, two\-constituent perturbation, or one\-band generation rather than shared\-frame, single\-slot control across multiple ordered conditions\. To our knowledge, none starts from a root verb, generates all event\-slot values, and constructs multiple plausibility targets while holding every non\-target slot fixed\. This slot\-controlled construction setting defines the STRIVE generation contribution rather than graded plausibility generation in general\.

##### Evaluating and refining plausibility\.

Kauf et al\. \([2024](https://arxiv.org/html/2608.04567#bib.bib18)\)compare sentence log probabilities with direct zero\-shot prompting and find log probabilities more reliable for estimating semantic plausibility\.Amouyal et al\. \([2024](https://arxiv.org/html/2608.04567#bib.bib2)\)test LM ratings as substitutes for human psycholinguistic pretesting and find that even GPT\-4 lacks sufficient fine\-grained discrimination\. VERA\(Liu et al\.,[2023](https://arxiv.org/html/2608.04567#bib.bib22)\)is a trained general\-purpose plausibility estimator for commonsense statements\. Using minimally different event pairs,Kauf et al\. \([2023](https://arxiv.org/html/2608.04567#bib.bib19)\)find that LMs distinguish possible from impossible events more consistently than likely from unlikely events\. EWoK\(Ivanova et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib17)\)evaluates conceptual world knowledge more broadly rather than event plausibility alone\. Self\-Refine\(Madaan et al\.,[2023](https://arxiv.org/html/2608.04567#bib.bib23)\)uses feedback from the generating model itself, while Cross\-Refine\(Wang et al\.,[2025](https://arxiv.org/html/2608.04567#bib.bib36)\)uses feedback from a separate critic model\. STRIVE instead uses evaluator verdicts on both condition placement and set\-level slot control to iteratively revise a matched four\-condition set, while evaluator models and reasoning settings are assessed against human judgments \(§[3\.2](https://arxiv.org/html/2608.04567#S3.SS2); §[4\.2](https://arxiv.org/html/2608.04567#S4.SS2)\)\.

## Appendix BRB Generator Prompt

The full RB generator prompt is shown in Figure[5](https://arxiv.org/html/2608.04567#A2.F5)\. It implements the five\-step procedure described in §[3\.1\.1](https://arxiv.org/html/2608.04567#S3.SS1.SSS1)and specifies the output JSON schema\.

RB Generation Prompt \(Part 1/9\)=== SYSTEM PROMPT ===You are an expert psycholinguist who understands event cognition, thematic roles, and typicalityeffects\.Your task is to generate event sentences that vary systematically in plausibility\.You reason step\-by\-step before producing output, and you return only valid JSON\.=== USER TEMPLATE ===TASKYou will be given a verb \(root form\) at the end of this prompt\. Your job is to generate exactlyfour event sentences that vary in plausibility for that verb\.Each sentence describes a concrete, observable scene that could be depicted in a singlephotograph\.IMPORTANT: These sentences will later be used to generate images\. A human participant will thenview each image and judge whether the depicted event is plausible or not\. Therefore, the fouragents must be visually distinguishable from one another in a photograph \-\- each agent’sprofessionshould be identifiable from their appearance alone \(clothing, gear, accessories, physique, etc\.\)\.\-\-\-EXPERIMENTAL CONTEXT\-\-\-These sentences will be converted into images shown to brain injury patients who must make quickplausible/implausible judgments\. The patients see ONLY the image \-\- no text, no captions, nolabels\.Therefore:\(1\) Each agent’s profession must be INSTANTLY recognizable from appearance alone\.\(2\) The plausibility distinction must be visually obvious and must not depend on subtlereasoning\.\(3\) The scene must be concrete and unambiguous\.If a profession cannot be identified from a photograph without a caption, do NOT use thatprofession\.\-\-\-SENTENCE STRUCTURE\-\-\-Every sentence MUST follow this exact slot order:\[Agent\] \[verb\] \[patient\] \[instrument\] \[location\]Slot definitions:\- Agent: A living human identified by profession or social role \(the doer of the action\)\.\- Verb: A chosen inflected form of the root verb \(e\.g\., "is catching", "catches", "caught"\)\.Use the SAME inflected form in ALL four sentences\.\- Patient: The object or entity being acted upon\.\- Instrument: The tool, body part, or means used to perform the action \(phrased with "with"\)\.\- Location: Where the event takes place \(phrased with "in", "on", "at", etc\.\)\.Example: "A goalie is catching a soccer ball with his hands on the soccer field\."Agent = "A goalie"Verb = "is catching"Patient = "a soccer ball"Instrument = "with his hands"Location = "on the soccer field"\-\-\-CORE CONCEPT: THEMATIC FITFigure 5:RB Generation Prompt\. Continued in Figure[6](https://arxiv.org/html/2608.04567#A2.F6)\.RB Generation Prompt \(Part 2/9\)\-\-\-Thematic fit is how naturally a particular AGENT fits as the doer of the described action\(verb \+ patient \+ instrument \+ location\)\. It is about whether this professional, in theirtypical work context, would perform this specific action on this patient using this instrumentin this location\.Example: Given the sentence frame "\_\_\_ is catching a soccer ball with his hands on the soccerfield":\- "A goalie" has HIGH thematic fit \(this is exactly their job\)\.\- "A referee" has MODERATE thematic fit \(present in the setting, could do it, not their mainrole\)\.\- "A hockey player" has LOW thematic fit \(related sports domain, but wrong sport and wrongsetting\)\.\- "A ballerina" has VERY LOW thematic fit \(completely unrelated domain\)\.The four conditions form a gradient of thematic fit from high to low:plausible\_easy \-\-\- plausible\_hard \-\-\- implausible\_hard \-\-\- implausible\_easy\(clear yes\) \(unsure yes\) \(unsure no\) \(clear no\)HIGH fit MODERATE fit LOW fit VERY LOW fit\-\-\-CONDITION DEFINITIONS\-\-\-1\) PLAUSIBLE\_EASY \-\- "clear yes" \-\- HIGH thematic fitThe agent is the prototypical, most expected professional for performing \[verb\] on \[patient\]with \[instrument\] in \[location\]\. The event is physically possible, reasonable, and typical\.A human would instantly accept this as a normal, everyday event\.Agent selection: Pick the profession whose job routinely involves exactly this actionin this setting\.2\) PLAUSIBLE\_HARD \-\- "unsure yes" \-\- MODERATE thematic fitThe agent is a professional who could plausibly perform \[verb\] on \[patient\] with \[instrument\]in \[location\], but it is NOT their primary role\. The event is physically possible andreasonable,but less typical or less frequent\.A human would pause but ultimately accept it as possible\.Agent selection: Pick a profession that is ALREADY PRESENT or COULD NATURALLY BE PRESENT in\[location\], and who has the physical ability to perform \[verb\] on \[patient\] with \[instrument\],but for whom this action is secondary or occasional rather than central to their role\.KEY TEST: Can you easily imagine a specific, realistic scenario where this professional doesthis? If yes, it qualifies\.3\) IMPLAUSIBLE\_HARD \-\- "unsure no" \-\- LOW thematic fitThe agent is a professional from a RELATED but WRONG domain\. There IS semantic overlap betweenthe agent’s profession and the plausible\_easy agent’s profession, but the agent stilldoes NOT belong in \[location\] performing \[verb\] on \[patient\] with \[instrument\]\.The event is unreasonable and atypical, but the domain similarity makes a human hesitatebefore rejecting it\.CRITICAL \-\- MULTI\-DIMENSIONAL OVERLAP REQUIREMENT:The implausible\_hard agent must share overlap with the plausible\_easy agent on AT LEAST TWOof the following five dimensions:\(a\) Occupational domain \(e\.g\., both belong to the same broad professional field\)\(b\) Typical work setting/environment \(e\.g\., both typically work in similar environments\)\(c\) Tools or instruments routinely used \(e\.g\., both routinely use similar types of tools\)\(d\) Type of patient/object acted upon \(e\.g\., both act on similar types of objects orentities\)\(e\) Physical actions routinely performed \(e\.g\., both routinely perform similar physicalactions\)Figure 6:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 3/9\)A single dimension of overlap \(e\.g\., both use sharp tools\) is NOT SUFFICIENT\. If the onlyconnection is through the verb itself, the agent belongs in implausible\_easy, not here\.NEGATIVE TEST: Remove the shared verb from consideration\. Is there STILL a reason to associatethis agent with the scene? If not, the overlap is too thin\.Agent selection: Pick a profession that shares at least two of the above dimensions with theplausible\_easy agent but who would NOT perform THIS specific action in THIS specific setting\.KEY TEST: Does this agent make you briefly think "wait, maybe\.\.\." before you conclude "no"?If yes, it qualifies\.4\) IMPLAUSIBLE\_EASY \-\- "clear no" \-\- VERY LOW thematic fitThe agent is a professional from a COMPLETELY UNRELATED domain\. There is no semantic connectionbetween the agent’s profession and the action, patient, instrument, or location\.The event is unreasonable and deeply atypical\.A human would instantly reject it\.Agent selection: Pick a profession that has NO topical, domain, or contextual overlap with anypart of the event\. The agent’s professional world should be maximally distant from the scene\.\-\-\-HARD CONSTRAINTS\-\-\-AGENT RULES:\- Every agent MUST be a living human identified by a real profession or social role\.Valid examples: any real\-world profession or social role \(e\.g\., "a \[profession\]"\)\.\- Age\-based social roles with clear visual markers are also acceptable when they carrydistinct role expectations\. Valid examples: "a schoolchild" \(school uniform, backpack,small stature\), "a retiree" or "an elderly person" \(gray hair, glasses, possibly a cane\)\.These roles are useful because people have schematic expectations about what a schoolchildor a retiree would or would not do, which creates natural thematic fit variation\.\- IMPORTANT: Plausibility must always stem from ROLE\-ACTION MISMATCH \(whether this roletypically performs this action\), NOT from physical incapability\. Do not use agents whoseimplausibility comes from being physically unable to perform the action \(e\.g\., an infantlifting heavy equipment\)\. The question is "does this role fit this event?" not "can thisperson physically do this?"\- FORBIDDEN agents: animals, objects, fictional beings, statues, robots, corpses, body parts,descriptions of states \(e\.g\., "a sleeping person"\), physical descriptors without a role\(e\.g\., "a tall person", "an overweight person"\), or any non\-human entity\.\- Assume each agent wears clothing and gear appropriate to their profession or social role\.PROFESSION IDENTIFIABILITY RULES:\- Every agent MUST have a profession that is visually identifiable from a photograph WITHOUTany caption, label, or scene context\. A naive viewer seeing ONLY the person in their typicalprofessional attire must be able to correctly guess their profession\.\- A profession is visually identifiable if it has at least one of the following:\(a\) A dedicated uniform or standardized work attire distinct from everyday civilian clothing\.\(b\) Profession\-specific gear or equipment that is worn or carried as part of the role\.\(c\) A highly distinctive physical presentation strongly associated with the role\.\- A profession is NOT visually identifiable if the person would look like a generic civilianin everyday clothing\. The test: Imagine this person standing alone against a plain whitebackground \-\- no scene context, no caption, no other people\. Could a viewer correctly guesstheir profession from appearance alone? If not, do NOT use that profession\.\- When in doubt, prefer professions with uniforms or standardized professional attire overthose with informal or variable dress codes\.Figure 7:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 4/9\)VISUAL DISTINCTIVENESS RULES:\- All four agents must be visually distinguishable from one another in a photograph\.\- Each agent’s profession should be identifiable by their typical professional appearance:clothing, uniform, gear, accessories, protective equipment, or other visual markers\.\- STRONGLY PREFER professions that have a recognizable "look" \-\- those with distinctiveuniforms, professional gear, or role\-specific attire that a viewer could identify withoutany caption or context\. The more visually iconic the profession, the better\.Avoid: professions that look like generic civilians in everyday clothing, such as"a freelance writer", "a software engineer", "a real estate agent", "an accountant"\.\- Pairs of agents that would look nearly identical in a photo are NOT acceptable\.If two agents would look nearly identical in a photograph due to similar workwear, replaceone with a profession that has a more distinctive visual identity while preserving thematic fit\.\- When two professions share the same broad domain, ensure they differ in at least onestrong visual cue \(uniform type, headgear, tools they carry, body build, etc\.\)\.CONSISTENCY RULES:\- The verb form, patient, instrument, and location MUST be IDENTICAL in all four sentences\.\- ONLY the agent changes between conditions\.IMAGEABILITY:\- Each sentence must depict a scene that a photographer could capture in one still image\.\- Avoid abstract, metaphorical, or unobservable events\.\-\-\-PROCEDURE \(follow these six steps in order\)\-\-\-STEP 1 \-\- DESIGN THE SCENEChoose a naturalistic, concrete scene for the root verb\.Select a specific verb form, patient, instrument, and location that together create a vivid,everyday scenario\. The scene should be one where a clear prototypical agent exists\.STEP 1b \-\- VERIFY SCENE TYPICALITYBefore selecting any agents, verify that the scene itself supports a prototypical event\.Construct a test sentence by replacing the agent with "someone":"Someone \[verb\_form\] \[patient\] \[instrument\] \[location\]\."Ask: Does this sentence describe a ROUTINE, EVERYDAY event that a person would recognize astypical and unsurprising? Would a person hearing this sentence think "yes, that happensregularly"?If the answer is NO \-\- the scene is atypical regardless of the agent \(e\.g\., "Someone is catchinga soccer ball with his hands in a library" \-\- catching soccer balls in a library is not routinefor anyone\) \-\- then REDESIGN the scene\. Choose a different patient, instrument, or location wherethe verb represents a routine, everyday activity with a clear prototypical agent\.This step prevents a common failure mode: choosing a location first \(e\.g\., library\) and thenforcing a setting\-familiar agent \(e\.g\., librarian\) into an action that is not actually typicalin that setting\. The scene must be action\-typical, not just setting\-familiar\.Only proceed to Step 2 after the "someone" test passes\.Figure 8:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 5/9\)STEP 2 \-\- SELECT FOUR AGENTS ALONG THE THEMATIC\-FIT GRADIENTWorking from the scene you designed:a\) plausible\_easy agent: Who is THE most expected professional performing \[verb\] on \[patient\]with \[instrument\] in \[location\]? \(prototypical, high thematic fit\)CORE\-JOB VERIFICATION: Confirm that \[verb\] \[patient\] is part of this agent’s PRIMARYprofessional duties \-\- not merely something they could do, but something they are PAID orEXPECTED to do routinely\. Ask: "If this agent were removed from \[location\], would \[verb\]\[patient\] stop happening there?" If the answer is no, this agent is setting\-prototypical\(familiar with the location\) but not action\-prototypical \(the action is not their core job\)\.A setting\-prototypical agent belongs in plausible\_hard, not plausible\_easy\.Example: A groundskeeper on a soccer field is setting\-prototypical, but "catchinga soccer ball" is not a groundskeeper’s job\. The groundskeeper belongs in plausible\_hardat best\.b\) plausible\_hard agent: Who else might do this in \[location\], even if it is not their mainjob? \(atypical but naturally present, moderate thematic fit\)c\) implausible\_hard agent: Who SEEMS related due to domain overlap but does NOT belong in thisspecific scene? This agent must share overlap on AT LEAST TWO of the five dimensions\(occupational domain, setting, tools, patient type, physical actions\) with the plausible\_easyagent\. Apply the negative test: removing the shared verb, is there still a reason toassociate this agent with the scene?d\) implausible\_easy agent: Who has ZERO connection to any part of this scene?\(maximally distant profession, very low thematic fit\)STEP 3 \-\- CHECK VISUAL DISTINCTIVENESS AND IDENTIFIABILITYFirst, for each agent individually, ask: Is this profession visually identifiable from aphotograph against a plain white background \-\- no scene context, no caption, no other people?Does this agent have a uniform, professional gear, or distinctive attire that makes theirprofession obvious? If not, replace the agent with a profession that has stronger visualidentity while preserving its thematic\-fit level\.Then, for each pair of agents, ask: Would these two look different in a photograph?Consider their typical professional attire, gear, accessories, and physical presentation\.\- Can a viewer tell plausible\_easy from plausible\_hard by appearance? If not, replace one\.\- Can a viewer tell implausible\_hard from implausible\_easy by appearance? If not, replace one\.\- Can a viewer tell ANY two agents apart? If any pair looks identical, replace one agentwith a profession that has a more distinctive visual identity while preserving itsthematic\-fit level\.If you make replacements, re\-verify that thematic fit is still correct for the new agent\.STEP 4 \-\- COMPOSE FOUR SENTENCESInsert each agent into the shared sentence frame:\[Agent\] \[verb form\] \[patient\] \[instrument\] \[location\]\.STEP 5 \-\- VALIDATECheck every item below before producing output:\- All four agents are living humans with identifiable professions? YES/NO\- All four agents are visually identifiable from a photograph without captions? YES/NO\- Verb form, patient, instrument, and location identical across all four sentences? YES/NO\- plausible\_easy: Would a person INSTANTLY say "yes, that’s normal"? YES/NO\- Scene typicality: Does "Someone \[verb\] \[patient\] \[instrument\] \[location\]" sound routine?YES/NO\- plausible\_easy CORE\-JOB: Is \[verb\] \[patient\] part of this agent’s primary job description?YES/NO\- plausible\_easy: Is the agent action\-prototypical \(not just setting\-familiar\)? YES/NO\- plausible\_hard: Would a person PAUSE then say "yes, I suppose that could happen"? YES/NO\- implausible\_hard: Would a person HESITATE then say "no, that doesn’t fit"? YES/NO\- implausible\_hard: Does this agent share at least TWO overlap dimensions with plausible\_easy?YES/NO\- implausible\_hard: Removing the shared verb, is there still a reason to associate this agentwith the scene? YES/NO\- implausible\_easy: Would a person INSTANTLY say "no, that’s absurd"? YES/NOFigure 9:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 6/9\)\- Ordering is correct: plausible\_easy \> plausible\_hard \>\> implausible\_hard \> implausible\_easy?YES/NO\- Each sentence is imageable as a single photograph? YES/NO\- All four agents are visually distinguishable from each other in a photograph? YES/NO\- Each agent’s profession can be identified from their typical appearance? YES/NOIf any check fails, revise before outputting\.STEP 6 \-\- OUTPUTReturn the JSON object specified in OUTPUT FORMAT below\.\-\-\-WORKED EXAMPLE \(root verb = "catch"\)\-\-\-Scene design:verb\_form = "is catching", patient = "a soccer ball",instrument = "with his hands", location = "on the soccer field"Scene typicality check:"Someone is catching a soccer ball with his hands on the soccer field\."Is this routine? YES \-\- catching soccer balls on a field is a typical, everyday sporting event\.Proceed to agent selection\.Agent selection along thematic\-fit gradient:a\) plausible\_easy \-\> "A goalie":Catching the ball is the goalie’s primary job on the soccer field\. Prototypical agent\.HIGH fit\.Visual: goalie jersey, gloves, distinct from other players\.b\) plausible\_hard \-\> "A referee":Referees are present on the soccer field and might catch a ball occasionally \(e\.g\., ballthrown to them, preventing a deflection\)\. Not their main role, but realistic\. MODERATE fit\.Visual: black referee uniform, whistle \-\- clearly different from a goalie\.c\) implausible\_hard \-\> "A hockey player":Hockey is a related team sport \(athletes, goals, ball/puck\), creating domain overlap thatcauses momentary confusion\. But a hockey player does not belong on a soccer field\. LOW fit\.Overlap dimensions: \(a\) occupational domain \-\- both are team sport athletes; \(e\) physicalactions \-\- both catch/block projectiles aimed at a goal\. That is 2 dimensions = sufficient\.Negative test: Removing "catching" from consideration, is there still a reason to associatea hockey player with a soccer field? Yes \-\- both are team sports with goals, balls/pucks,fields/rinks, and athletic roles\. The association persists beyond the verb\.Visual: hockey jersey, pads, helmet \-\- distinct from both goalie and referee\.d\) implausible\_easy \-\> "A ballerina":Ballet has zero connection to soccer, team sports, or outdoor fields\. No semantic overlapwith any part of the scene\. VERY LOW fit\.Visual: tutu, pointe shoes, hair in bun \-\- maximally distinct from all others\.Visual identifiability check:\- Goalie: goalie jersey \+ gloves = IDENTIFIABLE \(distinct sports uniform\)\- Referee: black uniform \+ whistle = IDENTIFIABLE \(recognizable official attire\)\- Hockey player: hockey gear \+ pads \+ helmet = IDENTIFIABLE \(sport\-specific equipment\)\- Ballerina: tutu \+ pointe shoes \+ bun = IDENTIFIABLE \(iconic performance attire\)Visual distinctiveness check \(pairwise\):\- Goalie \(goalie jersey \+ gloves\) vs Referee \(black uniform \+ whistle\): DISTINCT\- Hockey player \(hockey gear \+ pads\) vs Ballerina \(tutu \+ pointe shoes\): DISTINCT\- All six pairwise comparisons: DISTINCTFigure 10:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 7/9\)Output:\{"reasoning": "STEP 1 \- SCENE DESIGN: For ’catch’, a natural everyday scene is a soccer fieldwhere catching is a routine action\. verb\_form=’is catching’, patient=’a soccer ball’,instrument=’with his hands’, location=’on the soccer field’\. The goalie is the clearprototypical doer\.\\n\\nSTEP 1b \- SCENE TYPICALITY CHECK: ’Someone is catching a soccer ballwith his hands on the soccer field\.’ Is this a routine, unsurprising event? YES \-\- thishappens every match\. The scene is action\-typical, not just setting\-familiar\. Proceed toagent selection\.\\n\\nSTEP 2 \- AGENT SELECTION:\\n\(a\) plausible\_easy = A goalie\. CORE\-JOBVERIFICATION: catching shots on goal is the goalie’s primary professional duty \-\- they arepaid to do exactly this\. REMOVAL TEST: if the goalie is removed from the field, catchingshots at the goal stops happening\. Goalie is action\-prototypical, not just setting\-familiar\.HIGH fit\.\\n\(b\) plausible\_hard = A referee\. Already present on the field and couldrealistically catch a ball thrown to them by a player or block a stray deflection\. Not theirmain role but easy to imagine\. MODERATE fit\.\\n\(c\) implausible\_hard = A hockey player\.Overlap dimensions with goalie: \(1\) occupational\_domain \-\- both are team sport athletes; \(2\)physical\_actions \-\- both catch or block projectiles aimed at a goal\. Two dimensions =sufficient\. Negative test: remove the verb ’catching’ from consideration\. Is there still areason to associate a hockey player with a soccer field? Yes \-\- both are team sports playedon a field/rink with goals, athletic roles, and competitive structure\. The associationpersists beyond the verb\. LOW fit\.\\n\(d\) implausible\_easy = A ballerina\. Ballet has zeroconnection to soccer, team sports, or outdoor fields\. No semantic overlap on any dimension\.VERY LOW fit\.\\n\\nSTEP 3 \- VISUAL CHECK:\\nIdentifiability \(each agent against a plain whitebackground\):\\n\- Goalie: goalie jersey \+ padded gloves = IDENTIFIABLE\.\\n\- Referee: blackstriped uniform \+ whistle = IDENTIFIABLE\.\\n\- Hockey player: hockey jersey \+ shoulder/shinpads \+ helmet = IDENTIFIABLE\.\\n\- Ballerina: tutu \+ pointe shoes \+ bun =IDENTIFIABLE\.\\nPairwise \(6 pairs\): all distinct\. Goalie vs referee differ in jersey colorand gloves; hockey player vs ballerina are maximally different; cross\-pairs are obviouslydistinct\. No replacements needed\.\\n\\nSTEP 5 \- VALIDATION: All checks pass \-\- four livinghuman professions, all visually identifiable, slot consistency across all four sentences,scene typicality confirmed, plausible\_easy passes core\-job and action\-prototypicality,implausible\_hard has 2 overlap dimensions and survives the negative test, gradient orderingis correct, all four sentences imageable as single photographs, and pairwise visualdistinctiveness holds for all six pairs\.","scene": \{"verb\_form": "is catching","patient": "a soccer ball","instrument": "with his hands","location": "on the soccer field"\},"agents": \{"plausible\_easy": \{"agent": "A goalie","thematic\_fit": "HIGH","rationale": "Catching the ball is the goalie’s primary job on the soccer field\.Prototypical agent for this event\.","visual\_markers": "Goalie jersey in a bright color, padded goalie gloves, athleticshorts, cleats\."\},"plausible\_hard": \{"agent": "A referee","thematic\_fit": "MODERATE","rationale": "Referees are present on the field and could catch the ball in unusualbut realistic situations\. Not their main role\.","visual\_markers": "Black referee uniform with vertical stripes, whistle around neck,no gloves\."\},"implausible\_easy": \{"agent": "A ballerina","thematic\_fit": "VERY LOW","rationale": "Ballet is an unrelated domain with zero semantic overlap to soccer orFigure 11:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 8/9\)team sports\.","visual\_markers": "White or pink tutu, pointe shoes, hair in a tight bun, slender build\."\},"implausible\_hard": \{"agent": "A hockey player","thematic\_fit": "LOW","rationale": "Hockey is a related team sport \(athletes, goals, ball/puck\), creatingdomain overlap, but a hockey player does not belong on a soccer field\.","overlap\_dimensions": \["occupational\_domain", "physical\_actions"\],"visual\_markers": "Hockey jersey, bulky shoulder and shin pads, hockey helmet withvisor, ice skates\."\}\},"sentences": \{"plausible\_easy": "A goalie is catching a soccer ball with his hands on the soccer field\.","plausible\_hard": "A referee is catching a soccer ball with his hands on the soccer field\.","implausible\_easy": "A ballerina is catching a soccer ball with his hands on the soccerfield\.","implausible\_hard": "A hockey player is catching a soccer ball with his hands on thesoccer field\."\}\}\-\-\-OUTPUT FORMAT \(STRICT\)\-\-\-Return ONLY the following JSON object\. No markdown fences\. No text outside the JSON\.\{"reasoning": "<Use this field as a scratchpad\. Work through the entire 6\-step PROCEDURE here:design the scene, verify scene typicality with the ’someone’ test, select all four agentswith thematic fit reasoning \(including the core\-job verification and removal test forplausible\_easy, and the multi\-dimensional overlap audit plus negative test forimplausible\_hard\), check visual identifiability and pairwise distinctiveness, composesentences, and run all validation checks\. Write freely \-\- this is your workspace to thinkthrough the problem before committing to the output below\.\>","scene": \{"verb\_form": "<chosen inflected form of the root verb\>","patient": "<patient phrase\>","instrument": "<instrument phrase starting with ’with’\>","location": "<location phrase starting with a preposition\>"\},"agents": \{"plausible\_easy": \{"agent": "<agent phrase\>","thematic\_fit": "HIGH","rationale": "<1\-2 sentences: why this agent is prototypical for this event\>","visual\_markers": "<brief description of this agent’s distinctive professional appearance\>"\},"plausible\_hard": \{"agent": "<agent phrase\>","thematic\_fit": "MODERATE","rationale": "<1\-2 sentences: why this agent is plausible but atypical\>","visual\_markers": "<brief description of this agent’s distinctive professional appearance\>"\},"implausible\_easy": \{"agent": "<agent phrase\>","thematic\_fit": "VERY LOW","rationale": "<1\-2 sentences: why this agent is from a completely unrelated domain\>","visual\_markers": "<brief description of this agent’s distinctive professional appearance\>"\},"implausible\_hard": \{Figure 12:RB Generation Prompt \(continued from previous page\)\.RB Generation Prompt \(Part 9/9\)"agent": "<agent phrase\>","thematic\_fit": "LOW","rationale": "<1\-2 sentences: why this agent is from a related but wrong domain\>","overlap\_dimensions": \["<list the 2\+ overlapping dimensions from: occupational\_domain,typical\_setting, tools\_instruments, patient\_object\_type, physical\_actions\>"\],"visual\_markers": "<brief description of this agent’s distinctive professional appearance\>"\}\},"sentences": \{"plausible\_easy": "<full sentence\>","plausible\_hard": "<full sentence\>","implausible\_easy": "<full sentence\>","implausible\_hard": "<full sentence\>"\}\}\-\-\-VERB INPUT \(process this verb now\)\-\-\-The root verb for this generation is: "\{verb\}"Follow the PROCEDURE above for this verb\. Return ONLY the JSON object\. No other text\.Figure 13:RB Generation Prompt \(continued from previous page\)\.### B\.1Design Decisions

The five\-step procedure embeds several design decisions, each addressing a specific failure mode observed during prompt iteration\.

##### Scene typicality \(Step 1\)\.

LLMs frequently anchor on a location \(for example, a library\) and then force the verb into that location, producing scenes that are setting\-familiar but action\-atypical\. The result is a gradient where no agent is genuinely plausible because the underlying scene is itself implausible\. To prevent this, Step 1 includes a scene\-typicality test: substitutesomeonefor the agent and ask whether the resulting sentence describes a routine event\. If not, the scene must be redesigned before agent selection begins\.

##### Core\-job verification for PE \(Step 2\)\.

LLMs default to setting\-familiar agents \(a groundskeeper on a soccer field\) when the action\-prototypical agent \(a goalie\) sits one inferential step away, producing PE agents who are present in the location but for whom the action is not part of their job\. We therefore require the PE agent to satisfy a stronger standard than location\-presence: the action must be part of the agent’s primary professional duties\. The prompt operationalises this via a removal test, asking whether removing this agent from the location would stop\[verb\]​\[patient\]\[\\textit\{verb\}\]\[\\textit\{patient\}\]from happening there\.

##### Multi\-dimensional IH overlap \(Step 3\)\.

IH is the hardest of the four conditions to operationalise: the agent must be a near\-miss to PE rather than an obvious mismatch\. We give the model a principled selection rule by requiring overlap with PE on at least two of five dimensions \(occupational domain, setting, tools, patient type, physical actions\)\. We selected the two\-of\-five cutoff by inspecting seed examples during prompt development\. It operationalises a near\-miss for this instantiation and can be tuned for other datasets, target slots, or domains; we do not treat it as a universal semantic boundary\. The four conditions are first defined by their relation to the scene \(PE = typical doer; PH = present but not typical; IH = related domain but not present; IE = unrelated\); the≥2\\geq 2rule is an additional constraint within the IH cell, so a high\-overlap profession that is action\-prototypical for the scene falls in PE, not IH\. Forcatch a soccer ball\(PE = goalie, PH = referee\), a hockey player satisfies IH \(shares sport domain and defensive actions, but does not belong on a soccer field\); a fisherman, whose only link to the scene is the verbcatch, collapses to IE\.

##### Visual identifiability \(Step 4\)\.

STRIVE outputs are intended for downstream image\-based assessment, where the final stimulus is a photograph with no caption or scene context\. An agent whose profession cannot be identified from appearance alone may satisfy its intended plausibility condition without providing an experimental signal: the patient sees a generic person rather than a specific role\. We therefore require each agent to have a profession recognisable from a photograph against a plain white background, ruling out professions whose visual presentation collapses to ordinary civilian clothing \(software engineer, freelance writer, accountant\)\. Pairwise distinctiveness across all six combinations is verified at the same step\.

##### Pre\-output reasoning trace \(cross\-cutting\)\.

Per\-stimulus rationales capture the model’s reasoning about individual agent choices but do not give the model space to design the scene holistically, audit overlap dimensions across agents, or resolve visual conflicts before committing to specific outputs\. We therefore include a free\-form scratchpad at the top of the JSON schema where the entire procedure is worked through before any final commitments\. This trace is what F1 in §[1](https://arxiv.org/html/2608.04567#S1)identifies as the central lever: removing it cuts GPT\-5\.1’s RB GOLD rate from 28\.3% to 16\.7%, even with per\-stimulus rationales preserved\.

## Appendix CR2V Generator Prompt

Since the prompts are very long and occupy many pages, we show the portion of the prompt which dictates the verbalized sampling\(Zhang et al\.,[2026](https://arxiv.org/html/2608.04567#bib.bib42)\)in Figure[14](https://arxiv.org/html/2608.04567#A3.F14)\. The complete prompt would be released upon publication along with the code\.

R2V Generation Prompt \(Partial\)Distribution guidance \-\- aim for approximately:\- 3\-4 candidates with probability 0\.70\-1\.00 \(prototypical agents, HIGH fit\)\- 5\-6 candidates with probability 0\.35\-0\.69 \(plausible but atypical, MODERATE fit\)\- 5\-6 candidates with probability 0\.10\-0\.34 \(related domain but wrong, LOW fit\)\- 3\-4 candidates with probability 0\.00\-0\.09 \(completely unrelated, VERY LOW fit\)IMPORTANT DIVERSITY RULES FOR CANDIDATE GENERATION:\- Do NOT cluster candidates within one professional field\. Spread across diverse domains:medical, sports, trades, arts, military, food service, education, emergency services,religious, legal, scientific, transportation, agriculture, etc\.\- Every candidate MUST be visually identifiable from a photograph against a plain whitebackground \-\- no scene context, no caption, no other people\. Do NOT include any professionthat looks like a generic civilian\.\- Think carefully about the probability estimates\. A probability of 0\.50 means a humanwould be genuinely uncertain whether this agent fits the event\. A probability of 0\.25means a human would lean toward "no" but see why someone might think otherwise\.STEP 3 \-\- SELECT THE OPTIMAL FOUR\-AGENT COMBINATIONFrom your 20 candidates, select exactly four agents \-\- one for each plausibility condition\.Optimize for GRADIENT SPACING: the four selected agents should be well\-separated in theirthematic fit probabilities, not clustered together\.Selection criteria:a\) plausible\_easy: Select the candidate with the highest thematic fit probability\.Target range: 0\.80\-1\.00\. This must be THE prototypical agent for this event\.CORE\-JOB VERIFICATION: Confirm that \[verb\] \[patient\] is part of this agent’s PRIMARYprofessional duties \-\- not merely something they could do, but something they are PAID orEXPECTED to do routinely\. Ask: "If this agent were removed from \[location\], would \[verb\]\[patient\] stop happening there?" If the answer is no, this agent is setting\-prototypical\(familiar with the location\) but not action\-prototypical \(the action is not their core job\)\.A setting\-prototypical agent belongs in plausible\_hard, not plausible\_easy\.Example: A groundskeeper on a soccer field is setting\-prototypical, but "catchinga soccer ball" is not a groundskeeper’s job\. The groundskeeper belongs in plausible\_hardat best\.b\) plausible\_hard: Select a candidate with probability in the 0\.40\-0\.65 range\.This agent must be naturally present or could naturally be present in the location,and the action must be secondary or occasional for them, not their primary role\.c\) implausible\_hard: Select a candidate with probability in the 0\.15\-0\.35 range\.This agent MUST share overlap with the plausible\_easy agent on AT LEAST TWO of thefive overlap dimensions \(occupational domain, typical setting, tools/instruments,patient/object type, physical actions\)\. Apply the NEGATIVE TEST: removing the sharedverb, is there STILL a reason to associate this agent with the scene?d\) implausible\_easy: Select the candidate with the lowest thematic fit probability\.Target range: 0\.00\-0\.05\. This must be from a COMPLETELY UNRELATED domain\.STEP 4 \-\- CHECK COMBINATION QUALITYVerify the selected four agents satisfy ALL of the following:\- All four are living humans with identifiable professions? YES/NO\- All four are visually identifiable from a photograph against a plain white background? YES/NO\- All four are visually distinguishable from each other in a photograph? YES/NO\- Scene typicality: Does "Someone \[verb\] \[patient\] \[instrument\] \[location\]" sound routine?YES/NO\- plausible\_easy: Would a person INSTANTLY say "yes, that’s normal"? YES/NO\- plausible\_easy CORE\-JOB: Is \[verb\] \[patient\] part of this agent’s primary job description?YES/NO\- plausible\_easy: Is the agent action\-prototypical \(not just setting\-familiar\)? YES/NO\- plausible\_hard: Would a person PAUSE then say "yes, I suppose that could happen"? YES/NO\- implausible\_hard: Would a person HESITATE then say "no, that doesn’t fit"? YES/NO\- implausible\_hard: Does this agent share at least TWO overlap dimensions with plausible\_easy?YES/NO\- implausible\_hard: Removing the shared verb, is there still a reason to associate this agentFigure 14:R2V Generation Prompt \(Partial\)\.
## Appendix DIterative Refine Prompt

The common refinement prompt contains the same task definitions and validation rubric as RB, with two additions: an instruction to rectify rather than regenerate \(Figure[15](https://arxiv.org/html/2608.04567#A4.F15)\) and a feedback\-history placeholder \(Figure[16](https://arxiv.org/html/2608.04567#A4.F16)\)\. At each iteration, this field contains accumulated compact, instance\-specific diagnoses rather than the evaluator rubric or full output\. The previously generated output is supplied separately\. An abridged diagnosis has the following structure:

AGENTS:plausible\_hard:agent: "A librarian"verdict: INCORRECTactual\_condition: implausible\_easyissues: Librarian has no natural presence in a homeliving room\.implausible\_hard:agent: "A carpenter"verdict: CORRECToverlap\_count: 2overlap\_dimensions:\[typical\_setting, patient\_object\_type\]GRADIENT:correctly\_ordered: NOSCENE:verdict: NEEDS\_REDESIGNThe overlap fields report findings for a specific generated agent; they do not restate the IH definition\.

The complete prompt would be released upon publication along with the code\.

Iterative Prompt \(Instruction\)=== SYSTEM PROMPT ===You are an expert psycholinguist who understands event cognition, thematic roles, and typicalityeffects\.Your task is to revise a previously generated set of event sentences based on evaluator feedback\.You carefully analyze each criticism and make targeted fixes while preserving what already works\.You reason step\-by\-step before producing output, and you return only valid JSON\.=== USER TEMPLATE ===TASKYou will be given a verb \(root form\), a previous generation attempt, and evaluator feedbackat the end of this prompt\. Your job is to revise the previous output to fix ALL identifiedissues while preserving anything that was correct\.Each sentence describes a concrete, observable scene that could be depicted in a singlephotograph\.IMPORTANT: These sentences will later be used to generate images\. A human participant will thenview each image and judge whether the depicted event is plausible or not\. Therefore, the fouragents must be visually distinguishable from one another in a photograph \-\- each agent’sprofessionshould be identifiable from their appearance alone \(clothing, gear, accessories, physique, etc\.\)\.\-\-\-EXPERIMENTAL CONTEXT\-\-\-These sentences will be converted into images shown to brain injury patients who must make quickplausible/implausible judgments\. The patients see ONLY the image \-\- no text, no captions, nolabels\.Therefore:\(1\) Each agent’s profession must be INSTANTLY recognizable from appearance alone\.\(2\) The plausibility distinction must be visually obvious and must not depend on subtlereasoning\.\(3\) The scene must be concrete and unambiguous\.If a profession cannot be identified from a photograph without a caption, do NOT use thatprofession\.\-\-\-SENTENCE STRUCTURE\-\-\-Every sentence MUST follow this exact slot order:\[Agent\] \[verb\] \[patient\] \[instrument\] \[location\]Slot definitions:\- Agent: A living human identified by profession or social role \(the doer of the action\)\.\- Verb: A chosen inflected form of the root verb \(e\.g\., "is catching", "catches", "caught"\)\.Use the SAME inflected form in ALL four sentences\.\- Patient: The object or entity being acted upon\.\- Instrument: The tool, body part, or means used to perform the action \(phrased with "with"\)\.\- Location: Where the event takes place \(phrased with "in", "on", "at", etc\.\)\.Example: "A goalie is catching a soccer ball with his hands on the soccer field\."Agent = "A goalie"Verb = "is catching"Patient = "a soccer ball"Instrument = "with his hands"Location = "on the soccer field"\-\-\-CORE CONCEPT: THEMATIC FITFigure 15:Iterative Prompt \(Instruction\)\.Iterative Prompt \(Feedback placeholder\)"agent": "<agent phrase\>","thematic\_fit": "VERY LOW","rationale": "<1\-2 sentences: why this agent is from a completely unrelated domain\>","visual\_markers": "<brief description of this agent’s distinctive professional appearance\>","changed": false\},"implausible\_hard": \{"agent": "<agent phrase\>","thematic\_fit": "LOW","rationale": "<1\-2 sentences: why this agent is from a related but wrong domain\>","overlap\_dimensions": \["<list the 2\+ overlapping dimensions from: occupational\_domain,typical\_setting, tools\_instruments, patient\_object\_type, physical\_actions\>"\],"visual\_markers": "<brief description of this agent’s distinctive professional appearance\>","changed": true\}\},"sentences": \{"plausible\_easy": "<full sentence\>","plausible\_hard": "<full sentence\>","implausible\_easy": "<full sentence\>","implausible\_hard": "<full sentence\>"\}\}Note on the ’changed’ field: it is true for any agent that differs from the correspondingagent in the previous output, and false for any agent preserved unchanged\. The true/falsevalues shown above are illustrative only \-\- set them to reflect your actual changes\.\-\-\-REVISION INPUT \(process this now\)\-\-\-Root verb: "\{verb\}"PREVIOUS OUTPUT:\{previous\_output\}EVALUATOR FEEDBACK:\{evaluator\_feedback\}Follow the REVISION PROCEDURE \(Steps R1\-R4\) then the GENERATION PROCEDURE \(Steps 1\-6\)\.Return ONLY the JSON object\. No other text\.Figure 16:Iterative Prompt \(Feedback placeholder\)\.
## Appendix EEvaluator Prompt

The STRIVE evaluator assesses each generated set across five stages in fixed order:

1. 1\.Agent constraint validity\.Checks that all four agents are living humans with identifiable professions, that the verb form, patient, instrument, and location are identical across sentences, and that implausibility stems from role\-action misfit rather than physical incapability\.
2. 2\.Per\-condition plausibility correctness\.For each condition, the evaluator applies the adversarial counter\-argument check \(§[3\.2](https://arxiv.org/html/2608.04567#S3.SS2)\) and, for IH, the overlap audit\. Verdicts arecorrect,borderline, orincorrect;incorrectverdicts specify which condition the agent actually belongs in\.
3. 3\.Gradient ordering\.Checks that the overall plausibility ordering PE\>\>PH≫\\ggIH\>\>IE is maintained and that no adjacent conditions are collapsed\.
4. 4\.Visual evaluation\.First rates each agent’s profession identifiability \(identifiable/ambiguous/unidentifiable\); then checks all six pairwise agent combinations for visual distinguishability\. Verdict:all\_distinct,mostly\_distinct, orproblematic\.
5. 5\.Scene assessment\.Checks instrument correctness, prototype clarity, imageability, and gradient supportability\. For eachincorrectorborderlinecondition, the evaluator states whether a viable replacement agent exists for the current scene; if not, the scene itself is the root cause of failure\. Verdict:good,agent\_fixable, orneeds\_redesign\.

The full evaluator prompt is shown in Figure[17](https://arxiv.org/html/2608.04567#A5.F17)\.

Evaluator Prompt \(Part 1/8\)=== SYSTEM PROMPT ===You are an expert psycholinguist and cognitive scientist who evaluates event plausibilitysentences\.You assess whether generated sentences correctly represent four levels of plausibility based onthematic fit\.You also assess whether the four agents are visually distinguishable in photographs, since thesesentences will be used to generate images for a human judgment experiment\.You are rigorous, precise, and catch subtle errors that might seem acceptable at first glance\.When in doubt, be strict\. A false CORRECT verdict leads to an invalid experimental stimulus thatcould waste a patient session\. A false INCORRECT verdict only triggers one more refinementiteration\.You return only valid JSON\.=== USER TEMPLATE ===TASKYou will be given a verb \(root form\) and a set of four generated event sentences at the end ofthis prompt\. The sentences vary only in their AGENT \(a human professional or social role\) whilekeeping the verb form, patient, instrument, and location identical\. Your job is to judge whethereach sentence is correctly assigned to its plausibility condition, whether the four agents arevisually distinguishable and identifiable, and whether the scene itself can support the fullplausibility gradient\.\-\-\-EXPERIMENTAL CONTEXT\-\-\-These sentences will be converted into images shown to brain injury patients who must make quickplausible/implausible judgments\. The patients see ONLY the image \-\- no text, no captions, nolabels\.Therefore:\(1\) Each agent’s profession must be INSTANTLY recognizable from appearance alone\.\(2\) The plausibility distinction must be visually obvious and must not depend on subtlereasoning\.\(3\) The scene must be concrete and unambiguous\.Keep this context in mind throughout your evaluation\. An agent whose profession cannot beidentifiedfrom a photograph is a critical failure, even if the thematic fit is correct\.\-\-\-BACKGROUND: SENTENCE STRUCTURE AND THEMATIC FIT\-\-\-Each sentence follows the structure: \[Agent\] \[verb\] \[patient\] \[instrument\] \[location\]\.Thematic fit is how naturally a particular AGENT fits as the doer of the described action\(verb \+ patient \+ instrument \+ location\)\. The four conditions form a gradient:plausible\_easy \-\-\- plausible\_hard \-\-\- implausible\_hard \-\-\- implausible\_easy\(clear yes\) \(unsure yes\) \(unsure no\) \(clear no\)HIGH fit MODERATE fit LOW fit VERY LOW fit\-\-\-CONDITION DEFINITIONS \(use these to judge each agent\)\-\-\-1\) PLAUSIBLE\_EASY \-\- "clear yes" \-\- HIGH thematic fitThe agent is the prototypical, most expected professional for this action in this setting\.The event is physically possible, reasonable, and typical\.A human would INSTANTLY accept this as normal\.Test: Is this THE default professional you would expect in this scene?Figure 17:Evaluator Prompt\. Continued in Figure[18](https://arxiv.org/html/2608.04567#A5.F18)\.Evaluator Prompt \(Part 2/8\)CORE\-JOB TEST \(mandatory for plausible\_easy\):Is \[verb\] \[patient\] part of this agent’s PRIMARY professional duties \-\- something they arepaid or expected to do routinely? Or is the agent merely familiar with \[location\]?Ask: "If this agent were removed from \[location\], would \[verb\] \[patient\] stop happening there?"If the answer is no, the agent’s connection is to the SETTING, not the ACTION\. Asetting\-prototypical agent belongs in plausible\_hard, not plausible\_easy\.Example: A groundskeeper on a soccer field is setting\-prototypical, but "catchinga soccer ball" is not a groundskeeper’s job \-\- this agent is plausible\_hard at best\.2\) PLAUSIBLE\_HARD \-\- "unsure yes" \-\- MODERATE thematic fitThe agent could plausibly perform this action in this setting, but it is NOT their primaryrole\.They must be NATURALLY PRESENT or COULD NATURALLY BE PRESENT in the location\.A human would PAUSE but ultimately accept it\.Test: Can you easily imagine a specific, realistic scenario where this professional does this?CRITICAL: If the agent has no natural reason to be at the location, they belong in animplausible category, not here\.3\) IMPLAUSIBLE\_HARD \-\- "unsure no" \-\- LOW thematic fitThe agent is from a RELATED but WRONG domain\. There IS semantic overlap with theplausible\_easy agent, but the agent does NOT belong in this specific scene\.A human would HESITATE then reject it\.CRITICAL \-\- MULTI\-DIMENSIONAL OVERLAP REQUIREMENT:The implausible\_hard agent must share overlap with the plausible\_easy agent on AT LEAST TWOof the following five dimensions:\(a\) Occupational domain \(e\.g\., both belong to the same broad professional field\)\(b\) Typical work setting/environment \(e\.g\., both typically work in similar environments\)\(c\) Tools or instruments routinely used \(e\.g\., both routinely use similar types of tools\)\(d\) Type of patient/object acted upon \(e\.g\., both act on similar types of objects orentities\)\(e\) Physical actions routinely performed \(e\.g\., both routinely perform similar physicalactions\)A single dimension of overlap \(e\.g\., both use sharp tools\) is NOT SUFFICIENT\. If the onlyconnection is through the verb itself, the agent belongs in implausible\_easy, not here\.NEGATIVE TEST: Remove the shared verb from consideration\. Is there STILL a reason to associatethis agent with the scene? If not, the overlap is too thin\.Test: Does this agent share at least two of the above dimensions with plausible\_easywhile still being wrong for THIS specific scene?CRITICAL: If the agent could realistically perform this action in this setting, they aretoo plausible for this category\. If they have zero or only one dimension of domain connection,they belong in implausible\_easy\.4\) IMPLAUSIBLE\_EASY \-\- "clear no" \-\- VERY LOW thematic fitThe agent is from a COMPLETELY UNRELATED domain with zero semantic connection to anypart of the event\.A human would INSTANTLY reject it\.Test: Is this profession maximally distant from the scene?\-\-\-AGENT VALIDITYFigure 18:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 3/8\)\-\-\-Valid agents include living humans identified by profession or social role\.Age\-based social roles with clear visual markers are acceptable \(e\.g\., "a schoolchild" inschool uniform, "a retiree" with gray hair and glasses\) when they carry distinct roleexpectations\.Plausibility must stem from ROLE\-ACTION MISMATCH \(whether this role typically performs thisaction\),NOT from physical incapability\. Reject agents whose implausibility depends on being physicallyunable rather than role\-inappropriate\.Forbidden: animals, objects, fictional beings, statues, robots, corpses, body parts,descriptions of states, pure physical descriptors without a role\.\-\-\-EVALUATION CRITERIA\-\-\-You will evaluate six aspects IN THIS ORDER\. The order matters because later assessmentsdepend on earlier ones\.A\) AGENT CONSTRAINTS \(check first\)A1\) Is every agent a living human with an identifiable profession or social role?A2\) Are there any forbidden agents \(animals, objects, statues, robots, states, pure physicaldescriptors\)?A3\) Are the verb form, patient, instrument, and location identical across all four sentences?A4\) Does only the agent change?A5\) Does plausibility stem from role\-action mismatch \(not physical incapability\)?B\) PER\-CONDITION CORRECTNESS \(most important\)For each of the four conditions, judge whether the assigned agent TRULY belongs in thatplausibility level\. Consider:\- Would this agent realistically be in this location?\- Is this action part of \(or adjacent to\) their professional duties?\- How strong is the semantic overlap with the plausible\_easy agent?\- Could this agent be confused with an adjacent condition?ADVERSARIAL CHECK \(mandatory for every condition\):Before issuing your verdict, you MUST articulate:\(i\) The strongest argument for why this agent belongs ONE LEVEL HIGHER on the gradient\(more plausible than assigned\)\.\(ii\) The strongest argument for why this agent belongs ONE LEVEL LOWER on the gradient\(less plausible than assigned\)\.Only assign CORRECT if BOTH counter\-arguments are clearly weaker than the assigned placement\.If either counter\-argument is compelling, assign BORDERLINE or INCORRECT\.Record these counter\-arguments in the output\.For PLAUSIBLE\_EASY specifically, you must also perform the TYPICALITY AUDIT:\(i\) CORE\-JOB: Is \[verb\] \[patient\] in this agent’s job description? Or are they merelyfamiliar with \[location\]? Setting\-familiarity alone = plausible\_hard, not PE\.\(ii\) REMOVAL TEST: If this agent were removed from the scene, would the action stop?If no, they are not the prototypical doer\.Record the typicality audit results in the output\.For IMPLAUSIBLE\_HARD specifically, you must also perform the OVERLAP AUDIT:Evaluate each of the five overlap dimensions against the plausible\_easy agent:\(a\) Occupational domain: Do both agents belong to the same broad professional field?\(b\) Typical setting: Do both agents typically work in similar environments?\(c\) Tools/instruments: Do both agents routinely use similar tools?\(d\) Patient/object type: Do both agents act on similar types of objects or entities?\(e\) Physical actions: Do both agents routinely perform similar physical actions?Count the number of overlapping dimensions\. If fewer than 2, the agent cannot beimplausible\_hard \-\- it should be implausible\_easy\.Also apply the NEGATIVE TEST: Remove the shared verb from consideration\. Is there STILLFigure 19:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 4/8\)a reason to associate this agent with the scene? If not, the overlap is too thin\.Assign one of these verdicts per condition:\- CORRECT: The agent clearly belongs in this condition\.\- BORDERLINE: The agent is debatable but not clearly wrong\.\- INCORRECT: The agent belongs in a different condition\.If INCORRECT, specify which condition the agent actually belongs in\.C\) GRADIENT ORDERINGIs the overall ordering maintained?plausible\_easy \> plausible\_hard \>\> implausible\_hard \> implausible\_easyAre any adjacent conditions swapped or collapsed \(too similar to distinguish\)?D\) VISUAL EVALUATIONThese sentences will be used to generate images for a human judgment experiment\.A participant will view each image and must be able to identify the agent’s professionfrom their appearance alone\. Evaluate in this order:D0\) PROFESSION IDENTIFIABILITY \(evaluate FIRST, before pairwise checks\):For each agent, ask: "If a photograph showed ONLY this person in their typicalprofessional attire against a plain white background \-\- no scene context, no caption,no other people \-\- could a viewer correctly guess their profession or social role?"A profession is identifiable if it has:\- A dedicated uniform or standardized attire distinct from civilian clothing, OR\- Profession\-specific gear or equipment worn/carried as part of the role, OR\- A highly distinctive physical presentation associated with the role\.Rate each agent:IDENTIFIABLE: Uniform/gear makes profession obvious to a naive viewer\.AMBIGUOUS: Could be guessed with effort but easily confused with other roles\.UNIDENTIFIABLE: Looks like a generic civilian in everyday clothing\.Any agent rated UNIDENTIFIABLE is a critical failure for the experiment\.D1\) Does each agent’s profession have a recognizable visual identity?A profession has a recognizable visual identity if a naive viewer could identify it fromappearance alone \-\- through a distinctive uniform, professional gear, role\-specific attire,or iconic physical presentation\. Professions that rely on generic civilian clothing\(business casual, jeans, everyday wear\) are NOT visually identifiable\.D2\) Are all four agents visually distinguishable from EACH OTHER?Check all six pairwise combinations:\- plausible\_easy vs plausible\_hard\- plausible\_easy vs implausible\_hard\- plausible\_easy vs implausible\_easy\- plausible\_hard vs implausible\_hard\- plausible\_hard vs implausible\_easy\- implausible\_hard vs implausible\_easyFor each pair, ask: Would these two look noticeably different in a photograph basedon their typical professional attire, gear, accessories, or physical presentation?D3\) Flag any pair that would appear nearly identical in a photo\.Two agents are indistinguishable if they share the same general workwear, attire type,or lack of distinctive professional markers\. Agents from the same broad domain mustdiffer in at least one strong visual cue \(uniform type, headgear, carried tools, etc\.\)\.Assign a visual verdict:\- ALL\_DISTINCT: All six pairs are visually distinguishable AND all agents are IDENTIFIABLE\.\- MOSTLY\_DISTINCT: 4\-5 pairs are distinguishable, 1\-2 pairs are marginal, no agent isUNIDENTIFIABLE\.\- PROBLEMATIC: 3\+ pairs are visually indistinguishable, OR any agent is UNIDENTIFIABLE\.Figure 20:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 5/8\)E\) COMPREHENSIVE SCENE ASSESSMENT \(evaluate LAST, after all above\)This is the final and most holistic check\. Now that you have evaluated all four agents,the gradient, and visual distinctiveness, assess the SCENE itself\. A good scene mustsatisfy ALL of the following:E1\) INSTRUMENT CORRECTNESS: Is the instrument appropriate for the verb?\(e\.g\., you erase with an eraser, not a marker; you cut with a knife, not a spoon\)E2\) CLEAR PROTOTYPE: Does the scene have a profession for whom this action is THE core job?If no profession is strongly prototypical, the scene may need redesign\.E3\) IMAGEABILITY: Can the scene be captured in a single still photograph?E4\) GRADIENT SUPPORTABILITY \(most critical\): Can this scene support all four plausibilitylevels with distinct, correctly\-placed agents?This is the key question\. A scene FAILS gradient supportability if:\- The domain is so narrow that any related professional is also plausible\(e\.g\., "breaking a board in a martial arts dojo" \-\- all combat sport professionalsare plausible, leaving no room for implausible\_hard\)\.\- The domain is so broad that it is hard to find an implausible\_hard agent withgenuine multi\-dimensional semantic overlap \(no "wait, maybe\.\.\." moment\)\.\- Any condition was marked INCORRECT and you cannot think of a viable replacementagent that would correctly fill that condition within this scene\.To assess this, ask yourself for each INCORRECT or BORDERLINE condition:"Can I think of at least one profession that would correctly fill this conditionin this scene?" If yes, the scene supports the gradient \(it’s an agent selectionproblem\)\. If no, the scene itself is too narrow or too broad\.Scene verdict:\- GOOD: All four checks pass\. The scene supports the full gradient\.\- AGENT\_FIXABLE: The scene is sound but some agents are wrong\. Replacing agents can fix it\.\- NEEDS\_REDESIGN: The scene itself cannot support the full gradient, or the instrument iswrong, or there is no clear prototype\. The entire scene should be redesigned\.\-\-\-EVALUATION PROCEDURE\-\-\-STEP 1: Check all agent constraints and consistency\.STEP 2: For each condition, carefully reason about whether the agent truly belongs there\.Compare each agent against ALL four condition definitions, not just the one it wasassigned to\.Ask: "Where would I place this agent if I were assigning from scratch?"For EACH condition, perform the ADVERSARIAL CHECK: articulate the strongest argument forone level higher and one level lower before issuing your verdict\.For PLAUSIBLE\_EASY, also perform the TYPICALITY AUDIT: verify core\-job and removal test\.For IMPLAUSIBLE\_HARD, also perform the OVERLAP AUDIT: evaluate all five dimensions andapply the negative test\.STEP 3: Check the overall gradient ordering\.STEP 4: Assess visual evaluation of all four agents\.FIRST, rate each agent’s profession identifiability\(IDENTIFIABLE/AMBIGUOUS/UNIDENTIFIABLE\)\.THEN, for each agent, describe their typical professional appearance\.THEN, check all six pairwise combinations for visual similarity\.STEP 5: Perform comprehensive scene assessment\. Using your findings from Steps 2\-4,determine whether the scene itself is the root cause of any failures\.For each INCORRECT condition, ask: "Can I think of a profession that WOULD correctlyfill this condition in this scene?" If not, the scene needs redesign\.STEP 6: Produce your evaluation in the OUTPUT FORMAT below\.Figure 21:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 6/8\)\-\-\-OUTPUT FORMAT \(STRICT\)\-\-\-Return ONLY the following JSON object\. No markdown fences\. No text outside the JSON\.\{"constraint\_checks": \{"all\_human\_professionals": true/false,"consistency\_maintained": true/false,"role\_action\_mismatch\_valid": true/false,"constraint\_notes": "<brief note or ’All constraints met’\>"\},"condition\_evaluations": \{"plausible\_easy": \{"agent": "<the agent that was assigned\>","verdict": "CORRECT\|BORDERLINE\|INCORRECT","actual\_condition": "<if INCORRECT, which condition this agent actually belongs in;otherwise same as assigned\>","reasoning": "<1\-3 sentences explaining your judgment\>","counter\_arguments": \{"argument\_for\_higher": "<N/A for plausible\_easy since it is the highest level\>","argument\_for\_lower": "<strongest argument that this agent actually belongs inplausible\_hard or lower\>"\},"typicality\_audit": \{"core\_job\_test": "<Is \[verb\] \[patient\] part of this agent’s primary job duties?YES/NO with brief reasoning\>","removal\_test": "<If this agent were removed from \[location\], would \[verb\]\[patient\] stop happening? YES/NO with brief reasoning\>"\}\},"plausible\_hard": \{"agent": "<the agent that was assigned\>","verdict": "CORRECT\|BORDERLINE\|INCORRECT","actual\_condition": "<if INCORRECT, which condition this agent actually belongs in;otherwise same as assigned\>","reasoning": "<1\-3 sentences explaining your judgment\>","counter\_arguments": \{"argument\_for\_higher": "<strongest argument that this agent actually belongs inplausible\_easy\>","argument\_for\_lower": "<strongest argument that this agent actually belongs inimplausible\_hard or lower\>"\}\},"implausible\_hard": \{"agent": "<the agent that was assigned\>","verdict": "CORRECT\|BORDERLINE\|INCORRECT","actual\_condition": "<if INCORRECT, which condition this agent actually belongs in;otherwise same as assigned\>","reasoning": "<1\-3 sentences explaining your judgment\>","counter\_arguments": \{"argument\_for\_higher": "<strongest argument that this agent actually belongs inplausible\_hard or higher\>","argument\_for\_lower": "<strongest argument that this agent actually belongs inimplausible\_easy\>"\},"overlap\_audit": \{"occupational\_domain": \{"overlaps": true/false, "reasoning": "<brief\>"\},"typical\_setting": \{"overlaps": true/false, "reasoning": "<brief\>"\},"tools\_instruments": \{"overlaps": true/false, "reasoning": "<brief\>"\},"patient\_object\_type": \{"overlaps": true/false, "reasoning": "<brief\>"\},Figure 22:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 7/8\)"physical\_actions": \{"overlaps": true/false, "reasoning": "<brief\>"\},"overlap\_count": <number of true dimensions\>,"sufficient": true/false,"negative\_test": "<Does association persist after removing the shared verb? Explain\.\>"\}\},"implausible\_easy": \{"agent": "<the agent that was assigned\>","verdict": "CORRECT\|BORDERLINE\|INCORRECT","actual\_condition": "<if INCORRECT, which condition this agent actually belongs in;otherwise same as assigned\>","reasoning": "<1\-3 sentences explaining your judgment\>","counter\_arguments": \{"argument\_for\_higher": "<strongest argument that this agent actually belongs inimplausible\_hard or higher\>","argument\_for\_lower": "<N/A for implausible\_easy since it is the lowest level\>"\}\}\},"gradient\_ordering": \{"correctly\_ordered": true/false,"swapped\_pairs": "<describe any swaps, or ’None’\>","collapsed\_pairs": "<describe any conditions too similar to distinguish, or ’None’\>"\},"visual\_evaluation": \{"identifiability": \{"plausible\_easy": \{"rating": "IDENTIFIABLE\|AMBIGUOUS\|UNIDENTIFIABLE", "reasoning":"<brief\>"\},"plausible\_hard": \{"rating": "IDENTIFIABLE\|AMBIGUOUS\|UNIDENTIFIABLE", "reasoning":"<brief\>"\},"implausible\_hard": \{"rating": "IDENTIFIABLE\|AMBIGUOUS\|UNIDENTIFIABLE", "reasoning":"<brief\>"\},"implausible\_easy": \{"rating": "IDENTIFIABLE\|AMBIGUOUS\|UNIDENTIFIABLE", "reasoning":"<brief\>"\}\},"agent\_appearances": \{"plausible\_easy": "<typical professional appearance of this agent\>","plausible\_hard": "<typical professional appearance of this agent\>","implausible\_hard": "<typical professional appearance of this agent\>","implausible\_easy": "<typical professional appearance of this agent\>"\},"pairwise\_checks": \{"pe\_vs\_ph": \{"distinguishable": true/false, "note": "<brief reason\>"\},"pe\_vs\_ih": \{"distinguishable": true/false, "note": "<brief reason\>"\},"pe\_vs\_ie": \{"distinguishable": true/false, "note": "<brief reason\>"\},"ph\_vs\_ih": \{"distinguishable": true/false, "note": "<brief reason\>"\},"ph\_vs\_ie": \{"distinguishable": true/false, "note": "<brief reason\>"\},"ih\_vs\_ie": \{"distinguishable": true/false, "note": "<brief reason\>"\}\},"visual\_verdict": "ALL\_DISTINCT\|MOSTLY\_DISTINCT\|PROBLEMATIC","indistinguishable\_pairs": "<list any pairs that look too similar, or ’None’\>","unidentifiable\_agents": "<list any agents rated UNIDENTIFIABLE, or ’None’\>"\},"scene\_assessment": \{"instrument\_correct": true/false,"has\_clear\_prototype": true/false,"imageable": true/false,"gradient\_supportable": true/false,"gradient\_support\_reasoning": "<For each INCORRECT/BORDERLINE condition, state whether youcan think of a viable replacement agent for this scene\. If you cannot, explain why thescene is too narrow or too broad\.\>","scene\_verdict": "GOOD\|AGENT\_FIXABLE\|NEEDS\_REDESIGN",Figure 23:Evaluator Prompt \(continued from previous page\)\.Evaluator Prompt \(Part 8/8\)"scene\_notes": "<brief overall scene assessment summarizing all findings\>"\},"overall": \{"score": "<number of CORRECT conditions out of 4, e\.g\., ’3/4’\>","borderline\_count": <number of BORDERLINE conditions\>,"scene\_verdict": "GOOD\|AGENT\_FIXABLE\|NEEDS\_REDESIGN","visual\_verdict": "ALL\_DISTINCT\|MOSTLY\_DISTINCT\|PROBLEMATIC","summary": "<2\-3 sentence overall assessment covering thematic fit, visualdistinctiveness, identifiability, and scene quality\>"\}\}\-\-\-INPUT TO EVALUATE \(process this now\)\-\-\-Verb \(root form\): "\{verb\}"\{generated\_output\}Follow the EVALUATION PROCEDURE above \(Steps 1\-6\)\. Return ONLY the JSON object\. No other text\.Figure 24:Evaluator Prompt \(continued from previous page\)\.
## Appendix FHyperparameter Details

This appendix lists the hyperparameters used across the generation and evaluation runs reported in §[5](https://arxiv.org/html/2608.04567#S5)\.

### F\.1Models

All experiments use the model versions listed in Table[6](https://arxiv.org/html/2608.04567#A6.T6)\. Closed\-source models are accessed via provider SDKs; open\-weight models are served locally withvLLM\.

Table 6:Model identifiers\. The six rows above the mid\-rule are used as generators throughout the paper; Qwen3\-30B\-Thinking is used only as an open\-source evaluator\.
### F\.2Generation

Generators use sampling atT=0\.3T=0\.3across all models, with up to 16,384 output tokens\. Native reasoning/thinking modes are disabled on generators since they either pin temperature to fixed values or return summarised traces; the reasoning scratchpad introduced in §[3\.1](https://arxiv.org/html/2608.04567#S3.SS1)preserves full reasoning atT=0\.3T=0\.3\.

For R2V, the generator samples 20 candidate agents per stimulus set with verbalized plausibility scores, then commits one agent per condition from the tiers PE∈\[0\.80,1\.00\]\\in\[0\.80,1\.00\], PH∈\[0\.40,0\.65\]\\in\[0\.40,0\.65\], IH∈\[0\.15,0\.35\]\\in\[0\.15,0\.35\], and IE==lowest\-scoring candidate\. A fallback of 5 additional candidates is sampled if no tier\-appropriate agent passes quality checks\. R2R and R2VR run up tok=3k=3iterations and exit early once the in\-loop evaluator returns a GOOD scene verdict with all four per\-condition verdicts marked CORRECT\.

### F\.3Evaluation

GPT\-5\.1 evaluators run withreasoning\_effort = highat the provider default temperature; Qwen3\-30B\-Thinking evaluators run withT=0\.6T=0\.6and an 8,000\-token think budget\. Both apply the multi\-step prompt described in §[3\.2](https://arxiv.org/html/2608.04567#S3.SS2)that requires per\-condition adversarial counter\-arguments, an overlap audit at the IH boundary, and a holistic scene assessment\. Maximum output tokens are 24,576 for in\-loop \(iterative\) evaluation and 32,768 for standalone \(post\-hoc\) evaluation\.

Table 7:Decoding parameters by role\.
### F\.4Compute

Open\-weight runs executed on a SLURM\-managed GPU cluster with NVIDIA L40S and RTX 6000 GPUs\. Local models were served withvLLM\(\-\-dtype auto, max model context 65,536 tokens\) using tensor parallelism matched to model size: TP=1=1for the 4B and 14B models, TP=2=2for the 30B model, and TP=4=4for the 80B model\. Single\-shot generation jobs complete in 0\.5–3\.5 hours per model; iterative generation runs in 12–18 hours\. Closed\-source API calls require no GPU allocation\. Code, prompts, and SLURM scripts will be released upon acceptance\.

## Appendix GEvaluator Scaling

We extend the AAT validation from a single judge \(§[5\.2](https://arxiv.org/html/2608.04567#S5.SS2)\) to six judge models spanning closed\-source frontier \(GPT\-5\.1, Claude Sonnet 4\.6\), closed\-source mid\-tier \(GPT\-5\.4\-mini\), and three open\-source scales \(Qwen3\-Next\-80B\-A3B, Qwen3\-30B\-A3B, Qwen3\-4B\)\. Each is run ingreedymode \(no model\-native reasoning\) andreasoningmode \(thinking or extended reasoning, where available\)\. L\-H agreement is compared to the H\-H baseline \(κ=\.806\\kappa=\.806weighted,κ=\.529\\kappa=\.529unweighted\), and AAT non\-inferiority is tested atδ=0\.05\\delta=0\.05\. The experiment, summarised in Table[8](https://arxiv.org/html/2608.04567#A7.T8), addresses two questions: \(i\) which models are recommendable as evaluators, and \(ii\) whether test\-time reasoning is necessary or greedy decoding suffices\.

Reasoning helps almost everywhere\.Across 9 of 10 model\-metric comparisons where both modes are tested, reasoning beats greedy\. The lone exception \(Sonnet 4\.6 unweighted:\.528\.528vs\.521\.521\) is sub\-noise\. Gains are largest on mid\-tier and small models: GPT\-5\.4\-mini gains\+0\.030\+0\.030weighted and\+0\.029\+0\.029unweighted; Qwen3\-4B gains\+0\.028\+0\.028weighted\. Test\-time reasoning is therefore the default operating mode for the evaluator\.

Only GPT\-5\.1 with reasoning passes both AAT criteria\.On weightedκ\\kappa\(ordering\), every model where reasoning is tested except Qwen3\-4B passes AAT, so the gradient ordering is broadly recoverable across the lineup\. On unweightedκ\\kappa\(exact classification\), only GPT\-5\.1 with reasoning passes\. Exact classification is the criterion we use to recommend an evaluator for downstream use, and it separates GPT\-5\.1 from the rest of the field\.

Mid\-tier open\-source preserves ordering; 4B is below the floor\.Qwen3\-30B\-A3B with reasoning \(κ=\.811\\kappa=\.811weighted\) matches Sonnet 4\.6 greedy \(\.815\.815\)\. For pipelines that need ordering but not exact classification, e\.g\., as the iteration feedback source in §[5\.1](https://arxiv.org/html/2608.04567#S5.SS1)rather than the final tier\-assignment evaluator, this is a viable open\-source operating point\. Qwen3\-4B, by contrast, fails AAT on both metrics in both modes and is not usable as a graded\-plausibility evaluator\.

Table 8:Evaluator Scaling: Judge Model×\\timesReasoning\. Each cell:greedy∣\\midreasoning\. H\-H baselines: weightedκ=\.806\\kappa=\.806, unweightedκ=\.529\\kappa=\.529\. Boot 5th = 5th percentile of bootstrappedΔ=κ¯​\(L,H\)−κ¯​\(H,H\)\\Delta=\\bar\{\\kappa\}\(\\text\{L,H\}\)\-\\bar\{\\kappa\}\(\\text\{H,H\}\); AAT holds when boot 5th\>−δ\>\-\\delta\.n=120n=120sentences, 8 annotators,B=10,000B=10\{,\}000\.Bold= better of the greedy/reasoning pair\. Each cell:greedy/no\-thinking∣\\midreasoning/thinking\. ✓ = non\-inferiority \(boot 5th\>−0\.05\\text\{boot 5th\}\>\-0\.05\);×\\times= fails\. “—” = not tested due to model loading issue\.
## Appendix HProbing the Evaluation Floor

Before committing to an LLM\-as\-judge, we ask whether a cheaper information\-theoretic baseline can rank the four agents using only token probabilities from the same Qwen3\-30B backbone\. We then ask whether the residual IH errors of the strongest judge can be recovered by an ensemble over weaker ones\.

### H\.1Information\-Theoretic Baselines

We score each of the four candidate agents under three progressively stronger formulations, all on the samen=120n=120sentences \(30 items×\\times4 conditions\)\. For each condition, we count how often the top\-ranked agent matches the generator\-assigned label\.

##### evalv5: total sentence log\-probability\.

We score each candidate sentenceS=\(w1,…,wT\)S=\(w\_\{1\},\\dots,w\_\{T\}\)by

log⁡P​\(S\)=∑t=1Tlog⁡P​\(wt∣w<t\),\\textstyle\\log P\(S\)=\\sum\_\{t=1\}^\{T\}\\log P\(w\_\{t\}\\mid w\_\{<t\}\),and rank the four agents bylog⁡P​\(S\)\\log P\(S\)\. This conflates fit with the agent’s marginal pretraining frequency: a high\-frequency agent \(e\.g\.*doctor*\) wins regardless of event fit\.

##### evalv5\.1: agent prior removed\.

We strip the agent span and score the post\-agent tokenswa\+1:Tw\_\{a\+1\{:\}T\}conditioned on the agent prefixw1:aw\_\{1\{:\}a\}:

log⁡P​\(S∖agent\)=∑t=a\+1Tlog⁡P​\(wt∣w<t\)\.\\textstyle\\log P\(S\\setminus\\text\{agent\}\)=\\sum\_\{t=a\+1\}^\{T\}\\log P\(w\_\{t\}\\mid w\_\{<t\}\)\.This removes prior frequency but relies on left\-to\-right conditioning to propagate agent–event compatibility through the remaining tokens\.

##### Surprise \- frame\-first conditional\.

We restructure the prompt so the frame precedes the agent —“Someone \[verb\] \[patient\] \[instrument\] \[location\]\. This person is a”— and score each candidate agentaadirectly at the blank,

log⁡P​\(a∣frame\)=∑t∈alog⁡P​\(wt∣frame,w<t\)\.\\textstyle\\log P\(a\\mid\\text\{frame\}\)=\\sum\_\{t\\in a\}\\log P\(w\_\{t\}\\mid\\text\{frame\},w\_\{<t\}\)\.This is the strongest distributional formulation: the full event context conditions the agent with no prior\-frequency or ordering penalty\.

##### Findings\.

Frame\-first conditioning helps but does not close the gap \(Table[9](https://arxiv.org/html/2608.04567#A8.T9)\)\. evalv5\.2 improves 4\-point accuracy to 48% \(\+7 over evalv5\), with most of the gain on PE \(\+\+3\) and IE \(\+\+6\)\. PH stays pinned at 9/30 across all three variants and IH does not improve\. Even fully optimized, the best surprisal method sits ten points below the*weakest*LLM judge \(Qwen3\-4B greedy, 58%\) and twenty below the strongest \(GPT\-5\.1 with reasoning, 68%\)\. We read this as structural: “unusual but plausible” \(PH\) and “related but wrong domain” \(IH\) are constraint\-satisfaction judgments, not properties of distributional support, and no amount of conditioning extracts them from token probabilities alone\.

### H\.2Can an Ensemble Recover the Residual?

A natural follow\-up is whether the IH errors of any single judge can be recovered by majority vote across multiple judges\. Under Condorcet’s jury theorem, independent errors above 50% individual accuracy should aggregate into a strictly stronger jury\. We test this on the 11 evaluator configurations from §[G](https://arxiv.org/html/2608.04567#A7), split into a passing group \(6 configurations meeting our IH threshold\) and a failing group \(5 below it\), and report per\-item difficulty, pairwise error correlationϕ\\phi, and a majority\-vote simulation \(Table[10](https://arxiv.org/html/2608.04567#A8.T10)\)\.

Three results show that ensembling cannot help\. First, 15/30 IH items are unsolvable: less than 50% of passing models get them right, and 5 items are missed by*all 11*configurations — a fixed LLM–human disagreement that no aggregation can resolve\. Second, the failing\-group jury \(33%\) is worse than its best member \(47%\); the passing\-group jury also degrades by 9 points\. Error correlation is high enough \(ϕ¯=\.49\\bar\{\\phi\}=\.49within passing,\.27\.27cross\-group\) that majority vote amplifies shared mistakes rather than canceling independent ones\. Third, the pass–fail accuracy gap collapses on hard items \(8% vs\. 13%\) after being 26 points on easy items, indicating that capability matters only up to a ceiling fixed by item ambiguity\.

The two probes bracket the operating regime of a viable judge from opposite sides: distributional shortcuts cannot reach the LLM floor, and ensembling cannot lift its ceiling\. Reliability claims must therefore be made conditional on items being above the LLM–human ambiguity threshold rather than across the full distribution\.

Per\-condition match \(/30\)MethodSignal usedPEPHIHIE4\-pt \(/120\)Binary \(/120\)Information\-theoretic evaluation \(Qwen3\-30B\)evalv5Total sentencelog⁡P\\log P159121349 \(41%\)82 \(68%\)evalv5\.1Frame\-onlylog⁡P\\log P\(agent prior removed\)129141348 \(40%\)86 \(72%\)SurpriseFrame\-firstP​\(agent∣frame\)P\(\\text\{agent\}\\mid\\text\{frame\}\)189121958 \(48%\)86 \(72%\)LLM\-as\-judge evaluationWorstQwen3\-4B, greedy27772869 \(58%\)98 \(82%\)BestGPT\-5\.1, reasoning289172882 \(68%\)104 \(87%\)Table 9:Information\-theoretic metrics vs\. LLM\-as\-judge evaluation\. All surprisal methods use Qwen3\-30B via sentence log\-probability ranking \(n=120n=120sentences, 30 items×\\times4 conditions\)\. PE identification was the primary motivation; none achieve reliable PE detection \(human\-aligned PE≥\\geq27/30\)\. PH remains flat at 9/30 across all information\-theoretic methods — token probability cannot capture “unusual but plausible\.” IH shows no improvement — multi\-dimensional overlap reasoning requires deliberative evaluation\.Table 10:IH failure analysis summary\. Jury = majority vote over model subset\.ϕ\\phi= mean pairwise phi coefficient \(error correlation\)\. Condorcet assumes independent errors\.

## Appendix IVerb\-Level Evidence for Slot\-Dependent Difficulty

We expand the slot\-dependency analysis from §[6](https://arxiv.org/html/2608.04567#S6)with three complementary views\.

##### Per\-verb GOLD distribution \(Table[11](https://arxiv.org/html/2608.04567#A9.T11)\)\.

Aggregating across all models and methods in agent\-varying generation, per\-verb GOLD rates span42\.4%42\.4\\%\(throw\) to3\.4%3\.4\\%\(tickle,pay\)\. The distribution divides into three bands: Easy \(≥20%\\geq 20\\%GOLD,n=11n=11\), Moderate \(10–19%,n=34n=34\), and Hard \(<10%<10\\%,n=15n=15\)\. The Hard band concentrates two kinds of verbs: \(a\) actions tied to no specific occupation \(eat,drink,tickle,pay\), which leave no prototypical PE agent, and \(b\) actions with very narrow occupational scope \(tow,iron,erase\), which leave no room for a distinct IH agent\. Both failure modes are properties of the agent slot, not of the action itself\.

##### Aggregate slot comparison \(Table[5](https://arxiv.org/html/2608.04567#S7.T5)\)\.

With the patient slot varied, mean R2V GOLD rises from11\.3%11\.3\\%\(agent\) to16\.4%16\.4\\%\(patient\) across the five models with both\-slot coverage\. The gain concentrates where agent\-varying was weakest: Qwen3\-Next\-80B\-A3B improves from6\.7%6\.7\\%to21\.7%21\.7\\%\. RB shows a smaller and reverse aggregate shift \(10\.0%10\.0\\%agent,6\.7%6\.7\\%patient\), driven by Sonnet 4\.6’s outlier agent\-side performance \(28\.3%28\.3\\%\); without iteration, neither slot can be reliably exploited\.

##### Per\-verb slot overlap \(Figure[25](https://arxiv.org/html/2608.04567#A9.F25)\)\.

At the per\-\(model, verb\) level, only3%3\\%of pairs achieve GOLD in both slot conditions under R2V; the remaining GOLD pairs split between agent\-only \(8%8\\%\) and patient\-only \(13%13\\%\)\. Verbs that succeed in agent\-varying generation are largely a different set from those that succeed in patient\-varying generation\. A single\-slot ceiling on GOLD rate is a ceiling on what that slot can support for the verb, not a ceiling on what the framework can generate; comprehensive stimulus coverage requires varying multiple slots\.

Table 11:Per\-verb GOLD rate across all models and methods \(60 verbs, sorted by GOLD%\)\. Difficulty: Easy \(≥\\geq20%\), Moderate \(10–19%\), Hard \(<<10%\)\.![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/verb_framing_agent_vs_patient_gold.png)Figure 25:Per\-verb GOLD outcome under agent vs\. patient slot variation, for each model\. Each cell is one \(model, verb\) pair where both slot conditions were evaluated; cell color indicates which slot\(s\) achieved GOLD\. The both\-GOLD share is2%2\\%\(RB, top\) and3%3\\%\(R2V, bottom\), confirming that the agent\-GOLD and patient\-GOLD verb sets are largely disjoint\.

## Appendix JOutput Token Costs

Table[12](https://arxiv.org/html/2608.04567#A10.T12)reports the mean per\-verb generator output tokens for each method across all six generators, computed over the 60\-verb set\. Per\-verb counts sum all generator API call outputs in a run; for R2R and R2VR, this includes outputs across all refinement iterations but excludes the evaluator’s tokens\. Iterative methods average 2\.4–2\.8 rounds per verb, with early exit when the evaluator returns a GOOD scene verdict and four CORRECT condition verdicts \(§[4](https://arxiv.org/html/2608.04567#S4)\)\.

Table 12:Mean generator output tokens per verb on the 60\-verb set\. R2R/R2VR sum tokens across all refinement iterations \(mean 2\.4–2\.8 rounds per verb\)\. Open\-source iterative runs used the Qwen judge for both iteration feedback and evaluation \(§[4](https://arxiv.org/html/2608.04567#S4)\) and are omitted here to preserve like\-for\-like comparisons against the GPT\-judged closed\-source pipeline\.##### Iteration cost\.

On the closed\-source generators where R2R was run, total output tokens scale by 2\.72×\\times\(GPT\-5\.1\) to 3\.51×\\times\(Sonnet 4\.6\) relative to RB at matched reasoning configuration; R2VR scales by 2\.02×\\timesto 2\.25×\\timesrelative to R2V\. The mean per\-iteration generator output \(total÷\\divrounds\) is comparable to the corresponding single\-shot method \(e\.g\., GPT\-5\.1 R2R averages 2,634 tokens per iteration vs RB’s 2,633\), indicating that the cost inflation in iterative methods comes from running multiple rounds rather than heavier per\-call generation\.

##### Reasoning scratchpad cost\.

Within a method, the global reasoning scratchpad adds\+1,690\+1\{,\}690tokens to RB on average \(3\.74×\\timesthe no\-scratchpad baseline\) and\+1,856\+1\{,\}856tokens to R2V \(1\.88×\\times\)\. The larger relative increase for RB reflects that no\-scratchpad RB is also the shortest method \(mean 557–685 output tokens across generators\), so the added reasoning trace dominates total length\.

## Appendix KCross\-Evaluator GOLD Concordance

Whether an open\-source pipeline can substitute for the closed\-source one hinges on a directional question: when the Qwen3\-30B judge flags a stimulus set asGold, does the GPT\-5\.1 judge also flag it asGold? We answer this on the closed\-source iterative subset—R2R and R2VR on GPT\-5\.1 and Sonnet 4\.6,N=235N=235—the configuration that produces the paper’s primary high\-quality results \(§[5\.1](https://arxiv.org/html/2608.04567#S5.SS1)\)\. Single\-shot methods \(RB, R2V\) are excluded becauseGoldis rare under either judge in those runs, and the smallGoldpool inflates variance\.

Of the 235 cells, both judges agreed onGoldin 97 and on non\-Goldin 35\. The remaining 103 split asymmetrically: GPT alone flaggedGoldin 81 cells, Qwen alone in 22 \(Table[13](https://arxiv.org/html/2608.04567#A11.T13)\)\.

Table 13:Gold\-tier contingency on the closed\-source iterative subset \(N=235N=235\)\. Rows: GPT\-5\.1 judge; columns: Qwen3\-30B\-Thinking judge\.##### Primary statistic\.

Conditional on Qwen flaggingGold, GPT also flagsGoldin97/119=81\.5%97/119=81\.5\\%of cells \(Wilson 95% CI:\[73\.6%,87\.5%\]\[73\.6\\%,87\.5\\%\]\)\. The CI excludes 50%, ruling out chance agreement\. The reverse conditional is lower:P​\(Qwen=Gold∣GPT=Gold\)=97/178=54\.5%P\(\\text\{Qwen\}\\\!=\\\!\\textsc\{Gold\}\\mid\\text\{GPT\}\\\!=\\\!\\textsc\{Gold\}\)=97/178=54\.5\\%\. Qwen is therefore the stricter judge, and itsGoldset is a higher\-precision subset of GPT’s\. Restricting downstream stimulus selection to Qwen\-flaggedGoldyields a smaller pool, but one that GPT also endorses with high probability\.

##### Consistency across iterative methods\.

The two iterative methods show comparable concordance:P​\(GPT∣Qwen\)=80\.4%P\(\\text\{GPT\}\\mid\\text\{Qwen\}\)=80\.4\\%on R2R \(N=115N=115\) and82\.5%82\.5\\%on R2VR \(N=120N=120\)\. Figure[26](https://arxiv.org/html/2608.04567#A11.F26)shows the per\-verb structure of agreement and disagreement; theGold\-overlap pattern \(green cells\) is distributed across verbs rather than concentrated in a small subset, confirming that the 81\.5% is a population\-level property of the two judges rather than an artefact of a few easy verbs\.

![Refer to caption](https://arxiv.org/html/2608.04567v1/figures/assets/fig_judge_concordance.png)Figure 26:Per\-verbGoldoverlap between the GPT\-5\.1 and Qwen3\-30B\-Thinking judges on the closed\-source iterative subset\. Top: R2R \(115 complete cells\)\. Middle: R2VR \(120 complete cells\)\. Bottom: net preference per verb \(GPT\-only minus Qwen\-only count across the two generators\)\. Green: both judgesGold; blue: GPT only; red: Qwen only; light gray: neither; dark gray: missing\-judge cells\.

## Appendix LPlausibility\-Gate Pass Rates

The plausibility\-gate pass rate is the proportion of sets for which all four condition verdicts are correct, irrespective of scene and visual verdicts\. Under the hierarchical rule in Table[1](https://arxiv.org/html/2608.04567#S4.T1), this is the share that clears the plausibility prerequisite for a non\-FAIL tier before scene and visual verdicts determine the final quality tier\. Table[14](https://arxiv.org/html/2608.04567#A12.T14)compares this rate with GOLD under the GPT judge\.

Table 14:GOLD and plausibility\-gate pass rates \(%\) under the GPT judge\. The refinement gains persist before applying the scene and visual gates\.
## Appendix MHuman Annotation Study Details

##### Stimulus sampling\.

Thirty stimulus sets were sampled from the full evaluation pool via stratified sampling across quality tiers \(GOLD, SILVER, BRONZE\+FAIL\), ensuring coverage of the full quality range rather than selecting only successful outputs\.

##### Survey instrument\.

Each annotator rated all 120 sentences \(30 sets×\\times4 conditions\) on a four\-point plausibility scale: \(1\)Clearly Plausible, \(2\)Somewhat Plausible, \(3\)Somewhat Implausible, \(4\)Clearly Implausible, plus aCannot Decideoption\. Annotators additionally rated decision difficulty per sentence \(1 = Easy, 2 = Moderate, 3 = Hard\)\. At the set level, annotators rated the visual identifiability of each agent \(Yes / Maybe / No\) and flagged any agent pairs they judged visually indistinguishable\. No annotator selectedCannot Decidefor any item, confirming that all stimuli were interpretable\. Instructions to annotators and sample questions are shown in Figure[27](https://arxiv.org/html/2608.04567#A13.F27)

##### Gradient separability\.

For each sentence, we average the eight plausibility ratings and treat each matched four\-sentence set as the repeated\-measures unit \(n=30n=30\)\. Because the original ratings are ordinal and every set contains all four conditions, we use a Friedman test\(Friedman,[1937](https://arxiv.org/html/2608.04567#bib.bib13)\)rather than a parametric or independent\-samples test to assess whether ratings differ overall\. We then compare the three adjacent pairs using paired Wilcoxon signed\-rank tests\(Wilcoxon,[1945](https://arxiv.org/html/2608.04567#bib.bib40)\)with Holm correction\(Holm,[1979](https://arxiv.org/html/2608.04567#bib.bib15)\)for multiple comparisons\. Mean ratings increase from PE 1\.21 to PH 1\.89, IH 2\.58, and IE 3\.86\. The Friedman test finds an overall condition effect \(χ2​\(3\)=76\.5\\chi^\{2\}\(3\)=76\.5,p<10−15p<10^\{\-15\}\)\. All adjacent comparisons remain significant after Holm correction: PE–PH \(p=3\.3×10−4p=3\.3\\times 10^\{\-4\}\), PH–IH \(p=5\.3×10−4p=5\.3\\times 10^\{\-4\}\), and IH–IE \(p=5\.1×10−6p=5\.1\\times 10^\{\-6\}\)\.

##### Binary mapping\.

A supplementary binary analysis maps ratings 1–2 toplausibleand 3–4 toimplausible, reflecting the patient\-facing judgment task in which participants make a binary decision\. Per\-condition agreement is additionally reported using Gwet’s AC1 to guard against theκ\\kappaprevalence paradox at the extreme conditions \(PE and IE\)\.

##### Annotator details\.

Eight trained annotators participated; all were fluent English speakers familiar with psycholinguistic experimental design\. Annotation was conducted via an online survey platform\. Median completion time was 169 minutes \(5\.6 minutes per set\)\.

![Refer to caption](https://arxiv.org/html/2608.04567v1/x4.png)Figure 27:Survey welcome page with instructions and sample questions for per\-sentence plausibility classification and visual distinctiveness

## Appendix NExtended AAT Results

##### H\-H baseline\.

Krippendorff’sα=0\.805\\alpha=0\.805across all 8 annotators and 120 sentences confirms strong inter\-annotator agreement and that the stimuli are interpretable \(zeroCannot Decideresponses\)\. Mean pairwise weightedκw=0\.806\\kappa\_\{w\}=0\.806\(±0\.061\\pm 0\.061\)\.

##### Per\-condition exact match \(4\-point\)\.

Exact match between LLM and human majority: PE = 90%, PH = 30%, IH = 53%, IE = 93%\. The low PH match reflects genuine ambiguity: 75% of PH sentences received aplausiblehuman majority, but only 30% matched the LLM’s assigned category exactly\. 94% of all 4\-point disagreements are one\-step adjacent on the gradient; mean disagreement difficulty exceeds mean agreement difficulty \(Mann\-Whitneyp<0\.001p<0\.001\), confirming that disagreements cluster on items that human annotators themselves found hard\.

##### Per\-condition binary AAT \(AC1\)\.

All four conditions pass non\-inferiority atδ≤0\.05\\delta\\leq 0\.05\(AC1 bootstrap\)\. LLM agreement exceeds H\-H at PH and IH, the two conditions where unaided human intuition is least reliable\.

##### Supporting weightedκw\\kappa\_\{w\}and binary AAT\.

4\-point weightedκw\\kappa\_\{w\}: H\-H = 0\.806, L\-H = 0\.821,Δ=\+0\.015\\Delta=\{\+\}0\.015, passing atδ≤0\.01\\delta\\leq 0\.01\. Binaryκ\\kappa: H\-H = 0\.650, L\-H = 0\.686,Δ=\+0\.036\\Delta=\{\+\}0\.036, passing atδ≤0\.01\\delta\\leq 0\.01\.

##### F1 identifiability detail\.

Human identifiability distribution: Yes = 52\.4%, Maybe = 35\.9%, No = 11\.7%\. LLM distribution: IDENTIFIABLE = 75%, AMBIGUOUS = 22\.5%, UNIDENTIFIABLE = 2\.5%\. F1κ\\kappa: H\-H = 0\.277, L\-H = 0\.210,Δ=−0\.067\\Delta=\{\-\}0\.067, passing atδ≤0\.15\\delta\\leq 0\.15\. The LLM is systematically more optimistic about identifiability than human annotators, underrating the Maybe and No categories\.

Table 15:Generation ExamplesGOLDSILVERBRONZEFAIL FAIL rows:sentence highlighting is applied only when ‘scene\_verdict=GOOD‘\.

相似文章

基于外部子图生成的大语言模型逐步推理增强

arXiv cs.CL

本文提出了SGR框架,通过查询相关的子图生成将外部知识图谱与大语言模型相结合,融合基于Cypher的推理与协同推理集成,从而增强大语言模型的逐步推理能力。在CWQ、WebQSP、GrailQA和KQA Pro上的实验表明,该框架相比标准提示方法和知识增强基线具有更高的推理准确性。