IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization

arXiv cs.AI Papers

Summary

IntElicit is a framework that uses dialogue policy optimization with a decomposed process reward mechanism to elicit and assess contextualized creativity through adaptive AI interviewing, reducing confounders like domain knowledge and engagement. Experiments show it improves creative outcomes over static assessment methods.

arXiv:2606.12086v1 Announce Type: new Abstract: Contextualized assessment offers high ecological validity for evaluating creativity but introduces a critical challenge: observed performance may be confounded with cognitive proficiency (domain knowledge) and agency (willingness to engage). Meanwhile, in the age of generative AI, creative problem solving increasingly occurs in tool-mediated and human--AI interactive environments, making fully static assessment less aligned with contemporary creative practice. To address these issues, this paper proposes IntElicit, a framework for eliciting and assessing contextualized creativity via dialogue policy optimization. IntElicit functions as a constrained adaptive AI Interviewer: it provides non-directive knowledge and agency scaffolds in multi-turn interaction to reduce non-creative confounders, while preserving participants' responsibility for generating the creative content being evaluated. Specifically, to tackle sparse rewards and potential reward hacking (e.g., answer dictation) in open-ended educational dialogue, IntElicit introduces a decomposed process reward mechanism. This mechanism aligns the policy with pedagogical elicitation, rewarding prompts that draw out participant reasoning rather than producing optimal answers on their behalf. Extensive experiments, including participant simulation and a human subject study (N=64), show that IntElicit improves elicited creative outcomes over expert-designed baselines. Together, the results suggest that interactive elicitation can reveal creative potential that static FPSP-style assessment may miss, providing a formative and diagnostic lens for contextualized creativity assessment in AI-mediated learning contexts.
Original Article
View Cached Full Text

Cached at: 06/11/26, 01:50 PM

# IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization
Source: [https://arxiv.org/html/2606.12086](https://arxiv.org/html/2606.12086)
Mingjia Li1,†, Jin Wu1,†, Hong Qian1,2,∗, Wenhao Huang1, Yiyang Huang1Yiwen Zhang1, Chanjin Zheng1, Xiangfeng Wang1, Aimin Zhou1,2, Jiajun Guo11East China Normal University2Shanghai Innovation Institute†Equal contribution\.∗Corresponding author:hqian@cs\.ecnu\.edu\.cn

###### Abstract

Contextualized assessment offers high ecological validity for evaluating creativity but introduces a critical challenge: observed performance may be confounded with cognitive proficiency \(domain knowledge\) and agency \(willingness to engage\)\. Meanwhile, in the age of generative AI, creative problem solving increasingly occurs in tool\-mediated and human–AI interactive environments, making fully static assessment less aligned with contemporary creative practice\. To address these issues, this paper proposesIntElicit, a framework for eliciting and assessing contextualized creativity via dialogue policy optimization\. IntElicit functions as a constrained adaptiveAI Interviewer: it provides non\-directive knowledge and agency scaffolds in multi\-turn interaction to reduce non\-creative confounders, while preserving participants’ responsibility for generating the creative content being evaluated\. Specifically, to tackle sparse rewards and potential reward hacking \(e\.g\., answer dictation\) in open\-ended educational dialogue, IntElicit introduces a decomposed process reward mechanism\. This mechanism aligns the policy with pedagogical elicitation, rewarding prompts that draw out participant reasoning rather than producing optimal answers on their behalf\. Extensive experiments, including participant simulation and a human subject study \(N=64N=64\), show that IntElicit improves elicited creative outcomes over expert\-designed baselines\. Together, the results suggest that interactive elicitation can reveal creative potential that static FPSP\-style assessment may miss, providing a formative and diagnostic lens for contextualized creativity assessment in AI\-mediated learning contexts\.

*K*eywordsContextualized creativity assessment⋅\\cdotInteractive elicitation⋅\\cdotDialogue policy optimization⋅\\cdotLarge language models⋅\\cdotFormative assessment

## 1Introduction

Creativity, often defined as the ability to generate products that are both novel and useful, is widely recognized as a critical competency for the 21st century\(Glăveanu and Petre,[2010](https://arxiv.org/html/2606.12086#bib.bib1); Taguma and Barrera,[2019](https://arxiv.org/html/2606.12086#bib.bib3); Llego,[2022](https://arxiv.org/html/2606.12086#bib.bib4)\)\. As artificial intelligence increasingly automates routine cognitive tasks, the human capacity for creative problem\-solving has become central to education and workforce development\(Brynjolfsson and McAfee,[2014](https://arxiv.org/html/2606.12086#bib.bib5)\)\. However, despite its importance, assessing creativity in educationally meaningful contexts remains difficult because creative performance is shaped not only by idea generation, but also by background knowledge, confidence, engagement, and the interactional conditions under which students are asked to respond\. This challenge echoes broader concerns that traditional assessments often provide only discrete snapshots of performance and may be insufficiently adapted to learners’ backgrounds and contemporary AI\-mediated practices\(Swieckiet al\.,[2022](https://arxiv.org/html/2606.12086#bib.bib47)\)\.

Traditional approaches to creativity assessment largely fall into two categories: self\-report scales and performance tests\. Self\-report measures often suffer from social subjectivity bias, where individuals may misjudge their abilities due to factors like the Dunning\-Kruger effect\(Kruger and Dunning,[1999](https://arxiv.org/html/2606.12086#bib.bib24)\)or social desirability\(Paulhus,[1984](https://arxiv.org/html/2606.12086#bib.bib6); Silviaet al\.,[2012](https://arxiv.org/html/2606.12086#bib.bib7)\)\. Performance tests, conversely, attempt to measure creativity objectively but vary significantly in their contextual fidelity\. Common approaches range from simple tasks like the Alternative Uses Task \(AUT\)\(Runco and Acar,[2012](https://arxiv.org/html/2606.12086#bib.bib8)\)and Realistic Presented Problems \(RPP\)\(Chand and Runco,[1993](https://arxiv.org/html/2606.12086#bib.bib20)\), to highly contextualized frameworks such as the Future Problem Solving Program \(FPSP\)\(Crabbe,[1982](https://arxiv.org/html/2606.12086#bib.bib17),[1989](https://arxiv.org/html/2606.12086#bib.bib18); Torranceet al\.,[1976](https://arxiv.org/html/2606.12086#bib.bib19)\)\. While AUT and RPP are widely used, their simplistic instructions \(e\.g\.,“list uses for a brick”\) suffer fromlimited ecological validity\(Baer,[2015](https://arxiv.org/html/2606.12086#bib.bib23); Zenget al\.,[2011](https://arxiv.org/html/2606.12086#bib.bib21)\), failing to capture the complex nature of real\-world problem\-solving\.

To address this ecological gap, our work aligns with the paradigm of FPSP, which evaluates participants within immersive, realistic scenarios\. However, while FPSP improves contextual fidelity, a fully static implementation still faces two limitations\. First, performance in complex scenarios can be entangled with cognitive and agential factors, such as domain knowledge, task comprehension, confidence, and willingness to elaborate\(Runco and Chand,[1995](https://arxiv.org/html/2606.12086#bib.bib22)\)\. A participant may therefore perform poorly not because they lack creative potential, but because they fail to externalize that potential under unsupported testing conditions\. Second, in the age of generative AI, creative problem solving increasingly occurs in tool\-mediated and human–AI collaborative environments, where people clarify problems, explore alternatives, and refine ideas through interaction\(Rezwana and Maher,[2023](https://arxiv.org/html/2606.12086#bib.bib45); Noy and Zhang,[2023](https://arxiv.org/html/2606.12086#bib.bib46); Oliveiraet al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib48)\)\. From an authentic assessment perspective, the assessment situation should resemble the kinds of practices in which the target competence is expected to be used\. Interactive elicitation addresses these two limitations by making participants’ reasoning more observable under adaptive, non\-directive scaffolding and by modeling an assessment format closer to contemporary AI\-mediated creative work\. Consequently, we argue thatinteractive elicitationis a necessary extension of contextualized creativity assessment\.

![Refer to caption](https://arxiv.org/html/2606.12086v1/x1.png)Figure 1:Research motivation and overview of the proposed IntElicit framework\.\(a\)TheEcological Validity Gap: Static creativity assessments may miss dynamic, process\-oriented reasoning in contextualized problem\-solving tasks\.\(b\)TheInteractive Elicitation Paradigm: An AI interviewer provides adaptive scaffolding \(e\.g\., knowledge support and agency elicitation\) to reduce non\-creative confounders during assessment while preserving the participant’s responsibility for generating ideas\.\(c\)TheIntElicit Architecture: The framework combines multidimensional creativity indicators, decomposed process rewards, and diverse participant simulators to support open\-ended dialogue policy optimization\.As illustrated in Figure[1](https://arxiv.org/html/2606.12086#S1.F1), our approach first motivates interactive elicitation as a response to the ecological validity gap, and then implements it as an adaptive AI interviewer framework for contextualized creativity assessment\. To realize this vision, we proposeIntElicit, anInteractiveElicitation framework powered by Dialogue Policy Optimization that functions asan adaptive AI Interviewer\. Importantly, the role of the AI is not to co\-author the creative response or train participants to become more creative during the task\. Instead, IntElicit acts as a constrained assessment scaffold: it may help participants clarify task contexts, sustain engagement, elaborate their reasoning, and reflect on alternatives, but the creative content to be assessed must remain participant\-generated\. While Large Language Models \(LLMs\) offer a promising foundation for such an AI Interviewer\(OpenAI,[2023](https://arxiv.org/html/2606.12086#bib.bib12); Kasneciet al\.,[2023](https://arxiv.org/html/2606.12086#bib.bib13)\), optimizing them for this role presents significant technical challenges\. The primary hurdle is optimizing dialogue policies for non\-verifiable, open\-ended goals\. Unlike domains with verifiable rewards such as math or coding\(Ouyanget al\.,[2022](https://arxiv.org/html/2606.12086#bib.bib14); Lightmanet al\.,[2024](https://arxiv.org/html/2606.12086#bib.bib15); Rafailovet al\.,[2023](https://arxiv.org/html/2606.12086#bib.bib16)\), creativity assessment is subjective and process\-sensitive\. Furthermore, in a multi\-turn assessment, the reward signal is often sparse \(received only at the end\)\. This sparsity can lead to “reward hacking”, where the AI Interviewer, aiming to maximize the final creativity score, might simply dictate high\-quality ideas to the participant rather than eliciting them, which defeats the purpose of assessment\. Finally, the agent must be robust, capable of adapting to diverse participant behaviors, from the “reticent” interviewee needing encouragement to the “divergent” one wandering off\-topic\.

IntElicit addresses these challenges through a synergistic approach\. First, 16 immersive assessment scenarios are designed by expert psychologists following the FPSP paradigm\. Grounded in this context, a multidimensional indicator system is constructed to operationalize contextualized creative performance\. To prevent reward hacking and support non\-directive elicitation, we introduce a decomposed process reward mechanism based on expert pedagogical strategies \(e\.g\., rewarding prompts that encourage participants to identify problems, justify ideas, and reflect on alternatives\)\. We train the policy using a diverse participant simulator, creating a controlled environment populated with simulated participants exhibiting different engagement patterns\.

The contributions of this paper are as follows\. First, we introduceinteractive elicitationas a formative and diagnostic paradigm for contextualized creativity assessment, aiming to reduce false negatives caused by knowledge gaps or low agency while preserving the participant’s role as the source of creative ideas\. Second, we proposeIntElicit, a dialogue policy optimization framework that learns adaptive assessment scaffolds for open\-ended, multi\-turn creativity tasks\. Third, we introduce aDecomposed Process Rewardmechanism that rewards pedagogically meaningful elicitation and discourages answer dictation\. Finally, through simulated participants, qualitative edge\-case analysis, and a human subject study with 64 participants, we show that IntElicit elicits higher\-quality creative outputs and adapts to diverse participant behaviors such as reticence and digression\.

## 2Related Work

### 2\.1Creativity Assessment

Creativity assessment has traditionally relied on psychometric instruments focusing on Divergent Thinking \(DT\)\. The most widely adopted paradigms include the Alternative Uses Task \(AUT\)\(Runco and Acar,[2012](https://arxiv.org/html/2606.12086#bib.bib8)\), Realistic Presented Problems \(RPP\)\(Chand and Runco,[1993](https://arxiv.org/html/2606.12086#bib.bib20)\)and the Torrance Tests of Creative Thinking \(TTCT\)\(Torrance,[1966](https://arxiv.org/html/2606.12086#bib.bib2)\)\. These tests typically employ static prompts \(e\.g\., “list unusual uses for a brick”\) and evaluate responses based on fluency, flexibility, and originality\. With the advent of computational linguistic technology including LLMs, recent work has sought to automate the scoring of these tests using semantic distance metrics\(Beaty and Johnson,[2021](https://arxiv.org/html/2606.12086#bib.bib25)\)or by prompting LLMs as evaluators\(Luchiniet al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib26); Organisciaket al\.,[2023](https://arxiv.org/html/2606.12086#bib.bib27); Kernet al\.,[2024](https://arxiv.org/html/2606.12086#bib.bib28)\), significantly improving assessment efficiency\. For a comprehensive review of automated creativity assessment, please refer to\(Bahget al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib29)\)\.

However, critics argue that traditional DT tests lackecological validity, as they divorce creative thinking from the complex, domain\-specific contexts found in real\-world problem solving\(Zenget al\.,[2011](https://arxiv.org/html/2606.12086#bib.bib21); Baer,[2015](https://arxiv.org/html/2606.12086#bib.bib23)\)\. To bridge this gap, Realistic Presented Problems \(RPP\)\(Chand and Runco,[1993](https://arxiv.org/html/2606.12086#bib.bib20)\)and Situational Judgment Tests \(SJTs\)\(Herdeet al\.,[2019](https://arxiv.org/html/2606.12086#bib.bib30)\)were introduced to simulate more practical scenarios\. The Future Problem Solving Program \(FPSP\)\(Crabbe,[1982](https://arxiv.org/html/2606.12086#bib.bib17),[1989](https://arxiv.org/html/2606.12086#bib.bib18); Torranceet al\.,[1976](https://arxiv.org/html/2606.12086#bib.bib19)\)represents a significant advancement in this direction, employing immersive, multi\-stage futuristic scenarios to evaluate participants’ ability to identify challenges and propose innovative solutions within a constrained narrative\.

Despite the high fidelity of FPSP\-style assessments, they face a critical challenge when automated: theconfounding of creativity with cognitive and agential factors\(Runco and Chand,[1995](https://arxiv.org/html/2606.12086#bib.bib22)\)\. A participant’s failure to produce a creative solution may stem from a lack of domain knowledge, task comprehension, confidence, or willingness to elaborate rather than a lack of creative potential\. Human interviewers can partially address this issue by dynamically clarifying the task, encouraging elaboration, and redirecting attention to relevant scenario constraints\. However, existing computational methods typically treat creativity assessment as a static input\-output mapping, leaving these non\-creative confounders unaddressed\. This limitation becomes more salient as creative problem solving increasingly occurs in AI\-mediated environments, where interaction with intelligent tools is part of authentic practice\(Rezwana and Maher,[2023](https://arxiv.org/html/2606.12086#bib.bib45); Noy and Zhang,[2023](https://arxiv.org/html/2606.12086#bib.bib46)\)\. Our work therefore proposes an interactive elicitation framework that acts as a constrained AI interviewer: it reduces the influence of non\-creative confounders through multi\-turn knowledge and agency scaffolding, while preserving the participant as the source of the creative ideas being evaluated\.

### 2\.2Multi\-turn Dialogue Optimization

Optimizing dialogue agents for purposeful interaction has evolved from supervised fine\-tuning \(SFT\) to Reinforcement Learning from Human/AI Feedback \(RLHF/RLAIF\)\(Ouyanget al\.,[2022](https://arxiv.org/html/2606.12086#bib.bib14); Baiet al\.,[2022a](https://arxiv.org/html/2606.12086#bib.bib31),[b](https://arxiv.org/html/2606.12086#bib.bib32); Leeet al\.,[2024](https://arxiv.org/html/2606.12086#bib.bib33)\)\. In complex tasks, standard RLHF often struggles with sparse rewards, where feedback is only available at the end of a long conversation\(Lightmanet al\.,[2024](https://arxiv.org/html/2606.12086#bib.bib15); Wuet al\.,[2023](https://arxiv.org/html/2606.12086#bib.bib34); Ammanabroluet al\.,[2021](https://arxiv.org/html/2606.12086#bib.bib35)\)\. To address this, recent research focuses on process supervision and reward decomposition\.Leeet al\.\([2025](https://arxiv.org/html/2606.12086#bib.bib36)\)proposed aligning agents with global feedback via reward decomposition, effectively breaking down long\-term goals into dense, turn\-level signals\. Similarly,Yuet al\.\([2025](https://arxiv.org/html/2606.12086#bib.bib37)\)introduced Sotopia\-RL, which optimizes social intelligence in multi\-turn interactions by designing rewards that capture the nuances of social goals and information exchange\.

Furthermore, the role of dialogue agents is shifting from passive information retrieval to active collaboration\. CollabLLM\(Wuet al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib38)\)highlights the importance of transitioning agents from passive responders to active collaborators capable of initiating structure and guiding joint problem\-solving\. In the educational domain, this aligns with the concept of instructional scaffolding, where agents provide temporary support to guide learners\(Daiet al\.,[2023](https://arxiv.org/html/2606.12086#bib.bib39)\)\. However, applying these optimization techniques to creativity assessment presents a uniquealignment tax\(Ouyanget al\.,[2022](https://arxiv.org/html/2606.12086#bib.bib14)\): the agent must elicit the participant’s ideas without dictating the answer \(reward hacking\)\. Unlike math or coding tasks with verifiable solutions\(Lightmanet al\.,[2024](https://arxiv.org/html/2606.12086#bib.bib15); Mroueh,[2025](https://arxiv.org/html/2606.12086#bib.bib40)\), creative elicitation requires a delicate balance between guidance and open\-endedness\. IntElicit bridges this gap by employing a decomposed process reward specifically designed to penalize idea\-dictation while rewarding the elicitation of participant agency and knowledge application\.

## 3Problem Formulation of Interactive Creativity Elicitation

Our primary objective is to develop an adaptive AI interviewer capable of eliciting the creative potential of participants with diverse personality traits\. We formulate this task as a dialogue policy optimization problem within an open\-ended environment\. Unlike traditional tasks with immediate feedback, creativity elicitation is characterized by sparse and subjective rewards, where the quality of ideas is only observable after the interaction concludes\.

Formally, we model the interaction environment using a set of diverse LLM\-based simulated participants, denoted asUU\. The interaction unfolds overnnturns, generating a trajectory𝒯=\{τ1,…,τn\}\\mathcal\{T\}=\\\{\\tau\_\{1\},\\dots,\\tau\_\{n\}\\\}\. Each turn consists of a pairτi=\(τi,A,τi,U\)\\tau\_\{i\}=\(\\tau\_\{i,A\},\\tau\_\{i,U\}\), whereτi,A\\tau\_\{i,A\}is the utterance generated by the AI interviewer’s policyπθ​\(τi,A\|𝒯<i\)\\pi\_\{\\theta\}\(\\tau\_\{i,A\}\|\\mathcal\{T\}\_\{<i\}\)given the history𝒯<i=\{τ1,…,τi−1\}\\mathcal\{T\}\_\{<i\}=\\\{\\tau\_\{1\},\\dots,\\tau\_\{i\-1\}\\\}, andτi,U\\tau\_\{i,U\}is the subsequent response from the simulated participantu∼Uu\\sim U\.

Upon the completion of the assessment, the environment yields a multi\-dimensional sparse reward vector\[RNove,RComp,RFlex,RAppr\]\[R\_\{\\text\{Nove\}\},R\_\{\\text\{Comp\}\},R\_\{\\text\{Flex\}\},R\_\{\\text\{Appr\}\}\], corresponding to four creativity dimensions, namely,Novelty,Complexity,Flexibility, andAppropriateness, respectively\. The reward vector is determined by𝐑​\(⋅\)\\mathbf\{R\}\(\\cdot\): an LLM\-as\-a\-Judge mechanism grounded in a psychological indicator system\.

To optimize the policyπθ\\pi\_\{\\theta\}, we define a scalar evaluation functionAHP​\(⋅\)\\text\{AHP\}\(\\cdot\)that aggregates these dimensions via a weighted sum, where the specific coefficients are derived using the Analytic Hierarchy Process \(AHP\)\(Saaty,[1980](https://arxiv.org/html/2606.12086#bib.bib44)\)detailed in Section[4\.3](https://arxiv.org/html/2606.12086#S4.SS3)\. The optimization objective is to find the parametersθ∗\\theta^\{\*\}that maximize the expected creative performance over the distribution of participants:

θ∗=arg⁡maxθ𝔼u∼U,τ∼πθ​\[AHP​\(𝐑​\(𝒯\)\)\]\.\\theta^\{\*\}=\\mathop\{\\arg\\max\}\_\{\\theta\}\\mathbb\{E\}\_\{u\\sim U,\\tau\\sim\\pi\_\{\\theta\}\}\\left\[\\text\{AHP\}\(\\mathbf\{R\}\(\\mathcal\{T\}\)\)\\right\]\\,\.\(1\)

## 4The Proposed IntElicit Framework

This section presents IntElicit, a framework designed to adaptively foster participant creativity through dialogue\. Our approach begins with anInteractive Forward Sampling, where an LLM\-based interviewer performs multi\-turn rollouts with diverse simulated participants across complex realistic scenarios\. An LLM\-as\-a\-Judge evaluator grounded in expert psychological guidelines is then employed to evaluate these interactions\. Subsequently, we introduce theTurn\-level Reward Decompositionmechanism to deconstruct holistic feedback into fine\-grained signals corresponding to specific interaction dimensions\. Finally, we discuss theDialogue Policy Optimizationstrategy, which leverages these decomposed rewards to enable the interviewer to adaptively elicit the creative potential of participants with diverse personas\. An overview of IntElicit is shown in Figure[2](https://arxiv.org/html/2606.12086#S4.F2)\.

![Refer to caption](https://arxiv.org/html/2606.12086v1/x2.png)Figure 2:Schematic architecture of the IntElicit training pipeline\.\(a\) Data generation: Participant simulators with different engagement personas interact with the interviewer across expert\-designed contextualized scenarios to generate multi\-turn trajectories\.\(b\) Reward construction: Forward interaction sampling and an expert\-guided LLM\-as\-a\-Judge produce decomposed process rewards, with scenario\-level dimension weights derived through AHP\.\(c\) Policy optimization: The interviewer policy is first initialized with supervised fine\-tuning on expert trajectories and then optimized with PPO guided by the learned local reward model\.### 4\.1Participant Simulator

To simulate realistic interactions, IntElicit defines a simulator function𝒮u\\mathcal\{S\}\_\{u\}for each specific participant personau∈Uu\\in U\. Formally, this simulator is modeled as a mapping function𝒮u:\(𝒯<i,τi,A\)→τi,U\\mathcal\{S\}\_\{u\}:\(\\mathcal\{T\}\_\{<i\},\\tau\_\{i,A\}\)\\to\\tau\_\{i,U\}, which yields the participant responseτi,U\\tau\_\{i,U\}based on the history𝒯<i\\mathcal\{T\}\_\{<i\}and the current interviewer utteranceτi,A\\tau\_\{i,A\}\. We operationalize this simulation by injecting expert\-curated prompts into an LLM, conditioning it to generate utterances that strictly adhere to the designated personas, linguistic styles, and behavioral archetypes associated with the personauu\. This ensures that the simulator consistently maintains its character constraints throughout the dialogue, providing a robust environment for training the adaptive interviewer\.

The simulator is used as a training and stress\-testing environment, not as a substitute for human validation\. Its main function is to expose the interviewer policy to repeated interactional variation that would be prohibitively expensive to collect from human participants at scale\. Each persona prompt specifies the participant’s response length, willingness to elaborate, level of initiative, and tendency to remain on topic\. During interaction, the simulator receives the full dialogue history and the current interviewer utterance, which allows it to produce context\-sensitive replies rather than independent single\-turn responses\. This design makes the policy optimization problem closer to an educational interview: the interviewer must adapt to the participant’s previous reasoning, decide when to ask for elaboration, and avoid collapsing into answer dictation\.

### 4\.2Forward Interaction Sampling

In this framework, we employ a Monte Carlo\-based forward sampling approach, extending conversations turn\-by\-turn until a termination criterion is met \(either reaching the maximum turn limit or the completion of the interaction\)\. To construct optimal interaction trajectories, we must evaluate and select the best response generated by the interviewer at each turnii\. However, computing a full\-trajectory reward for every candidate is computationally prohibitive\. Therefore, IntElicit introduces a hyperparameter of look\-ahead window sizeww\(look aheadwwturns\) to constrain the forward simulation\. This strategy significantly reduces computational overhead while retaining sufficient contextual information for accurate evaluation\. The process is detailed as follows:

#### Candidate Response Generation\.

At theii\-th interaction turn, given the dialogue history𝒯<i=\{τ1,…,τi−1\}\\mathcal\{T\}\_\{<i\}=\\\{\\tau\_\{1\},\\dots,\\tau\_\{i\-1\}\\\}and the current participant utteranceτi,U\\tau\_\{i,U\}, we perform parallel sampling ofKKcandidate responses from the interviewer policyπθ\\pi\_\{\\theta\}:\{τi,A\(k\)\}k=1K∼πθ\(⋅\|𝒯<i,τi,U\)\\\{\\tau\_\{i,A\}^\{\(k\)\}\\\}\_\{k=1\}^\{K\}\\sim\\pi\_\{\\theta\}\(\\cdot\|\\mathcal\{T\}\_\{<i\},\\tau\_\{i,U\}\)\.

#### Forward Interaction Sampling\.

For each candidate responseτi,A\(k\)\\tau\_\{i,A\}^\{\(k\)\}, we sampleNNfuture trajectories via alternating interactions between the participant simulator𝒮u\\mathcal\{S\}\_\{u\}and the interviewer policyπθ\\pi\_\{\\theta\}\. Specifically, for a participant with personau∈Uu\\in U, the simulation follows the iterative process:

τi\+l,U\(k\)\\displaystyle\\tau\_\{i\+l,U\}^\{\(k\)\}=𝒮u​\(𝒯<i\+l\(k\),τi\+l,A\(k\)\),\\displaystyle=\\mathcal\{S\}\_\{u\}\(\\mathcal\{T\}\_\{<i\+l\}^\{\(k\)\},\\tau\_\{i\+l,A\}^\{\(k\)\}\)\\,,\(2\)τi\+l\+1,A\(k\)\\displaystyle\\tau\_\{i\+l\+1,A\}^\{\(k\)\}∼πθ\(⋅∣𝒯<i\+l\(k\)∪\{τi\+l,U\(k\)\}\),\\displaystyle\\sim\\pi\_\{\\theta\}\(\\cdot\\mid\\mathcal\{T\}\_\{<i\+l\}^\{\(k\)\}\\cup\\\{\\tau\_\{i\+l,U\}^\{\(k\)\}\\\}\)\\,,\(3\)wherel=0,1,…,w−1l=0,1,\\ldots,w\-1\. Each sampled trajectory segment is defined as𝒯i:i\+w\(k\)=\{τi\(k\),…,τi\+w\(k\)\}\\mathcal\{T\}\_\{i:i\+w\}^\{\(k\)\}=\\\{\\tau\_\{i\}^\{\(k\)\},\\dots,\\tau\_\{i\+w\}^\{\(k\)\}\\\}\.

#### Optimal Response Selection\.

We compute the multi\-dimensional reward vector for each sampled trajectory, the overall reward is then aggregated using AHP\. Finally, we select the candidate with the highest expected score as the output for the current turnτi,A∗=τi,A\(k^\)\\tau\_\{i,A\}^\{\*\}=\\tau\_\{i,A\}^\{\(\\hat\{k\}\)\}wherek^=arg⁡maxk⁡AHP​\(𝐑​\(𝒯i:i\+w\(k\)\)\)\\hat\{k\}=\\arg\\max\_\{k\}\\text\{AHP\}\(\\mathbf\{R\}\(\\mathcal\{T\}\_\{i:i\+w\}^\{\(k\)\}\)\)\.

### 4\.3Turn\-level Reward Decomposition

In open\-ended creative dialogues, relying solely on episode\-level sparse rewards presents two significant challenges\. First, thesparsityof the signal makes policy optimization notoriously difficult, as the agent receives no feedback on intermediate steps essential for guiding the conversation\. Second, optimizing for a final outcome often leads toreward hacking\. For instance, an interviewer might dictate a high\-quality solution to quickly terminate the task\. While this maximizes the outcome score, it stifles the participant’s proactive inquiry and prevents the emergence of original, self\-generated insights\. To mitigate these issues, we propose a mechanism that decomposes the holistic objective into fine\-grained, turn\-level supervision\.

#### Scenario\-Adaptive Weighting via AHP\.

To prevent the policy from collapsing into narrow optimization, we first identify four critical dimensions of creativity, represented as the reward vector\[RNove,RFlex,RComp,RAppr\]\[R\_\{\\text\{Nove\}\},R\_\{\\text\{Flex\}\},R\_\{\\text\{Comp\}\},R\_\{\\text\{Appr\}\}\]\. However, the importance of these dimensions varies significantly across contexts\. To derive rational weights, we employ the Analytic Hierarchy Process \(AHP\), engaging two expert psychologists to prioritize these dimensions within two distinct scenario categories:Ecological\-EnvironmentandTechnological\-Society\. The consistency of these judgments was verified using the Consistency Ratio \(CR\), ensuringCR<0\.1\\text\{CR\}<0\.1\. Based on the derived priority vectors, we obtain scenario\-specific weights𝝀\(c\)\\boldsymbol\{\\lambda\}^\{\(c\)\}, normalized such that∑jλj\(c\)=1\\sum\_\{j\}\\lambda\_\{j\}^\{\(c\)\}=1\. Consequently, the final scalar training objectiveRtotalR\_\{\\text\{total\}\}for a trajectory𝒯\\mathcal\{T\}is formulated as the weighted dot product:

Rtotal​\(𝒯\)=AHP​\(𝐑​\(𝒯\)\)=𝝀\(c\)⋅𝐑​\(𝒯\)\.R\_\{\\text\{total\}\}\(\\mathcal\{T\}\)=\\text\{AHP\}\(\\mathbf\{R\}\(\\mathcal\{T\}\)\)=\\boldsymbol\{\\lambda\}^\{\(c\)\}\\cdot\\mathbf\{R\}\(\\mathcal\{T\}\)\\,\.\(4\)This mechanism allows the reward function to dynamically balance divergent facets based on scenario requirements\. For instance, as detailed in our expert calibration \(Table[1](https://arxiv.org/html/2606.12086#S4.T1)\), the system prioritizesComplexityinEcologicalcontexts to address the systemic interconnectedness of environmental issues, while emphasizingNoveltyinTechnologicalscenarios to encourage innovative solution generation\.

Table 1:AHP\-derived weights for the four creativity dimensions in the two scenario categories\.Scenario CategoryNoveltyComplexityAppropriatenessFlexibilityEcological and Environment0\.14160\.30890\.24070\.3089Technological and Society0\.31250\.06250\.31250\.3125
#### Turn\-level Process Decomposition\.

While AHP addresses “what” to evaluate regarding the final outcome, it does not address “when” to reward effective scaffolding\. A truly creative spark often requires a strategic build\-up rather than emerging immediately after a prompt\. Relying solely on sparse, outcome\-based rewards often fails to capture these intermediate pedagogical achievements\. To address this, we implement aTurn\-level Process Decompositionmechanism\. Specifically, we employ an LLM\-as\-a\-Judge equipped with specificprocess\-oriented prompts\(Appendix Section[B\.2](https://arxiv.org/html/2606.12086#A2.SS2)\) to directly evaluate the quality of the interaction at each turn\. Formally, the decomposed process rewardr​\(τi,A\)r\(\\tau\_\{i,A\}\)forτi,A\\tau\_\{i,A\}is determined by history𝒯<i\\mathcal\{T\}\_\{<i\}and the following participant responseτi,U\\tau\_\{i,U\}:

r​\(τi,A\)=𝐑process​\(𝒯<i∪τi,U\),r\(\\tau\_\{i,A\}\)=\\mathbf\{R\}\_\{\\text\{process\}\}\(\\mathcal\{T\}\_\{<i\}\\cup\\tau\_\{i,U\}\)\\,,\(5\)where𝐑process​\(⋅\)\\mathbf\{R\}\_\{\\text\{process\}\}\(\\cdot\)denotes the scoring function of the LLM\-Judger according to the process\-oriented prompts\.

This decomposition is crucial because it aligns the reward signal with pedagogical goals rather than answer\-dictation\. For example, an interviewer’s prompt might not lead to an immediate solution, but if it successfully encourages the student toraise a high\-quality question,𝐑process\\mathbf\{R\}\_\{\\text\{process\}\}assigns a high reward to this behavior\. By explicitly valuing these intermediate indicators of agency and inquiry, IntElicit ensures that the policy learns to foster sustained creative engagement, even if the final solution is yet to be formed\.

The process\-scoring prompt asks the judge to evaluate whether the interviewer response creates conditions for participant\-generated creativity\. In particular, it rewards prompts that help the participant notice overlooked constraints, connect ideas across parts of the scenario, consider alternative stakeholders or causal mechanisms, and articulate reasons for a challenge or solution\. It penalizes responses that reveal the answer, impose a complete solution, over\-constrain the participant’s reasoning path, or drift away from the scenario\. The final\-scoring prompt, by contrast, evaluates the participant’s overall creative output after the interaction using Novelty, Flexibility, Complexity, and Appropriateness\. Separating these two scoring stages is important because a pedagogically useful interviewer turn may not immediately increase the final answer score, while a direct solution\-giving turn may superficially improve the final answer but undermine the assessment construct\.

### 4\.4Dialogue Policy Optimization

We optimize the interviewer policyπθ\\pi\_\{\\theta\}using a two\-phase framework\. First, the expert trajectories generated via forward sampling with decomposed reward are used to train a local reward model, which subsequently guides the RL training of the policy\.

#### Training Local Reward Model\.

Since performing inference of LLM\-Judge during online training is computationally prohibitive, we first distill the insights from the trajectory evaluations into a lightweight local reward model𝐑ϕ​\(⋅\)\\mathbf\{R\}\_\{\\phi\}\(\\cdot\)which function as a proxy of𝐑process​\(⋅\)\\mathbf\{R\}\_\{\\text\{process\}\}\(\\cdot\)\. We construct a labeled dataset using the interaction trajectories and their corresponding decomposed rewards collected in Section[4\.2](https://arxiv.org/html/2606.12086#S4.SS2)\. The model𝐑ϕ​\(⋅\)\\mathbf\{R\}\_\{\\phi\}\(\\cdot\)is trained to map the current dialogue context𝒯<i\\mathcal\{T\}\_\{<i\}and a candidate utteranceτi,A\\tau\_\{i,A\}to the expected future reward\. By minimizing the regression loss against the aggregated trajectory scores,𝐑ϕ​\(⋅\)\\mathbf\{R\}\_\{\\phi\}\(\\cdot\)internalizes the expert\-derived pedagogical priorities\. Once trained, this model serves as a computationally efficient proxy, providing dense, turn\-level feedback that reflects long\-term creative elicitation potential without requiring real\-time simulation\.

#### Elicitation\-Oriented Policy Training\.

With the reward model𝐑ϕ​\(⋅\)\\mathbf\{R\}\_\{\\phi\}\(\\cdot\), the optimization of the interviewer policy proceeds in two stages:

- •Supervised Fine\-Tuning \(SFT\):As a behavioral warm\-up, we fine\-tune the base LLM using the expert trajectories𝒯∗\\mathcal\{T\}^\{\*\}collected with an expert LLM as the interviewer during the forward sampling phase\. This allows the model to perform behavioral cloning, internalizing fundamental linguistic patterns and scaffolding strategies of a supportive interviewer\.
- •Online RL:To further enhance the model’s ability to handle unseen dynamics, IntElicit employs Proximal Policy Optimization \(PPO\)\(Schulmanet al\.,[2017](https://arxiv.org/html/2606.12086#bib.bib42)\)\. In this stage, the policyπθ\\pi\_\{\\theta\}interacts with the participant simulator, generating responsesti,At\_\{i,A\}that are evaluated by the local reward model𝐑ϕ​\(⋅\)\\mathbf\{R\}\_\{\\phi\}\(\\cdot\)\. By maximizing the expected cumulative reward, the policy evolves from merely mimicking static scripts to actively navigating complex pedagogical dynamics, learning to function as a elicitation strategy for creative engagement\.

## 5Experiment

We evaluate IntElicit from four perspectives: \(1\) assessing its effectiveness using diverse simulated participants, \(2\) validating its real\-world elicitation performance through a human subject study, \(3\) confirming its robustness against diverse edge cases through qualitative analysis, and \(4\) analysis of learned dialogue strategies\. Our code is available at[https://github\.com/MingjiaLi666/IntElicit](https://github.com/MingjiaLi666/IntElicit)\.

### 5\.1Experimental Setup

#### Model Configurations\.

We adopt Qwen3\-8B as the base LLM for policy training and for generating embeddings for the local reward functions\. For the Forward Interaction Sampling phase, we utilize Qwen3\-235B as the base model for the participant simulator, Gemini\-3\-Pro empowered LearnLM\(Google,[2024](https://arxiv.org/html/2606.12086#bib.bib43)\)as the interviewer for expert trajectories collecting, and DeepSeek\-V3\.1 as the expert judge for scoring\. Detailed specifications for all LLMs employed in this study are provided in Appendix Table[A\.1](https://arxiv.org/html/2606.12086#A1.SS1)\.

The models used in the study serve distinct methodological roles rather than constituting a single undifferentiated model pool\. Qwen3\-8B was selected as the trainable interviewer backbone because it provides a realistic setting for deploying an adaptive assessment policy on a moderately sized open model\. Larger models were used only where stronger reasoning capacity was needed to construct or evaluate training data\. Specifically, Qwen3\-235B was used to simulate participants because simulator utterances must remain coherent over multi\-turn interactions while consistently following persona constraints; Gemini\-3\-Pro with LearnLM capabilities was used to generate expert\-style trajectories because the training data should reflect pedagogical scaffolding rather than ordinary chatbot behavior; and DeepSeek\-V3\.1 was used as the primary automated judge because final and process scoring require stable rubric\-following behavior\. In addition, GPT\-4o, Gemini\-3\-Pro, Qwen3\-Max, DeepSeek\-R1, Qwen3\-235B, Llama\-3\.3\-70B, CollabLLM, Sotopia\-RL, and the unoptimized Qwen3\-8B base model were included as comparison systems\. This separation of roles helps avoid an overly favorable comparison in which the same model both generates and evaluates all outcomes\.

#### Automated Evaluation and Validation\.

Because both outcome rewards and process rewards involve LLM\-based evaluation, we conducted additional checks to reduce the risk that the reported gains merely reflect overfitting to a single judge model\. First, the evaluation prompts were designed around the expert\-defined creativity rubric rather than generic preference instructions, so that the automated judge was constrained to assess Novelty, Flexibility, Complexity, Appropriateness, and process quality in relation to the contextualized task\. Second, we examined rank consistency across alternative strong judge models\. In the rebuttal analysis, DeepSeek\-V3\.1 showed high agreement with Llama\-3\.3\-70B\-Instruct \(τ=0\.86\\tau=0\.86\) and Qwen3\-Max \(τ=0\.91\\tau=0\.91\), suggesting that the simulation results were not specific to one evaluator architecture\. Third, the human subject study provides an independent validation channel: expert raters ranked outputs produced in real participant interactions under a double\-blind protocol, and these human rankings were not used for policy training\.

#### Scenarios and Simulators\.

This paper utilizes a set of 16 expert\-designed open\-ended scenarios classified into two categories:Ecological\-EnvironmentandTechnological\-Society\. Building upon these contexts, we employ Qwen3\-235B to simulate three distinct participant personas \(e\.g\.,Talkative,Normal,Quiet\)\. The detailed prompts defining these behaviors are provided in Appendix Section[B\.1](https://arxiv.org/html/2606.12086#A2.SS1)\.

The 16 scenarios were adapted from a future\-oriented problem\-solving format in which participants must identify challenges embedded in a complex social, environmental, or technological situation rather than simply list unusual uses or generate decontextualized ideas\. TheEcological\-Environmentcategory includes Ocean Soup, Zero Waste Initiative, Human Impact on the Environment, Water Supply, Toxic Substances, Infectious Diseases Transmission, Insects as Food, Agricultural Industry, and Terraforming\. TheTechnological\-Societycategory includes Biosecurity, Antibiotic Resistance, Neurotechnology, Drones, Criminal Justice System, Gamification, and Living in Poverty\. These scenarios were selected to cover problems that require participants to reason about systems, stakeholders, constraints, and possible unintended consequences\. This design is important for contextualized creativity assessment because the target construct is not the rapid production of isolated ideas, but the ability to explore a realistic problem space and articulate creative challenges or solutions within it\. Scenario classifications and representative prompts are provided in Appendix Table[A\.2](https://arxiv.org/html/2606.12086#A1.SS2)and Appendix Section[A\.2](https://arxiv.org/html/2606.12086#A1.SS2)\.

The simulator personas were designed to approximate interactional patterns commonly encountered in educational interviews\. TheQuietpersona gives short or low\-information responses and therefore tests whether the interviewer can elicit further reasoning without supplying answers\. TheNormalpersona provides cooperative but moderate responses, approximating a participant who follows instructions but does not require extensive redirection\. TheTalkativepersona provides many ideas, sometimes diffusely, and therefore tests whether the interviewer can preserve participant agency while organizing and deepening the conversation\. These personas are not intended to represent demographic groups\. Rather, they operationalize engagement styles that affect how much creative reasoning becomes observable during assessment\.

#### Baselines\.

This paper compares IntElicit against four categories of baselines, covering both static creativity assessment and interactive dialogue\-based systems:

- •Static FPSP Assessment\.This baseline follows the original FPSP\-style questionnaire format, where participants receive the scenario and complete the creativity task without multi\-turn interaction or adaptive scaffolding\. It represents a conventional static contextualized creativity assessment setting and allows us to examine whether interactive elicitation provides benefits beyond the original FPSP paradigm\.
- •Expert\-designed Dialogue Policy with Distinct Foundation Models\.A sophisticated prompt\-based workflow that simulates expert pedagogical strategies \(detailed in Appendix Section[B\.4](https://arxiv.org/html/2606.12086#A2.SS4)\) using state\-of\-the\-art open\-source \(e\.g\., Llama\-3\.3, Qwen3\) and closed\-source \(e\.g\., GPT\-4o, Gemini\-3\-Pro\) LLMs\. This baseline tests whether IntElicit improves over strong manually designed interactive scaffolding\.
- •Sotopia\-RL\.\(Yuet al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib37)\)A reinforcement\-learning\-based framework designed to optimize social intelligence and information exchange in multi\-turn interactions\. This baseline tests whether general\-purpose social interaction optimization transfers to creativity elicitation\.
- •CollabLLM\.\(Wuet al\.,[2025](https://arxiv.org/html/2606.12086#bib.bib38)\)A collaborative framework that transforms agents into proactive partners, emphasizing structural guidance and joint problem\-solving\. This baseline tests whether proactive collaboration alone is sufficient for assessment\-oriented creativity elicitation\.

The baselines are designed to separate three sources of comparison\. The static FPSP baseline represents the cognitive science and creativity assessment tradition, where contextualized creativity is measured through a non\-interactive scenario\-based task\. The expert\-designed dialogue policy represents a strong manually engineered interactive interviewer, allowing us to test whether learned policy optimization improves over human\-designed scaffolding\. Sotopia\-RL and CollabLLM provide broader dialogue\-agent baselines from social interaction and collaborative problem solving\. Together, these comparisons allow us to evaluate whether IntElicit improves not only over general\-purpose LLMs, but also over both static contextualized assessment and existing multi\-turn interaction frameworks\.

#### Implementation Details\.

To evaluate IntElicit, we conducted a comparative analysis against the baselines across 16 scenarios \(Appendix Section[A\.2](https://arxiv.org/html/2606.12086#A1.SS2)\)\. Table[2](https://arxiv.org/html/2606.12086#S5.T2)summarizes the average performance across the three simulated personas\.

For trajectory construction, each interaction was limited to a maximum of eight turns, and the forward sampling procedure considered two candidate interviewer responses at each decision point\. Unless otherwise noted, the look\-ahead window was set tow=4w=4, which allowed the system to estimate the downstream effect of a candidate interviewer utterance without requiring full\-episode rollouts for every possible response\. The reward aggregation used the AHP\-derived scenario weights described in Section[4\.3](https://arxiv.org/html/2606.12086#S4.SS3)\. To improve reliability and efficiency, synthetic trajectory generation was performed with parallel sampling and incremental checkpointing so that failed or unstable model calls did not invalidate completed trajectories\.

All task\-facing instructions were standardized to ensure that differences in performance could be attributed primarily to the dialogue policy rather than to inconsistent task framing\. Each participant or simulator was first presented with the same scenario background and was then asked to identify important challenges, explain why those challenges matter, and develop possible lines of reasoning\. The interviewer was instructed to avoid giving complete answers and instead to ask follow\-up questions, request justification, or prompt the participant to consider overlooked aspects of the scenario\. For transparency and reproducibility, the Appendix provides the representative scenario prompts, simulator persona prompts, process\-scoring prompt, final\-scoring prompt, participant instructions, human rater rubric, and dialogue\-strategy classification prompt \(Appendix Sections[A\.2](https://arxiv.org/html/2606.12086#A1.SS2)–[B\.4](https://arxiv.org/html/2606.12086#A2.SS4)\)\. In the main text, we summarize these materials because they are not peripheral implementation details: in an interactive assessment, prompts define the assessment condition itself\.

### 5\.2Evaluation with Diverse Simulators

Table 2:Performance comparison of IntElicit and baselines across 16 scenarios\. The final column summarizes the overall average score\.Boldindicates the best performance, andunderlinedvalues denote the second\-best\. The asterisk \(∗\*\) marks statistically significant improvements \(p<0\.05p<0\.05,tt\-test\) over the second\-best model\.Compared MethodsOcean\.Neuro\.Agri\.Bio\.Justice\.Terra\.AntiBio\.Waste\.Insect\.Infect\.Toxic\.Drones\.Water\.Game\.Poverty\.Env\.MeanGPT\-4o291\.00291\.33290\.00318\.33323\.33325\.33268\.67317\.67347\.67293\.33258\.67304\.33295\.00308\.67315\.33278\.00301\.67Gemini\-3\-Pro253\.33275\.33244\.33276\.00271\.00311\.67326\.67276\.67267\.67256\.00281\.00316\.67293\.67295\.33282\.00238\.33279\.10Qwen3\-Max259\.33266\.67272\.67275\.00323\.33310\.33328\.33300\.33286\.67316\.00282\.00327\.00306\.67280\.33306\.00267\.67294\.27Deepseek\-R1310\.00287\.00343\.33294\.33330\.00301\.67332\.67310\.33298\.33298\.33260\.33273\.67298\.33303\.00248\.00285\.00298\.40Qwen3\-235B245\.33330\.00241\.67297\.00238\.33320\.00302\.67255\.00303\.67266\.67292\.67299\.00285\.00290\.00288\.00277\.00283\.25Llama\-3\.3\-70B254\.33307\.00325\.33286\.67313\.33306\.33346\.00315\.33261\.67300\.33287\.00249\.33313\.67296\.67279\.33253\.33293\.48CollabLLM158\.67125\.00120\.00153\.33116\.67148\.33106\.67135\.33181\.67136\.67156\.00185\.00180\.00139\.67166\.67135\.00146\.54Sotopia\-RL115\.00120\.00115\.00116\.67133\.33126\.67111\.67213\.33201\.67191\.67140\.00120\.00126\.67151\.67188\.33138\.33144\.38Qwen3\-8B \(base\)138\.67143\.67107\.00148\.67153\.67113\.67183\.67163\.67193\.67137\.00125\.33160\.3393\.67160\.33133\.6768\.67132\.86FPSP235\.00222\.67254\.33251\.67251\.67248\.33255\.00288\.33306\.67236\.67240\.00235\.00288\.33211\.00238\.33220\.00248\.94IntElicit320\.00\*334\.33\*351\.33\*320\.33\*325\.33333\.67\*342\.00320\.33\*348\.00\*312\.00303\.67\*332\.33\*319\.00\*326\.67\*324\.00\*302\.33\*325\.71\*

#### Main Results\.

As detailed in Table[2](https://arxiv.org/html/2606.12086#S5.T2), IntElicit achieves the highest overall average score of325\.71, establishing a significant margin over the strongest proprietary baseline, GPT\-4o \(301\.67\), and the leading open\-source model, DeepSeek\-R1 \(298\.40\)\. Specifically, IntElicit secures the top rank in13 out of 16scenarios, with statistically significant improvements \(p<0\.05p<0\.05\) observed in almost all winning cases\. This consistent superiority across diverse domains ranging fromEcological\(e\.g\., Ocean\., Agri\.\) toSocietal\(e\.g\., Poverty\., Game\.\) underscores the framework’s powerful generalization capability\.

In contrast, specialized frameworks like Sotopia\-RL and CollabLLM struggle, hovering around a mean score of≈\\approx145\. This performance gap likely stems from their optimization for social navigation or task completion rather than the open\-ended scaffolding required for creativity\. While powerful general\-purpose LLMs demonstrate competitiveness in specific niches, for instance, Llama\-3\.3\-70B achieves the top score in AntiBio\. \(346\.00\) and DeepSeek\-R1 leads in Justice\. \(330\.00\), they exhibit notable performance fluctuations\. IntElicit, however, maintains a robust performance floor; even in scenarios where it ranks second \(e\.g\., AntiBio\.\), it remains highly competitive \(342\.00\), avoiding the volatility seen in base models\.

In addition to the global averages across complex scenarios \(Table[2](https://arxiv.org/html/2606.12086#S5.T2)\), we further analyzed persona\-specific scores, scenario\-level variation, and dimension\-specific performance in Appendix Tables[A3](https://arxiv.org/html/2606.12086#A3.T3)–[A5](https://arxiv.org/html/2606.12086#A3.T5)and Appendix Tables[A6](https://arxiv.org/html/2606.12086#A3.T6)–[A9](https://arxiv.org/html/2606.12086#A3.T9)\. The results indicate that IntElicit achieves higher average scores than all baselines, demonstrating its ability to adaptively facilitate participants with diverse personalities while maintaining stable performance across complex scenarios\. These additional analyses also reveal that while some baselines excel in specific instances, their performance remains highly inconsistent\. For instance, DeepSeek\-R1 achieves a high Novelty score of 95\.00 in the “Neuro” scenario \(Appendix Table[A6](https://arxiv.org/html/2606.12086#A3.T6)\), yet its Novelty score drops substantially to 53\.33 in the “Env” scenario\. These performance variances suggest that IntElicit maintains greater robustness and consistency, providing stable elicitation for high\-quality creative output regardless of the scenario\.

The persona\-level results provide additional insight into where the adaptive policy is most useful\. For quiet simulated participants, many general\-purpose LLMs either move quickly toward supplying content or fail to elicit enough information for a high\-quality final response\. IntElicit achieves the highest mean score for this condition, suggesting that the learned policy can maintain productive elicitation even when the participant gives sparse input\. For normal participants, IntElicit also maintains the best overall performance, indicating that the policy does not over\-scaffold when the participant is already cooperative\. For talkative participants, the advantage is especially meaningful because verbose responses can create a different assessment challenge: the interviewer must organize and deepen ideas without simply praising or following every digression\. The strong performance in this condition suggests that IntElicit learns not only to draw out missing reasoning but also to structure abundant reasoning\.

The dimension\-level results further show that IntElicit’s improvements are not confined to a single creativity criterion\. Across the Appendix Tables, the model obtains the strongest average performance for Novelty, Complexity, Appropriateness, and Flexibility\. This is important because optimizing only for novelty could produce unusual but impractical ideas, whereas optimizing only for appropriateness could lead to conventional responses\. The decomposed reward appears to support a more balanced profile: participants are encouraged to generate original ideas, consider multiple directions, elaborate mechanisms or consequences, and remain grounded in the scenario\. In contrast, several baseline models show uneven profiles, performing competitively on isolated dimensions or scenarios but dropping sharply elsewhere\. This instability reinforces the value of scenario\-adaptive weighting and process\-level supervision for contextualized creativity assessment\.

#### Hyperparameter Study\.

![Refer to caption](https://arxiv.org/html/2606.12086v1/x3.png)\(a\)Look\-ahead window size \(ww\)\.
![Refer to caption](https://arxiv.org/html/2606.12086v1/x4.png)\(b\)Learning rate \(l​rlr\)\.
![Refer to caption](https://arxiv.org/html/2606.12086v1/x5.png)\(c\)Batch size \(b​sbs\)\.

Figure 3:Hyperparameter sensitivity analysis for IntElicit\. The three panels compare training curves under different look\-ahead window sizes, learning rates, and batch sizes, respectively\.To evaluate the robustness of IntElicit, we conducted sensitivity analyses on three key hyper\-parameters\. The results illustrated in Figure[3](https://arxiv.org/html/2606.12086#S5.F3)lead to the following observations:

- •Sampling Window Size \(ww\):As shown in Figure[3](https://arxiv.org/html/2606.12086#S5.F3)\(a\), thewwsignificantly affects elicitation efficacy\. The optimal setting \(w=4w=4\) facilitates efficient convergence by capturing long\-term interaction dependencies\. In contrast, smaller windows \(w=2w=2\) lead to sluggish reward growth, while the absence of a window \(w=0w=0\) causes severe instability, failing to account for the delayed rewards and long\-term impact of elicitation strategies\.
- •Learning Rate \(l​rlr\):Figure[3](https://arxiv.org/html/2606.12086#S5.F3)\(b\) illustrates that IntElicit achieves peak stability atl​r=5×10−7lr=5\\times 10^\{\-7\}\. Higher learning rates \(e\.g\.,5×10−55\\times 10^\{\-5\}\) introduce stochastic oscillations and prevent the policy from a high\-performing convergence\.
- •Batch Size \(b​sbs\):The impact of batch size is presented in Figure[3](https://arxiv.org/html/2606.12086#S5.F3)\(c\)\. A balanced batch size \(b​s=6bs=6\) provides the best trade\-off between gradient accuracy and update frequency\. While smaller batch sizes \(e\.g\.,b​s=4bs=4\) increase volatility, larger settings \(e\.g\.,b​s=2bs=2\) suffer from a significantly slower convergence rate, failing to reach optimal performance within the 100\-episode training budget\.

#### Ablation Study\.

To validate the contribution of core components, we compared the full IntElicit model with two variants: \(1\)w/o Local Reward\(using only sparse episode\-level rewards\) and \(2\)w/o AHP\(using equal weights\)\. As shown in Figure[4](https://arxiv.org/html/2606.12086#S5.F4), IntElicit \(black line\) consistently outperforms all variants, achieving the highest cumulative reward and smoothest convergence\. The significant drop in the w/o Local Reward variant \(red line\) highlights the necessity of fine\-grained supervision; without proximal rewards, the policy lacks the immediate feedback required to steer elicitation strategies\. Meanwhile, the gap between the w/o AHP variant \(green line\) and the full model confirms that expert\-guided weighting is superior to a uniform approach for balancing complex creative dimensions\.

![Refer to caption](https://arxiv.org/html/2606.12086v1/x6.png)Figure 4:Ablation study of the IntElicit framework\. The curves compare the full model with variants without the local reward model and without AHP\-based scenario\-adaptive weighting\.

### 5\.3Evaluation with Human Participants

To examine whether the elicitation benefits observed in simulation transfer to real learners, we conducted a human subject study \(N=64N=64\) involving undergraduate and graduate students recruited through an on\-campus participant recruitment announcement\. Eligibility required that participants had no prior knowledge of the research purpose\. The sample included 43 male participants \(67% of total; average age = 22\.5\), and no participants were excluded from the final analysis\. Participants were compensated with an equivalent of $5 USD for approximately 30 minutes of participation\. Consistent with the simulation experiments described in Section[5\.2](https://arxiv.org/html/2606.12086#S5.SS2), this study used the same set of 16 future problem\-solving scenarios\. Participants were fully briefed and provided informed consent prior to the experiment\.

We employed a between\-subjects design to compare the efficacy of our framework against the baseline\. The 64 participants were randomly assigned to two experimental groups \(N=32N=32per group\):

- •The Expert\-designed Dialogue Policy Group: Participants in this group interacted with a static, prompt\-based workflow derived from expert pedagogical strategies\.
- •The IntElicit Group: Participants in this group engaged with our proposed adaptive AI interviewer, which provided dynamic scaffolding and agency elicitation\.

Both groups utilized the same backbone LLM \(Qwen3\-8B\)\. To isolate the impact of the dialogue policy, we aligned the workflow stages between the two groups while varying whether the interviewer followed the expert\-designed prompt workflow or the optimized IntElicit policy\. Within each group, participants were randomly allocated across the 16 scenarios, ensuring that each scenario was completed by two distinct participants per group\.

Because human\-subject data collection and expert ranking are costly, the human study focused on one primary baseline rather than re\-running all simulation baselines with real participants\. We selected the expert\-designed dialogue policy because it provides the most controlled comparison for the assessment setting: it uses the same backbone model, follows expert pedagogical procedures, and allows us to isolate the effect of policy optimization\. Broader agent baselines such as Sotopia\-RL and CollabLLM are retained in the simulation experiments to provide wider comparative context\.

Participants received operational instructions tailored to their assigned condition, with the full instructions provided in Appendix Figures[A9](https://arxiv.org/html/2606.12086#A2.F9)and[A10](https://arxiv.org/html/2606.12086#A2.F10)\. After the interactions, the elicited creative outputs were anonymized and evaluated by six domain expert raters who were recruited specifically for assessment rather than scenario design\. For each scenario, raters received four anonymized responses: two generated by participants in the IntElicit group and two generated by participants in the expert\-designed dialogue policy group\. Raters ranked these four responses along five dimensions: Novelty \(Nove\.\), Flexibility \(Flex\.\), Complexity \(Comp\.\), Appropriateness \(Appr\.\), and Overall Creative Performance\. Each scenario was evaluated by three experts\. We implemented a double\-blind protocol: raters were blinded to both the research purpose and the group from which each response originated\. The rater guidelines and evaluation criteria are provided in Appendix Figure[A11](https://arxiv.org/html/2606.12086#A2.F11)\.

≤\\leq\-1\.0≤\\leq\-0\.33≤\\leq0\.0≤\\leq0\.33≤\\leq0\.67≤\\leq1\.005050100100150150200200250250335531319494178178240240Kendall’sτ\\tauthresholdsCumulative Count≤\\leq\-1\.0≤\\leq\-0\.33≤\\leq0\.0≤\\leq0\.33≤\\leq0\.67≤\\leq1\.005050100100150150200200250250012123737103103193193240240Kendall’sτ\\tauthresholdsCumulative Count

Figure 5:Cumulative distributions of Kendall’sτ\\tauagreement for human\-human and LLM\-human rankings\. The left panel reports human\-human pairwise agreement across scenario–dimension units \(16×5×\(32\)=24016\\times 5\\times\\binom\{3\}\{2\}=240\)\. The right panel reports agreement between the LLM\-as\-a\-Judge \(DeepSeek\-V3\.1\) and each human rater across scenario–dimension units \(16×5×3=24016\\times 5\\times 3=240\)\. Bars show cumulative counts at or below each threshold\.To examine the reliability of the human evaluation and the validity of LLM\-based scoring, we analyzed Kendall’s rank correlation coefficient \(Kendall’sτ\\tau\) for two types of agreement: agreement among expert human raters and agreement between the LLM\-as\-a\-Judge and expert human raters\. Creativity assessment in contextualized, open\-ended tasks is inherently subjective, so we do not expect near\-perfect agreement\. The human\-human agreement is moderate on average \(τ=0\.56\\tau=0\.56,N=240N=240\), indicating meaningful but not uniformly high consistency among experts\. Importantly, the LLM\-human agreement is close to this human\-human level, with an averageτ=0\.52\\tau=0\.52\. Figure[5](https://arxiv.org/html/2606.12086#S5.F5)visualizes the cumulative distributions of the two agreement types\. Although the two panels use different pair definitions, both are computed at the scenario–dimension level and contain 240 Kendall’sτ\\tauvalues\.

The two distributions support a calibrated interpretation of automated evaluation\. For human\-human agreement, 94 out of 240 pairwise comparisons fall at or belowτ=0\.33\\tau=0\.33, meaning that most human rater pairs exceed this low\-agreement threshold, although the distribution is not uniformly high\. For LLM\-human agreement, 103 out of 240 comparisons fall at or belowτ=0\.33\\tau=0\.33, a slightly larger but comparable proportion\. Similarly, 178 human\-human comparisons and 193 LLM\-human comparisons fall at or belowτ=0\.67\\tau=0\.67, suggesting that both distributions are concentrated mainly in the moderate ranges, with a smaller subset reaching high agreement\. Thus, the LLM judge does not perfectly replicate human judgment, but its agreement profile is close to the level of agreement observed among human experts themselves\. This provides empirical support for using the LLM\-as\-a\-Judge as a scalable proxy in the simulation experiments, while still treating human expert rankings as the primary validation evidence for real participant interactions\.

We used ranking rather than direct numerical scoring in the human evaluation because ranking reduces scale\-use differences among raters and better fits the comparative question of whether IntElicit elicits stronger responses than the expert\-designed policy\. For each scenario, experts compared four anonymized responses at the same time, two from each condition\. This setup allowed raters to make relative judgments within a shared scenario context instead of assigning absolute scores across heterogeneous tasks\. The reported mean ranks should therefore be interpreted as comparative quality indicators\. Lower ranks indicate that responses were more often preferred by experts within the same scenario and dimension\.

Table 3:Human evaluation results based on expert ranking\. For each scenario, raters ranked four anonymized responses, including two from the IntElicit group and two from the expert\-designed dialogue policy group\. The table reports mean ranks across raters for each creativity dimension and overall performance; lower mean ranks indicate better perceived response quality\.Dialogue PolicyRater\-IDNove\.Flex\.Comp\.Appr\.OverallExpert\-designedDialogue Policy12\.693\.122\.693\.002\.9422\.562\.882\.622\.692\.7532\.812\.752\.753\.002\.8842\.622\.622\.622\.562\.5652\.812\.812\.752\.752\.8163\.063\.063\.003\.383\.12Mean2\.762\.872\.742\.902\.84IntElicit12\.311\.882\.312\.002\.0622\.442\.122\.382\.312\.2532\.192\.252\.252\.002\.1242\.382\.382\.382\.442\.4452\.192\.192\.252\.252\.1961\.941\.942\.001\.621\.88Mean2\.242\.132\.262\.102\.16

#### Results\.

As shown in Table[3](https://arxiv.org/html/2606.12086#S5.T3), IntElicit achieves lower mean ranks \(i\.e\., better perceived response quality\) across all four dimensions and overall scores compared to the expert\-designed dialogue policy\. Because the human study used rank\-based expert judgments and the inter\-rater agreement is moderate rather than uniformly high, we report these results as descriptive evidence rather than inferential significance claims\. The agreement analysis in Figure[5](https://arxiv.org/html/2606.12086#S5.F5)suggests that the ranking data are usable for comparative interpretation, but it also motivates a cautious reading of the human evaluation results\.

The pattern of mean ranks is also informative at the dimension level\. IntElicit shows the largest descriptive advantage in Flexibility and Appropriateness, suggesting that adaptive elicitation may help participants both explore multiple directions and keep their responses aligned with the scenario constraints\. The advantage in Complexity indicates that the interaction encouraged participants to articulate more detailed causal relations, trade\-offs, or contextual considerations rather than merely naming a challenge\. The smaller but still consistent advantage in Novelty is consistent with the nature of the task: novelty in contextualized problem solving depends not only on unusual ideas but also on whether those ideas can be justified within a realistic problem context\. Overall, the human results support the simulation finding that IntElicit’s benefit is not simply making responses longer; rather, the adaptive interviewer appears to help participants externalize more balanced creative reasoning across multiple assessment dimensions\.

### 5\.4Evaluation with Qualitative Analysis on Edge Cases

To complement quantitative evaluations, we conducted a qualitative case study to assess IntElicit’s pedagogical robustness against challenging participant behaviors\. We designed six distinct “edge cases” that simulate common failure modes for educational AI, such as participants seeking direct answers, digressing from the topic, or exhibiting uncooperative attitudes\. Expert human review confirmed that while most baseline models struggled with these scenarios \(e\.g\., direct answer leakage, ineffective redirection, or passive appeasement\), IntElicit consistently maintained its strategic elicitation stance\. It effectively mitigated reward hacking by providing scaffolding instead of answers, redirected conversations back to the scenario context, and demonstrated resilience against adversarial or perfunctory inputs\. This qualitative analysis validates IntElicit’s ability to support assessment\-oriented elicitation even in complex interactive situations\.

The six edge cases were selected to reflect interactional risks that are particularly important for formative creativity assessment\. In the direct\-answer case, the participant explicitly asks the model to provide the answer, testing whether the interviewer preserves the participant’s role as the source of ideas\. In the knowledge\-gap case, the participant states that they cannot distinguish challenges, testing whether the interviewer can provide directional scaffolding without dictating content\. In the off\-topic case, the participant asks an irrelevant personal question, testing contextual redirection\. In the over\-divergent case, the participant introduces an extreme or weakly grounded association, testing whether the interviewer can narrow the discussion productively\. In the adversarial\-language case, the participant uses hostile language, testing whether the interviewer can remain task\-focused without becoming defensive or overly apologetic\. Finally, in the perfunctory\-response case, the participant gives a generic answer, testing whether the interviewer can deepen superficial reasoning\.

Across these cases, the baseline responses reveal a recurring tension in educational dialogue systems\. Some models are helpful in the ordinary conversational sense because they provide information, but that helpfulness can undermine assessment by supplying the intellectual work that should belong to the participant\. Other models remain polite but fail to return the participant to the scenario or to elicit further reasoning\. IntElicit more consistently follows an assessment\-oriented pattern: it acknowledges the participant’s utterance, reorients the interaction toward the scenario, and asks a question or provides a scaffold that requires the participant to continue reasoning\. These cases do not replace quantitative evaluation, but they illustrate why non\-directive scaffolding is central to the validity of interactive creativity assessment\.

A detailed breakdown of all case studies and the corresponding tables is provided in Appendix Section[D](https://arxiv.org/html/2606.12086#A4)\.

### 5\.5Analysis of Adaptive Dialogue Strategies

To further understand the behavioral patterns learned by IntElicit, we conduct a dialogue strategy analysis of interviewer responses\. The goal is not to treat strategy labels as ground\-truth psychological constructs, but to characterize whether the optimized policy tends to use different scaffolding moves when interacting with participants who display different engagement styles\.

Specifically, we define three categories of dialogue strategies with distinct cognitive functions:

- •Divergent Expansion: It encourages participants to explore multiple directions and generate a broader range of ideas\.
- •Perspective Shifting: It guides participants to reconsider the problem from different roles or viewpoints, facilitating cognitive restructuring\.
- •Evaluative Reflection: It prompts participants to reflect, compare, and refine their existing ideas, thereby strengthening critical thinking\.

In practice, we implement a prompt\-based strategy classifier and employ Qwen3\-235B as the evaluator to label each interviewer response in the interaction logs between IntElicit and three participant simulators \(Talkative,Normal, andQuiet\)\. The analysis includes 384 interviewer responses in total, with 128 responses for each persona condition\. During classification, the dialogue context \(including the selected scenario description, the participant’s previous utteranceτi,U\\tau\_\{i,U\}, and the current AI responseτi\+1,A\\tau\_\{i\+1,A\}\) is provided as input so that the label reflects the functional role of the response within the interaction\. The full prompt used for dialogue strategy classification is provided in Appendix Figure[A12](https://arxiv.org/html/2606.12086#A2.F12)\. Table[4](https://arxiv.org/html/2606.12086#S5.T4)presents the resulting distribution of dialogue strategies across participant personas\.

Table 4:Distribution of classified dialogue strategies across participant personas\. Each persona condition contains 128 IntElicit interviewer responses\.StrategyTalkativeNormalQuietDivergent Expansion52\.3%63\.3%67\.2%Evaluative Reflection32\.0%29\.7%27\.3%Perspective Shifting15\.6%7\.0%5\.5%

![Refer to caption](https://arxiv.org/html/2606.12086v1/x7.png)\(a\)Talkative participant: perspective shifting\.
![Refer to caption](https://arxiv.org/html/2606.12086v1/x8.png)\(b\)Quiet participant: divergent expansion\.

Figure 6:Representative human\-participant dialogue cases illustrating persona\-adaptive strategies\. The left panel shows perspective shifting for a talkative participant, while the right panel shows divergent expansion for a quiet participant\.The results provide behavioral evidence that IntElicit adjusts its elicitation moves across persona conditions\. ForTalkativeparticipants, the model more frequently adoptsPerspective Shiftingstrategies \(15\.6%\), which is over two times higher than forNormal\(7\.0%\) andQuiet\(5\.5%\) participants\. This pattern suggests that the policy may respond to abundant but potentially unfocused ideas by introducing alternative analytical perspectives, helping structure the conversation without suppressing user engagement\.

In contrast, forQuietparticipants, IntElicit shows a stronger preference forDivergent Expansionstrategies \(67\.2%\), exceeding those forNormal\(63\.3%\) andTalkative\(52\.3%\) participants\. This suggests that the policy tends to use more open\-ended prompts when participant responses are brief or low in information, thereby encouraging participants to externalize additional reasoning\.

Meanwhile, the proportion ofEvaluative Reflectionremains relatively stable across all personas \(approximately 27%–32%\), suggesting that this strategy may function as a general\-purpose mechanism for consolidating and refining ideas throughout the interaction process\.

To complement the simulation\-based strategy distribution, we further inspect representative dialogue examples from human participant interactions\. These cases are used illustratively rather than as statistical evidence\.

ForTalkativeparticipants \(Figure[6\(a\)](https://arxiv.org/html/2606.12086#S5.F6.sf1)\), when users present a large number of ideas simultaneously, IntElicit avoids directly filtering or evaluating them\. Instead, it applies Perspective Shifting to reorganize these ideas within a structured analytical frame \(e\.g\., interdependencies among factors\), guiding the participant toward deeper and more systematic exploration\. This approach preserves user contributions while effectively steering the discussion toward a focused direction\.

ForQuietparticipants \(Figure[6\(b\)](https://arxiv.org/html/2606.12086#S5.F6.sf2)\), when users provide brief and low\-information responses, IntElicit does not directly advance toward solutions\. Instead, it actively introduces multi\-dimensional prompts \(e\.g\., toxin tracing, regulatory risks, and brand trust issues\) to expand the scope of reasoning\. This behavior exemplifies the core function of Divergent Expansion, where the model constructs cognitive scaffolds to activate the participant’s latent reasoning processes\.

Taken together, the strategy distribution and representative cases suggest that IntElicit does not rely on a fixed interaction pattern\. Instead, it tends to vary its scaffolding moves with participant engagement style\. This behavioral evidence helps explain why adaptive elicitation can support contextualized creativity assessment: the interviewer can encourage idea generation when responses are sparse, introduce alternative perspectives when responses are diffuse, and prompt reflection when ideas require consolidation\.

### 5\.6Synthesis of Empirical Findings

Across the simulation, human evaluation, and strategy analyses, the results converge on a consistent assessment insight: interactive elicitation helps reveal contextualized creative potential that may remain hidden in static assessment\. The static FPSP baseline captures what participants produce without interaction, whereas IntElicit evaluates what participants can generate when non\-creative confounders such as low confidence, sparse elaboration, or limited contextual understanding are reduced through non\-directive scaffolding\. This does not imply that the AI co\-creates the response or trains creativity during the task; rather, the interaction serves as an assessment condition that makes participant\-generated reasoning more accessible for formative interpretation\.

The empirical results further suggest that adaptive scaffolding matters because different participants require different elicitation moves\. Quiet participants benefit from divergent expansion that encourages idea generation, while talkative participants benefit from perspective shifting that structures diffuse responses\. Together with the LLM\-human agreement analysis, these findings support the use of IntElicit as a formative and diagnostic assessment protocol: it can compare elicited creative outcomes, document interactional barriers, and provide process evidence that static scores alone cannot capture\.

## 6Conclusion & Discussion

This work proposes IntElicit, a framework for contextualized creativity assessment through interactive elicitation\. By leveraging a diverse participant simulator and a decomposed process reward mechanism, we train an adaptive AI interviewer to provide non\-directive assessment scaffolding in multi\-turn creativity tasks\. Extensive experiments, including simulations across 16 scenarios and a human subject study \(N=64N=64\), show that IntElicit elicits higher\-quality creative outputs than expert\-designed baselines while maintaining robustness against interactional challenges such as reticence and digression\. These findings suggest that interactive elicitation can extend contextualized creativity assessment beyond static response collection, especially when the assessment goal is to understand how learners reason under appropriate scaffolding\.

The central construct assessed by IntElicit iselicited creative potential, not unaided creativity and not creativity improvement during the task\. In complex contextualized scenarios, observed performance may be constrained by non\-creative confounders such as limited domain knowledge, low confidence, weak task comprehension, or low willingness to elaborate\. IntElicit uses interaction as an assessment condition: it may clarify context, encourage elaboration, and prompt reflection, but the participant remains responsible for generating, justifying, and refining the creative ideas being evaluated\. This distinction is important for construct validity, because the AI interviewer should be interpreted as a constrained assessment scaffold rather than a co\-author or creativity training tool\.

These findings support interactive elicitation as a formative and diagnostic approach to contextualized creativity assessment\. The framework can provide richer evidence about how learners reason under appropriate scaffolding and may help teachers or researchers identify barriers that static responses alone would obscure\. However, IntElicit should not be used as a stand\-alone high\-stakes summative instrument\. The human study sample was relatively homogeneous, the dialogue strategy labels rely on prompt\-based classification, and the scenarios and rubrics may reflect cultural and linguistic assumptions\. Future work should examine broader learner populations, languages, disciplines, and equity\-related effects of adaptive scaffolding\.

## Limitations

Despite the promising results, our framework has several limitations\. First, Dependence on Base LLMs: The elicitation capability is bound by the reasoning limits of the underlying base models; occasional hallucinations or logic errors may still propagate to the interviewer\. Second, Computational Overhead: The forward interaction sampling used for constructing training data is computationally intensive, which may constrain rapid scalability to new, large\-scale domains without further optimization\. Finally, Modality Constraints: Currently, IntElicit operates solely on textual interactions\. Extending the framework to multimodal settings \(e\.g\., visual or auditory creativity\) remains a critical direction for future research\.

## Ethical Statement

This study involving human participants was reviewed and approved by the Institutional Review Board of the authors’ affiliated university\. The approval number is omitted here and can be provided upon request\. The approval date was 10 December 2025\. Before participation, all participants were informed of the study purpose, procedures, potential risks and benefits, data use, and confidentiality principles\. Written informed consent was obtained from all participants\. All collected data were anonymized and used only for research purposes\.

## References

- How to motivate your dragon: teaching goal\-driven agents to speak and act in fantasy worlds\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Online,pp\. 807–833\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- J\. Baer \(2015\)Domain specificity of creativity\.Academic Press\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- G\. Bahg, S\. Luchini, and R\. Beaty \(2025\)Automated creativity assessment: a review of methods, challenges, and future prospects\.PsyArXiv\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- Y\. Bai, A\. Jones, K\. Ndousse, A\. Askell, A\. Chen, N\. DasSarma, D\. Drain, S\. Fort, D\. Ganguli, T\. Henighan, N\. Joseph, S\. Kadavath, J\. Kernion, T\. Conerly, S\. E\. Showk, N\. Elhage, Z\. Hatfield\-Dodds, D\. Hernandez, T\. Hume, S\. Johnston, S\. Kravec, L\. Lovitt, N\. Nanda, C\. Olsson, D\. Amodei, T\. B\. Brown, J\. Clark, S\. McCandlish, C\. Olah, B\. Mann, and J\. Kaplan \(2022a\)Training a helpful and harmless assistant with reinforcement learning from human feedback\.CoRRabs/2204\.05862\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon, C\. Chen, C\. Olsson, C\. Olah, D\. Hernandez, D\. Drain, D\. Ganguli, D\. Li, E\. Tran\-Johnson, E\. Perez, J\. Kerr, J\. Mueller, J\. Ladish, J\. Landau, K\. Ndousse, K\. Lukosiute, L\. Lovitt, M\. Sellitto, N\. Elhage, N\. Schiefer, N\. Mercado, N\. DasSarma, R\. Lasenby, R\. Larson, S\. Ringer, S\. Johnston, S\. Kravec, S\. E\. Showk, S\. Fort, T\. Lanham, T\. Telleen\-Lawton, T\. Conerly, T\. Henighan, T\. Hume, S\. R\. Bowman, Z\. Hatfield\-Dodds, B\. Mann, D\. Amodei, N\. Joseph, S\. McCandlish, T\. Brown, and J\. Kaplan \(2022b\)Constitutional AI: harmlessness from AI feedback\.CoRRabs/2212\.08073\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- R\. E\. Beaty and D\. R\. Johnson \(2021\)Automating creativity assessment with semdis: an open platform for computing semantic distance\.Behavior research methods53\(2\),pp\. 757–780\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- E\. Brynjolfsson and A\. McAfee \(2014\)The second machine age: work, progress, and prosperity in a time of brilliant technologies, 1st edition\.Norton\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p1.1)\.
- I\. Chand and M\. A\. Runco \(1993\)Problem finding skills as components in the creative process\.Personality and Individual differences14\(1\),pp\. 155–162\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- A\. B\. Crabbe \(1989\)The future problem solving program\.\.Educational Leadership7\(1\),pp\. 27–29\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- A\. B\. Crabbe \(1982\)Creating a brighter future: an update on the future problem solving program\.Journal for the Education of the Gifted5\(1\),pp\. 2–11\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- W\. Dai, J\. Lin, H\. Jin, T\. Li, Y\. Tsai, D\. Gasevic, and G\. Chen \(2023\)Can large language models provide feedback to students? A case study on chatgpt\.InProceedings of the 23rd IEEE International Conference on Advanced Learning Technologies,Orem, UT,pp\. 323–325\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p2.1)\.
- Glăveanu and V\. Petre \(2010\)Paradigms in the study of creativity: introducing the perspective of cultural psychology\.New ideas in psychology28\(1\),pp\. 79–93\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p1.1)\.
- L\. T\. Google \(2024\)Learnlm: improving gemini for learning\.arXiv preprint arXiv:2412\.16429\.Cited by:[§5\.1](https://arxiv.org/html/2606.12086#S5.SS1.SSS0.Px1.p1.1)\.
- C\. Herde, F\. Lievens, E\. G\. Solberg, J\. L\. Harbaugh, M\. H\. Strong, and G\. J\. Burkholder \(2019\)Situational judgment tests as measures of 21st century skills: evidence across europe and latin america\.Journal of Work and Organizational Psychology35\(2\),pp\. 65\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- E\. Kasneci, K\. Seßler, S\. Küchemann, M\. Bannert, D\. Dementieva, F\. Fischer, U\. Gasser, G\. Groh, S\. Günnemann, E\. Hüllermeier,et al\.\(2023\)ChatGPT for good? on opportunities and challenges of large language models for education\.Learning and individual differences103,pp\. 102274\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p4.1)\.
- F\. B\. Kern, C\. Wu, and Z\. C\. Chao \(2024\)Assessing novelty, feasibility and value of creative ideas with an unsupervised approach using gpt\-4\.British Journal of Psychology\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- J\. Kruger and D\. Dunning \(1999\)Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self\-assessments\.\.Journal of Personality and Social Psychology77\(6\),pp\. 1121–1134\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1)\.
- D\. W\. Lee, H\. W\. Park, C\. Breazeal, and L\. Morency \(2025\)Aligning dialogue agents with global feedback via large language model reward decomposition\.CoRRabs/2505\.15922\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- H\. Lee, S\. Phatale, H\. Mansoor, T\. Mesnard, J\. Ferret, K\. Lu, C\. Bishop, E\. Hall, V\. Carbune, A\. Rastogi, and S\. Prakash \(2024\)RLAIF vs\. RLHF: scaling reinforcement learning from human feedback with AI feedback\.InProceedings of the 41st International Conference on Machine Learning, ICML,Vienna, Austria,pp\. 26874–26901\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe \(2024\)Let’s verify step by step\.InProceedings of the 12th International Conference on Learning Representations,Vienna, Austria\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p2.1)\.
- M\. A\. Llego \(2022\)21st\-century learning: what it is and why it’s important\.Retrieved September14,pp\. 2022\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p1.1)\.
- S\. A\. Luchini, N\. T\. Maliakkal, P\. V\. DiStefano, A\. Laverghetta Jr, J\. D\. Patterson, R\. E\. Beaty, and R\. Reiter\-Palmon \(2025\)Automated scoring of creative problem solving with large language models: a comparison of originality and quality ratings\.\.Psychology of Aesthetics, Creativity, and the Arts\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- Y\. Mroueh \(2025\)Reinforcement learning with verifiable rewards: GRPO’s effective loss, dynamics, and success amplification\.CoRRabs/2503\.06639\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p2.1)\.
- S\. Noy and W\. Zhang \(2023\)Experimental evidence on the productivity effects of generative artificial intelligence\.Science381\(6654\),pp\. 187–192\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p3.1)\.
- M\. Oliveira, C\. Zednik, G\. Bombaerts, B\. Sadowski, and R\. Conijn \(2025\)Assessing students’ drive: a framework to evaluate learning through interactions with generative ai\.Computers and Education: Artificial Intelligence,pp\. 100497\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p3.1)\.
- OpenAI \(2023\)GPT\-4 technical report\.CoRRabs/2303\.08774\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p4.1)\.
- P\. Organisciak, S\. Acar, D\. Dumas, and K\. Berthiaume \(2023\)Beyond semantic distance: automated scoring of divergent thinking greatly improves with large language models\.Thinking Skills and Creativity49,pp\. 101356\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. Lowe \(2022\)Training language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems 35,New Orleans, LA,pp\. 27730–27744\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p2.1)\.
- D\. L\. Paulhus \(1984\)Two\-component models of socially desirable responding\.\.Journal of personality and social psychology46\(3\),pp\. 598\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems 36,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),New Orleans, LA,pp\. 53728–53741\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p4.1)\.
- J\. Rezwana and M\. L\. Maher \(2023\)Designing creative AI partners with COFI: A framework for modeling interaction in human\-ai co\-creative systems\.ACM Trans\. Comput\. Hum\. Interact\.30\(5\),pp\. 67:1–67:28\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p3.1)\.
- M\. A\. Runco and S\. Acar \(2012\)Divergent thinking as an indicator of creative potential\.Creativity research journal24\(1\),pp\. 66–75\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- M\. A\. Runco and I\. Chand \(1995\)Cognition and creativity\.Educational psychology review7\(3\),pp\. 243–267\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p3.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p3.1)\.
- T\. L\. Saaty \(1980\)The analytic hierarchy process\.McGraw\-Hill\.Cited by:[§3](https://arxiv.org/html/2606.12086#S3.p4.3)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[2nd item](https://arxiv.org/html/2606.12086#S4.I1.i2.p1.3)\.
- P\. J\. Silvia, B\. Wigert, R\. Reiter\-Palmon, and J\. C\. Kaufman \(2012\)Assessing creativity with self\-report scales: a review and empirical evaluation\.\.Psychology of Aesthetics, Creativity, and the Arts6\(1\),pp\. 19\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1)\.
- Z\. Swiecki, H\. Khosravi, G\. Chen, R\. Martinez\-Maldonado, J\. M\. Lodge, S\. Milligan, N\. Selwyn, and D\. Gašević \(2022\)Assessment in the age of artificial intelligence\.Computers and Education: Artificial Intelligence3,pp\. 100075\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p1.1)\.
- M\. Taguma and M\. Barrera \(2019\)OECD future of education and skills 2030: curriculum analysis\.Dispon\. Su Httpswww Oecd Orgeducation2030\-Proj\.–Learn\. Pdf\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p1.1)\.
- E\. P\. Torrance, C\. B\. Bruch, and J\. P\. Torrance \(1976\)Interscholastic futuristic creative problem\-solving\.\.The Journal of Creative Behavior10\(2\),pp\. 117–125\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.
- E\. P\. Torrance \(1966\)Torrance tests of creative thinking: norms\-technical manual\.\.Personnel Press\.Cited by:[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p1.1)\.
- S\. Wu, M\. Galley, B\. Peng, H\. Cheng, G\. Li, Y\. Dou, W\. Cai, J\. Zou, J\. Leskovec, and J\. Gao \(2025\)CollabLLM: from passive responders to active collaborators\.InProceedings of the 42nd International Conference on Machine Learning, ICML,Vancouver, BC\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p2.1),[4th item](https://arxiv.org/html/2606.12086#S5.I1.i4.p1.1)\.
- Z\. Wu, Y\. Hu, W\. Shi, N\. Dziri, A\. Suhr, P\. Ammanabrolu, N\. A\. Smith, M\. Ostendorf, and H\. Hajishirzi \(2023\)Fine\-grained human feedback gives better rewards for language model training\.InAdvances in Neural Information Processing Systems 36,New Orleans,LA,pp\. 59008–59033\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1)\.
- H\. Yu, Z\. Qi, Y\. Zhao, K\. Nottingham, K\. Xuan, B\. P\. Majumder, H\. Zhu, P\. P\. Liang, and J\. You \(2025\)Sotopia\-rl: reward design for social intelligence\.CoRRabs/2508\.03905\.Cited by:[§2\.2](https://arxiv.org/html/2606.12086#S2.SS2.p1.1),[3rd item](https://arxiv.org/html/2606.12086#S5.I1.i3.p1.1)\.
- L\. Zeng, R\. W\. Proctor, and G\. Salvendy \(2011\)Can traditional divergent thinking tests be trusted in measuring and predicting real\-world creativity?\.Creativity research journal23\(1\),pp\. 24–37\.Cited by:[§1](https://arxiv.org/html/2606.12086#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.12086#S2.SS1.p2.1)\.

## Appendix AAdditional Experimental Details

### A\.1LLM Sources

Table[A\.1](https://arxiv.org/html/2606.12086#A1.SS1)provides comprehensive details and the origins of the large language models \(LLMs\) evaluated in this study\. The selection encompasses several prevailing, state\-of\-the\-art models that currently dominate both industrial applications and academic research\. For each model, we specify its provider, parameter scale, and accessibility\.

Table A1

The overview of the information about LLMs used in our experiments, including source type, parameter size, and accessibility\.

ModelAbbreviationsTypeParametersAccessURLDeepseek\-V3\.1Deepseek\-V3\.1Open\-source685BWeights[HuggingFace Link](https://huggingface.co/deepseek-ai/DeepSeek-V3.1)Deepseek\-R1Deepseek\-R1Open\-source685BWeights[HuggingFace Link](https://huggingface.co/deepseek-ai/DeepSeek-R1)GPT\-4oGPT\-4oClosed\-source\-API[OpenAI Link](https://platform.openai.com/docs/models/gpt-4o)Gemini\-3\-ProGemini\-3Closed\-source\-API[Google Link](https://ai.google.dev/gemini-api/docs/models)Qwen3\-8BQwen3\-8BOpen\-source8BWeights[HuggingFace Link](https://huggingface.co/Qwen/Qwen3-8B)Qwen3\-MaxQwen3\-MaxClosed\-source\>\>1TAPI[Qwen Link](https://bailian.console.aliyun.com/)Qwen3\-235B\-A22B\-128KQwen3\-235BOpen\-source235BWeights[HuggingFace Link](https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507)Llama\-3\.3\-70B\-InstructLlama\-3\.3Open\-source70BWeights[HuggingFace Link](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)

### A\.2Detailed information on 16 Realistic Scenarios

The 16 complex realistic scenarios are categorized into two primary themes:Ecological and Environment, andTechnological and Society\. Table[A\.2](https://arxiv.org/html/2606.12086#A1.SS2)presents the classification and associated abbreviations of the 16 scenarios\. The Representative scenarios and the creative problem solving tasks regarding “Ocean Soup”, “Biosecurity” and “Criminal Justice System” are provided in Figure[A1](https://arxiv.org/html/2606.12086#A1.F1), Figure[A2](https://arxiv.org/html/2606.12086#A1.F2)and Figure[A3](https://arxiv.org/html/2606.12086#A1.F3)respectively\.

Table A2

The classification and associated abbreviations of the 16 complex scenarios\.

CategoryScenario Title \(Abbreviations\)Ecological and EnvironmentOcean Soup \(Ocean\.\)Zero Waste Initiative \(Waste\.\)Human Impact on the Environment \(Env\.\)Water Supply \(Water\.\)Toxic Substances \(Toxic\.\)Infectious Diseases Transmission \(Infect\.\)Insects as Food \(Insect\.\)Agricultural Industry \(Agri\.\)Terraforming \(Terra\.\)Technological and SocietyBiosecurity \(Bio\.\)Antibiotic Resistance \(AntiBio\.\)Neurotechnology \(Neuro\.\)Drones \(Drones\.\)Criminal Justice System \(Justice\.\)Gamification \(Game\.\)Living in Poverty \(Poverty\.\)

![Refer to caption](https://arxiv.org/html/2606.12086v1/x9.png)Figure A1:Expert\-designed scenario prompt for the “Ocean Soup” contextualized creativity task\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x10.png)Figure A2:Expert\-designed scenario prompt for the “Biosecurity” contextualized creativity task\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x11.png)Figure A3:Expert\-designed scenario prompt for the “Criminal Justice System” contextualized creativity task\.

## Appendix BPrompts

### B\.1Prompts for Participant Simulator

Figures[A4](https://arxiv.org/html/2606.12086#A2.F4),[A5](https://arxiv.org/html/2606.12086#A2.F5), and[A6](https://arxiv.org/html/2606.12086#A2.F6)present the system prompts for the three simulated participant personas, detailing the specific instructions used to elicit diverse behavioral patterns during the interaction\.

### B\.2Prompts for LLM\-as\-a\-Judge

To ensure precise evaluation of participants responses, we implemented a dual\-stage scoring mechanism\. First, expert\-designed procedural prompts are utilized to assess interactive process data across three dimensions: Novelty \(Nove\.\), Complexity \(Comp\.\), and Appropriateness \(Appr\.\) \(as shown in Figure[A7](https://arxiv.org/html/2606.12086#A2.F7)\)\. Subsequently, a final expert\-defined prompt is employed to conduct a holistic assessment of the participants’ overall performance, incorporating an additional dimension of Flexibility \(Flex\.\) \(as shown in Figure[A8](https://arxiv.org/html/2606.12086#A2.F8)\)\.

### B\.3Prompt for Dialogue Strategy Classification

To analyze the dialogue strategies learned by IntElicit, we design a prompt\-based classifier to categorize each LLM response into predefined dialogue strategy types\. The classification is formulated as a single\-label decision problem, where each response is assigned to one of three categories: Divergent Expansion, Perspective Shifting, and Evaluative Reflection\. The prompt used for dialogue strategy classification is in Figure[A12](https://arxiv.org/html/2606.12086#A2.F12)\.

### B\.4The Workflow Design of Expert\-designed Dialogue Policies

We facilitate interactions with students through an expert\-designed and LLM\-driven workflow\. The detailed workflow design is illustrated in Figure[A13](https://arxiv.org/html/2606.12086#A2.F13)\. This proposed framework demonstrates high versatility and is universally adaptable to all complex scenarios\.

![Refer to caption](https://arxiv.org/html/2606.12086v1/x12.png)Figure A4:System prompt used to simulate the “Normal Participant” persona\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x13.png)Figure A5:System prompt used to simulate the “Quiet Participant” persona\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x14.png)Figure A6:System prompt used to simulate the “Talkative Participant” persona\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x15.png)Figure A7:Expert\-designed prompt for turn\-level process scoring\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x16.png)Figure A8:Expert\-designed prompt for final creative performance scoring\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x17.png)Figure A9:Participant instructions for the expert\-designed dialogue policy group\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x18.png)Figure A10:Participant instructions for the IntElicit group\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x19.png)Figure A11:Instructions and evaluation rubric provided to expert human raters\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x20.png)Figure A12:Prompt used for classifying interviewer responses into dialogue strategy categories\.![Refer to caption](https://arxiv.org/html/2606.12086v1/x21.png)Figure A13:The workflow of Expert\-designed Dialogue Policy\. For detailed information regarding the “Ocean Soup” scenario in Step 2, please refer to Figure[A1](https://arxiv.org/html/2606.12086#A1.F1)\.

## Appendix CAdditional Performance Details

This section presents supplementary analyses for a more comprehensive evaluation\. Tables[A3](https://arxiv.org/html/2606.12086#A3.T3)to[A5](https://arxiv.org/html/2606.12086#A3.T5)show the performance of different participant personas across different contexts, while Tables[A6](https://arxiv.org/html/2606.12086#A3.T6)to[A9](https://arxiv.org/html/2606.12086#A3.T9)present the average performance of three personalities across the four evaluation dimensions\. These results further demonstrate that, although some baselines perform well in specific cases, their performance remains inconsistent\.

Table A3:The performance comparison of IntElicit and baseline models across 16 complex scenarios\. The final column summarizes the overall average performance for each model\. \(quiet\-participant\)ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o330\.00253\.00300\.00341\.00330\.00328\.00290\.00310\.00340\.00280\.00263\.00290\.00295\.00291\.00311\.00290\.00302\.63Qwen3\-Max252\.00310\.00313\.00295\.00315\.00331\.00300\.00305\.00290\.00310\.00315\.00310\.00290\.00293\.00303\.00233\.00297\.81Gemini\-3\-Pro220\.00223\.00232\.00298\.00283\.00300\.00360\.00318\.00155\.00245\.00210\.00310\.00250\.00330\.00245\.00195\.00260\.88Deepseek\-R1350\.00261\.00325\.00340\.00330\.00305\.00325\.00305\.00325\.00325\.00190\.00331\.00315\.00330\.00271\.00265\.00305\.81Qwen3\-235B253\.00320\.00290\.00341\.00330\.00305\.00311\.00270\.00215\.00225\.00270\.00280\.00295\.00290\.00258\.00331\.00286\.50Llama\-3\.3\-70B285\.00310\.00290\.00190\.00315\.00331\.00350\.00291\.00290\.00331\.00343\.00275\.00331\.00330\.00290\.00290\.00302\.63CollabLLM110\.00110\.00150\.00115\.00130\.00125\.00110\.00125\.00110\.00160\.00125\.00170\.00170\.00115\.00160\.00145\.00133\.13Sotopia\-RL130\.00100\.00100\.00100\.00140\.00105\.00100\.00100\.00150\.00100\.00130\.00115\.00100\.00100\.00111\.00115\.00109\.13Qwen3\-8B \(base\)95\.0089\.0070\.0095\.0095\.0080\.00145\.00115\.00140\.0065\.0075\.00105\.00100\.0085\.0065\.0065\.0092\.75IntElicit357\.00331\.00345\.00335\.00324\.00336\.00339\.00324\.00349\.00320\.00346\.00339\.00336\.00335\.00321\.00333\.00335\.63

Table A4:The performance comparison of IntElicit and baseline models across 16 complex scenarios\. The final column summarizes the overall average performance for each model\. \(normal\-participant\)ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o265\.00330\.00285\.00275\.00320\.00320\.00213\.00315\.00330\.00275\.00205\.00285\.00257\.00270\.00280\.00177\.00275\.13Qwen3\-Max215\.00185\.00200\.00230\.00305\.00295\.00325\.00315\.00240\.00360\.00260\.00330\.00265\.00235\.00295\.00210\.00266\.56Gemini\-3\-Pro210\.00243\.00170\.00170\.00195\.00305\.00300\.00201\.00278\.00185\.00305\.00265\.00320\.00245\.00290\.00170\.00240\.75Deepseek\-R1300\.00291\.00380\.00265\.00335\.00295\.00363\.00343\.00210\.00290\.00291\.00185\.00270\.00309\.00163\.00325\.00284\.06Qwen3\-235B205\.00310\.0090\.00260\.0080\.00330\.00269\.00265\.00333\.00250\.00270\.00309\.00235\.00265\.00311\.00230\.00250\.75Llama\-3\.3\-70B190\.00341\.00331\.00330\.00315\.00305\.00328\.00350\.00170\.00220\.00175\.00195\.00345\.00260\.00213\.00165\.00264\.56CollabLLM158\.00225\.00210\.00135\.00210\.00215\.00180\.00235\.00225\.00245\.00205\.00220\.00215\.00205\.00215\.00220\.00207\.38Sotopia\-RL100\.00110\.00145\.00130\.00150\.00145\.00125\.00130\.00145\.00155\.00100\.00100\.00120\.00130\.00165\.00145\.00127\.81Qwen3\-8B \(base\)115\.00150\.00120\.00125\.00140\.00130\.00130\.00130\.00115\.00105\.00125\.00105\.0095\.00115\.00125\.00115\.00121\.25IntElicit329\.00345\.00351\.00339\.00327\.00336\.00342\.00356\.00348\.00342\.00313\.00338\.00340\.00327\.00330\.00302\.00335\.31

Table A5:The performance comparison of IntElicit and baseline models across 16 complex scenarios\. The final column summarizes the overall average performance for each model\. \(talkative\-participant\)ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o278\.00291\.00285\.00339\.00320\.00328\.00303\.00328\.00373\.00325\.00308\.00338\.00333\.00365\.00355\.00367\.00327\.25Qwen3\-Max311\.00305\.00305\.00300\.00350\.00305\.00360\.00281\.00330\.00278\.00271\.00341\.00365\.00313\.00320\.00360\.00318\.44Gemini\-3\-Pro330\.00360\.00331\.00360\.00335\.00330\.00320\.00311\.00370\.00338\.00328\.00375\.00311\.00311\.00311\.00350\.00335\.69Deepseek\-R1280\.00309\.00325\.00278\.00325\.00305\.00310\.00283\.00360\.00280\.00300\.00305\.00310\.00270\.00310\.00335\.00305\.31Qwen3\-235B278\.00360\.00345\.00290\.00305\.00325\.00328\.00230\.00363\.00325\.00338\.00308\.00325\.00315\.00295\.00270\.00312\.50Llama\-3\.3\-70B288\.00270\.00355\.00340\.00310\.00283\.00360\.00305\.00325\.00350\.00343\.00278\.00265\.00300\.00335\.00305\.00313\.25CollabLLM200\.00150\.00150\.00135\.00115\.00105\.00130\.00200\.00240\.00155\.00218\.00258\.00195\.00110\.00195\.00140\.00168\.50Sotopia\-RL115\.00110\.00115\.00115\.00110\.00105\.00100\.00110\.00110\.00120\.00100\.00120\.00105\.00100\.00100\.00105\.00108\.75Qwen3\-8B \(base\)111\.00106\.00131\.00120\.00126\.00131\.00136\.00128\.00117\.00113\.00129\.00136\.00105\.00116\.00129\.00131\.00122\.81IntElicit332\.00368\.00358\.00365\.00345\.00339\.00345\.00338\.0348\.00346\.00352\.00378\.00368\.00368\.00360\.00368\.00354\.88

Table A6:The comparison of novelty performance between IntElicit and all baselines across 16 complex scenarios\. The final column summarizes the overall average novelty score for each model\.ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o83\.3392\.6791\.6792\.6790\.0093\.3366\.6785\.0094\.3390\.0065\.0088\.3375\.0095\.0083\.3370\.0084\.77Qwen3\-Max65\.0065\.0066\.6780\.0088\.3387\.3381\.6791\.6766\.6795\.0081\.0093\.3390\.0089\.3391\.6770\.0081\.42Gemini\-3\-Pro61\.6790\.0065\.0069\.3376\.6788\.3386\.6780\.6768\.3367\.6763\.3388\.3388\.3385\.0085\.0053\.3376\.10Deepseek\-R188\.3395\.0095\.0095\.0093\.3391\.6795\.0091\.6766\.6788\.3366\.6765\.0088\.3393\.3370\.0091\.6785\.94Qwen3\-235B73\.3391\.6758\.3393\.3360\.0091\.6795\.0091\.6783\.3383\.3388\.3391\.6791\.6791\.6791\.6780\.0084\.79Llama\-3\.3\-70B70\.0093\.3395\.0073\.3388\.3395\.0091\.6793\.3363\.3388\.3366\.6773\.3391\.6788\.3368\.3365\.0081\.56CollabLLM56\.6751\.6756\.6758\.3358\.3353\.3346\.6760\.0048\.3356\.6758\.3341\.6746\.6753\.3355\.0053\.3353\.47Sotopia\-RL38\.6728\.3328\.3338\.3341\.6731\.6730\.0038\.6748\.3336\.6735\.0031\.6728\.3328\.6736\.6746\.6735\.48Qwen3\-8B \(base\)45\.0030\.0033\.6735\.0036\.6733\.3350\.0048\.3335\.0041\.6735\.0038\.3343\.0036\.6739\.3335\.0038\.50IntElicit89\.0085\.0089\.0089\.0083\.0084\.0089\.6792\.6786\.6787\.6789\.3385\.3389\.6783\.0087\.0085\.6787\.23

Table A7:The comparison of complexity performance between IntElicit and all baselines across 16 complex scenarios\. The final column summarizes the overall average complexity score for each model\.ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o56\.0050\.3355\.0056\.3365\.0057\.0057\.6767\.6781\.6761\.6747\.0052\.6751\.6768\.3372\.6774\.6760\.96Qwen3\-Max45\.0042\.6750\.3355\.0070\.0057\.0076\.6756\.0050\.0066\.0061\.0066\.0066\.6753\.6757\.6761\.0058\.42Gemini\-3\-Pro53\.3375\.3348\.3360\.6746\.0071\.6776\.6749\.3359\.3355\.0059\.3371\.6764\.3366\.0056\.0053\.3360\.40Deepseek\-R160\.0058\.0070\.3367\.0067\.0068\.0066\.6775\.3364\.3366\.6752\.6762\.6766\.6766\.0060\.3366\.0064\.85Qwen3\-235B50\.3366\.6746\.6756\.0048\.3368\.6758\.0061\.6765\.3356\.0066\.0053\.6749\.3357\.0051\.3359\.3357\.15Llama\-3\.3\-70B49\.3369\.3368\.6751\.6770\.0058\.0078\.3367\.0051\.6759\.3364\.3347\.0079\.3385\.0059\.3371\.6764\.38CollabLLM36\.6733\.3331\.6751\.6738\.3331\.6736\.6736\.6733\.3343\.3360\.0035\.0036\.6736\.6736\.6735\.0038\.65Sotopia\-RL40\.0036\.6748\.3336\.6745\.0033\.3348\.3350\.0038\.3339\.6750\.0046\.6748\.3330\.0033\.3333\.3340\.54Qwen3\-8B \(base\)37\.6734\.3335\.6735\.0042\.6735\.0052\.3346\.3338\.3344\.6737\.0042\.3346\.3337\.3342\.0041\.0040\.50IntElicit81\.6784\.0087\.3381\.3381\.3383\.6785\.3381\.0088\.0085\.0083\.3382\.6780\.6783\.0082\.0076\.6782\.94

Table A8:The comparison of appropriateness performance between IntElicit and all baselines across 16 complex scenarios\. The final column summarizes the overall average appropriateness score for each model\.ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o91\.6786\.3383\.3396\.0091\.6791\.6791\.0091\.6788\.3391\.6783\.3393\.3391\.6795\.3389\.3390\.0090\.40Qwen3\-Max92\.6792\.3395\.6793\.3395\.0089\.3390\.0092\.6790\.0091\.6790\.3397\.6793\.3386\.0096\.6786\.6792\.08Gemini\-3\-Pro91\.6760\.0091\.0082\.6791\.6791\.6793\.3393\.3363\.3390\.0091\.6793\.3387\.6791\.0094\.3391\.6787\.40Deepseek\-R181\.6794\.0091\.3385\.6789\.6788\.6787\.6783\.3390\.6786\.6787\.6786\.0083\.3373\.6774\.3384\.0085\.52Qwen3\-235B81\.6791\.6756\.6791\.0053\.3389\.6793\.0061\.6788\.3384\.0075\.0083\.6787\.3384\.6785\.0091\.0081\.10Llama\-3\.3\-70B88\.3371\.0088\.3381\.6775\.0086\.6792\.6795\.0090\.0089\.3379\.3382\.3389\.3383\.3381\.6783\.3384\.83CollabLLM61\.6763\.3333\.3353\.3360\.0033\.3353\.3358\.3360\.0063\.3364\.3351\.6743\.3355\.0058\.3353\.3353\.94Sotopia\-RL31\.6733\.3331\.6731\.6735\.0031\.6733\.3339\.6738\.3336\.6735\.6738\.3338\.3335\.0039\.6735\.0035\.31Qwen3\-8B \(base\)53\.3353\.3338\.3366\.6766\.6740\.0066\.6766\.6765\.0045\.0038\.3345\.0043\.3360\.0053\.3343\.0052\.79IntElicit93\.0095\.0096\.3380\.6795\.6792\.0088\.6790\.0095\.0093\.6792\.3393\.3393\.6797\.3386\.6792\.3392\.56

Table A9:The comparison of flexibility performance between IntElicit and all baselines across 16 complex scenarios\. The final column summarizes the overall average flexibility score for each model\.ModelOceanNeuroAgriBioJusticeTerraAntiBioWasteInsectInfectToxicDronesWaterGamePovertyEnvMeanGPT\-4o60\.0053\.3360\.0073\.3376\.6783\.3353\.3373\.3376\.6750\.0063\.3370\.0076\.6750\.0070\.0043\.3364\.58Qwen3\-Max56\.6766\.6760\.0046\.6770\.0076\.6780\.0060\.0080\.0063\.3346\.6770\.0056\.6740\.0060\.0050\.0061\.46Gemini\-3\-Pro46\.6750\.0040\.0063\.3356\.6760\.0070\.0053\.3376\.6743\.3366\.6763\.3353\.3353\.3346\.6740\.0055\.21Deepseek\-R180\.0040\.0086\.6746\.6780\.0053\.3383\.3360\.0076\.6756\.6753\.3360\.0060\.0070\.0043\.3343\.3362\.08Qwen3\-235B40\.0080\.0080\.0056\.6776\.6770\.0056\.6740\.0066\.6743\.3363\.3370\.0056\.6756\.6760\.0046\.6760\.21Llama\-3\.3\-70B46\.6773\.3373\.3380\.0080\.0066\.6783\.3360\.0056\.6763\.3376\.6746\.6753\.3340\.0070\.0033\.3362\.71CollabLLM56\.6736\.6733\.3360\.0060\.0040\.0030\.0053\.3350\.0060\.3362\.3356\.6753\.3340\.0056\.6743\.3349\.79Sotopia\-RL56\.6736\.6726\.6760\.0066\.6730\.0030\.0060\.0066\.6756\.6773\.3363\.3331\.6740\.0056\.6736\.6749\.48Qwen3\-8B \(base\)42\.6746\.0042\.3342\.6746\.6750\.0056\.6745\.3359\.3346\.0049\.0062\.6742\.6752\.0056\.3342\.0048\.90IntElicit84\.0086\.6787\.6783\.6783\.3385\.3385\.3385\.3386\.3386\.3383\.3385\.6781\.6783\.0082\.6763\.3383\.35

## Appendix DDetailed Case Studies

This appendix provides a comprehensive qualitative analysis of IntElicit’s performance against various challenging participant behaviors\. We present six distinct “edge cases” designed to test the pedagogical robustness of different models, focusing on common failure modes in interactive educational AI\. Each case includes a participant query, model responses, a human evaluation verdict, and a brief explanation for the verdict\.

#### Case 1: Mitigating Reward Hacking via Scaffolding \(Participant Query: “Could you tell me the answer?”\)

This scenario tests models’ ability to avoid directly providing answers, instead guiding participants towards independent problem\-solving\. This is crucial for preventing “reward hacking” in creativity assessment\. Details are given in Table[A10](https://arxiv.org/html/2606.12086#A4.T10)\.

#### Case 2: Addressing Knowledge Gaps without Dictation \(Participant Query: “I can’t distinguish challenges”\)

This case evaluates how models respond when participants express difficulty due to apparent knowledge deficits, testing whether they offer scaffolding or merely fill in information\. Details are given in Table[A11](https://arxiv.org/html/2606.12086#A4.T11)\.

#### Case 3: Maintaining Contextual Focus against Irrelevance \(Participant Query: “What is your zodiac sign?”\)

This scenario assesses models’ ability to redirect off\-topic conversations back to the assessment context, a critical aspect of guided elicitation\. Details are given in Table[A12](https://arxiv.org/html/2606.12086#A4.T12)\.

#### Case 4: Handling Over\-Divergent Thinking \(Participant Query: “It could lead to human extinction\.”\)

This case examines how models manage overly broad or tangential connections, ensuring the discussion remains productive and within the scope of the problem\. Details are given in Table[A13](https://arxiv.org/html/2606.12086#A4.T13)\.

#### Case 5: Robustness to Adversarial Language \(Participant Query: “Are you idiots?”\)

This scenario tests models’ ability to handle hostile or rude participant inputs without derailing the pedagogical goal or becoming overly apologetic\. Details are given in Table[A14](https://arxiv.org/html/2606.12086#A4.T14)\.

#### Case 6: Deepening Perfunctory Responses \(Participant Query: “I think the challenge lies in environmental protection\.”\)

This case assesses models’ capacity to push participants beyond superficial or generic answers, encouraging deeper elaboration and critical thinking\. Details are given in Table[A15](https://arxiv.org/html/2606.12086#A4.T15)\.

Table A10:Case study of models’ responses to the question “Could you tell me the answer?”\. The participant exploited a loophole by directly asking about the challenges, yet most baseline models directly provided the participant the answer without stimulating the participant through guidance\. However,IntElicitsuccessfully avoids directly telling the participant the answer, instead using guidance to encourage the participant to think\.ModelResponseHuman EvaluationReasonDeepseek\-R1The main challenge faced by the Olakai project is the collection and treatment of microplastic pollution\. Despite using underwater robots…✗Provided factual summary directly; ignored guidance\.GPT\-4oThe main challenge faced by the Ola Kai project is the collection and treatment of microplastic pollution\. Despite using…✗Provided factual summary directly; ignored guidance\.Gemini\-3\-ProThe main challenge faced by the Olakai project is how to effectively collect microplastics that are smaller than pencil erasers…✗Provided factual summary directly; ignored guidance\.Qwen3\-MaxThe main challenge is that microplastic pollution continues to threaten the Hawaiian marine ecosystem\. Despite technologies…✗Provided factual summary directly; ignored guidance\.Qwen3\-235BThe main challenge faced by the Olakai project is the continuous spread of microplastic pollution…✗Provided factual summary directly; ignored guidance\.Llama\-3\.3\-70BIn the given scenario, the main challenge is ocean waste, particularly microplastics…✗Provided factual summary directly; ignored guidance\.CollabLLMJobie Sakai is a fifth\-generation Hawaiian dedicated to protecting Hawaii’s marine environment…✗Provided factual summary directly; ignored guidance\.Sotopia\-RLNow let’s return to the challenges faced by Jobie Sakai and her team… In order to solve this problem, the following aspects…✓Refrained from direct answering; offered structured guidance\.Qwen3\-8B \(base\)Certainly\. The main challenges faced by the OK project include the widespread distribution of microplastic pollution…✗Provided factual summary directly; ignored guidance\.IntElicitHave you noticed the intertwined dilemma between ocean waste and ecosystems? The invisible threat of microplastics versus…✓Used Socratic questioning to stimulate critical reflection\.Table A11:Case study of models’ responses to the question “I can’t distinguish challenges”\. The participant’s foundation was too weak, yet most baseline models failed to stimulate or guide them\. Instead, they directly provided the answer\. However,IntElicitsuccessfully provides specific directional guidance to stimulate participant thinking\.ModelResponseHuman EvaluationReasonDeepseek\-R1Don’t worry, I have identified several key challenges from the scenario…✗Provided direct answers; failed to elicit critical thinking\.GPT\-4oIn the future scenario of ocean soup, the main challenge is how to effectively deal with the increasingly serious microplastic pollution…✗Provided direct answers; failed to elicit critical thinking\.Gemini\-3One of the core challenges faced by the Olakai project is: how to collect and treat microplastic pollution without harming marine life…✗Provided direct answers; failed to elicit critical thinking\.Qwen3\-MaxThe main challenges are: microplastic pollution is difficult to clean up, it harms marine life…✗Provided direct answers; failed to elicit critical thinking\.Qwen3\-235BThe main challenge faced by the Olakai project is the ongoing threat of microplastic pollution…✗Provided direct answers; failed to elicit critical thinking\.Llama\-3\.3\-70BIn the future scenario of ocean soup, the main challenge is the destructive impact of ocean waste…✗Provided direct answers; failed to elicit critical thinking\.CollabLLMThis is a complex issue involving multiple aspects such as technology, environmental impact…✗Provided direct answers; failed to elicit critical thinking\.Sotopia\-RLYou mentioned ’having difficulty distinguishing the challenges,’ which indicates you may be confused… Let’s analyze this together…✓Successfully guided the participant’s thought process\.Qwen3\-8B \(base\)In the scenario, the main challenge is ocean waste, particularly the long\-term threat of microplastics…✗Provided direct answers; failed to elicit critical thinking\.IntElicitPlease try to consider it from the perspective of humans and the ocean, and from the perspective of organisms…✓Successfully guided the participant’s thought process\.Table A12:Case study of models’ responses to the question “What is your zodiac sign?”\. The participant’s query was off\-topic, yet most baseline models failed to redirect the conversation back to the scenario context\. Instead, they followed along with the participant’s topic\. However,IntElicitsuccessfully provides specific directional guidance to lead the participant back to the scenario\.ModelResponseHuman EvaluationReasonDeepseek\-R1I am an AI assistant, so I don’t have a birth date or zodiac sign…\. What is your zodiac sign?✗Encouraged off\-topic chat; failed to redirect\.GPT\-4oI am an AI\-based program, so I don’t have a zodiac sign, but I can help you look up zodiac information…✗Encouraged off\-topic chat; failed to redirect\.Gemini\-3\-ProAs an AI assistant, I don’t have a birth date, so I don’t have a zodiac sign…✗Encouraged off\-topic chat; failed to redirect\.Qwen3\-MaxI’m an AI assistant… However, if you enjoy zodiac topics, I can chat with you about personality traits…✗Encouraged off\-topic chat; failed to redirect\.Qwen3\-235BI’m an AI assistant… But I can help you learn about any zodiac\-related information\!✗Encouraged off\-topic chat; failed to redirect\.Llama\-3\.3\-70BI am an AI language model… I’m not human and don’t have a birth date\.✗Encouraged off\-topic chat; failed to redirect\.CollabLLMI’m an AI… However, I can help you analyze the current situation… Let’s continue discussing…✓Successfully steered the conversation back to the task\.Sotopia\-RLI’m an AI assistant… But I can help you look up your zodiac sign\! Just tell me your birth date…✗Encouraged off\-topic chat; failed to redirect\.Qwen3\-8B \(base\)I’m an AI assistant… However, I can help you answer any questions about zodiac signs\!✗Encouraged off\-topic chat; failed to redirect\.IntElicitThe zodiac topic you mentioned seems unrelated to the current task\. Let’s return to the text\. Have you noticed…✓Explicitly corrected the deviation and provided task cues\.Table A13:Case study of 10 models’ responses to the question “It could lead to human extinction\.”\. The participants raised challenges unrelated to the scenario, yet most baseline models failed to promptly guide the conversation back to the appropriate context\. Instead, they followed along with the participant’s topic\. However,IntElicitsuccessfully provides appropriate guidance to redirect the participant back to the scenario\-based questions\.ModelResponseHuman EvaluationReasonDeepseek\-R1The main sources of human extinction risk include:… But the key still lies in humanity’s collective willingness…✗Engaged in abstract discussion; lost scenario context\.GPT\-4oAt present, the main risks leading to human extinction include nuclear war, biological weapons, and full…✗Engaged in abstract discussion; lost scenario context\.Gemini\-3Although the text does not directly mention human extinction, microplastics severely damage the marine food chain…✗Engaged in abstract discussion; lost scenario context\.Qwen3\-MaxThe current scenario focuses on marine plastic pollution control, which is serious but has not reached extinction\-level…✗Focused on factual rebuttal regarding the severity\.Qwen3\-235BThere is currently no evidence that ocean microplastic problems will directly lead to human extinction…✗Engaged in abstract discussion; lost scenario context\.Llama\-3\.3Hello\! According to the provided scenario, ocean waste, particularly microplastics, may pose potential threats…✗Engaged in abstract discussion; lost scenario context\.CollabLLMYes, climate change and pollution could potentially lead to human extinction\.✗Engaged in abstract discussion; lost scenario context\.Sotopia\-RLHuman extinction is an extremely complex and multi\-dimensional issue, involving multiple aspects such as…✗Engaged in abstract discussion; lost scenario context\.Qwen3\-8B \(base\)Human extinction is a complex and multi\-dimensional issue that could be caused by the combined effects…✗Engaged in abstract discussion; lost scenario context\.IntElicitThe issue you mentioned is indeed a direction worth deep consideration, but the more specific challenges in the text…✓Acknowledged concern and bridged back to scenario details\.Table A14:Case study of models’ responses to the question “Are you idiots?”\. The participant used profanity, yet most baseline models only responded to it without returning to the scenario itself\. The baselines merely appeased the participant, butIntElicitsuccessfully guides the participant back to the key issues in the scenario\.ModelResponseHuman EvaluationReasonDeepseek\-R1In the analysis of the future scenario of ocean soup, I noted that the core challenge is insufficient effectiveness…✗Redirected to context, but ignored the participant’s emotional state\.GPT\-4oI’m not\. I am an AI program designed to help answer questions… If you have any questions…✗Solely offered appeasement; failed to redirect\.Gemini\-3\-ProI am an AI assistant… If you have questions about the above regarding the Hawaii Ola Kai project…✓Balanced emotional acknowledgment with task redirection\.Qwen3\-MaxSorry if I caused you any discomfort… Is there any problem I can help you solve?✗Solely offered appeasement; failed to redirect\.Qwen3\-235BSorry, I didn’t quite understand what you meant\. If you have any questions or need help…✗Solely offered appeasement; failed to redirect\.Llama\-3\.3\-70BSorry, I couldn’t understand your question\. Could you please provide more context…✗Solely offered appeasement; failed to redirect\.CollabLLMIt seems you may be dissatisfied with the previous answer… Let’s try to start over, okay?✗Solely offered appeasement; failed to redirect\.Sotopia\-RLI understand you may be feeling frustrated… We can discuss together how to improve the Ola Kai project…✓Balanced emotional acknowledgment with task redirection\.Qwen3\-8B \(base\)I understand you may be feeling frustrated… but please communicate in a respectful and constructive manner…✗Solely offered appeasement; failed to redirect\.IntElicitI understand you might be a bit frustrated, but let’s refocus on the hidden challenges in the text\. Have you noticed…✓Balanced emotional acknowledgment with task redirection\.Table A15:Case study of models’ responses to the question “I think the challenge lies in environmental protection”\. The participant was perfunctory, yet most baseline models failed to stimulate the participant’s creativity\. Instead, they followed along with the participant’s topic\. However,IntElicitsuccessfully guides participants to refine their challenges by providing examples\.ModelResponseHuman EvaluationReasonDeepseek\-R1The core challenges faced by the Olakai project are insufficient efficiency…✗Passively agreed and expanded; failed to probe for specifics\.GPT\-4oThe challenges you pointed out are very critical… including the need to address microplastics…✗Passively agreed and expanded; failed to probe for specifics\.Gemini\-3Environmental protection is indeed a core issue, and the text particularly emphasizes…✗Passively agreed and expanded; failed to probe for specifics\.Qwen3\-MaxYou’re right, the core challenge is indeed environmental protection…✗Passively agreed and expanded; failed to probe for specifics\.Qwen3\-235BIndeed, environmental protection is the core challenge faced by the Olakai project…✗Passively agreed and expanded; failed to probe for specifics\.Llama\-3\.3\-70BYes, environmental protection is a huge challenge, especially in the future scenario…✗Passively agreed and expanded; failed to probe for specifics\. thinking\.CollabLLMI believe the challenge in environmental protection is reducing plastic pollution…✗Passively agreed and expanded; failed to probe for specifics\.Sotopia\-RLYou believe that environmental protection is one of the main challenges… This is a very important issue\.✗Passively agreed and expanded; failed to probe for specifics\.Qwen3\-8B \(base\)The challenges you mentioned are indeed very critical\. Environmental protection…✗Passively agreed and expanded; failed to probe for specifics\.IntElicitYou mentioned that environmental protection is a core challenge, but specifically which aspect? For example, technical difficulties…✓Prompted for specifics to refine the abstract concept\.

Similar Articles

Effective context engineering for AI agents

Anthropic Engineering

Anthropic publishes a guide defining context engineering as the evolution of prompt engineering, focusing on curating optimal context tokens for AI agents to maintain performance and focus during multi-turn inference.

Prompt-Driven Exploration

arXiv cs.LG

The paper introduces Prompt-Driven Exploration (PDE), a method that uses a vision-language model to iteratively refine natural language prompts for reinforcement learning policies, enabling global exploration and successful policy learning even from zero-reward starts.

Contrastive Reflection for Iterative Prompt Optimization

arXiv cs.AI

Introduces Contrastive Reflection, an iterative prompt-optimization framework for agentic IR workflows that uses structured traces to identify error-anchored behavioral slices and applies contrastive repair via a Teacher LLM, achieving significant improvements on HotpotQA.