Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty
Summary
This paper introduces Think-Probe-Respond, a method to improve large language models' judgment of research idea novelty by probing latent judgments and reducing bias towards medium novelty ratings, achieving a 22.30% performance improvement.
View Cached Full Text
Cached at: 08/27/26, 09:25 AM
# Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty Source: [https://arxiv.org/html/2608.25660](https://arxiv.org/html/2608.25660) LLMlarge language modelMAEMean Absolute ErrorRINoBenchResearchIdeaNovelty JudgmentBenchmarkSOTAstate\-of\-the\-artTPRThink\-Probe\-RespondCoTChain\-of\-Thought Tim SchopfTobias SchreiederAffiliation:TU Dresden & ScaDS\.AI Dresden/Leipzig, GermanyEmail:[tobias\.schreieder@tu\-dresden\.de](mailto:[email protected])Akiko AizawaAffiliation:National Institute of Informatics, Tokyo, JapanEmail:[aizawa@nii\.ac\.jp](mailto:[email protected]) ###### Abstract Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas\. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially\. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as “medium novel”\. To mitigate this, we proposeThink\-Probe\-Respond \(TPR\), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response\. Across strong baselines, TPR improves novelty judgment performance by 22\.30% and successfully mitigates the prevalent “medium novelty” bias\. Figure 1:Overview of our[tpr](https://arxiv.org/html/2608.25660#id5)approach\. The snowflake \(\) indicates frozen parameters, while the flame \(\) indicates trainable parameters\.## 1Introduction Judging the novelty of research ideas is crucial to scientific progress, enabling original contributions and shaping future scientific directions\. However, manual novelty judgment requires substantial expertise and a broad understanding of the relevant literature, making it time\-consuming, subjective, and difficult to scale[Picard et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib1)\. As scientific output continues to grow rapidly[Fortunato et al\. \(2018\)](https://arxiv.org/html/2608.25660#bib.bib4), automated support for novelty judgment becomes important for helping researchers assess, refine, and compare research ideas\. Recent work increasingly relies on[llm](https://arxiv.org/html/2608.25660#id1)to judge research idea novelty\([Lu et al\., 2024](https://arxiv.org/html/2608.25660#bib.bib3);[Li et al\., 2024a](https://arxiv.org/html/2608.25660#bib.bib31);[Si et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib36);[Su et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib41);[Lu et al\., 2026](https://arxiv.org/html/2608.25660#bib.bib34);[Gottweis et al\., 2026](https://arxiv.org/html/2608.25660#bib.bib35);[Mostafa et al\., 2026](https://arxiv.org/html/2608.25660#bib.bib38);[Wu et al\., 2026](https://arxiv.org/html/2608.25660#bib.bib39),inter alia\)\. However, a critical discrepancy persists: while these models produce plausible, human\-like novelty arguments[Afzal et al\. \(2026\)](https://arxiv.org/html/2608.25660#bib.bib42), their final novelty judgments remain fundamentally disconnected from their own reasoning rationales and fail to align with human judgments[Si et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib36);[Schopf and Färber \(2026\)](https://arxiv.org/html/2608.25660#bib.bib32)\. In this work, we investigate the underlying causes of this miscalibration and demonstrate that[llm](https://arxiv.org/html/2608.25660#id1)exhibit a systematic bias toward conservative “medium novelty” judgments, even when their internal reasoning supports substantially different judgments\. Motivated by this finding, we propose[tpr](https://arxiv.org/html/2608.25660#id5)\([tpr](https://arxiv.org/html/2608.25660#id5)\), a lightweight method for mitigating novelty miscalibration\.[tpr](https://arxiv.org/html/2608.25660#id5)probes latent novelty beliefs from hidden states during reasoning and conditions the final response on these extracted beliefs\. Experimental results show that[tpr](https://arxiv.org/html/2608.25660#id5)improves novelty judgment performance by 22\.30% on average over strong baselines while producing less biased novelty judgments\. ## 2Related Work Early efforts for automated research idea novelty judgment have evolved from citation\- and lexical\-based methods[Uzzi et al\. \(2013\)](https://arxiv.org/html/2608.25660#bib.bib28);[Wang et al\. \(2017\)](https://arxiv.org/html/2608.25660#bib.bib29);[Amplayo et al\. \(2019\)](https://arxiv.org/html/2608.25660#bib.bib43);[Wang et al\. \(2019\)](https://arxiv.org/html/2608.25660#bib.bib5);[Sarica et al\. \(2020\)](https://arxiv.org/html/2608.25660#bib.bib6)to semantic embedding approaches[Gómez\-Pérez et al\. \(2022\)](https://arxiv.org/html/2608.25660#bib.bib7)\. While these methods improve semantic matching, they largely remain limited to surface\-level similarity estimation[Mysore et al\. \(2022\)](https://arxiv.org/html/2608.25660#bib.bib44)\. More recent work adopt[llm](https://arxiv.org/html/2608.25660#id1)for automated novelty judgments\([Wu et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib30);[Si et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib36);[Liu et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib26);[Wang et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib2);[Baek et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib45);[Lin et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib46);[Zhang et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib11);[Tang et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib12);[Li et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib47);[Feng et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib27);[Hou et al\., 2026](https://arxiv.org/html/2608.25660#bib.bib37),inter alia\)\. In contrast to prior work, we investigate a previously overlooked failure mode of[llm](https://arxiv.org/html/2608.25660#id1): the systematic miscalibration between their generated novelty rationales and final novelty judgments\. ## 3Benchmark We conduct all experiments on[rino](https://arxiv.org/html/2608.25660#id3)[Schopf and Färber \(2026\)](https://arxiv.org/html/2608.25660#bib.bib32), the only publicly available benchmark for research idea novelty judgment with human\-annotated novelty scores\. It comprises 1,381 expert\-authored ideas, each paired with related works, a human\-annotated novelty score on a five\-point Likert scale \(rubric in Table[5](https://arxiv.org/html/2608.25660#A1.T5)\), and an expert\-written justification supporting the assigned novelty score\. Given a research idea and its related work, models must predict the novelty score ranging from 1 \(not novel\) to 5 \(highly novel\) and generate a textual justification grounded in comparisons with prior work\. Table 1:Evaluation results of novelty judgments on the[rino](https://arxiv.org/html/2608.25660#id3)test set\. The reported metrics includeF1F\_\{1\}macro averaged and for each rubric category \(1\-5\),[mae](https://arxiv.org/html/2608.25660#id2),Alignment \(ALI\),Recall,Additional Ratio\(in %\), andHallucination Rate\(in %\) forKnown Aspects \(KA\)andNovelty Aspects \(NA\)respectively \(for more details on the metrics, see Appendix[B](https://arxiv.org/html/2608.25660#A2)\)\. ## 4Miscalibration in LLM Novelty Judgments Although[llm](https://arxiv.org/html/2608.25660#id1)novelty rationales align closely with human reasoning[Afzal et al\. \(2026\)](https://arxiv.org/html/2608.25660#bib.bib42), their final novelty judgments often diverge significantly from human consensus[Si et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib36);[Schopf and Färber \(2026\)](https://arxiv.org/html/2608.25660#bib.bib32)\. To investigate this miscalibration and examinewhy[llm](https://arxiv.org/html/2608.25660#id1)struggle to produce accurate novelty judgments, we prompt six[sota](https://arxiv.org/html/2608.25660#id4)\([sota](https://arxiv.org/html/2608.25660#id4)\)[llm](https://arxiv.org/html/2608.25660#id1)\(Gemini 2\.5 Pro[Comanici et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib19), Gemini 3 Pro[Google \(2025\)](https://arxiv.org/html/2608.25660#bib.bib21), Claude Sonnet 4\.5[Anthropic \(2025b\)](https://arxiv.org/html/2608.25660#bib.bib22), Claude Opus 4\.5[Anthropic \(2025a\)](https://arxiv.org/html/2608.25660#bib.bib20), GPT\-5 mini[Singh et al\. \(2026\)](https://arxiv.org/html/2608.25660#bib.bib33), GPT\-5\.4[OpenAI \(2026\)](https://arxiv.org/html/2608.25660#bib.bib10)\) to judge research idea novelty \(prompt in Figure[3](https://arxiv.org/html/2608.25660#A6.F3)\)\. As Table[1](https://arxiv.org/html/2608.25660#S3.T1)shows, all models perform poorly on the novelty judgment task, yielding[mae](https://arxiv.org/html/2608.25660#id2)values around one and consistently low macro\-F1F\_\{1\}scores, with the best\-performing model achieving just 17\.1\. LLMs Avoid Extreme Novelty JudgmentsThe dominant failure pattern is a strong middle\-ground bias\. Across models, predictions concentrate on novelty classes 3 and 4, while the lowest and highest categories are rarely predicted correctly\. With the exception of Gemini 3 Pro, which occasionally assigns classes 1 and 5, models fail almost entirely on these extreme cases\. Table[8](https://arxiv.org/html/2608.25660#A6.T8)shows examples of this[llm](https://arxiv.org/html/2608.25660#id1)behavior\. Latent Belief vs\. Expressed Novelty JudgmentThis middle\-ground bias contrasts with the quality of the generated justifications\. Models achieve high recall with respect to human\-annotated justification arguments, indicating that they often identify the same overlaps, differences, and novelty aspects as human experts\. Overall, these results suggest that[llm](https://arxiv.org/html/2608.25660#id1)often“know”substantially more about the true novelty level than their generated numerical novelty scores reveal\. TakeawayThe miscalibration of[llm](https://arxiv.org/html/2608.25660#id1)as novelty judges is therefore best understood as a mismatch between latent belief and expressed judgment\.The models’ reasoning often contains evidence for accurate novelty judgment, but their final outputs are biased towards safe middle categories\. ## 5Probing Latent LLM Judgments Building on the finding that[llm](https://arxiv.org/html/2608.25660#id1)internalize beliefs about research idea novelty that closely mirror those of human experts—and therefore generate comparable novelty rationales—yet are biased towards predicting medium novelty categories, we propose the[tpr](https://arxiv.org/html/2608.25660#id5)approach for research idea novelty judgment\.[tpr](https://arxiv.org/html/2608.25660#id5)explicitly exploits the model’s internal beliefs during reasoning about novelty to yield less biased and more accurate quantitative judgments\. These judgments are then reused as conditioning signals to generate textual justifications that are coherent and well\-aligned with the predicted numerical scores\. As illustrated in Figure[1](https://arxiv.org/html/2608.25660#S0.F1),[tpr](https://arxiv.org/html/2608.25660#id5)consists of three stages\. \(1\)Think:We instruct an[llm](https://arxiv.org/html/2608.25660#id1)to judge the novelty of a research idea and to think step by step before producing a final response\. For reasoning models that generate think tokens by default, we omit any explicit “think step by step” instruction\. Importantly, we provide only textual descriptions of the novelty categories—without numerical scores—and instruct the model to evaluate novelty solely based on these descriptions without generating numerical judgments\. This design encourages the model to think about the novelty of research ideas qualitatively, avoiding anchoring its internal representations to explicit numeric outcomes that could bias the reasoning process\. Figure[4](https://arxiv.org/html/2608.25660#A6.F4)shows the prompt used in this approach\. \(2\)Probe:Given an LLM withLLhidden layers, letH=h\(1\),…,h\(L\)H=h^\{\(1\)\},\\dots,h^\{\(L\)\}represent the stack of hidden states\. For a generated sequence of reasoning \(“think”\) tokensT=t1,…,tnT=t\_\{1\},\\dots,t\_\{n\}, we terminate generation upon the production of the final think tokentnt\_\{n\}and extract the hidden statehtn\(L\)h^\{\(L\)\}\_\{t\_\{n\}\}\(Section[6](https://arxiv.org/html/2608.25660#S6)motivates the choice oftnt\_\{n\}\)\. This representationhtn\(L\)h^\{\(L\)\}\_\{t\_\{n\}\}is then used as the input feature vector for a logistic regression probing classifier\.111While probing intermediate layers is a viable alternative, we follow prior work showing that representations from the final layer typically yield strong probing performance and offer practical advantages, as extraction of the last hidden state of a given model is more easily facilitated in common open\-[llm](https://arxiv.org/html/2608.25660#id1)frameworks than earlier layers[Maiya et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib48)\.\(3\)Respond:We append the textual description of the predicted novelty class to the[llm](https://arxiv.org/html/2608.25660#id1)\-generated output and resume generation\. Conditioned on both its prior reasoning and the predicted novelty judgment, the[llm](https://arxiv.org/html/2608.25660#id1)generates the final response, which is used as the justification of the novelty judgment\. [tpr](https://arxiv.org/html/2608.25660#id5)is computationally efficient and lightweight\. All[llm](https://arxiv.org/html/2608.25660#id1)parameters remain frozen and are used only at inference\. Training is limited to a simple logistic regression classifier, which can be learned efficiently on CPU\. ### 5\.1Experimental Setup We evaluate[tpr](https://arxiv.org/html/2608.25660#id5)against a diverse set of approaches\. These includeZero\-shotprompting, as used in Section[4](https://arxiv.org/html/2608.25660#S4);Few\-shotprompting adapted from[Shahid et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib56), which provides one example per novelty class;[cot](https://arxiv.org/html/2608.25660#id6)\([cot](https://arxiv.org/html/2608.25660#id6)\)prompting, where the model is instructed to reason step by step before producing a novelty judgment; several prompt\-based methods derived fromMoose[Yang et al\. \(2024\)](https://arxiv.org/html/2608.25660#bib.bib49),ResearchAgent[Baek et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib45),AI Scientist[Lu et al\. \(2024\)](https://arxiv.org/html/2608.25660#bib.bib3), andAI Researcher[Si et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib36); as well as aFineTuneapproach that uses LoRA[Hu et al\. \(2022\)](https://arxiv.org/html/2608.25660#bib.bib23)for[llm](https://arxiv.org/html/2608.25660#id1)fine\-tuning \(training details are provided in Appendix[C](https://arxiv.org/html/2608.25660#A3)\)\. Since we require access to model parameters, we conduct experiments exclusively with open\-source[llm](https://arxiv.org/html/2608.25660#id1)spanning multiple model families\. We evaluate reasoning models that generate explicit think tokens by default, namely Qwen3 \(4B, 14B, 32B\)[Yang et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib25)and GPT\-OSS\-20B[OpenAI et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib8), as well as non\-reasoning models for which we explicitly instruct step\-by\-step reasoning to elicit think tokens under[tpr](https://arxiv.org/html/2608.25660#id5), including Gemma 3 \(4B, 12B, 27B\)[Google et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib24)and Llama 3\.1 \(8B, 70B\)[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2608.25660#bib.bib9)\. ### 5\.2Evaluation Results Figure 2:Macro\-F1F\_\{1\}scores for different approaches and[llm](https://arxiv.org/html/2608.25660#id1)on the[rino](https://arxiv.org/html/2608.25660#id3)test set\.Figure[2](https://arxiv.org/html/2608.25660#S5.F2)presents macro\-F1F\_\{1\}scores for different[llm](https://arxiv.org/html/2608.25660#id1)and approaches \(see Appendix[E](https://arxiv.org/html/2608.25660#A5)for metric selection details\)\. Across all evaluated models,[tpr](https://arxiv.org/html/2608.25660#id5)achieves the highest performance, surpassing the best competing approach by an average of 22\.30%\. Remarkably, this performance is achieved using only a lightweight logistic regression probing classifier, significantly surpassing the computationally expensive FineTune approach\. While FineTune often outperforms prompt\-based approaches, it cannot match the performance of[tpr](https://arxiv.org/html/2608.25660#id5)\. Consistency Across Prompt\-Based ApproachesThe effectiveness of prompt\-based methods varies substantially across models, with no single prompting strategy consistently outperforming others\. In contrast,[tpr](https://arxiv.org/html/2608.25660#id5)demonstrates robust, model\-agnostic performance, delivering strong results across all[llm](https://arxiv.org/html/2608.25660#id1)investigated Model Size Is Not DeterminativeProbing[llm](https://arxiv.org/html/2608.25660#id1)beliefs during reasoning yields strong novelty judgments regardless of model size\. Larger models do not necessarily perform better: for instance, the biggest Llama\-3\.1\-70B ranks second\-worst, whereas the smallest Gemma3\-4B ranks second\-best\. Moreover, all open\-source models using[tpr](https://arxiv.org/html/2608.25660#id5)surpass the novelty judgment macro\-F1F\_\{1\}scores of proprietary models in Table[1](https://arxiv.org/html/2608.25660#S3.T1)\. This indicates that even smaller models using[tpr](https://arxiv.org/html/2608.25660#id5)can outperform prompting substantially larger[llm](https://arxiv.org/html/2608.25660#id1)\. Reasoning vs\. Non\-Reasoning Models[tpr](https://arxiv.org/html/2608.25660#id5)is effective for both reasoning and non\-reasoning models, showing that think tokens encode useful information about novelty judgments, whether generated automatically or via explicit instruction\. Interestingly, reasoning models exhibit slightly lower[tpr](https://arxiv.org/html/2608.25660#id5)performance on average than non\-reasoning models\. Fine\-tuning reasoning models such as Qwen3 provides modest gains, yet[tpr](https://arxiv.org/html/2608.25660#id5)still outperforms FineTune, albeit with a smaller margin than for non\-reasoning models\. Overall, the largest gains appear when[tpr](https://arxiv.org/html/2608.25660#id5)is applied to non\-reasoning models with explicit "think step\-by\-step" instructions\. [tpr](https://arxiv.org/html/2608.25660#id5)vs\.[cot](https://arxiv.org/html/2608.25660#id6)Across all[llm](https://arxiv.org/html/2608.25660#id1),[tpr](https://arxiv.org/html/2608.25660#id5)substantially outperforms the[cot](https://arxiv.org/html/2608.25660#id6)approach\. While[cot](https://arxiv.org/html/2608.25660#id6)marginally benefits non\-reasoning models, its effect on reasoning models is inconsistent, and its overall performance lags behind both FineTune and[tpr](https://arxiv.org/html/2608.25660#id5)\. These results suggest that merely instructing[llm](https://arxiv.org/html/2608.25660#id1)to reason is insufficient; explicitly probing the hidden representations formed during the reasoning process is crucial for achieving high\-quality novelty judgments\. [tpr](https://arxiv.org/html/2608.25660#id5)Enables Balanced Predictions Across Novelty ClassesTable[2](https://arxiv.org/html/2608.25660#S5.T2)reports class\-wiseF1F\_\{1\}scores for different[llm](https://arxiv.org/html/2608.25660#id1)using our[tpr](https://arxiv.org/html/2608.25660#id5)approach, while Table[1](https://arxiv.org/html/2608.25660#S3.T1)shows the corresponding scores for proprietary LLMs under zero\-shot prompting\. Compared to the prompting approach,[tpr](https://arxiv.org/html/2608.25660#id5)produces substantially more balanced predictions across all novelty classes\. In particular, prompted proprietary[llm](https://arxiv.org/html/2608.25660#id1)often avoid extreme novelty judgments \(classes 1 and 5\), whereas[tpr](https://arxiv.org/html/2608.25660#id5)enables models to recognize both very low and very high research idea novelties\. The exception is the Qwen3 model family, which rarely assigns class 1 even under[tpr](https://arxiv.org/html/2608.25660#id5), indicating a persistent tendency to avoid “no novelty” judgments\. Overall, these results show that[tpr](https://arxiv.org/html/2608.25660#id5)not only improves macro\-levelF1F\_\{1\}performance but also encourages more uniform coverage of the full novelty spectrum, mitigating the middle\-class novelty judgment bias observed in Section[4](https://arxiv.org/html/2608.25660#S4)\. Table 2:F1F\_\{1\}scores per novelty class using[tpr](https://arxiv.org/html/2608.25660#id5)with different[llm](https://arxiv.org/html/2608.25660#id1)\.An additional evaluation of the textual justifications is provided in Appendix[F](https://arxiv.org/html/2608.25660#A6)\. ## 6Probing over Time Table 3:MacroF1F\_\{1\}scores for research idea novelty judgments on the[rino](https://arxiv.org/html/2608.25660#id3)test set obtained by probing last\-layer hidden states of various[llm](https://arxiv.org/html/2608.25660#id1)at different generation steps: when generating the first and last think tokens \(t1t\_\{1\},tnt\_\{n\}\), the first and last response tokens \(r1r\_\{1\},rnr\_\{n\}\), and intermediate steps during both the thinking and response generation phases \(reported as percentages of token generation within each phase; e\.g\.,t50%t\_\{50\\%\}denotes probing halfway through the thinking phase, after 50% of think tokens have been generated\)\.We investigate at which generation stages[llm](https://arxiv.org/html/2608.25660#id1)encode the most salient information for research idea novelty judgments when applying[tpr](https://arxiv.org/html/2608.25660#id5)\. For this analysis, we allow the models to generate their full reasoning and response exactly as instructed by the prompt in Figure[4](https://arxiv.org/html/2608.25660#A6.F4),withoutinserting the probing classifier’s prediction as an intermediate control signal\. This setup enables a clean examination of when novelty\-related information naturally emerges in the model’s representations\. We probe last\-layer hidden states at multiple time steps during both the thinking and response generation phases and report the results in Table[3](https://arxiv.org/html/2608.25660#S6.T3)\. Novelty Signals Peak at the End of the Thinking PhaseAcross nearly all models, probing at theend of the thinking phase\(tnt\_\{n\}\) yields the strongest or near\-strongest novelty judgment performance\. This pattern holds consistently for both reasoning and non\-reasoning models, indicating that novelty\-related beliefs are most fully consolidated once the model has completed its internal reasoning process\. In contrast, probing earlier thinking tokens \(e\.g\.,t25%t\_\{25\\%\}ort50%t\_\{50\\%\}\) generally results in substantially lower performance, suggesting that novelty representations emerge progressively rather than being present at the onset of reasoning\. Probing Is Robust Across Reasoning LengthsAs shown in Table[4](https://arxiv.org/html/2608.25660#S6.T4), models vary widely in the number of tokens generated during the thinking phase, from short sequences to long reasoning chains\. Despite these differences, probing attnt\_\{n\}provides consistently strong novelty signals\. Importantly, there is no clear correlation between the absolute length of reasoning chains and probing performance: models with shorter or longer chains achieve comparableF1F\_\{1\}scores at the final thinking token\. This indicates that the position within the reasoning sequence \(final token\) matters more than absolute length for capturing novelty\-related beliefs\. Table 4:Number of tokens generated by different[llm](https://arxiv.org/html/2608.25660#id1)during research idea novelty prediction using our[tpr](https://arxiv.org/html/2608.25660#id5)approach\. Note, we use GPT\-OSS\-20B with reasoning level“low”, resulting in a small number of think tokens\.Response Generation Dilutes Novelty RepresentationsWhile some models achieve local maxima at intermediate response steps \(e\.g\.,r50%r\_\{50\\%\}orr75%r\_\{75\\%\}\), probing during the response generation phase produces competitive but generally weaker results than probing attnt\_\{n\}\. This suggests that once the model transitions to response generation, the representations become increasingly influenced by surface realization and linguistic planning, diluting the underlying novelty signal\. This trend holds for both reasoning and non\-reasoning models\. While some models exhibit local peaks at intermediate response steps, the final thinking token remains the most reliable and stable probing point overall\. TakeawayNovelty judgments are primarily formed during the reasoning phase and are most reliably captured toward their later stages\. ## 7Conclusion We showed that[llm](https://arxiv.org/html/2608.25660#id1)are miscalibrated judges of research idea novelty\. Although their rationales often align with human reasoning, their final judgments are biased towards medium novelty\. To mitigate this, we proposed[tpr](https://arxiv.org/html/2608.25660#id5), a lightweight approach that probes latent novelty judgments from hidden states during reasoning and conditions the final response on the probed judgment\. Experiments demonstrate that[tpr](https://arxiv.org/html/2608.25660#id5)improves novelty judgment performance over strong baselines by 22\.30% and reduces the prevalent medium\-novelty bias\. ## 8Limitations Our experiments are conducted on[rino](https://arxiv.org/html/2608.25660#id3), which focuses on machine learning research ideas and may not fully reflect novelty judgments in other scientific domains\. In addition,[tpr](https://arxiv.org/html/2608.25660#id5)requires access to model hidden states, limiting its direct applicability to closed\-source[llm](https://arxiv.org/html/2608.25660#id1)\. Finally, novelty judgments are inherently subjective, and even expert annotations may reflect individual preferences or incomplete knowledge of the literature\. ## 9Ethical Considerations We emphasize that this work is intended for research and educational purposes only\. Users should not use our models or approaches to make formal or high\-stakes judgments of research ideas, as novelty judgments are inherently subjective and context\-dependent\. Our work is intended to advance AI\-assisted scientific discovery by enabling models to reason about and explain novel contributions in research\. However, automated predictions of research idea novelty should not replace human expert judgment\. Our approaches are intended as tools to support, rather than replace human judgments of research ideas\. ## Acknowledgments This work is supported by a scholarship of the German Academic Exchange Service \(DAAD\) \- 57557629 and by the BMFTR through a Software Campus project with identification number 16\|S23070\. The authors acknowledge the financial support by the Federal Ministry of Research, Technology and Space of Germany \(BMFTR\) and by Sächsische Staatsministerium für Wissenschaft, Kultur und Tourismus in the programme Center of Excellence for AI\-research „Center for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzig“, project identification number: ScaDS\.AI\. We used AI\-based assistance tools to support language editing, minor formatting, and coding tasks\. These tools did not contribute to the intellectual content or scientific conclusions\. All content was reviewed by the authors, who assume full responsibility for the publication\. ## References - Afzalet al\.\(2025\)A\. Afzal, F\. Matthes, G\. Chechik, and Y\. ZiserKnowing before saying: LLM representations encode information about chain\-of\-thought success before completion\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 12791–12806\.External Links:[Link](https://aclanthology.org/2025.findings-acl.662/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.662),ISBN 979\-8\-89176\-256\-5Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Afzalet al\.\(2026\)O\. M\. Afzal, P\. Nakov, T\. Hope, and I\. GurevychBeyond “not novel enough”: enriching scholarly critique with LLM\-assisted feedback\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2648–2671\.External Links:[Link](https://aclanthology.org/2026.eacl-long.121/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.121),ISBN 979\-8\-89176\-380\-7Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1),[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Amplayoet al\.\(2019\)R\. K\. Amplayo, S\. Hwang, and M\. SongEvaluating research novelty detection: counterfactual approaches\.InProceedings of the Thirteenth Workshop on Graph\-Based Methods for Natural Language Processing \(TextGraphs\-13\),D\. Ustalov, S\. Somasundaran, P\. Jansen, G\. Glavaš, M\. Riedl, M\. Surdeanu, and M\. Vazirgiannis \(Eds\.\),Hong Kong,pp\. 124–133\.External Links:[Link](https://aclanthology.org/D19-5315/),[Document](https://dx.doi.org/10.18653/v1/D19-5315)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Anthropic \(2025a\)AnthropicIntroducing Claude Opus 4\.5 — anthropic\.com\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-5](https://www.anthropic.com/news/claude-opus-4-5)\[Accessed 05\-01\-2026\]Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Anthropic \(2025b\)AnthropicIntroducing Claude Sonnet 4\.5 — anthropic\.com\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-4\-5](https://www.anthropic.com/news/claude-sonnet-4-5)\[Accessed 05\-01\-2026\]Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Baeket al\.\(2025\)J\. Baek, S\. K\. Jauhar, S\. Cucerzan, and S\. J\. HwangResearchAgent: iterative research idea generation over scientific literature with large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 6709–6738\.External Links:[Link](https://aclanthology.org/2025.naacl-long.342/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.342),ISBN 979\-8\-89176\-189\-6Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1),[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Comaniciet al\.\(2025\)G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen, L\. Marris, S\. Petulla, C\. Gaffney, A\. Aharoni, N\. Lintz, T\. C\. Pais, H\. Jacobsson, I\. Szpektor, N\. Jiang, K\. Haridasan, A\. Omran, N\. Saunshi, D\. Bahri, G\. Mishra, E\. Chu, T\. Boyd, B\. Hekman, A\. Parisi, C\. Zhang, K\. Kawintiranon, T\. Bedrax\-Weiss, O\. Wang, Y\. Xu, O\. Purkiss, U\. Mendlovic, I\. Deutel, N\. Nguyen, A\. Langley, F\. Korn, L\. Rossazza, A\. Ramé, S\. Waghmare, H\. Miller, N\. Byrd, A\. Sheshan, R\. Hadsell, S\. Bhardwaj, P\. Janus, T\. Rissa, D\. Horgan, A\. Abdagic, L\. Belenki, J\. Allingham, A\. Singh, T\. Guidroz, S\. Srinivasan, H\. Schmit, K\. Chiafullo, A\. Elisseeff, N\. Jha, P\. Kolhar, L\. Berrada, F\. Ding, X\. Si, S\. B\. Mallick, F\. Och, S\. Erell, E\. Ni, T\. Latkar, S\. Yang, P\. Sirkovic, Z\. Feng, R\. Leland, R\. Hornung, G\. Wu, C\. Blundell, H\. Alvari, P\. Huang, C\. Yip, S\. Deur, L\. Liu, G\. Surita, P\. Duque, D\. Damen, J\. Jia, A\. Guez, M\. Mircea, A\. Sinha, A\. Magni, P\. Stradomski, T\. Marian, V\. Galić, W\. Chen, H\. Husain, A\. Singhal, D\. Grewe, F\. Aubet, S\. Song, L\. Blanco, L\. Rechis, L\. Ho, R\. Munoz, K\. Zheng, J\. Hamrick, K\. Mather, H\. Taitelbaum, E\. Rutherford, Y\. Lei, K\. Chen, A\. Shukla, E\. Moreira, E\. Doi, B\. Isik, N\. Shabat, D\. Rogozińska, K\. Kolipaka, J\. Chang, E\. Vušak, S\. Venkatachary, S\. Noghabi, T\. Bharti, Y\. Jun, A\. Zaks, S\. Green, J\. Challagundla, W\. Wong, M\. Mohammad, D\. Hirsch, Y\. Cheng, I\. Naim, L\. Proleev, D\. Vincent, A\. Singh, M\. Krikun, D\. Krishnan, Z\. Ghahramani, A\. Atias, R\. Aggarwal, C\. Kirov, D\. Vytiniotis, C\. Koh, A\. Chronopoulou, P\. Dogra, V\. Ion, G\. Tyen, J\. Lee, F\. Weissenberger, T\. Strohman, A\. Balakrishna, J\. Rae, M\. Velic, R\. de Liedekerke, O\. Elyada, W\. Yuan, C\. Liu, L\. Shani, S\. Kishchenko, B\. Alessio, Y\. Li, R\. Song, S\. Kwei, O\. Jankowski, A\. Pappu, Y\. Namiki, Y\. Ma, N\. Tripuraneni, C\. Cherry, M\. Ikonomidis, Y\. Ling, C\. Ji, B\. Westberg, A\. Wright, D\. Yu, D\. Parkinson, S\. Ramaswamy, J\. Connor, S\. H\. Yeganeh, S\. Grover, G\. Kenwright, L\. Litchev, C\. Apps, A\. Tomala, F\. Halim, A\. Castro\-Ros, Z\. Li, A\. Boral, P\. Sho, M\. Yarom, E\. Malmi, D\. Klinghoffer, R\. Lin, A\. Ansell, P\. K\. S, S\. Zhao, S\. Zuo, A\. Santoro, H\. Cheng, S\. Demmessie, Y\. Liu, N\. Brichtova, A\. Culp, N\. Braun, D\. Graur, W\. Ng, N\. Mehta, A\. Phillips, P\. Sundberg, V\. Godbole, F\. Liu, Y\. Katariya, D\. Rim, M\. Seyedhosseini, S\. Ammirati, J\. Valfridsson, M\. Malihi, T\. Knight, A\. Toor, T\. Lampe, A\. Ittycheriah, L\. Chiang, C\. Yeung, A\. Fréchette, J\. Rao, H\. Wang, H\. Srivastava, R\. Zhang, R\. Rhodes, A\. Brand, D\. Weesner, I\. Figotin, F\. Gimeno, R\. Fellinger, P\. Marcenac, J\. Leal, E\. Marcus, V\. Cotruta, R\. Cabrera, S\. Luo, D\. Garrette, V\. Axelrod, S\. Baltateanu, D\. Barker, D\. Chen, H\. Toma, B\. Ingram, J\. Riesa, C\. Kulkarni, Y\. Zhang, H\. Liu, C\. Wang, M\. Polacek, W\. Wu, K\. Hui, A\. N\. Reyes, Y\. Su, M\. Barnes, I\. Malhi, A\. Siddiqui, Q\. Feng, M\. Damaschin, D\. Pighin, A\. Steiner, S\. Yang, R\. S\. Boppana, S\. Ivanov, A\. Kandoor, A\. Shah, A\. Mujika, D\. Huang, C\. A\. Choquette\-Choo, M\. Patel, T\. Yu, T\. Creswell, Jerry, Liu, C\. Barros, Y\. Razeghi, A\. Roy, P\. Culliton, B\. Xiong, J\. Pan, T\. Strohmann, T\. Powell, B\. Seal, D\. DeCarlo, P\. Shyam, K\. Katircioglu, X\. Wang, C\. Hardin, I\. Odisho, J\. Broder, O\. Chang, A\. Nair, A\. Shtefan, M\. O’Brien, M\. Agarwal, S\. Potluri, S\. Goyal, A\. Jhindal, S\. Thakur, Y\. Stuken, J\. Lyon, K\. Toutanova, F\. Feng, A\. Wu, B\. Horn, A\. Wang, A\. Cullum, G\. Taubman, D\. Shrivastava, C\. Shi, H\. Tomlinson, R\. Patel, T\. Tu, A\. M\. Oflazer, F\. Pongetti, M\. Yang, A\. A\. Taïga, V\. Perot, N\. W\. Pierse, F\. Han, Y\. Drori, I\. Iturrate, A\. Chakrabarti, L\. Yeung, D\. Dopson, Y\. Chen, A\. Kulshreshtha, T\. Guo, P\. Pham, T\. Schuster, J\. Chen, A\. Polozov, J\. Xing, H\. Zhou, P\. Kacham, D\. Kukliansky, A\. Miech, S\. Yaroshenko, E\. Chi, S\. Douglas, H\. Fei, M\. Blondel, P\. Myla, L\. Madmoni, X\. Wu, D\. Keysers, K\. Kjems, I\. Albuquerque, L\. Yu, J\. D’sa, M\. Plantan, V\. Ionescu, J\. S\. Elias, A\. Gupta, M\. R\. Vuyyuru, F\. Alcober, T\. Zhou, K\. Ji, F\. Hartmann, S\. Puttagunta, H\. Song, E\. Amid, A\. Stefanoiu, A\. Lee, P\. Pucciarelli, E\. Wang, A\. Raul, S\. Petrov, I\. Tian, V\. Anklin, N\. Nti, V\. Gomes, M\. Schumacher, G\. Vesom, A\. Panagopoulos, K\. Bousmalis, D\. Andor, J\. Jacob, Y\. Zhang, B\. Rosgen, M\. Kecman, M\. Tung, A\. Belias, N\. Goodman, P\. Covington, B\. Wieder, N\. Saxena, E\. Davoodi, M\. Huang, S\. Maddineni, V\. Roulet, F\. Campbell\-Ajala, P\. G\. Sessa, Xintian, Wu, G\. Lai, P\. Collins, A\. Haig, V\. Sakenas, X\. Xu, M\. Giustina, L\. E\. Shafey, P\. Charoenpanit, S\. Garg, J\. Ainslie, B\. Severson, M\. G\. Arenas, S\. Pathak, S\. Rajayogam, J\. Feng, M\. Bakker, S\. Li, N\. Wichers, J\. Rogers, X\. Geng, Y\. Li, R\. Jagerman, C\. Jia, N\. Olmert, D\. Sharon, M\. Mauger, S\. Mariserla, H\. Ma, M\. Mohabey, K\. Kim, A\. Andreev, S\. Pollom, J\. Love, V\. Jain, P\. Agrawal, Y\. Schroecker, A\. Fortin, M\. Warmuth, J\. Liu, A\. Leach, I\. Blok, G\. P\. Girirajan, R\. Aharoni, B\. Uria, A\. Sozanschi, D\. Goldberg, L\. Ionita, M\. T\. Ribeiro, M\. Zlocha, V\. Birodkar, S\. Lachgar, L\. Yuan, H\. Choudhury, M\. Ginsberg, F\. Zheng, G\. Dibb, E\. Graves, S\. Lokhande, G\. Rasskin, G\. Muraru, C\. Quick, S\. Tata, P\. Sermanet, A\. Chawla, I\. Karo, Y\. Wang, S\. Zhang, O\. Keller, A\. Dragan, G\. Su, I\. Chou, X\. Liu, Y\. Tao, S\. Prabhakara, M\. Wilson, R\. Liu, S\. Wang, G\. Evans, D\. Du, A\. Castaño, G\. Prasad, M\. E\. Mahdy, S\. Gerlach, M\. Reid, J\. Kahn, A\. Zait, T\. S\. Pillai, T\. Ulrich, G\. Wang, J\. Wassenberg, E\. Farkash, K\. Yalasangi, C\. Wang, M\. Bauza, S\. Bucher, T\. Liu, J\. Yan, G\. Leung, V\. Sindhwani, P\. Barnes, A\. Singh, I\. Jurin, J\. Chang, N\. K\. Bhumihar, S\. Eiger, G\. Citovsky, B\. Withbroe, Z\. Li, S\. Xue, N\. D\. Santo, G\. Stoyanov, Y\. Raimond, S\. Zheng, Y\. Gao, V\. Listík, S\. Kwasiborski, R\. Saputro, A\. Ozturel, G\. Mallya, K\. Majmundar, R\. West, P\. Caron, J\. Wei, L\. Castrejon, S\. Vikram, D\. Ramachandran, N\. Dhawan, J\. Park, S\. Smoot, G\. van den Driessche, Y\. Blau, C\. Malik, W\. Liang, R\. Hirsch, C\. N\. dos Santos, E\. Weinstein, A\. van den Oord, S\. Lall, N\. FitzGerald, Z\. Jiang, X\. Yang, D\. Webster, A\. Elqursh, A\. Pope, G\. Rotival, D\. Raposo, W\. Zhu, J\. Dean, S\. Alabed, D\. Tran, A\. Gupta, Z\. Gleicher, J\. Austin, E\. Rosseel, M\. Umekar, D\. Das, Y\. Sun, K\. Chen, K\. Misiunas, X\. Zhou, Y\. Di, A\. Loo, J\. Newlan, B\. Li, V\. Ramasesh, Y\. Xu, A\. Chen, S\. Gandhe, R\. Soricut, N\. Gupta, S\. Hu, S\. El\-Sayed, X\. Garcia, I\. Brusilovsky, P\. Chen, A\. Bolt, L\. Huang, A\. Gurney, Z\. Zhang, A\. Pritzel, J\. Wilkiewicz, B\. Seybold, B\. K\. Shamanna, F\. Fischer, J\. Dean, K\. Gill, R\. Mcilroy, A\. Bhowmick, J\. Selier, A\. Yang, D\. Cheng, V\. Magay, J\. Tan, D\. Varma, C\. Walder, T\. Kocisky, R\. Nakashima, P\. Natsev, M\. Kwong, I\. Gog, C\. Zhang, S\. Dieleman, T\. Jimma, A\. Ryabtsev, S\. Brahma, D\. Steiner, D\. Du, A\. Žužul, M\. Žanić, M\. Raghavachari, W\. Gierke, Z\. Zheng, D\. Petrova, Y\. Dauphin, Y\. Liu, I\. Kessler, S\. Hand, C\. Duvarney, S\. Kim, H\. Lee, L\. Hussenot, J\. Hui, J\. Smith, D\. Jain, J\. Xia, G\. S\. Tomar, K\. Amiri, D\. Phan, F\. Fuchs, T\. Weyand, N\. Tomasev, A\. Cordell, X\. Liu, J\. Mallinson, P\. Joshi, A\. Crawford, A\. Suggala, S\. Chien, N\. Fernando, M\. Sanchez\-Vargas, D\. Williams, P\. Crone, X\. Luo, I\. Karpov, J\. Shan, T\. Thurk, R\. Strudel, P\. Voigtlaender, P\. Patil, T\. Dozat, A\. Khodaei, S\. Singla, P\. Ambroszczyk, Q\. Wu, Y\. Chang, B\. Roark, C\. Hegde, T\. Ding, A\. Filos, Z\. Wu, A\. S\. Pinto, S\. Liu, S\. Khanna, A\. Pandey, S\. Mcloughlin, Q\. Li, S\. Haves, A\. Zhou, E\. Buchatskaya, I\. Leal, P\. de Boursac, N\. Akazawa, N\. Anderson, T\. Chen, K\. Somandepalli, C\. Liang, S\. Goenka, S\. Winkler, A\. Grushetsky, Y\. Ding, J\. Smith, F\. Ye, J\. Pont\-Tuset, E\. Li, R\. Li, T\. Golany, D\. Wegner, T\. Jiang, O\. Barak, Y\. Shangguan, E\. Vértes, R\. Wong, J\. Bornschein, A\. Tudor, M\. Bevilacqua, T\. Schaul, A\. S\. Rawat, Y\. Zhao, K\. Axiotis, L\. Meng, C\. McLean, J\. Lai, J\. Beattie, N\. Kushman, Y\. Liu, B\. Kutzman, F\. Lang, J\. Ye, P\. Netrapalli, P\. Mishra, M\. Khan, M\. Goel, R\. Willoughby, D\. Tian, H\. Zhuang, J\. Chen, Z\. Tsai, T\. Kementsietsidis, A\. Khare, J\. Keeling, K\. Xu, N\. Waters, F\. Altché, A\. Popat, B\. Mittal, D\. Saxton, D\. E\. Badawy, M\. Mathieu, Z\. Zheng, H\. Zhou, N\. Ranka, R\. Shin, Q\. Duan, T\. Salimans, I\. Mihailescu, U\. Shaham, M\. Chang, Y\. Assael, N\. Dikkala, M\. Izzard, V\. Cohen\-Addad, C\. Graves, V\. Feinberg, G\. Chung, D\. Strouse, D\. Karmon, S\. Sharifzadeh, Z\. Ashwood, K\. Pham, J\. Blanton, A\. Vasiloff, J\. Barber, M\. Geller, A\. Zhou, F\. Zubach, T\. Huang, L\. Zhang, H\. Gupta, M\. Young, J\. Proskurnia, R\. Votel, V\. Gabeur, G\. Barcik, A\. Tripathi, H\. Yu, G\. Yan, B\. Changpinyo, F\. Pavetić, A\. Coyle, Y\. Fujii, J\. G\. Mendez, T\. Zhou, H\. Rajamani, B\. Hechtman, E\. Cao, D\. Juan, Y\. Tan, V\. Dalibard, Y\. Du, N\. Clay, K\. Yao, W\. Jia, D\. Vijaykumar, Y\. Zhou, X\. Bai, W\. Hung, S\. Pecht, G\. Todorov, N\. Khadke, P\. Gupta, P\. Lahoti, A\. Autef, K\. Duddu, J\. Lee\-Thorp, A\. Bykovsky, T\. Misiunas, S\. Flennerhag, S\. Thangaraj, J\. McGiffin, Z\. Nado, M\. Kunesch, A\. Noever, A\. Hertz, M\. Liang, V\. Stone, E\. Palmer, S\. Daruki, A\. Pramanik, S\. Põder, A\. Kyker, M\. Khan, E\. Sluzhaev, M\. Ritter, A\. Ruderman, W\. Zhou, C\. Nagpal, K\. Vodrahalli, G\. Necula, P\. Barham, E\. Pavlick, J\. Hartford, I\. Shafran, L\. Zhao, M\. Mikuła, T\. Eccles, H\. Shimokawa, K\. Garg, L\. Vilnis, H\. Chen, I\. Shumailov, K\. Lee, A\. Abdelhamed, M\. Xie, V\. Cohen, E\. Hlavnova, D\. Malkin, C\. Sitawarin, J\. Lottes, P\. Coquinot, T\. Yu, S\. Kumar, J\. Zhang, A\. Mahendru, Z\. Ahmed, J\. Martens, T\. Chen, A\. Boag, D\. Peng, C\. Devin, A\. Klimovskiy, M\. Phuong, D\. Vainstein, J\. Xie, B\. Ramabhadran, N\. Howard, X\. Yu, G\. Goswami, J\. Cui, S\. Shleifer, M\. Pinto, C\. Yeh, M\. Yang, S\. Javanmardi, D\. Ethier, C\. Lee, J\. Orbay, S\. Kotecha, C\. Bromberg, P\. Shaw, J\. Thornton, A\. G\. Rosenthal, S\. Gu, M\. Thomas, I\. Gemp, A\. Ayyar, A\. Ushio, A\. Selvan, J\. Wee, C\. Liu, M\. Majzoubi, W\. Yu, J\. Abernethy, T\. Liechty, R\. Pan, H\. Nguyen, Qiong, Hu, S\. Perrin, A\. Arora, E\. Pitler, W\. Wang, K\. Shivakumar, F\. Prost, B\. Limonchik, J\. Wang, Y\. Gao, T\. Cour, S\. Buch, H\. Gui, M\. Ivanova, P\. Neubeck, K\. Chan, L\. Kim, H\. Chen, N\. Goyal, D\. Chung, L\. Liu, Y\. Su, A\. Petrushkina, J\. Shen, A\. Joulin, Y\. Xu, S\. X\. Lin, Y\. Kulizhskaya, C\. Chelba, S\. Vasudevan, E\. Collins, V\. Bashlovkina, T\. Lu, D\. Fritz, J\. Park, Y\. Zhou, C\. Su, R\. Tanburn, M\. Sushkov, M\. Rasquinha, J\. Li, J\. Prendki, Y\. Li, P\. LV, S\. Sharma, H\. Fitoussi, H\. Huang, A\. Dai, P\. Dao, M\. Burrows, H\. Prior, D\. Qin, G\. Pundak, L\. L\. Sjoesund, A\. Khurshudov, Z\. Zhu, A\. Webson, E\. Kemp, T\. Tan, S\. Agrawal, S\. Sargsyan, L\. Cheng, J\. Stephan, T\. Kwiatkowski, D\. Reid, A\. Byravan, A\. H\. Michaely, N\. Heess, L\. Zhou, S\. Goenka, V\. Carpenter, A\. Levskaya, B\. Wang, R\. Roberts, R\. Leblond, S\. Chikkerur, S\. Ginzburg, M\. Chang, R\. Riachi, Chuqiao, Xu, Z\. Borsos, M\. Pliskin, J\. Pawar, M\. Lustman, H\. Kirkwood, A\. Anand, A\. Chaudhary, N\. Kalb, K\. Milan, S\. Augenstein, A\. Goldie, L\. Prince, K\. Raman, Y\. Sun, V\. Xia, A\. Cohen, Z\. Huo, J\. Camp, S\. Ellis, L\. Zilka, D\. V\. Torres, L\. Patel, S\. Arora, B\. Chan, J\. Adler, K\. Ayoub, J\. Liang, F\. Jamil, J\. Jiang, S\. Baumgartner, H\. Sun, Y\. Karov, Y\. Akulov, H\. Zheng, I\. Cai, C\. Fantacci, J\. Rubin, A\. R\. Acha, M\. Wang, N\. D’Souza, R\. Sathyanarayana, S\. Dai, S\. Rowe, A\. Simanovsky, O\. Goldman, Y\. Kuang, X\. Pan, A\. Rosenberg, T\. Rojas\-Esponda, P\. Dutta, A\. Zeng, I\. Jurenka, G\. Farquhar, Y\. Bansal, S\. Iqbal, B\. Roelofs, G\. Joung, P\. Beak, C\. Ryu, R\. Poplin, Y\. Wu, J\. Alayrac, S\. Buthpitiya, O\. Ronneberger, C\. Habtegebriel, W\. Li, P\. Cavallaro, A\. Wei, G\. Bensky, T\. Denk, H\. Ganapathy, J\. Stanway, P\. Joshi, F\. Bertolini, J\. Lo, O\. Ma, Z\. Charles, G\. Sampemane, H\. Sahni, X\. Chen, H\. Askham, D\. Gaddy, P\. Young, J\. Tan, M\. Eyal, A\. Bražinskas, L\. Zhong, Z\. Wu, M\. Epstein, K\. Bailey, A\. Hard, K\. Lee, S\. Goldshtein, A\. Ruiz, M\. Badawi, M\. Lochbrunner, J\. Kearns, A\. Brown, F\. Pardo, T\. Weber, H\. Yang, P\. Jiang, B\. Akin, Z\. Fu, M\. Wainwright, C\. Zou, M\. Gaba, P\. Manzagol, W\. Kan, Y\. Song, K\. Zainullina, R\. Lin, J\. Ko, S\. Deshmukh, A\. Jindal, J\. Svensson, D\. Tyam, H\. Zhao, C\. Kaeser\-Chen, S\. Baird, P\. Moradi, J\. Hall, Q\. Guo, V\. Tsang, B\. Liang, F\. Pereira, S\. Ganesh, I\. Korotkov, J\. Adamek, S\. Thiagarajan, V\. Tran, C\. Chen, C\. Tar, S\. Jain, I\. Dasgupta, T\. Bilal, D\. Reitter, K\. Zhao, G\. Vezzani, Y\. Gehman, P\. Mehta, L\. Beltrone, X\. Dotiwalla, S\. Guadarrama, Z\. Abbas, S\. Karp, P\. Georgiev, C\. Ferng, M\. Brockschmidt, L\. Peng, C\. Hirnschall, V\. Verma, Y\. Bi, Y\. Xiao, A\. Dabush, K\. Xu, P\. Wallis, R\. Parker, Q\. Wang, Y\. Xu, I\. Safarli, D\. Tewari, Y\. Zhang, S\. Kim, A\. Gesmundo, M\. Thomas, S\. Levi, A\. Chowdhury, K\. Rao, P\. Garst, S\. Conway\-Rahman, H\. Ran, K\. McKinney, Z\. Xiao, W\. Yu, R\. Agrawal, A\. Stjerngren, C\. Ionescu, J\. Chen, V\. Sharma, J\. Chiu, F\. Liu, K\. Franko, C\. Sanford, X\. Cai, P\. Michel, S\. Ganapathy, J\. Labanowski, Z\. Garrett, B\. Vargas, S\. Sun, B\. Gale, T\. Buschmann, G\. Desjardins, N\. Ghelani, P\. Jain, M\. Verma, C\. Asawaroengchai, J\. Eisenschlos, J\. Harlalka, H\. Kazawa, D\. Metzler, J\. Howland, Y\. Jian, J\. Ades, V\. Shah, T\. Gangwani, S\. Lee, R\. Ring, S\. M\. Hernandez, D\. Reich, A\. Sinha, A\. Sathe, J\. Kovac, A\. Gill, A\. Kannan, A\. D’olimpio, M\. Sevenich, J\. Whang, B\. Kim, K\. C\. Sim, J\. Chen, J\. Zhang, S\. Lall, Y\. Matias, B\. Jia, A\. Friesen, S\. Nasso, A\. Thapliyal, B\. Perozzi, T\. Yu, A\. Shekhawat, S\. Huda, P\. Grabowski, E\. Wang, A\. Sreevatsa, H\. Dib, M\. Hassen, P\. Schuh, V\. Milutinovic, C\. Welty, M\. Quinn, A\. Shah, B\. Wang, G\. Barth\-Maron, J\. Frye, N\. Axelsson, T\. Zhu, Y\. Ma, I\. Giannoumis, H\. Sedghi, C\. Ye, Y\. Luan, K\. Aydin, B\. Chandra, V\. Sampathkumar, R\. Huang, V\. Lavrenko, A\. Eleryan, Z\. Hong, S\. Hansen, S\. M\. Carthy, B\. Samanta, D\. Ćevid, X\. Wang, F\. Li, M\. Voznesensky, M\. Hoffman, A\. Terzis, V\. Sehwag, G\. Fidel, L\. He, M\. Cai, Y\. He, A\. Feng, M\. Nikoltchev, S\. Phatale, J\. Chase, R\. Lawton, M\. Zhang, T\. Ouyang, M\. Tragut, M\. H\. Manshadi, A\. Narayanan, J\. Shen, X\. Gao, T\. Bolukbasi, N\. Roy, X\. Li, D\. Golovin, L\. Panait, Z\. Qin, G\. Han, T\. Anthony, S\. Kudugunta, V\. Patraucean, A\. Ray, X\. Chen, X\. Yang, T\. Bhatia, P\. Talluri, A\. Morris, A\. Ražnatović, B\. Brownfield, J\. An, S\. Peng, P\. Kane, C\. Zheng, N\. Duduta, J\. Kessinger, J\. Noraky, S\. Liu, K\. Rong, P\. Veličković, K\. Rush, A\. Goldin, F\. Wei, S\. M\. R\. Garlapati, C\. Pantofaru, O\. Kwon, J\. Ni, E\. Noland, J\. D\. Trapani, F\. Beaufays, A\. G\. Roy, Y\. Chow, A\. Turker, G\. Cideron, L\. Mei, J\. Clark, Q\. Dou, M\. Bošnjak, R\. Leith, Y\. Du, A\. Yazdanbakhsh, M\. Nasr, C\. Kwak, S\. S\. Sheth, A\. Kaskasoli, A\. Anand, B\. Lakshminarayanan, S\. Jerome, D\. Bieber, C\. Chu, A\. Senges, T\. Shen, M\. Sridhar, N\. Ndebele, B\. Beyret, S\. Mohamed, M\. Chen, M\. Freitag, J\. Guo, L\. Liu, P\. Roit, H\. Chen, S\. Yan, T\. Stone, J\. Co\-Reyes, J\. Cole, S\. Scellato, S\. Azizi, H\. Hashemi, A\. Jin, A\. Iyer, M\. Valentine, A\. György, A\. Ahuja, D\. H\. Diaz, C\. Lee, N\. Clement, W\. Kong, D\. Garmon, I\. Watts, K\. Bhatia, K\. Gupta, M\. Miecnikowski, H\. Vallet, A\. Taly, E\. Loper, S\. Joshi, J\. Atwood, J\. Chick, M\. Collier, F\. Iliopoulos, R\. Trostle, B\. Gunel, R\. Leal\-Cavazos, A\. M\. Hrafnkelsson, M\. Guzman, X\. Ju, A\. Forbes, J\. Emond, K\. Chauhan, B\. Caine, L\. Xiao, W\. Zeng, A\. Moufarek, D\. Murphy, M\. Meng, N\. Gupta, F\. Riedel, A\. Das, E\. Lawal, S\. Narayan, T\. Sosea, J\. Swirhun, L\. Friso, B\. Neyshabur, J\. Lu, S\. Girgin, M\. Wunder, E\. Yvinec, A\. Pyne, V\. Carbune, S\. Rijhwani, Y\. Guo, T\. Doshi, A\. Briukhov, M\. Bain, A\. Hitron, X\. Wang, A\. Gupta, K\. Chen, C\. Du, W\. Zhang, D\. Shah, A\. Akula, M\. Dylla, A\. Kachra, W\. Kuo, T\. Zou, L\. Wang, L\. Xu, J\. Zhu, J\. Snyder, S\. Menon, O\. Firat, I\. Mordatch, Y\. Yuan, N\. Ponomareva, R\. Blevins, L\. Moore, W\. Wang, P\. Chen, M\. Scholz, A\. Dwornik, J\. Lin, S\. Li, D\. Antognini, T\. I, X\. Song, M\. Miller, U\. Kalra, A\. Raveret, O\. Akerlund, F\. Wu, A\. Nystrom, N\. Godbole, T\. Liu, H\. DeBalsi, J\. Zhao, B\. Liu, A\. Caciularu, L\. Lax, U\. Khandelwal, V\. Langston, E\. Bailey, S\. Lattanzi, Y\. Wang, N\. Kovelamudi, S\. Mondal, G\. Guruganesh, N\. Hua, O\. Roval, P\. Wesołowski, R\. Ingale, J\. Halcrow, T\. Sohn, C\. Angermueller, B\. Raad, E\. Stickgold, E\. Lu, A\. Kosik, J\. Xie, T\. Lillicrap, A\. Huang, L\. L\. Zhang, D\. Paulus, C\. Farabet, A\. Wertheim, B\. Wang, R\. Joshi, C\. Ko, Y\. Wu, S\. Agrawal, L\. Lin, X\. Sheng, P\. Sung, T\. Breland\-King, C\. Butterfield, S\. Gawde, S\. Singh, Q\. Zhang, R\. Apte, S\. Shetty, A\. Hutter, T\. Li, E\. Salesky, F\. Lebron, J\. Kanerva, M\. Paganini, A\. Nguyen, R\. Vallu, J\. Peter, S\. Velury, D\. Kao, J\. Hoover, A\. Bortsova, C\. Bishop, S\. Jakobovits, A\. Agostini, A\. Agarwal, C\. Liu, C\. Kwong, S\. Tavakkol, I\. Bica, A\. Greve, A\. GP, J\. Marcus, L\. Hou, T\. Duerig, R\. Moroshko, D\. Lacey, A\. Davis, J\. Amelot, G\. Wang, F\. Kim, T\. Strinopoulos, H\. Wan, C\. L\. Lan, S\. Krishnan, H\. Tang, P\. Humphreys, J\. Bai, I\. H\. Shtacher, D\. Machado, C\. Pang, K\. Burke, D\. Liu, R\. Aravamudhan, Y\. Song, E\. Hirst, A\. Singh, B\. Jou, L\. Bai, F\. Piccinno, C\. K\. Fu, R\. Alazard, B\. Meiri, D\. Winter, C\. Chen, M\. Zhang, J\. Heitkaemper, J\. Lambert, J\. Lee, A\. Frömmgen, S\. Rogulenko, P\. Nair, P\. Niemczyk, A\. Bulyenov, B\. Xu, H\. Shemtov, M\. Zadimoghaddam, S\. Toropov, M\. Wirth, H\. Dai, S\. Gollapudi, D\. Zheng, A\. Kurakin, C\. Lee, K\. Bullard, N\. Serrano, I\. Balazevic, Y\. Li, J\. Schalkwyk, M\. Murphy, M\. Zhang, K\. Sequeira, R\. Datta, N\. Agrawal, C\. Sutton, N\. Attaluri, M\. Chiang, W\. Farhan, G\. Thornton, K\. Lin, T\. Choma, H\. Nguyen, K\. Dasgupta, D\. Robinson, I\. Comşa, M\. Riley, A\. Pillai, B\. Mustafa, B\. Golan, A\. Zandieh, J\. Lespiau, B\. Porter, D\. Ross, S\. Rajayogam, M\. Agarwal, S\. Venugopalan, B\. Shahriari, Q\. Yan, H\. Xu, T\. Tobin, P\. Dubov, H\. Shi, A\. Recasens, A\. Kovsharov, S\. Borgeaud, L\. Dery, S\. Vasanth, E\. Gribovskaya, L\. Qiu, M\. Mahdieh, W\. Skut, E\. Nielsen, C\. Zheng, A\. Yu, C\. G\. Bostock, S\. Gupta, A\. Archer, C\. Rawles, E\. Davies, A\. Svyatkovskiy, T\. Tsai, Y\. Halpern, C\. Reisswig, B\. Wydrowski, B\. Chang, J\. Puigcerver, M\. H\. Taege, J\. Li, E\. Schnider, X\. Li, D\. Dena, Y\. Xu, U\. Telang, T\. Shi, H\. Zen, K\. Kastner, Y\. Ko, N\. Subramaniam, A\. Kumar, P\. Blois, Z\. Dai, J\. Wieting, Y\. Lu, Y\. Zeldes, T\. Xie, A\. Hauth, A\. Ţifrea, Y\. Li, S\. El\-Husseini, D\. Abolafia, H\. Zhou, W\. Ding, S\. Ghalebikesabi, C\. Guía, A\. Maksai, Á\. Weisz, S\. Arik, N\. Sukhanov, A\. Świetlik, X\. Jia, L\. Yu, W\. Wang, M\. Brand, D\. Bloxwich, S\. Kirmani, Z\. Chen, A\. Go, P\. Sprechmann, N\. Kannen, A\. Carin, P\. Sandhu, I\. Edkins, L\. Nooteboom, J\. Gupta, L\. Maggiore, J\. Azizi, Y\. Pritch, P\. Yin, M\. Gupta, D\. Tarlow, D\. Smith, D\. Ivanov, M\. Babaeizadeh, A\. Goel, S\. Kambala, G\. Chu, M\. Kastelic, M\. Liu, H\. Soltau, A\. Stone, S\. Agrawal, M\. Kim, K\. Soparkar, S\. Tadepalli, O\. Bunyan, R\. Soh, A\. Kannan, D\. Kim, B\. J\. Chen, A\. Halumi, S\. Roy, Y\. Wang, O\. Sercinoglu, G\. Gibson, S\. Bhatnagar, M\. Sano, D\. von Dincklage, Q\. Ren, B\. Mitrevski, M\. Olšák, J\. She, C\. Doersch, Jilei, Wang, B\. Liu, Q\. Tan, T\. Yakar, T\. Warkentin, A\. Ramirez, C\. Lebsack, J\. Dillon, R\. Mathews, T\. Cobley, Z\. Wu, Z\. Chen, J\. Simon, S\. Nath, T\. Sainath, A\. Bendebury, R\. Julian, B\. Mankalale, D\. Ćurko, P\. Zacchello, A\. R\. Brown, K\. Sodhia, H\. Howard, S\. Caelles, A\. Gupta, G\. Evans, A\. Bulanova, L\. Katzen, R\. Goldenberg, A\. Tsitsulin, J\. Stanton, B\. Schillings, V\. Kovalev, C\. Fry, R\. Shah, K\. Lin, S\. Upadhyay, C\. Li, S\. Radpour, M\. Maggioni, J\. Xiong, L\. Haas, J\. Brennan, A\. Kamath, N\. Savinov, A\. Nagrani, T\. Yacovone, R\. Kappedal, K\. Andriopoulos, L\. Lao, Y\. Li, G\. Rozhdestvenskiy, K\. Hashimoto, A\. Audibert, S\. Austin, D\. Rodriguez, A\. Ruoss, G\. Honke, D\. Karkhanis, X\. Xiong, Q\. Wei, J\. Huang, Z\. Leng, V\. Premachandran, S\. Bileschi, G\. Evangelopoulos, T\. Mensink, J\. Pavagadhi, D\. Teplyashin, P\. Chang, L\. Xue, G\. Tanzer, S\. Goldman, K\. Patel, S\. Li, J\. Wiesner, I\. Zheng, I\. Stewart\-Binks, J\. Han, Z\. Li, L\. Luo, K\. Lenc, M\. Lučić, F\. Xue, R\. Mullins, A\. Guseynov, C\. Chang, I\. Galatzer\-Levy, A\. Zhang, G\. Bingham, G\. Hu, A\. Hartman, Y\. Ma, J\. Griffith, A\. Irpan, C\. Radebaugh, S\. Yue, L\. Fan, V\. Ungureanu, C\. Sorokin, H\. Teufel, P\. Li, R\. Anil, D\. Paparas, T\. Wang, C\. Lin, H\. Peng, M\. Shum, G\. Petrovic, D\. Brady, R\. Nguyen, K\. Macherey, Z\. Li, H\. Singh, M\. Yenugula, M\. Iinuma, X\. Chen, K\. Kopparapu, A\. Stern, S\. Dave, C\. Thekkath, F\. Perot, A\. Kumar, F\. Li, Y\. Xiao, M\. Bilotti, M\. H\. Bateni, I\. Noble, L\. Lee, A\. Vázquez\-Reina, J\. Salazar, X\. Yang, B\. Wang, E\. Gruzewska, A\. Rao, S\. Raghuram, Z\. Xu, E\. Ben\-David, J\. Mei, S\. Dalmia, Z\. Zhang, Y\. Liu, G\. Bansal, H\. Pankov, S\. Schwarcz, A\. Burns, C\. Chan, S\. Sanghai, R\. Liang, E\. Liang, A\. He, A\. Stuart, A\. Narayanan, Y\. Zhu, C\. Frank, B\. Fatemi, A\. Sabne, O\. Lang, I\. Bhattacharya, S\. Settle, M\. Wang, B\. McMahan, A\. Tacchetti, L\. B\. Soares, M\. Hadian, S\. Cabi, T\. Chung, N\. Putikhin, G\. Li, J\. Chen, A\. Tarango, H\. Michalewski, M\. Kazemi, H\. Masoom, H\. Sheftel, R\. Shivanna, A\. Vadali, R\. Comanescu, D\. Reid, J\. Moore, A\. Neelakantan, M\. Sander, J\. Herzig, A\. Rosenberg, M\. Dehghani, J\. Choi, M\. Fink, R\. Hayes, E\. Ge, S\. Weng, C\. Ho, J\. Karro, K\. Krishna, L\. N\. Thiet, A\. Skerry\-Ryan, D\. Eppens, M\. Andreetto, N\. Sarma, S\. Bonacina, B\. K\. Ayan, M\. Nawhal, Z\. Shan, M\. Dusenberry, S\. Thakoor, S\. Gubbi, D\. D\. Nguyen, R\. Tsarfaty, S\. Albanie, J\. Mitrović, M\. Gandhi, B\. Chen, A\. Epasto, G\. Stephanov, Y\. Jin, S\. Gehman, A\. Amini, J\. Weber, F\. Behbahani, S\. Xu, M\. Allamanis, X\. Chen, M\. Ott, C\. Sha, M\. Jastrzebski, H\. Qi, D\. Greene, X\. Wu, A\. Toki, D\. Vlasic, J\. Shapiro, R\. Kotikalapudi, Z\. Shen, T\. Saeki, S\. Xie, A\. Cassirer, S\. Bharadwaj, T\. Kiyono, S\. Bhojanapalli, E\. Rosenfeld, S\. Ritter, J\. Mao, J\. G\. Oliveira, Z\. Egyed, B\. Bandemer, E\. Parisotto, K\. Kinoshita, J\. Pluto, P\. Maniatis, S\. Li, Y\. Guo, G\. Ghiasi, J\. Tarbouriech, S\. Chatterjee, J\. Jin, Katrina, Xu, J\. Palomaki, S\. Arnold, M\. Sewak, F\. Piccinini, M\. Sharma, B\. Albrecht, S\. Purser\-haskell, A\. Vaswani, C\. Chen, M\. Wisniewski, Q\. Cao, J\. Aslanides, N\. M\. Phu, M\. Sieb, L\. Agubuzu, A\. Zheng, D\. Sohn, M\. Selvi, A\. Andreassen, K\. Subudhi, P\. Eruvbetine, O\. Woodman, T\. Mery, S\. Krause, X\. Ren, X\. Ma, J\. Luo, D\. Chen, W\. Fan, H\. Griffiths, C\. Schuler, A\. Li, S\. Zhang, J\. Sarr, S\. Luo, R\. Patana, M\. Watson, D\. Naboulsi, M\. Collins, S\. Sidhwani, E\. Hoogeboom, S\. Silver, E\. Caveness, X\. Zhao, M\. Rodriguez, M\. Deines, L\. Bai, P\. Griffin, M\. Tagliasacchi, E\. Xue, S\. R\. Babbula, B\. Pang, N\. Ding, G\. Shen, E\. Peake, R\. Crocker, S\. S\. Raghvendra, D\. Swisher, W\. Han, R\. Singh, L\. Wu, V\. Pchelin, T\. Munkhdalai, D\. Alon, G\. Bacon, E\. Robles, J\. Bulian, M\. Johnson, G\. Powell, F\. T\. Ferreira, Y\. Li, F\. Benzing, M\. Velimirović, H\. Soyer, W\. Kong, Tony, Nguyên, Z\. Yang, J\. Liu, J\. van Amersfoort, D\. Gillick, B\. Sun, N\. Rauschmayr, K\. Zhang, S\. Zhan, T\. Zhou, A\. Frolov, C\. Yang, D\. Vnukov, L\. Rouillard, H\. Li, A\. Mandhane, N\. Fallen, R\. Venkataraman, C\. H\. Hu, J\. Brennan, J\. Lee, J\. Chang, M\. Sundermeyer, Z\. Pan, R\. Ke, S\. Tong, A\. Fabrikant, W\. Bono, J\. Gu, R\. Foley, Y\. Mao, M\. Delakis, D\. Bhaswar, R\. Frostig, N\. Li, A\. Zipori, C\. Hope, O\. Kozlova, S\. Mishra, J\. Djolonga, C\. Schiff, M\. A\. Merey, E\. Briakou, P\. Morgan, A\. Wan, A\. Hassidim, R\. Skerry\-Ryan, K\. Sengupta, M\. Jasarevic, P\. Kallakuri, P\. Kunkle, H\. Brennan, T\. Lieber, H\. Mansoor, J\. Walker, B\. Zhang, A\. Xie, G\. Žužić, A\. Chukwuka, A\. Druinsky, D\. Cho, R\. Yao, F\. Naeem, S\. Butt, E\. Kim, Z\. Jia, M\. Jordan, A\. Lelkes, M\. Kurzeja, S\. Wang, J\. Zhao, A\. Over, A\. Chakladar, M\. Prasetya, N\. Jha, S\. Ganapathy, Y\. Cong, P\. Shroff, C\. Saroufim, S\. Miryoosefi, M\. Hammad, T\. Nasir, W\. Xi, Y\. Gao, Y\. Maeng, B\. Hora, C\. Cheng, P\. Haghani, Y\. Lewenberg, C\. Lu, M\. Matysiak, N\. Raisinghani, H\. Wang, L\. Baugher, R\. Sukthankar, M\. Giang, J\. Schultz, N\. Fiedel, M\. Chen, C\. Lee, T\. Dey, H\. Zheng, S\. Paul, C\. Smith, A\. Ly, Y\. Wang, R\. Bansal, B\. Perz, S\. Ricco, S\. Blank, V\. Keshava, D\. Sharma, M\. Chow, K\. Lad, K\. Jalan, S\. Osindero, C\. Swanson, J\. Scott, A\. Ilić, X\. Li, S\. R\. Jonnalagadda, A\. S\. Soudagar, Y\. Xiong, B\. Batsaikhan, D\. Jarrett, N\. Kumar, M\. Shah, M\. Lawlor, A\. Waters, M\. Graham, R\. May, S\. Ramos, S\. Lefdal, Z\. Cankara, N\. Cano, B\. O’Donoghue, J\. Borovik, F\. Liu, J\. Grimstad, M\. Alnahlawi, K\. Tsihlas, T\. Hudson, N\. Grigorev, Y\. Jia, T\. Huang, T\. P\. Igwe, S\. Lebedev, X\. Tang, I\. Krivokon, F\. Garcia, M\. Tan, E\. Jia, P\. Stys, S\. Vashishth, Y\. Liang, B\. Venkatraman, C\. Gu, A\. Kementsietsidis, C\. Zhu, J\. Jung, Y\. Bai, M\. J\. Hosseini, F\. Ahmed, A\. Gupta, X\. Yuan, S\. Ashraf, S\. Nigam, G\. Vasudevan, P\. Awasthi, A\. M\. Gilady, Z\. Mariet, R\. Eskander, H\. Li, H\. Hu, G\. Garrido, P\. Schlattner, G\. Zhang, R\. Saxena, P\. Dević, K\. Muralidharan, A\. Murthy, Y\. Zhou, M\. Choi, A\. Wongpanich, Z\. Wang, P\. Shah, Y\. Xu, Y\. Huang, S\. Spencer, A\. Chen, J\. Cohan, J\. Wang, J\. Tompson, J\. Wu, R\. Haroun, H\. Li, B\. Huergo, F\. Yang, T\. Yin, J\. Wendt, M\. Bendersky, R\. Chaabouni, J\. Snaider, J\. Ferret, A\. Jindal, T\. Thompson, A\. Xue, W\. Bishop, S\. M\. Phal, A\. Sharma, Y\. Sung, P\. Radhakrishnan, M\. Shomrat, R\. Ingle, R\. Vij, J\. Gilmer, M\. D\. Istin, S\. Sobell, Y\. Lu, E\. Nottage, D\. Sadigh, J\. Willcock, T\. Zhang, S\. Xu, S\. Brown, K\. Lee, G\. Wang, Y\. Zhu, Y\. Tay, C\. Kim, A\. Gutierrez, A\. Sharma, Y\. Xian, S\. Seo, C\. Cui, E\. Pochernina, C\. Baetu, K\. Jastrzębski, M\. Ly, M\. Elhawaty, D\. Suh, E\. Sezener, P\. Wang, N\. Yuen, G\. Tucker, J\. Cai, Z\. Yang, C\. Wang, A\. Muzio, H\. Qian, J\. Yoo, D\. Lockhart, K\. R\. McKee, M\. Guo, M\. Mehrotra, A\. Mendonça, S\. V\. Mehta, S\. Ben, C\. Tekur, J\. Mu, M\. Zhu, V\. Krakovna, H\. Lee, A\. Maschinot, S\. Cevey, H\. Choe, A\. Bai, H\. Srinivasan, D\. Gasaway, N\. Young, P\. Siegler, D\. Holtmann\-Rice, V\. Piratla, K\. Baumli, R\. Yogev, A\. Hofer, H\. van Hasselt, S\. Grant, Y\. Chervonyi, D\. Silver, A\. Hogue, A\. Agarwal, K\. Wang, P\. Singh, F\. Flynn, J\. Lipschultz, R\. David, L\. Bellot, Y\. Yang, L\. Le, F\. Graziano, K\. Olszewska, K\. Hui, A\. Maurya, N\. Parotsidis, W\. Chen, T\. Oguntebi, J\. Kelley, A\. Baddepudi, J\. Mauerer, G\. Shaw, A\. Siegman, L\. Yang, S\. Shetty, S\. Roy, Y\. Song, W\. Stokowiec, R\. Burnell, O\. Savant, R\. Busa\-Fekete, J\. Miao, S\. Ghosh, L\. MacDermed, P\. Lippe, M\. Dektiarev, Z\. Behrman, F\. Mentzer, K\. Nguyen, M\. Wei, S\. Verma, C\. Knutsen, S\. Dasari, Z\. Yan, P\. Mitrichev, X\. Wang, V\. Shejwalkar, J\. Austin, S\. Sunkara, N\. Potti, Y\. Virin, C\. Wright, G\. Liu, O\. Riva, E\. Pot, G\. Kochanski, Q\. Le, G\. Balasubramaniam, A\. Dhar, Y\. Liao, A\. Bloniarz, D\. Shukla, E\. Cole, J\. Lee, S\. Zhang, S\. Kafle, S\. Vashishtha, P\. Mahmoudieh, G\. Chen, R\. Hoffmann, P\. Srinivasan, A\. D\. Lago, Y\. B\. Shalom, Z\. Wang, M\. Elabd, A\. Sharma, J\. Oh, S\. Kothawade, M\. Le, M\. Monteiro, S\. Yang, K\. Alarakyia, R\. Geirhos, D\. Mincu, H\. Garnes, H\. Kobayashi, S\. Mariooryad, K\. Krasowiak, Zhixin, Lai, S\. Mourad, M\. Wang, F\. Bu, O\. Aharoni, G\. Chen, A\. Goyal, V\. Zubov, A\. Bapna, E\. Dabir, N\. Kothari, K\. Lamerigts, N\. D\. Cao, J\. Shar, C\. Yew, N\. Kulkarni, D\. Mahaarachchi, M\. Joshi, Z\. Zhu, J\. Lichtarge, Y\. Zhou, H\. Muckenhirn, V\. Selo, O\. Vinyals, P\. Chen, A\. Brohan, V\. Mehta, S\. Cogan, R\. Wang, T\. Geri, W\. Ko, W\. Chen, F\. Viola, K\. Shivam, L\. Wang, M\. C\. Elish, R\. A\. Popa, S\. Pereira, J\. Liu, R\. Koster, D\. Kim, G\. Zhang, S\. Ebrahimi, P\. Talukdar, Y\. Zheng, P\. Poklukar, A\. Mikhalap, D\. Johnson, A\. Vijayakumar, M\. Omernick, M\. Dibb, A\. Dubey, Q\. Hu, A\. Suman, V\. Aggarwal, I\. Kornakov, F\. Xia, W\. Lowe, A\. Kolganov, T\. Xiao, V\. Nikolaev, S\. Hemingray, B\. Li, J\. Iljazi, M\. Rybiński, B\. Sandhu, P\. Lu, T\. Luong, R\. Jenatton, V\. Govindaraj, Hui, Li, G\. Dulac\-Arnold, W\. Park, H\. Wang, A\. Modi, J\. Pouget\-Abadie, K\. Greller, R\. Gupta, R\. Berry, P\. Ramachandran, J\. Xie, L\. McCafferty, J\. Wang, K\. Gupta, H\. Lim, B\. Bratanič, A\. Brock, I\. Akolzin, J\. Sproch, D\. Karliner, D\. Kim, A\. Goedeckemeyer, N\. Shazeer, C\. Schmid, D\. Calandriello, P\. Bhatia, K\. Choromanski, C\. Montgomery, D\. Dua, A\. Ramalho, H\. King, Y\. Gao, L\. Nguyen, D\. Lindner, D\. Pitta, O\. Johnson, K\. Salama, D\. Ardila, M\. Han, E\. Farnese, S\. Odoom, Z\. Wang, X\. Ding, N\. Rink, R\. Smith, H\. T\. Lehri, E\. Cohen, N\. Vats, T\. He, P\. Gopavarapu, A\. Paszke, M\. Patel, W\. V\. Gansbeke, L\. Loher, L\. Castro, M\. Voitovich, T\. von Glehn, N\. George, S\. Niklaus, Z\. Eaton\-Rosen, N\. Rakićević, E\. Jue, S\. Perel, C\. Zhang, Y\. Bahat, A\. Pouget, Z\. Xing, F\. Huot, A\. Shenoy, T\. Bos, V\. Coriou, B\. Richter, N\. Noy, Y\. Wang, S\. Ontanon, S\. Qin, G\. Makarchuk, D\. Hassabis, Z\. Li, M\. Sharma, K\. Venkatesan, I\. Kemaev, R\. Daniel, S\. Huang, S\. Shah, O\. Ponce, Warren, Chen, M\. Faruqui, J\. Wu, S\. Andačić, S\. Payrits, D\. McDuff, T\. Hume, Y\. Cao, M\. Tessler, Q\. Wang, Y\. Wang, I\. Rendulic, E\. Agustsson, M\. Johnson, T\. Lando, A\. Howard, S\. G\. S\. Padmanabhan, M\. Daswani, A\. Banino, M\. Kilgore, J\. Heek, Z\. Ji, A\. Caceres, C\. Li, N\. Kassner, A\. Vlaskin, Z\. Liu, A\. Grills, Y\. Hou, R\. Sukkerd, G\. Cheon, N\. Shetty, L\. Markeeva, P\. Stanczyk, T\. Iyer, Y\. Gong, S\. Gao, K\. Gopalakrishnan, T\. Blyth, M\. Reynolds, A\. Bhoopchand, M\. Bilenko, D\. Gharibian, V\. Zayats, A\. Faust, A\. Singh, M\. Ma, H\. Jiao, S\. Vijayanarasimhan, L\. Aroyo, V\. Yadav, S\. Chakera, A\. Kakarla, V\. Meshram, K\. Gregor, G\. Botea, E\. Senter, D\. Jia, G\. Kovacs, N\. Sharma, S\. Baur, K\. Kang, Y\. He, L\. Zhuo, M\. Kostelac, I\. Laish, S\. Peng, L\. O’Bryan, D\. Kasenberg, G\. R\. Rao, E\. Leurent, B\. Zhang, S\. Stevens, A\. Salazar, Y\. Zhang, I\. Lobov, J\. Walker, A\. Porter, M\. Redshaw, H\. Ke, A\. Rao, A\. Lee, H\. Lam, M\. Moffitt, J\. Kim, S\. Qiao, T\. Koo, R\. Dadashi, X\. Song, M\. Sundararajan, P\. Xu, C\. Kawamoto, Y\. Zhong, C\. Barbu, A\. Reddy, M\. Verzetti, L\. Li, G\. Papamakarios, H\. Klimczak\-Plucińska, M\. Cassin, K\. Kavukcuoglu, R\. Swavely, A\. Vaucher, J\. Zhao, R\. Hemsley, M\. Tschannen, H\. Ge, G\. Menghani, Y\. Yu, N\. Ha, W\. He, X\. Wu, M\. Song, R\. Sterneck, S\. Zinke, D\. A\. Calian, A\. Marsden, A\. C\. Ruiz, M\. Hessel, A\. Gueta, B\. Lee, B\. Farris, M\. Gupta, Y\. Li, M\. Saleh, V\. Misra, K\. Xiao, P\. Mendolicchio, G\. Buttimore, V\. Krayvanova, N\. Nayakanti, M\. Wiethoff, Y\. Pande, A\. Mirhoseini, N\. Lao, J\. Liu, Y\. Hua, A\. Chen, Y\. Malkov, D\. Kalashnikov, S\. Gupta, K\. Audhkhasi, Y\. Zhai, S\. Kopalle, P\. Jain, E\. Ofek, C\. Meyer, K\. Baatarsukh, H\. Strejček, J\. Qian, J\. Freedman, R\. Figueira, M\. Sokolik, O\. Bachem, R\. Lin, D\. Kharrat, C\. Hidey, P\. Xu, D\. Duan, Y\. Li, M\. Ersoy, R\. Everett, K\. Cen, R\. Santamaria\-Fernandez, A\. Taubenfeld, I\. Mackinnon, L\. Deng, P\. Zablotskaia, S\. Viswanadha, S\. Goel, D\. Yates, Y\. Deng, P\. Choy, M\. Chen, A\. Sinha, A\. Mossin, Y\. Wang, A\. Szlam, S\. Hao, P\. K\. Rubenstein, M\. Toksoz\-Exley, M\. Aperghis, Y\. Zhong, J\. Ahn, M\. Isard, O\. Lacombe, F\. Luisier, C\. Anastasiou, Y\. Kalley, U\. Prabhu, E\. Dunleavy, S\. Bijwadia, J\. Mao\-Jones, K\. Chen, R\. Pasumarthi, E\. Wood, A\. Dostmohamed, N\. Hurley, J\. Simsa, A\. Parrish, M\. Pajarskas, M\. Harvey, O\. Skopek, Y\. Kochinski, J\. Rey, V\. Rieser, D\. Zhou, S\. J\. Lee, T\. Acharya, G\. Li, J\. Jiang, X\. Zhang, B\. Gipson, E\. Mahintorabi, M\. Gelmi, N\. Khajehnouri, A\. Yeh, K\. Lee, L\. Matthey, L\. Baker, T\. Pham, H\. Fu, A\. Pak, P\. Gupta, C\. Vasconcelos, A\. Sadovsky, B\. Walker, S\. Hsiao, P\. Zochbauer, A\. Marzoca, N\. Velan, J\. Zeng, G\. Baechler, D\. Driess, D\. Jain, Y\. Huang, L\. Tao, J\. Maggs, N\. Levine, J\. Schneider, E\. Gemzer, S\. Petit, S\. Han, Z\. Fisher, D\. Zelle, C\. Biles, E\. Ie, A\. Fadeeva, C\. Liu, J\. V\. Franco, A\. Collister, H\. Zhang, R\. Wang, R\. Zhao, L\. Kieliger, K\. Shuster, R\. Zhu, B\. Gong, L\. Chan, R\. Sun, S\. Basu, R\. Zimmermann, J\. Hayes, A\. Bapna, J\. Snoek, W\. Yang, P\. Datta, J\. A\. Abdallah, K\. Kilgour, L\. Li, S\. Mah, Y\. Jun, M\. Rivière, A\. Karmarkar, T\. Spalink, T\. Huang, L\. Gonzalez, D\. Tran, A\. Nowak, J\. Palowitch, M\. Chadwick, E\. Talius, H\. Mehta, T\. Sellam, P\. Fränken, M\. Nicosia, K\. He, A\. Kini, D\. Amos, S\. Basu, H\. Jobe, E\. Shaw, Q\. Xu, C\. Evans, D\. Ikeda, C\. Yan, L\. Jin, L\. Wang, S\. Yadav, I\. Labzovsky, R\. Sampath, A\. Ma, C\. Schumann, A\. Siddhant, R\. Shah, J\. Youssef, R\. Agarwal, N\. Dabney, A\. Tonioni, M\. Ambar, J\. Li, I\. Guyon, B\. Li, D\. Soergel, B\. Fang, G\. Karadzhov, C\. Udrescu, T\. Trinh, V\. Raunak, S\. Noury, D\. Guo, S\. Gupta, M\. Finkelstein, D\. Petek, L\. Liang, G\. Billock, P\. Sun, D\. Wood, Y\. Song, X\. Yu, T\. Matejovicova, R\. Cohen, K\. Andra, D\. D’Ambrosio, Z\. Deng, V\. Nallatamby, E\. Songhori, R\. Dangovski, A\. Lampinen, P\. Botadra, A\. Hillier, J\. Cao, N\. Baddi, A\. Kuncoro, T\. Yoshino, A\. Bhagatwala, M\. Ranzato, R\. Schaeffer, T\. Liu, S\. Ye, O\. Sarvana, J\. Nham, C\. Kuang, I\. Gao, J\. Baek, S\. Mittal, A\. Wahid, A\. Gergely, B\. Ni, J\. Feldman, C\. Muir, P\. Lamblin, W\. Macherey, E\. Dyer, L\. Kilpatrick, V\. Campos, M\. Bhutani, S\. Fort, Y\. Ahmad, A\. Severyn, K\. Chatziprimou, O\. Ferludin, M\. Dimarco, A\. Kusupati, J\. Heyward, D\. Bahir, K\. Villela, K\. Millican, D\. Marcus, S\. Bahargam, C\. Unlu, N\. Roth, Z\. Wei, S\. Gopal, D\. Ghoshal, E\. Lee, S\. Lin, J\. Lees, D\. Lee, A\. Hosseini, C\. Fan, S\. Neel, M\. Wu, Y\. Altun, H\. Cai, E\. Piqueras, J\. Woodward, A\. Bissacco, S\. Haykal, M\. Bordbar, P\. Sundaram, S\. Hodkinson, D\. Toyama, G\. Polovets, A\. Myers, A\. Sinha, T\. Levinboim, K\. Krishnakumar, R\. Chhaparia, T\. Sholokhova, N\. B\. Gundavarapu, G\. Jawahar, H\. Qureshi, J\. Hu, N\. Momchev, M\. Rahtz, R\. Wu, A\. P\. S, K\. Dhamdhere, M\. Guo, U\. Gupta, A\. Eslami, M\. Schain, M\. Blokzijl, D\. Welling, D\. Orr, L\. Bolelli, N\. Perez\-Nieves, M\. Sirotenko, A\. Prasad, A\. Kar, B\. D\. B\. Pigem, T\. Terzi, G\. Weisz, D\. Ghosh, A\. Mavalankar, D\. Madeka, K\. Daugaard, H\. Adam, V\. Shah, D\. Berman, M\. Tran, S\. Baker, E\. Andrejczuk, G\. Chole, G\. Raboshchuk, M\. Mirzazadeh, T\. Kagohara, S\. Wu, C\. Schallhart, B\. Orlando, C\. Wang, A\. Rrustemi, H\. Xiong, H\. Liu, A\. Vezer, N\. Ramsden, S\. Chang, S\. Mudgal, Y\. Li, N\. Vieillard, Y\. Hoshen, F\. Ahmad, A\. Slone, A\. Hua, N\. Potikha, M\. Rossini, J\. Stritar, S\. Prakash, Z\. Wang, X\. Dong, A\. Nazari, E\. Nehoran, K\. Tekelioglu, Y\. Li, K\. Badola, T\. Funkhouser, Y\. Li, V\. Yerram, R\. Ganeshan, D\. Formoso, K\. Langner, T\. Shi, H\. Li, Y\. Yamamori, A\. Panda, A\. Saade, A\. S\. Scarpati, C\. Breaux, C\. Carey, Z\. Zhou, C\. Hsieh, S\. Bridgers, A\. Butryna, N\. Gupta, V\. Tulsyan, S\. Woo, E\. Eltyshev, W\. Grathwohl, C\. Parks, S\. Benjamin, R\. Panigrahy, S\. Dodhia, D\. D\. Freitas, C\. Sauer, W\. Song, F\. Alet, J\. Tolins, C\. Paduraru, X\. Zhou, B\. Albert, Z\. Zhang, L\. Shu, M\. Bansal, S\. Nguyen, A\. Globerson, O\. Xiao, J\. Manyika, T\. Hennigan, R\. Rong, J\. Matak, A\. Bakalov, A\. Sharma, D\. Sinopalnikov, A\. Pierson, S\. Roller, G\. Brown, M\. Gao, T\. Fukuzawa, A\. Ghafouri, K\. Vassigh, I\. Barr, Z\. Wang, A\. Korsun, R\. Jayaram, L\. Ren, T\. Zaman, S\. Khan, Y\. Lunts, D\. Deutsch, D\. Uthus, N\. Katz, M\. Samsikova, A\. Khalifa, N\. Sethi, J\. Sun, L\. Tang, U\. Alon, X\. Luo, D\. Yu, A\. Nayyar, B\. Petrini, W\. Truong, V\. Hellendoorn, N\. Chinaev, C\. Alberti, W\. Wang, J\. Hu, V\. Mirrokni, A\. Balashankar, A\. Aharon, A\. Mehta, A\. Iscen, J\. Kready, L\. Manning, A\. Mohananey, Y\. Chen, A\. Tripathi, A\. Wu, I\. Petrovski, D\. Hwang, M\. Baeuml, S\. Chandrakaladharan, Y\. Liu, R\. Coaguila, M\. Chen, S\. Ma, P\. Tafti, S\. Tatineni, T\. Spitz, J\. Ye, P\. Vicol, M\. Rosca, A\. Puigdomènech, Z\. Yahav, S\. Ghemawat, H\. Lin, P\. Kirk, Z\. Nabulsi, S\. Brin, B\. Bohnet, K\. Caluwaerts, A\. S\. Veerubhotla, D\. Zheng, Z\. Dai, P\. Petrov, Y\. Xu, R\. Mehran, Z\. Xu, L\. Zintgraf, J\. Choi, S\. A\. Hombaiah, R\. Thoppilan, S\. Reddi, L\. Lew, L\. Li, K\. Webster, K\. Sawhney, L\. Lamprou, S\. Shakeri, M\. Lunayach, J\. Chen, S\. Bagri, A\. Salcianu, Y\. Chen, Y\. Donchev, C\. Magister, S\. Nørly, V\. Rodrigues, T\. Izo, H\. Noga, J\. Zou, T\. Köppe, W\. Zhou, K\. Lee, X\. Long, D\. Eisenbud, A\. Chen, C\. Schenck, C\. M\. To, P\. Zhong, E\. Taropa, M\. Truong, O\. Levy, D\. Martins, Z\. Zhang, C\. Semturs, K\. Zhang, A\. Yakubovich, P\. Moreno, L\. McConnaughey, D\. Lu, S\. Redmond, L\. Weerts, Y\. Bitton, T\. Refice, N\. Lacasse, A\. Conmy, C\. Tallec, J\. Odell, H\. Forbes\-Pollard, A\. Socala, J\. Hoech, P\. Kohli, A\. Walton, R\. Wang, M\. Sazanovich, K\. Zhu, A\. Kapishnikov, R\. Galt, M\. Denton, B\. Murdoch, C\. Sikora, K\. Mohamed, W\. Wei, U\. First, T\. McConnell, L\. C\. Cobo, J\. Qin, T\. Avrahami, D\. Balle, Y\. Watanabe, A\. Louis, A\. Kraft, S\. Ariafar, Y\. Gu, E\. Rives, C\. Yoon, A\. Rusu, J\. Cobon\-Kerr, C\. Hahn, J\. Luo, Yuvein, Zhu, N\. Ahuja, R\. Benenson, R\. L\. Kaufman, H\. Yu, L\. Hightower, J\. Zhang, D\. Ni, L\. A\. Hendricks, G\. Wang, G\. Yona, L\. Jain, P\. Barrio, S\. Bhupatiraju, S\. Velusamy, A\. Dafoe, S\. Riedel, T\. Thomas, Z\. Yuan, M\. Bellaiche, S\. Panthaplackel, K\. Kloboves, S\. Jauhari, C\. Akbulut, T\. Davchev, E\. Gladchenko, D\. Madras, A\. Chuklin, T\. Hill, Q\. Yuan, M\. Madhavan, L\. Leonhard, D\. Scandinaro, Q\. Chen, N\. Niu, A\. Douillard, B\. Damoc, Y\. Onoe, F\. Pedregosa, F\. Bertsch, C\. Leichner, J\. Pagadora, J\. Malmaud, S\. Ponda, A\. Twigg, O\. Duzhyi, J\. Shen, M\. Wang, R\. Garg, J\. Chen, U\. Evci, J\. Lee, L\. Liu, K\. Kojima, M\. Yamaguchi, A\. Rajendran, A\. Piergiovanni, V\. K\. Rajendran, M\. Fornoni, G\. Ibagon, H\. Ragan, S\. M\. Khan, J\. Blitzer, A\. Bunner, G\. Sun, T\. Kosakai, S\. Lundberg, N\. Elue, K\. Guu, S\. Park, J\. Park, A\. Narayanaswamy, C\. Wu, J\. Mudigonda, T\. Cohn, H\. Mu, R\. Kumar, L\. Graesser, Y\. Zhang, R\. Killam, V\. Zhuang, M\. Giménez, W\. A\. Jishi, R\. Ley\-Wild, A\. Zhai, K\. Osawa, D\. Cedillo, J\. Liu, M\. Upadhyay, M\. Sieniek, R\. Sharma, T\. Paine, A\. Angelova, S\. Addepalli, C\. Parada, K\. Majumder, A\. Lamp, S\. Kumar, X\. Deng, A\. Myaskovsky, T\. Sabolić, J\. Dudek, S\. York, F\. de Chaumont Quitry, J\. Nie, D\. Cattle, A\. Gunjan, B\. Piot, W\. Khawaja, S\. Bang, S\. Wang, S\. Khodadadeh, R\. R, P\. Rawlani, R\. Powell, K\. Lee, J\. Griesser, G\. Oh, C\. Magalhaes, Y\. Li, S\. Tokumine, H\. N\. Vogel, D\. Hsu, A\. BC, D\. Jindal, M\. Cohen, Z\. Yang, J\. Yuan, D\. de Cesare, T\. Bruguier, J\. Xu, M\. Roy, A\. Jacovi, D\. Belov, R\. Arya, P\. Meadowlark, S\. Cohen\-Ganor, W\. Ye, P\. Morris\-Suzuki, P\. Banzal, G\. Song, P\. Ponnuramu, F\. Zhang, G\. Scrivener, S\. Zaiem, A\. R\. Rochman, K\. Han, B\. Ghazi, K\. Lee, S\. Drath, D\. Suo, A\. Girgis, P\. Shenoy, D\. Nguyen, D\. Eck, S\. Gupta, L\. Yan, J\. Carreira, A\. Gulati, R\. Sang, D\. Mirylenka, E\. Cooney, E\. Chou, M\. Ling, C\. Fan, B\. Coleman, G\. Tubone, R\. Kumar, J\. Baldridge, F\. Hernandez\-Campos, A\. Lazaridou, J\. Besley, I\. Yona, N\. Bulut, Q\. Wellens, A\. Pierigiovanni, J\. George, R\. Green, P\. Han, C\. Tao, G\. Clark, C\. You, A\. Abdolmaleki, J\. Fu, T\. Chen, A\. Chaugule, A\. Chandorkar, A\. Rahman, W\. Thompson, P\. Koanantakool, M\. Bernico, J\. Ren, A\. Vlasov, S\. Vassilvitskii, M\. Kula, Y\. Liang, D\. Kim, Y\. Huang, C\. Ye, D\. Lepikhin, and W\. HelmholzGemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Fenget al\.\(2025\)T\. Feng, Y\. Sun, and J\. YouGraphEval: a lightweight graph\-based LLM framework for idea evaluation\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=5RUM1aIdok)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Fortunatoet al\.\(2018\)S\. Fortunato, C\. T\. Bergstrom, K\. Börner, J\. A\. Evans, D\. Helbing, S\. Milojević, A\. M\. Petersen, F\. Radicchi, R\. Sinatra, B\. Uzzi, A\. Vespignani, L\. Waltman, D\. Wang, and A\. BarabásiScience of science\.Science359\(6379\),pp\. eaao0185\.External Links:[Document](https://dx.doi.org/10.1126/science.aao0185),[Link](https://www.science.org/doi/abs/10.1126/science.aao0185),https://www\.science\.org/doi/pdf/10\.1126/science\.aao0185Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p1.1)\. - Gómez\-Pérezet al\.\(2022\)J\. M\. Gómez\-Pérez, A\. García\-Silva, R\. Leone, M\. Albani, M\. Fontaine, C\. Poncet, L\. Summerer, A\. Donati, I\. Roma, and S\. ScaglioniArtificial intelligence and natural language processing and understanding in space: a methodological framework and four esa case studies\.External Links:2210\.03640,[Link](https://arxiv.org/abs/2210.03640)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Googleet al\.\(2025\)Google, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p2.1)\. - Google \(2025\)GoogleA new era of intelligence with Gemini 3 — blog\.google\.Note:[https://blog\.google/products/gemini/gemini\-3/\#note\-from\-ceo](https://blog.google/products/gemini/gemini-3/#note-from-ceo)\[Accessed 05\-01\-2026\]Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Gottesman and Geva \(2024\)D\. Gottesman and M\. GevaEstimating knowledge in large language models without generating a single token\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3994–4019\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.232/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.232)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Gottweiset al\.\(2026\)J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, P\. Sirkovic, A\. Myaskovsky, G\. Glowaty, F\. Weissenberger, A\. Orlandi, D\. Popovici, A\. Palepu, K\. Rong, R\. Tanno, K\. Saab, F\. Zhang, J\. Blum, A\. Carroll, K\. Kulkarni, N\. Tomašev, D\. Zverinski, I\. Rendulic, E\. Vedadi, F\. Hasler, L\. Rimanic, M\. Boia, I\. Budiselic, B\. Feinstein, M\. Bellaiche, T\. Sheffer, J\. Freyberg, J\. Ratcliff, O\. Bertolli, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Matias, J\. Manyika, D\. Hassabis, Y\. Xu, P\. Kohli, A\. Pawlosky, A\. Karthikesalingam, and V\. NatarajanAccelerating scientific discovery with co\-scientist\.Nature\.External Links:ISSN 1476\-4687,[Document](https://dx.doi.org/10.1038/s41586-026-10644-y),[Link](https://doi.org/10.1038/s41586-026-10644-y)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p2.1)\. - Gurnee and Tegmark \(2024\)W\. Gurnee and M\. TegmarkLanguage models represent space and time\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jE8xbmvFin)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Heet al\.\(2024\)L\. He, P\. Chen, E\. Nie, Y\. Li, and J\. R\. BrennanDecoding probing: revealing internal linguistic structures in neural language models using minimal pairs\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 4488–4497\.External Links:[Link](https://aclanthology.org/2024.lrec-main.402/)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Houet al\.\(2026\)J\. Hou, H\. Deng, W\. Jiao, X\. Liu, X\. Ke, and M\. ZhangNoveltyAgent: autonomous novelty reporting agent with point\-wise novelty analysis and self\-validation\.External Links:2603\.20884,[Link](https://arxiv.org/abs/2603.20884)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Huet al\.\(2022\)E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Appendix C](https://arxiv.org/html/2608.25660#A3.p1.1),[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Jinet al\.\(2025\)M\. Jin, Q\. Yu, J\. Huang, Q\. Zeng, Z\. Wang, W\. Hua, H\. Zhao, K\. Mei, Y\. Meng, K\. Ding, F\. Yang, M\. Du, and Y\. ZhangExploring concept depth: how large language models acquire knowledge and concept at different layers?\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 558–573\.External Links:[Link](https://aclanthology.org/2025.coling-main.37/)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Juet al\.\(2024\)T\. Ju, W\. Sun, W\. Du, X\. Yuan, Z\. Ren, and G\. LiuHow large language models encode context knowledge? a layer\-wise probing study\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 8235–8246\.External Links:[Link](https://aclanthology.org/2024.lrec-main.722/)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Kleringset al\.\(2025\)A\. Klerings, J\. Brinkmann, D\. Ruffinelli, and S\. P\. PonzettoSteering language models in multi\-token generation: a case study on tense and aspect\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 8621–8639\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.435/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.435),ISBN 979\-8\-89176\-332\-6Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Liet al\.\(2023\)K\. Li, A\. K\. Hopkins, D\. Bau, F\. Viégas, H\. Pfister, and M\. WattenbergEmergent world representations: exploring a sequence model trained on a synthetic task\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=DeG07_TcZvT)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Liet al\.\(2025\)L\. Li, W\. Xu, J\. Guo, R\. Zhao, X\. Li, Y\. Yuan, B\. Zhang, Y\. Jiang, Y\. Xin, R\. Dang, Y\. Rong, D\. Zhao, T\. Feng, and L\. BingChain of ideas: revolutionizing research via novel idea development with LLM agents\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 8971–9004\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.477/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.477),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Liet al\.\(2024a\)L\. Li, W\. Xu, J\. Guo, R\. Zhao, X\. Li, Y\. Yuan, B\. Zhang, Y\. Jiang, Y\. Xin, R\. Dang, D\. Zhao, Y\. Rong, T\. Feng, and L\. BingChain of ideas: revolutionizing research via novel idea development with llm agents\.External Links:2410\.13185,[Link](https://arxiv.org/abs/2410.13185)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Liet al\.\(2024b\)Z\. Li, Y\. Cao, and J\. C\.K\. CheungDo llms build world representations? probing through the lens of state abstraction\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 98009–98032\.External Links:[Document](https://dx.doi.org/10.52202/079017-3110),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/b1b16c4b875eb84d3585cb70d23970ca-Paper-Conference.pdf)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Linet al\.\(2025\)E\. Lin, Z\. Peng, and Y\. FangEvaluating and enhancing large language models for novelty assessment in scholarly publications\.InProceedings of the 1st Workshop on AI and Scientific Discovery: Directions and Opportunities,P\. Jansen, B\. Dalvi Mishra, H\. Trivedi, B\. Prasad Majumder, T\. Hope, T\. Khot, D\. Downey, and E\. Horvitz \(Eds\.\),Albuquerque, New Mexico, USA,pp\. 46–57\.External Links:[Link](https://aclanthology.org/2025.aisd-main.5/),[Document](https://dx.doi.org/10.18653/v1/2025.aisd-main.5),ISBN 979\-8\-89176\-224\-4Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Liuet al\.\(2025\)Y\. Liu, Z\. Yang, S\. Poria, T\. Nguyen, and E\. CambriaHarnessing large language models for scientific novelty detection\.External Links:2505\.24615,[Link](https://arxiv.org/abs/2505.24615)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Luet al\.\(2024\)C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. HaThe ai scientist: towards fully automated open\-ended scientific discovery\.External Links:2408\.06292,[Link](https://arxiv.org/abs/2408.06292)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Luet al\.\(2026\)C\. Lu, C\. Lu, R\. T\. Lange, Y\. Yamada, S\. Hu, J\. Foerster, D\. Ha, and J\. CluneTowards end\-to\-end automation of ai research\.Nature651\(8107\),pp\. 914–919\.External Links:ISSN 1476\-4687,[Document](https://dx.doi.org/10.1038/s41586-026-10265-5),[Link](https://doi.org/10.1038/s41586-026-10265-5)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Maaset al\.\(2011\)A\. L\. Maas, R\. E\. Daly, P\. T\. Pham, D\. Huang, A\. Y\. Ng, and C\. PottsLearning word vectors for sentiment analysis\.InProceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies,D\. Lin, Y\. Matsumoto, and R\. Mihalcea \(Eds\.\),Portland, Oregon, USA,pp\. 142–150\.External Links:[Link](https://aclanthology.org/P11-1015/)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Maiyaet al\.\(2025\)S\. Maiya, Y\. Liu, R\. Debnath, and A\. KorhonenImproving preference extraction in LLMs by identifying latent knowledge through classifying probes\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 9061–9081\.External Links:[Link](https://aclanthology.org/2025.acl-long.444/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.444),ISBN 979\-8\-89176\-251\-0Cited by:[footnote 1](https://arxiv.org/html/2608.25660#footnote1)\. - Marks and Tegmark \(2024\)S\. Marks and M\. TegmarkThe geometry of truth: emergent linear structure in large language model representations of true/false datasets\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=aajyHYjjsk)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Mostafaet al\.\(2026\)A\. Mostafa, T\. H\. Nguyen, and Z\. AhmadiWhat is novel? a knowledge\-driven framework for bias\-aware literature originality evaluation\.External Links:2602\.06054,[Link](https://arxiv.org/abs/2602.06054)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Mysoreet al\.\(2022\)S\. Mysore, A\. Cohan, and T\. HopeMulti\-vector models with textual guidance for fine\-grained scientific document similarity\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 4453–4470\.External Links:[Link](https://aclanthology.org/2022.naacl-main.331/),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.331)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - OpenAIet al\.\(2025\)OpenAI, :, S\. Agarwal, L\. Ahmad, J\. Ai, S\. Altman, A\. Applebaum, E\. Arbus, R\. K\. Arora, Y\. Bai, B\. Baker, H\. Bao, B\. Barak, A\. Bennett, T\. Bertao, N\. Brett, E\. Brevdo, G\. Brockman, S\. Bubeck, C\. Chang, K\. Chen, M\. Chen, E\. Cheung, A\. Clark, D\. Cook, M\. Dukhan, C\. Dvorak, K\. Fives, V\. Fomenko, T\. Garipov, K\. Georgiev, M\. Glaese, T\. Gogineni, A\. Goucher, L\. Gross, K\. G\. Guzman, J\. Hallman, J\. Hehir, J\. Heidecke, A\. Helyar, H\. Hu, R\. Huet, J\. Huh, S\. Jain, Z\. Johnson, C\. Koch, I\. Kofman, D\. Kundel, J\. Kwon, V\. Kyrylov, E\. Y\. Le, G\. Leclerc, J\. P\. Lennon, S\. Lessans, M\. Lezcano\-Casado, Y\. Li, Z\. Li, J\. Lin, J\. Liss, Lily, Liu, J\. Liu, K\. Lu, C\. Lu, Z\. Martinovic, L\. McCallum, J\. McGrath, S\. McKinney, A\. McLaughlin, S\. Mei, S\. Mostovoy, T\. Mu, G\. Myles, A\. Neitz, A\. Nichol, J\. Pachocki, A\. Paino, D\. Palmie, A\. Pantuliano, G\. Parascandolo, J\. Park, L\. Pathak, C\. Paz, L\. Peran, D\. Pimenov, M\. Pokrass, E\. Proehl, H\. Qiu, G\. Raila, F\. Raso, H\. Ren, K\. Richardson, D\. Robinson, B\. Rotsted, H\. Salman, S\. Sanjeev, M\. Schwarzer, D\. Sculley, H\. Sikchi, K\. Simon, K\. Singhal, Y\. Song, D\. Stuckey, Z\. Sun, P\. Tillet, S\. Toizer, F\. Tsimpourlas, N\. Vyas, E\. Wallace, X\. Wang, M\. Wang, O\. Watkins, K\. Weil, A\. Wendling, K\. Whinnery, C\. Whitney, H\. Wong, L\. Yang, Y\. Yang, M\. Yasunaga, K\. Ying, W\. Zaremba, W\. Zhan, C\. Zhang, B\. Zhang, E\. Zhang, and S\. ZhaoGpt\-oss\-120b & gpt\-oss\-20b model card\.External Links:2508\.10925,[Link](https://arxiv.org/abs/2508.10925)Cited by:[§B\.2](https://arxiv.org/html/2608.25660#A2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p2.1)\. - OpenAI \(2026\)OpenAIGPT\-5\.4 thinking system card\.OpenAI\.Note:Accessed: 2026\-05\-20External Links:[Link](https://deploymentsafety.openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf)Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Pedregosaet al\.\(2011\)F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and E\. DuchesnayScikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[Appendix D](https://arxiv.org/html/2608.25660#A4.p1.1)\. - Picardet al\.\(2025\)C\. Picard, K\. M\. Edwards, A\. C\. Doris, B\. Man, G\. Giannone, M\. F\. Alam, and F\. AhmedFrom concept to manufacturing: evaluating vision\-language models for engineering design\.Artificial Intelligence Review58\(9\),pp\. 288\.External Links:[Document](https://dx.doi.org/10.1007/s10462-025-11290-y),ISBN 1573\-7462,[Link](https://doi.org/10.1007/s10462-025-11290-y)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p1.1)\. - Saricaet al\.\(2020\)S\. Sarica, J\. Luo, and K\. L\. WoodTechNet: technology semantic network based on patent data\.Expert Systems with Applications142,pp\. 112995\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.eswa.2019.112995),[Link](https://www.sciencedirect.com/science/article/pii/S0957417419307122)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Schopf and Färber \(2026\)T\. Schopf and M\. FärberIs this idea novel? an automated benchmark for judgment of research ideas\.InProceedings of the Fifteenth Language Resources and Evaluation Conference \(LREC 2026\),S\. Piperidis, N\. Bel, H\. van den Heuvel, N\. Ide, S\. Krek, and A\. Toral \(Eds\.\),Palma, Mallorca, Spain,pp\. 4716–4727\.External Links:[Document](https://dx.doi.org/10.63317/4c3gy3f7epnj)Cited by:[Appendix B](https://arxiv.org/html/2608.25660#A2.p1.1),[§1](https://arxiv.org/html/2608.25660#S1.p2.1),[§3](https://arxiv.org/html/2608.25660#S3.p1.1),[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Shahidet al\.\(2025\)S\. Shahid, M\. Radensky, R\. Fok, P\. Siangliulue, D\. S\. Weld, and T\. HopeLiterature\-grounded novelty assessment of scientific ideas\.InProceedings of the Fifth Workshop on Scholarly Document Processing \(SDP 2025\),T\. Ghosal, P\. Mayr, A\. Singh, A\. Naik, G\. Rehm, D\. Freitag, D\. Li, S\. Schimmler, and A\. De Waard \(Eds\.\),Vienna, Austria,pp\. 96–113\.External Links:[Link](https://aclanthology.org/2025.sdp-1.9/),[Document](https://dx.doi.org/10.18653/v1/2025.sdp-1.9),ISBN 979\-8\-89176\-265\-7Cited by:[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Siet al\.\(2025\)C\. Si, D\. Yang, and T\. HashimotoCan llms generate novel research ideas? a large\-scale human study with 100\+ nlp researchers\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 94003–94092\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/ea94957d81b1c1caf87ef5319fa6b467-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1),[§2](https://arxiv.org/html/2608.25660#S2.p1.1),[§4](https://arxiv.org/html/2608.25660#S4.p1.1),[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Singhet al\.\(2026\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. de Avila Belbute Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Y\. Guan, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Korbak, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. WangOpenAI gpt\-5 system card\.External Links:2601\.03267,[Link](https://arxiv.org/abs/2601.03267)Cited by:[§4](https://arxiv.org/html/2608.25660#S4.p1.1)\. - Suet al\.\(2025\)H\. Su, R\. Chen, S\. Tang, Z\. Yin, X\. Zheng, J\. Li, B\. Qi, Q\. Wu, H\. Li, W\. Ouyang, P\. Torr, B\. Zhou, and N\. DongMany heads are better than one: improved scientific idea generation by a LLM\-based multi\-agent system\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 28201–28240\.External Links:[Link](https://aclanthology.org/2025.acl-long.1368/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1368),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Tanget al\.\(2025\)J\. Tang, L\. Xia, Z\. Li, and C\. HuangAI\-researcher: autonomous scientific innovation\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=kQWyOYUAC4)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Uzziet al\.\(2013\)B\. Uzzi, S\. Mukherjee, M\. Stringer, and B\. JonesAtypical combinations and scientific impact\.Science342\(6157\),pp\. 468–472\.External Links:[Document](https://dx.doi.org/10.1126/science.1240474),[Link](https://www.science.org/doi/abs/10.1126/science.1240474),https://www\.science\.org/doi/pdf/10\.1126/science\.1240474Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Valoiset al\.\(2025\)P\. H\. V\. Valois, L\. S\. Souza, E\. K\. Shimomoto, and K\. FukuiFrame representation hypothesis: multi\-token llm interpretability and concept\-guided text generation\.Transactions of the Association for Computational Linguistics13,pp\. 1436–1458\.External Links:ISSN 2307\-387X,[Document](https://dx.doi.org/10.1162/TACL.a.48),[Link](https://doi.org/10.1162/TACL.a.48),https://direct\.mit\.edu/tacl/article\-pdf/doi/10\.1162/TACL\.a\.48/2561673/tacl\.a\.48\.pdfCited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. - Wanget al\.\(2017\)J\. Wang, R\. Veugelers, and P\. StephanBias against novelty in science: a cautionary tale for users of bibliometric indicators\.Research Policy46\(8\),pp\. 1416–1436\.External Links:ISSN 0048\-7333,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.respol.2017.06.006),[Link](https://www.sciencedirect.com/science/article/pii/S0048733317301038)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Wanget al\.\(2019\)K\. Wang, B\. Dong, and J\. MaTowards computational assessment of idea novelty\.InProceedings of the 52nd Hawaii International Conference on System Sciences,External Links:ISBN 978\-0\-9981331\-2\-6,[Link](https://ssrn.com/abstract=3393611)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Wanget al\.\(2025\)W\. Wang, L\. Gu, L\. Zhang, Y\. Luo, Y\. Dai, C\. Shen, L\. Xie, B\. Lin, X\. He, and J\. YeSciPIP: an llm\-based scientific paper idea proposer\.External Links:2410\.23166,[Link](https://arxiv.org/abs/2410.23166)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. RushTransformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Q\. Liu and D\. Schlangen \(Eds\.\),Online,pp\. 38–45\.External Links:[Link](https://aclanthology.org/2020.emnlp-demos.6/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6)Cited by:[Appendix D](https://arxiv.org/html/2608.25660#A4.p1.1)\. - Wuet al\.\(2025\)W\. Wu, C\. Zhang, and Y\. ZhaoAutomated novelty evaluation of academic paper: a collaborative approach integrating human and large language model knowledge\.Journal of the Association for Information Science and Technology76\(11\),pp\. 1452–1469\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/asi.70005),[Link](https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.70005),https://asistdl\.onlinelibrary\.wiley\.com/doi/pdf/10\.1002/asi\.70005Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Wuet al\.\(2026\)W\. Wu, Y\. Zhao, Y\. Wang, S\. Li, J\. Shao, Y\. Long, and C\. ZhangNovBench: evaluating large language models on academic paper novelty assessment\.External Links:2604\.11543,[Link](https://arxiv.org/abs/2604.11543)Cited by:[§1](https://arxiv.org/html/2608.25660#S1.p2.1)\. - Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p2.1)\. - Yanget al\.\(2024\)Z\. Yang, X\. Du, J\. Li, J\. Zheng, S\. Poria, and E\. CambriaLarge language models for automated open\-domain scientific hypotheses discovery\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 13545–13565\.External Links:[Link](https://aclanthology.org/2024.findings-acl.804/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.804)Cited by:[§5\.1](https://arxiv.org/html/2608.25660#S5.SS1.p1.1)\. - Zhanget al\.\(2025\)Y\. Zhang, H\. Diddee, S\. Holm, H\. Liu, X\. Liu, V\. Samuel, B\. Wang, and D\. IppolitoNoveltyBench: evaluating creativity and diversity in language models\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=XZm1ekzERf)Cited by:[§2](https://arxiv.org/html/2608.25660#S2.p1.1)\. - Zuret al\.\(2025\)A\. Zur, A\. Geiger, E\. S\. Lubana, and E\. BigelowAre language models aware of the road not taken? token\-level uncertainty and hidden state dynamics\.External Links:2511\.04527,[Link](https://arxiv.org/abs/2511.04527)Cited by:[Appendix A](https://arxiv.org/html/2608.25660#A1.p1.1)\. ## Appendix AAbout Probing[llm](https://arxiv.org/html/2608.25660#id1)During Generation Probing approaches quantify the extent to which[llm](https://arxiv.org/html/2608.25660#id1)representations encode specific knowledge\. While extensive research investigates internal knowledge across diverse domains such as sentiment[Maas et al\. \(2011\)](https://arxiv.org/html/2608.25660#bib.bib50)and factual knowledge[Marks and Tegmark \(2024\)](https://arxiv.org/html/2608.25660#bib.bib16), spatial and temporal understanding[Gurnee and Tegmark \(2024\)](https://arxiv.org/html/2608.25660#bib.bib17), and world models[Li et al\. \(2023\)](https://arxiv.org/html/2608.25660#bib.bib18), such existing studies predominantly focus on layer\-wise localization of internal knowledge\([He et al\., 2024](https://arxiv.org/html/2608.25660#bib.bib51);[Li et al\., 2024b](https://arxiv.org/html/2608.25660#bib.bib13);[Ju et al\., 2024](https://arxiv.org/html/2608.25660#bib.bib52);[Jin et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib57),inter alia\)\. Work on probing[llm](https://arxiv.org/html/2608.25660#id1)during different generation steps is scarce and primarily addresses steering text generation\([Valois et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib14);[Zur et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib15);[Klerings et al\., 2025](https://arxiv.org/html/2608.25660#bib.bib53),inter alia\)\. Although some works investigate the encoded information in[llm](https://arxiv.org/html/2608.25660#id1)before generating the first token[Gottesman and Geva \(2024\)](https://arxiv.org/html/2608.25660#bib.bib58);[Afzal et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib54), they do not distinguish between functional phases during generation\. We address this gap by providing the first comparison of[llm](https://arxiv.org/html/2608.25660#id1)representations during thereasoning\(“thinking”\) phase versus theresponse generationphase via probing, demonstrating that[llm](https://arxiv.org/html/2608.25660#id1)encode more information about research idea novelty judgments whilethinkingthan when producing the actual response\. Table 5:Novelty Judgment Rubric ## Appendix BEvaluation Metrics We briefly summarize the[rino](https://arxiv.org/html/2608.25660#id3)evaluation metrics, which we adopt in this work\. For full details, see[Schopf and Färber \(2026\)](https://arxiv.org/html/2608.25660#bib.bib32)\. The metrics evaluate both numerical novelty scores and textual justifications\. ### B\.1Novelty Score Metrics We evaluate predicted novelty scores using macro\-F1F\_\{1\}, class\-wiseF1F\_\{1\}, and[mae](https://arxiv.org/html/2608.25660#id2)\. Macro\-F1F\_\{1\}measures overall classification performance across the five novelty categories, class\-wiseF1F\_\{1\}shows performance for each individual score, and[mae](https://arxiv.org/html/2608.25660#id2)measures the average distance between predicted and human gold scores on the ordinal 1–5 scale\. ### B\.2Justification Metrics For textual justifications,[rino](https://arxiv.org/html/2608.25660#id3)distinguishes betweenknown aspects, which describe overlaps with prior work, andnovelty aspects, which describe new contributions of the research idea\. Following[rino](https://arxiv.org/html/2608.25660#id3), these metrics are computed using an[llm](https://arxiv.org/html/2608.25660#id1)\-as\-a\-judge approach that compares model\-generated justifications against human gold\-standard justifications\. In this work, we use the GPT\-OSS\-120B[OpenAI et al\. \(2025\)](https://arxiv.org/html/2608.25660#bib.bib8)model for evaluation\. AlignmentAlignment measures whether the model\-generated justification follows reasoning consistent with the human gold justification and supports a similar novelty judgment\. Scores range from 0 to 1, where higher is better: 1 indicates strong agreement with the human rationale, while 0 indicates no alignment\. RecallRecall measures how many known\-aspect and novelty\-aspect arguments from the human gold justification are captured by the model\-generated justification\. Scores range from 0 to 100, where higher is better: 100 means all relevant gold arguments are covered, while 0 means none are covered\. Additional RatioAdditional ratio measures how many extra known\-aspect and novelty\-aspect arguments the model adds beyond the gold justification, while still being grounded in the related works or research idea\. Scores are non\-negative percentages, where 0% means no additional grounded arguments are added and higher values indicate more extra grounded content\. This metric is not inherently good or bad: moderate or high values can indicate useful additional evidence, but very high values may also reflect overly verbose justifications\. Hallucination RateHallucination rate measures the proportion of generated known\-aspect and novelty\-aspect arguments that are not supported by the related works or research idea\. Scores range from 0% to 100%, where lower is better: 0% indicates justifications that are fully grounded in the research idea and related works, while higher values indicate more unsupported or hallucinated content\. ## Appendix CLoRA Fine\-tuning Details For the FineTune approach in Section[5](https://arxiv.org/html/2608.25660#S5), we fine\-tune the base[llm](https://arxiv.org/html/2608.25660#id1)using Low\-Rank Adaptation \(LoRA;[Hu et al\., 2022](https://arxiv.org/html/2608.25660#bib.bib23)\)\. We apply LoRA to all major projection layers in the transformer, including the query, key, value, and output projections of the attention mechanism, as well as the gate, up, and down projections in the feed\-forward network\. We use a rank ofr=16r=16, a scaling factorα=32\\alpha=32, and a LoRA dropout of0\.10\.1\. Training is performed for two epochs using a per\-device batch size of 1 and gradient accumulation over 8 steps, resulting in an effective batch size of 8\. We employ a learning rate of2×10−42\\times 10^\{\-4\}with a short warmup of 20 steps\. To reduce memory consumption, gradient checkpointing is enabled, and training is conducted in bfloat16 precision\. ## Appendix DExperimental Details All experiments were conducted on two NVIDIA A100 \(80GB\) GPUs\. The probing classifier was implemented using scikit\-learn[Pedregosa et al\. \(2011\)](https://arxiv.org/html/2608.25660#bib.bib40)\. Hidden states were extracted using the Hugging Face Transformers library[Wolf et al\. \(2020\)](https://arxiv.org/html/2608.25660#bib.bib55)\. ## Appendix EOn the Choice of Novelty Score Metrics We use macro\-F1F\_\{1\}as the primary metric for evaluating novelty score predictions, rather than[mae](https://arxiv.org/html/2608.25660#id2)\. Although[mae](https://arxiv.org/html/2608.25660#id2)is useful for measuring the average ordinal distance between predicted and gold scores, it is less informative in our setting because the novelty scale is small \(1–5\) and model predictions are strongly concentrated around the middle categories\. As shown in Table[6](https://arxiv.org/html/2608.25660#A5.T6), different models and approaches obtain very similar[mae](https://arxiv.org/html/2608.25660#id2)values, typically around one\. This makes[mae](https://arxiv.org/html/2608.25660#id2)unsuitable for evaluation in our setting\. Table 6:[mae](https://arxiv.org/html/2608.25660#id2)scores for different approaches and[llm](https://arxiv.org/html/2608.25660#id1)on the[rino](https://arxiv.org/html/2608.25660#id3)test set\.The limitation arises because a model that repeatedly predicts a middle score, such as 3, can achieve a deceptively low[mae](https://arxiv.org/html/2608.25660#id2): many gold labels are only one or two points away on a five\-point Likert scale\. Conversely, a less biased model that also predicts more extreme novelty categories may occasionally incur larger absolute errors, even if it produces more accurate novelty judgments overall\. Optimizing for[mae](https://arxiv.org/html/2608.25660#id2)can therefore favor conservative middle\-ground predictions, precisely the behavior we aim to mitigate\. Macro\-F1F\_\{1\}better reflects our evaluation objective\. It treats all novelty classes equally, regardless of their frequency, and explicitly rewards models for correctly predicting low, medium, and high novelty judgments\. This is crucial for evaluating whether a model can correctly judge extreme novelty categories, such as“not novel”and“highly novel”, rather than merely staying close to the center of the scale\. We therefore report[mae](https://arxiv.org/html/2608.25660#id2)for completeness in Table[6](https://arxiv.org/html/2608.25660#A5.T6), but use macro\-F1F\_\{1\}and class\-wiseF1F\_\{1\}as the main indicators of novelty judgment performance\. ## Appendix FJustification Evaluation Beyond novelty judgment performance, we also evaluate the quality of textual justifications generated by open\-source[llm](https://arxiv.org/html/2608.25660#id1)using[tpr](https://arxiv.org/html/2608.25660#id5), with the results shown in Table[7](https://arxiv.org/html/2608.25660#A6.T7)\. Table 7:Evaluation of textual justifications generated by different[llm](https://arxiv.org/html/2608.25660#id1)using[tpr](https://arxiv.org/html/2608.25660#id5)\.AlignmentThe alignment scores of open\-source models using[tpr](https://arxiv.org/html/2608.25660#id5)are comparable to those of substantially larger proprietary[llm](https://arxiv.org/html/2608.25660#id1)in Table[1](https://arxiv.org/html/2608.25660#S3.T1)\. In particular, the strongest reasoning\-capable models achieve ALI values around 0\.5, close to the range observed for Claude and GPT models under zero\-shot prompting\. This indicates that[tpr](https://arxiv.org/html/2608.25660#id5)does not merely improve numerical novelty prediction, but also enables open\-source models to generate justifications whose reasoning remains broadly aligned with human gold\-standard rationales\. RecallSimilarly to propriety[llm](https://arxiv.org/html/2608.25660#id1), smaller open\-source models exhibit relatively high recall, indicating substantial overlap between model\-generated and human\-annotated justification arguments\. Comparing the results to the ones in Table[1](https://arxiv.org/html/2608.25660#S3.T1), recall in open\-source models is slightly lower than in the top\-performing OpenAI models but remains competitive with Gemini Pro models\. Additional RatioTheAdditional Ratiois generally higher for reasoning\-capable models than for non\-reasoning models, indicating that reasoning models produce more elaborate justifications\. While proprietary OpenAI[llm](https://arxiv.org/html/2608.25660#id1)exhibit even higher additional ratios, the best\-performing Qwen3 models remain competitive with Gemini Pro models, highlighting that[tpr](https://arxiv.org/html/2608.25660#id5)enables smaller open\-source models to generate rich novelty judgment justifications\. Hallucination RateReasoning models show low hallucination rates similar to proprietary models, whereas non\-reasoning models, particularly Gemma\-3\-4B, display higher hallucination rates when attempting to justify novelty judgments\. This suggests that[tpr](https://arxiv.org/html/2608.25660#id5)is most effective in producing grounded justifications when paired with reasoning[llm](https://arxiv.org/html/2608.25660#id1)\. TakeawayOverall,[tpr](https://arxiv.org/html/2608.25660#id5)allows open\-source LLMs to generate high\-quality, human\-aligned novelty justifications, similar to what is achievable with large proprietary models\. However, performance depends strongly on the model:reasoning\-capable[llm](https://arxiv.org/html/2608.25660#id1)consistently produce more accurate, elaborate, and reliable justifications, whereas smaller non\-reasoning[llm](https://arxiv.org/html/2608.25660#id1)may struggle to generate appropriate novelty judgment justifications\. Table 8:Selected comparison of gold novelty judgments and LLM\(Claude Opus 4\.5\)\-generated novelty judgments\.Greendenotes alignment of model and gold justifications\.Redhighlights miscalibration, where the model’s novelty judgment diverges from the gold judgment despite exhibiting a rationale aligned with the gold justification\. Symbols indicate novelty overestimation \(↑\\uparrow\), correct prediction \(✓\\checkmark\), and underestimation \(↓\\downarrow\)\.systemprompt="Youareanexpertresearcherexperiencedinjudgingthenoveltyofaresearchidea\." userprompt=f""" Youareanexpertinmachinelearningresearchevaluation\.Youwillbegiventwoinputs: 1\.Aresearchideawithobjective,problemstatement,andsolutionapproach\. 2\.Alistofrelatedworks,eachwithatitleandabstract\. Yourtaskisto\*\*assessthenoveltyoftheresearchidea\*\*comparedtotherelatedworks\. \#\#\#Instructions: \-Analyzetheresearchideaandsummarizeitskeycontributions\. \-Compareitwiththerelatedworkstoidentifyoverlapsanddifferences\. \-Specifically,assesswhethertheideaintroduces\*\*significantnewaspects\*\*notpresentinexistingwork,orifitislargelyavariationonknownapproaches\. \-Provideyouroutputasa\*\*JSONobjectonly\*\*,with: \-"reasoning":ashortparagraph\(2\-4sentences\)explainingthereasoningbehindthenoveltyscore\. \-"novelty\_score":anintegerbetween1\-5where:\{novelty\_rubric\} \#\#\#Inputs: \*\*ResearchIdea:\*\* \{research\_idea\} \*\*RelatedWorks:\*\* \{related\_works\} \#\#\#OutputFormat: “‘json \{\{ "reasoning":<shortexplanation\>, "novelty\_score":<1\|2\|3\|4\|5\> \}\} """ Figure 3:Prompts for the zero\-shot approach to judging the novelty of research ideas\. Here, an[llm](https://arxiv.org/html/2608.25660#id1)receives a research idea, its related works, and the[rino](https://arxiv.org/html/2608.25660#id3)novelty rubric, and is asked to generate both a numerical novelty score and a textual justification\.systemprompt="YouareReviewerGPT,anintelligentassistantthathelpsresearchersevaluatethenoveltyoftheirideas\." userprompt=f""" Youaregivensomepaperssimilartotheproposedidea\(<IDEA\>and</IDEA\>\)\.Yourtaskistoevaluatetheidea’snoveltyusingtherelatedpapers\(<PAPER\>and</PAPER\>\)only\. \#\#\#Noveltytypes: \{novelty\_class\_descriptions\} \#\#\#Instructions: \-Usetheexamplereviewbelowtowriteareviewfortheprovidedideabycomparingittotherelatedpapers\. \-Don’tassumeanypriorknowledgeabouttheidea\. \-Makesurethegeneratedreviewfollowstheformatinexamplereviewprovidedbelow\. \-Thereviewshouldbeconcise\-around60to100words\. \#\#\#ResearchIdea: \{research\_idea\} \#\#\#RelatedPapers: \{related\_papers\} \#\#\#ExampleReview: \{example\_review\} \#\#\#OutputFormat: <REVIEW\>concisereview</REVIEW\> Thinkstepbystepbeforegeneratingthereview\! """ Figure 4:Instruction used for our[tpr](https://arxiv.org/html/2608.25660#id5)approach as introduced in Section[5](https://arxiv.org/html/2608.25660#S5)\. The instruction to reason step by step is included only for models that do not generate think tokens by default and is omitted otherwise\.
Similar Articles
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
This paper explores teaching language models to forecast the empirical success of research ideas by comparing pairs of ideas. Using a dataset of 11,488 idea pairs from PapersWithCode, the authors show that fine-tuning (SFT) boosts accuracy to 77.1%, outperforming GPT-5, and reinforcement learning with verifiable rewards achieves 71.35% with interpretable reasoning.
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers
This paper systematically evaluates human creativity tests for LLMs and finds they fail to predict scientific ideation. It introduces the DRAT, a new test that combines convergent and divergent thinking to reliably predict scientific ideation ability in language models.
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.