代码的 Token 签名:跨大型语言模型的编程行为比较

arXiv cs.CL 论文

摘要

本文提出了 CLIC,一个视觉分析系统,通过 token 频率分析来比较大型语言模型的编程行为,为 LLM 选择和提示工程提供见解。

arXiv:2609.22097v1 Announce Type: new Abstract: The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely unexplored: how do LLMs differ in their coding behavior? We propose CLIC (Code Learning for Identification and Comparison), a visual analytics approach that characterizes LLM coding behavior through token-frequency analysis. CLIC represents each code sample as a feature vector of token frequencies and trains an interpretable decision tree to separate two LLMs' code sets. Beyond classification accuracy, we define two new metrics: robustness, which measures whether the two LLMs remain distinguishable as their most-discriminative tokens are progressively removed, and concentration, which measures whether the difference is driven by a few dominant tokens or spread across many. Interpreting numerous pairwise comparisons (across LLM pairs, tasks, and tokenization levels) and tracing the full analytical chain form an inherently multi-scale, hypothesis-driven exploration task. We therefore develop an interactive visual analytics system to navigate the comparison landscape, identify pairs of interest, and drill down into discriminative tokens and their code contexts. Case studies comparing 10 LLMs across 22 Kaggle ML tasks reveal actionable insights for LLM selection and prompt engineering.
查看原文
查看缓存全文

缓存时间: 2026/09/22 09:00

# Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models
Source: [https://arxiv.org/html/2609.22097](https://arxiv.org/html/2609.22097)
\\onlineid

0\\vgtccategoryResearch\\vgtcpapertypeAnalytics & Decisions\\authorfooterJ\. Wang, Y\. Chen, M\. Pan, U\. Saini, Y\. Cai are with Visa Research\. E\-mail: \{junpenwa, yuzchen, menpan, udasaini, yicai\}@visa\.com\\teaser![The CLIC visual analytics system, composed of three coordinated views: the Metric View (matrix and scatterplot), the Decision Tree View (node-link diagram with feature importance), and the Feature Ablation View (bar and area charts showing progressive feature removal).](https://arxiv.org/html/2609.22097v1/1system.png)\\sysnameis composed of three coordinated views\. The\\metricview\(a\) alternates between a matrix and scatterplot visualization to present three metrics quantifying the separability of LLM pairs\. The\\treeview\(b\) visualizes the trained decision tree used to separate two code sets, along with the top discriminative features\. The\\ablationview\(c\) records the dynamics when progressively removing the most important features one after another and supports time travel back to an iteration of interest\.\\tl\_set:Ne\\reqboxonereqboxone\\tl\_set:Ne\\reqboxtworeqboxtwo\\tl\_set:Ne\\reqboxthreereqboxthree\\tl\_set:Ne\\sharpcascadingsharpcascading\\tl\_set:Ne\\deepdistributeddeepdistributed\\tl\_set:Ne\\fragilediffusefragilediffuse\\tl\_set:Ne\\lexicalbrittlelexicalbrittle\\tl\_set:Ne\\sumllmboxsumllmbox\\tl\_set:Ne\\simcirclesimcircle\\tl\_set:Ne\\diffcirclediffcircle

Introduction

###### Abstract

The evaluation of large language models \(LLMs\) on coding tasks has primarily focused on performance metrics such as pass@kk\. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance\-based evaluation alone\. Yet a key question remains largely unexplored:how do LLMs differ in theircoding behavior?We propose\\sysname\(Code Learning for Identification and Comparison\), a visual analytics approach that characterizes LLM coding behavior through token\-frequency analysis\.\\sysnamerepresents each code sample as a feature vector of token frequencies and trains an interpretable decision tree to separate two LLMs’ code sets\. Beyond classification accuracy, we define two new metrics:robustness, which measures whether the two LLMs remain distinguishable as their most\-discriminative tokens are progressively removed, andconcentration, which measures whether the difference is driven by a few dominant tokens or spread across many\. Interpreting numerous pairwise comparisons—across LLM pairs, tasks, and tokenization levels—and tracing the full analytical chain form an inherently multi\-scale, hypothesis\-driven exploration task\. We therefore develop an interactive visual analytics system to navigate the comparison landscape, identify pairs of interest, and drill down into discriminative tokens and their code contexts\. Case studies comparing 10 LLMs across 22 Kaggle ML tasks reveal actionable insights for LLM selection and prompt engineering\.

###### keywords

LLM Comparison, Agentic Coding, Artificial Intelligence, Visualization, Visual Analytics\.

With the rapid proliferation of LLMs, systematic comparison among them has become increasingly critical\[[23](https://arxiv.org/html/2609.22097#bib.bib4),[24](https://arxiv.org/html/2609.22097#bib.bib31),[44](https://arxiv.org/html/2609.22097#bib.bib3),[74](https://arxiv.org/html/2609.22097#bib.bib2),[62](https://arxiv.org/html/2609.22097#bib.bib72)\]\. Domain practitioners regularly ask: Which LLM adheres more closely to a prompt? Which is more cost\-efficient in real\-world use? To address these needs, numerous benchmarks\[[44](https://arxiv.org/html/2609.22097#bib.bib3),[29](https://arxiv.org/html/2609.22097#bib.bib30),[20](https://arxiv.org/html/2609.22097#bib.bib29)\]have been proposed to evaluate LLMs across many dimensions\. When it comes to LLM\-powered coding tasks, however, most existing work remains narrowly focused on functional correctness or execution performance\[[8](https://arxiv.org/html/2609.22097#bib.bib1),[55](https://arxiv.org/html/2609.22097#bib.bib35),[21](https://arxiv.org/html/2609.22097#bib.bib34)\]\. For example, MLE\-Bench\[[8](https://arxiv.org/html/2609.22097#bib.bib1)\]ranks LLMs and agentic coding frameworks by a performance metric, pass@kk\[[9](https://arxiv.org/html/2609.22097#bib.bib33)\], on its leaderboard\[[39](https://arxiv.org/html/2609.22097#bib.bib36)\]\.

As LLM capabilities continue to advance, many models now meet baseline performance requirements on standard coding tasks, further reducing the discriminative power ofperformance\-basedevaluations\. Yet an equally important question remains largely unexplored:how do LLMs differ in their coding behavior?For instance, among LLMs with similar performance, does one consistently produce more verbose code than another? Does one exhibit greater caution in file operations? Suchbehavior\-baseddifferences, while often overlooked, can meaningfully affect code reliability, maintainability, and the quality of developer interaction\. Understanding them is essential for accurately characterizing LLMs, guiding prompt engineering, and selecting LLMs\. To make this concrete, consider two scenarios\.First, when iterating on an agentic coding framework—e\.g\., refining a prompt or adding a memory module\[[75](https://arxiv.org/html/2609.22097#bib.bib69)\]—developers need to know how the resulting code*differs*from before, since the modification rarely manifests as a single scalar improvement and may shift the model’s behavior in unintended ways\.Second, when multiple LLMs achieve comparable performance on a target task, the choice between them hinges on which model’s coding behavior aligns best with the team’s conventions, downstream tooling, and reliability needs\. Both scenarios demand a principled, interpretable comparison of*how*two LLMs code, not merely*whether*they pass tests\. They are also the daily concerns of our target users:ML scientists and engineerswho use LLMs for code generation\.

To fill this research gap, we proposeCodeLearning forIdentification andComparison \(\\sysname\)\.\\sysnamecompares LLMs’ coding behavior by analyzing token frequencies in LLM\-generated code\. Concretely, it represents each code sample as a feature vector of token frequencies \(computed at five distinct semantic levels\)\. Two collections of code from two LLMs thus become two sets of feature vectors\. We train a binary decision tree to separate them, and itsaccuracyquantifies how distinguishable the two LLMs’ coding behaviors are\. The cornerstone of this pipeline is the token\-frequency representation itself\. Whiletoken frequencymay at first appear to be a low\-level lexical property, in our setting, it is in fact a behaviorally rich, well\-validated representation that surfaces concrete coding behaviors\. We focus on it for four reasons\.First, tokens are the atomic units produced by LLMs, and their frequencies directly reflect a model’s generation preferences\.Second, many code tokens carry clear semantics, and their frequency differences correspond to distinct coding behaviors, e\.g\., a high occurrence of theprinttoken reflects*verbose*code\.Third, token\-frequency analysis is the canonical feature representation of NLP authorship attribution and code stylometry\[[7](https://arxiv.org/html/2609.22097#bib.bib75),[6](https://arxiv.org/html/2609.22097#bib.bib71)\], repeatedly shown to reliably distinguish source\-code authors\. This long\-validated stylometric tradition grounds our use of token frequency as a meaningful lens on coding behavior\.Fourth, token\-frequency analysis is intuitive and easy to communicate, enabling clear and interpretable comparisons across LLMs\.

The decision treeaccuracyalone, however, does not fully characterize coding behavior differences\. Two LLM pairs with the sameaccuracycan differ substantially in how that separability is structured\. One pair’s separability may reside entirely in a single token: removing it causes theaccuracyto collapse to∼50%\{\\sim\}50\\%\(random guessing\)\. The other pair’s separability may persist even after the top token is removed, as a new token rises to take its place\. This contrast motivates ourrobustnessmetric, which measures whether two LLMs remain distinguishable as their most\-discriminative tokens are progressively removed\. A separate question concerns how the discriminative signal is*distributed*at each step: is it driven by a single dominant token, or spread across many? This motivates ourconcentrationmetric, which measures the skewness of the token\-importance distribution—high when a few tokens dominate, low when many tokens contribute roughly equally\.

While the three metrics quantify behavioral differences, applying them at scale—across many LLM pairs, coding tasks, and tokenization levels—yields a multi\-granular, hypothesis\-driven workflow: an analyst first explores the comparison landscape to identify pairs of interest, then drills into the discriminative tokens, and finally examines code samples to validate a hypothesis\. Findings at one level raise new questions at the next, demanding coordinated, interactive views rather than any single static representation\[[11](https://arxiv.org/html/2609.22097#bib.bib67),[46](https://arxiv.org/html/2609.22097#bib.bib66)\]\. We therefore develop a visual analytics system \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\) to support this exploration\. The system presents metrics across LLM pairs through matrix and scatterplot visualizations \(\\metricview\), exposes the trained decision tree via a node\-link diagram \(\\treeview\), and illustrates feature\-importance dynamics through coordinated bar and area charts \(\\ablationview\)\.

We conduct multiple case studies comparing 10 LLMs on Python code generated for 22 Kaggle ML competitions, showing how\\sysnamereveals behavioral differences that inform LLM selection and steer prompt engineering\. Although we focus on Python code and ML tasks, the proposed methodology is task\-agnostic and can be directly extended to other programming languages, a direction we leave for future work\.

To summarize, our contributions are threefold:

1. 1\.We propose five levels of code tokenization and a learning\-to\-compare algorithm to compare LLMs’ code at each level\.
2. 2\.We introduce two new metrics,robustnessandconcentration, to thoroughly characterize the separability between two LLMs\.
3. 3\.We develop a visual analytics system to explore the LLM comparison results and steer the drill\-down exploration process\.

## 1Related Work

LLM Code Generation and Benchmark\.LLMs have shown remarkable capability on coding tasks\[[13](https://arxiv.org/html/2609.22097#bib.bib19)\], giving rise to a rich ecosystem of coding frameworks\[[67](https://arxiv.org/html/2609.22097#bib.bib45),[42](https://arxiv.org/html/2609.22097#bib.bib21),[72](https://arxiv.org/html/2609.22097#bib.bib52),[52](https://arxiv.org/html/2609.22097#bib.bib53),[35](https://arxiv.org/html/2609.22097#bib.bib5)\]and benchmarks\[[9](https://arxiv.org/html/2609.22097#bib.bib33),[22](https://arxiv.org/html/2609.22097#bib.bib50),[8](https://arxiv.org/html/2609.22097#bib.bib1),[69](https://arxiv.org/html/2609.22097#bib.bib26)\]to evaluate them\. Early work introduced dedicated code LLMs such as Codex\[[9](https://arxiv.org/html/2609.22097#bib.bib33)\]and CodeT5\[[66](https://arxiv.org/html/2609.22097#bib.bib28)\], and tools like GitHub Copilot\[[15](https://arxiv.org/html/2609.22097#bib.bib38)\]brought LLM\-powered code generation to everyday practice\. Building on these foundations, more recent agentic frameworks automate multi\-step coding workflows\. For example, OpenHands\[[64](https://arxiv.org/html/2609.22097#bib.bib6)\]introduces a chain\-based iterative solution search\. AIDE\[[67](https://arxiv.org/html/2609.22097#bib.bib45)\]extends it to a tree\-based search, which enables the agent to alternately explore multiple solution branches\. R&D\-Agent\[[72](https://arxiv.org/html/2609.22097#bib.bib52)\]further introduces a graph\-based solution search, enabling solution fusion to leverage knowledge gained across branches\. For evaluation, HumanEval\[[9](https://arxiv.org/html/2609.22097#bib.bib33)\]introduces 164 hand\-written programming problems and the pass@kkmetric, establishing the dominant evaluation paradigm of functional correctness\. SWE\-bench\[[22](https://arxiv.org/html/2609.22097#bib.bib50)\]extends this to real\-world software engineering by testing LLMs on resolving GitHub issues, while MLE\-Bench\[[8](https://arxiv.org/html/2609.22097#bib.bib1)\], which we use in this work, evaluates agentic LLM coding on Kaggle ML competitions\. A comprehensive survey by Jiang et al\.\[[21](https://arxiv.org/html/2609.22097#bib.bib34)\]catalogs the broader landscape of evaluation across code generation, completion, and repair\. Despite this breadth, these benchmarks\[[4](https://arxiv.org/html/2609.22097#bib.bib18),[8](https://arxiv.org/html/2609.22097#bib.bib1),[9](https://arxiv.org/html/2609.22097#bib.bib33),[22](https://arxiv.org/html/2609.22097#bib.bib50)\]and frameworks\[[67](https://arxiv.org/html/2609.22097#bib.bib45),[72](https://arxiv.org/html/2609.22097#bib.bib52),[14](https://arxiv.org/html/2609.22097#bib.bib17)\]share a common focus: whether generated code executes correctly\.How LLMs differ in the way they write code—their stylistic and behavioral tendencies—remains largely unexamined\.

LLM Visual Comparison\.The growing adoption of LLMs has motivated a body of work on visually comparing their behaviors\[[10](https://arxiv.org/html/2609.22097#bib.bib16),[48](https://arxiv.org/html/2609.22097#bib.bib59),[43](https://arxiv.org/html/2609.22097#bib.bib60),[65](https://arxiv.org/html/2609.22097#bib.bib68),[12](https://arxiv.org/html/2609.22097#bib.bib65)\]\. LLM Comparator\[[23](https://arxiv.org/html/2609.22097#bib.bib4),[24](https://arxiv.org/html/2609.22097#bib.bib31)\]juxtaposes LLM responses to help analysts discover cases where one model performs better than another\. VisEval\[[10](https://arxiv.org/html/2609.22097#bib.bib16)\]evaluates and compares multiple LLMs on the task of translating natural\-language queries over tabular data into executable visualization specifications, revealing common failure patterns such as incorrect data mapping and poor design choices\. EvalLM\[[26](https://arxiv.org/html/2609.22097#bib.bib11)\]lets users define custom criteria and leverage LLMs to evaluate model outputs, supporting iterative prompt refinement\. ChainForge\[[3](https://arxiv.org/html/2609.22097#bib.bib10)\]enables cross\-LLM prompt comparison through an interactive interface, organizing responses in structured views to support hypothesis testing\.While these systems focus on comparing LLM capabilities in language or image generation, none focuses on comparing their coding capabilities\.Most recently, Wang et al\.\[[55](https://arxiv.org/html/2609.22097#bib.bib35)\]developed a visual analytics system to study and improve LLM\-powered coding agents, exposing agent behavior through tree\-structured exploration and code examination\. Their work focuses on distinguishing LLMs in their iterative coding strategy exploration process\.Our work instead focuses on coding style differences across LLMs, particularly their token usage patterns\.

VIS4AI and Comparative Analysis\.Leveraging visualization to make AI models more interpretable \(VIS4AI\[[59](https://arxiv.org/html/2609.22097#bib.bib27),[34](https://arxiv.org/html/2609.22097#bib.bib51),[19](https://arxiv.org/html/2609.22097#bib.bib43)\]\) has been a prevalent research direction\. Many works have demonstrated that visualization can effectively help humans better understand\[[57](https://arxiv.org/html/2609.22097#bib.bib42),[32](https://arxiv.org/html/2609.22097#bib.bib8),[31](https://arxiv.org/html/2609.22097#bib.bib13)\], diagnose\[[45](https://arxiv.org/html/2609.22097#bib.bib7),[58](https://arxiv.org/html/2609.22097#bib.bib9),[68](https://arxiv.org/html/2609.22097#bib.bib41),[28](https://arxiv.org/html/2609.22097#bib.bib32),[51](https://arxiv.org/html/2609.22097#bib.bib14)\], improve\[[18](https://arxiv.org/html/2609.22097#bib.bib57),[5](https://arxiv.org/html/2609.22097#bib.bib12),[56](https://arxiv.org/html/2609.22097#bib.bib39)\], and steer\[[71](https://arxiv.org/html/2609.22097#bib.bib25),[38](https://arxiv.org/html/2609.22097#bib.bib40),[30](https://arxiv.org/html/2609.22097#bib.bib22)\]AI models\. Among these works, two groups are closely related to ours\. The first is interpretable tree visualization\. Decision trees are inherently interpretable and there are many visualizations designed for them\[[41](https://arxiv.org/html/2609.22097#bib.bib63),[27](https://arxiv.org/html/2609.22097#bib.bib62)\]\. For example, BaobabView\[[54](https://arxiv.org/html/2609.22097#bib.bib24)\]is a visual analytics system for interactively constructing, editing, and analyzing decision trees\. It uses a flow\-based tree visualization where link width and color encode how data instances of different classes move through the splits\. Later works adapt node\-link diagrams to visualize decision trees and facilitate the analysis of an ensemble of trees\[[33](https://arxiv.org/html/2609.22097#bib.bib23),[61](https://arxiv.org/html/2609.22097#bib.bib15),[49](https://arxiv.org/html/2609.22097#bib.bib61)\]\.We employ a similar visualization to present the decision tree for its intuitive clarity\.The second related VIS4AI direction is visual comparative analysis\[[16](https://arxiv.org/html/2609.22097#bib.bib64),[17](https://arxiv.org/html/2609.22097#bib.bib58)\]\. The most closely related to our work is learning\-from\-disagreement \(LFD\)\[[60](https://arxiv.org/html/2609.22097#bib.bib47),[63](https://arxiv.org/html/2609.22097#bib.bib46)\], which trains a classifier to distinguish instances on which two ML models disagree, and probes important features to characterize the models’ relative strengths and weaknesses\.Similar to LFD, our work also uses a binary classifier to distinguish two LLMs’ coding behavior, but through the lens of their token usage\. Furthermore, we introduce two new metrics,robustnessandconcentration, built on the trained classifier to more comprehensively quantify the separability\.

## 2Requirement Analysis

\\sysname

was developed through a user\-centered design process in close collaboration with three ML scientists \(target users\)\. All hold Ph\.D\. degrees in computer science and have 3–5 years of full\-time industry experience\. Their primary responsibilities involve building and deploying ML models in the financial industry\. In their work, they extensively use LLMs for code generation and are well\-versed in state\-of\-the\-art \(SOTA\) agentic coding frameworks\. Over two months, we conducted weekly meetings to iteratively elicit their pain points in LLM selection and generated\-code comparison\. Alongside this user\-centered process, our requirement analysis was also actively informed by three threads of prior literature, each directly shaping one core aspect of the final requirements\.First, classifier\-based two\-sample testing\[[36](https://arxiv.org/html/2609.22097#bib.bib70)\]and recent LLM\-code stylometry\[[6](https://arxiv.org/html/2609.22097#bib.bib71)\]motivated us to formulate requirements for measuring code separability with*quantifiable metrics*\.Second, corpus\-contrastive analyses\[[40](https://arxiv.org/html/2609.22097#bib.bib73),[25](https://arxiv.org/html/2609.22097#bib.bib74)\]drove us to formulate requirements for attributing the observed separability back to*individual code tokens*and their usage contexts\.Third, code\-stylometry research\[[7](https://arxiv.org/html/2609.22097#bib.bib75)\]led us to formulate requirements for decomposing code into distinct*semantic levels*and analyzing each level separately\. Synthesizing the experts’ pain points with these literature insights and broader prior work on LLM code generation\[[13](https://arxiv.org/html/2609.22097#bib.bib19),[53](https://arxiv.org/html/2609.22097#bib.bib20),[55](https://arxiv.org/html/2609.22097#bib.bib35)\], we identified three key requirements\.

- \\reqboxone R1 Quantifying Code Separability: To characterize how LLMs differ in their generated code, the experts preferred quantitative metrics to holistically measure the*separability*between two code sets, i\.e\., how easily the originating LLM can be inferred from the code and whether a few features dominate the separability\. Moreover, since practitioners can easily modify LLM prompts to regenerate code, it is necessary to assess the robustness of this separability\. Thus, we need: - ⋄\\diamond\\reqboxone R1\.1 An intuitive metric quantifying theseparabilitybetween two code sets, reflecting how distinguishable the underlying LLMs are\. - ⋄\\diamond\\reqboxone R1\.2 A metric that captures therobustnessof the observed separability under progressive ablation of the key discriminative features\. - ⋄\\diamond\\reqboxone R1\.3 A metric that characterizes theskewnessof the discriminative feature distribution, i\.e\., how the separability is structured internally\.
- \\reqboxtwo R2 Identifying Key Discriminative Features: Beyond global separability measures, it is necessary to pinpoint the specific features \(code tokens\) that drive the observed differences\. Understanding which tokens are more discriminative and how they collectively contribute to the separability enables more targeted LLM characterization and informed prompt engineering\. This requires\\sysnameto: - ⋄\\diamond\\reqboxtwo R2\.1 Identify the features that contribute the most to the separation of two code sets and the code contexts where the features were used\. - ⋄\\diamond\\reqboxtwo R2\.2 Reveal the distributional pattern of discriminative features, characterizing whether the separability is driven by a dominant feature or emerges from the joint effect of multiple features\. - ⋄\\diamond\\reqboxtwo R2\.3 Support what\-if analysis\[[68](https://arxiv.org/html/2609.22097#bib.bib41)\]by excluding the most important features, uncovering features of equivalent discriminative power masked by more dominant ones\.
- \\reqboxthree R3 Multi\-Level Code Analysis: A piece of code is composed of distinct semantic levels, e\.g\., comments, functional code, and imported packages\. Each reflects a unique aspect of LLM coding behavior, and the separability may disappear when one level is peeled off\. It is therefore necessary to attribute observed differences to specific levels for fine\-grained characterization\. For instance, divergences may primarily reside in code comments, reflecting differences in natural language capability\. One LLM may consistently favor tree\-based approaches while another defaults to neural network solutions, reflecting differences in model preference\. To reveal this,\\sysnameneeds to: - ⋄\\diamond\\reqboxthree R3\.1 Decompose code into distinct semantic levels and quantify the separability and key discriminative features at each level\. - ⋄\\diamond\\reqboxthree R3\.2 Compare separability across levels to identify which aspects of code generation most distinguish the two LLMs\.

## 3Code Generation

Coding Tasks\.We investigate the differences in LLMs’ coding behaviors by comparing their generated code\. To ground the comparison in a concrete and reproducible setting, we use Python code generation for ML tasks as our study domain\. Rather than relying on a single task, we evaluate on 22 ML tasks \(22 Kaggle competitions from MLE\-Bench\[[8](https://arxiv.org/html/2609.22097#bib.bib1)\]\) spanning distinct data modalities \(tabular, image, text, and audio\) and a range of difficulty levels\. This internal diversity exposes\\sysnameto qualitatively different problems within a single, well\-validated benchmark\[[8](https://arxiv.org/html/2609.22097#bib.bib1)\], and the per\-task analysis \(in Appendix\) shows that the discriminative tokens shift task\-by\-task rather than collapsing into a single task\-specific signal\. We emphasize that this is a methodological choice:\\sysname’s comparison mechanism makes no assumptions specific to ML tasks, and it applies directly to other software\-development scenarios, which we discuss as future work in Sec\.[8](https://arxiv.org/html/2609.22097#S8)\.

Coding Agents\.To isolate differences attributable to the LLM rather than the agentic coding frameworks, we use a single coding agent throughout this work, varying only the backend LLM\. Specifically, we choose AI\-Driven Exploration \(AIDE\)\[[67](https://arxiv.org/html/2609.22097#bib.bib45),[1](https://arxiv.org/html/2609.22097#bib.bib37)\], as it is open\-source and achieves SOTA performance\. It takes natural language descriptions of an ML problem as input and generates code to iteratively solve it\. Within each iteration, the agent candraftan initial solution,debugexisting buggy code, orimprovefunctional code\. Since the latter two depend on code from prior iterations and thus introduce variability beyond the LLM itself, we only use thedraftoperation to ensure all LLMs generate code under identical conditions\.

LLM Settings\.Our work uses 10 LLMs:o3,gpt\-4o,gpt\-4\.1,gpt\-4\.1\-mini,gpt\-5,gpt\-5\-mini,gpt\-5\.1,gpt\-5\.2,grok\-4\(grokfor short\), andclaude\-sonnet\-4\.5\(claudefor short\)111All trademarks are the property of their respective owners, are used for identification purposes only, and do not necessarily imply product endorsement or affiliation with Visa\.\. They are chosen along two complementary axes that maximize the comparative leverage of\\sysname—\(i\)*diversity across providers*\(OpenAI, xAI, and Anthropic\), exposing stylistic differences attributable to different model families and training regimes, and \(ii\)*similarity within providers*\(e\.g\.,gpt\-5vs\.gpt\-5\-mini\), exposing fine\-grained behavioral differences between closely related models\. Given the rapid pace at which frontier LLMs are released, exhaustive coverage in any single study is intrinsically infeasible\. The contribution of\\sysnameis methodological rather than a benchmark of LLM coverage: the comparison mechanism is agnostic to the specific LLMs being compared and can be readily transferred to any pair of code corpora\. For each LLM, AIDE attempts to solve the 22 Kaggle competitions1,0001,000times from scratch withidenticalprompts, generating1,0001,000code samples per task\. We call this a single LLMrun\. We perform two runs per LLM to verify behavioral consistency: if an LLM’s coding style is stable, the two runs should be hard to distinguish from each other\. By default, we set the temperature to 1 for all LLMs\. In total, our experiment generates10​\(L​L​M​s\)×22​\(t​a​s​k​s\)×2​\(r​u​n​s\)×1,000​\(c​o​d​e\)=440,00010\\ \(LLMs\)\{\\times\}22\\ \(tasks\)\{\\times\}2\\ \(runs\)\{\\times\}1,000\\ \(code\)\{=\}440\{,\}000code samples, which form the corpus analyzed later in Sec\.[6\.1](https://arxiv.org/html/2609.22097#S6.SS1)\.

## 4Code Comparison with a Decision Tree

Our approach represents each code sample as a token\-frequency vector and trains an interpretable binary classifier to separate two LLMs’ code sets, yielding both*a separability score*and*the most discriminative tokens*\. Concretely, given2×1,0002\{\\times\}1,000code samples from two LLM runs, we tokenize each sample222Breaking down the code into a set of atomic words\. The termsword,token, andfeatureare used interchangeably\. Note that our tokenizer differs from LLM tokenizers, which may split a single word into multiple sub\-word tokens\.and compute the frequencies of the top\-500 tokens \(empirically determined\), representing each sample as a 500\-dimensional feature vector\. This high\-dimensional representation can be analyzed by many paradigms, e\.g\., dimensionality reduction, embedding\-based comparison, or statistical feature analysis\. We analyze it with an interpretable discriminative classifier, because it is superior to others in directly returning the discriminative tokens that drive the separation between two code sets \(\\reqboxtwoR2\)\. Within this paradigm, we choose to use decision tree models for three deliberate reasons:

1. 1\.Faithful, glass\-box interpretability: unlike post\-hoc surrogates, a decision tree is itself the model—every prediction can be traced exactly to a sequence of token\-frequency conditions, which is essential for attributing differences to specific lexical evidence\.
2. 2\.Established precedent for analogous tasks: decision trees have been repeatedly adopted as the interpretable backbone for visual analytics systems\[[60](https://arxiv.org/html/2609.22097#bib.bib47),[37](https://arxiv.org/html/2609.22097#bib.bib48),[73](https://arxiv.org/html/2609.22097#bib.bib49)\], validating their suitability for the contrastive, token\-level interpretation our work requires\.
3. 3\.Alignment with target users: our experts \(consulted during the requirement analysis stage\) explicitly indicated familiarity and trust in tree\-based reasoning based on their experiments, reducing the cognitive overhead of adopting\\sysnamein real workflows\.

We emphasize, however, that this choice is intentionally*modular*: the binary classifier requires only a model whose decisions can be locally attributed to input features, so any inherently interpretable classifier \(e\.g\., linear models or generalized additive models\) can be substituted\. We leave a full comparison across these alternatives as future work\.

We train a decision tree on the2,0002\{,\}000vectors using an 80/20 stratified train/test split \(see pseudocode in the Appendix\)\. All reportedaccuracyvalues in this paper are evaluated on the held\-out test set, ensuring the reported separability reflects generalization without overfitting\.

The trained decision tree yields two key outputs: itsaccuracyquantifies how distinguishable the two LLMs’ coding behaviors are, and itstoken importance scores\(the tokens’ cumulative Gini impurity decrease across all tree nodes\) reveal which tokens drive the separation\.

### 4\.1Three Metrics Quantifying Separability

We introduce three complementary metrics—accuracy,robustness, andconcentration—each capturing a different dimension of separability\.

Accuracy\.The classifier’s test accuracy intuitively reflects the separability\.A higher accuracyindicates that the two code sets are more distinct and easily separable, whilean accuracy near 50%suggests the code sets are similar and difficult to distinguish \(\\reqboxoneR1\.1\)\. We choose accuracy because it aligns naturally with our balanced classifier setting, in which the two compared code sets contribute a comparable number of instances and play*symmetric*roles—neither is treated as a privileged “positive” class\. Metrics such as precision, recall, and F1\-score additionally encode*which*class is mistaken for which, but this asymmetric information is not meaningful for our goal of quantifying separability between two interchangeable code sets\. Accuracy is therefore selected not because it is universally superior, but because it provides the most parsimonious and interpretable aggregate\. This choice is also*modular*within\\sysnameand can be substituted by other metrics when necessary\.

Robustness\.The accuracy metric surfaces only the initial separability, but not its depth\. Two code sets may differ only in the frequency of a single token, and on this basis alone, the classifier can reach 100% accuracy\. However, removing this token and retraining the classifier withn−1n\{\-\}1features would cause a dramatic accuracy drop, indicating that the separability is fragile\. Our robustness metric is designed to quantify how resilient the differences between two code sets are \(\\reqboxoneR1\.2\), i\.e\., we iteratively remove the most important feature, retrain the classifier, and measure its accuracy\. The metric is initially defined as the number of features that can be removed before the accuracy drops below a threshold, e\.g\., 60%\. However, this threshold would significantly impact the robustness value\. Through iterative refinement, we instead define robustness asthe area under the accuracy curveobtained by progressively removing the most discriminative features\. Since our classifier is binary, an accuracy ofxx% is equivalent to\(100−x\)\(100\{\-\}x\)%\. Therefore, we accumulate the absolute deviation from the 50% accuracy baseline:

Robustness=∑s=1S\|accuracy​\(s\)−0\.5\|,\\vskip\-4\.0pt\\text\{Robustness\}=\\sum\_\{s=1\}^\{S\}\\bigl\|\\,\\text\{accuracy\}\(s\)\-0\.5\\,\\bigr\|,\\vskip\-3\.0pt\(1\)whereaccuracy​\(s\)\\text\{accuracy\}\(s\)is the classifier accuracy after thess\-th most important feature has been removed andSSis the total number of feature removal steps \(S≤500S\\leq 500\)\. Robustness measures the*persistence*of separability, not its initial level\.A large robustnessmeans the discriminative signal is present in many features, and the classifier stays meaningfully above random guessing across many removal steps\.A small robustnessmeans the signal is fragile—this can arise either because accuracy was always near 50% \(the two code sets were never separable\), or because accuracy was initially high but collapsed after removing a few features\.

Concentration\.While robustness captures how long the classifier remains accurate, it says nothing about how the discriminative signal is structured within the remaining tokens \(\\reqboxoneR1\.3\)\. Two code sets could be separable due to a single dominant token or the joint effect of multiple tokens\. We address this with a complementary metric derived from the entropy \(H⁡\(s\)H\(s\)\) of the token importance distribution\. At each removal stepss, we compute the entropy of the top\-15 \(empirically determined\) token importance values:

H\(s\)=−∑i=115pi\(s\)log2pi\(s\),pi\(s\)=wi\(s\)/∑j=115wj\(s\),\\vskip\-4\.0ptH\(s\)=\-\\sum\_\{i=1\}^\{15\}p\_\{i\}\(s\)\\,\\log\_\{2\}p\_\{i\}\(s\),\\quad p\_\{i\}\(s\)=\{w\_\{i\}\(s\)\}/\{\\displaystyle\\sum\_\{j=1\}^\{15\}w\_\{j\}\(s\)\},\\vskip\-3\.0pt\(2\)wherewi​\(s\)w\_\{i\}\(s\)is the importance of theii\-th ranked token at stepss\.H⁡\(s\)H\(s\)ranges from00\(all importance concentrated in a single token\) tolog2⁡15≈3\.91\\log\_\{2\}15\\approx 3\.91bits \(uniform importance across the 15 tokens\)\. We define concentration as the integral of the gap between the maximum\-entropy ceiling and the observed entropy curve over removal steps, i\.e\.,the area above the entropy curve:

Concentration=∑s=1S/2\(log2⁡15−H⁡\(s\)\)\.\\vskip\-4\.0pt\\text\{Concentration\}=\\sum\_\{s=1\}^\{S/2\}\\bigl\(\\log\_\{2\}15\-H\(s\)\\bigr\)\.\(3\)Note that the accumulation stops at the first half of steps \(S/2S/2\)\. This is because, towards the end of feature removal, the remaining features are reduced to just a few and become dominant by default, as no other features remain\. The entropy curve and its drop towards the end of feature removal are shown in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c2 and explained in Sec\.[5\.3](https://arxiv.org/html/2609.22097#S5.SS3)\.

A high concentrationindicates that at every removal step, a few tokens always dominate the remaining importance distribution\. After the most important token is stripped away, another rises\. The two LLMs are therefore distinguished by a*cascade of dominant tokens*—their coding styles differ through a sequence of clear, individually decisive signals\.A low concentrationindicates that after each removal, the remaining importance distribution is relatively flat\. The two LLMs differ in a*diffuse way*—no token jumps out as individually decisive\.

All three metrics are needed for a complete picture\. Accuracy captures the*initial level*of separability; robustness captures its*persistence*across feature removals; concentration captures the*structure*of the discriminative signal across steps\. Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1)characterizes the four regimes defined by robustness and concentration, and should be read in conjunction with accuracy to determine whether low robustness reflects genuine low separability or merely a fragile high\-accuracy classifier\.

Figure 1:The four regimes defined by robustness and concentration\.
### 4\.2Five Code Tokenizers for Multi\-Level Analysis

Code consists of different semantic levels—comments, functional code, and package usages—each reflecting a distinct aspect of LLM coding behavior\.Analyzing separability at each level independently allows observed differences to be attributed to specific parts of the code \(\\reqboxthreeR3\)\.We introduce five tokenizers, each targeting a different level\.

Thecompletetokenizer extracts all tokens from the entire code without filtering, capturing the full vocabulary including code identifiers, comments, and string literals\. It uses a simple regex pattern \(\\b\[a\-zA\-Z\_\]\[a\-zA\-Z0\-9\_\]\*\\b\) applied directly to the code\. This tokenizer provides a comprehensive view of an LLM’s vocabulary usage and encompasses all tokens extracted by the other four tokenizers\.

Thecommenttokenizer tokenizes only the inline comments \(starting with\#\) and docstrings \(triple\-quoted strings\)\. It uses Python’stokenizemodule to identify inline comments and abstract syntax tree \(AST\)\[[2](https://arxiv.org/html/2609.22097#bib.bib44)\]parsing to extract docstrings from modules, functions, and classes\. This tokenizer focuses on an LLM’s documentation and explanation style, revealing how LLMs communicate intent and provide context through natural language rather than code\.

Thecodetokenizer is complementary to thecommenttokenizer, extracting all tokens from code while excluding comments\. This includes code identifiers \(variable/function/class names\) and string literals within the code \(such as those inprintstatements\)\. It uses the same tools as thecommenttokenizer to accurately identify and remove comments/docstrings before applying regex extraction\. This tokenizer captures an LLM’s programming vocabulary and naming conventions\.

Thefunctionaltokenizer extracts only identifiers that contribute to the functional logic of the code, producing a strict subset of thecodetokenizer’s output\. It uses a custom AST visitor to traverse the syntax tree and collect identifiers from functional constructs \(e\.g\., variables, functions, classes, etc\.\) while explicitly skipping non\-functional code through predefined rules, such asprintstatements,loggingcalls, and display output \(e\.g\.,tqdm\)\. This tokenizer isolates an LLM’s core algorithmic and computational vocabulary from auxiliary code\.

Thepackagetokenizer extracts only the base names of imported packages\. It parses import statements \(import Xorfrom X import Y\) using AST to identify module names, keeping only the base module name \(e\.g\.,sklearnfromsklearn\.model\_selection\)\. It reveals an LLM’s library usage patterns and ecosystem preferences\.

Our Appendix includes more details on the implementation of the five tokenizers and an example tokenization result from a code snippet\.

## 5Visual Analytics System

\\sysname

is a visual analytics system comprising three coordinated views that together support a top\-down, hypothesis\-driven exploration \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\): \(a\) the\\metricviewprovides an overview of pairwise LLM comparisons across tasks and tokenizers, enabling users to identify LLM pairs of interest; \(b\) the\\treeviewvisualizes the trained classifier and supports interactive exploration of discriminative tokens through code examples; \(c\) the\\ablationviewenables progressive feature removal to assess the robustness and concentration of observed differences\. Tight view coordination allows users to navigate fluidly from aggregate metrics down to individual tokens and their code contexts\.

### 5\.1\\metricview: Overview of LLM Comparisons

Design Rationale\.The\\metricviewdirectly exposes quantitative comparisons across LLM pairs \(\\reqboxoneR1\)\. Since we conduct exhaustive pairwise comparisons, a symmetric matrix—where rows and columns represent LLM runs and cell color encodes the metric—is a natural choice\. However, the matrix cannot reveal the joint structure of robustness and concentration illustrated in Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1)\. We therefore complement it with a scatterplot that plots robustness against concentration, characterizinghowthe separability between different LLM runs is structured\.

Matrix Mode\.The matrix view \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsa\) employs a symmetric matrix layout where both rows and columns represent LLM runs with a naming convention:LLM\_T\(emperature\)\_R\(un\), e\.g\.,grok\_T1\_R1means the first run of thegrokmodel using temperature 1\. Each cell encodes a metric value through color intensity using a sequential color scale, where darker shades indicate higher values\. Since the matrix is symmetric, the upper\-right \(\\mathord\{\\text\{\\hbox to0pt\{\\vbox to0pt\{\\pgfpicture\\makeatletter\\hbox\{\\hskip 0\.0pt\\lower 0\.0pt\\hbox to0\.0pt\{\\lxSVG@begingroup@\{\_scopebegin\} \\lxSVG@begingroup@\{stroke\} \\lxSVG@begingroup@\{fill\} \\lxSVG@setlinewidth\{\\the\\pgflinewidth\}\\lxSVG@begingroup@\{stroke\-width\} \\lx@inpgf@ignorespaces\\nullfont\\hbox to0\.0pt\{\\lxSVG@begingroup@\{\_scopebegin\} \{\{\\lx@inpgf@ignorespaces\}\{\{\}\}\{\} \{\\lx@inpgf@ignorespaces\}\{\} \{\\lx@inpgf@ignorespaces\}\{\} \{\\lx@inpgf@ignorespaces\}\\lxSVG@begingroup@\{\_scopebegin\} \\color\[rgb\]\{0\.5,0\.5,0\.5\}\\lxSVG@fill\\lxSVG@drawpath@unclipped\{M 0 0 L 0 0 L 0 0 Z\}\{stroke:none\} \\lx@inpgf@ignorespaces \\lxSVG@closescope \} \\lxSVG@closescope \{\\lx@inpgf@ignorespaces\}\{\\lx@inpgf@ignorespaces\}\{\\lx@inpgf@ignorespaces\}\\hss\}\\lxSVG@discardpath\\lxSVG@closescope \\hss\}\}\\lxSVG@closescope\\endpgfpicture\}\}\}\}\) and lower\-left \(\\mathord\{\\text\{\\hbox to0pt\{\\vbox to0pt\{\\pgfpicture\\makeatletter\\hbox\{\\hskip 0\.0pt\\lower 0\.0pt\\hbox to0\.0pt\{\\lxSVG@begingroup@\{\_scopebegin\} \\lxSVG@begingroup@\{stroke\} \\lxSVG@begingroup@\{fill\} \\lxSVG@setlinewidth\{\\the\\pgflinewidth\}\\lxSVG@begingroup@\{stroke\-width\} \\lx@inpgf@ignorespaces\\nullfont\\hbox to0\.0pt\{\\lxSVG@begingroup@\{\_scopebegin\} \{\{\\lx@inpgf@ignorespaces\}\{\{\}\}\{\} \{\\lx@inpgf@ignorespaces\}\{\} \{\\lx@inpgf@ignorespaces\}\{\} \{\\lx@inpgf@ignorespaces\}\\lxSVG@begingroup@\{\_scopebegin\} \\color\[rgb\]\{0\.5,0\.5,0\.5\}\\lxSVG@fill\\lxSVG@drawpath@unclipped\{M 0 0 L 0 0 L 0 0 Z\}\{stroke:none\} \\lx@inpgf@ignorespaces \\lxSVG@closescope \} \\lxSVG@closescope \{\\lx@inpgf@ignorespaces\}\{\\lx@inpgf@ignorespaces\}\{\\lx@inpgf@ignorespaces\}\\hss\}\\lxSVG@discardpath\\lxSVG@closescope \\hss\}\}\\lxSVG@closescope\\endpgfpicture\}\}\}\}\) triangles can encode different metrics, selectable via control widgets at the top left\. This view presents metrics for one tokenizer at a time; users can switch among the five tokenizers via a dropdown widget in the view header\.

![Refer to caption](https://arxiv.org/html/2609.22097v1/3scatterplot.png)Figure 2:The scatterplot realizes the design in Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1): \(a\) raw values; \(b\) normalized values\. Each point represents a pair of LLM runs\. Comparing 20 LLM runs across 5 tokenizers results in 950 \(20×19/2×520\{\\times\}19/2\{\\times\}5\) pairs/points\.Scatterplot Mode\.Each point in the scatterplot \(Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)a\) represents a pair of LLM runs \(i\.e\., a cell in the matrix mode\)\. The x\- and y\-axes encode robustness and concentration, respectively\. This mode presents five tokenization levels concurrently, and the points are color\-coded by their tokenizer type\. When all tokenizers are shown, metric values vary considerably across tokenizers\. We therefore enable the points to be normalized within each set \(Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)b\), so that a pair’s separability can be compared across tokenizers relative to other pairs \(\\reqboxthreeR3\.2\)\. For example, in Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)\-b1, the separability betweengpt\-4o\_T1\_R1andgpt\-5\.2\_T1\_R1is always “\\sharpcascadingSharp & Cascading” \(Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1), top\-right\) across all five tokenizers\. Dashed lines are drawn to connect points representing tokenizers of the same LLM pair for easy identification\.

Interactions\.Users can toggle between the matrix and scatterplot from the view title \(see associated video\)\. Both modes support cell or point selection through clicking, which triggers coordinated updates across views, i\.e\., broadcasting the selection to downstream views\.

Since LLM coding behavior can vary substantially across ML tasks, a task selection dropdown allows users to focus on individual tasks \(see details of the 22 tasks in Appendix\)\. To support a holistic comparison, we also provide an “All Tasks” mode that trains classifiers on all 22 tasks jointly, separating two code sets of22×1,00022\{\\times\}1,000samples each\.

### 5\.2\\treeview: Interactive Token Exploration

Design Rationale\.While the\\metricviewanswerswhichLLM pairs differ, the\\treeviewexplainshowthey differ by visualizing the trained classifier structure\. Decision trees are inherently interpretable: each internal node denotes a discriminative feature and its splitting threshold, while leaf nodes indicate classification outcomes\. By directly visualizing the tree structure, users can understand the hierarchical decision logic that separates two code sets \(\\reqboxtwoR2\.1\)\.

Decision Tree Visualization\.We visualize the tree with a node\-link diagram for its intuitive clarity \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b1\)\. The encoding is as follows:

- •Node label: Shows the splitting token and threshold for internal nodes \(e\.g\.,try≤\\leq0\.5: samples with notrygo to the Yes branch\) and the classified set for leaf nodes; sample counts are shown in both\.
- •Node color: The diverging color bar at the bottom encodes the sample distribution in a node—blue for Set 1, green for Set 2\. The gray and white background indicate internal and leaf nodes, respectively\.
- •Edge width: Width is proportional to code sample count\.

Feature Importance Panel\.A horizontal bar chart \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2\) next to the tree visualization shows the top\-15 most important tokens, ranked by their contribution to the separability \(\\reqboxtwoR2\.2\)\. Each bar encodes a token’s importance score—its cumulative Gini impurity decrease across all nodes, ranging from 0 to 1\. This panel serves dual purposes: \(1\) providing a ranked summary of discriminative tokens without requiring tree traversal, and \(2\) enabling direct token selection for code context retrieval or feature ablation experiments \(see Interactions below\)\.

Interactions\.This view supports three primary interactions:

Tree Navigation:Users can expand \(\) or collapse \(\) subtrees by clicking individual nodes, allowing progressive disclosure of tree complexity\. The tree is initially rendered with five levels—sufficient for most analyses—with deeper subtrees expandable on demand\. Panning and zooming further support exploration of large trees\.

![Refer to caption](https://arxiv.org/html/2609.22097v1/4code.png)Figure 3:The side\-by\-side comparison of code samples containing the tokenprintfrom the two code sets \(left:o3\_T1\_R1, right:claude\_T1\_R1\)\. The statistics show that, on average,claudeusesprint22\.08 times in each of its code samples\. In contrast,o3uses it much less frequently\.Token\-to\-Code Linking:Clicking any token in the Feature Importance Panel \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2\) triggers aCode Context View\(Fig\.[3](https://arxiv.org/html/2609.22097#S5.F3)\) that retrieves code snippets containing the selected token\. The view shows side\-by\-side comparisons of both LLM runs’ code samples, with the queried token highlighted\. Occurrence statistics \(count, average, variance\) are displayed for each LLM run, enabling users to ground abstract feature importance scores in concrete code examples\. A search bar \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2\) further enables users to search for a specific token, supporting hypothesis\-driven exploration, e\.g\., searchingtryorexceptto test whether the two LLMs differ in exception handling\.

Feature Ablation:Clicking the feature importance bars in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2 initiates dynamic feature removal\. The selected token is added to an exclusion list, and the system retrains the classifier without tokens in the list\. The tree view updates in real\-time with the new tree structure, allowing users to observe how classification logic adapts when discriminative tokens are removed \(\\reqboxtwoR2\.3\)\. Excluded tokens are tracked in the\\ablationview\(Sec\.[5\.3](https://arxiv.org/html/2609.22097#S5.SS3)\)\. This interaction supports “what\-if” analysis: users can test whether observed differences are robust or fragile\. Feature ablation can also uncover features hidden by an equally dominant one: if two features each perfectly separate the two code sets, the tree uses only the first and assigns it an importance of 1, leaving the second invisible in the panel\. Removing the first reveals the second\.

### 5\.3\\ablationview: Progressive Feature Removal

Design Rationale\.The\\ablationview\(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsc\) operationalizes the robustness and concentration metrics through progressive feature removal at scale\. While accuracy captures the initial separability level, it does not reveal the depth or structure\. The\\treeviewsupports single\-step removal, but computing robustness and concentration requires removing hundreds of features in sequence, a process that would be prohibitively tedious to perform manually\. The\\ablationviewautomates this and visualizes the metrics with cumulative details\.

Tri\-Axis Encoding\.The view employs a tri\-axis combination chart \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsc\) whose x\-axis represents the token removal sequence\. Three y\-axes encode complementary metrics:

- •Left y\-axis \(bars\):Each bar \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c1\) encodes the importance of the token removed \(i\.e\., the top\-1\) at that step, representing the upper bound of a token’s importance\. The importance typically decreases as more tokens are removed, but rises toward the end when only a few tokens remain and each dominates by default\. Concentration is therefore computed over only the first half of removals \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c2\)\.
- •Right primary y\-axis \(orange\):An orange curve encodes classification accuracy after each removal, showing how separability degrades as discriminative tokens are progressively eliminated\. A dashed orange line marks the 50% baseline, and the orange area is the robustness\.
- •Right secondary y\-axis \(green\):A green curve encodes the entropy of the top\-15 feature importances after each removal, reflecting how concentrated the discriminative signal is\. A dashed green line marks the maximum entropy \(log2⁡15\\log\_\{2\}15\), and the green area is the concentration\.

Interactions\.The view supports manual removal, automated batch removal, and retrospective inspection of any removal step\.

Manual Removal:Tokens removed via the\\treeview’s Feature Importance Panel are immediately reflected here\. Each removal appends a new bar, accuracy point, and entropy point to the chart, with animated transitions extending the visualization rightward\. Users can thus observe whether each removed feature is critical or redundant\.

Automated Batch Removal:A control panel \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c3\) lets users specify a target removal count \(1–500 features\) and trigger batch removal via a “Start” button\. The system iteratively removes the currently most important feature, retrains the classifier, and updates the chart until the target count is reached\. A progress overlay \(e\.g\., “Removing feature 47/500…”\) keeps users informed, with the option to “Stop” or “Reset” at any time\. This mode is essential for cases where users want to remove many tokens, in which manual removal is impractical\.

Time Travel:Clicking any bar restores the decision tree to the state at that removal step\. For example, clicking the 10th bar loads the tree after the first 9 features have been removed\. Users can identify where accuracy dropped sharply and examine the corresponding tree to understand which remaining features carry more discriminative power\.

Zoom and Pan\.With hundreds of removal steps, the chart becomes compressed; users can zoom into specific regions \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c1\) to locate significant accuracy oscillation that may be invisible in the full view\.

Checkpoint Caching\.Since users often revisit the same pair of LLM runs across sessions, the system automatically caches each removal step keyed by \(LLM run pair, task, tokenizer\)\. Becauseexhaustivefeature removal has already been performed offline when computing robustness and concentration, these precomputed results allow the system to restore any removal step instantly when a pair is selected\.

## 6Case Study and Actionable Insights

This section compares code sets generated by different LLMs, using different tokenizers, at different LLM temperatures, and with different prompts\. Through these comparisons, we reveal fundamental differences between LLMs’ coding styles and demonstrate how the comparative insights help users better control LLMs’ coding behavior\.

Figure 4:The\\ablationviewwhen comparing different pairs of LLM runs in different tokenization levels\.### 6\.1Comparing Different LLMs

To gain an overall impression, we first compare the 10 LLMs \(Sec\.[3](https://arxiv.org/html/2609.22097#S3)\) using the “All Tasks” mode with thecompletetokenizer\. The 20 LLM runs appear as 20 rows/columns in the matrix of Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsa\. Each cell represents a comparison between two code sets, each with22,00022\{,\}000code samples \(22​tasks×1,000​code samples per LLM run22\\ \\text\{tasks\}\{\\times\}1\{,\}000\\ \\text\{code samples per LLM run\}\)\.

The matrix shows that code from the*same*LLM’s two independent runs is not separable: light blue cells near the diagonal in the upper\-right triangle indicate≈\\approx50% accuracy, and light orange cells near the diagonal in the lower\-left triangle confirm low robustness\. We click a representative cell,gpt\-4o\_T1\_R1vs\.gpt\-4o\_T1\_R2\(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-a1\), and examine it in the\\ablationview\. As shown in Fig\.[4](https://arxiv.org/html/2609.22097#S6.F4)a, the accuracy curve is always around 50%, reflecting low separability and a low robustness value of 1\.11\. The entropy curve is always aroundlog2⁡15\\log\_\{2\}15\(the max value\), indicating that no dominant features contribute to the separation, resulting in a very low concentration value of 4\.97\.

In contrast, for any two runs from*different*LLMs, the separability is very high, as reflected by the dark blue cell color in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsa\. Switching to the scatterplot mode, we select a point at the far top\-right corner with “\\sharpcascadingSharp & Cascading” separability, representinggpt\-4o\_T1\_R1vs\.gpt\-5\.2\_T1\_R1\(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-a2\)\. The corresponding\\ablationviewis shown in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Modelsc, in which the orange and green areas are very large \(robustness:192\.49192\.49, concentration:220\.97220\.97\), echoing the persistent separability and concentrated discriminative signal\.

We then use the\\treeviewto identify which tokens drive this separation\. As shown in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2, the separability is dominated by tokenexist\_ok, the root splitting feature of the decision tree \(Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b1\)\. Whenexist\_okappears fewer than 0\.5 times \(theYesbranch\), the code is classified as Set 1 \(gpt\-4o\_T1\_R1\); otherwise as Set 2 \(gpt\-5\.2\_T1\_R1\)\. Based on this single token’s frequency, the tree can almost correctly separate the35,20035\{,\}200code samples from the two sets, i\.e\.,2×22,000×80%2\{\\times\}22\{,\}000\{\\times\}80\\%\(the training data is 80% of all code samples\), as the two second\-level tree nodes have almost pure blue or green color\.

Inspecting token occurrences via theCode Context View\(Fig\.[3](https://arxiv.org/html/2609.22097#S5.F3)\),exist\_okappears exclusively in two patterns:

os\.makedirs\(WORKING\_DIR,exist\_ok=True\)

WORKING\_DIR\.mkdir\(parents=True,exist\_ok=True\)

Among the22,00022\{,\}000code samples generated bygpt\-4o\_T1\_R1, only 20 containexist\_ok; forgpt\-5\.2\_T1\_R1, the count is 21,922\. This stark contrast makesexist\_oka fingerprint ofgpt\-5\.2\. As long as it occurs, we can confidently conclude that the code is fromgpt\-5\.2\.

The code context of this token also leads us to hypothesize thatgpt\-5\.2is more cautious about data writing, i\.e\., usingmakedirsormkdirto proactively create the directory in case it does not exist\. To verify this, we use the search bar in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-b2 to retrieve the occurrence counts formakedirsandmkdir\. As shown in Tab\.[1](https://arxiv.org/html/2609.22097#S6.T1),gpt\-5\.2uses these two tokens in21,92221\{,\}922out of its22,00022\{,\}000code samples, whilegpt\-4odoes so only in2525out of its22,00022\{,\}000samples\. This stark frequency contrast convincingly confirms our hypothesis\. The case also illustrates\\sysname’s hypothesis\-driven workflow: a discriminative token \(exist\_ok\) surfaces a behavioral signal, and the search bar lets users immediately pursue the underlying cause, confirming a systematic divergence in how the two LLMs handle directory creation\.

Table 1:The frequency ofmakedirsandmkdirin two code sets\.gpt\-4o\_T1\_R1gpt\-5\.2\_T1\_R1\# occurrences\# files\# occurrences\# filesmakedirs332516,14916,138mkdir005,7915,784total3325/22,00021,94021,922/22,000To verify whether the two code sets only differ inexist\_ok, we remove it and retrain the tree with the remaining 499 features\. As shown in Fig\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\-c1, the tokentr\_idxbecomes the most important feature, with an importance of 0\.9122\. This cascading dominance reflects the concentrated discriminative signal, echoing the high concentration\.

Per\-Task Comparison\.We have also compared LLM runs using the code of individual tasks, i\.e\., differentiating2×1,0002\{\\times\}1\{,\}000code samples\.The general trend is consistent with the “All Tasks” modewhere we differentiate2×22,0002\{\\times\}22\{,\}000code samples \(the more comprehensive comparison\)\. However, the most discriminative tokens vary across individual tasks, as each task has a particular problem type and data modality\. For example, the tokentorchvisionappears more prominently in image\-related tasks but rarely appears in tasks of tabular data\. As the per\-task comparison does not introduce extra insights beyond the “All Tasks” mode, we include its details in our Appendix\.

### 6\.2Comparing LLMs using Different Tokenizers

Next, we switch to different tokenizers to compare the findings with those from thecompletetokenizer in Sec\.[6\.1](https://arxiv.org/html/2609.22097#S6.SS1)\.

Thefunctionalandcodetokenizers behave very similarly to thecompletetokenizer: code from the same LLM is inseparable, while code from different LLMs is highly separable\. The corresponding three sets of 190 \(20×19/220\{\\times\}19/2\) LLM run pairs also follow similar distributions in the scatterplot of Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)a \(the yellow, red, and blue points\)\. There is a slight shift toward the right \(more robust\) from the yellow \(functional\), to the red \(code\), to the blue \(complete\) clusters\. We believe this is due to the inclusive relationship among the token sets produced by the three tokenizers, i\.e\.,functional⊂\\subsetcode⊂\\subsetcomplete\. All LLM pairs in the three tokenization levels have robust separability, while the concentration of discriminative signals varies across pairs\. Some LLM pairs, e\.g\.,o3andclaude, can be easily separated by a single token \(printin Fig\.[3](https://arxiv.org/html/2609.22097#S5.F3)\), i\.e\., the separability is “\\sharpcascadingSharp & Cascading”, while others, e\.g\.,gpt\-5andgpt\-5\-mini, need multiple tokens to work collaboratively, i\.e\., the “\\deepdistributedDeep & Distributed” separability in Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1)\. The LLM pairs with the latter type of separability often share similar LLM versions, indicating thatalthough the separability between them is still robust, the discriminative signal is not sharp\.

For thecommenttokenizer, we notice that the accuracies for pairs \(1\)gpt\-4\.1vs\.gpt\-4\.1\-mini, \(2\)gpt\-5vs\.gpt\-5\-mini, and \(3\)gpt\-5vs\.gpt\-5\.1are relatively low, as reflected by the three lighter blue blocks in Fig\.[5](https://arxiv.org/html/2609.22097#S6.F5)a\. These LLM pairs show more similar commenting styles, and we believe their similarity was driven by their similar LLM versions\. From them, we randomly select an LLM run pair to examine its details, i\.e\.,gpt\-5\_T1\_R1vs\.gpt\-5\-mini\_T1\_R1\. As shown by the top\-6 important features in Fig\.[5](https://arxiv.org/html/2609.22097#S6.F5)b, the discriminative signal is not dominated by any single feature\. The corresponding tree visualization in Fig\.[5](https://arxiv.org/html/2609.22097#S6.F5)c echoes this\. The root node uses the tokenreadto split the two code sets\. However, after this splitting, the two second\-level tree nodes still show a significant mix of code samples \(the mix of blue and green bars\), reflecting the weak discriminative power ofread\. The corresponding\\ablationviewis shown in Fig\.[4](https://arxiv.org/html/2609.22097#S6.F4)b, where the orange and green areas reflect the relatively lower robustness and much lower concentration values\. Extending the analysis to more pairs, we found that thecommentlevel often shows relatively lower robustness and concentration compared to thecodelevel, as shown by the green points in Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)a, indicating there are fewer dominant tokens in the comments\. This is expected, as comments are in natural language and share paraphrasable vocabularies, spreading the discriminative signal across many tokens\. In contrast, code must obey strict syntax, so each LLM’s preferences can manifest as individually decisive token choices\.In short, the discriminative power of comment tokens is weaker than that of code tokens, as code needs to follow much stricter syntax\.

For thepackagelevel, the accuracy, robustness, and concentration values are generally lower\. This is because the number of unique tokens \(i\.e\., unique packages\) is much smaller, resulting in much shorter feature vectors when training the decision trees\. While exploring the\\metricview\(screenshots in the Appendix\), we notice that the classifier separatinggpt\-4o\_T1\_R1andgrok\_T1\_R1shows noticeably lower accuracy\. Fig\.[4](https://arxiv.org/html/2609.22097#S6.F4)c shows the\\ablationviewwhen comparing this pair\. There are only 97 unique features from the two code sets \(other levels have 500 features\)\. Among them, 37 have the same frequency \(features 61–97 in Fig\.[4](https://arxiv.org/html/2609.22097#S6.F4)c\), contributing nothing to the separability\. This further suggests thatgpt\-4oandgrok\-4share similar package preferences\.To conclude, thepackagetokenizer produces the most “\\fragilediffuseFragile & Diffuse” separations \(Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)a\): its feature space is small \(fewer unique packages\), and capable LLMs tend to converge on similar optimal libraries for the same tasks, leaving little stylistic divergence\.

Note that for the two runs generated by the same LLM, their separability is always “\\fragilediffuseFragile & Diffuse” across all five tokenizers, because these runs are indistinguishable\. Fig\.[2](https://arxiv.org/html/2609.22097#S5.F2)\-a1 shows one example, i\.e\.,o3\_T1\_R1vs\.o3\_T1\_R2\. The five points representing the five tokenizers for the pair are connected by dashed lines for easy identification\.

Regime Coverage\.We have seen separability types covering three regimes explained in Fig\.[1](https://arxiv.org/html/2609.22097#S4.F1)\. Only the “\\lexicalbrittleLexical & Brittle” one is not covered\. This regime works as a sanity check on the generalizability of our trained classifiers\. Specifically, the feature importance can only be computed from the training data, but the classifier accuracy is derived from the test data\. A “\\lexicalbrittleLexical & Brittle” separability would mean that dominant features exist \(high concentration\) during training and significantly reduce the impurity when separating the two code sets, yet the test accuracy remains around 50% \(low robustness\)\. This contradiction leads to the conclusion that the trained classifier does not generalize well to the test data\. The absence of LLM run pairs in this regime suggests consistency between the training and test data distributions and supports the generalizability of our trained classifiers\.

![Refer to caption](https://arxiv.org/html/2609.22097v1/6comments.png)Figure 5:\(a\) Thecommentlevel accuracy matrix\. \(b\) Diffuse feature importance \(i\.e\., high entropy\)\. \(c\) Decision tree with mixed sets of code\.
### 6\.3Comparing LLMs with Different Temperatures

To investigate the impact of LLMs’ temperature on their code generation behavior, we varied the temperature \(0\.3, 0\.5, 0\.7, and 1\) of two LLMs,grokandclaude, and generated additional code following our dual\-run mechanism \(Sec\.[3](https://arxiv.org/html/2609.22097#S3)\)\. This code corpus contains2​\(L​L​M​s\)×22​\(t​a​s​k​s\)×2​\(r​u​n​s\)×1,000​\(c​o​d​e\)×4​\(t​e​m​p​e​r​a​t​u​r​e​s\)=352,0002\\ \(LLMs\)\{\\times\}\\allowbreak 22\\ \(tasks\)\{\\times\}\\allowbreak 2\\ \(runs\)\{\\times\}\\allowbreak 1\{,\}000\\ \(code\)\{\\times\}\\allowbreak 4\\ \(temperatures\)\{=\}352\{,\}000code samples\. We focus on thecompletetokenizer and use the “All Tasks” mode, as they provide the most comprehensive characterization\.

Fig\.[6](https://arxiv.org/html/2609.22097#S6.F6)shows the\\metricview\. Note that we adjusted the color mapping to emphasize value differences in the lower range \(see the legends\)\. As expected, accuracy is high when the compared code sets are from different LLMs\. When comparing the same LLM with different temperatures,grokshows no obvious difference, as the accuracy values are always around 50%\. In contrast,claudeshows differences that become more pronounced as the temperature gap widens \(i\.e\., between 0\.3 and 1\)\. However, the difference is not dramatic, as the accuracy is always below 65% according to the color legend\. The robustness metric follows a similar pattern to accuracy\. The concentration metric, however, shows no notable variation across temperatures and remains very small, indicating no single token dominates the separability\.

We further explored individual tokens using the\\treeviewand\\ablationview, but no dominant discriminative features emerged, corroborating the low concentration values above\. Taken together, while temperature has limited impact on coding style,claudeexhibits richer temperature\-sensitive behavior thangrok, suggesting that temperature is a more effective lever for steeringclaude\.

Figure 6:Comparinggrokandclaudewith different temperatures\.
### 6\.4Comparing LLMs with Different Prompts

Comparingo3withclaude, we found that tokenprintdominates the separation\. As shown in the corresponding decision tree \(Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)b\), whenprintoccurs less than 9\.5 times, the code is more likely to be fromo3, indicating thatclaudeis much more verbose thano3\(Fig\.[3](https://arxiv.org/html/2609.22097#S5.F3)echoes this\)\. In practice, users may need to useclaudebut control its verbosity\. This motivates us to steerclaude’s coding behavior\. Specifically, we added the following to the prompt ofclaudeand generated a new run of code\. Following our naming convention, this new run isclaude\_T1\_R3\.

“When writing the code, keep output minimal\. Avoid verbose prints, debug messages, chatty logs, or explanatory console text\. Only include print/log statements if they are strictly necessary for the program’s functionality \(e\.g\., required output format\)\.”

Comparingclaude\_T1\_R3toclaude\_T1\_R1, the classifier reaches 100% accuracy \(see values inside the cells of Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)a\) andprintis the dominant feature \(Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)c\)\. Whenprintoccurs less than 7\.5 times, the code is more likely to be fromclaude\_T1\_R3, indicating that the prompt successfully suppressed the verbosity ofclaude\. This new set of code is also sufficiently separable fromclaude\_T1\_R1even after removingprint, with robustness of 106\.3 and concentration of 55\.1\. For reference, when comparingclaude\_T1\_R2toclaude\_T1\_R1, these two values are only 1\.0 and 7\.4, respectively \(Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)a\)\. The much higher robustness and concentration reflect that the new prompt not only changes the verbosity but also other intricate coding behaviors\.

Figure 7:Comparing LLMs when using different coding prompts\.When comparingclaude\_T1\_R3too3\_T1\_R1,printis no longer the dominant feature \(Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)d\), because both sets are less verbose\. In fact,printwas never used in the top three levels of the tree as a splitting feature\. The entropy of the top\-15 important features is also much higher \(2\.36 bits\), indicating no single feature dominates the separability\. Notably,claude\_T1\_R3is more distant fromo3\_T1\_R1than fromclaude\_T1\_R1, because the robustness \(174\.3\>106\.3174\.3\{\>\}106\.3\) and concentration \(81\.8\>55\.181\.8\{\>\}55\.1\) are both larger \(Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)a\)\. This implies that reducing verbosity inclaudedoes not simply bring it closer too3but instead exposes deeper stylistic differences between the two LLMs\. The classifier accuracy alone cannot reflect this relative distance, as it is always very close to 100% \(the dark blue cells in Fig\.[7](https://arxiv.org/html/2609.22097#S6.F7)a\)\.

## 7Feedback from Target Users

We further evaluate\\sysnamewith ML scientists and engineers—our target users—in two complementary settings: expert think\-alouds forin\-depthdesign feedback, and a large\-scale practitioner seminar forbroadsignal\.

Think\-Aloud \+ Guided Explorations\.In the first setting, we conducted think\-aloud sessions with the three experts \(E1E\_\{1\}–E3E\_\{3\}\) who contributed to our requirement analysis\. We walked them through\\sysnameand encouraged free exploration\. The sessions totaled over six hours across two weeks\. All experts quickly grasped the mechanics of\\sysnameand engaged enthusiastically\.E1E\_\{1\}, who proposedarea under the accuracy curvefor robustness, expressed strong appreciation for seeing the idea realized, noting that this formulation is hyperparameter\-free\.E2E\_\{2\}, whose work centers on feature engineering, confirmed that progressive feature ablation is an effective lens for studying separability\. He was particularly impressed by the prompt engineering study \(Sec\.[6\.4](https://arxiv.org/html/2609.22097#S6.SS4)\), calling it a practical method for steering LLMs through targeted prompting\.E3E\_\{3\}offered a more theoretical perspective, confirming the soundness of token\-frequency analysis and suggesting that squaring the deviation from 50% in Eq\.[1](https://arxiv.org/html/2609.22097#S4.E1)could give greater weight to steps with high accuracy\. The three experts consistently recognized the unique value of the system\. As a strong indication of their interest, all offered to contribute code from their own projects for future comparison\.E1E\_\{1\}was also excited about extending the framework to compare LLM\-generated and human\-written code—a direction that could reveal fundamental differences in coding between humans and LLMs\.

Seminar \+ Feedback Survey\.In the second setting, we gave a one\-hour seminar to 40\+ full\-time ML scientists\. The seminar began with a structured overview of\\sysname, followed by a live demonstration using real LLM comparison tasks, with participants encouraged to ask questions freely\. Afterward, 35 of them voluntarily completed an anonymized 20\-question survey, covering demographics, self\-assessed understanding of\\sysname, perceived utility of each component, and overall impressions\. Fig\.[8](https://arxiv.org/html/2609.22097#S8.F8)shows the response distributions across 10 key questions \(see full responses in Appendix\)\. Overall, participants demonstrated strong appreciation for the system and expressed positive assessments of\\sysname\. The participants also offered valuable comments\. One senior scientist connected our frequency\-based analysis to TF\-IDF\[[50](https://arxiv.org/html/2609.22097#bib.bib54)\], suggesting that token frequency should be normalized by code length to account for verbosity\. Two scientists noted that our regex\-based tokenizer differs from LLM tokenizers\[[47](https://arxiv.org/html/2609.22097#bib.bib56),[70](https://arxiv.org/html/2609.22097#bib.bib55)\]that may split a single word into multiple tokens \(e\.g\., fromtokenizationtotokenandization\), whereas ours never breaks a word, preserving word\-level semantics\. Multiple scientists also commented that humans are needed to connect code tokens to LLMs’ personas, e\.g\., from the frequency ofprinttoverbose\. This observation highlights a meaningful role for\\sysname: it serves as the bridge between raw code tokens and human\-interpretable LLM personas, making the implicit explicit\.

## 8Discussion, Limitations, and Future Work

While we focus on Python code and ML tasks, the approach generalizes readily to other programming languages and task domains\. Below, we discuss limitations, open design choices, and future directions\.

Syntax Error Handling\.When tokenizing the code, we use AST\-based parsing to accurately extract code components, e\.g\., comments\. However, AST\-based parsing cannot handle code with syntax errors\. Our pipeline handles these cases gracefully: samples that fail to parse contribute zero tokens to the frequency analysis\. In practice, the proportion of unparseable samples is small, as all 10 LLMs are highly capable\. This strategy therefore does not significantly impact our results\.

Figure 8:Post\-session survey results for 10 key questions\.Online Retraining of Decision Trees\.In the\\treeview, users can interactively remove a feature and trigger an immediate retraining of the decision tree withn−1n\{\-\}1features\. At the current scale, retraining completes in seconds and introduces negligible latency\. When computing the robustness and concentration metrics, we have also precomputed all trees offline; the cache mechanism \(Sec\.[5\.3](https://arxiv.org/html/2609.22097#S5.SS3)\) restores any step instantly, eliminating wait time during interactive explorations\.

Hyperparameters\.Three hyperparameters govern our analysis pipeline: the feature vector length \(the 500 most frequent tokens\), the number of features used in the entropy calculation \(top\-15\), and the step fraction used to compute concentration \(1/21/2ofSSin Eq\.[3](https://arxiv.org/html/2609.22097#S4.E3)\)\. These values were determined empirically through iterative exploration but are not unique\. For instance, concentration could be accumulated over the first2/32/3of steps instead of1/21/2\. Such choices do not affect the overall conclusions; as long as all LLM pairs are evaluated under the same hyperparameter settings, the resulting metrics remain fully comparable\.

Beyond Token Frequencies and Python Tasks\.\\sysnamecharacterizes LLM coding behavior through token frequencies for four reasons justified in Sec\.Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models\. We acknowledge, however, that token frequencies do not directly modelstructured properties of programssuch as code length, syntactic complexity, or control\-flow shape\. Extending the analytical pipeline to incorporate per\-sample code\-length normalization \(TF\-IDF\[[50](https://arxiv.org/html/2609.22097#bib.bib54)\]\), AST\-based structural features, or control\-flow graphs is therefore a natural future direction\. A second, complementary axis of extension concerns the*task domain*: our studies are grounded in Python code generation for Kaggle competitions\. While the analytical pipeline of\\sysnameis task\-agnostic, an empirical characterization of how behavioral fingerprints differ across other software\-development scenarios—API integration or multi\-file maintenance—would meaningfully broaden our work\. Both axes can be pursued without modifying the rest of the framework:\\sysname’s modular pipeline provides a clean foundation for such extensions\.

Additionally, the behavioral fingerprints produced by\\sysnameopen several compelling directions beyond pairwise comparison\.*First*, the fine\-grained style signals provide a forensic lens on model lineage\. A model that has been distilled from a teacher model should closely replicate the teacher’s coding patterns and remain nearly indistinguishable inside\\sysname\. This makes\\sysnamea practical tool for surfacing undisclosed distillation relationships\.*Second*, practitioners currently choose LLMs based largely on availability, reputation, or personal familiarity—there is little actionable guidance on*how*LLMs actually differ\.\\sysnameconstructs a*behavioral persona*for each LLM, giving users a principled basis for model selection\.*Third*, the discriminative tokens surfaced by\\sysnameidentify precisely where two LLMs diverge, providing direct, evidence\-based guidance for prompt engineering\. Users can leverage these signals to steer code generation toward a desired coding style\.

## 9Conclusion

In this paper, we present\\sysname, a visual analytics approach for comparing LLM coding behaviors through token\-frequency analysis\.\\sysnametrains an interpretable decision tree on token\-frequency vectors across five tokenization levels and characterizes the separability between two LLMs into four regimes based on two orthogonal metrics: robustness and concentration\. Case studies comparing code across LLMs, tokenization levels, LLM temperatures, and LLM prompts demonstrate the practical value of the approach\. We hope our work encourages the community to look beyond performance\-based LLM evaluation and adopt behavior\-based characterization as a complementary lens\.

## References

- \[1\]AIDE GitHub\.Note:[https://github\.com/WecoAI/aideml](https://github.com/WecoAI/aideml)Accessed: 2025\-03\-05Cited by:[§3](https://arxiv.org/html/2609.22097#S3.p2.1)\.
- \[2\]V\. A\. Alfred, S\. L\. Monica, and D\. U\. Jeffrey\(2007\)Compilers principles, techniques & tools\.pearson Education\.Cited by:[§4\.2](https://arxiv.org/html/2609.22097#S4.SS2.p3.1)\.
- \[3\]I\. Arawjo, C\. Swoopes, P\. Vaithilingam, M\. Wattenberg, and E\. L\. Glassman\(2024\)Chainforge: a visual toolkit for prompt engineering and llm hypothesis testing\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[4\]J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le,et al\.\(2021\)Program synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[5\]A\. Bilal, A\. Jourabloo, M\. Ye, X\. Liu, and L\. Ren\(2017\)Do convolutional neural networks learn class hierarchy?\.IEEE transactions on visualization and computer graphics24\(1\),pp\. 152–162\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[6\]T\. Bisztray, B\. Cherif, R\. A\. Dubniczky, N\. Gruschka, B\. Borsos, M\. A\. Ferrag, A\. Kovacs, V\. Mavroeidis, and N\. Tihanyi\(2025\)I know which llm wrote your code last summer: llm generated code stylometry for authorship attribution\.InProceedings of the 18th ACM Workshop on Artificial Intelligence and Security,pp\. 28–39\.Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p5.1)\.
- \[7\]A\. Caliskan\-Islam, R\. Harang, A\. Liu, A\. Narayanan, C\. Voss, F\. Yamaguchi, and R\. Greenstadt\(2015\)De\-anonymizing programmers via code stylometry\.In24th USENIX Security Symposium \(USENIX Security\),pp\. 255–270\.Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p5.1)\.
- \[8\]J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan,et al\.\(2024\)Mle\-bench: evaluating machine learning agents on machine learning engineering\.arXiv preprint arXiv:2410\.07095\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1),[§3](https://arxiv.org/html/2609.22097#S3.p1.1),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[9\]M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.\(2021\)Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[10\]N\. Chen, Y\. Zhang, J\. Xu, K\. Ren, and Y\. Yang\(2024\)Viseval: a benchmark for data visualization in the era of large language models\.IEEE Transactions on Visualization and Computer Graphics31\(1\),pp\. 1301–1311\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[11\]K\. A\. Cook and J\. J\. Thomas\(2005\)Illuminating the path: the research and development agenda for visual analytics\.Technical reportPacific Northwest National Laboratory \(PNNL\), Richland, WA \(US\)\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p7.1)\.
- \[12\]A\. Coscia, L\. Holmes, W\. Morris, J\. S\. Choi, S\. Crossley, and A\. Endert\(2024\)Iscore: visual analytics for interpreting how language models automatically score summaries\.InProceedings of the 29th international conference on intelligent user interfaces,pp\. 787–802\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[13\]Y\. Dong, X\. Jiang, J\. Qian, T\. Wang, K\. Zhang, Z\. Jin, and G\. Li\(2025\)A survey on code generation with llm\-based agents\.arXiv preprint arXiv:2508\.00083\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1),[§2](https://arxiv.org/html/2609.22097#S2.p1.2)\.
- \[14\]H\. Fang, B\. Han, N\. Erickson, X\. Zhang, S\. Zhou, A\. Dagar, J\. Zhang, A\. C\. Turkmen, C\. Hu, H\. Rangwala,et al\.\(2025\)Mlzero: a multi\-agent system for end\-to\-end machine learning automation\.arXiv preprint arXiv:2505\.13941\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[15\]GitHub Copilot\.Note:[https://github\.com/features/copilot](https://github.com/features/copilot)Accessed: 2025\-03\-05Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[16\]M\. Gleicher, D\. Albers, R\. Walker, I\. Jusufi, C\. D\. Hansen, and J\. C\. Roberts\(2011\)Visual comparison for information visualization\.Information Visualization10\(4\),pp\. 289–309\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[17\]M\. Gleicher\(2017\)Considerations for visualizing comparison\.IEEE transactions on visualization and computer graphics24\(1\),pp\. 413–423\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[18\]L\. Gou, L\. Zou, N\. Li, M\. Hofmann, A\. K\. Shekar, A\. Wendt, and L\. Ren\(2020\)VATLD: a visual analytics system to assess, understand and improve traffic light detection\.IEEE transactions on visualization and computer graphics27\(2\),pp\. 261–271\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[19\]F\. Hohman, M\. Kahng, R\. Pienta, and D\. H\. Chau\(2018\)Visual analytics in deep learning: an interrogative survey for the next frontiers\.IEEE transactions on visualization and computer graphics25\(8\),pp\. 2674–2693\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[20\]T\. Ivanov and V\. Penchev\(2024\)AI benchmarks and datasets for llm evaluation\.arXiv preprint arXiv:2412\.01020\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[21\]J\. Jiang, F\. Wang, J\. Shen, S\. Kim, and S\. Kim\(2024\)A survey on large language models for code generation\.ACM Transactions on Software Engineering and Methodology\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[22\]C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. Narasimhan\(2024\)SWE\-bench: can language models resolve real\-world GitHub issues?\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[23\]M\. Kahng, I\. Tenney, M\. Pushkarna, M\. X\. Liu, J\. Wexler, E\. Reif, K\. Kallarackal, M\. Chang, M\. Terry, and L\. Dixon\(2024\)Llm comparator: interactive analysis of side\-by\-side evaluation of large language models\.IEEE Transactions on Visualization and Computer Graphics\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[24\]M\. Kahng, I\. Tenney, M\. Pushkarna, M\. X\. Liu, J\. Wexler, E\. Reif, K\. Kallarackal, M\. Chang, M\. Terry, and L\. Dixon\(2024\)Llm comparator: visual analytics for side\-by\-side evaluation of large language models\.InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems,pp\. 1–7\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[25\]J\. S\. Kessler\(2017\)Scattertext: a browser\-based tool for visualizing how corpora differ\.InProceedings of ACL 2017, System Demonstrations,pp\. 85–90\.Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2)\.
- \[26\]T\. S\. Kim, Y\. Lee, J\. Shin, Y\. Kim, and J\. Kim\(2024\)Evallm: interactive evaluation of large language model prompts on user\-defined criteria\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–21\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[27\]T\. T\. Le and J\. H\. Moore\(2021\)Treeheatr: an r package for interpretable decision tree visualizations\.Bioinformatics37\(2\),pp\. 282–284\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[28\]G\. Li, J\. Wang, H\. Shen, K\. Chen, G\. Shan, and Z\. Lu\(2020\)Cnnpruner: pruning convolutional neural networks with visual analytics\.IEEE Transactions on Visualization and Computer Graphics27\(2\),pp\. 1364–1373\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[29\]L\. Li, B\. Dong, R\. Wang, X\. Hu, W\. Zuo, D\. Lin, Y\. Qiao, and J\. Shao\(2024\)Salad\-bench: a hierarchical and comprehensive safety benchmark for large language models\.arXiv preprint arXiv:2402\.05044\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[30\]Y\. Li, J\. Wang, P\. Aboagye, C\. M\. Yeh, Y\. Zheng, L\. Wang, W\. Zhang, and K\. Ma\(2024\)Visual analytics for efficient image exploration and user\-guided image captioning\.IEEE transactions on visualization and computer graphics30\(6\),pp\. 2875–2887\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[31\]Y\. Li, J\. Wang, X\. Dai, L\. Wang, C\. M\. Yeh, Y\. Zheng, W\. Zhang, and K\. Ma\(2023\)How does attention work in vision transformers? a visual analytics attempt\.IEEE transactions on visualization and computer graphics29\(6\),pp\. 2888–2900\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[32\]M\. Liu, J\. Shi, Z\. Li, C\. Li, J\. Zhu, and S\. Liu\(2016\)Towards better analysis of deep convolutional neural networks\.IEEE transactions on visualization and computer graphics23\(1\),pp\. 91–100\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[33\]S\. Liu, J\. Xiao, J\. Liu, X\. Wang, J\. Wu, and J\. Zhu\(2017\)Visual diagnosis of tree boosting methods\.IEEE transactions on visualization and computer graphics24\(1\),pp\. 163–173\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[34\]S\. Liu, W\. Yang, J\. Wang, and J\. Yuan\(2025\)Visualization for artificial intelligence\.Springer\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[35\]Z\. Liu, Y\. Cai, X\. Zhu, Y\. Zheng, R\. Chen, Y\. Wen, Y\. Wang, S\. Chen,et al\.\(2025\)ML\-master: towards ai\-for\-ai via integration of exploration and reasoning\.arXiv preprint arXiv:2506\.16499\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[36\]D\. Lopez\-Paz and M\. Oquab\(2017\)Revisiting classifier two\-sample tests\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2)\.
- \[37\]Y\. Ming, H\. Qu, and E\. Bertini\(2018\)RuleMatrix: visualizing and understanding classifiers with rules\.IEEE Transactions on Visualization and Computer Graphics25\(1\),pp\. 342–352\.Cited by:[item 2](https://arxiv.org/html/2609.22097#S4.I1.i2.p1.1)\.
- \[38\]Y\. Ming, P\. Xu, F\. Cheng, H\. Qu, and L\. Ren\(2019\)ProtoSteer: steering deep sequence model with prototypes\.IEEE transactions on visualization and computer graphics26\(1\),pp\. 238–248\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[39\]MLE\-Bench Leaderboard\.Note:[https://github\.com/openai/mle\-bench](https://github.com/openai/mle-bench)Accessed: 2026\-03\-05Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[40\]B\. L\. Monroe, M\. P\. Colaresi, and K\. M\. Quinn\(2008\)Fightin’ words: lexical feature selection and evaluation for identifying the content of political conflict\.Political Analysis16\(4\),pp\. 372–403\.Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2)\.
- \[41\]T\. Mühlbacher, L\. Linhardt, T\. Möller, and H\. Piringer\(2017\)Treepod: sensitivity\-aware selection of pareto\-optimal decision trees\.IEEE transactions on visualization and computer graphics24\(1\),pp\. 174–183\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[42\]J\. Nam, J\. Yoon, J\. Chen, J\. Shin, S\. Ö\. Arık, and T\. Pfister\(2025\)Mle\-star: machine learning engineering agent via search and targeted refinement\.arXiv preprint arXiv:2506\.15692\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[43\]B\. Pan, Y\. Fu, K\. Wang, J\. Lu, L\. Pan, Z\. Qian, Y\. Chen, G\. Wang, Y\. Zhou, L\. Zheng,et al\.\(2025\)VIS\-shepherd: constructing critic for llm\-based data visualization generation\.arXiv preprint arXiv:2506\.13326\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[44\]Y\. Qin, K\. Song, Y\. Hu, W\. Yao, S\. Cho, X\. Wang, X\. Wu, F\. Liu, P\. Liu, and D\. Yu\(2024\)Infobench: evaluating instruction following ability in large language models\.arXiv preprint arXiv:2401\.03601\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[45\]A\. Rathore, S\. Dev, J\. M\. Phillips, V\. Srikumar, Y\. Zheng, C\. M\. Yeh, J\. Wang, W\. Zhang, and B\. Wang\(2024\)VERB: visualizing and interpreting bias mitigation techniques geometrically for word representations\.ACM Transactions on Interactive Intelligent Systems14\(1\),pp\. 1–34\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[46\]J\. C\. Roberts\(2007\)State of the art: coordinated & multiple views in exploratory visualization\.InFifth international conference on coordinated and multiple views in exploratory visualization \(CMV 2007\),pp\. 61–71\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p7.1)\.
- \[47\]R\. Sennrich, B\. Haddow, and A\. Birch\(2016\)Neural machine translation of rare words with subword units\.InProceedings of the 54th annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 1715–1725\.Cited by:[§7](https://arxiv.org/html/2609.22097#S7.p3.1)\.
- \[48\]R\. Sevastjanova, S\. Vogelbacher, A\. Spitz, D\. Keim, and M\. El\-Assady\(2023\)Visual comparison of text sequences generated by large language models\.In2023 IEEE Visualization in Data Science \(VDS\),pp\. 11–20\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[49\]M\. Sondag, C\. Meinecke, D\. Collaris, T\. Von Landesberger, and S\. Van Den Elzen\(2025\)Cluster\-based random forest visualization and interpretation\.IEEE Transactions on Visualization and Computer Graphics\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[50\]K\. Sparck Jones\(1972\)A statistical interpretation of term specificity and its application in retrieval\.Journal of documentation28\(1\),pp\. 11–21\.Cited by:[§7](https://arxiv.org/html/2609.22097#S7.p3.1),[§8](https://arxiv.org/html/2609.22097#S8.p5.1)\.
- \[51\]H\. Strobelt, S\. Gehrmann, M\. Behrisch, A\. Perer, H\. Pfister, and A\. M\. Rush\(2018\)S eq 2s eq\-v is: a visual debugging tool for sequence\-to\-sequence models\.IEEE transactions on visualization and computer graphics25\(1\),pp\. 353–363\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[52\]E\. Toledo, K\. Hambardzumyan, M\. Josifoski, R\. Hazra, N\. Baldwin, A\. Audran\-Reiss, M\. Kuchnik, D\. Magka, M\. Jiang, A\. M\. Lupidi,et al\.\(2025\)AI research agents for machine learning: search, exploration, and generalization in mle\-bench\.arXiv preprint arXiv:2507\.02554\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[53\]L\. Twist, J\. M\. Zhang, M\. Harman, D\. Syme, J\. Noppen, and D\. Nauck\(2025\)LLMs love python: a study of llms’ bias for programming languages and libraries\.arXiv preprint arXiv:2503\.17181\.Cited by:[§2](https://arxiv.org/html/2609.22097#S2.p1.2)\.
- \[54\]S\. Van Den Elzen and J\. J\. Van Wijk\(2011\)Baobabview: interactive construction and analysis of decision trees\.In2011 IEEE conference on visual analytics science and technology \(VAST\),pp\. 151–160\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[55\]J\. Wang, Y\. Chen, M\. Pan, C\. M\. Yeh, and M\. Das\(2026\)Illuminating llm coding agents: visual analytics for deeper understanding and enhancement\.IEEE Transactions on Visualization and Computer Graphics\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1),[§2](https://arxiv.org/html/2609.22097#S2.p1.2),[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[56\]J\. Wang, L\. Gou, H\. Shen, and H\. Yang\(2018\)Dqnviz: a visual analytics approach to understand deep q\-networks\.IEEE transactions on visualization and computer graphics25\(1\),pp\. 288–298\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[57\]J\. Wang, L\. Gou, H\. Yang, and H\. Shen\(2018\)GANViz: a visual analytics approach to understand the adversarial game\.IEEE transactions on visualization and computer graphics24\(6\),pp\. 1905–1917\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[58\]J\. Wang, L\. Gou, W\. Zhang, H\. Yang, and H\. Shen\(2019\)Deepvid: deep visual interpretation and diagnosis for image classifiers via knowledge distillation\.IEEE transactions on visualization and computer graphics25\(6\),pp\. 2168–2180\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[59\]J\. Wang, S\. Liu, and W\. Zhang\(2024\)Visual analytics for machine learning: a data perspective survey\.IEEE transactions on visualization and computer graphics30\(12\),pp\. 7637–7656\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[60\]J\. Wang, L\. Wang, Y\. Zheng, C\. M\. Yeh, S\. Jain, and W\. Zhang\(2022\)Learning\-from\-disagreement: a model comparison and visual analytics framework\.IEEE Transactions on Visualization and Computer Graphics29\(9\),pp\. 3809–3825\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1),[item 2](https://arxiv.org/html/2609.22097#S4.I1.i2.p1.1)\.
- \[61\]J\. Wang, W\. Zhang, L\. Wang, and H\. Yang\(2021\)Investigating the evolution of tree boosting models with visual analytics\.In2021 IEEE 14th Pacific visualization symposium \(PacificVis\),pp\. 186–195\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[62\]L\. Wang, J\. Wang, C\. M\. Yeh, Y\. Zheng, J\. Sun, X\. Fan, X\. Dai, Y\. Fan, and Y\. Cai\(2026\)Understanding llm evaluator behavior: a structured multi\-evaluator framework for merchant risk assessment\.arXiv preprint arXiv:2602\.05110\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[63\]L\. Wang, J\. Wang, Y\. Zheng, S\. Jain, C\. M\. Yeh, Z\. Zhuang, J\. Ebrahimi, and W\. Zhang\(2022\)Learning from disagreement for event detection\.In2022 IEEE International Conference on Big Data \(Big Data\),pp\. 2411–2418\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[64\]X\. Wang, B\. Li, Y\. Song, F\. F\. Xu, X\. Tang, M\. Zhuge, J\. Pan, Y\. Song, B\. Li, J\. Singh,et al\.\(2024\)Openhands: an open platform for ai software developers as generalist agents\.arXiv preprint arXiv:2407\.16741\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[65\]X\. Wang, C\. Liang, S\. Zheng, J\. Liang, G\. Li, Y\. Zhang, and C\. H\. Liu\(2024\)Visualization generation with large language models: an evaluation\.arXiv preprint arXiv:2401\.11255\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p2.1)\.
- \[66\]Y\. Wang, W\. Wang, S\. Joty, and S\. C\.H\. Hoi\(2021\)CodeT5: identifier\-aware unified pre\-trained encoder\-decoder models for code understanding and generation\.InEMNLP,Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[67\]Weco AI\(2024\)AIDE: Human\-Level Performance in Data Science Competitions,[https://www\.weco\.ai/blog/technical\-report](https://www.weco.ai/blog/technical-report)\.Note:Accessed: 2025\-01\-10External Links:[Link](https://www.weco.ai/blog/technical-report)Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1),[§3](https://arxiv.org/html/2609.22097#S3.p2.1)\.
- \[68\]J\. Wexler, M\. Pushkarna, T\. Bolukbasi, M\. Wattenberg, F\. Viégas, and J\. Wilson\(2019\)The what\-if tool: interactive probing of machine learning models\.IEEE transactions on visualization and computer graphics26\(1\),pp\. 56–65\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1),[3rd item](https://arxiv.org/html/2609.22097#S2.I1.i2.I1.i3.p1.2)\.
- \[69\]H\. Wijk, T\. Lin, J\. Becker, S\. Jawhar, N\. Parikh, T\. Broadley, L\. Chan, M\. Chen, J\. Clymer, J\. Dhyani,et al\.\(2024\)Re\-bench: evaluating frontier ai r&d capabilities of language model agents against human experts\.arXiv preprint arXiv:2411\.15114\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[70\]Y\. Wu, M\. Schuster, Z\. Chen, Q\. V\. Le, M\. Norouzi, W\. Macherey, M\. Krikun, Y\. Cao, Q\. Gao, K\. Macherey,et al\.\(2016\)Google’s neural machine translation system: bridging the gap between human and machine translation\.arXiv preprint arXiv:1609\.08144\.Cited by:[§7](https://arxiv.org/html/2609.22097#S7.p3.1)\.
- \[71\]W\. Yang, X\. Wang, J\. Lu, W\. Dou, and S\. Liu\(2020\)Interactive steering of hierarchical clustering\.IEEE Transactions on Visualization and Computer Graphics27\(10\),pp\. 3953–3967\.Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p3.1)\.
- \[72\]X\. Yang, X\. Yang, S\. Fang, Y\. Zhang, J\. Wang, B\. Xian, Q\. Li, J\. Li, M\. Xu, Y\. Li, H\. Pan, Y\. Zhang, W\. Liu, Y\. Shen, W\. Chen, and J\. Bian\(2025\)R&D\-Agent: An LLM\-Agent Framework Towards Autonomous Data Science\.arXiv preprint arXiv:2505\.14738\.External Links:[Link](https://arxiv.org/abs/2505.14738)Cited by:[§1](https://arxiv.org/html/2609.22097#S1.p1.1)\.
- \[73\]J\. Yuan, B\. Barr, K\. Overton, and E\. Bertini\(2022\)Visual exploration of machine learning model behavior with hierarchical surrogate rule sets\.IEEE Transactions on Visualization and Computer Graphics30\(2\),pp\. 1470–1488\.Cited by:[item 2](https://arxiv.org/html/2609.22097#S4.I1.i2.p1.1)\.
- \[74\]Z\. Zeng, J\. Yu, T\. Gao, Y\. Meng, T\. Goyal, and D\. Chen\(2023\)Evaluating large language models at evaluating instruction following\.arXiv preprint arXiv:2310\.07641\.Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p3.1)\.
- \[75\]X\. Zhao, J\. Wang, Y\. Chen, M\. Pan, C\. M\. Yeh, J\. Sun, Y\. Zheng, M\. Das, and T\. Chen\(2026\)Demystify the role of memory in machine learning engineering agents\.InFindings of the Association for Computational Linguistics: ACL 2026,San Diego, USA\.External Links:[Link](https://aclanthology.org/2026.findings-acl.1/)Cited by:[Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models](https://arxiv.org/html/2609.22097#p4.1)\.

相似文章

大型语言模型的频谱特征

arXiv cs.CL

本文介绍了一种基于频谱形状的度量,利用重尾自正则化理论来表征、比较和管理大型语言模型。该方法无需数据、计算高效且尺度不变,支持跨多种模型集合的血统追踪、无监督聚类和性能量化。

大规模语言模型的概率归因

arXiv cs.CL

本文提出了一种与模型无关的基于概率的令牌归因度量,利用贝叶斯规则反转下一个令牌的对数概率,捕捉模型对令牌序列的内部表示,并通过熵分析提高可解释性。