Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations
Summary
This paper introduces an unsupervised, training-free approach to discover prompt-conditional stylistic axes in LLM hidden activations using sampling and PCA, validated through human studies.
View Cached Full Text
Cached at: 09/18/26, 08:47 AM
# Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations
Source: [https://arxiv.org/html/2609.19150](https://arxiv.org/html/2609.19150)
Ajit Mallavarapu Cornell University / Ithaca, NY am3574@cornell\.edu &Ziwei Gu Harvard University / Cambridge, MA ziweigu@g\.harvard\.edu
###### Abstract
Large language models \(LLMs\) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data\. We present a training\-free, prompt\-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis \(PCA\) to the pooled hidden activations, and label the resulting axes automatically from the pole generations\. We validate the discovered axes against 245 human\-elicited stylistic annotations in a two\-phase study\. On our strongest model \(Qwen\-3\.5\-4B\-Instruct\), the top two axes match spontaneously requested human dimensions with 72\.8% precision and 43\.6% macro\-recall, and 75\.6% of validity ratings judge the axes’ polar generations accurate to their labels, with 90\.9% adjacent inter\-annotator agreement\. Discoverability is strongly model\-dependent: both Qwen models and Llama\-3\.2\-3B expose human\-salient axes, while DeepSeek\-7B\-Chat drops to 35\.3% precision, its leading components dominated by structural rather than stylistic variance\. Simple PCA over a model’s own decoding variance is thus an effective, low\-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross\-model differences in how that structure is organized\.
Sampling Reveals Style: Unsupervised, Training\-Free Discovery of Prompt\-Conditional Stylistic Axes in LLM Activations
Ajit MallavarapuCornell University / Ithaca, NYam3574@cornell\.eduZiwei GuHarvard University / Cambridge, MAziweigu@g\.harvard\.edu
## 1Introduction
Large language models \(LLMs\) encode a vast array of human knowledge within their internal representations\. As these models scale, understanding the conceptual structures they naturally form becomes increasingly critical for interpretability and safety\(Zouet al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib4)\)\. However, discovering these concepts without relying on manual annotations or predefined taxonomies remains a formidable challenge\. The internal states of state\-of\-the\-art models are notoriously complex, making it difficult to isolate distinct, human\-interpretable concepts from dense, high\-dimensional spaces\(Cunninghamet al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib9); Templetonet al\.,[2026](https://arxiv.org/html/2609.19150#bib.bib10)\)\.
In this work, we ask whether the stylistic dimensions most salient for a given prompt can be discovered directly from a model’s activations, with no supervision at any stage\. Our pipeline repeatedly samples completions of a single prompt at elevated temperature, applies Principal Component Analysis \(PCA\) to the pooled hidden activations, and labels the surviving components with an LLM judge\. The result is a training\-free, prompt\-conditional method that surfaces the few stylistic axes dominating the model’s own decoding variance for that instruction\.
To rigorously evaluate the semantic coherence and interpretability of the discovered axes, we conducted a two\-phase human evaluation\. In the first phase, annotators were asked to spontaneously envision ideal stylistic controls for a given prompt, and our unsupervised pipeline recovered those requested dimensions with strong agreement on our strongest model\. In the second phase, human judges evaluated text generated at the geometric extremes of the discovered axes and found that the pole generations closely embodied the automatically generated labels, with 75\.6% of ratings marking them as accurate or highly accurate and 90\.9% adjacent inter\-annotator agreement\. Together, these results suggest that the discovered axes do not merely capture statistical regularities, but correspond to representations that are semantically meaningful and legible to human annotators\.
Crucially, applying our framework across various architectures revealed a surprising representational phenomenon, which we term structural entanglement\. Specifically, we observed a profound degree of latent entanglement in the DeepSeek model we evaluate\. In this model, stylistic variance exhibits strong, inextricable coupling with structural and syntactic variance in the leading components of the representation space\. This entanglement challenges the assumption that stylistic and task variance are linearly separable, highlighting differences in how distinct training regimes or architectures organize their internal representations\.
Our main contributions are threefold:
1. 1\.A training\-free, prompt\-conditional method for discovering stylistic axes from LLM activations via PCA and automatic labeling\.
2. 2\.A two\-phase human evaluation showing that the discovered axes align with spontaneous human intent and are judged semantically valid\.
3. 3\.An empirical characterization of model\-specific sensitivity, including a failure case inDeepSeek\-7B\-Chatwhere structural variance overwhelms stylistic variance\.
## 2Methodology
This section details our unsupervised pipeline for discovering and validating latent stylistic axes within Large Language Models \(LLMs\)\. Our approach, theLatent Engine\(illustrated in Figure[1](https://arxiv.org/html/2609.19150#S2.F1)\), moves beyond supervised contrastive pairs\(Rimskyet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib2); Konenet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib8)\)by leveraging internal model geometry to surface human\-salient concepts\.
Activation HarvestingLatent Component SelectionSemantic GroundingPPBase PromptC=\{p1,…,pN\}C=\\\{p\_\{1\},\\dots,p\_\{N\}\\\}Variation Cloud\(N=30N=30\)Stochastic Generation\(Auto\-Mutation\)Target Model\(e\.g\., Qwen\-3\.5\)Penultimate LayerActivationsMatrixA∈ℝN×dA\\in\\mathbb\{R\}^\{N\\times d\}Sequence\-Avg PoolingPCA DecompositionRetainPC1,PC2PC\_\{1\},PC\_\{2\}Llama\-3 JudgeAuto\-LabelingInterpretable Axis\[−\]\[\-\]Sarcasm\[\+\]\[\+\]De\-noised AxesPolarityCheck
Figure 1:TheLatent Enginearchitecture for Unsupervised Concept Discovery\. The pipeline isolates stylistic variance through stochastic manifold generation into a variation cloud \(N=30N=30\), extracts representations from the penultimate hidden layer, mathematically isolates the primary axes of variance via PCA by extracting the top two principal components \(PC1,PC2PC\_\{1\},PC\_\{2\}\), and autonomously labels these orthogonal components using a lightweight linguistic judge\.### 2\.1Stochastic Manifold Generation \(Variation Clouds\)
To discover latent dimensions without supervised contrastive pairs or explicit stylistic prompting, we isolate the natural geometric variations of a model’s activations during stochastic decoding\. For a given base promptxx, the system generates a “variation cloud”C=\{p1,p2,…,pN\}C=\\\{p\_\{1\},p\_\{2\},\\dots,p\_\{N\}\\\}whereN=30N=30\.
- •Autoregressive Stochasticity:Unlike methods that rely on explicit stylistic mutators, our framework generates variations by samplingNNindependent trajectories from the base model using a single static input prompt\.
- •Sampling Parameters:To maximize stylistic exploration without degrading token\-level coherence, we utilize elevated temperature settings \(T=0\.9T=0\.9,p=0\.95p=0\.95\)\. This ensures that the latent axes emerge naturally from the model’s own stochastic decoding distribution, capturing the model’s native stylistic manifold for that specific instruction\.
### 2\.2Activation Harvesting & PCA Engine
We extract the internal representations of the target model as it processes the variation cloudCC\.
- •Target Layers:Activations are harvested dynamically from the penultimate hidden layer to maintain consistent relative representational depth across varying architectures, thereby avoiding the structural noise prevalent in earlier layers and the final vocabulary\-projection biases of the ultimate layer\.
- •Activation Harvesting Window:For each variationpi∈Cp\_\{i\}\\in C, we concatenate the static base promptxxwith the stochastically generated response tokens\. We then execute a forward pass to extract the hidden states specifically across the generated sequence window \(from the first generated response token to the final token prior to the end\-of\-sequence identifier\)\. This guarantees that our extraction framework isolates downstream token choices and stylistic variance rather than deterministic prompt representations\.
- •Sequence Pooling:We apply sequence\-average pooling over the lengthLiL\_\{i\}of the generated response token hidden states for variantii\. This compresses the dynamic token sequence matrix into a single sequence\-averaged activation vector𝐡i∈ℝd\\mathbf\{h\}\_\{i\}\\in\\mathbb\{R\}^\{d\}, whereddis the hidden dimension of the model\. Stacking these vectors across all generations yields the target variation matrix𝐀∈ℝN×d\\mathbf\{A\}\\in\\mathbb\{R\}^\{N\\times d\}\.
- •Latent Extraction:We perform Principal Component Analysis \(PCA\) on the target variation matrix𝐀\\mathbf\{A\}to identify the orthogonal axes capturing the greatest variance within the style manifold\.
### 2\.3Latent Component Isolation
Following the PCA decomposition of the variation manifold, the pipeline isolates the top two principal components \(PC1,PC2PC\_\{1\},PC\_\{2\}\)\. By definition, standard PCA yields strictly orthogonal singular vectors ordered by their explained variance\. By retaining these primary components, the Latent Engine successfully isolates the most dominant, linearly independent axes of variance that emerged naturally during the stochastic decoding process\. These primary vectors capture the core geometric directions of the model’s native generation space, surfacing highly targeted latent vectors for subsequent semantic evaluation and surgical stylistic manipulation\.
### 2\.4Auto\-Labeling and Polarity Alignment
To make the discovered axes interpretable, the pipeline synthesizes “pole generations” by applying the extracted principal components as activation steering vectors during a new forward pass at fixed intervention strengths \(α=±0\.6\\alpha=\\pm 0\.6\)\. These synthesized extremes are then evaluated in an autonomous grounding step:
- •Linguistic Judge:A quantized Meta\-Llama\-3\-8B\-Instruct model examines generations at the extremes of each PCA axis and generates a single, concise label \(e\.g\., “Formality,” “Sarcasm”\)\. Crucially, this labeling process was conducted completely blind; the judge was provided exclusively with the raw text strings and was stripped of any metadata identifying the target model’s family or architecture\.
- •Dynamic Polarity Alignment:Because PCA axes possess arbitrary sign orientations, we utilize the same LLM judge to establish intuitive polarity\. The judge evaluates the extreme generations to classify which text more strongly embodies the newly discovered label\. The axis is then automatically oriented so that the positive direction \(\+X\+X\) consistently corresponds to an increase in the labeled attribute\.
### 2\.5Target Architectures
To evaluate the universality of this latent organization, we port the pipeline across four distinct model families: Qwen\-3\.5\(Huiet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib22)\), Qwen\-2\.5\(Team,[2026](https://arxiv.org/html/2609.19150#bib.bib23)\), Llama\-3\.2\(Grattafioriet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib24)\), and DeepSeek\(Liuet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib25)\)\. This selection allows us to measureLatent Entanglement—the degree to which stylistic variance is natively separable from strict structural and formatting priors across different alignment regimes\.
### 2\.6Grounded Geometric Semantic Evaluation
Finally, we validate the discovered axes through a two\-phase human evaluation protocol:
1. 1\.Phase 1 \(Spontaneous Recall\):Annotators list the “magic sliders” they desire for a given prompt, testing if the system’s discovered axes match spontaneous human intent\. To compute precision and recall against these unconstrained free\-text responses, we embed the labels utilizingall\-mpnet\-base\-v2\(Songet al\.,[2020](https://arxiv.org/html/2609.19150#bib.bib32)\)and apply a semantic matching threshold ofτ=0\.65\\tau=0\.65to reliably capture stylistic synonyms without introducing false positives\. To mathematically account for this linguistic variance, we anchor our baseline semantic alignment threshold atτ=0\.65\\tau=0\.65\. While a naive evaluation framework might demand near\-perfect embedding similarity \(τ≥0\.85\\tau\\geq 0\.85\) between the model’s auto\-generated label and the human target concept, empirical observation of the latent manifolds reveals a persistenthuman\-machine lexical gap\. Models frequently isolate stylistic control vectors that are geometrically coherent but exhibit slight lexical drift from typical human responses\. For instance, an unsupervised axis that a human annotator classifies asSarcasmis frequently labeled by the autonomous pipeline asSnarkinessorCynicism\(cos\(θ\)≈0\.68\\text\{cos\}\(\\theta\)\\approx 0\.68\)\. Similarly, vectors capturingPragmatismoften surface under the adjacent label ofDirectness\. Demanding exact token matches heavily penalizes the system for identifying valid, contiguous semantic manifolds that simply utilize adjacent vocabulary blocks\. Qualitative manual audits confirm thatτ=0\.65\\tau=0\.65strikes an optimal mathematical elbow: it is sufficiently permissive to capture legitimate synonymic variants while remaining strictly bounded to prevent the cross\-contamination of disjoint concepts \(e\.g\., preventingHumorfrom erroneously merging intoAggression\)\. Under this baseline configuration, Precision is evaluated against theD=40D=40auto\-generated dimensions \(20prompts×2components20\\text\{ prompts\}\\times 2\\text\{ components\}\), while Macro\-Recall is computed against the total cumulative pool ofM=245M=245unique human\-elicited ground\-truth expressions \(averaging a pooled total of 12\.25 traits across all annotators per task framework, whereas individual annotators averaged 3\.9 distinct stylistic requests per prompt\)\. To ensure that our downstream evaluations are not an artifact of an arbitrary “magic number,” we report a comprehensive sensitivity sweep \(τ∈\[0\.40,1\.00\]\\tau\\in\[0\.40,1\.00\]\), Table[4](https://arxiv.org/html/2609.19150#A5.T4)in Appendix[E](https://arxiv.org/html/2609.19150#A5), demonstrating that our chosen parameters settle securely within a stable performance plateau\.
2. 2\.Phase 2 \(Polar Validity\):Annotators rate the accuracy of the system’s auto\-generated labels against actual polar text generations \(−X\-Xvs\+X\+X\) on a 1–5 Likert scale\. We evaluate annotator consensus using Mean Exact Agreement and Mean Adjacent Agreement \(where consensus is achieved if annotator ratings fall within±1\\pm 1point of each other on the Likert scale\)\. This effectively bypasses the ordinal skew paradox \(the paradox of high agreement\) common in traditional reliability metrics\.
## 3Results
Evaluation PhaseMetricScoreInterpretationPhase 1: Spontaneous Recall\(Unconstrained\)Human Semantic IAA47\.2%Baseline semantic consensus among human annotators\.System Precision \(Qwen\-3\.5\)72\.8%When the system extracts a concept, 72\.8% of the time it exactly maps to a human\-desired trait\.System Macro\-Recall43\.6%The unsupervised pipeline recovers nearly half of all uniquely requested human concepts zero\-shot\.Phase 2: Grounded Validity\(Constrained\)Global Median Score4\.0 / 5\.0The typical system\-discovered axis is rated “Mostly Accurate\.” \(IQR: 1\.0\)Top\-2 Box Accuracy75\.6%\>75% of all extreme polar generations were rated highly accurate to their auto\-generated label\.Statistical Significancep<0\.001p<0\.001Wilcoxon Signed\-Rank Test confirms axes strictly outperform a neutral baseline\.Mean Exact Agreement57\.5%Absolute majority of annotators selected the exact same 1\-5 rating cell\.Adjacent Agreement90\.9%9 out of 10 annotator ratings land within±1\\pm 1point, establishing near\-universal human consensus\.Table 1:Human Grounding and Geometric Validity of Unsupervised Latent Axes\. Precision and Recall metrics for Qwen\-3\.5 represent the robust macro\-average across 9 independent stochastic seeds \(N=30N=30prompt variations per seed\)\. Meta\-Llama\-3\-8B\-Instruct was utilized as the semantic labeler across all runs\.### 3\.1Human Validation and Grounded Geometry
Our evaluation isolates the effectiveness of the unsupervised latent discovery pipeline by testing against spontaneous human desire \(Phase 1\) and grounded geometric validity \(Phase 2\)\. Before examining model\-specific differences \(detailed in Section 3\.3\), we first report the aggregate baseline performance across our primary evaluation architecture, summarized in Table[1](https://arxiv.org/html/2609.19150#S3.T1)\.
In Phase 1 \(Unconstrained Generation\), annotators evaluated whether the autonomously discovered latent axes matched their spontaneous conceptualization of stylistic control\. Human annotators demonstrated a highly consistent semantic Inter\-Annotator Agreement \(IAA\) of 47\.2%, establishing a stable ground\-truth baseline for human vocabulary alignment\. Because our variation clouds are generated using stochastic decoding, all reported system metrics represent the macro\-average across 9 independent experimental seeds to ensure robust reproducibility\. Against this baseline, our top\-performing architecture \(Qwen\-3\.5\) achieved a robust72\.8% mean Precision, indicating that nearly three\-quarters of the automatically discovered latent axes perfectly mapped onto verified, human\-desired dimensions\. Additionally, we recorded a Phase 1mean Macro Recall of 43\.6%\. This is an exceptionally high recovery rate given that the theoretical maximum recall for a Top\-2 axis extraction is approximately 51\.3%111The theoretical maximum recall is constrained by the structural limits of the top\-kkevaluation setting\. In our methodology, the unsupervised pipeline was restricted to surfacing only the top two most salient principal components \(k=2k=2\) per prompt\. In contrast, the open\-ended human annotation phase yielded an average of 3\.9 distinct stylistic dimensions per prompt\. Because the system can retrieve at most 2 items against an average ground\-truth pool of 3\.9 items, the absolute mathematical upper bound for recall is exactly2/3\.9≈51\.3%2/3\.9\\approx 51\.3\\%\. Therefore, our achieved recall of 43\.6% captures roughly 85% of the available theoretical capacity \(43\.6/51\.343\.6/51\.3\) due to the vast diversity of unconstrained human wording\.\.
In Phase 2 \(Polar Validity\), annotators judged the semantic validity of the generated extremes on a 1–5 Likert scale\. To accurately match auto\-generated axis labels with human intents as “stylistic synonyms” prior to evaluation, we applied an embedding cosine\-similarity threshold ofτ=0\.65\\tau=0\.65\. Under this threshold constraint, the aggregate results demonstrate exceptional alignment: the system achieved aGlobal Median Score of 4\.0 out of 5\.0\(IQR=1\.0IQR=1\.0,Q25=4\.0Q\_\{25\}=4\.0,Q75=5\.0Q\_\{75\}=5\.0\)\. Furthermore, the pipeline recorded aTop\-2 Box Accuracy of 75\.6%, meaning more than three\-quarters of all human ratings judged the auto\-discovered axes as either “Mostly Accurate” \(4\) or “Highly Accurate” \(5\)\.
For a qualitative showcase of the auto\-discovered stylistic axes, including unedited generation snippets from the geometric extremes \(−X,\+X\-X,\+X\) of the latent manifold, please refer to Appendix[D](https://arxiv.org/html/2609.19150#A4)\.
This positive alignment is statistically strong \(Wilcoxon Signed\-Rank Test comparing the ordinal ratings against a neutral null\-hypothesis baseline of 3\.0,p=1\.493×10−42p=1\.493\\times 10^\{\-42\}\) \(see Figure[4](https://arxiv.org/html/2609.19150#A6.F4)in Appendix[F](https://arxiv.org/html/2609.19150#A6)\), indicating a very low probability that these dimension alignments represent random noise\. Finally, the system achieved an Exact Mode Agreement of 57\.5% and an exceptional90\.9% Adjacent Agreement222Metric Definitions:Adjacent Agreementis an inter\-rater reliability metric for ordinal scales that considers scores within one point \(e\.g\., a 4 and a 5\) as consensus, thereby preventing penalization for minor subjective differences in intensity judgments\.\. This confirms that the PCA\-derived axes genuinely track the targeted semantic properties—establishing a strong, statistically significant baseline of human agreement before introducing model\-specific variations\.
Figure 2:Human\-Baseline Alignment Across Architectures\.Macro Precision and Recall scores for PCA\-discovered stylistic axes evaluated against the Phase 1 human spontaneous recall baseline\. While Qwen and Llama architectures successfully expose human\-salient latent controls, DeepSeek\-7B exhibits a severe interpretability collapse, yielding only 35\.3% precision\.
### 3\.2Stochastic Stability of Discovered Axes
A key concern in unsupervised representation engineering is whether the discovered latent geometries are stable mechanisms or stochastic artifacts of the generation process\. We performed repeated extractions across9 experimental seedsfor the variation clouds, yielding a total of180 extracted axes\. Crucially, these variation clouds were generated strictly through elevated temperature sampling \(TT=0\.90\) from the static base prompt, free of any external mutators\. Across these conditions, theLatent Enginedemonstrated high robustness, achieving an82\.5% Top\-2 coverage rediscovery rate\. This indicates that the primary stylistic dimensions in the manifold are structurally entrenched in the model’s internal representations \(specifically, within the targeted penultimate hidden layer\) rather than transient noise\.
### 3\.3Cross\-Architecture Discoverability and the DeepSeek Anomaly
Applying the Latent Engine across our target architectures revealed highly model\-dependent discoverability\. The Qwen models demonstrated superior human\-baseline alignment, maintaining a precision of approximately 70%\.meta\-llama\-3\.2\-3Bshowed moderate but reliable extraction capabilities\.
However,deepseek\-7b\-chatpresented a significant anomaly\. While the Qwen models successfully isolated clean stylistic axes,deepseek\-7b\-chatexperienced a severe performance drop, yielding only∼\\sim35% precision\. Qualitative analysis of the extracted axes for DeepSeek revealed severe entanglement between stylistic manipulation and the model’s rigid structural and formatting traces\.
## 4Discussion: The DeepSeek Performance Drop
Figure 3:Orthogonal Disentanglement vs\. Latent Entanglement\.PCA projection of penultimate layer activations generated via Stochastic Manifold Generation\.\(Left\)Qwen\-3\.5 cleanly separates the stochastically discovered concepts within the primary principal components, enabling the orthogonal isolation of steering vectors\.\(Right\)DeepSeek\-7B fails to linearly separate the variations, visually demonstrating theStructural Entanglementphenomenon, where stylistic variance and rigid syntactic priors remain inextricably linked\.The sharp degradation of unsupervised stylistic extraction indeepseek\-7b\-chathighlights a fundamental architectural divergence in how certain models encode concepts, a phenomenon we term theLatent Entanglement of Style and Structure\. We hypothesize that this performance drop is a direct consequence of a representational phenomenon we termStructural Entanglement\. Early DeepSeek architectures, such asdeepseek\-7b\-chat, were heavily optimized for strict instruction formatting, bilingual alignment, and structural compliance\. When our pipeline generates the “variation cloud” of prompts to isolate style,deepseek\-7b\-chataggressively allocates its highest\-variance latent capacity to maintaining these rigid structural boundaries and output formats rather than exploring stylistic fluency\. Consequently, structural and syntactic variance geometrically dominates the primary principal components \(PC1PC\_\{1\}andPC2PC\_\{2\}\)\. Because our unsupervised extraction strictly isolates these top two components for semantic labeling, the true stylistic variance is crowded out and pushed into much lower, unextracted singular dimensions\. The auto\-labeling judge evaluates the structural variance captured in the primary components and consequently assigns format\-focused or incoherent descriptors rather than human\-aligned stylistic ones, causing Precision@2 to plummet to 35\.3%\. This finding suggests that certain alignment regimes natively intertwine “how to format it” with “how to say it\.” This raises critical questions about the limits of zero\-shot representation engineering using linear methods in architectures with heavily entrenched structural priors, suggesting that future steering methods for these models may require non\-linear or layer\-specific manifold learning strategies\. Analogous to recent interpretability analyses demonstrating that specialized cognitive priors occupy substantial representational capacity in later architectures\(Galichin and others,[2025](https://arxiv.org/html/2609.19150#bib.bib33)\), the structural dominance we observe suggests this entanglement reflects the model’s strict alignment regime rather than an artifact of our pipeline\.
## 5Related Work
#### Steering along known concepts\.
Considerable evidence suggests that high\-level concepts are encoded as linear directions in LLM activation space\(Parket al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib6); Tiggeset al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib5)\), though some features are irreducibly multi\-dimensional\(Engelset al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib27)\)\. Activation steering exploits this structure by adding concept vectors to hidden states at inference time, deriving them from contrastive prompt pairs\(Turneret al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib3); Zouet al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib4); Rimskyet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib2)\), style\-labeled corpora\(Konenet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib8)\), or optimization toward target completions\(Subramaniet al\.,[2022](https://arxiv.org/html/2609.19150#bib.bib1)\)\. Extensions condition steering on the instruction or input\(Stolfoet al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib16); Leeet al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib26)\), compose multiple stylistic properties\(Scalenaet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib18)\), and personalize outputs to user preferences\(Caoet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib17); Boet al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib7)\), though steering reliability varies substantially across concepts and inputs\(Tanet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib15)\); seeWehneret al\.\([2025](https://arxiv.org/html/2609.19150#bib.bib31)\)for a survey\. All of these methods presuppose that the concept of interest is specified in advance, and they remain the natural choice in that setting\. We study the complementary, upstream problem: identifying which stylistic dimensions are salient in a model’s generations for a given prompt, without labels, curated corpora, or a predefined taxonomy\. Our pipeline requires no supervision at any stage; even the polarity of each discovered axis is established automatically by the labeling judge\.
#### Unsupervised feature discovery\.
Two lines of work discover features without predefined concepts\. Sparse autoencoders decompose activations into large dictionaries of interpretable features\(Cunninghamet al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib9); Templetonet al\.,[2026](https://arxiv.org/html/2609.19150#bib.bib10)\), but they require corpus\-scale training, and recent evaluations question whether their features are canonical units or useful for downstream control\(Wuet al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib12); Kantamneniet al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib13); Leasket al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib14); Jørgensen and Hansen,[2026](https://arxiv.org/html/2609.19150#bib.bib30)\)\. SAE features have also been used to quantify how far representations are universal across model families\(Lanet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib28)\), a question our cross\-model analysis revisits at the level of stylistic axes\. Closer to our setting,Macket al\.\([2026](https://arxiv.org/html/2609.19150#bib.bib11)\)elicit latent behaviors by optimizing unsupervised perturbations to a model’s activations\. Our approach differs from both: it is training\-free and prompt\-conditional, decomposing the model’s own decoding variance via PCA to surface the few stylistic axes that dominate variation under a given instruction, rather than learning a global feature dictionary or optimizing for behavioral change\.
#### Automated labeling and its validation\.
A final line of work uses LLMs to generate and score natural\-language explanations of model internals at scale\(Templetonet al\.,[2026](https://arxiv.org/html/2609.19150#bib.bib10); Pauloet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib29)\)\. Such labels require careful validation, as generated explanations can appear plausible without being faithful\(Huanget al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib20)\), and LLM judges systematically favor their own generations\(Panicksseryet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib19)\)\. The latter is a direct concern for our pipeline, where a Llama\-family judge labels axes for a Llama target model; we therefore validate labels with a two\-phase human study and discuss residual circularity in the Limitations section\.
## Limitations
While our Unsupervised Concept Discovery pipeline demonstrates strong alignment with human stylistic intent, we acknowledge several limitations in the current framework:
#### Assumption of Linear Separability
The Latent Engine relies on Principal Component Analysis \(PCA\), which inherently assumes that stylistic variance and structural variance are linearly separable in the model’s representation space\(Engelset al\.,[2025](https://arxiv.org/html/2609.19150#bib.bib27)\)\. As demonstrated by the “DeepSeek Anomaly,” this assumption fails for architectures heavily optimized for rigid structural and syntactic compliance, resulting in latent entanglement\. Extending this unsupervised discovery to models with heavily entrenched formatting priors will likely require non\-linear manifold learning techniques \(e\.g\., Kernel PCA or Sparse Autoencoders\(Cunninghamet al\.,[2023](https://arxiv.org/html/2609.19150#bib.bib9)\)\) to successfully disentangle features\.
#### Dependence on the Linguistic Judge
The autonomous labeling phase is bottlenecked by the semantic capacity of the quantizedMeta\-Llama\-3\-8B\-Instructjudge model\. If the judge lacks the vocabulary to accurately describe a highly nuanced, purely geometric latent axis, it may assign a reductive or proximate label\. While we mitigated this using an embedding synonym threshold \(τ=0\.65\\tau=0\.65\), the framework remains constrained by the teacher model’s linguistic ceiling and potential inherent biases\. Moreover, the judge shares a model family with one evaluated target \(meta\-llama\-3\.2\-3B\)\. While we implemented a strictly blind evaluation protocol to prevent explicit bias, LLM evaluators can still implicitly recognize and systematically favor outputs of their own family\(Panicksseryet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib19)\)\. Consequently, judge\-assigned labels for Llama\-family axes may still be slightly inflated; however, our two\-phase human validation mitigates this circularity\.
#### Confounding Structural Proxies
Our unsupervised extraction does not explicitly control for low\-level structural covariates, such as response length, lexical diversity, or formatting markers \(e\.g\., bullet points\)\. Because PCA isolates the directions of maximum variance, it is possible that some discovered principal components partially align with these simpler text\-surface proxies rather than purely abstract stylistic shifts\.
#### Dimensional Truncation and Recall Bounds
To prioritize precision and interpretability, our pipeline isolates only the two highest\-variance components \(PC1PC\_\{1\}andPC2PC\_\{2\}\) from the latent manifold\. However, human stylistic desire is vastly open\-ended\. Restricting the extraction to a strict top\-kktruncation \(k=2k=2\) mathematically caps the maximum possible recall \(bounded at approximately 51\.3% in our evaluation setup, as annotators averaged 3\.9 distinct stylistic requests per prompt\)\. Consequently, many valid but lower\-variance stylistic features present in the model’s manifold are inevitably left undiscovered by this specific configuration\. Future work could explore dynamic\-variance thresholding to determine an optimal, variable number of components to extract per prompt rather than utilizing a fixed cutoff\.
#### Discovery versus Steering Efficacy
Our evaluation validates the semantic interpretability of discovered axes, not their efficacy as steering interventions\. In preliminary experiments, supervised contrastive vectors\(Rimskyet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib2)\)produced stronger steering effects than our PCA\-derived directions; our contribution is surfacing salient axes without supervision, not matching supervised steering strength\.
#### Scope of Evaluation
Finally, our empirical evaluations were conducted exclusively on English\-language generative tasks and focused strictly on stylistic, rhetorical, and affective dimensions of control\. The efficacy of stochastic manifold generation for discovering localized factual knowledge, cross\-lingual representations, or strict safety\-alignment concepts remains an open question for future research\.
## Conclusion
In this work, we introduced a fully unsupervised framework for discovering human\-salient stylistic axes within Large Language Models\. By leveraging stochastic manifold generation and Principal Component Analysis, we demonstrated that interpretable stylistic axes can be extracted from a single prompt without any supervision\. Our human evaluation shows that the discovered axes align with unconstrained human intent and that their polar generations are judged semantically valid by annotators\. Overall, our results suggest that prompt\-conditional stylistic structure is linearly accessible with remarkably simple tools\. A natural next step is to close the loop between discovery and control: pairing unsupervised discovery with light supervised contrastive vectors\(Rimskyet al\.,[2024](https://arxiv.org/html/2609.19150#bib.bib2)\)of the surfaced axes is a promising route to stylistic control without predefined taxonomies\.
## References
- Steerable chatbots: personalizing llms with preference\-based activation steering\.arXiv preprint arXiv:2505\.04260\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- Y\. Cao, T\. Zhang, B\. Cao, Z\. Yin, L\. Lin, F\. Ma, and J\. Chen \(2024\)Personalized steering of large language models: versatile steering vectors through bi\-directional preference optimization\.Advances in Neural Information Processing Systems37,pp\. 49519–49551\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2023\)Sparse autoencoders find highly interpretable features in language models\.arXiv preprint arXiv:2309\.08600\.Cited by:[§1](https://arxiv.org/html/2609.19150#S1.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1),[Assumption of Linear Separability](https://arxiv.org/html/2609.19150#Sx1.SS0.SSS0.Px1.p1.1)\.
- J\. Engels, E\. Michaud, I\. Liao, W\. Gurnee, and M\. Tegmark \(2025\)Not all language model features are one\-dimensionally linear\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 84591–84622\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1),[Assumption of Linear Separability](https://arxiv.org/html/2609.19150#Sx1.SS0.SSS0.Px1.p1.1)\.
- A\. Galichinet al\.\(2025\)I have covered all the bases here: interpreting reasoning features in large language models via sparse autoencoders\.arXiv preprint arXiv:2503\.18878\.Cited by:[§4](https://arxiv.org/html/2609.19150#S4.p1.2)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§2\.5](https://arxiv.org/html/2609.19150#S2.SS5.p1.1)\.
- J\. Huang, A\. Geiger, K\. D’Oosterlinck, Z\. Wu, and C\. Potts \(2023\)Rigorously assessing natural language explanations of neurons\.InProceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 317–331\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px3.p1.1)\.
- B\. Hui, J\. Yang, Z\. Cui, J\. Yang, D\. Liu, L\. Zhang, T\. Liu, J\. Zhang, B\. Yu, K\. Lu,et al\.\(2024\)Qwen2\. 5\-coder technical report\.arXiv preprint arXiv:2409\.12186\.Cited by:[§2\.5](https://arxiv.org/html/2609.19150#S2.SS5.p1.1)\.
- M\. G\. Jørgensen and L\. K\. Hansen \(2026\)Steering llms? actually, sparse autoencoders can outperform simple baselines\.arXiv preprint arXiv:2605\.31183\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- S\. Kantamneni, J\. Engels, S\. Rajamanoharan, M\. Tegmark, and N\. Nanda \(2025\)Are sparse autoencoders useful? a case study in sparse probing\.arXiv preprint arXiv:2502\.16681\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- K\. Konen, S\. Jentzsch, D\. Diallo, P\. Schütt, O\. Bensch, R\. El Baff, D\. Opitz, and T\. Hecking \(2024\)Style vectors for steering generative large language models\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 782–802\.Cited by:[§2](https://arxiv.org/html/2609.19150#S2.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- M\. Lan, P\. Torr, A\. Meek, A\. Khakzar, D\. Krueger, and F\. Barez \(2024\)Quantifying feature space universality across large language models via sparse autoencoders\.arXiv preprint arXiv:2410\.06981\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- P\. Leask, B\. Bussmann, M\. Pearce, J\. Bloom, C\. Tigges, N\. Al Moubayed, L\. Sharkey, and N\. Nanda \(2025\)Sparse autoencoders do not find canonical units of analysis\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 53617–53642\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- B\. W\. Lee, I\. Padhi, K\. Natesan Ramamurthy, E\. Miehling, P\. Dognin, M\. Nagireddy, and A\. Dhurandhar \(2025\)Programming refusal with conditional activation steering\.InInternational conference on learning representations,Vol\.2025,pp\. 90960–90985\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§2\.5](https://arxiv.org/html/2609.19150#S2.SS5.p1.1)\.
- A\. Mack, N\. Panickssery, and A\. M\. Turner \(2026\)Mechanistically eliciting latent behaviors in language models\.arXiv preprint arXiv:2606\.29604\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- A\. Panickssery, S\. R\. Bowman, and S\. Feng \(2024\)Llm evaluators recognize and favor their own generations\.Advances in Neural Information Processing Systems37,pp\. 68772–68802\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px3.p1.1),[Dependence on the Linguistic Judge](https://arxiv.org/html/2609.19150#Sx1.SS0.SSS0.Px2.p1.1)\.
- K\. Park, Y\. J\. Choe, and V\. Veitch \(2023\)The linear representation hypothesis and the geometry of large language models\.arXiv preprint arXiv:2311\.03658\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- G\. Paulo, A\. Mallen, C\. Juang, and N\. Belrose \(2024\)Automatically interpreting millions of features in large language models\.arXiv preprint arXiv:2410\.13928\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px3.p1.1)\.
- N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. Turner \(2024\)Steering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15504–15522\.Cited by:[§2](https://arxiv.org/html/2609.19150#S2.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1),[Discovery versus Steering Efficacy](https://arxiv.org/html/2609.19150#Sx1.SS0.SSS0.Px5.p1.1),[Conclusion](https://arxiv.org/html/2609.19150#Sx2.p1.1)\.
- D\. Scalena, G\. Sarti, and M\. Nissim \(2024\)Multi\-property steering of large language models with dynamic activation composition\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 577–603\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- K\. Song, X\. Tan, T\. Qin, J\. Lu, and T\. Liu \(2020\)Mpnet: masked and permuted pre\-training for language understanding\.arXiv preprint arXiv:2004\.09297\.Cited by:[item 1](https://arxiv.org/html/2609.19150#S2.I4.i1.p1.1)\.
- A\. Stolfo, V\. Balachandran, S\. Yousefi, E\. Horvitz, and B\. Nushi \(2025\)Improving instruction\-following in language models through activation steering\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 55790–55823\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- N\. Subramani, N\. Suresh, and M\. E\. Peters \(2022\)Extracting latent steering vectors from pretrained language models\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 566–581\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- D\. Tan, D\. Chanin, A\. Lynch, B\. Paige, D\. Kanoulas, A\. Garriga\-Alonso, and R\. Kirk \(2024\)Analysing the generalisation and reliability of steering vectors\.Advances in Neural Information Processing Systems37,pp\. 139179–139212\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- Q\. Team \(2026\)Qwen3\. 5\-omni technical report\.arXiv preprint arXiv:2604\.15804\.Cited by:[§2\.5](https://arxiv.org/html/2609.19150#S2.SS5.p1.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones,et al\.\(2026\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.arXiv preprint arXiv:2605\.29358\.Cited by:[§1](https://arxiv.org/html/2609.19150#S1.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px3.p1.1)\.
- C\. Tigges, O\. J\. Hollinsworth, A\. Geiger, and N\. Nanda \(2024\)Language models linearly represent sentiment\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 58–87\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid \(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- J\. Wehner, S\. Abdelnabi, D\. Tan, D\. Krueger, and M\. Fritz \(2025\)Taxonomy, opportunities, and challenges of representation engineering for large language models\.arXiv preprint arXiv:2502\.19649\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
- Z\. Wu, A\. Arora, A\. Geiger, Z\. Wang, J\. Huang, D\. Jurafsky, C\. D\. Manning, and C\. Potts \(2025\)Axbench: steering llms? even simple baselines outperform sparse autoencoders\.arXiv preprint arXiv:2501\.17148\.Cited by:[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px2.p1.1)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.\(2023\)Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§1](https://arxiv.org/html/2609.19150#S1.p1.1),[§5](https://arxiv.org/html/2609.19150#S5.SS0.SSS0.Px1.p1.1)\.
## Appendix APrompt Templates for Auto\-Labeling
To enable the quantized Meta\-Llama\-3\-8B\-Instruct judge to autonomously label the discovered latent components, we utilized a strict, zero\-shot comparative prompt\. This prompt isolates the extracted geometric extremes \(−X\-Xand\+X\+Xgenerations\) and restricts the judge to outputting a concise stylistic descriptor without extraneous reasoning\. The exact template is detailed below\.
System Prompt:You are an expert linguistic analyst evaluating text\. Identify the SINGLE most prominent stylistic difference between Text 1 and Text 2\.RULES:1\. DO NOT name the subject matter\.2\. Output EXACTLY ONE WORD representing the style\.3\. NO conversational filler\. NO sentences\.User Prompt:Text 1 \(Negative Pole\):\[\-X GENERATION\]Text 2 \(Positive Pole\):\[\+X GENERATION\]
## Appendix BImplementation Details
#### Hardware & Compute Infrastructure
Due to strict compute resource constraints, all stochastic manifold generation, PCA extractions, and model evaluation pipelines were executed utilizing a Kaggle dual NVIDIA T4 GPU \(2x16GB VRAM\) environment\. These hardware limitations necessitated our highly targeted extraction approach and the use of native 4\-bit quantization for the labeling judge\.
#### Software & Libraries
TheLatent Engineframework was implemented in Python 3\.10 and relies on the following core libraries for distributed inference and geometric decomposition: PyTorch, HuggingFace Transformers,scikit\-learnfor Principal Component Analysis, andbitsandbytesfor native 4\-bit quantization\.
#### Model Configuration & Quantization
To maximize computational throughput within the dual\-T4 constraints, the pipeline utilizes an asymmetric, two\-model architecture\. The base target model for generation and activation extraction \(Qwen\-3\.5\) operates in native 16\-bit floating\-point precision \(FP16\)\. The teacher labeler model \(Meta\-Llama\-3\-8B\-Instruct\) was loaded alongside the base model using 4\-bit NormalFloat \(NF4\) quantization\.
#### Hyperparameters & Reproducibility
For the stochastic manifold generation phase, target architectures were sampled using an elevated temperature ofT=0\.90T=0\.90with nucleus sampling capped atp=0\.90p=0\.90\. This configuration was chosen empirically to maximize the stylistic distribution of the variation cloud without deteriorating token\-level syntactic coherence\. The variation cloud size was strictly bounded toN=30N=30generations per base prompt\. During extraction, representations were harvested using sequence\-average pooling from Layer 21, avoiding structural representation traps found in earlier layers\.
To account for generation stochasticity, all reported Phase 1 metrics \(Section[3\.1](https://arxiv.org/html/2609.19150#S3.SS1)\) reflect the macro\-average across 9 independent experimental seeds\. The semantic alignment thresholds \(τ=0\.65\\tau=0\.65\) were computed using cosine similarity derived from theall\-mpnet\-base\-v2sentence encoder\.
## Appendix CEvaluation Stimuli
Table[2](https://arxiv.org/html/2609.19150#A3.T2)details the complete evaluation matrix utilized in this study, comprising 20 distinct base prompts equally distributed across four semantic domains\.
Table 2:Stimuli Matrix\.The 20 base prompts utilized for evaluation, organized by domain\.
## Appendix DQualitative Showcase of Discovered Axes
To demonstrate the semantic validity and geometric structure of the discovered axes, Table[3](https://arxiv.org/html/2609.19150#A4.T3)presents unedited generation snippets sampled from the geometric extremes \(−X\-Xand\+X\+X\) of the latent manifold for a representative base prompt\. These examples illustrate how the PCA\-derived axes map to the autonomous labels assigned by the Llama\-3 judge, effectively shifting the output without requiring explicit human vocabulary\.
Table 3:Unedited Polar Generations\.Snippets extracted from the geometric extremes of the top principal components evaluated on Qwen\-3\.5\. The negative pole \(α=−0\.6\\alpha=\-0\.6\) introduces conversational phrasing and contractions, while the positive pole \(α=\+0\.6\\alpha=\+0\.6\) defaults to rigid corporate structures\.
## Appendix EThreshold Sensitivity Analysis
Table 4:Sensitivity sweep over the semantic matching parameterτ\\tau\.The macro precision maps strictly to the discrete hits within a single canonical evaluation dimension cohort \(D=40D=40\)\. While our top\-performing architecture \(Qwen\-3\.5\) achieves a higher aggregate baseline of 72\.8% macro\-averaged across 9 seeds, this specific sample run \(62\.5% atτ=0\.65\\tau=0\.65\) isolates a single seed to profile the smooth mathematical decay of the alignment framework without multi\-seed averaging artifacts\.
## Appendix FExtended Evaluation Metrics
Figure 4:Non\-Parametric Density Profile of Phase 2 Human Evaluations\.A violin plot with an embedded box plot illustrating the distribution of ordinal validity ratings \(N=320N=320\) for the auto\-discovered stylistic axes\. Because Likert data is strictly ordinal, this non\-parametric visualization captures the true central tendency without assuming a normal distribution\. The density profile reveals a severe left\-skew concentrated at the upper bounds of the scale: both the median and the first quartile \(Q25Q\_\{25\}\) are anchored at 4\.0 \(“Mostly Accurate”\), with the third quartile \(Q75Q\_\{75\}\) at 5\.0 \(“Highly Accurate”\)\. The overlaid statistical markers confirm the robustness of this alignment: a Wilcoxon Signed\-Rank test \(p=1\.49×10−42p=1\.49\\times 10^\{\-42\}\) decisively rejects the null hypothesis of neutral or random rating assignment, while Krippendorff’sα=0\.844\\alpha=0\.844indicates a strong inter\-annotator agreement\. This provides granular structural support for the global top\-2 box accuracy reported in Section 3\.1\.Similar Articles
Interpreting Style Representations via Style-Eliciting Prompts
This paper proposes a framework to interpret style representations by using style-eliciting prompts—natural language instructions that steer LLMs to generate text with specific stylistic attributes. The method outperforms baseline LLM prompting techniques in both describing and imitating writing styles.
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
Proposes LiSCP, a lightweight stylistic consistency profiling method for robust detection of LLM-generated textual content, focusing on feature stability under adversarial manipulation. Achieves superior performance on in-domain and cross-domain detection with notable robustness.
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
An unsupervised method called activation-matched finetuning is proposed to detect hidden behaviors in large language models by comparing activations with a reference model, reliably identifying triggers without prior knowledge.
Measuring, Localizing, and Ablating Alignment Signatures in LLMs
This paper investigates how post-training of LLMs introduces AI-like stylistic regularities and proposes PASTA, a training-free method to localize and ablate these alignment signatures, reducing AI detection rates while maintaining coherence across 11 models and 6 detectors.
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
This paper discovers a shared valence axis (V-axis) across modern LLMs and human EEG signals, showing that a single direction from LLM internal representations aligns with neural responses to emotional stimuli. It also identifies the saturation regularity, explaining why LLM-derived supervision fails to improve EEG decoding and how leveraging residual diversity boosts performance.