How smoothing the affinity matrix affects neighborhood preservation in t-SNE

arXiv cs.LG Papers

Summary

This paper studies how smoothing the affinity matrix in t-SNE via a row-wise power transform affects neighborhood preservation, finding that sharpening improves nearest neighbor retention while smoothing enhances broader local neighborhoods.

arXiv:2608.17190v1 Announce Type: new Abstract: Dimensionality reduction methods are instrumental to visualize high-dimensional data, and t-SNE stands as one of the most widely used methods due to its emphasis on local neighborhood preservation. A central component of t-SNE is the affinity matrix, which expresses pairwise similarities in the form of symmetrized probabilities, over which the optimization problem of t-SNE is defined. We study how the sharpness of this probability distribution affects neighborhood preservation at different scales. We introduce a row-wise power transform controlled by a parameter gamma that can smooth or sharpen each row of the affinity matrix while preserving sparsity and rank order. We show that this transform is equivalent to rescaling the Gaussian bandwidth and thus to changing the perplexity. However, as the sharpness of the probability distribution varies per point, a fixed gamma leads to point-dependent effective perplexities, making it distinct from changing the global perplexity. Empirically, we find that sharpening improves preservation of the very nearest neighbors, while smoothing improves preservation of broader local neighborhoods, outperforming alternative affinity constructions including multiscale methods in the mid-local range.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:24 AM

# How smoothing the affinity matrix affects neighborhood preservation in t-SNE
Source: [https://arxiv.org/html/2608.17190](https://arxiv.org/html/2608.17190)
###### Abstract

Dimensionality reduction methods are instrumental to visualize high\-dimensional data, and t\-SNE stands as one of the most widely used methods due to its emphasis on local neighborhood preservation\. A central component of t\-SNE is the*affinity matrix*, which expresses pairwise similarities in the form of symmetrized probabilities, over which the optimization problem of t\-SNE is defined\. We study how the ‘sharpness’ of this probability distribution affects neighborhood preservation at different scales\. We introduce a row\-wise power transform controlled by a parameterγ\\gammathat can smooth or sharpen each row of the affinity matrix while preserving sparsity and rank order\. We show that this transform is equivalent to rescaling the Gaussian bandwidth and thus to changing the perplexity\. However, as the sharpness of the probability distribution varies per point, a fixedγ\\gammaleads to point\-dependent effective perplexities, making it distinct from changing the global perplexity\. Empirically, we find that sharpening improves preservation of the very nearest neighbors, while smoothing improves preservation of broader local neighborhoods, outperforming alternative affinity constructions including multiscale methods in the mid\-local range\.

###### Keywords:

t\-SNE affinity matrix neighborhood preservation

## 1Introduction

Motivation\. Machine learning models are highly effective for inference and prediction, but users often still want to inspect data directly to identify patterns and gain insights that complement model outputs\. Visualization can support this kind of exploration, but this is difficult when the data are high\-dimensional, meaning that each observation is described by many features\. Dimensionality reduction \(DR\) methods mitigate this by projecting the data into a two dimensional \(2D\) space, after which we can visualize the data in a scatter plot\.

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/qiu_local_global_SP.png)

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/TRACE.jpg)

Figure 1:\(a\) Redrawn from Novak et al\. \(2026\)\[[17](https://arxiv.org/html/2608.17190#bib.bib3)\], showing local versus global neighborhood preservation \(here measured as the AUC under the RnX and log\-RnX curve respectively, see their paper for details\) for a range of state\-of\-the\-art DR methods\. One dataset shown here, the paper contains another seven\. As can be seen, t\-SNE manages to retain local structure substantially better than other DR methods\. \(b\) From Heiter et al\. \(2024\)\[[8](https://arxiv.org/html/2608.17190#bib.bib4)\], illustrating on a t\-SNE projection of mouse cell data\[[19](https://arxiv.org/html/2608.17190#bib.bib5)\]that neighborhood preservation strongly varies across the embeddings\. Color corresponds to neighborhood overlap withk=200k=200\.Unfortunately, projecting data into 2D typically means that only part of the original structure is retained\. A substantial amount of research has been conducted on developing effective DR methods for visualizing high\-dimensional data\. Among these methods, t\-distributed Stochastic Neighbor Embedding \(t\-SNE\)\[[21](https://arxiv.org/html/2608.17190#bib.bib1)\]is one of the most widely adopted nonlinear dimensionality reduction techniques\[[7](https://arxiv.org/html/2608.17190#bib.bib2)\]\. A major reason for its widespread use is its focus on local neighborhood structure: points that are close in the high\-dimensional space are encouraged to remain close in the embedding\. t\-SNE is known to preserve local neighborhood structure well\[[17](https://arxiv.org/html/2608.17190#bib.bib3)\], see Figure[1](https://arxiv.org/html/2608.17190#S1.F1)\(a\) for an illustration from the literature\. Although t\-SNE often performs better than other methods on local preservation benchmarks, neighborhood preservation scores are often low in the absolute sense, and errors may also be spread unevenly over the embedding\. See Figure[1](https://arxiv.org/html/2608.17190#S1.F1)\(b\) for an illustration of this effect\. It has been argued repeatedly that low reliability of \(local\) structure preservation creates problems in interpreting apparent structures on the embedding, as they may be created artificially by the DR method, rather than derived from a reliable signal in the data\[[4](https://arxiv.org/html/2608.17190#bib.bib6)\]\.

Aim of the study\. As also shown in Figure[1](https://arxiv.org/html/2608.17190#S1.F1)\(a\), other methods \(e\.g\., PaCMAP\[[23](https://arxiv.org/html/2608.17190#bib.bib8)\], Autoencoders\[[17](https://arxiv.org/html/2608.17190#bib.bib3)\]\) outperform t\-SNE on global structure preservation and most research appears to have focused on improving global structure preservation of t\-SNE\. However, it is equally interesting to ask*whether it would be possible to improve local structure preservation further*, which may for example make the inspection of cluster substructures more reliable and useful\. To the best of our knowledge, there has been no explicit investigation whether it is possible to improve neighborhood preservation of or beyond t\-SNE\. To this end, we study how t\-SNE encodes local neighborhood relations in its high\-dimensional affinity matrix and explore the impact of modifying the affinity scores\.

Methodology\. To better understand what guides t\-SNE’s local preservation behavior, we focus on its high\-dimensional affinity matrix\. t\-SNE first converts high\-dimensional distances into neighborhood probabilities, forming the affinity matrixPP\. The embedding is then optimized so that the low\-dimensional probabilities reproduce these high\-dimensional affinities as closely as possible\. AfterPPhas been constructed, the optimization is driven only by the affinities\. This makesPPa central component of t\-SNE: it encodes which neighborhood relations are emphasized and how strongly they influence the embedding\. Consequently, changes in the affinity matrix can change which local structures are preserved or lost\. In this work, we investigate how changes in the smoothness of this affinity matrix affect t\-SNE’s neighborhood preservation\.

Standard t\-SNE uses point\-specific Gaussian bandwidths to construct the affinity matrix, and these bandwidths are chosen so that every point has the same perplexity\. Thus, t\-SNE adapts the local distance scale while fixing the perplexity globally\. The perplexity is2H⁡\(Pi\)2^\{H\(P\_\{i\}\)\}withH⁡\(Pi\)H\(P\_\{i\}\)the Shannon entropy of the conditional probability row and generally interpreted as the effective number of neighbors\[[21](https://arxiv.org/html/2608.17190#bib.bib1)\]\. However, we observe that typically only a few neighbors dominate the probability distribution for each point, i\.e\., that the distributions are ‘sharp’, causing a few nearest neighbors to dominate the attractive forces\. Moreover, the sharpness of the distributions varies per point\. We influence this sharpness by applying a power transform controlled by a parameterγ\\gammathat can smooth \(or sharpen further\) each row of the affinity matrix, while preserving the original neighbor support and rank order\. This gives a controlled way of studying the effect of affinity sharpness on the final embedding\.

Theoretical and empirical results\. We show that the transform is equivalent to changing the effective Gaussian bandwidth, and hence the perplexity, but unlike increasing the global perplexity, it produces point\-dependent effective perplexities\. Empirically, we find that changingγ\\gammaaffects neighborhood preservation differently at different scales: sharpening improves preservation of the very nearest neighbors, while smoothing improves both preservation of broader local neighborhoods and global structure\. The effect is most pronounced at lower perplexities, where the original affinity rows are sharpest, and the effect becomes smaller as perplexity increases\. These improvements cannot be replicated by perplexity tuning alone, because smoothing assigns different effective neighborhood sizes to different points rather than just increasing perplexity of all points\. We also show thatγ\\gammasmoothing outperforms alternative affinity constructions in the mid\-local neighborhood range\. Smoothing does not affect scalability as the number of \(non\-zero\) neighbors in the affinity matrix is left unchanged\.

Paper outline\. We first discuss related work in Section[2](https://arxiv.org/html/2608.17190#S2)\. In Section[3](https://arxiv.org/html/2608.17190#S3), we revisit t\-SNE, introduce the affinity smoothing operation controlled byγ\\gamma, and characterize its effect on effective perplexity across data points\. Section[4](https://arxiv.org/html/2608.17190#S4)presents the empirical investigation of smoothing’s effects on neighborhood preservation\. Section[5](https://arxiv.org/html/2608.17190#S5)provides a summary of the main conclusions\. Supplementary material, including an Appendix in PDF format, source figures, and replication code, is provided at[https://github\.com/aida\-ugent/smooth\_affinity\_tsne\.git](https://github.com/aida-ugent/smooth_affinity_tsne.git)\.

## 2Related work

t\-distributed Stochastic Neighbor Embedding \(t\-SNE\)\[[21](https://arxiv.org/html/2608.17190#bib.bib1)\]uses high\-dimensional neighborhood probabilities from Gaussian kernels and then optimizes a low\-dimensional embedding whose probabilities match those using a heavy\-tailed Studentttdistribution\. Kobak and Berens\[[10](https://arxiv.org/html/2608.17190#bib.bib9)\]describe t\-SNE as particularly effective in revealing local structure\. Broader transcriptomic and cytometry benchmarks also report strong local\-neighborhood preservation for t\-SNE compared with methods such as UMAP\[[16](https://arxiv.org/html/2608.17190#bib.bib7)\], TriMap\[[1](https://arxiv.org/html/2608.17190#bib.bib10)\], and PaCMAP\[[23](https://arxiv.org/html/2608.17190#bib.bib8)\]\. Similarly, Novak et al\.\[[17](https://arxiv.org/html/2608.17190#bib.bib3)\]compare popular DR methods using separate local and global structure\-preservation scores and find that t\-SNE achieves the strongest local structure preservation among the evaluated methods\. At the same time, these studies also show that local preservation is not perfect: even strong t\-SNE embeddings can miss many high\-dimensional neighbors\[[4](https://arxiv.org/html/2608.17190#bib.bib6)\]\. This motivates analyses of which parts of the t\-SNE pipeline control local\-neighborhood preservation\.

A large body of work exists on improving the practical use of t\-SNE\. Barnes\-Hut t\-SNE reduces the cost of the original algorithm using tree\-based approximations\[[22](https://arxiv.org/html/2608.17190#bib.bib11)\], while FIt\-SNE uses interpolation and fast Fourier transform techniques to further accelerate the computation\[[15](https://arxiv.org/html/2608.17190#bib.bib12)\]\. openTSNE provides a modular implementation that makes large\-scale and experimental use of t\-SNE more accessible\[[18](https://arxiv.org/html/2608.17190#bib.bib13)\]\. Other work has studied optimization choices such as initialization, learning rate, early exaggeration, and iteration schedules\[[3](https://arxiv.org/html/2608.17190#bib.bib14),[5](https://arxiv.org/html/2608.17190#bib.bib15),[10](https://arxiv.org/html/2608.17190#bib.bib9)\]\. These works show that the quality and usability of t\-SNE strongly depend on how the optimization is performed\. Our focus is complementary to these studies: we keep the optimization procedure fixed and study how the high\-dimensional affinity distribution affects the embedding\.

The work closest to ours concerns the construction of high\-dimensional affinities and the choice of neighborhood scale\. In standard t\-SNE, the neighborhood scale is controlled by the target perplexity, making perplexity a central design choice in the affinity construction\. Several works have questioned whether this single\-scale construction is sufficient\. Multiscale Stochastic Neighbor Embedding constructs multiscale similarities by averaging softmax\-based similarities over multiple bandwidths, with the goal of improving embedding quality across scales and reducing dependence on a single scale parameter\[[13](https://arxiv.org/html/2608.17190#bib.bib16)\]\. Perplexity\-free t\-SNE follows a related motivation by constructing neighborhoods across Gaussian kernels with growing widths, avoiding the need for a user\-specified perplexity\[[6](https://arxiv.org/html/2608.17190#bib.bib17)\]\. Multi\-scale kernels are also used in practical protocols for single\-cell transcriptomics\[[10](https://arxiv.org/html/2608.17190#bib.bib9)\]\. Some implementations instead fix a single Gaussian bandwidth for all points\[[18](https://arxiv.org/html/2608.17190#bib.bib13)\]\. Other variants modify the similarity kernels themselves\. Twice Student t\-SNE replaces the Gaussian high\-dimensional similarities with Student\-type ones, changing how neighborhoods are represented before optimization\[[6](https://arxiv.org/html/2608.17190#bib.bib17)\]\. Heavy\-tailed kernel variants instead modify the low\-dimensional Student\-t kernel by changing its degrees of freedom, showing that heavier tails can reduce crowding and reveal finer cluster structure\[[11](https://arxiv.org/html/2608.17190#bib.bib18)\]\.

Our analysis is closest to this affinity\-construction literature, but asks a narrower question: what is the effect of changing only the row\-wise distribution of probability mass after the standard t\-SNE affinities have been constructed? We use a row\-wise power transform as a controlled perturbation that preserves neighbor support and rank order while changing affinity sharpness\. This allows us to isolate whether local neighborhood preservation depends only on which neighbors are present, or also on how strongly those neighbors are weighted\.

## 3Methodology

In this section we detail the proposed smoothing operation\. The section is structured as follows\. We first review t\-SNE and the construction of the conditional probabilities that we later transform\. We then introduce a power transform that can smooth \(or sharpen\) each conditional row, without changing which neighbors are present or how they are ordered\. Finally, we show that this transform is equivalent to changing the bandwidth of the Gaussian distribution and introduce the concept of*effective perplexity*, being the perplexity of the resulting distribution parameterized by the input perplexityρ\\rhoand the smoothing parameterγ\\gamma\. This yields a generalized variant of t\-SNE that enables us to study how the smoothness of the affinities affects neighborhood preservation in embeddings\.

### 3\.1Standard t\-SNE affinity construction

t\-SNE is built up through three main concepts: \(1\) a mapping of the high\-dimensional pairwise distances to \(symmetrized\) probabilities, collected in the affinity matrix, \(2\) definition of the similarity kernel in the low\-dimensional space using the t\-distribution, which is where the prefix*t*comes from, and \(3\) the definition of the overall objective, namely the KL\-divergence between the low\- and high\-dimensional probabilities\. We review each in turn below\.

From distances to conditional probabilities\. The first step maps pairwise distances in the high\-dimensional space to probabilities that represent similarities: LetX=\{x1,x2,…,xn\}X=\\\{x\_\{1\},x\_\{2\},\\ldots,x\_\{n\}\\\}be the high\-dimensional data, wherexi∈ℝmx\_\{i\}\\in\\mathbb\{R\}^\{m\}\. For each pointxix\_\{i\}, a conditional probability distribution is defined over the other points\. The conditional probability of choosingxjx\_\{j\}as a neighbor ofxix\_\{i\}is

pj\|i=exp⁡\(−βi​di​j2\)∑m≠iexp⁡\(−βi​di​m2\),βi=12​σi2\.\\displaystyle p\_\{j\|i\}=\\frac\{\\exp\(\-\\beta\_\{i\}d\_\{ij\}^\{2\}\)\}\{\\sum\_\{m\\neq i\}\\exp\(\-\\beta\_\{i\}d\_\{im\}^\{2\}\)\},\\qquad\\beta\_\{i\}=\\frac\{1\}\{2\\sigma\_\{i\}^\{2\}\}\.\(1\)withpi\|i=0p\_\{i\|i\}=0anddi​j=‖xi−xj‖d\_\{ij\}=\\\|x\_\{i\}\-x\_\{j\}\\\|,σi\\sigma\_\{i\}is the bandwidth of the Gaussian kernel centered atxix\_\{i\}\. A smallσi\\sigma\_\{i\}produces a sharper distribution, while a largerσi\\sigma\_\{i\}produces a smoother distribution\.

In t\-SNE, a separateσi\\sigma\_\{i\}is identified for every point, so that each conditional distribution has a fixed user\-defined perplexity:Perp⁡\(Pi\)=ρ\\mathrm\{Perp\}\(P\_\{i\}\)=\\rho\. The perplexity of the conditional distributionPiP\_\{i\}is defined as the exponential of its Shannon entropy wherePi=\{pj\|i\}j≠iP\_\{i\}=\\\{p\_\{j\|i\}\\\}\_\{j\\neq i\}denotes rowiiof the conditional probability \(affinity\) matrix and it contains the neighborhood relations for pointii:

Perp\(Pi\)=2H⁡\(Pi\),H\(Pi\)=−∑jpj\|ilog2pj\|i\.\\displaystyle\\mathrm\{Perp\}\(P\_\{i\}\)=2^\{H\(P\_\{i\}\)\},\\qquad H\(P\_\{i\}\)=\-\\sum\_\{j\}p\_\{j\|i\}\\log\_\{2\}p\_\{j\|i\}\.\(2\)
Perp⁡\(Pi\)\\mathrm\{Perp\}\(P\_\{i\}\)is commonly interpreted as the effective number of neighbors represented by the distribution\[[21](https://arxiv.org/html/2608.17190#bib.bib1)\], and equalsρ\\rhofor all points\.

The affinity matrix\. After computing the conditional probabilities, t\-SNE symmetrizes them to obtain a joint high\-dimensional*affinity matrix*PPby settingpi​j=\(pj\|i\+pi\|j\)/\(2​n\)p\_\{ij\}=\(p\_\{j\|i\}\+p\_\{i\|j\}\)/\(2n\)fori≠ji\\neq j, andpi​i=0p\_\{ii\}=0\. The entriespi​jp\_\{ij\}define the strength of the high\-dimensional neighborhood relation between pointsiiandjj\. Larger values ofpi​jp\_\{ij\}indicate stronger attractive relations in the embedding optimization\.

Similarity in the embedding\. t\-SNE then defines a corresponding probability distribution in the low\-dimensional embedding\. LetY=Y=\{y1,y2,…,yn\}\\\{y\_\{1\},y\_\{2\},\\ldots,y\_\{n\}\\\}, whereyi∈ℝ2y\_\{i\}\\in\\mathbb\{R\}^\{2\}\. The low\-dimensional similarity between two embedded points is defined using a heavy\-tailed Studentttkernel with one degree of freedom:

qi​j=\(1\+‖yi−yj‖2\)−1∑k≠l\(1\+‖yk−yl‖2\)−1,qi​i=0\.\\displaystyle q\_\{ij\}=\\frac\{\\left\(1\+\\\|y\_\{i\}\-y\_\{j\}\\\|^\{2\}\\right\)^\{\-1\}\}\{\\sum\_\{k\\neq l\}\\left\(1\+\\\|y\_\{k\}\-y\_\{l\}\\\|^\{2\}\\right\)^\{\-1\}\},\\qquad q\_\{ii\}=0\.\(3\)The heavy\-tailed kernel allows moderately distant points in the embedding to still have non\-negligible similarity, which helps reduce the crowding problem\.

The optimization objective\. Finally, the embeddingYYis found by minimizing the Kullback–Leibler divergence between the high\-dimensional distributionPPand the low\-dimensional distributionQQ:

C=KL\(P∥Q\)=∑i≠jpi​jlogpi​jqi​j\.\\displaystyle C=KL\(P\\\|Q\)=\\sum\_\{i\\neq j\}p\_\{ij\}\\log\\frac\{p\_\{ij\}\}\{q\_\{ij\}\}\.\(4\)This objective encourages pairs with largepi​jp\_\{ij\}to also have largeqi​jq\_\{ij\}, meaning that points with strong high\-dimensional affinities should remain close in the embedding\. Therefore, the high\-dimensional affinity matrixPPdetermines which local neighbor relations are emphasized during optimization\.

### 3\.2Row\-wise smoothing of conditional probabilities

Standard t\-SNE fixes the target perplexity of every conditional probability row\. However, this does not necessarily mean that the resulting affinity rows assign probability mass evenly across neighbors\.

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_01_affinity_sharpness.PNG)

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_02a_eff_perp.PNG)

Figure 2:\(a\) Affinity row sharpness atρ=30\\rho=30for a single point from MNIST \(selected as average wrt\. summed mass of the top\-5 neighbors\)\. Standard t\-SNE \(γ=1\.0\\gamma=1\.0\) drops steeply after the first few neighbors, while smoothing \(γ=0\.5\\gamma=0\.5\) redistributes probability mass more evenly across the neighborhood\. \(b\) Distribution of Effective perplexity across points on MNIST for standard t\-SNE \(ρ=30\\rho=30\), smoothed t\-SNE \(γ=0\.7\\gamma=0\.7\), and sharpened t\-SNE \(γ=1\.5\\gamma=1\.5\)From early empirical analysis, we observed that the affinity distributions are typically very skewed: as illustrated in Figure[2](https://arxiv.org/html/2608.17190#S3.F2)\(a\), at the default perplexity of 30 just a few neighbors dominate\. This led us to the idea of smoothing the distribution to spread probability mass more equally over the neighbors\.

We apply the following power transform to the conditional probabilities:

p~j\|i=pj\|iγ∑m≠ipm\|iγ,\\displaystyle\\tilde\{p\}\_\{j\|i\}=\\frac\{p\_\{j\|i\}^\{\\gamma\}\}\{\\sum\_\{m\\neq i\}p\_\{m\|i\}^\{\\gamma\}\},\(5\)whereγ≥0\\gamma\\geq 0controls the sharpness of the row distribution\. When0≤γ<10\\leq\\gamma<1, the transform smooths the distribution: large probabilities are reduced relative to smaller probabilities, and probability mass is redistributed toward lower\-probability neighbors\. Whenγ\>1\\gamma\>1, the transform sharpens the distribution by concentrating more mass on the largest probabilities\. Whenγ=1\\gamma=1, the original t\-SNE conditional probabilities are recovered\.

As shown in Figure[2](https://arxiv.org/html/2608.17190#S3.F2)\(a\), applying the power transform withγ=0\.5\\gamma=0\.5redistributes the mass more evenly and all neighbors now receive meaningful probability mass\. Note that it is common to use3⋅ρ3\\cdot\\rhoneighbors in the optimization and define all further neighbors to havepj\|i=0p\_\{j\|i\}=0\[[22](https://arxiv.org/html/2608.17190#bib.bib11),[18](https://arxiv.org/html/2608.17190#bib.bib13)\]\. We do not make any change with respect to this, and hence runtime will be largely unaffected by smoothing\.

This transformation preserves the rank order within each row, because the power transform is monotonic forγ\>0\\gamma\>0and therefore changes only the relative probability mass\. After applying the row\-wise transform, the modified conditional probabilities are symmetrized in the same way as in standard t\-SNE, namely by settingp~i​j=\(p~j\|i\+p~i\|j\)/\(2​n\)\\tilde\{p\}\_\{ij\}=\(\\tilde\{p\}\_\{j\|i\}\+\\tilde\{p\}\_\{i\|j\}\)/\(2n\)wheni≠ji\\neq j, andp~i​i=0\\tilde\{p\}\_\{ii\}=0\. The resulting matrixP~\\tilde\{P\}is then used in the standard t\-SNE objective:C~=KL\(P~∥Q\)\\tilde\{C\}=KL\(\\tilde\{P\}\\\|Q\)\. We leave the low\-dimensional kernel and optimization objective unchanged\.

### 3\.3Bandwidth and effective\-perplexity interpretation

Smoothing changes the effective bandwidth\. The smoothing transform in Eq\.[5](https://arxiv.org/html/2608.17190#S3.E5)has a direct interpretation in terms of the Gaussian bandwidths used in the high\-dimensional affinity construction\. Starting from the standard t\-SNE conditional probabilities in Eq\.[1](https://arxiv.org/html/2608.17190#S3.E1), applying the row\-wise power transform gives

p~j\|i=\(exp⁡\(−βi​di​j2\)∑r≠iexp⁡\(−βi​di​r2\)\)γ∑m≠i\(exp⁡\(−βi​di​m2\)∑r≠iexp⁡\(−βi​di​r2\)\)γ\.\\displaystyle\\tilde\{p\}\_\{j\|i\}=\\frac\{\\left\(\\frac\{\\exp\(\-\\beta\_\{i\}d\_\{ij\}^\{2\}\)\}\{\\sum\_\{r\\neq i\}\\exp\(\-\\beta\_\{i\}d\_\{ir\}^\{2\}\)\}\\right\)^\{\\gamma\}\}\{\\sum\_\{m\\neq i\}\\left\(\\frac\{\\exp\(\-\\beta\_\{i\}d\_\{im\}^\{2\}\)\}\{\\sum\_\{r\\neq i\}\\exp\(\-\\beta\_\{i\}d\_\{ir\}^\{2\}\)\}\\right\)^\{\\gamma\}\}\.The denominator of the original conditional distribution,Zi=∑r≠iexp⁡\(−βi​di​r2\)Z\_\{i\}=\\sum\_\{r\\neq i\}\\exp\(\-\\beta\_\{i\}d\_\{ir\}^\{2\}\), is constant within rowii\. Therefore, the factorZi−γZ\_\{i\}^\{\-\\gamma\}appears in both the numerator and the denominator of the transformed row and cancels out during renormalization\. This gives

p~j\|i=exp⁡\(−γ​βi​di​j2\)∑m≠iexp⁡\(−γ​βi​di​m2\)\.\\tilde\{p\}\_\{j\|i\}=\\frac\{\\exp\(\-\\gamma\\beta\_\{i\}d\_\{ij\}^\{2\}\)\}\{\\sum\_\{m\\neq i\}\\exp\(\-\\gamma\\beta\_\{i\}d\_\{im\}^\{2\}\)\}\.Thus, applying the power transform to the conditional probabilities is equivalent to replacing the original Gaussian precisionβi\\beta\_\{i\}byβ~i=γ​βi\\tilde\{\\beta\}\_\{i\}=\\gamma\\beta\_\{i\}\. Sinceβi=1/\(2​σi2\)\\beta\_\{i\}=1/\(2\\sigma\_\{i\}^\{2\}\), we can writeβ~i=12​σ~i2=γ​12​σi2\\tilde\{\\beta\}\_\{i\}=\\frac\{1\}\{2\\tilde\{\\sigma\}\_\{i\}^\{2\}\}=\\gamma\\frac\{1\}\{2\\sigma\_\{i\}^\{2\}\}\. Solving for the new effective bandwidth gives

σ~i=σiγ\.\\tilde\{\\sigma\}\_\{i\}=\\frac\{\\sigma\_\{i\}\}\{\\sqrt\{\\gamma\}\}\.\(6\)
Therefore, when0≤γ<10\\leq\\gamma<1, the effective bandwidth increases, i\.e\.,σ~i\>σi\\tilde\{\\sigma\}\_\{i\}\>\\sigma\_\{i\}\. This means that smoothing makes the effective Gaussian kernel wider and the resulting conditional distribution less concentrated\. Conversely, whenγ\>1\\gamma\>1, the effective bandwidth decreases, making the distribution sharper\.

Smoothing yields point\-dependent effective perplexities\. In standard t\-SNE, the bandwidthsσi\\sigma\_\{i\}are chosen so that each conditional probability rowPi=\{pj\|i\}j≠iP\_\{i\}=\\\{p\_\{j\|i\}\\\}\_\{j\\neq i\}matches the same target perplexityρ\\rho\. After applying the transform in Eq\.[5](https://arxiv.org/html/2608.17190#S3.E5), the*effective perplexity*of the smoothed rowP~i=\{p~j\|i\}j≠i\\tilde\{P\}\_\{i\}=\\\{\\tilde\{p\}\_\{j\|i\}\\\}\_\{j\\neq i\}is

Perp\(P~i\)=2H⁡\(P~i\),H\(P~i\)=−∑jp~j\|ilog2p~j\|i\.\\displaystyle\\mathrm\{Perp\}\(\\tilde\{P\}\_\{i\}\)=2^\{H\(\\tilde\{P\}\_\{i\}\)\},\\qquad H\(\\tilde\{P\}\_\{i\}\)=\-\\sum\_\{j\}\\tilde\{p\}\_\{j\|i\}\\log\_\{2\}\\tilde\{p\}\_\{j\|i\}\.Substituting Eq\.[5](https://arxiv.org/html/2608.17190#S3.E5)into the entropy of the smoothed row gives

H\(P~i\)=−∑jpj\|iγ∑m≠ipm\|iγlog2\(pj\|iγ∑m≠ipm\|iγ\)\.H\(\\tilde\{P\}\_\{i\}\)=\-\\sum\_\{j\}\\frac\{p\_\{j\|i\}^\{\\gamma\}\}\{\\sum\_\{m\\neq i\}p\_\{m\|i\}^\{\\gamma\}\}\\log\_\{2\}\\left\(\\frac\{p\_\{j\|i\}^\{\\gamma\}\}\{\\sum\_\{m\\neq i\}p\_\{m\|i\}^\{\\gamma\}\}\\right\)\.
This expression depends on the individual probabilitiespj\|ip\_\{j\|i\}, not only on the original entropyH⁡\(Pi\)H\(P\_\{i\}\)that is fixed by the constant perplexityρ\\rho\. Applying a constantγ≠1\\gamma\\neq 1across rows can therefore produce different smoothed entropies, and hence different effective perplexities\.

This gives the adaptive\-perplexity interpretation of the transform: row\-wise smoothing lets perplexity differ across rows, unlike standard t\-SNE\. Smoothing is thus also a different operation than running standard t\-SNE with a larger perplexity, which would again match every row to the same new target perplexity\.

Figure[2](https://arxiv.org/html/2608.17190#S3.F2)\(b\) confirms this empirically on MNIST\. It shows the distribution of effective perplexities across points for standard t\-SNE \(γ=1\\gamma=1\) and smoothed and sharpened distributions, all starting from perplexityρ=30\\rho=30\. Standard t\-SNE concentrates all points at exactly the target perplexity of3030\. In contrast, smoothing withγ=0\.7\\gamma=0\.7produces a broad distribution of effective perplexities centered around5353, and sharpening withγ=1\.5\\gamma=1\.5produces a broad distribution centered around1212\.

Since the change in effective perplexity depends on the original row distribution, it naturally varies across points: points with sharper affinity rows receive a larger increase, while points in sparser regions tend to be more affected as well \(see online Appendix Fig\. 2 for details\)\. Together, these results confirm thatγ\\gammacontrols a point\-adaptive redistribution of probability mass, producing a heterogeneous spread of effective neighborhood sizes rather than a uniform shift\.

## 4Experiments

We designed the experiments to address the following research questions:

1. 1\.Does smoothing affect neighborhood preservation, and at what neighborhood scale? \(Section[4\.1](https://arxiv.org/html/2608.17190#S4.SS1)\)
2. 2\.Is it possible to achieve the same effect by changing the global perplexity? \(Section[4\.2](https://arxiv.org/html/2608.17190#S4.SS2)\)
3. 3\.How doγ\\gammaandρ\\rhoaffect neighborhood preservation of near\-local and mid\-local neighborhood scales? \(Section[4\.3](https://arxiv.org/html/2608.17190#S4.SS3)\)
4. 4\.Does smoothing also improve global structure preservation? \(Section[4\.4](https://arxiv.org/html/2608.17190#S4.SS4)\)
5. 5\.How doesγ\\gammasmoothing compare to alternative affinity constructions in terms of neighborhood preservation? \(Section[4\.5](https://arxiv.org/html/2608.17190#S4.SS5)\)

We used three datasets: MNIST\[[12](https://arxiv.org/html/2608.17190#bib.bib22)\]\(n=70,000n=70\{,\}000,m=784m=784\), UCI Adult\[[2](https://arxiv.org/html/2608.17190#bib.bib23)\]\(n=48,842n=48\{,\}842,m=14m=14\), and the mouse cortex dataset from Tasic et al\.\[[20](https://arxiv.org/html/2608.17190#bib.bib19)\]\(n=23,822n=23\{,\}822,m=50m=50\)\. For MNIST, pixel values are scaled to\[0,1\]\[0,1\]and reduced to 50 principal components, following Kobak and Berens\[[10](https://arxiv.org/html/2608.17190#bib.bib9)\]\. For Adult, numerical variables are standardized and categorical variables are one\-hot encoded\. For mouse cortex, we use the preprocessed data from Kobak and Berens\[[10](https://arxiv.org/html/2608.17190#bib.bib9)\]\. Due to space constraints, we show results for MNIST only\. Full results are available in the online appendix and are consistent across the three datasets\.

All embeddings are computed using openTSNE\[[18](https://arxiv.org/html/2608.17190#bib.bib13)\], and we implementedγ\\gammaas a parameter of the affinity construction within the forked version at[https://github\.com/aida\-ugent/smooth\_affinity\_tsne\.git](https://github.com/aida-ugent/smooth_affinity_tsne.git)\. For fair comparison, all variants use the same initialization, early exaggeration, learning rate, number of iterations, and optimization settings\. The only difference between standard t\-SNE and the smoothed variants is the row\-wise transformation\. All reported metric values are averages over five runs\. Standard deviations are omitted as too small to visualize: below0\.0030\.003\(typically0\.00050\.0005\) for every reportedN​O​@​kNO@kand AUC across all three datasets\. For global structure preservation they are larger, so bands are shown in Figure[7](https://arxiv.org/html/2608.17190#S4.F7)\(a\)\. All results and per\-run values are in the online repository\.

### 4\.1Effect of smoothing on local neighborhood preservation

To measure neighborhood structure preservation, we use neighborhood overlap atkk, denotedN​O​@​kNO@k\(also known asQN​X​\(k\)Q\_\{NX\}\(k\)in the DR evaluation literature\[[14](https://arxiv.org/html/2608.17190#bib.bib21)\]\), which measures the fraction of high\-dimensionalkk\-nearest neighbors that are also among thekk\-nearest neighbors in the embedding\.

N​Oi​\(k\)=\|𝒩iH​D​\(k\)∩𝒩iL​D​\(k\)\|k,N​O​@​k=1n​∑i=1nN​Oi​\(k\)\.NO\_\{i\}\(k\)=\\frac\{\|\\mathcal\{N\}^\{HD\}\_\{i\}\(k\)\\cap\\mathcal\{N\}^\{LD\}\_\{i\}\(k\)\|\}\{k\},\\qquad NO@k=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}NO\_\{i\}\(k\)\.
Figure[3](https://arxiv.org/html/2608.17190#S4.F3)showsN​O​@​kNO@kfork=1,…,200k=1,\\ldots,200at perplexityρ=30\\rho=30for MNIST\. Smoothed variants \(γ<1\\gamma<1\) score better at higher values ofkk, with the gap growing askkincreases, while sharpened variants \(γ\>1\\gamma\>1\) score better at smallkk\.

Figure[4](https://arxiv.org/html/2608.17190#S4.F4)shows the embeddings produced by standard t\-SNE and smoothed t\-SNE \(γ=0\.5\\gamma=0\.5\) on MNIST\. The overall cluster structure is preserved in both cases, but the smooth version shows stronger inter\-cluster separation: the dense overlapping region in the center is broken up, digit 0 \(blue\) becomes clearly isolated, and several neighboring digit classes are pulled apart\. This is consistent with smoothing giving more weight to broader neighborhoods, which pushes clusters further apart in the embedding\.

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_04_neighborhood_overlap_gamma_sweep.png)Figure 3:N​O​@​kNO@kfork=1,…,200k=1,\\ldots,200at perplexityρ=30\\rho=30with varyingγ\\gamma, on MNIST\. Sharpened variants \(γ\>1\\gamma\>1\) perform better for smallkk, while smoothed variants \(γ<1\\gamma<1\) perform better for largerkk\(starting fromk≈2/3​ρk\\approx 2/3\\rho\)\.![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_05_embedding.png)Figure 4:Standard t\-SNE \(ρ=30\\rho=30\) and smoothed t\-SNE \(ρ=30\\rho=30,γ=0\.5\\gamma=0\.5\) embeddings of MNIST\. Smoothing produces clearer inter\-cluster separation while preserving the overall digit cluster structure\.![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_07_neighborhood_overlap_comparison.PNG)

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_10a_no_vs_perplexity.png)

Figure 5:\(a\) NO@k comparison of standard t\-SNE \(ρ=30\\rho=30\), smoothed t\-SNE \(ρ=30\\rho=30,γ=0\.7\\gamma=0\.7\), and standard t\-SNE atρ=53\\rho=53, matching the median effective perplexity of smoothed t\-SNE\. \(b\)N​O​@​30NO@30for standard and smoothed t\-SNE \(with constantγ=0\.7\\gamma=0\.7\), for varying values ofρ\\rho\. We observe that both standard t\-SNE with same perplexity and matched perplexity have lower broad\-local NO\. Comparing across perplexities for a specificN​O​@​kNO@k, we see that standard t\-SNE with higher perplexities also does not achieve the same improved NO\.
### 4\.2Smoothing versus increasing perplexity

A natural question is whether the increased neighborhood preservation at largerkkcan also be achieved by increasing the perplexity\. We showed in Section[3](https://arxiv.org/html/2608.17190#S3)that smoothing produces point\-dependent effective perplexities, but the question is whether that matters for neighborhood preservation\.

To this end, Figure[5](https://arxiv.org/html/2608.17190#S4.F5)\(a\) shows three variants: standard t\-SNE atρ=30\\rho=30, smoothed t\-SNE atρ=30\\rho=30withγ=0\.7\\gamma=0\.7, and standard t\-SNE atρ=53\\rho=53, chosen to match the median effective perplexity across points of the smoothed variant \(see Appendix D\)\. Fork\>25k\>25, smoothed t\-SNE consistently outperforms standard t\-SNE atρ=30\\rho=30andρ=53\\rho=53\. This confirms that the point\-dependent effective perplexities produced by smoothing lead to different and better mid\-local preservation than simply increasing the global perplexity\.

However, we are unsure which perplexity would be best for whichN​O​@​kNO@k\. To explore this through an example, Figure[5](https://arxiv.org/html/2608.17190#S4.F5)\(b\) showsN​O​@​30NO@30, while varying perplexity for standard and smoothed t\-SNE \(with constantγ=0\.7\\gamma=0\.7; not tuned\)\. We observe that smoothing strongly affects which perplexity is optimal for this specific value ofkk\. For standard t\-SNE the best perplexity is around 40, while for smoothed t\-SNE withγ=0\.7\\gamma=0\.7the best value is achieved with perplexity 20\. We also confirm that a higher value ofN​O​@​30NO@30is achieved for smoothed t\-SNE \(0\.3600\.360\) than standard t\-SNE \(0\.3530\.353\), confirming that smoothed t\-SNE cannot be emulated by simply changing the perplexity\.

### 4\.3Sensitivity toγ\\gammaandρ\\rho

This experiment studies the results’ sensitivity to smoothing parameterγ\\gammaand perplexityρ\\rho\. We considerρ∈\{30,50,100,200\}\\rho\\in\\\{30,50,100,200\\\}andγ∈\{0,0\.5,0\.7,1,1\.2,1\.5,2\.0\}\\gamma\\in\\\{0,0\.5,0\.7,1,1\.2,1\.5,2\.0\\\}\. The extreme caseγ=0\\gamma=0produces a uniform distribution over all neighbors, i\.e\. discards the rank order within each row\.

We examined the effect of these parameters on near\-local structure \(k=1,…,10k=1,\\ldots,10\) and mid\-local structure \(k=11,…,90k=11,\\ldots,90; arbitrary cut\-offs\), shown in Figure[6](https://arxiv.org/html/2608.17190#S4.F6)\. We summarize neighborhood preservation for each range using the area under theN​O​@​kNO@kcurve \(AUC\)\. For near\-local preservation \(Figure[6](https://arxiv.org/html/2608.17190#S4.F6)a\), sharpening consistently improves performance across all perplexity settings\. The best value of AUC is0\.4500\.450atρ=30\\rho=30,γ=2\\gamma=2, while the worst is0\.0430\.043atρ=200\\rho=200,γ=0\\gamma=0\. Higherρ\\rhoalso hurts near\-local preservation, regardless ofγ\\gamma\.

For mid\-local preservation \(Figure[6](https://arxiv.org/html/2608.17190#S4.F6)b\), the pattern depends on the initial perplexity\. At lower perplexities, moderate smoothing outperforms standard t\-SNE: the best value is0\.3560\.356atρ=30\\rho=30,γ=0\.7\\gamma=0\.7\(highest overall\)\. However, at higher perplexity, the optimalγ\\gammashifts toward standard or even slightly sharpened variants\. Atρ=200\\rho=200, the best value is0\.3470\.347atγ=1\.5\\gamma=1\.5\. As expected,γ=0\\gamma=0performs worst across all perplexity settings, reaching only0\.1670\.167atρ=200\\rho=200\.

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_06a_neighborhood_overlap_auc_1_10.png)

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_06b_neighborhood_overlap_auc_11_90.png)

Figure 6:Sensitivity of t\-SNE behavior toγ\\gammaandρ\\rhoon the MNIST dataset\. Each cell shows the metric value for a combination ofρ\\rho\(rows\) andγ\\gamma\(columns\)\. \(a\) Near\-local preservation improves consistently with both sharpening and lowering the perplexity\. \(b\) Mid\-local preservation peaks at moderate smoothing aroundγ=0\.7\\gamma=0\.7; extreme smoothing withγ=0\\gamma=0performs worst\.![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_08_global_spearman_vs_gamma.PNG)

![Refer to caption](https://arxiv.org/html/2608.17190v1/figures/mnist_figures/mnist_nh_affinity_comparison.png)

Figure 7:\(a\) Effect ofγ\\gammaon global structure preservation \(Spearmanρ\\rhobetween high\- and low\-dimensional distances\) across four perplexity values on MNIST\. In general, lowerγ\\gammaand higherρ\\rhoimprove global structure preservation\. Shaded bands indicate ±1 SD across reruns\. \(b\)N​O​@​kNO@kfork=1,…,200k=1,\\ldots,200on MNIST, comparingγ\\gammavariants against alternative affinity constructions\. Sharpening \(γ=1\.5\\gamma=1\.5\) leads at smallkk; smoothing \(γ=0\.7\\gamma=0\.7\) outperforms all methods in the mid\-local range;γ=0\.0\\gamma=0\.0/UniformandFixedSigmaNNdominate at largekk\.
### 4\.4Effect of smoothing on global structure preservation

Figure[7](https://arxiv.org/html/2608.17190#S4.F7)\(a\) shows the effect ofγ\\gammaon global structure preservation, measured through the Spearman rank correlation between high\- and low\-dimensional distances\[[9](https://arxiv.org/html/2608.17190#bib.bib20)\]\. Generally, smoothing improves global structure preservation while sharpening degrades it\. As expected, larger perplexities also lead to stronger global structure preservation, but the effect of smoothing is also stronger at higher perplexities\. The best result is achieved atρ=200\\rho=200,γ=0\.75\\gamma=0\.75\.

### 4\.5Comparison with alternative affinity constructions

We compare affinity smoothing against five affinity constructions available in openTSNE\[[18](https://arxiv.org/html/2608.17190#bib.bib13)\]\.PerplexityBasedNN\[[21](https://arxiv.org/html/2608.17190#bib.bib1)\]is the standard t\-SNE affinity, equivalent to ourγ=1\.0\\gamma=1\.0\.MultiscaleMixtureandMultiscale\[[10](https://arxiv.org/html/2608.17190#bib.bib9)\]both incorporate multiple bandwidth scales: the former uses a mixture of Gaussians at different perplexities, while the latter averages single\-scale probability distributions across scales\.Uniformassigns equal probability mass to allkknearest neighbors, which should be equivalent toγ=0\\gamma=0\.FixedSigmaNNuses a fixed Gaussian bandwidth across all points rather than a perplexity\-adaptive one, which also results in point\-dependent effective perplexities\. We additionally include tt\-SNE\[[6](https://arxiv.org/html/2608.17190#bib.bib17)\], by reimplementing the respective affinity computation in openTSNE\. In tt\-SNE the standard high\-dimensional Gaussian kernel is replaced with a Student\-ttkernel\.

For all variants we use perplexityρ=30\\rho=30as default\. For the multiscale models we used perplexities\{30,50,100,200\}\\\{30,50,100,200\\\}\. ForFixedSigmaNN, we setσ\\sigmato the median distance to theρ\\rho\-th nearest neighbor across all points\. All variants are evaluated under an identical optimization pipeline so that only the affinities differ\.

Figure[7](https://arxiv.org/html/2608.17190#S4.F7)\(b\) shows the results on MNIST\. At smallkk, sharpened t\-SNE \(γ=1\.5\\gamma=1\.5\) gives the highestN​O​@​kNO@k, reaching approximately0\.50\.5compared to0\.420\.42for standard t\-SNE, before dropping sharply for largerkk\. In the mid\-local range \(k≈25k\\approx 25–9090\), smoothed t\-SNE \(γ=0\.7\\gamma=0\.7\) outperforms all other methods, including the multiscale variants, even thoughMultiscaleMixtureandMultiscaleincorporate information from multiple perplexity scales\. Beyondk≈90k\\approx 90,γ=0\.0\\gamma=0\.0,Uniform, andFixedSigmaNNtake over and reachN​O​@​k≈0\.40NO@k\\approx 0\.40atk=200k=200\.γ=0\.0\\gamma=0\.0andUniformindeed produce identical curves\. tt\-SNE tracks above standard t\-SNE fromk≈50k\\approx 50onwards, consistent with its smoother high\-dimensional affinity distribution, and shows the flattest curve overall with the least change acrosskk\. In contrast,γ=0\.0\\gamma=0\.0/UniformandFixedSigmaNNshow the largest variation, rising steeply from low values at smallkkto the highest values at largekk\. The curves reveal a consistent trade\-off governed by how sharply each affinity construction concentrates probability mass on the nearest neighbors\.

## 5Conclusion

In this work we analyzed the role of affinity sharpness in t\-SNE, focusing on how the row\-wise distribution of probability mass in the high\-dimensional affinity matrix affects neighborhood preservation at different scales\. We introduced a row\-wise power transform controlled by the parameterγ\\gamma, which smooths or sharpens each conditional probability row\. We showed analytically that this transform is equivalent to rescaling the Gaussian bandwidth, but produces point\-dependent effective perplexities, making it different from changing the global perplexity\.

Our experiments reveal a clear scale\-dependent trade\-off\. Sharpening the affinity rows improves preservation of the very nearest neighbors, while smoothing improves preservation of broader local neighborhoods\. Standard t\-SNE, with its fixed\-perplexity construction, appears to concentrate probability mass too heavily on the first few neighbors, which limits its ability to preserve broader local structure\. Smoothing redistributes this mass and gives more influence to mid\-ranked neighbors during optimization, improving mid\-local and global structure preservation\. The two parameters are complementary: perplexity controls how many neighbors are included, whileγ\\gammacontrols how mass is distributed among them\. Smoothing has the most effect at lower perplexities, where the affinity distributions are very sharp\. We further show thatγ=0\.7\\gamma=0\.7gives the strongest mid\-local preservation among the affinity constructions we evaluated, including multiscale methods, while remaining competitive at other scales; whether others could do better there remains open\.

These findings suggest that affinity sharpness is a meaningful and previously underexplored axis of t\-SNE behavior\. The choice ofγ\\gammaprovides practitioners with a lightweight way to shift the preservation focus depending on their visualization goal: sharper affinities for nearest\-neighbor fidelity, smoother affinities for broader local and global structure\.

A limitation of this study is the main focus on a single quality measure \(NO@k/QN​X​\(k\)Q\_\{NX\}\(k\)\), although its simplicity gives a clear view of one specific aspect of embedding quality\. The improvements are consistent but modest in absolute terms, and establishing when they translate into a practical benefit is itself difficult and beyond our present scope\. Secondly, future work could study how to optimize the hyperparametersρ,γ\\rho,\\gammato maximize specific NO@k\.

#### Acknowledgements

This research was funded by the Flemish Government \(AI Research Program\), the FWO \(G073924N\), the EU \(ERC, VIGILIA, 101142229\), and the BOF of Ghent University \(BOF20/IBF/117\)\. Views and opinions expressed are however those of the author\(s\) only and do not necessarily reflect those of the EU or the ERC Executive Agency\. Neither the EU nor the granting authority can be held responsible for them\. For the purpose of Open Access the authors have applied a CC BY public copyright license to any Author Accepted Manuscript version\.

#### Disclosure of Interests\.

The authors have no competing interests to declare that are relevant to the content of this article\.

## References

- \[1\]E\. Amid and M\. K\. Warmuth\(2019\)TriMap: Large\-scale Dimensionality Reduction Using Triplets\.Note:arXiv:1910\.00204Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p1.1)\.
- \[2\]B\. Becker and R\. Kohavi\(1996\)Adult data\.Note:UCI Machine Learning RepositoryCited by:[§4](https://arxiv.org/html/2608.17190#S4.p2.1)\.
- \[3\]A\. C\. Belkinaet al\.\(2019\)Automated optimized parameters for T\-distributed stochastic neighbor embedding improve visualization and analysis of large datasets\.Nature Comm\.10,pp\. 5415\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p2.1)\.
- \[4\]T\. Chari and L\. Pachter\(2023\)The specious art of single\-cell genomics\.PLOS Comp\. Bio\.19,pp\. e1011288\.Cited by:[§1](https://arxiv.org/html/2608.17190#S1.p2.1),[§2](https://arxiv.org/html/2608.17190#S2.p1.1)\.
- \[5\]P\. Chourasia, S\. Ali, and M\. Patterson\(2022\)Informative initialization and kernel selection improves t\-SNE for biological sequences\.Note:arXiv:2211\.09263Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p2.1)\.
- \[6\]C\. de Bodt, D\. Mulders, M\. Verleysen, and J\. A\. Lee\(2018\)Perplexity\-free t\-SNE and twice Student tt\-SNE\.Comp\. Intell\.\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p3.1),[§4\.5](https://arxiv.org/html/2608.17190#S4.SS5.p1.1)\.
- \[7\]C\. de Bodtet al\.\(2025\)Low\-dimensional embeddings of high\-dimensional data\.External Links:2508\.15929Cited by:[§1](https://arxiv.org/html/2608.17190#S1.p2.1)\.
- \[8\]E\. Heiter, L\. Martens, R\. Seurinck, M\. Guilliams, T\. De Bie, Y\. Saeys, and J\. Lijffijt\(2024\)Pattern or Artifact? Interactively Exploring Embedding Quality with TRACE\.InProc\. of ECML\-PKDD,pp\. 379–382\.Cited by:[Figure 1](https://arxiv.org/html/2608.17190#S1.F1)\.
- \[9\]P\. Joia, F\. V\. Paulovich, D\. Coimbra, J\. A\. Cuminato, and L\. G\. Nonato\(2011\)Local affine multidimensional projection\.IEEE TVCG17\(12\),pp\. 2563–2571\.Cited by:[§4\.4](https://arxiv.org/html/2608.17190#S4.SS4.p1.1)\.
- \[10\]D\. Kobak and P\. Berens\(2019\)The art of using t\-SNE for single\-cell transcriptomics\.Nature Comm\.10,pp\. 5416\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p1.1),[§2](https://arxiv.org/html/2608.17190#S2.p2.1),[§2](https://arxiv.org/html/2608.17190#S2.p3.1),[§4\.5](https://arxiv.org/html/2608.17190#S4.SS5.p1.1),[§4](https://arxiv.org/html/2608.17190#S4.p2.1)\.
- \[11\]D\. Kobak, G\. Linderman, S\. Steinerberger, Y\. Kluger, and P\. Berens\(2020\)Heavy\-tailed kernels reveal a finer cluster structure in t\-SNE visualisations\.InProc\. of ECML\-PKDD,pp\. 124–139\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p3.1)\.
- \[12\]Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner\(1998\)Gradient\-based learning applied to document recognition\.Proc\. of the IEEE86\(11\),pp\. 2278–2324\.Cited by:[§4](https://arxiv.org/html/2608.17190#S4.p2.1)\.
- \[13\]J\. A\. Lee, D\. H\. Peluffo\-Ordonez, and M\. Verleysen\(2014\)Multiscale stochastic neighbor embedding: Towards parameter\-free dimensionality reduction\.Comp\. Intell\.\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p3.1)\.
- \[14\]J\. A\. Lee and M\. Verleysen\(2009\)Quality assessment of dimensionality reduction: rank\-based criteria\.Neurocomp\.72\(7–9\),pp\. 1431–1443\.Cited by:[§4\.1](https://arxiv.org/html/2608.17190#S4.SS1.p1.1)\.
- \[15\]G\. C\. Linderman, M\. Rachh, J\. G\. Hoskins, S\. Steinerberger, and Y\. Kluger\(2019\)Fast interpolation\-based t\-SNE for improved visualization of single\-cell RNA\-seq data\.Nature Methods16\(3\),pp\. 243–245\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p2.1)\.
- \[16\]L\. McInnes, J\. Healy, and J\. Melville\(2018\)UMAP: uniform manifold approximation and projection for dimension reduction\.Note:arXiv:1802\.03426Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p1.1)\.
- \[17\]D\. Novak, C\. De Bodt, P\. Lambert, J\. A\. Lee, S\. Van Gassen, and Y\. Saeys\(2026\)Interpretable models for scRNA\-seq data embedding with multi\-scale structure preservation\.Note:bioRxiv 2023\.11\.23\.568428 v4Cited by:[Figure 1](https://arxiv.org/html/2608.17190#S1.F1),[§1](https://arxiv.org/html/2608.17190#S1.p2.1),[§1](https://arxiv.org/html/2608.17190#S1.p3.1),[§2](https://arxiv.org/html/2608.17190#S2.p1.1)\.
- \[18\]P\. G\. Poličar, M\. Stražar, and B\. Zupan\(2024\)OpenTSNE: a modular python library for t\-SNE dimensionality reduction and embedding\.JoSS109\(3\),pp\. 1–30\.Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p2.1),[§2](https://arxiv.org/html/2608.17190#S2.p3.1),[§3\.2](https://arxiv.org/html/2608.17190#S3.SS2.p4.1),[§4\.5](https://arxiv.org/html/2608.17190#S4.SS5.p1.1),[§4](https://arxiv.org/html/2608.17190#S4.p3.1)\.
- \[19\]I\. Scheyltjenset al\.\(2022\)Single\-cell RNA and protein profiling of immune cells from the mouse brain and its border tissues\.Nature Protocols17,pp\. 2354–2388\.Cited by:[Figure 1](https://arxiv.org/html/2608.17190#S1.F1)\.
- \[20\]B\. Tasicet al\.\(2018\)Shared and distinct transcriptomic cell types across neocortical areas\.Nature563,pp\. 72–78\.Cited by:[§4](https://arxiv.org/html/2608.17190#S4.p2.1)\.
- \[21\]L\. van der Maaten and G\. Hinton\(2008\)Visualizing data using t\-SNE\.JMLR9,pp\. 2579–2605\.Cited by:[§1](https://arxiv.org/html/2608.17190#S1.p2.1),[§1](https://arxiv.org/html/2608.17190#S1.p5.1),[§2](https://arxiv.org/html/2608.17190#S2.p1.1),[§3\.1](https://arxiv.org/html/2608.17190#S3.SS1.p4.1),[§4\.5](https://arxiv.org/html/2608.17190#S4.SS5.p1.1)\.
- \[22\]L\. van der Maaten\(2013\)Barnes\-Hut\-SNE\.Note:arXiv:1301\.3342Cited by:[§2](https://arxiv.org/html/2608.17190#S2.p2.1),[§3\.2](https://arxiv.org/html/2608.17190#S3.SS2.p4.1)\.
- \[23\]Y\. Wang, H\. Huang, C\. Rudin, and Y\. Shaposhnik\(2020\)Understanding How Dimension Reduction Tools Work: An Empirical Approach to Deciphering t\-SNE, UMAP, TriMAP, and PaCMAP for Data Visualization\.Note:arXiv:2012\.04456Cited by:[§1](https://arxiv.org/html/2608.17190#S1.p3.1),[§2](https://arxiv.org/html/2608.17190#S2.p1.1)\.

Similar Articles

Exploring Oversmoothing with Householder Matrices

arXiv cs.LG

This paper introduces HouseGNN, a graph neural network that uses Householder reflections and GroupSort to address oversmoothing, proving that each layer preserves node-wise Euclidean norm and showing improved behavior at depth.

Oversmoothing as Representation Degeneracy in Neural Sheaf Diffusion

arXiv cs.LG

This paper analyzes oversmoothing in Neural Sheaf Diffusion (NSD) as a representation degeneracy phenomenon using quiver theory and Geometric Invariant Theory. It proposes moment-map-inspired regularizers and explores non-uniform stalk dimensions to mitigate this issue in heterophilic graph benchmarks.

Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization

arXiv cs.LG

Introduces Depth-wise Gradient Augmentation, a general optimization paradigm that transforms block-wise optimizer updates along depth dimension. The method, Gradient Smoothing, improves optimization and generalization across diverse architectures including transformers and diffusion models.