Tag
This paper proposes ZCA whitening as a geometric pre-processing step for WEAT to address embedding anisotropy, showing that calibration changes significance status for over 30% of results and that uncalibrated bias measurements may be unreliable.
This paper studies using sparse PPMI graph averaging to refine Random Indexing embeddings, showing it improves accuracy on a fairytales analogy benchmark but trails neural baselines on text8 and SimLex-999.
This paper explores how semantic relations are encoded in the geometry of language model semantic spaces, finding that asymmetric relations occupy distinct regions and that lexical information matters more for causal models while contextual information matters more for masked and diffusion models.
This paper surveys mechanisms for calculating word embeddings, investigates popular toolkits and embedding matrices, and experiments with selected implementations to understand their properties.
This paper proposes a framework using Supervised Semantic Differential to represent psychological constructs as directions in a shared word-embedding space, enabling comparison across different measurement instruments and research traditions.
This paper measures the semantic structure and evolution of conspiracy theories using 169.9M Reddit comments from r/politics (2012-2022), introducing the concept of "semantic objects" bounded by semantic neighborhoods to track how conspiracy theory meanings change over time beyond simple keyword-based approaches.
This paper presents adversarial and virtual adversarial training methods adapted for text classification by applying perturbations to word embeddings in RNNs rather than raw inputs. The approach achieves state-of-the-art results on semi-supervised and supervised text classification benchmarks while reducing overfitting.