Tag
This paper proposes an alternative signal-centric method for processing sonar data in remote sensing applications, using CSV format and acoustic processing to reduce processing time by 91.18% and improve machine learning-driven object detection.
This paper demonstrates that removing stopwords in legal text analysis distorts doctrinal and ideological signals, recommending that stopwords be retained for measurement validity.
The paper proposes SSLD, a method that improves the DSATUR heuristic for graph coloring by using semidefinite programming to preprocess an initial color class, demonstrating better performance on benchmark instances.
An LLM benchmark's scoring pipeline could erroneously treat unsupported tool-call responses as stable empty outputs due to preprocessing, highlighting the importance of distinguishing between parsing failures and genuine outputs in evaluation.
Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.
This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.
This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.
This paper introduces FourierQK, a method that applies FFT-based frequency-domain preprocessing to learned query and key projections in transformer attention, achieving significant validation loss reductions on character-level language modelling. The approach preserves the full attention score structure and demonstrates reproducible gains over standard dot-product attention.
Introduces Evolutionary Feature Engineering (EFE), a framework that uses LLM-based evolution to automatically discover preprocessing transformations for structured data, improving time-series forecasting and tabular prediction accuracy while preserving interpretability.
This paper demonstrates that careful preprocessing—especially context length selection, normalization, and regularization—can make simple linear models like Ridge regression competitive with or superior to large Transformer, MLP, and CNN models on time-series forecasting benchmarks.
This paper systematically investigates the optimal order of preprocessing techniques for sentiment analysis on Twitter data, finding that tokenisation is most impactful and spelling correction least, with the best order being tokenisation, cleaning, stemming, then stopword removal.
The author shares how running multiple persistent AI agent profiles under Hermes led to high API costs, solved by implementing tiered model policies per profile, pre-processing inputs, and using an API gateway for cost visibility, reducing daily costs from $14-18 to $7-10.
This paper introduces Triadic Suffix Tokenization (TST), a deterministic tokenization scheme that partitions digits into three-digit triads with explicit magnitude markers to improve numerical reasoning in large language models. The method addresses inconsistent number fragmentation in standard tokenizers by providing transparent order-of-magnitude relationships at the token level, with two implementation variants offering scalable vocabulary expansion.