preprocessing

Tag

Cards List
#preprocessing

Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications

arXiv cs.LG ↗ · 2026-09-21 Cached

This paper proposes an alternative signal-centric method for processing sonar data in remote sensing applications, using CSV format and acoustic processing to reduce processing time by 91.18% and improve machine learning-driven object detection.

0 favorites 0 likes
#preprocessing

Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

arXiv cs.CL ↗ · 2026-09-18 Cached

This paper demonstrates that removing stopwords in legal text analysis distorts doctrinal and ideological signals, recommending that stopwords be retained for measurement validity.

0 favorites 0 likes
#preprocessing

One Color Preprocessing Improves DSATUR

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper proposes SSLD, a method that improves the DSATUR heuristic for graph coloring by using semidefinite programming to preprocess an initial color class, demonstrating better performance on benchmark instances.

0 favorites 0 likes
#preprocessing

How an unsupported tool-call response could become “perfectly stable” in an LLM benchmark

Reddit r/artificial ↗ · 2026-09-03

An LLM benchmark's scoring pipeline could erroneously treat unsupported tool-call responses as stable empty outputs due to preprocessing, highlighting the importance of distinguishing between parsing failures and genuine outputs in evaluation.

0 favorites 0 likes
#preprocessing

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence ↗ · 2026-07-30

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

0 favorites 0 likes
#preprocessing

Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse

arXiv cs.CL ↗ · 2026-07-14 Cached

This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.

0 favorites 0 likes
#preprocessing

When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation

arXiv cs.CL ↗ · 2026-07-10 Cached

This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.

0 favorites 0 likes
#preprocessing

FourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention

arXiv cs.CL ↗ · 2026-07-09 Cached

This paper introduces FourierQK, a method that applies FFT-based frequency-domain preprocessing to learned query and key projections in transformer attention, achieving significant validation loss reductions on character-level language modelling. The approach preserves the full attention score structure and demonstrates reproducible gains over standard dot-product attention.

0 favorites 0 likes
#preprocessing

Evolutionary Feature Engineering for Structured Data

arXiv cs.LG ↗ · 2026-07-03 Cached

Introduces Evolutionary Feature Engineering (EFE), a framework that uses LLM-based evolution to automatically discover preprocessing transformations for structured data, improving time-series forecasting and tabular prediction accuracy while preserving interpretability.

0 favorites 0 likes
#preprocessing

How Good Can Linear Models Be for Time-Series Forecasting?

Hugging Face Daily Papers ↗ · 2026-06-25 Cached

This paper demonstrates that careful preprocessing—especially context length selection, normalization, and regularization—can make simple linear models like Ridge regression competitive with or superior to large Transformer, MLP, and CNN models on time-series forecasting benchmarks.

0 favorites 0 likes
#preprocessing

Best Preprocessing Techniques for Sentiment Analysis

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper systematically investigates the optimal order of preprocessing techniques for sentiment analysis on Twitter data, finding that tokenisation is most impactful and spelling correction least, with the best order being tokenisation, cleaning, stemming, then stopword removal.

0 favorites 0 likes
#preprocessing

Hermes got expensive when I let every profile think like a senior engineer.

Reddit r/AI_Agents ↗ · 2026-05-19

The author shares how running multiple persistent AI agent profiles under Hermes led to high API costs, solved by implementing tiered model policies per profile, pre-processing inputs, and using an API gateway for cost visibility, reducing daily costs from $14-18 to $7-10.

0 favorites 0 likes
#preprocessing

A Triadic Suffix Tokenization Scheme for Numerical Reasoning

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper introduces Triadic Suffix Tokenization (TST), a deterministic tokenization scheme that partitions digits into three-digit triads with explicit magnitude markers to improve numerical reasoning in large language models. The method addresses inconsistent number fragmentation in standard tokenizers by providing transparent order-of-magnitude relationships at the token level, with two implementation variants offering scalable vocabulary expansion.

0 favorites 0 likes
← Back to home

Submit Feedback