anonymization

Tag

Cards List
#anonymization

Feedback on a Privacy Middleware for Sending Sensitive Data to LLMs

Reddit r/AI_Agents ↗ · 2d ago

A concept for a privacy middleware that anonymizes sensitive data before sending to LLMs to protect privacy, with local placeholder mapping, seeking feedback on feasibility.

0 favorites 0 likes
#anonymization

Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes

arXiv cs.CL ↗ · 4d ago Cached

This study audits demographic leakage in German LLM-generated resumes, finding that despite anonymization, classifiers can reliably distinguish between male and female names due to subtle linguistic differences, raising fairness concerns in AI hiring pipelines.

0 favorites 0 likes
#anonymization

Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation

arXiv cs.CL ↗ · 4d ago Cached

This paper investigates the privacy-personalization trade-off in LLMs by reducing stylistic signals in user-specific text generation, finding that anonymization lowers stylistic fidelity while preserving semantic meaning.

0 favorites 0 likes
#anonymization

Decades-Old Anonymized Medical Data May Cause AI Misdiagnoses Now

Reddit r/ArtificialInteligence ↗ · 2026-09-18 Cached

A study shows that anonymized patient records from decades past in AI training datasets can cause models to misdiagnose current patients by recalling historical health states, increasing the risk of errors in diagnoses.

0 favorites 0 likes
#anonymization

On the Impact of Anonymization on the Performance of Large Language Models

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper presents a systematic empirical study on the trade-off between privacy and performance in large language models when anonymizing input data, finding that anonymization degrades performance with effects varying by model capability and task type.

0 favorites 0 likes
#anonymization

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

arXiv cs.CL ↗ · 2026-08-10 Cached

Introduces GRASP, a method that uses Group Relative Policy Optimization to train a small on-device language model for adversarial anonymization, improving the privacy-utility trade-off over DPO-distilled baselines while running at a fraction of the cost of frontier teacher models.

0 favorites 0 likes
#anonymization

TIM PG

Product Hunt ↗ · 2026-08-08 Cached

TIM PG is a strictly offline Windows utility that automatically masks personal data from the clipboard before pasting into AI tools, ensuring secure local data privacy without cloud reliance.

0 favorites 0 likes
#anonymization

Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling

arXiv cs.CL ↗ · 2026-07-15 Cached

Introduces Knowledgeless Language Models (KLLMs), pretrained on corpora with anonymized entities to suppress parametric recall and enhance evidence-grounded reasoning, achieving substantial improvements on contextual QA, fact verification, and hallucination detection benchmarks.

0 favorites 0 likes
#anonymization

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

arXiv cs.CL ↗ · 2026-06-10 Cached

This paper investigates whether LLMs can identify their own model family from stylometric fingerprints in role-constrained political analysis texts, even after prompt-level anonymization. The findings confirm that anonymization is insufficient and have implications for EU AI Act compliance and multi-agent system validation.

0 favorites 0 likes
#anonymization

LLM Anonymization Against Agentic Re-Identification

Hugging Face Daily Papers ↗ · 2026-06-01 Cached

AURA is an LLM-powered anonymization framework that balances privacy protection against agentic web-search re-identification while preserving contextual utility through adaptive privacy scopes and mask-reconstruct methods.

0 favorites 0 likes
#anonymization

A Case Study on the Impact of Anonymization Along the RAG Pipeline

arXiv cs.CL ↗ · 2026-04-20 Cached

This case study empirically investigates where anonymization should be applied in Retrieval-Augmented Generation (RAG) pipelines to balance privacy and utility, examining the impact of anonymization at different stages (dataset vs. generated answer) to inform privacy risk mitigation strategies.

0 favorites 0 likes
← Back to home

Submit Feedback