deduplication

Tag

Cards List
#deduplication

GVD: Governed Versioning and Deduplication for Document Repositories

arXiv cs.AI · 5d ago Cached

GVD introduces a unified framework for governed versioning and deduplication in document repositories, using bidirectional rule alignment and conflict resolution policies with local encoder models to achieve high performance on enterprise data.

0 favorites 0 likes
#deduplication

We tried to make our AI verifier read less. How do you cut cost without silently missing evidence?

Reddit r/AI_Agents · 2026-09-15

The article describes testing to reduce reading cost in an AI verification pipeline while preserving high recall and minority evidence, noting challenges with current approaches and inviting community input.

0 favorites 0 likes
#deduplication

@ZyphraAI: Zyphra Research releases PUFFER, a novel incremental fuzzy deduplication system for LLM-scale datasets. PUFFER runs ent…

X AI KOLs Timeline · 2026-09-01 Cached

Zyphra Research releases PUFFER, an incremental fuzzy deduplication system for LLM-scale datasets that runs entirely on CPU with 11-35x speedup over existing methods, open-sourced under Apache 2.0.

0 favorites 0 likes
#deduplication

Truck Driver Builds AI News Aggregator

Reddit r/artificial · 2026-08-25

A truck driver with no coding background built an AI news aggregator using Next.js and Supabase that summarizes and deduplicates news from multiple sources.

0 favorites 0 likes
#deduplication

Every news-reading agent I built refetched the same articles forever. Here is the memory layer I settled on.

Reddit r/AI_Agents · 2026-08-23

The author describes building a memory layer for news-reading AI agents to handle deduplication and improve continuity across sessions, using SQLite and vector storage with a Python API, CLI, and MCP server.

0 favorites 0 likes
#deduplication

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

arXiv cs.CL · 2026-08-05 Cached

Proposes a scalable subdocument deduplication framework for LLM pretraining that separates duplicate detection from copy retention, using frequency- and length-aware policies. Experiments on FineWeb-Edu and a code web corpus show improved model performance.

0 favorites 0 likes
#deduplication

Frame selection is the whole game: notes on making LLMs watch video

Hacker News Top · 2026-08-03 Cached

Engineering notes on optimizing frame selection for feeding video to LLMs, covering scene detection, deduplication strategies, and token budget management.

0 favorites 0 likes
#deduplication

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv cs.AI · 2026-08-03 Cached

This paper presents a production extraction layer that converts heterogeneous documents into an ontology-aligned knowledge graph using a locally hosted tuned Qwen LLM, with ontology-guided prompts, multi-stage deduplication, and embedding-based resolution. Evaluation on intelligence corpora improved search recall from about 70 to 95 percent with no false merges.

0 favorites 0 likes
#deduplication

Vacuum 16T

Reddit r/LocalLLaMA · 2026-08-02

A deliberately empty 16.5-trillion-parameter model uploaded to Hugging Face exposes that parameter counts are computed from safetensors headers alone, and that Xet's content-defined deduplication reduces upload bandwidth enormously while storage quota still bills the full logical size.

0 favorites 0 likes
#deduplication

Hugging Face Storage Buckets (Website)

TLDR AI · 2026-07-31 Cached

Hugging Face introduces Storage Buckets, a scalable object storage service for AI teams with per-TB pricing, Xet deduplication, built-in CDN, and no git overhead, designed for datasets, model checkpoints, and ML artifacts.

0 favorites 0 likes
#deduplication

@jchudnov: Flying to #ICML2026 to present Internal Data Repetition Destroys Language Models, an Oral at Foundations of Deep Gen Mo…

X AI KOLs Following · 2026-07-04 Cached

A new study quantifies the damage caused by repeated data in language model pretraining, showing that even aggressive deduplication leaves harmful repetition that can waste up to a third of compute FLOPs.

0 favorites 0 likes
#deduplication

Internal Data Repetition Destroys Language Models

arXiv cs.LG · 2026-06-25 Cached

This paper systematically studies the damage caused by exact document repetition during language model pretraining, showing that repeating a moderately sized subset a moderate number of times maximally harms performance, and that repetition can waste up to 33% of compute (as measured by compute-equivalent loss).

0 favorites 0 likes
#deduplication

@pauliusztin_: I spent months optimizing GraphRAG retrieval. But it turned out I was optimizing the wrong thing.... The biggest knowle…

X AI KOLs Timeline · 2026-06-10 Cached

A detailed guide on optimizing knowledge graph ingestion for AI agents, presenting a five-step pipeline (extraction, resolution, embedding, deduplication, routing) to prevent graph corruption and improve retrieval quality.

0 favorites 0 likes
#deduplication

@pauliusztin_: 2 months ago, I started building unified memory layers with knowledge graphs. Here’s the most common question I’ve been…

X AI KOLs Timeline · 2026-05-23 Cached

This thread discusses best practices for building unified memory layers with knowledge graphs, emphasizing the separation of entity resolution (naming) from deduplication (identity) to avoid graph corruption. It also highlights using orchestration tools like PrefectIO to manage expensive LLM extraction pipelines with checkpointing and caching.

0 favorites 0 likes
#deduplication

@ClementDelangue: Great to see @CommonCrawl using and recommending @huggingface Buckets for large constantly evolving training datasets! …

X AI KOLs Following · 2026-05-22 Cached

Hugging Face announces Storage Buckets, a storage solution for large, evolving training datasets with built-in CDN and deduplication, recommended by CommonCrawl.

0 favorites 0 likes
#deduplication

Velonus – Open-source AppSec scanner that deduplicates SAST noise

Hacker News Top · 2026-05-15 Cached

Velonus is an open-source AppSec scanner for Python that runs five security tools in one command, normalizes findings, and deduplicates noise, with support for SARIF output and CI integration.

0 favorites 0 likes
← Back to home

Submit Feedback