data-processing

Tag

Cards List
#data-processing

Don't throw away the raw conversation after extracting facts

Reddit r/AI_Agents · 12h ago

A tip emphasizing not to discard raw conversation data after extracting facts, as the original context may be valuable for future use or analysis.

0 favorites 0 likes
#data-processing

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

arXiv cs.AI · 6d ago Cached

OriginBlame is a record- and token-level data provenance system that propagates author identity through AI training data pipelines, enabling precise forget sets for machine unlearning. It eliminates over-deletion from dataset-level systems and improves unlearning effectiveness.

0 favorites 0 likes
#data-processing

@QingQ77: A command-line CSV processing tool written in Rust ("CSV Magician"), capable of extremely fast, low-memory processing of CSV files of several GB, with built-in expression language, terminal visualization, and social science-oriented extensions such as dictionary statistics, graph theory, and even scraping. https://github.com/medialab/…

X AI KOLs Timeline · 2026-07-12 Cached

A command-line CSV processing tool 'xan' written in Rust, capable of fast, low-memory processing of large CSV files, with built-in expression language, terminal visualization, and social science analysis extensions.

0 favorites 0 likes
#data-processing

@robertnishihara: Try Ray 2.56!

X AI KOLs Following · 2026-06-30 Cached

Ray 2.56 is released with stability improvements for Ray Data and a re-architecture of Ray Serve for better LLM serving performance.

0 favorites 0 likes
#data-processing

@raydistributed: We just released Ray 2.56! This includes - Ray Data stability improvements: reduced object store spilling, automatic ba…

X AI KOLs Following · 2026-06-30

Ray 2.56 has been released with improvements to Ray Data, Ray Serve for LLMs, GPU-domain-aware placement groups, and Kubernetes integration.

0 favorites 0 likes
#data-processing

Implementing a Custom Query Language with Python and Apache Spark

Lobsters Hottest · 2026-06-23 Cached

A technical guide on implementing a custom query language (EHQL) using Python and Apache Spark, with a focus on grammar definition and parsing using Lark.

0 favorites 0 likes
#data-processing

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

Hugging Face Daily Papers · 2026-06-19 Cached

DataClaw0 proposes an agentic data tailoring paradigm that uses learnable data processing to structure high-entropy multimodal streams, achieving robust alignment via SFT and GRPO on a novel benchmark.

0 favorites 0 likes
#data-processing

Data Visualization from the Comfort of your Terminal

Lobsters Hottest · 2026-06-17 Cached

xan is a high-performance command-line CSV processing tool written in Rust, featuring SIMD parsing, parallel computation, and terminal-based data visualization capabilities.

0 favorites 0 likes
#data-processing

A columnar database for analytics in pure Clojure

Lobsters Hottest · 2026-06-12 Cached

Flatiron is a pure Clojure columnar analytics library for in-memory tables with a SQL-like DSL, designed for performance with primitive arrays and batch processing.

0 favorites 0 likes
#data-processing

Refiner: Robotics library from the ex-Hugging Face pre-training team

Reddit r/LocalLLaMA · 2026-06-11 Cached

Refiner is an open-source engine from Macrodata Labs for converting raw robotics and multimodal data into high-quality datasets for model training, with local and cloud execution.

0 favorites 0 likes
#data-processing

@gui_penedo: Today we’re announcing Macrodata Labs. Over the last few years, @HKydlicek and I have been turning a large part of the …

X AI KOLs Following · 2026-06-11 Cached

Today, Macrodata Labs announced its launch, along with Refiner, an open-source framework for processing robotics datasets. The framework aims to help teams extract more signal from demonstrations and sensor data.

0 favorites 0 likes
#data-processing

@lianapatel_: Beyond excited to share we're releasing LOTUSPlan, a new API & optimizer for higher performance LLM-powered data proces…

X AI KOLs Following · 2026-06-10 Cached

LOTUSPlan is a new API and optimizer for LLM-based data processing that reduces cost by up to 2.4× and improves accuracy by 4.6× through lazy execution and global planning. Developed at Berkeley and Stanford, it supports tasks like agent trace analysis, RAG, and document extraction.

0 favorites 0 likes
#data-processing

With screen-aware AI the privacy question isn't just ""what does it see."" It's where what it sees goes.

Reddit r/ArtificialInteligence · 2026-05-29

An article exploring privacy concerns with AI tools that read screens, questioning whether screen content leaves the user's machine and the need for local-only processing or clear disclosures.

0 favorites 0 likes
#data-processing

@lhoestq: You don't know you actually need local Common Crawl

X AI KOLs Timeline · 2026-05-22 Cached

Learn how to set up and use Common Crawl data locally for web data processing tasks.

0 favorites 0 likes
#data-processing

@VikParuchuri: We'll process ~1B pages this week. The team at @datalabto has done incredible work orchestrating our models across thou…

X AI KOLs Following · 2026-05-11 Cached

The DataLab team is orchestrating AI models across thousands of GPUs to process approximately one billion pages this week, highlighting significant large-scale document processing capabilities.

0 favorites 0 likes
#data-processing

@rwayne: Absolutely impressive for building local knowledge bases with academic papers—the bottleneck has always been cleanly converting PDFs to Markdown. OpenDataLoader-PDF achieves a 0.907 accuracy rate, ranking first on the open-source PDF parsing leaderboard, all under Apache 2.0. Key metrics from a test set of 200 real papers: Overall score 0…

X AI KOLs Timeline · 2026-05-10

OpenDataLoader-PDF is an open-source PDF parsing tool that achieves a high accuracy rate of 0.907 in tests with real academic papers. It efficiently converts complex PDF documents (including tables, formulas, and scanned images) into Markdown and JSON, making it ideal for local knowledge bases and RAG applications.

0 favorites 0 likes
#data-processing

@cmpatino_: I’ve been using ml-intern for a while, and it genuinely changed my workflow. It's super good at: - Model/Dataset discov…

X AI KOLs Following · 2026-04-21 Cached

Developer praises ml-intern tool for streamlining model/dataset discovery, post-training iteration and data workflows.

0 favorites 0 likes
← Back to home

Submit Feedback