Tag
A tip emphasizing not to discard raw conversation data after extracting facts, as the original context may be valuable for future use or analysis.
OriginBlame is a record- and token-level data provenance system that propagates author identity through AI training data pipelines, enabling precise forget sets for machine unlearning. It eliminates over-deletion from dataset-level systems and improves unlearning effectiveness.
A command-line CSV processing tool 'xan' written in Rust, capable of fast, low-memory processing of large CSV files, with built-in expression language, terminal visualization, and social science analysis extensions.
Ray 2.56 is released with stability improvements for Ray Data and a re-architecture of Ray Serve for better LLM serving performance.
Ray 2.56 has been released with improvements to Ray Data, Ray Serve for LLMs, GPU-domain-aware placement groups, and Kubernetes integration.
A technical guide on implementing a custom query language (EHQL) using Python and Apache Spark, with a focus on grammar definition and parsing using Lark.
DataClaw0 proposes an agentic data tailoring paradigm that uses learnable data processing to structure high-entropy multimodal streams, achieving robust alignment via SFT and GRPO on a novel benchmark.
xan is a high-performance command-line CSV processing tool written in Rust, featuring SIMD parsing, parallel computation, and terminal-based data visualization capabilities.
Flatiron is a pure Clojure columnar analytics library for in-memory tables with a SQL-like DSL, designed for performance with primitive arrays and batch processing.
Refiner is an open-source engine from Macrodata Labs for converting raw robotics and multimodal data into high-quality datasets for model training, with local and cloud execution.
Today, Macrodata Labs announced its launch, along with Refiner, an open-source framework for processing robotics datasets. The framework aims to help teams extract more signal from demonstrations and sensor data.
LOTUSPlan is a new API and optimizer for LLM-based data processing that reduces cost by up to 2.4× and improves accuracy by 4.6× through lazy execution and global planning. Developed at Berkeley and Stanford, it supports tasks like agent trace analysis, RAG, and document extraction.
An article exploring privacy concerns with AI tools that read screens, questioning whether screen content leaves the user's machine and the need for local-only processing or clear disclosures.
Learn how to set up and use Common Crawl data locally for web data processing tasks.
The DataLab team is orchestrating AI models across thousands of GPUs to process approximately one billion pages this week, highlighting significant large-scale document processing capabilities.
OpenDataLoader-PDF is an open-source PDF parsing tool that achieves a high accuracy rate of 0.907 in tests with real academic papers. It efficiently converts complex PDF documents (including tables, formulas, and scanned images) into Markdown and JSON, making it ideal for local knowledge bases and RAG applications.
Developer praises ml-intern tool for streamlining model/dataset discovery, post-training iteration and data workflows.