Tag
The paper introduces a capability-driven data infrastructure with curriculum scheduling to train generalist image generation models using heterogeneous supervision for diverse generative tasks.
A report based on a survey of 300 data and technology executives examines how legacy data systems limit the effectiveness and scaling of AI agents in enterprises, highlighting that 'data leaders' who give agents broader data access experience greater trust and success.
This paper introduces Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR principles for AI-agent-mediated nutrition research, addressing data identity, search, and crosswalk challenges. It demonstrates strong benchmark results and improved reproducibility compared to open-web reconstruction.
This paper argues that Latin America lacks the datasets and benchmarks layers needed to build its own AI, and proposes DataHub, an open, incentive-driven platform with a task-first ontology to index and contribute regional datasets.
At the Future of Memory and Storage conference, NVIDIA announced open sourcing its cuFile APIs for GPU-direct storage and featured its Vera CPU delivering up to 3.21x higher throughput than x86 in compression and encryption pipelines, addressing AI's growing storage demands.
Observations from conversations with healthcare payors indicate a shift in focus from AI models to data readiness, PHI handling, and integration across disparate systems.
An interview with Pierre Zemb, staff engineer at Clever Cloud, discussing his work building data layers on FoundationDB and his previous experience at OVHcloud.
A Twitter thread highlights key takeaways from a Latent.Space podcast episode with Databricks co-founders, covering why Databricks beat Snowflake, the rise of metaharners, Neon's success, HTAP via LTAP, MosaicML's fate, and maintaining startup culture in a large company.
Robotics teams are rebuilding the data stack from scratch to overcome the 'data layer tax' that slows down iteration and scaling in robot learning, as existing infrastructure doesn't handle multi-rate and multimodal data.
Anthropic's science blog argues that AI progress in biology lags behind coding because biological data infrastructure is not designed for agents. A case study shows that adding a deterministic retrieval layer (gget virus) boosts accuracy to nearly 100%.
Despite $2.5 trillion in projected global AI spending in 2026, MIT's NANDA Initiative reports 95% of enterprise generative AI projects deliver zero measurable ROI, with a practitioner's first-hand analysis of 14 engagements pointing to misallocated budgets favoring model work over data infrastructure as the root cause.
Jerry Liu, CEO of LlamaIndex, discusses on the Venture with Grace podcast why data infrastructure is crucial for the agentic AI boom, emphasizing that AI agents need access to the right data at the right time.
Judgment Labs is launching today with a $32M funding round, providing infrastructure to improve AI agents using production data.
The article argues that AI inference poses unique challenges to cloud data infrastructure, likening its demand to high-concurrency OLTP systems rather than traditional human-speed applications. It emphasizes the need to optimize storage and data access layers to handle the 'AI data tsunami' driven by autonomous agents.
Anthropic researcher Laura Luebbert argues that biological data infrastructure needs to be redesigned for AI agents, using a case study where even strong models failed to reliably retrieve sequence data from NCBI Virus until a deterministic retrieval layer was added.