structured-data

Tag

Cards List
#structured-data

Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents

arXiv cs.AI · 2d ago Cached

This paper introduces persistent discovery context, a lightweight memory layer for data-centric agents that stores prior intent-to-object mappings to enhance retrieval quality across tasks, demonstrating improvements in structured data environments.

0 favorites 0 likes
#structured-data

@freeCodeCamp: Getting structured data your app can actually trust can be tricky. In this tutorial, Vineeth explains how to design sch…

X AI KOLs Timeline · 2026-08-29 Cached

This tutorial from freeCodeCamp explains how to design schemas, validate outputs, and handle failures to reliably extract structured data from LLMs, covering techniques like constrained outputs, retry loops, and streaming.

0 favorites 0 likes
#structured-data

@jerryjliu0: We've introduced native, agentic spreadsheet extraction into LlamaParse. Spreadsheets are a wildly different format fro…

X AI KOLs Timeline · 2026-08-27 Cached

LlamaParse has introduced native agentic spreadsheet extraction using a tuned model and harness for schema-guided extraction, enabling conversion of dense sheets like balance sheets into clean structured fields.

0 favorites 0 likes
#structured-data

@tom_doerr: HyperExtract converts unstructured documents into structured knowledge graphs, hypergraphs, and lists using large langu…

X AI KOLs Timeline · 2026-08-27 Cached

HyperExtract is an LLM-powered framework that converts unstructured documents into structured knowledge graphs, hypergraphs, and lists, simplifying knowledge extraction and management.

0 favorites 0 likes
#structured-data

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

arXiv cs.CL · 2026-08-24 Cached

This paper introduces a structured approach to extracting persona information for LLM-based digital twins, showing that structured representations improve predictive accuracy over raw transcripts. An automatic pipeline is proposed to adapt structures to different tasks.

0 favorites 0 likes
#structured-data

@llama_index: Most parsers treat tracked changes as noise. Result: a deleted clause comes back as live text, and your pipeline reads …

X AI KOLs Timeline · 2026-08-18 Cached

LlamaParse now handles revision tracking in documents, providing clean markdown of the final state and structured data for edits, deletions, and comments, addressing issues where parsers misinterpret tracked changes.

0 favorites 0 likes
#structured-data

@XAMTO_AI: AnakinScraper OSS – Open-source web scraping API built for AI. Convert any website to clean Markdown or structured JSON with a single click, directly for use in RAG and AI agents. Highlights: Anti-detection browser (Camoufox Firefox), Smart proxy…

X AI KOLs Timeline · 2026-08-18 Cached

AnakinScraper OSS is an open-source web scraping API designed for AI applications, converting websites to clean Markdown or structured JSON for use in RAG pipelines and AI agents.

0 favorites 0 likes
#structured-data

HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

arXiv cs.CL · 2026-08-17 Cached

HERMES is a scalable multi-agent framework for extracting structured knowledge from ultra-long scientific documents in geoscience, achieving high accuracy and sixfold efficiency improvement over manual methods.

0 favorites 0 likes
#structured-data

@llama_index: You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from…

X AI KOLs Following · 2026-08-03 Cached

LiteParse now supports extracting structured data from PDFs—form fields, checkbox states, annotations, images, vector graphics, and word-level bounding boxes—without a vision model, plus complexity signals to route harder pages to tools like LlamaParse.

0 favorites 0 likes
#structured-data

@DanKornas: Adding a conversational interface to a website should not require rebuilding its content layer. NLWeb is a collection o…

X AI KOLs Timeline · 2026-07-28 Cached

NLWeb is an open-source collection of protocols and Python tools that allows developers to add natural-language interfaces to websites without rebuilding the content layer, using existing Schema.org or RSS data.

0 favorites 0 likes
#structured-data

Microformats – building blocks for data-rich web pages

Lobsters Hottest · 2026-07-25 Cached

An article explaining how to consume and use microformats 2 data on personal websites, covering parser selection, fetching considerations, and data storage strategies.

0 favorites 0 likes
#structured-data

Authentication isn't authorization — how should authz work when agents talk to agents?

Reddit r/AI_Agents · 2026-07-16

The article argues that in agent-to-agent communication, authentication alone is insufficient for authorization; instead, structured, inspectable claims about intent, identity, and authority are needed, with the human remaining the final authority.

0 favorites 0 likes
#structured-data

@github: With GitHub Issue Fields, you can add structured, typed metadata (like priority, effort, dates, and custom values) to i…

X AI KOLs Timeline · 2026-07-12 Cached

GitHub Issue Fields are now generally available, enabling users to add structured metadata such as priority, effort, and dates to issues across repositories.

0 favorites 0 likes
#structured-data

Launch HN: Context.dev (YC S26) – API to get structured data from any website

Hacker News Top · 2026-07-09 Cached

Context.dev is a YC-backed API that allows developers and AI agents to scrape, crawl, and extract structured data from any website, with features like markdown, HTML, sitemaps, screenshots, and brand intelligence, aiming to simplify web data integration.

0 favorites 0 likes
#structured-data

@lionel_mora: Following the amazing reaction to the Marble Curriculum yesterday, we've decided to make it open source Everything a ch…

X AI KOLs Timeline · 2026-07-08 Cached

The Marble Curriculum, a comprehensive open-source dataset covering 1,590 primary school concepts with 3,221 connections across 8 subjects, has been released to enable building learning paths and AI-driven educational tools.

0 favorites 0 likes
#structured-data

How do you get a local agent to read CVEs and SEC filings without losing your mind

Reddit r/AI_Agents · 2026-07-03

A practical solution using AnySearch to enable local AI agents to efficiently query multiple specialized sources (CVEs, SEC filings) and return structured JSON/Markdown, avoiding rate limits and broken SDKs.

0 favorites 0 likes
#structured-data

Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

arXiv cs.CL · 2026-07-03 Cached

Object Aligner is an open-source Python library that deterministically scores two JSON objects by recursively aligning their trees, using Hungarian algorithm for unordered collections and sequence alignment for ordered ones. It introduces referential alignment for graphs/hypergraphs and can be used as a reward function in LLM prompt optimization.

0 favorites 0 likes
#structured-data

Evolutionary Feature Engineering for Structured Data

arXiv cs.LG · 2026-07-03 Cached

Introduces Evolutionary Feature Engineering (EFE), a framework that uses LLM-based evolution to automatically discover preprocessing transformations for structured data, improving time-series forecasting and tabular prediction accuracy while preserving interpretability.

0 favorites 0 likes
#structured-data

SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction

arXiv cs.AI · 2026-07-02 Cached

SchemaRAG is a retrieval-augmented generation framework that dynamically reduces the output schema space for LLM-driven structured information extraction, achieving improved performance and efficiency on healthcare and e-commerce datasets.

0 favorites 0 likes
#structured-data

From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

arXiv cs.CL · 2026-06-26 Cached

Presents a systematic methodology for converting Hindi WordNet into 1.25 million instruction-response pairs to fine-tune a 12B-parameter language model using LoRA, demonstrating improved pedagogical effectiveness for specialized conversational systems in low-resource languages.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback