structured-extraction

Tag

Cards List
#structured-extraction

@DataScienceDojo: Google's π₯𝐚𝐧𝐠𝐞𝐱𝐭𝐫𝐚𝐜𝐭 has crossed 37k stars on GitHub. The core idea: point an LLM at unstructured text and g…

X AI KOLs Timeline β†— Β· yesterday Cached

Google's open-source tool 'langextract' uses LLMs to extract structured data from unstructured text with grounded character positions, crossing 37k GitHub stars.

0 favorites 0 likes
#structured-extraction

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

arXiv cs.AI β†— Β· 2026-07-10 Cached

This paper explores distilling a large reasoning teacher model (8B) into a small student model (0.6B) for on-device structured text enrichment, achieving substantial speedup while recovering significant quality. The study finds that the reasoning nature of the teacher, rather than its scale, drives improvement in summary quality, but a same-size instruction teacher yields more faithful outputs on certain articles.

0 favorites 0 likes
#structured-extraction

From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control

arXiv cs.CL β†— Β· 2026-07-07 Cached

This paper presents an automated pipeline that uses the DeepSeek V3 model to decompose Italian tax-court judgments into individual legal issues structured in XML following the IRAC framework, and includes a hallucination-detection filter using the Linkoln parser to validate citations, validated by expert annotators.

0 favorites 0 likes
#structured-extraction

datalab-to/lift

Hugging Face Models Trending β†— Β· 2026-06-19 Cached

Datalab releases lift, a model that extracts structured JSON from PDFs and images using schema-constrained decoding, with local and hosted inference options.

0 favorites 0 likes
#structured-extraction

I made a small local model (llama3.2 3B) reliably extract structured JSON from documents - the hard part wasn't the model, it was everything around it

Reddit r/AI_Agents β†— Β· 2026-06-05

A developer shares lessons from building a local document-to-JSON extractor using llama3.2 3B on Ollama, highlighting that deterministic post-processing and schema-constrained outputs matter more than model size, while seeking feedback on hallucination and context truncation issues with long documents.

0 favorites 0 likes
#structured-extraction

@Michaelzsguo: Today I upgraded my Hermes agents with TencentDB Agent Memory. I did not connect it to a cloud LLM. Instead, I wired it…

X AI KOLs Timeline β†— Β· 2026-05-24 Cached

The author upgraded their Hermes agents with TencentDB Agent Memory, using a local Qwen 3.5-4B model via llama-server for structured JSON extraction and multi-step tool use, implementing a resilient layered memory pipeline with cursor-based checkpointing.

0 favorites 0 likes
#structured-extraction

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable) [P]

Reddit r/MachineLearning β†— Β· 2026-05-22

Numind released NuExtract3, a 4B open-weight vision-language model based on Qwen3.5-4B, designed for converting document images to Markdown, OCR, and structured data extraction. It is Apache-2.0 licensed and self-hostable with quantized versions for low VRAM.

0 favorites 0 likes
#structured-extraction

numind/NuExtract3

Hugging Face Models Trending β†— Β· 2026-04-29 Cached

NuExtract3 is a 4B vision-language reasoning model for document understanding, enabling structured extraction and image-to-Markdown conversion.

0 favorites 0 likes
← Back to home

Submit Feedback