Tag
SCTab-Diff is a semantics-consistent tabular diffusion framework that uses weak semantic priors to generate high-fidelity synthetic tabular data, improving distributional fidelity and semantic consistency over existing methods.
The author conducted a small experiment on LlamaIndex's extractbench benchmark, comparing Claude Opus 5 and Qwen models, finding similar performance with cost advantages and a correlation between extraction difficulty and document length.
The author argues that agent memory layers should skip LLM-based extraction for deciding what to remember, instead using simple storage, embeddings, and retrieval, exemplified by their open-source memU tool.