This thesis presents efficient and accurate systems for querying unstructured data, focusing on methods to improve data retrieval and analysis.
<p>Abstract: (shortened)</p>
<p>In recent years, automatic analysis over this unstructured data has become possible via machine learning (ML). Analysts can use ML to extract structured information from these unstructured sources, such as object types and location from a video. The structured information can subsequently be used in downstream analysis, e.g., the urban planner can count the number of cars that passed by an intersection.</p>
<p>Unfortunately, using ML for these analyses is challenging. Deploying ML is prohibitively expensive for many organizations: naively analyzing a year of video from a small town can cost millions in cloud compute credits. ML methods are also unreliable, returning incorrect results, which can lead to downstream errors. Finally, deploying ML for analytics requires knowledge of deep learning, data systems, programming, and other technical skills.</p>
<p>In light of these challenges, we make two observations: many applications can tolerate approximations, if there are guarantees on accuracy, and methods for answering unstructured data queries range by up to 10 orders of magnitude in cost.
In this dissertation, we develop systems and algorithms for efficient and reliable unstructured data analytics, leveraging the two observations. Instead of returning exact answers, we return approximate answers generated by cheap approximations to expensive ML methods.</p>
<p>Our systems can return statistically valid answers on a wide range of query types, including selection, aggregation, and limit queries. Furthermore, our systems can be up to orders of magnitude cheaper than standard methods of answering queries</p>
<p><a href="https://lobste.rs/s/v8atna/efficient_accurate_systems_for_querying">Comments</a></p>
Source: https://stacks.stanford.edu/file/fk030tb6783/thesis-augmented.pdf
This appears to be a corrupted or improperly extracted text file containing extensive garbled characters and encoding artifacts, making the content largely unreadable. The original document was likely a PDF thesis or paper from Stanford University, as indicated by the URL.
Given the severe text corruption, a meaningful translation from Chinese to English is not possible. The provided text does not contain coherent Chinese language content suitable for translation. It consists mainly of random character sequences, symbols, and broken encoding patterns.
If you have access to the original, uncorrupted PDF file, please provide that for an accurate translation.
The paper proposes agentic data cracking, a method to adaptively structure unstructured data during LLM reasoning to reduce token consumption and costs, achieving significant cost cuts while maintaining accuracy on benchmarks.
This paper proposes using Hyperdimensional Computing, specifically Holographic Reduced Representations, to embed tabular data rows for structured querying, enabling interpretable similarity thresholds and zero-match detection, outperforming a baseline method on row retrieval tasks.
This paper presents a system that automatically discovers an executable schema from raw multi-source data and uses it for knowledge graph construction and query-time retrieval, improving over baselines on QA benchmarks.
This paper introduces HybridDeepResearch, a benchmark for evaluating AI agents on tasks that require integrating web search and SQL querying, revealing that even state-of-the-art models struggle with maintaining constraints across structured and unstructured data.