An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case
Summary
This paper proposes an agentic framework combining rules and LLMs for embedding and annotating document layouts in plant science, demonstrating scalable trait extraction with improved annotation coverage.
View Cached Full Text
Cached at: 08/18/26, 09:49 AM
# An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case Source: [https://arxiv.org/abs/2608.14587](https://arxiv.org/abs/2608.14587) [View PDF](https://arxiv.org/pdf/2608.14587) > Abstract:Background: Recent advances in information retrieval \(IR\) leverage both dense and sparse representations, large language models \(LLMs\), and specialized retrieval models to improve ranking accuracy, relevance, and cross\-lingual performance\. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation further enhance retrieval effectiveness by capturing fine\-grained contextual and structural information\. Emerging agentic LLM frameworks extend these capabilities by enabling planning, iterative reasoning, tool use, and multi\-agent collaboration, thereby broadening applications across diverse domains\. These frameworks also emphasize rigorous evaluation, ethical considerations, and trustworthiness, ensuring responsible deployment in real\-world settings\. We propose a modular, agent\-based pipeline for botanical trait extraction\. Optical character recognition \(OCR\) converts PDFs into machine\-readable text, while segmentation and indexing organize content by genus and species\. Rule\-based parsers extract structured botanical traits, and ensembles of large language models \(LLMs\) expand trait vocabularies and resolve ambiguities\. This approach ensures accurate species recognition, scalable annotation, and explainable integration of textual botanical descriptions, enabling robust and interpretable data extraction across large botanical corpora\. Results: Using three regional botanical datasets, our system extracted 55,737 trait annotations across 4,961 species, averaging 9\.1 traits per species\. Integration of LLM\-based enrichment improved coverage for 75% of traits, increasing total annotations by 59%\. While the choice of OCR engine had a minor effect on species recognition, overall annotation counts remained stable, demonstrating the robustness, scalability, and reliability of the pipeline for large\-scale botanical trait extraction\. ## Submission history From: Nicolas Turenne \[[view email](https://arxiv.org/show-email/c0eda81f/2608.14587)\] **\[v1\]**Fri, 19 Jun 2026 11:50:27 UTC \(2,002 KB\)
Similar Articles
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
This paper presents an LLM-as-a-Judge evaluation framework for agentic AI in drug discovery, validated through human alignment studies with expert annotators. It optimizes the judge to improve alignment with human judgment and provides insights for reusable evaluation in scientific domains.
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
This paper proposes an agentic LLM framework for automated structural analysis of 3D frame systems from natural language inputs, achieving 90% accuracy on ten representative 3D frames through a multi-agent pipeline.
Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes
This paper demonstrates that frontier LLM-based agents, when provided with source publications, annotation guides, and ontologies, can perform phenotype annotation at a level comparable to trained human curators, substantially outperforming previous NLP tools and overcoming a key bottleneck in comparative morphological data integration.
Learning to Construct Practical Agentic Systems
This paper proposes principled approaches for designing and optimizing practical agentic LLM systems, introducing a framework with pseudo-tools and fixed workflows to improve modularity, cost-efficiency, and accuracy across diverse tasks.
Exploring Agentic Workflows for Generating High Quality Math Visual Aids
This paper introduces an agentic workflow that uses LLMs and VLMs to iteratively generate and improve high-quality mathematical diagrams for K-12 education, addressing the reliability gap in AI-generated visual aids.