Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations
Summary
The paper presents FAIRY, a full-stack smart-agriculture agent system developed and deployed for managing soybean farm operations, with evaluation in a real-world setting.
View Cached Full Text
Cached at: 09/02/26, 05:57 AM
# Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations Source: [https://arxiv.org/html/2609.00106](https://arxiv.org/html/2609.00106) DOI:[XXXXXXX\.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)Conference:Accepted to the 34th ACM International Conference on Advances in Geographic Information Systems; Nov; 2026ISBN:978\-1\-4503\-XXXX\-X/2018/06CCS:Computing methodologies Intelligent agentsCCS:Applied computing AgricultureAo Qu[https://orcid.org/0009-0008-8230-3211](https://orcid.org/0009-0008-8230-3211)Note:Equal contribution\.email:[aoqu@stu\.hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,ChinaPanagiotis Michelakis[https://orcid.org/0009-0003-0498-0499](https://orcid.org/0009-0003-0498-0499)Note:Currently with new affiliation\.email:[panosg@synkrasis\-labs\.com](mailto:[email protected])Affiliation:School of Electrical and Computer Engineering,National Technical University of Athens,Athens,Greece,Linyuan Han[https://orcid.org/0009-0009-2041-7202](https://orcid.org/0009-0009-2041-7202)email:[linyuanhan26@stu\.hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China,Yiannis Hadjiyianni[https://orcid.org/0009-0003-2413-6375](https://orcid.org/0009-0003-2413-6375)email:[yiannisha@synkrasis\-labs\.com](mailto:[email protected])Affiliation:School of Electrical and Computer Engineering,National Technical University of Athens,Athens,Greece,Kun Ouyang[https://orcid.org/0009-0006-4754-9625](https://orcid.org/0009-0006-4754-9625)email:[kunouyang@stu\.hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China,Konstantinos Siskos[https://orcid.org/0009-0000-0596-2522](https://orcid.org/0009-0000-0596-2522)email:[siskos@synkrasis\-labs\.com](mailto:[email protected])Affiliation:School of Electrical and Computer Engineering,National Technical University of Athens,Athens,Greece,Feng Li[https://orcid.org/0009-0008-3995-0931](https://orcid.org/0009-0008-3995-0931)Note:Corresponding authors\.email:[feng\.li@hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China,Ran Meng[https://orcid.org/0000-0003-4756-9934](https://orcid.org/0000-0003-4756-9934)email:[mengran@hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China,Jingchi Jiang[https://orcid.org/0000-0003-2167-4082](https://orcid.org/0000-0003-2167-4082)email:[jiangjingchi@hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China,Dimitrios Stamoulis[https://orcid.org/0000-0003-1682-9350](https://orcid.org/0000-0003-1682-9350)Note:Project Lead:FAIRY Platform, smart\-farm multi\-agent system and world models\.email:[dimi@hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,ChinaandJie Liu[https://orcid.org/0000-0001-6209-6886](https://orcid.org/0000-0001-6209-6886)email:[jieliu@hit\.edu\.cn](mailto:[email protected])Affiliation:Faculty of Computing,Harbin Institute of Technology,State Key Laboratory of Smart Farm Technologies and Systems,Harbin,China © none ###### Abstract\. This paper presents FAIRY, a full\-stack smart\-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology’s smart\-agriculture site\. We develop FAIRY to execute and evaluate agentic agronomic operations on full\-season*spatiotemporal*workflows that span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage\. FAIRY integrates APIs and infrastructure across production\-grade machinery, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop\-process models, agronomic records, and multi\-season yield histories\. The system is built around the novel “everything is an event” execution paradigm, which represents spatiotemporal world evolution, remote sensing and UAV observations, sensor readings, crop\-growth transitions, machinery actions, and management interventions as state\-changing events in a shared farm process engine\. On top of this event\-drivenworld model, FAIRY implements a complete agentic stack: a knowledge library of atomic agronomic skills; multi\-agent controller and orchestration backends; frontier\- and edge\-model execution; full\-path trace logging; and deployment profiling on local nodes\. We use FAIRY to evaluate nine state\-of\-the\-art agent controllers across one hundred full\-season soybean scenarios that preserve the operational coupling between spatial observations in a 64\-ridge field, temporal decision sequences, agronomic constraints, delayed effects, and final yield\. We develop an evaluation suite that combines agentic success,full\-path spatiotemporalcorrectness, token cost, and edge\-device runtime\. Our results show that ourspatiotemporallygrounded Kendall correctness \(KTC\) improves alignment with downstream yield outcomes compared to existing order\-only and exact\-match metrics\. Our analyses further show that hierarchical agronomic skills and expert operational context substantially improve long\-horizon behavior compared to geospatial in\-context learning and LLM\-as\-an\-Expert schemes\. ###### Keywords: Spatiotemporal agents, Smart agriculture, Agent evaluation, Geospatial workflows, Full\-season farm operations, World models ## 1\.Introduction Recent advances in agentic AI have made tool\-augmented language models increasingly useful for geospatial analysis, where tasks often require data discovery, API selection, model invocation, map operations, and multi\-step interpretation over spatial and temporal data\. Prior geospatial agent systems have shown this potential across remote sensing, Earth observation, urban analysis, forestry, climate studies, agriculture monitoring, and satellite\-vision workflows\([Lee et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib31);[Bhattaram et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib30)\)\. As reflected by recent releases in commercial geospatial platforms, such as Google Earth AI and Microsoft Planetary Computer offerings for agriculture and environmental monitoring, these settings are a natural fit for agents because geospatial workflows require selecting appropriate imagery or products, applying spatial filters, invoking detection or classification models, reasoning over intermediate outputs, and producing map\-based results\. Figure 1\.FAIRY deploys and evaluates agentic controllers over a full soybean season \(planting through grain storage\) on a 64\-ridge operating research farm, coupling spatial observations, temporal decisions, delayed agronomic effects, and final yield\.Full\-process overview of the FAIRY agentic farm system\.Despite these advances, deploying agentic systems for agricultural operations and complex agronomic spatiotemporal workflows remains difficult\. Farm operations differ in three concrete ways\([Yan et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib36);[Zhang et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib19);[Xu et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib20);[Seo and Lee, 2026](https://arxiv.org/html/2609.00106#bib.bib21);[Zuzuárregui et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib22);[Qu et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib1)\):Feedback is delayed:a poor planting decision surfaces as weak emergence days later; a missed irrigation window appears as yield loss months later\.Observations are partial and spatial:sensors are zone\-level across ridges, drones are weather\-gated, and ground inspection is sparse, so the agent reasons over a partially observed*spatial*field\.Consequences propagate:operational errors compound across the season, so an action can be syntactically valid and operationally wrong, such as harvesting before grain moisture is suitable\. At the same time, practical deployments of geospatial agents often inherit orchestration and evaluation practices originally developed for domains such as web automation, coding, or OS control, where feedback is immediate and the action space is bounded by application APIs\. However, unlike cloud\-centric remote\-sensing and geospatial big\-data pipelines, agricultural applications introduce a broader evaluation surface: Crop state evolves through biological growth, weather, soil\-water dynamics, sensing constraints, and management interventions, so agent decisions are coupled to physical processes that unfold over days, weeks, and seasons\. Ultimately, agent performance depends not only on successful tool calls, but also on spatial coverage, temporal alignment, data\-product choice, and the operational meaning of the produced analysis\. In this work, we present a full\-stack smart\-agriculture agent system developed for an operating soybean research farm at a university smart\-agriculture research site \(Figure[1](https://arxiv.org/html/2609.00106#S1.F1)\) spanning a 64\-ridge field\. Real\-life field operations span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage\. The site has production\-grade machinery; a digitized sensing layer covering fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, and a weather station; a calibrated, physics\-grounded process model; and multi\-season historical harvest and yield records\. We are preparing a controlled portion of this field to be operated by agents alongside the existing human\-operated workflow, with the goal of comparing agent\-managed and human\-managed operations in upcoming harvest seasons\. This makes the evaluation questionpractical: we need to understand how agent systems behave on real farm operations before deciding what to deploy in the field\. To this end, we develop FAIRY as a full\-stack agentic research engine for the operating farm\. FAIRY first integrates the farm\-facing infrastructure needed for deployment: production\-grade machinery, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop\-process models, agronomic records, and multi\-season yield histories\. The system then organizes these components through an “everything is an event” execution model, where weather updates, remote\-sensing observations, UAV inspections, sensor readings, crop\-growth transitions, machinery actions, and management interventions are represented as state\-changing events in a shared farm process engine\. On top of this event\-drivenworld model, FAIRY implements the agentic stack used in our evaluation: a retrieval\-based knowledge library of atomic agronomic skills, nine agent controller backends with optional agent\-to\-agent \(A2A\) orchestration, frontier\- and edge\-model execution, full\-path trace logging, and edge\-device deployment profiling\. Before deploying agents on the controlled farm portion, we need an evaluation setting that can exercise farm operations beyond the limited number of observed historical seasons\. The calibrated farm world model allows us to construct realistic scenarios that remain tied to our field geometry, crop\-process assumptions, sensing layer, and operational workflow, while also covering conditions that have not yet occurred in the deployment record\. Specifically, we build one hundred scenarios across three levels of complexity: atomic tasks, episode chains, and full\-season scenarios\. We develop a comprehensive evaluation suite which covers task success, temporal correctness, coverage, full\-path correctness, and yield preservation\. We conduct extensive metric\-calibration studies that assess which trace\-level metric best tracks yield, identifying temporally grounded Kendall correctness \(KTC\) as the best\-calibrated predictor and exposing the failure modes of order\-only and exact\-match alternatives\. Finally, we include an operational demonstration that runs the same event\-driven runtime end\-to\-end, from a user request through drone and satellite observation to a geospatial visualization in the FAIRY interface\. Our results lead to four practical lessons for deploying agents in farm operations\. First, domain expertise is the dominant lever: expert operational context reduces the full\-season yield shortfall from∼\\sim22% under zero context to∼\\sim5% for Qwen on held\-out L3 scenarios\. Second, agent and farm performance diverge under longer horizons: short tasks are nearly solved, with≥\\geq99% temporal correctness and near\-zero yield loss, while full\-season scenarios remain materially below the human oracle\. Third, multi\-agent A2A orchestration introduces coordination cost that degrades both correctness and yield in our setting\. Fourth, agriculture\-tuned LLM context is only a partial substitute for human\-expert context\. We report these as field observations from an operating farm; the engine, scenarios, and evaluation are documented for reproduction\. ## 2\.The FAIRY System: Architecture Overview An overview of FAIRY is shown in Figure[1](https://arxiv.org/html/2609.00106#S1.F1)\. The system couples an event\-driven engine, a physics\-grounded process stack, a tool/sensing layer, a knowledge library, and the agentic backend\. ### 2\.1\.Farm site and APIs Farm equipment and sensing infrastructure\.The study site is an industry\-grade university soybean research farm with a modeled268m×71m268\\,\\mathrm\{m\}\\times 71\\,\\mathrm\{m\}field organized into 64 ridges, which serve as the atomic*spatial*units for observation and intervention\. We expose farm capabilities to the agent as function\-calling tools derived from the installed APIs and operational interfaces\. As summarized in Table[1](https://arxiv.org/html/2609.00106#S2.T1), the site combines fixed soil, canopy, weather, light, radiation, and chlorophyll sensing; UAV and calibration assets for multispectral, thermal, and LiDAR observation; ground robotic platforms; production machinery; spraying equipment; and ridge\-level irrigation and fertilization facilities\. Table 1\.FAIRY agent APIs closely model and connect to the underlying farm equipment and on\-field sensing assets\.Satellite imagery and map products\.We use Google Earth Engine \(GEE\) to retrieve Sentinel\-2 Surface Reflectance Harmonized imagery over the target area of interest \(AOI\), which we resolve either from a GEE asset or from a local AOI ZIP/shapefile\. The retrieval pipeline filters scenes by AOI, date range, and cloudy\-pixel percentage, sorts candidate scenes by cloud coverage, and exports the selected image as a 10 m GeoTIFF in EPSG:4326\. We export bands B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12 through GEE export and Google Drive/API download logic\. Downstream, we clip the raster to the AOI and georegister map placement from the GeoTIFF bounds in EPSG:4326\. For visualization, we generate a satellite backdrop PNG from B4/B3/B2 RGB bands using percentile stretching and transparency masking, then overlay the result in the Leaflet\-based map UI\. Satellite crop classification\.We implement crop classification with an in\-house XGBoost multiclass model\. The classifier uses thegbtreebooster with amulti:softmaxobjective and takes the ten Sentinel\-2 bands together with vegetation indices including NDVI, kNDVI, GW1, GW2, LSWI, NDWI, EVI, EVI2, MSAVI, GNDVI, NDRE, GWCCI, and REP\. We extract labeled pixels from crop polygons using geometry masks and use a stratified 70/30 train\-test split\. The model supports both training from scratch and loading a pretrained checkpoint with optional fine\-tuning; in this study, we do not enable additional Heilongjiang regional fine\-tuning\. The operational output classes are rice, maize, soybean, other, and background\. Drone imagery and APIs\.We collect field imagery at regular intervals with farm personnel and are also experimenting with direct UAV control through DJI Cloud API\. The current deployment uses a DJI M300\-RTK equipped with a DJI P1 full\-frame RGB camera, producing 8192×\\times5460 imagery for field inspection\. We generate orthomosaic products through DJI Terra, with OpenDroneMap also supported as an agent\-callable processing backend\. For field identification, we extract field\-edge information using a multi\-self\-adaptive\-threshold Canny operator and use the excess green index \(ExG\) to filter plant\-covered regions\. For missing\-seedling detection, we combine overall plant density with local plant density estimated through sliding\-window counting to identify sparse or missing\-seedling areas\. The drone and examples of the imagery and collected results are illustrated in Figure[2](https://arxiv.org/html/2609.00106#acmlabel2)\. Figure 2\.On\-field drone inspections and UAV APIs\.Example visualizations of on\-field drone inspections\. ### 2\.2\.Growth process dynamics Weather model\.The farm world model is driven by weather and remote\-sensing observations that define the external conditions for crop growth and management\. Weather is generated through a WGEN/Richardson\-style daily weather model\([Siler and Singh, 2022](https://arxiv.org/html/2609.00106#bib.bib24)\), providing precipitation, temperature, radiation, wind, and humidity at the resolution required by the farm event loop\. The observation layer incorporates the remote\-sensing and drone\-based products available during operation\. We treat these observations as state updates that provide spatial evidence about canopy vigor, water stress, and anomaly regions for follow\-up inspection, irrigation, fertilization, or pest\-management decisions\. Growth model\.Plant growth dynamics are represented through a coupled soil, phenology, canopy, and biotic\-pressure stack following\([Wang et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib16)\)\. The soil component maintains water and temperature states using a bucket\-style water balance over precipitation, irrigation, runoff, drainage, and evapotranspiration\. Phenology follows a GDD\-based soybean development model with seed\-type\-specific maturity targets\([Akyuz et al\., 2017](https://arxiv.org/html/2609.00106#bib.bib23)\), supporting cultivar\-specific assumptions and stage\-dependent management rules\. The canopy and biomass model follows Monteith\-style radiation\-use and light\-interception principles\([Monteith, 1977](https://arxiv.org/html/2609.00106#bib.bib25);[Monsi and Saeki, 2005](https://arxiv.org/html/2609.00106#bib.bib26)\)\. The biotic\-pressure component tracks weed, insect, and disease pressure as weather\- and stage\-dependent processes with treatment effects\([Steduto et al\., 2009](https://arxiv.org/html/2609.00106#bib.bib27)\)\. Action effects and yield\.Together, these components define a process\-level state space in whichaction effectsdepend on crop stage, soil condition, recent weather, and observed canopy state\. Management actions modify the crop trajectory through delayed and stage\-dependent effects: Planting establishes stand fraction and emergence timing; irrigation changes soil\-water availability over subsequent days; fertilization affects nutrient stress and canopy development\. At maturity, we employ a yield\-recovery model to convert accumulated crop state into harvested grain\([Humburg, 2019](https://arxiv.org/html/2609.00106#bib.bib28)\)\. Physics\-engine calibration\.We validate the engine against historical data by reconstructing 18 plot\-level soybean scenarios from the 2025 growing season \(05\-10 to 05\-27\) and harvest campaign \(09\-13 to 09\-23\)\. Different scenario IDs span different planting dates, density treatments, and cultivar type\. Each FAIRY world model is initialized to closely match historical plot\-specific management operations, cultivar type, planting density, fertilization schedule, and measured daily weather\. We use observed phenology and yield records as evaluation targets, while the engine independently simulates emergence, crop growth, soil water and nutrient stress, biomass accumulation, harvest recovery, and final yield\. As shown in Figure[3](https://arxiv.org/html/2609.00106#acmlabel3), our engine yields closely track the observed yields across cultivars and treatments, with nearly all scenarios within yield prediction accuracies of2%2\\%MAE\. We leave integration with commercial tools, such as WOFOST, to future work\. Figure 3\.FAIRY Physics Engine yield vs\. real\-world farm yield across 18 Farm IDs\.FAIRY Physics Engine closely matches real\-world farm yields from historical harvest data\. ### 2\.3\.Stateful event\-driven engine We build the farm execution engine on the ARE framework\([Froger et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib18)\), which represents agent environments through stateful apps, an event queue, notifications, scenarios, and logged execution traces\. Farm operation is naturally event\-based: tools execute ridge preparation, planting, irrigation, fertigation, pesticide application, harvest, and storage as timestamped farm events, each with arguments \(e\.g\., operation duration, target ridges\) and preconditions \(e\.g\., weather and equipment readiness\)\. The environment maintains the farm state, event queue, notifications, and operation history, while the controller observes and acts through the tool interface\. Tool invocation is emulated inside the digital twin, so operations retain their farm\-state effects and timing dependencies without real\-clock execution\. This evaluation should therefore be interpreted as a deployment\-readiness study rather than a completed agent\-managed harvest trial\. The physical infrastructure, sensing stack, machinery interfaces, and historical operation records are real, while the agent actions are replayed through the calibrated digital twin to evaluate safety and operational correctness before field deployment\. ### 2\.4\.Agent families and LLMs We integrate nine controller architectures and evaluate them over the scenario suite: ReAct\([Yao et al\., 2023](https://arxiv.org/html/2609.00106#bib.bib2)\), Plan\-and\-Act\([Erdogan et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib6)\), Reflexion\([Shinn et al\., 2023](https://arxiv.org/html/2609.00106#bib.bib3)\), AutoGen\([Wu et al\., 2024](https://arxiv.org/html/2609.00106#bib.bib4)\), MMRL\([Tan et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib11)\), ReWOO\([Xu et al\., 2023](https://arxiv.org/html/2609.00106#bib.bib7)\), LATS\([Zhou et al\., 2024](https://arxiv.org/html/2609.00106#bib.bib12)\), CRITIC\([Gou et al\., 2024](https://arxiv.org/html/2609.00106#bib.bib5)\), and GoT\([Besta et al\., 2024](https://arxiv.org/html/2609.00106#bib.bib13)\)\. Each controller supports two modes:*direct*tool access and*agent\-to\-agent \(A2A\)*\([A2A Project, 2025](https://arxiv.org/html/2609.00106#bib.bib9)\), where weather, sensing, machinery, and operations queries are routed through specialist app agents\. Underlying farm state and tool capabilities are kept consistent across controllers and modes\. We run frontier API backends \(Qwen3\.6\-35B\-A3B\([Yang et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib10)\), DeepSeek\-V4\-Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.00106#bib.bib8)\)\)\. To reflect realistic deployment constraints, we also support edge execution through a vLLM API running on the NVIDIA Thor SDK with recent Qwen\-3 and Gemma\-4 model variants\. ### 2\.5\.Knowledge library Farm management relies on tacit timing rules and operating discipline that are usually implicit in human expertise\. FAIRY encodes this knowledge as aknowledge libraryof agronomic skills that controllers can retrieve at decision time, mirroring hierarchical skill/prompt libraries used for geospatial and remote\-sensing agents\([Bhattaram et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib30);[Badmus et al\., 2027](https://arxiv.org/html/2609.00106#bib.bib29);[Singh et al\., 2024](https://arxiv.org/html/2609.00106#bib.bib35)\)\. Each skill is a structured record with an identifier, a natural\-language title and description, a set of keywords, and a reusable operation*workflow*template\. At decision time, a retriever scores library entries against the current task and injects the top\-kk\(k=3k\{=\}3by default\) into the controller context\. We expose two retrieval mechanisms: a lightweight*lexical*retriever scoring token overlap between the query and each entry’s searchable text, and a*semantic*retriever over sentence embeddings of the entire full\-path workflow; both return ranked skills with scores\. This library underlies one of four interchangeable in\-context regimes the controller can run under:Zero Context\(operation primitives only\),LLM\-as\-an\-Expert\(an agriculture\-tuned LLM\([Wang et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib16)\)generates scenario context from the task, tools, and world state\),Skills Library\(retrieved skills injected into context\), andExpert Instruct\(human\-written scenario instructions\)\. We vary how the library is*organized*for retrieval \(i\.e\., a flat pool versus a tier\-grouped organization that separates atomic from composite operations\) and the*ranking mechanism*used to select entries \(i\.e\., manual selection, lexical text similarity, and path similarity\)\. Table 2\.Long\-horizon analysis on atomic tasks \(L1\), episode chains \(L2\), and full\-season scenarios \(L3\-test\) for Qwen3\.6\-35B\-A3B and DeepSeek\-V4\-Flash\. Yield Loss is the % drop from the human\-oracle biological yield; Succ\. is BFCL tool\-call success; KTC is temporally grounded correctness; Token cost is reported as the average tokens per task and per agent API call\. ### 2\.6\.Realistic full\-season scenarios We curate scenarios from representative on\-field procedures based on historical farm operations\. We consider three levels of increasing complexity:L1atomic tasks require one or two tool calls \(e\.g\., checking weather before a drone flight\);L2episodes chain observation, diagnosis, and intervention \(e\.g\., detecting an anomaly from a drone survey and applying targeted treatment\); andL3full\-season scenarios span planting through post\-harvest storage, with multiple interventions, delayed consequences, and accumulated effects on crop state and recovered yield\. For each scenario we follow the ARE annotation protocol\([Froger et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib18)\): three domain experts independently specify oracle solutions using the same operation primitives exposed to the agents; if the first two diverge we inspect the scenario to resolve ambiguity and revise, and the third expert confirms consistency\. Example scenarios by level:L1 \(atomic\)\.“The drone detected signs of aphids on ridges 15–25; verify the issue and apply an appropriate pesticide treatment\.”L2 \(episode\)\.“Remote sensing shows a localized low\-NDVI area during V4; diagnose drought vs\. pest vs\. nutrient deficiency and, if nutritional, treat via ridge\-level fertigation\.”L3 \(full season\)\.“Manage the full season from planting onward with weekly monitoring, addressing nutrient, water, pest, and disease issues from sensor/drone evidence, selecting a harvest window after maturity, and completing drying and storage\.” Replaying the oracle workflows in the engine produces human\-oracle farm\-state trajectories and target crop yields against which agent runs are compared\. The evaluation reported here uses a full\-season test set of 70 L3 scenarios, alongside the L1/L2 splits\. We also construct a focused 20 L3\-mini test set used for ablations, as well as a held\-out validation set of 20 L3 scenarios for confirming the generalization of different library schemes\. The oracle workflows are not assumed to be globally optimal; they represent expert\-validated operational references for reproducible comparison\. In future work, we will quantify inter\-expert disagreement and test the sensitivity of KTC and Yield Loss rankings to alternative oracle choices\. ### 2\.7\.Evaluation suite and metrics We evaluate each run along two axes:*how the agent acted*\(trace\-level correctness\) and*what it achieved*\(yield outcome\)\. Both compare an agent run against the human\-oracle workflow for the same scenario, represented as an ordered set of timestamped farm events, each carrying a tool, its arguments, target ridges, and a season\-day\. Trace\-level correctness\.As our primary correctness metric we useKTC\(Kendall\-style Temporal Correctness\), a temporally grounded full\-path score that extends the path\-correctness view of\([Michelakis et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib17)\)to farm\-event order\. After matching the agent’s executed operations to the oracle’s, KTC measures the order agreement of the matched operations as a normalized Kendall rank correlation,KTC=\(τ\+1\)/2∈\[0,1\]\\mathrm\{KTC\}=\(\\tau\+1\)/2\\in\[0,1\], rewarding required operations executed in causally valid order and penalizing reordering\. We complement KTC with two reference trace metrics computed from the same \(agent, oracle\) pair:BFCL, a tool\-call success rate measured as set overlap of executed\(tool,args\)\(\\text\{tool\},\\text\{args\}\)signatures \(order\- and time\-agnostic\);Path Correctness\(CORE\), a normalized edit distance between the agent and oracle operation sequences\([Michelakis et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib17)\)\. Agronomic outcome\.We reportYield Loss, the percentage shortfall of the agent’s biological yield relative to the human oracle \(1−agentbio/oraclebio1\-\\text\{agent\}\_\{\\text\{bio\}\}/\\text\{oracle\}\_\{\\text\{bio\}\}\)\. We pair correctness with outcome because they catch different failure modes: trace metrics flag plausible\-but\-wrong sequences a final\-state check would miss\([Froger et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib18)\), while yield loss captures operational misjudgments \(e\.g\., harvesting on day 87 instead of 89\) that a trace\-only metric may treat as minor deviations\. We additionally report token cost and runtime per scenario\. Figure 4\.FAIRY geospatial user interface\. The web\-based viewer displays the full\-path agent workflow together with map\-based visualizations\.Operational demonstration of the agentic farm\. ### 2\.8\.FAIRY system user interface We adapt the ARE user interface\([Froger et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib18)\)from a generic agent\-workflow environment into a geospatial inspection interface for agricultural workflows, as shown in Figure[4](https://arxiv.org/html/2609.00106#S2.F4)\. Notably, we extend the UI with domain\-specific visualization support for satellite and UAV outputs, including zoomable OpenStreetMap\-based map views, overlaid satellite imagery, crop\-classification rasters with adjustable opacity, and drone\-analysis tabs for UAV patrol paths and orthomosaic\-derived products\. We connect each application module to dedicated state and tool panels, so users can inspect the execution chain, intermediate outputs, and final geospatial products within the same workflow interface\. Table 3\.Agent performance onL3\-test\-minifull\-season scenarios with A2A enabled\.Table 4\.Agent performance onL3\-validationfull\-season scenarios across library schemes\. Flat structures correspond to directly retrieving against L3\-level contexts, and hierarchical structures correspond to multi\-level \(L1, L2, L3\) retrieval contexts\.Figure 5\.Agent\-family and backbone comparison onL3\-minifull\-season scenarios\. The left and right panels report Qwen3\.6\-35B\-A3B\-FP8 and DeepSeek\-V4\-Flash, respectively\. For each agent family, colored bars denote different context settings, and rows show full\-season yield, FAIRY KTC score, and token cost per step\.Figure summarizing agentic performance results across different controller and backbone implementations\. ## 3\.Results We organize results around the questions an operator would ask before deploying: what makes agents reliable \(context and skills\), how reliability scales with horizon, what multi\-agent orchestration costs, which metric we should trust, and what deployment costs\. Expert context dominates the controller spread\.Table[2](https://arxiv.org/html/2609.00106#S2.T2)reports the held\-outN=70N\{=\}70L3 evaluation under four in\-context regimes for both backbones\. Context is the dominant lever\. For Qwen, moving from Zero Context to Expert Instruct cuts Yield Loss from22\.3%22\.3\\%to4\.6%4\.6\\%and raises KTC from87\.9%87\.9\\%to92\.1%92\.1\\%; the Skills Library regime is close behind \(4\.9%4\.9\\%Yield Loss,93\.7%93\.7\\%KTC\) without any hand\-written per\-scenario policy\. DeepSeek shows the same ordering \(Zero→\\rightarrowExpert:12\.6%→2\.9%12\.6\\%\\rightarrow 2\.9\\%Yield Loss\)\. Notably,*LLM\-as\-an\-Expert*recovers only part of the gap \(Qwen15\.3%15\.3\\%Yield Loss versus4\.6%4\.6\\%for human instructions\), indicating that agriculture\-tuned LLM context captures high\-level decisions but not the full procedural discipline\. Table 5\.Local vLLM performance on atomic tasks \(L1\), episode chains \(L2\) and full\-season scenarios \(L3\-mini\) under expert\-instructed context\. Runtime \(seconds\) is measured on NVIDIA Thor\.Figure 6\.Left: Assessing full\-path agent metrics vs\. biological yield preservation\. Right: Evaluating the cost vs\. downstream objective performance \(yield\) tradeoff across scenario runs\.Figure summarizing the alignment of agentic evaluation metrics with downstream agronomic performance\.Knowledge\-library structure and retrieval\.Table[4](https://arxiv.org/html/2609.00106#S2.T4)ablates how the knowledge library is*organized*\(flat vs\. tier\-grouped\) and how entries are*ranked*\(manual, text similarity, path similarity\) on the L3\-mini set\. Tier\-grouped organization with similarity\-based retrieval gives the strongest correctness \(KTC up to97\.2%97\.2\\%for Qwen,96\.8%96\.8\\%for DeepSeek\) and the best task success, while flat/manual retrieval is weaker and more variable\. The effect on yield is smaller than the context effect, which is expected: retrieval quality refines*how*an already\-grounded agent acts, whereas context grounding determines whether it acts correctly at all\. Horizon analysis: short tasks are nearly solved, full seasons are not\.Table[2](https://arxiv.org/html/2609.00106#S2.T2)reports atomic \(L1\) and episodic \(L2\) tasks\. Under expert context, both horizons are essentially solved: L1 reaches≥\\geq99% KTC with near\-zero yield loss, and L2 reaches≥\\geq98% KTC with<<1% yield loss\. The contrast with the L3 results is the central horizon finding: short tasks mostly test whether the scaffold can select tools and satisfy local preconditions, whereas full\-season scenarios test whether decisions remain coherent after their effects propagate through soil state, growth, stress accumulation, treatment residuals, harvest timing, and storage\. Yield loss is where the horizon bites: it is small on L1/L2 but reaches double digits on L3 under zero context\. Figure 7\.FAIRY Operational Demonstration: drone patrol and Sentinel\-2 crop classification coordinated through the event\-driven runtime, rendered as a geospatial inspection result in the system UI\.Operational demonstration of the proposed FAIRY framework on our research farm\.Multi\-agent orchestration introduces coordination cost\.Table[3](https://arxiv.org/html/2609.00106#S2.T3)compares direct tool access against A2A routing, where the controller coordinates with weather, sensing, machinery, and operations specialists\. Averaged over contexts, A2A degrades both correctness and yield on both backbones: Yield Loss rises from7\.1%7\.1\\%to31\.2%31\.2\\%\(DeepSeek\) and from10\.7%10\.7\\%to29\.4%29\.4\\%\(Qwen\)\. Token cost per scenario drops because work is offloaded to specialists, but the main controller frequently loses track of what each specialist observed or executed\. The effect is controller\-dependent \(tree search is comparatively robust\), consistent with coordination\-overhead observations in other expert\-level multi\-agent studies\([Lee et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib31);[Badmus et al\., 2027](https://arxiv.org/html/2609.00106#bib.bib29)\)\. Early analysis suggests errors due to task decomposition issues, so we plan a detailed failure taxonomy as future work\. Which trace metric predicts yield?Figure[6](https://arxiv.org/html/2609.00106#acmlabel6)\(left\) and Table[6](https://arxiv.org/html/2609.00106#S3.T6)regress each metric against biological yield preservation across the full sweep\. Overall, we observe that KTC is the best\-calibrated predictor as it lies almost on the unit line\. Moreover, we note that BFCL’s highR2R^\{2\}performance comes from robustly identifying catastrophic “never harvested” runs \(zero recovered yield\), rather than grading quality among completing runs: it detects failure but may be poorly calibrated\. Overall, the results show that existing practices of order\-only and exact\-match metrics might not fully capture downstream tasks\. Overall, we consider the following for our practical deployment and future work: reportKTCfor trace\-level correctness andYield Lossfor outcome, and treat exact\-match success as a failure detector rather than a quality measure\. Table 6\.Alignment between agent\-evaluation metrics and the downstream agronomic objective, measured by yield\. Lower RMSE and slope closer to 1 indicate better calibration, while higherR2R^\{2\}indicates stronger explanatory fit\.Edge deployment profiling\.Because field deployment cannot assume frontier\-API availability, we assess performance and cost under edge deployment considerations\. Table[5](https://arxiv.org/html/2609.00106#S3.T5)shows that local vLLM execution can support L1 and L2 farm tasks with low yield loss across several mid\-size models, but full\-season L3 scenarios separate models more clearly: Qwen3\.6\-35B\-A3B\-FP8 and Qwen3\.6\-27B\-FP8 keep yield loss near 2%, while smaller or assistant\-tuned variants degrade sharply\. In practice, this suggests that edge deployment is feasible for farm\-agent evaluation, but full\-season autonomy still requires models with enough planning capacity\. Per\-controller robustness\.Figure[5](https://arxiv.org/html/2609.00106#S2.F5)summarizes the performance across the 9 different agentic back\-ends on the L3\-mini set across the four in\-context regimes\. Overall, we observe that in\-context operational grounding is the dominant factor across backbones: moving from zero context to the skills library or expert instruction reduces average yield loss from 19\.9% to 2\.9%/2\.1% for Qwen3\.6\-35B\-A3B and from 12\.8% to 2\.2%/3\.0% for DeepSeek\-V4\-Flash\. In practice, this means that the farm deploymentcannotrely on generic, state\-of\-the\-art agent scaffolds alone; agents need explicit agronomic skills or expert operational context to preserve yield under full\-season decision sequences\. ## 4\.FAIRY Operational Demonstration We demonstrate FAIRY on a real\-farm inspection episode at the HIT smart\-agriculture site\. As shown in Figure[7](https://arxiv.org/html/2609.00106#S3.F7), the demonstration focuses on the following L2 episode: a user issues a field\-monitoring request, and the agent must coordinate satellite and drone observations, crop\-identification tools, post\-processing, and the FAIRY UI to return an interpretable inspection result\. The integrated episode exercises the full observation workflow: the agent first triggers the drone API, retrieves and classifies satellite imagery while the drone is in flight; it then retrieves the drone result, generates the orthomosaic, plots the patrol path, runs field/crop analysis, integrates satellite and drone findings, and updates the UI display tool to render base\-map context, classification results, the patrol path, and image\-analysis products\. Overall, this demonstration provides us with a working deployment prototype and end\-to\-end path from user request, through event\-driven agent execution, to farm\-facing geospatial visualization\. ## 5\.Discussion From expert instructions to reusable skills\.Encoding tacit expert knowledge as scenario\-specific instructions is effective for evaluation but does not scale: each new task or seasonal edge case would need another written policy\. Organizing this knowledge into retrievable, composableskill librariesis the scalable alternative, consistent with the use of structured knowledge pools in power\-grid\([Badmus et al\., 2027](https://arxiv.org/html/2609.00106#bib.bib29)\)and Earth\-observation\([Bhattaram et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib30);[Shabbir et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib32);[Chen et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib33);[Feng et al\., 2026](https://arxiv.org/html/2609.00106#bib.bib34)\)agents; our library ablation is a first step in this direction\. Full\-path evaluation paired with farm objectives\.A planting, irrigation, or spraying decision can look locally plausible while still causing downstream yield loss, so the downstream objective must be evaluated alongside the action trace\. KTC and Yield Loss diagnose different failure modes: a high\-KTC, high\-yield\-loss run shows that small operational differences can have physical consequences, while a lower\-KTC, low\-yield\-loss run shows deviation from the reference trace that nonetheless preserved the outcome\. Our metric\-calibration study sharpens this: among trace metrics, temporally grounded correctness tracks yield, whereas flat order metrics and success measures do not\. We note that the current analysis aggregates results at the scenario level\. A natural next step is ridge\-level spatial analysis, including whether errors cluster around sparsely instrumented ridges, low\-observability zones, or operations that depend on UAV and satellite coverage\. Agent orchestration should preserve operational state\.Planning, memory, retrieval, verification, and specialist decomposition can improve the execution trace, but they do not remove the need to maintain agronomic and operational assumptions across time\. Splitting the farm interface into specialists matches the structure of real farm systems, yet introduces state\-sharing requirements; if the controller loses track of what each specialist observed or executed, orchestration becomes error\-prone\. The design requirement is not only better decomposition but preserving shared farm state, timing constraints, and task context across agent boundaries\. LLM\-as\-an\-Expert context\.Even when a domain LLM\([Wang et al\., 2025](https://arxiv.org/html/2609.00106#bib.bib16)\)is given the scenario, tools, world state, and a matched response template, it does not fully match human experts on procedural ordering\. This does not make the domain LLM unhelpful, as it recovers high\-level crop\-management choices, but it motivates structured knowledge libraries that complement both LLM\-distilled and human\-written guidance\. Quality of human\-annotated solutions\.Because every trace metric compares against a human\-oracle workflow, the oracle must be a fixed, shared artifact; regenerating it across software versions or machines introduces drift that silently changes every metric\. We recommend version\-controlling and serializing the oracle workflows alongside each scenario\. Moreover, a timing\-aware metric can only be validated where mistiming actually costs yield, making such information available in the oracle traces particularly important\. The next stage of deployment is to run an agent\-managed plot alongside the human\-operated workflow and compare both operational traces and harvested outcomes under real seasonal conditions\. ## 6\.Conclusion We presented FAIRY, a deployable smart\-agriculture agentic engine, and used it for a full\-season spatiotemporal evaluation of contemporary agent practices on an operating soybean research farm\. Expert context and retrievable agronomic skills are the dominant levers for long\-horizon reliability; multi\-agent orchestration and deployment cost are practical constraints; and among trace\-level metrics, temporally grounded correctness is the one that tracks real yield\. These results inform our next stage: operating a dedicated plot under agent management alongside the human\-operated farm\. We hope FAIRY provides the community a deployable, spatiotemporally grounded reference point for evaluating agents in other physical\-process settings\. Our entire working prototype can be found here:[https://github\.com/Fengrui\-Lab/FAIRY](https://github.com/Fengrui-Lab/FAIRY) ###### Acknowledgements\. This work was supported by the National Natural Science Foundation of China \(grant No\. 62350710797\)\. We gratefully acknowledge the support of the National Key Research and Development Program \[2025YFE0209200\] and the Key Research and Development Program of Heilongjiang Province, China \[2024ZX01A07, JD2023GJ01\]\. This work was also supported by the NSFC grant \(No\. 42471362\)\. DS gratefully acknowledges the support of the NSFC Excellent Young Scientists Fund Program \(Overseas\)\. ## References - A2A Project \(2025\)A2A ProjectAgent2Agent protocol \(A2A\)\.Note:[https://github\.com/a2aproject/A2A](https://github.com/a2aproject/A2A)GitHub repositoryCited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Akyuzet al\.\(2017\)F\. A\. Akyuz, H\. Kandel, and D\. MorlockDeveloping a growing degree day model for north dakota and northern minnesota soybean\.Agricultural and Forest Meteorology239,pp\. 134–140\.Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p2.1)\. - Badmus and Pandey \(2026\)E\. O\. Badmus and A\. PandeyPowerDAG: reliable agentic ai system for automating distribution grid analysis\.External Links:2603\.17418,[Link](https://arxiv.org/abs/2603.17418)Cited by:[Table 4](https://arxiv.org/html/2609.00106#S2.T4.4.1.9.1)\. - Badmuset al\.\(2027\)E\. O\. Badmus, P\. Sang, D\. Stamoulis, and A\. PandeyPowerChain: a verifiable agentic ai system for automating distribution grid analyses\.Electric Power Systems Research262,pp\. 113555\.External Links:[Document](https://dx.doi.org/10.1016/j.epsr.2026.113555),[Link](https://doi.org/10.1016/j.epsr.2026.113555)Cited by:[§2\.5](https://arxiv.org/html/2609.00106#S2.SS5.p1.1),[Table 4](https://arxiv.org/html/2609.00106#S2.T4.4.1.8.1),[§3](https://arxiv.org/html/2609.00106#S3.p5.1),[§5](https://arxiv.org/html/2609.00106#S5.p1.1)\. - Bestaet al\.\(2024\)M\. Besta, N\. Blach, A\. Kubicek, R\. Gerstenberger, M\. Podstawski, L\. Gianinazzi, J\. Gajda, T\. Lehmann, H\. Niewiadomski, P\. Nyczyk, and T\. HoeflerGraph of thoughts: solving elaborate problems with large language models\.InProceedings of the Thirty\-Eighth AAAI Conference on Artificial Intelligence and Thirty\-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’24/IAAI’24/EAAI’24\.External Links:ISBN 978\-1\-57735\-887\-9,[Link](https://doi.org/10.1609/aaai.v38i16.29720),[Document](https://dx.doi.org/10.1609/aaai.v38i16.29720)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Bhattaramet al\.\(2025\)A\. Bhattaram, J\. Chung, S\. Chung, R\. Gupta, J\. Ramamoorthy, K\. Gullapalli, D\. Marculescu, and D\. StamoulisGeoFlow: agentic workflow automation for geospatial tasks\.InProceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems,SIGSPATIAL ’25,New York, NY, USA,pp\. 1150–1153\.External Links:ISBN 9798400720864,[Link](https://doi.org/10.1145/3748636.3763217),[Document](https://dx.doi.org/10.1145/3748636.3763217)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p1.1),[§2\.5](https://arxiv.org/html/2609.00106#S2.SS5.p1.1),[Table 4](https://arxiv.org/html/2609.00106#S2.T4.4.1.7.1),[§5](https://arxiv.org/html/2609.00106#S5.p1.1)\. - Chenet al\.\(2026\)Z\. Chen, H\. Wang, J\. Yao, J\. Zhang, P\. Ghamisi, J\. Zhou, P\. M\. Atkinson, and B\. ZhangCangLing\-knowflow: a unified knowledge\-and\-flow\-fused agent for comprehensive remote sensing applications\.External Links:2512\.15231,[Link](https://arxiv.org/abs/2512.15231)Cited by:[§5](https://arxiv.org/html/2609.00106#S5.p1.1)\. - DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-v4: towards highly efficient million\-token context intelligence\.Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Erdoganet al\.\(2025\)L\. E\. Erdogan, N\. Lee, S\. Kim, S\. Moon, H\. Furuta, G\. Anumanchipalli, K\. Keutzer, and A\. GholamiPLAN\-and\-act: improving planning of agents for long\-horizon tasks\.InProceedings of the 42nd International Conference on Machine Learning,ICML’25\.Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Fenget al\.\(2026\)P\. Feng, Z\. Lv, J\. Ye, X\. Wang, X\. Huo, J\. Yu, W\. Xu, W\. Zhang, L\. Bai, C\. He, and W\. LiEarth\-agent: unlocking the full landscape of earth observation with agents\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/5b4a459db23e6db9be2a128380953d96-Abstract-Conference.html)Cited by:[§5](https://arxiv.org/html/2609.00106#S5.p1.1)\. - Frogeret al\.\(2025\)R\. Froger, P\. Andrews, M\. Bettini, A\. Budhiraja, R\. S\. Cabral, V\. Do, E\. Garreau, J\. Gaya, H\. Laurençon, M\. Lecanu, K\. Malkan, D\. Mekala, P\. Ménard, G\. M\. Bertran, U\. Piterbarg, M\. Plekhanov, M\. Rita, A\. Rusakov, V\. Vorotilov, M\. Wang, I\. Yu, A\. Benhalloum, G\. Mialon, and T\. ScialomARE: scaling up agent environments and evaluations\.External Links:2509\.17158,[Link](https://arxiv.org/abs/2509.17158)Cited by:[§2\.3](https://arxiv.org/html/2609.00106#S2.SS3.p1.1),[§2\.6](https://arxiv.org/html/2609.00106#S2.SS6.p1.1),[§2\.7](https://arxiv.org/html/2609.00106#S2.SS7.p3.1),[§2\.8](https://arxiv.org/html/2609.00106#S2.SS8.p1.1),[Table 6](https://arxiv.org/html/2609.00106#S3.T6.2.3.1.1)\. - Gouet al\.\(2024\)Z\. Gou, Z\. Shao, Y\. Gong, yelong shen, Y\. Yang, N\. Duan, and W\. ChenCRITIC: large language models can self\-correct with tool\-interactive critiquing\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Sx038qxjek)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Humburg \(2019\)D\. HumburgChapter 38: determining harvest losses in soybeans\.InSouth Dakota State University iGrow Soybean Best Management Practices,Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p3.1)\. - Leeet al\.\(2025\)C\. Lee, V\. Paramanayakam, A\. Karatzas, Y\. Jian, M\. Fore, H\. Liao, F\. Yu, R\. Li, I\. Anagnostopoulos, and D\. StamoulisMulti\-agent geospatial copilots for remote sensing workflows\.InIGARSS 2025 \- 2025 IEEE International Geoscience and Remote Sensing Symposium,Vol\.,pp\. 1084–1089\.External Links:[Document](https://dx.doi.org/10.1109/IGARSS55030.2025.11243915)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p1.1),[§3](https://arxiv.org/html/2609.00106#S3.p5.1)\. - Liet al\.\(2025\)K\. Li, J\. Wang, Z\. Wang, H\. Qiao, W\. Zhang, D\. Meng, and X\. CaoDesigning domain\-specific agents via hierarchical task abstraction mechanism\.External Links:2511\.17198,[Link](https://arxiv.org/abs/2511.17198)Cited by:[Table 4](https://arxiv.org/html/2609.00106#S2.T4.4.1.10.1)\. - Michelakiset al\.\(2025\)P\. Michelakis, Y\. Hadjiyianni, and D\. StamoulisCORE: full\-path evaluation of llm agents beyond final state\.External Links:2509\.20998,[Link](https://arxiv.org/abs/2509.20998)Cited by:[§2\.7](https://arxiv.org/html/2609.00106#S2.SS7.p2.1)\. - Monsi and Saeki \(2005\)M\. Monsi and T\. SaekiOn the factor light in plant communities and its importance for matter production\.Annals of botany95\(3\),pp\. 549–567\.Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p2.1)\. - Monteith \(1977\)J\. L\. MonteithClimate and the efficiency of crop production in britain\.Philosophical transactions of the royal society of London\. B, Biological Sciences281\(980\),pp\. 277–294\.Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p2.1)\. - Patilet al\.\(2025\)S\. G\. Patil, H\. Mao, F\. Yan, C\. C\. Ji, V\. Suresh, I\. Stoica, and J\. E\. GonzalezThe berkeley function calling leaderboard \(bfcl\): from tool use to agentic evaluation of large language models\.InProceedings of the 42nd International Conference on Machine Learning,ICML’25\.Cited by:[Table 6](https://arxiv.org/html/2609.00106#S3.T6.2.2.1.1)\. - Quet al\.\(2026\)A\. Qu, P\. Michelakis, Y\. Hadjiyianni, F\. Li, J\. Jiang, D\. Stamoulis, and J\. LiuFull\-season agent evaluation in soybean farm operations under real\-world agricultural process dynamics\.InSecond Workshop on Agents in the Wild: Safety, Security, and Beyond,External Links:[Link](https://openreview.net/forum?id=m49CywnvoV)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\. - Seo and Lee \(2026\)S\. Seo and K\. LeeDensity\-driven multidrone coordination for efficient farm coverage and management in smart agriculture\.IEEE Transactions on Control Systems Technology34\(2\),pp\. 711–724\.External Links:ISSN 2374\-0159,[Link](http://dx.doi.org/10.1109/TCST.2025.3631091),[Document](https://dx.doi.org/10.1109/tcst.2025.3631091)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\. - Shabbiret al\.\(2026\)A\. Shabbir, M\. A\. Munir, A\. Dudhane, M\. U\. Sheikh, M\. H\. Khan, P\. Fraccaro, J\. B\. Moreno, F\. S\. Khan, and S\. KhanThinkGeo: evaluating tool\-augmented agents for remote sensing tasks\.External Links:2505\.23752,[Link](https://arxiv.org/abs/2505.23752)Cited by:[§5](https://arxiv.org/html/2609.00106#S5.p1.1)\. - Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Siler and Singh \(2022\)T\. B\. Siler and M\. P\. SinghOptimal soybean maturity group selection is influenced by planting date in northern production systems\.Crop Science62\(6\),pp\. 2462–2475\.Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p1.1)\. - Singhet al\.\(2024\)S\. Singh, M\. Fore, and D\. StamoulisGeoLLM\-engine: a realistic environment for building geospatial copilots\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\) Workshops,pp\. 585–594\.Cited by:[§2\.5](https://arxiv.org/html/2609.00106#S2.SS5.p1.1),[Table 4](https://arxiv.org/html/2609.00106#S2.T4.4.1.6.1)\. - Stedutoet al\.\(2009\)P\. Steduto, T\. C\. Hsiao, D\. Raes, and E\. FereresAquaCrop—the fao crop model to simulate yield response to water: i\. concepts and underlying principles\.Agronomy journal101\(3\),pp\. 426–437\.Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p2.1)\. - Tanet al\.\(2026\)R\. Tan, B\. Peng, Z\. Yang, H\. Cheng, O\. Mees, T\. Zhao, A\. Tupini, I\. Meijier, Q\. Wu, Y\. Yang, L\. Liden, Y\. Gu, S\. Zhang, X\. Liu, L\. Wang, M\. Pollefeys, Y\. J\. Lee, and J\. GaoMultimodal reinforcement learning with adaptive verifier for ai agents\.External Links:2512\.03438,[Link](https://arxiv.org/abs/2512.03438)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Wanget al\.\(2025\)H\. Wang, Y\. Guan, F\. Meng, C\. Zhao, L\. Yan, Y\. Yang, and J\. JiangAgri\-CM3\{\}^\{3\}: a Chinese massive multi\-modal, multi\-level benchmark for agricultural understanding and reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 11729–11754\.External Links:[Link](https://aclanthology.org/2025.acl-long.576/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.576),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.2](https://arxiv.org/html/2609.00106#S2.SS2.p2.1),[§2\.5](https://arxiv.org/html/2609.00106#S2.SS5.p2.1),[§5](https://arxiv.org/html/2609.00106#S5.p4.1)\. - Wuet al\.\(2024\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, R\. W\. White, D\. Burger, and C\. WangAutoGen: enabling next\-gen LLM applications via multi\-agent conversations\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=BAakY1hNKS)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Xuet al\.\(2023\)B\. Xu, Z\. Peng, B\. Lei, S\. Mukherjee, Y\. Liu, and D\. XuReWOO: decoupling reasoning from observations for efficient augmented language models\.External Links:2305\.18323,[Link](https://arxiv.org/abs/2305.18323)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Xuet al\.\(2025\)Z\. Xu, J\. Xu, M\. Zhang, P\. Wang, C\. Deng, and C\. LiuMultimodal agricultural agent architecture \(ma3\): a new paradigm for intelligent agricultural decision\-making\.External Links:2504\.04789,[Link](https://arxiv.org/abs/2504.04789)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\. - Yanet al\.\(2026\)L\. Yan, H\. Wang, C\. Tang, H\. Liu, T\. Sun, L\. Liu, Y\. Guan, and J\. JiangAgrieval: a comprehensive chinese agricultural benchmark for large language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 34205–34213\.Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\. - Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2210.03629)Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Zhanget al\.\(2026\)Z\. Zhang, J\. Zhang, H\. Liu, Q\. Lv, J\. Yang, K\. Cai, and K\. WangAgriWorld:a world tools protocol framework for verifiable agricultural reasoning with code\-executing llm agents\.External Links:2602\.15325,[Link](https://arxiv.org/abs/2602.15325)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\. - Zhouet al\.\(2024\)A\. Zhou, K\. Yan, M\. Shlapentokh\-Rothman, H\. Wang, and Y\. WangLanguage agent tree search unifies reasoning, acting, and planning in language models\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§2\.4](https://arxiv.org/html/2609.00106#S2.SS4.p1.1)\. - Zuzuárreguiet al\.\(2025\)M\. A\. Zuzuárregui, M\. M\. Toslak, and S\. CarpinOne for all: llm\-based heterogeneous mission planning in precision agriculture\.External Links:2506\.10106,[Link](https://arxiv.org/abs/2506.10106)Cited by:[§1](https://arxiv.org/html/2609.00106#S1.p2.1)\.
Similar Articles
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
AgentFAIR is a multi-agent framework that uses LLM evaluators and a critic to assess FAIR compliance of geospatial datasets, achieving sub-principle agreement of 89% and Fleiss' κ=0.71 in expert studies, at a cost of $0.054 per dataset.
Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation
This paper introduces Agri-SAGE, a closed-loop framework that integrates multi-agent LLM reasoning with biophysical simulation (APSIM) to generate and validate context-aware agricultural advisories. The framework outperforms static baselines in retrospective analysis, with Tree-of-Thoughts achieving peak yields and Reflexion reducing computational cost via episodic memory.
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
This paper systematically evaluates seven frontier AI agents on long-horizon tasks, revealing they function more as engineering optimizers than autonomous researchers, with recommendations for improving training and experience management.
Skill-based Agentic Evaluation for Real-time Data Science Tasks
This paper introduces a ground-truth-as-code framework for evaluating data-science agents on continuously updated data, achieving improved agreement with human evaluators and reduced token consumption.
An Empirical Study of Automating Agent Evaluation
This paper introduces EvalAgent, a system that automates the evaluation of AI agents by encoding domain-specific expertise, addressing the limitations of standard coding assistants in this task. It also presents AgentEvalBench, a benchmark for testing evaluation pipelines, and demonstrates significant improvements in evaluation reliability.