MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity

arXiv cs.AI Papers

Summary

MineTRACE is a web-based evidence-grounded interactive reasoning system that integrates geochemical, geophysical, and geological data to provide transparent mineral prospectivity scores and natural language interaction, supporting efficient and verifiable mineral exploration.

arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language. A transparent expert tree, informed by geological knowledge and known deposits, combines multi-source evidence into interpretable prospectivity scores. For a new location, the conversational assistant retrieves the score and supporting evidence from the analysis pipeline and presents them in natural language. The scorer achieves spatial AUC values of up to 0.917 across different test scenarios, while end-to-end evaluation assesses query accuracy and response grounding. MineTRACE makes public geoscience data easier to access, interpret, and verify, supporting more efficient and transparent mineral exploration.
Original Article
View Cached Full Text

Cached at: 09/03/26, 06:00 AM

# MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
Source: [https://arxiv.org/html/2609.02060](https://arxiv.org/html/2609.02060)
Jinwen Liu11footnotemark:1Affiliation:The University of Western AustraliaDaniel Su11footnotemark:1Affiliation:The University of Western AustraliaYisu ChenAffiliation:Wilfrid Laurier Universityyiran\.zhang@research\.uwa\.edu\.au, \{wei\.liu,yihao\.ding\}@uwa\.edu\.auQiang SunAffiliation:The University of Western AustraliaChris GonzalezAffiliation:The University of Western AustraliaEun\-Jung HoldenAffiliation:The University of Western AustraliaMarco FiorentiniAffiliation:The University of Western AustraliaWei LiuAffiliation:The University of Western AustraliaYihao DingAffiliation:The University of Western Australia

###### Abstract

Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps\. We presentMineTRACE, a web\-based system for evidence\-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn\. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language\. A transparent expert tree, informed by geological knowledge and known deposits, combines multi\-source evidence into interpretable prospectivity scores\. For a new location, the conversational assistant retrieves the score and supporting evidence from the analysis pipeline and presents them in natural language\. The scorer achieves spatial AUC values of up to 0\.917 across different test scenarios, while end\-to\-end evaluation assesses query accuracy and response grounding\.MineTRACEmakes public geoscience data easier to access, interpret, and verify, supporting more efficient and transparent mineral exploration\. Our video is available via[https://geo\.nlp\-tlp\.org/video](https://geo.nlp-tlp.org/video), live demo is available via[https://geo\.nlp\-tlp\.org](https://geo.nlp-tlp.org/)\.

## 1Introduction

![Refer to caption](https://arxiv.org/html/2609.02060v1/images/overview.png)Figure 1:Overview ofMineTRACE\. Public geochemical, geophysical, and geological evidence is integrated by a transparent expert\-tree scorer to assess mineral prospectivity at unknown locations and regions\. The scorer returns a unified evidence record that powers the interactive map workspace and grounded natural\-language assistant, helping users inspect prospective areas and make evidence\-informed exploration decisions\.Mineral exploration is essential for expanding the resource base needed to secure critical mineral supplies for modern technologies and economic development\([Müller et al\., 2025](https://arxiv.org/html/2609.02060#bib.bib2)\)\. Identifying prospective targets, however, requires geologists to integrate heterogeneous geochemical, geophysical, and geological evidence across large and unevenly sampled regions\. Mineral prospectivity mapping helps organise this evidence into spatial scores or heatmaps, but practical exploration decisions require more than knowing where a score is high\. Geologists must also understand which evidence supports the assessment, whether different evidence sources agree, and whether the available data are sufficient to justify further investigation\.

Existing approaches support only parts of this process\. Data\-driven models can produce accurate prospectivity maps[Ding et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib1), but often hide the evidence behind a final score\. Knowledge\-driven methods are more transparent[Dong and Zhang \(2024\)](https://arxiv.org/html/2609.02060#bib.bib7), yet are commonly presented as static maps or offline analyses\. Commercial platforms allow users to view multiple data layers, while conversational assistants can simplify interaction[Liu et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib16), but these components are rarely connected\. As a result, users still lack a unified workflow that links each prospectivity assessment to its supporting evidence and allows the result to be inspected interactively within the same interface\.

We presentMineTRACE, a user\-friendly, evidence\-centric system designed to bridge the gap between mineral prospectivity predictions and practical exploration decisions\. Our aim is to move beyond static scores and heatmaps by making prospectivity assessments interactive, interpretable, and traceable\. To support practical exploration workflows, the system allows users to directly query new locations or regions, inspect the multi\-source evidence behind each assessment, compare prospective targets, and obtain evidence\-grounded explanations through natural language\.

MineTRACEsupports eight commodities and is built on large\-scale public geochemical, geophysical, and geological data from Western Australia\. At its core, a transparent expert tree combines nearby observations into a structured evidence record containing the prospectivity score, contributing signals, expert contributions, evidence sources, and local sample coverage\. The same record is shared across the map, evidence panels, spatial queries, and conversational interface, ensuring that all system outputs remain consistent with the underlying prospectivity analysis\. The contributions of this paper are summarized as follows:

- •A user\-friendly and explainable system for mineral prospectivity analysis\.MineTRACEintegrates exploration, point and region queries, target comparison, evidence inspection, and natural\-language interaction, allowing users to obtain prospectivity assessments and understand the evidence behind them within a workflow\.
- •A traceable multi\-source evidence representation\.MineTRACEuses a transparent expert tree to integrate geochemical, geophysical, and geological observations into a structured evidence record, explicitly linking each prospectivity score to its contributing signals, expert contributions, source data, and local coverage\.
- •A large\-scale evaluation of prospectivity and interaction quality\.We evaluateMineTRACEon approximately 9\.35 million assays and 3,420 known mineral sites across eight commodities, while also assessing prospectivity ranking, spatial generalisation, multi\-source evidence fusion, query accuracy, and response grounding\.

## 2Related Work

Data\-driven approaches model mineral prospectivity as an anomaly detection problem, using methods such as autoencoder–GMMs[Wang and Chen \(2025\)](https://arxiv.org/html/2609.02060#bib.bib3)and transformer\-based models[Yu et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib4)to capture complex geochemical patterns and achieve strong predictive performance; however, they produce only a single score per location and lack interpretability of the underlying evidence contributing to that score\. Knowledge\-driven approaches instead aim to improve interpretability through structured, tree\-based models[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.02060#bib.bib5);[Rai et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib6);[Dong and Zhang \(2024\)](https://arxiv.org/html/2609.02060#bib.bib7)that provide feature importance and decision rules for their predictions, but these methods are typically limited to a single data modality, do not generalise well across regions, lack explicit reasoning over multi\-source evidence, and are not designed as interactive systems\. Beyond geoscience, recent interactive systems have demonstrated how domain\-specific information can be extracted and inspected in a more transparent manner, including key information extraction from domain\-specific documents[Ding et al\. \(2025\)](https://arxiv.org/html/2609.02060#bib.bib14), context\-aware spatiotemporal information extraction[Zhang et al\. \(2026a\)](https://arxiv.org/html/2609.02060#bib.bib15), and visual inspection of multi\-turn LLM reasoning[Zhang et al\. \(2026b\)](https://arxiv.org/html/2609.02060#bib.bib20), but none of these systems specifically targets mineral prospectivity assessment\. Commercial platforms such as MINML111[https://minml\.co\.uk/](https://minml.co.uk/)attempt to integrate multi\-source geoscience data and provide interactive interfaces with prospectivity rankings, claiming evidence supported outputs, without publicly available quantitative evaluation\. Our MINETRACE provides a unified, evidence\-grounded system that links multi\-source data, interpretable scoring, and interactive natural\-language exploration, with all outputs traceable to underlying evidence and quantitatively evaluated\.

## 3System Design

This section presents the design ofMineTRACE\. We first introduce the overall system architecture, followed by the data and evidence layer, the interpretable prospectivity engine, and the interactive interfaces and their integration with the system\.

### 3\.1System Overview

Figure 2:Architecture design forMineTRACE\. Django is a Python web framework; PostGIS is a database for geospatial data; ETL means extract–transform–load for data preparation; MCP \(Model Context Protocol\) lets external agents call software tools\.MineTRACEis organized as a web system around a single scoring path \(Figure[2](https://arxiv.org/html/2609.02060#S3.F2)\)\. Users interact with a browser workspace that combines the map, evidence panel, and conversational assistant\. User requests are routed through a platform backend to a separate scoring service, which applies the expert\-tree model and returns a prospectivity score together with its evidence record\. The scoring service reads processed geoscience data from a PostGIS store and loads fitted expert weights produced by offline training and preprocessing jobs\. This design keeps the interface simple while ensuring that every map result and assistant response is generated from the same scoring path\.

### 3\.2Data and Evidence Layer

#### Data Sources\.

MineTRACEuses three open data products published by the Geological Survey of Western Australia \(GSWA\)\.*CM02 Near Surface Geochemistry \(Geochem\.\)*provides 9\.35 million assays from five sampling media: sediment, rock chip, drill\-hole maximum grade, shallow drilling, and surface soil\.*CM01 Mineralization Sites \(Sites\)*provides 3,420 known mine and deposit sites across eight commodities, which we use to construct supervision labels\. The*CM08 Critical Minerals Basemap \(Basemap\)*provides seven geophysical rasters and five geological vector layers, including magnetics, gravity, radiometrics, geochronology, faults, geological units, and Cenozoic cover\. All products are registered to GDA2020 and cover about2\.52\.5million km2across Western Australia\.222Available through the DMIRS Data and Software Centre:[https://dasc\.dmirs\.wa\.gov\.au](https://dasc.dmirs.wa.gov.au/)\.

#### Data to Evidence Processing\.

Raw data are first harmonised into a unified statewide evidence store\. Geochemical assays are cleaned, mapped to a fixed element schema, and separated by sampling medium; geophysical rasters are aligned to a common spatial reference; and geological, structural, and known\-deposit layers are spatially indexed and converted into queryable attributes and distance features\. For each query location or region, the system retrieves nearby observations from this store and assembles a structured*evidence record*containing multi\-scale geochemical statistics, geophysical values, geological and structural context, proximity to known deposits, and local sample coverage\. Missing evidence is handled explicitly: when fewer than three assay samples are available within 10 km, the geochemical component abstains rather than extrapolating from insufficient observations\. This record provides a shared evidence representation for downstream scoring, visual inspection, and direct conversational interaction\. Full preprocessing details are provided in Appendix[A](https://arxiv.org/html/2609.02060#A1)\.

![Refer to caption](https://arxiv.org/html/2609.02060v1/mapoverview.png)\(a\)Interactive interface and post\-training gold projects\.
\(b\)Post\-training gold project details\.
Figure 3:Interactive interface and post\-training gold\-project case study inMineTRACE\.\(a\)The interface illustrates the interaction modes in Section[3\.4](https://arxiv.org/html/2609.02060#S3.SS4): commodity switching, evidence\-layer toggling, map\-click prospectivity queries, evidence inspection, and natural\-language interaction with the assistant\. The six numbered points mark gold projects that became public after the model’s training cutoff\.\(b\)Project summaries for these six points\. As discussed in Section[5](https://arxiv.org/html/2609.02060#S5), they are used as a qualitative case study rather than a formal benchmark: they were not training positives, and their scores and statewide percentiles show howMineTRACEassesses recent gold targets\.

### 3\.3Interpretable Prospectivity Engine

The interpretable prospectivity engine systematically converts evidence around a query location into a prospectivity score and a traceable, auditable explanation\. It has three stages: evidence representation, expert scoring, and score aggregation\.

#### Evidence Representation\.

For each query point, the spatial index retrieves nearby samples within 5, 10, and 50,km\. For a fixed panel of 14 target and pathfinder elements, it computes robust statistics in log space, including local enrichment, high\-percentile anomalies, local\-to\-regional contrast, element ratios, spatial coherence, and sample coverage\. Geological and geophysical context is added through raster values, mapped attributes, and distances to relevant structures\. Locations with insufficient nearby observations are marked as low coverage rather than treated as confident predictions\. The detailed feature catalogue at Appendix[B](https://arxiv.org/html/2609.02060#A2)\.

#### Expert scoring\.

The resulting features are evaluated by named experts, each representing a fixed exploration heuristic and producing a score in\[0,1\]\[0,1\]\. This design preserves the meaning of each evidence source and allows the final score to be attributed to clear geological arguments\.Geochemical expertsevaluate target enrichment, pathfinder enrichment, and correlated element anomalies\.Geophysical expertsevaluate magnetic, gravity, radiometric, and geochronological responses\.Geological expertsevaluate fault proximity, geophysical worms, favourable host lithologies, and Cenozoic\-cover context\. Commodity\-specific pathfinder lists and expert definitions are given in Appendix[C](https://arxiv.org/html/2609.02060#A3)\.

#### Transparent Score aggregation\.

The final prospectivity score is a weighted mean of the active expert scores\. Expert weights are fitted offline according to how well each expert separates known deposits from background locations\. Experts without the required evidence abstain and are excluded from the aggregation, while a separate coverage value indicates how much local evidence supports the score\. No end\-to\-end neural model is trained\. Offline fitting estimates expert and source weights, background statistics for z\-scoring, and the favourable direction of each feature\. Pathfinder definitions remain fixed from domain knowledge\. Each prediction is stored as an evidence record containing the contributing features, z\-scores, expert scores, expert weights, active sources, coverage information, and final commodity score\. The fitted parameters are frozen at deployment, and the deployed scorer matches the reference implementation to within10−610^\{\-6\}on a fixed golden set\. This ensures that offline evaluation and interactive queries use the same audited scoring process\.

### 3\.4Interactive Interfaces

The interface exposes the same evidence record through three complementary interaction modes\.

#### Direct spatial querying\.

A user selects a commodity and scores any location by clicking a point or drawing a region\. The system returns a prospectivity value, a tier, local sample coverage, nearest known\-deposit context, and an evidence panel\. The panel renders the model’s own decomposition: ranked feature signals, z\-scores, source medium, contributing experts, expert weights, and active evidence layers\. This lets the user inspect why the location scores as it does\.

#### Contextual map exploration\.

The workspace serves a precomputed prospectivity surface for each commodity as a toggleable heatmap\. Known deposits, magnetic and gravity layers, faults, worms, and other context layers can be displayed alongside the score surface\. Away from local data, the surface is marked as interpolated rather than treated as direct evidence\. This helps users read a target in geological context and compare a geochemical anomaly with independent evidence\.

#### Conversational tool use\.

A user can also ask in natural language\. The assistant scores and compares locations, ranks regions for a metal, switches the active commodity, moves the map, and toggles evidence layers by calling system tools\. Because it answers from the same evidence records shown in the evidence panels, the assistant is a queryable interface to the interpretable model rather than an independent narrator\. In particular, since LLM assistants can otherwise drift toward plausible but unsupported claims[Kashyap et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib18), it is instructed to report only tool\-computed scores, distances, tiers, and rankings\.

## 4Evaluation

### 4\.1Scoring Quality

#### Evaluation Setup\.

We evaluate whether the scorer ranks known mineralisation above background across eight commodities\. Positive samples are sites classified as*Mine*or*Deposit*\. We consider four negative\-sampling strategies of increasing difficulty: random locations across Western Australia, far\-random locations more than 50 km from any known site, validated sites associated with other commodities \(*NonMine*\), and a south\-train/north\-test spatial holdout \(*Spatial*\)\. For each setting, we hold out 30 positive sites, sample 200 negatives, fit the expert weights on the remaining data, and report ROC\-AUC\. Because random and far\-random negatives are often located in sparsely sampled areas, we treat the NonMine and Spatial settings as the more realistic tests\.

Table 1:Prospectivity ranking across eight commodities, measured by ROC\-AUC\. Random and Far\-Random are contextual tests, while NonMine and Spatial provide more conservative evaluations\. Highlighted cells indicate the best commodity\-level result in each column\.Figure 4:Evidence\-family ablation under ROC\-AUC;Δ\\Deltais the fused model’s gain over the best single family\.
#### Results and insights\.

Table[1](https://arxiv.org/html/2609.02060#S4.T1)shows that performance varies across commodities and test settings\. Ni performs best on the realistic tests, reaching 0\.946 NonMine AUC and 0\.917 Spatial AUC, while Co also remains strong at 0\.929 and 0\.805\. Au and Sn perform well against validated non\-target sites, whereas Cu and Mn degrade substantially under spatial holdout, indicating weaker regional generalisation\. Figure[4](https://arxiv.org/html/2609.02060#S4.F4)further shows that geochemical, geophysical, and geological evidence are each informative, but their importance differs by commodity\. Although this ablation uses the easier far\-random setting, combining all three evidence families consistently gives the highest AUC for every commodity\. These results collectively support both multi\-source fusion and the interface design that exposes evidence contributions and local coverage alongside each score\.

### 4\.2End\-to\-End Human Evaluation

#### Evaluation Setup\.

To simulate realistic use, we designed a suite of 30 natural\-language test questions and had five human evaluators assess the deployed system end to end, complementing benchmark\-style evaluations of multi\-turn reasoning and cross\-domain model capability[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.02060#bib.bib19);[Joshi et al\. \(2026\)](https://arxiv.org/html/2609.02060#bib.bib17)\. Each evaluator independently posed all 30 questions to the live system and rated every response*Good*or*Bad*against a pre\-specified expected behaviour, recording a failure category for each*Bad*response\. We report end\-to\-end functional success and cross\-trial consistency; the full question list, capability groups, and failure\-category definitions are given in Appendix[D](https://arxiv.org/html/2609.02060#A4)\.

#### Results and insights\.

Across the 150 question–evaluator trials, 92% of responses were rated Good \(83–100% per evaluator\), and consistency is high: 21 of the 30 questions were rated Good by all five evaluators and 29 of 30 by a majority \(Table[2](https://arxiv.org/html/2609.02060#S4.T2)\)\. Every capability exceeds 87% except interpolation disclosure \(70%\)\. Grounding is strong, only one of the 150 responses was flagged for a fabricated number, and the dominant failures were map\-pin placement and occasional wrong\-tool calls rather than unsupported claims\. With the scoring evaluation and the blind discovery test \(Section[5](https://arxiv.org/html/2609.02060#S5)\), this shows thatMineTRACEperforms its functions reliably while keeping its answers grounded\.

Table 2:End\-to\-end human evaluation\.*Trials*is the number of question–evaluator judgements;*Success*is the fraction rated Good\.

## 5Case Study: A Blind Test on New Gold Discoveries

The evaluation above uses known deposits\. A more demanding question for a prospectivity system is whether it can highlight mineralisation that was still unknown when the model was built\. We therefore conduct a retrospective blind case study on six Western Australian gold discoveries first announced by ASX\-listed explorers in 2023–2025\([Kula Gold Limited, 2025](https://arxiv.org/html/2609.02060#bib.bib8);[Solstice Minerals Limited, 2025](https://arxiv.org/html/2609.02060#bib.bib9);[Great Western Exploration Limited, 2024](https://arxiv.org/html/2609.02060#bib.bib10);[Duketon Mining Limited, 2024](https://arxiv.org/html/2609.02060#bib.bib11);[Yandal Resources Ltd, 2024](https://arxiv.org/html/2609.02060#bib.bib12);[Yandal Resources Ltd, 2025](https://arxiv.org/html/2609.02060#bib.bib13)\)\. BecauseMineTRACEwas built from public data compiled to 2021–2022 \(Appendix[A](https://arxiv.org/html/2609.02060#A1)\), and none of these discoveries appears in the training labels, the case study tests whether the deployed system assigns high prospectivity to targets it could not have memorised\.

For each discovery, we use the coordinate reported in the company announcement or the closest available prospect centroid, score that point with the deployed model, and compare the score with the statewide gold prospectivity surface\. Figure[3](https://arxiv.org/html/2609.02060#S3.F3)shows the resulting locations and scores\. Five of the six discoveries fall in the top 13% of Western Australia, three fall in the top 10%, and the strongest case, Astro at the Barlee project, lies above the 98th percentile\. The only mid\-ranked site is Edjudina Range, which still scores near the 65th percentile\. Because these targets were not part of the supervision set, these results suggest that the model is not merely rediscovering labelled deposits: it concentrates prospectivity in areas where new gold systems were later reported\.

We use Mustang as a worked example to show why an evidence\-grounded interface is useful \(Figure[3\(a\)](https://arxiv.org/html/2609.02060#S3.F3.sf1)\)\. Mustang had no effective prior drilling in the model data, and its nearest catalogued gold deposit is approximately 53 km away\. Its score therefore cannot be explained simply by proximity to a known label\. The local gold\-geochemical evidence is also weak, because the drilling that reported gold mineralisation post\-dates the public data used by the system\. Instead, the evidence panel shows that the moderate\-to\-high score is driven by independent geological and geophysical support: a strong gravity\-gradient signal, favourable metasedimentary host context, Yilgarn Craton membership, and structural\-context features involving gravity worms and mapped faults\. These signals are consistent with the company’s description of Mustang as an early\-stage gold prospect in the Southwest Terrane Greenstones of the southwestern Yilgarn Craton, where mineralisation is associated with a significant shear\-zone setting\([Kula Gold Limited, 2025](https://arxiv.org/html/2609.02060#bib.bib8)\)\. In other words,MineTracedid not “know” that Mustang contained gold; it surfaced a defensible mineral\-systems argument that a geologist could inspect before the later drilling result was public\.

This case study also clarifies the system’s intended role\.The six discoveries are early\-stage exploration results, not defined resources, and the sample is too small for statistical validation\. Absolute scores are moderate in several cases, and the high percentile rankings partly reflect the generally low predicted gold prospectivity across Western Australia\. Its value is therefore in the workflow it demonstrates:MineTracesurfaces plausible targets, exposes the evidence behind each score\. For exploration decision support, this traceability matters as much as the ranking itself\.

## 6Conclusion

We presentedMineTRACE, a user\-friendly and evidence\-grounded interactive reasoning system for mineral prospectivity\. The system turns public exploration data into traceable evidence records that connect final scores to measured features, named experts, expert weights, source layers, and coverage metadata\. The same evidence structure powers the map interface, heatmaps, evidence panels, APIs, MCP endpoint, and a grounded conversational assistant\. This design turns a prospectivity model from a static score generator into an interactive reasoning workflow\.MineTRACEdemonstrates that interpretable domain models can support natural\-language scientific interaction when their evidence structure is preserved end to end\.

## Limitations

MineTRACEis a decision\-support system, not an autonomous or field\-validated one: it ranks and explains evidence from public data but does not replace field verification or expert geological judgement, and its scores are bounded by data coverage, assay quality, spatial sampling bias, and the completeness of public reporting, such that areas with sparse local samples receive weak or interpolated support that the interface must flag\. The expert tree, while transparent, remains a simplified model of mineral systems; its fixed expert set aids auditability but may miss commodity\-specific processes, and performance is uneven across commodities, being strong for Ni and Co but weak for Cu and Mn under spatial holdout\. Similarly, although the conversational assistant is grounded through tools, grounding is not correctness, since it is only as reliable as the scores and evidence records it retrieves; we verify that numerical claims are tool\-derived and that unsupported requests are declined, but larger user studies are still needed to measure how geologists use the system in practice\.

## Acknowledgement

This research was gratefully supported by the Australian Research Council \(ARC\) Training Centre for Critical Resources for the Future \(CCRF\) under grant number IC230100035\. The authors acknowledge the computational resources and technical support provided by the Kaya High Performance Computing facility at the University of Western Australia \(UWA\), which were essential to the data processing and analysis conducted in this study\. We also thank the Western Australian Mineral Exploration \(WAMEX\) database, administered by the Geological Survey of Western Australia, for providing free and public access to the exploration report data used throughout this work\.

## References

- Dinget al\.\(2025\)Y\. Ding, S\. C\. Han, Z\. Li, and H\. ChungSynjac: synthetic\-data\-driven joint\-granular adaptation and calibration for domain specific scanned document key information extraction\.Information Fusion,pp\. 104074\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Dinget al\.\(2026\)Y\. Ding, Y\. Zhang, C\. Gonzalez, E\. Holden, and W\. LiuGeoChemAD: benchmarking unsupervised geochemical anomaly detection for mineral exploration\.arXiv preprint arXiv:2603\.13068\.Cited by:[§1](https://arxiv.org/html/2609.02060#S1.p2.1)\.
- Dong and Zhang \(2024\)Y\. Dong and Z\. ZhangDeep forest modeling: an interpretable deep learning method for mineral prospectivity mapping\.Journal of Geophysical Research: Machine Learning and Computation1\(4\),pp\. e2024JH000311\.Cited by:[§1](https://arxiv.org/html/2609.02060#S1.p2.1),[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Duketon Mining Limited \(2024\)Duketon Mining LimitedSeptember 2024 quarterly report\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://announcements.asx.com.au/asxpdf/20241024/pdf/069hcpnnf7mvhq.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1)\.
- Great Western Exploration Limited \(2024\)Great Western Exploration LimitedFirebird aircore drilling results received\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://company-announcements.afr.com/asx/gte/7e1ccabc-cab5-11ee-be79-0abdb9403284.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1)\.
- Joshiet al\.\(2026\)H\. Joshi, G\. S\. Kashyap, R\. Ali, E\. Shabbir, N\. Jain, S\. Jain, J\. Gao, and U\. NaseemCan argus judge them all? comparing vlms across domains\.External Links:2507\.01042,[Link](https://arxiv.org/abs/2507.01042)Cited by:[§4\.2](https://arxiv.org/html/2609.02060#S4.SS2.SSS0.Px1.p1.1)\.
- Kashyapet al\.\(2026\)G\. S\. Kashyap, M\. Dras, and U\. NaseemWe think, therefore we align llms to helpful, harmless and honest before they go wrong\.External Links:2509\.22510,[Link](https://arxiv.org/abs/2509.22510)Cited by:[§3\.4](https://arxiv.org/html/2609.02060#S3.SS4.SSS0.Px3.p1.1)\.
- Kula Gold Limited \(2025\)Kula Gold LimitedMustang gold prospect – results update\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://www.kulagold.com.au/wp-content/uploads/2025/04/02935171.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1),[§5](https://arxiv.org/html/2609.02060#S5.p3.1)\.
- Liuet al\.\(2026\)Y\. Liu, W\. Zhang, C\. Cao, W\. Lu, F\. Yuan, D\. Guo, K\. Peng, Q\. Sun, K\. Zhang, Y\. Liu,et al\.PRISMA: reinforcement learning guided two\-stage policy optimization in multi\-agent architecture for open\-domain multi\-hop question answering\.arXiv preprint arXiv:2601\.05465\.Cited by:[§1](https://arxiv.org/html/2609.02060#S1.p2.1)\.
- Mülleret al\.\(2025\)D\. Müller, D\. I\. Groves, M\. Santosh, and C\. YangCritical metals: their mineral systems and exploration\.Geosystems and Geoenvironment4\(1\),pp\. 100323\.Cited by:[§1](https://arxiv.org/html/2609.02060#S1.p1.1)\.
- Raiet al\.\(2026\)A\. K\. Rai, U\. Tripathi, G\. Kumar, S\. Siddique, V\. Singh, R\. Sathikumar, and J\. BagchiGold prospectivity mapping in the eastern part of mahakoshal fold belt, india: a comparative study of random forest and xgboost leveraging knowledge\-guided feature engineering\.Geosystems and Geoenvironment,pp\. 100532\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Solstice Minerals Limited \(2025\)Solstice Minerals LimitedEdjudina range gold discovery ready for first rc drilling\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://solsticeminerals.com.au/upload/documents/investor/asx/250502003137_250502EdjudinaRangeGoldDiscoveryReadyForFirstRCDrillingfinal.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1)\.
- Wang and Chen \(2025\)X\. Wang and Y\. ChenUnsupervised detection of multivariate geochemical anomalies using a high\-performance deep autoencoder gaussian mixture model\.Journal of Geochemical Exploration271,pp\. 107671\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Yandal Resources Ltd \(2024\)Yandal Resources LtdEmerging gold discovery within the new england granite prospect\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://announcements.asx.com.au/asxpdf/20241021/pdf/069bssr29yf7qq.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1)\.
- Yandal Resources Ltd \(2025\)Yandal Resources LtdArrakis gold discovery confirmed with 54 m @ 1\.2 g/t au from 108 m\.Note:ASX AnnouncementAccessed: 2026\-07\-08External Links:[Link](https://announcements.asx.com.au/asxpdf/20250922/pdf/06pgs3znt24cpy.pdf)Cited by:[§5](https://arxiv.org/html/2609.02060#S5.p1.1)\.
- Yuet al\.\(2026\)S\. Yu, H\. Deng, X\. Liu, Y\. Zheng, Z\. Liu, J\. Chen, and X\. MaoExpectation–maximization\-derived self\-distillation meets transformer: a robust unsupervised deep learning approach for geochemical anomaly recognition\.Mathematical Geosciences58\(2\),pp\. 279–312\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Zhanget al\.\(2024\)S\. Zhang, E\. Carranza, C\. Fu, Z\. Wen\-zhi, and Q\. XiangInterpretable machine learning for geochemical anomaly delineation in the yuanbo nang district, gansu province, china\. minerals, 14 \(5\): 500\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Zhanget al\.\(2026a\)W\. Zhang, Y\. Liu, Q\. Sun, Y\. Ding, S\. Li, Y\. Liu, J\. B\. Hong, and W\. LiuSTIndex: a context\-aware multi\-dimensional spatiotemporal information extraction system\.InCompanion Proceedings of the ACM Web Conference 2026,pp\. 69–72\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Zhanget al\.\(2026b\)Y\. Zhang, M\. Lin, M\. Dras, and U\. NaseemBeyond the black box: demystifying multi\-turn llm reasoning with vista\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 41745–41747\.Cited by:[§2](https://arxiv.org/html/2609.02060#S2.p1.1)\.
- Zhanget al\.\(2025\)Y\. Zhang, M\. Wang, X\. Li, K\. Ren, C\. Zhu, and U\. NaseemTurnBench\-MS: a benchmark for evaluating multi\-turn, multi\-step reasoning in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 19892–19924\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1084/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1084),ISBN 979\-8\-89176\-335\-7Cited by:[§4\.2](https://arxiv.org/html/2609.02060#S4.SS2.SSS0.Px1.p1.1)\.

## Appendix AData and Preprocessing Details

The corpus comprises 9,352,545 assay samples across five sampling media \(Table[3](https://arxiv.org/html/2609.02060#A1.T3)\) and 3,420 positive sites across eight commodities: 1,057 Cu, 581 Ni, 492 Sn, 383 Co, 306 Ta, 287 Au, 215 Mn, and 99 W\. Geochemistry is drawn from GSWA*CM02 Near Surface Geochemistry*, deposit labels from*Mineralization Sites*, and the geophysical rasters and geological vectors from the 2021*CM08 Critical Minerals Basemap*; all products are openly licensed and standardised to GDA2020\.

For preprocessing, we map allCM02 Geochem\.records to a common schema containing coordinates, sampling medium, and 123 element or oxide fields\. We convert the−9999\-9999sentinel to missing, set negative below\-detection values to zero, split the samples into five medium\-specific tables, and apply the transformationlog⁡\(1\+x\)\\log\(1\{\+\}x\)to reduce skew\. For CM01, we merge the commodity\-specific exports into one supervision table per target commodity, treating polymetallic sites as positive for each associated commodity\. We retain sites classified as*Mine*or*Deposit*and exclude*Prospect*and*Occurrence*records\. All processed assay, site, raster, and vector layers are stored in PostGIS\. When no assay samples fall within the query radius, the system returns no geochemical score rather than extrapolating into unsupported areas\.

Table 3:Per\-medium assay\-sample counts \(GSWA*CM02 Near Surface Geochemistry*\)\.
## Appendix BNeighbourhood Feature Catalogue

Every feature is a statistic over samples within a neighbourhood radiuss∈\{5,10,50\}s\\in\\\{5,10,50\\\}km of the query location, computed onlog⁡\(1\+x\)\\log\(1\{\+\}x\)concentrations to reduce the effect of heavy censoring in assay data\. Element\-level features are instantiated for each element in a fixed 14\-element panel \(target metals and common pathfinders\); target metals outside this panel \(Ta, Mn\) therefore have no element\-level features for the target element itself\. Detailed feature descriptions are available at[https://github\.com/grantzyr/GeoResearchPlatformPublic](https://github.com/grantzyr/GeoResearchPlatformPublic)\. Section[3\.3](https://arxiv.org/html/2609.02060#S3.SS3)describes how the model composes them into a score\.

## Appendix CPathfinders and Expert Definitions

#### Commodity pathfinders\.

Each commodity is scored on its target element and a fixed suite of pathfinder elements and element ratios drawn from exploration geochemistry\. The pathfinder weights encode how diagnostic each element is of the target system and are fixed from domain knowledge, whereas the favourable direction of every feature is learned from data during offline fitting\. Target metals outside the 14\-element panel \(Ta, Mn\) are scored through these pathfinders and ratios rather than through a target\-element anomaly\.

#### Expert definitions\.

The ten experts are organised into three families \(Table[4](https://arxiv.org/html/2609.02060#A3.T4)\)\. The three geochemical experts each aggregate evidence across all active assay media \(Table[3](https://arxiv.org/html/2609.02060#A1.T3)\): every medium is fitted independently and weighted by its training ROC\-AUC, so the geochemical expert count is fixed at three regardless of how many media cover a query\. The four geophysical and three geological experts are single\-source leaves over the corresponding raster and vector layers\. Within every expert, features are z\-scored against a fitted background, combined by a weighted mean under a learned favourable direction \(±1\\pm 1\) per feature, and squashed to\[0,1\]\[0,1\]; an expert whose required evidence is absent abstains and is dropped from the aggregation\.

FamilyExpertEvidence evaluated and primary inputsGeochemicalTarget enrichmentLocal enrichment of the target element above regional background: local\-to\-regional log contrasts \(5–50 km, 10–50 km\) and high\-percentile and fraction\-above statistics at 5 and 10 km\.PathfinderAnomalies in the commodity’s pathfinder elements and element ratios, weighted by the domain pathfinder weights\.Element correlationJoint co\-enrichment of element pairs among the target and its pathfinders, scored as the product of their local contrasts\.GeophysicalMagneticMagnetic intensity and gradient\.GravityBouguer gravity and gradient\.RadiometricK, Th, and U channels, their gradients, and the K/Th, Th/U, and U/K ratios\.GeochronologyIsotopic crustal\-age proxies \(Lu–Hf, Sm–Nd\) and their gradients\.GeologicalFaultDistance to the nearest mapped fault and fault density at 5 and 10 km\.WormDistance to magnetic and gravity worms \(multi\-scale potential\-field edges\)\.GeologyHost\-rock lithology class \(granitic, felsic, mafic, ultramafic, metasedimentary, sedimentary, metamorphic, hydrothermal\), crustal age, craton membership \(Yilgarn, Pilbara\), and Cenozoic\-cover flag\.Table 4:The ten named experts across three families\. Geochemical experts aggregate over all active assay media; geophysical and geological experts are single\-source leaves over the geophysical rasters and geological vector layers, respectively\.

## Appendix DHuman Evaluation Details

The 30 questions used in Section[4](https://arxiv.org/html/2609.02060#S4)exercise seven user\-facing capabilities, each targeting a distinct part of the exploration workflow \(Table[5](https://arxiv.org/html/2609.02060#A4.T5)\)\. Each*Bad*response was tagged with one of ten failure categories: fabricated data or score; coverage/interpolation not disclosed; map\-location or pin error; layer\-toggle error; wrong tool call or workflow not followed; scope\-handling error \(wrong refusal or acceptance\); unsound comparison or recommendation; unclear evidence explanation; ambiguity not clarified; and other\.

Table 5:The 30 human\-evaluation questions grouped by capability \(cf\. Table[2](https://arxiv.org/html/2609.02060#S4.T2)\)\.

Similar Articles

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Hugging Face Daily Papers

TRACE is a taxonomy-guided environment with 1,000 visual reasoning tasks across 11 domains. Training Qwen2.5-VL-3B and Qwen2.5-VL-7B on 64,000 TRACE instances improves their macro-average performance across 24 external benchmarks by 3.51 and 4.06 percentage points respectively.

WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning

arXiv cs.CL

WildTrace is a benchmark of 481 tasks using naturally occurring evidence trails from long documents like technical reports and narratives. It evaluates 18 frontier systems and shows that even the best (75.3%) struggles with reasoning-intensive geometries such as counterfactual branching and causal attribution.