BaseCamp —— 一个自动化 DNA 测序数据管道的智能体 AI 框架

arXiv cs.AI 论文

摘要

本文介绍了 BaseCamp,这是一个智能体 AI 框架,通过使用专门的 AI 智能体处理诸如质量控制、比对和变异检测等任务,实现对端到端 DNA 测序管道中决策层的自动化。

arXiv:2609.28557v1 Announce Type: new Abstract: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decomposes the pipeline into six specialized AI agents, covering sample intake and quality control, alignment, variant calling, annotation, cross-stage monitoring, and reporting. Critically, BaseCamp agents do not perform sequence analysis: established tools execute alignment, calling, and annotation, while the agents select among them, configure them, interpret their output, and decide what follows. This confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility existing tooling guarantees. Agent reasoning is powered by a consortium of fine-tuned, domain-specialized large language models coordinated by a central reasoning LLM, executing locally so no sequencing data leaves the operating environment, under human-in-the-loop orchestration. Evaluation shows agent-generated configurations are concordant with expert practice, that an explicit filtering ledger renders inspectable what filtering otherwise removes without trace, and that cross-stage anomaly detection surfaces conditions execution monitoring misses. BaseCamp offers a generalizable blueprint for agentic automation of scientific data pipelines.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:27

# BaseCamp — An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
Source: [https://arxiv.org/html/2609.28557](https://arxiv.org/html/2609.28557)
Journal:Journal NameXueping LiangEmail:[xuliang@fiu\.edu](mailto:[email protected])Address:Deloitte & Touche LLP, USAAsanga GunaratnaEmail:[asanga\.gunaratna@complianceoslab\.app](mailto:[email protected])Address:AI Motion Labs, Melbourne, AustraliaTharaka HewaEmail:[tharaka\.hewa@oulu\.fi](mailto:[email protected])Address:Center for Wireless Communications, University of Oulu, FinlandAbdul RahmanEmail:[abdulrahman@deloitte\.com](mailto:[email protected])Address:Deloitte & Touche LLP, USAPeter FoytikEmail:[pfoytik@odu\.edu](mailto:[email protected])Address:Old Dominion University, Norfolk, VA, USASafdar H\. BoukEmail:[sbouk@odu\.edu](mailto:[email protected])Address:Old Dominion University, Norfolk, VA, USASachini RajapakseEmail:[sachini\.rajapakse@iciclelabs\.ai](mailto:[email protected])Address:IcicleLabs\.AIIsurunima KularathnaEmail:[isurunima\.kularathna@iciclelabs\.ai](mailto:[email protected])Address:IcicleLabs\.AIPramoda KarunarathnaEmail:[pramoda\.karunarathna@iciclelabs\.ai](mailto:[email protected])Address:IcicleLabs\.AIChalani RajapakseEmail:[chalani\.rajapakse@iciclelabs\.ai](mailto:[email protected])Address:IcicleLabs\.AINg Wee KeongEmail:[awkng@ntu\.edu\.sg](mailto:[email protected])Address:Nanyang Technological University, SingaporeKasun De ZoysaEmail:[kasun@ucsc\.cmb\.ac\.lk](mailto:[email protected])Address:University of Colombo, Sri LankaAmin HassEmail:[amin\.hassanzadeh@accenture\.com](mailto:[email protected])Address:Accenture Technology Labs, Arlington, VA, USAWathsala HerathEmail:[wathsala\.herath@agentsway\.ai](mailto:[email protected])Address:Agentsway\.AIRoss GoreEmail:[rgore@odu\.edu](mailto:[email protected])Address:Old Dominion University, Norfolk, VA, USARavi MukkamalaEmail:[rmukkama@odu\.edu](mailto:[email protected])Address:Old Dominion University, Norfolk, VA, USANihal SiriwardanageaEmail:[nihal@gsi2\.com](mailto:[email protected])Address:GSI Scandinavia ABGihan SiriwardanageaEmail:[gihan\.siriwardanagea@stud\.lsmu\.lt](mailto:[email protected])Address:Lithuanian University of Health SciencesAruna WithanageEmail:[aruna@effectz\.ai](mailto:[email protected])Address:Effectz\.AINilaan LoganathanEmail:[nilaan@effectz\.ai](mailto:[email protected])Address:Effectz\.AISachin ShettyEmail:[sshetty@odu\.edu](mailto:[email protected])Address:Old Dominion University, Norfolk, VA, USAAddress:Florida International University, USA

###### Abstract

DNA sequencing pipelines — spanning raw read quality control, reference alignment, variant calling, and annotation — are now reliably executed by workflow management systems that orchestrate established bioinformatics tools reproducibly at scale\. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a given sample and platform, interpreting quality reports in context, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review\. These decisions are repetitive, judgment\-intensive, inconsistently exercised across operators, and frequently undocumented\. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of end\-to\-end DNA sequencing pipelines\. The framework decomposes the pipeline into six specialized AI agents — covering sample intake and quality control, alignment, variant calling, annotation and interpretation, cross\-stage monitoring, and reporting\. Critically, BaseCamp agents do not perform sequence analysis: established, independently validated tools execute alignment, calling, and annotation, while the agents select among them, configure them, interpret their output, and decide what follows\. This division confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility that existing tooling guarantees\. To ensure contextual accuracy and adherence to responsible and explainable AI principles, the framework employs a consortium of fine\-tuned, domain\-specialized large language models coordinated by a central reasoning LLM, with all inference executing locally so that no sequencing data leaves the operating environment\. A human\-in\-the\-loop orchestration model allows bioinformaticians to supervise, validate, and intervene across pipeline stages via a Model Context Protocol \(MCP\)–enabled interface\. Evaluation demonstrates that agent\-generated configurations are concordant with expert practice, that an explicit filtering ledger renders inspectable what filtering otherwise removes without trace, and that cross\-stage anomaly detection surfaces conditions — samples passing every individual stage check yet jointly implausible — that execution monitoring does not detect at all\. Although grounded in a genomics operational environment, BaseCamp is domain\-independent and offers a generalizable blueprint for agentic AI–driven automation of scientific data pipelines\.

###### Keywords:

Agentic AI , DNA Sequencing , Bioinformatics Pipeline Automation , Variant Calling , Responsible AI , Explainable AI , LLM , Model Context Protocol

## 1Introduction

DNA sequencing has become a foundational tool across genomics, agricultural breeding, clinical diagnostics, and biomedical research, with sequencing throughput increasing and per\-genome cost falling steadily over the past decade\[[65](https://arxiv.org/html/2609.28557#bib.bib41),[60](https://arxiv.org/html/2609.28557#bib.bib42)\]\. Converting raw sequencer output into interpretable genetic information requires a multi\-stage computational pipeline spanning raw read quality control, adapter trimming, reference alignment, duplicate marking, variant calling, filtering, and functional annotation\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[41](https://arxiv.org/html/2609.28557#bib.bib45),[42](https://arxiv.org/html/2609.28557#bib.bib46)\]\. Mature, widely adopted tools exist for each individual stage\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[17](https://arxiv.org/html/2609.28557#bib.bib48),[48](https://arxiv.org/html/2609.28557#bib.bib49),[37](https://arxiv.org/html/2609.28557#bib.bib50)\], and pipeline orchestration frameworks such as Nextflow and Snakemake\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52)\]enable these tools to be chained into reproducible computational workflows\. Despite this tooling maturity, the decision\-making layer surrounding pipeline execution — selecting QC thresholds appropriate to a given sample and sequencing platform, choosing alignment and variant\-calling parameters, interpreting borderline variant calls, triaging pipeline failures, and prioritizing which findings warrant expert review — remains predominantly manual, dependent on bioinformatician availability, and inconsistent across operators and sequencing runs\.

The emergence of agentic AI systems represents a fundamental shift in how such operational complexity can be addressed\[[54](https://arxiv.org/html/2609.28557#bib.bib37),[69](https://arxiv.org/html/2609.28557#bib.bib15)\]\. Unlike deterministic workflow managers that execute a fixed sequence of tool invocations, agentic AI systems are capable of autonomous reasoning, multi\-step planning, and context\-aware decision\-making across interconnected pipeline stages\[[51](https://arxiv.org/html/2609.28557#bib.bib35)\]\. Networks of specialized AI agents can continuously interpret quality control metrics, reason over alignment statistics, evaluate variant call confidence in light of sequencing depth and platform\-specific error profiles, and flag anomalous runs for review — tasks that currently depend on the availability and judgment of an experienced bioinformatician and are prone to inconsistency when performed under time pressure across large sample batches\[[15](https://arxiv.org/html/2609.28557#bib.bib34)\]\. By delegating these reasoning\-intensive responsibilities to autonomous agents operating under human supervision, sequencing facilities and research groups can achieve levels of throughput, consistency, and reproducibility that manual oversight cannot sustain at scale\.

Despite this potential, the application of agentic AI to DNA sequencing pipelines remains largely unexplored in practice\. Most existing applications of large language models in genomics focus on isolated capabilities: genomic language models that learn sequence\-level representations directly from DNA\[[34](https://arxiv.org/html/2609.28557#bib.bib53),[49](https://arxiv.org/html/2609.28557#bib.bib54)\], or biomedical question\-answering systems that retrieve and summarize genomic literature\[[35](https://arxiv.org/html/2609.28557#bib.bib55)\], rather than end\-to\-end reasoning over the operational pipeline that produces variant calls in the first place\. These point solutions improve specific sub\-tasks but do not address the coordination structure of the sequencing workflow itself — the sequence of decisions connecting raw reads to a reviewed, reported variant set\. The integration of multiple autonomous agents into a cohesive, human\-supervised pipeline capable of reasoning across the full sequencing lifecycle, from raw read intake to annotated variant report, represents an open challenge that existing agentic AI transition approaches have not yet addressed in this domain\[[11](https://arxiv.org/html/2609.28557#bib.bib36)\]\.

A central barrier to this transition is the absence of a principled framework for decomposing the sequencing pipeline into agent\-executable analytical roles, designing the coordination interfaces between agents and existing bioinformatics tooling, and maintaining meaningful human oversight at each decision point\[[55](https://arxiv.org/html/2609.28557#bib.bib33),[8](https://arxiv.org/html/2609.28557#bib.bib28)\]\. Conventional pipeline automation is well suited to deterministic, rule\-bound execution, but is poorly suited to the contextual judgment that pipeline operation frequently requires — for example, whether a borderline variant call warrants confidence given local depth and mapping quality, or whether an unusual quality control profile reflects a genuine sample issue rather than a benign platform artifact\. At the same time, sequencing pipelines routinely inform decisions with direct downstream consequences — clinical variant interpretation, breeding selection, or diagnostic reporting — where errors are costly and where responsible AI design, explainability, and human\-in\-the\-loop oversight are essential requirements rather than optional enhancements\[[12](https://arxiv.org/html/2609.28557#bib.bib31),[28](https://arxiv.org/html/2609.28557#bib.bib27)\]\.

In response to these challenges, this paper proposes BaseCamp, a novel agentic AI framework for automating end\-to\-end DNA sequencing data pipelines\. BaseCamp introduces a multi\-agent architecture in which the manual decision\-making layer of the sequencing pipeline is systematically decomposed into specialized AI agents, each responsible for a clearly defined analytical role\[[59](https://arxiv.org/html/2609.28557#bib.bib22)\]\. Agents are coordinated through a human\-in\-the\-loop orchestration model in which bioinformaticians supervise, validate, and intervene at defined checkpoints via a Model Context Protocol \(MCP\)–enabled interface\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\]\. To ensure contextual accuracy and responsible, explainable decision\-making, the framework employs a consortium of fine\-tuned, domain\-specialized large language models coordinated by a central reasoning LLM\[[70](https://arxiv.org/html/2609.28557#bib.bib11),[3](https://arxiv.org/html/2609.28557#bib.bib20)\], enabling agents to reason jointly over pipeline logs, quality metrics, alignment statistics, and variant call outputs\. This architecture follows the same consortium\-plus\-reasoning\-LLM design applied in our prior work on retail supply chain automation\[[10](https://arxiv.org/html/2609.28557#bib.bib40)\]and tourism SME workflow automation\[[14](https://arxiv.org/html/2609.28557#bib.bib38)\], extending it to a domain in which the underlying computation is dominated by established bioinformatics tools rather than by natural\-language business processes, and in which the value of agentic reasoning lies in the judgment layer surrounding those tools rather than in replacing them\.

As a proof of concept, BaseCamp was implemented and evaluated using real\-world sequencing datasets\. The sequencing operational environment exhibits characteristics that are well suited to agentic AI automation, including high sample throughput, well\-defined but sample\-dependent quality thresholds, multiple sequential decision points at which contextual judgment is required, and a strong requirement for reproducibility and audit trail — properties that render it a representative and rigorously challenging validation environment for the proposed framework\. The primary contributions of this paper are fourfold:

1. 1\.A novel agentic AI framework, BaseCamp, for end\-to\-end automation of DNA sequencing data pipelines, encompassing quality control, alignment, variant calling, annotation, and pipeline exception handling under a unified multi\-agent architecture\.
2. 2\.A human\-in\-the\-loop orchestration model in which bioinformaticians act as supervisors and decision\-makers over autonomous agent workflows via MCP\-enabled interfaces, ensuring accountability, transparency, and operational control throughout the sequencing lifecycle\.
3. 3\.A responsible and explainable AI design incorporating a consortium of fine\-tuned, domain\-specialized LLMs coordinated by a reasoning LLM, supporting explainable, context\-aware, and reproducible decision\-making across all agent interactions\.
4. 4\.A real\-world proof of concept demonstrating the practical applicability of BaseCamp within a genomics sequencing operational environment, validating the framework’s effectiveness in reducing manual pipeline configuration overhead and improving anomaly detection across sequencing runs\.

The remainder of this paper is organized as follows\. Section 2 provides background on key concepts underlying the BaseCamp framework, including large language models, agentic AI systems, LLM consortiums, and human\-in\-the\-loop orchestration\. Section 3 reviews related work in bioinformatics pipeline automation, LLM\-based agents for genomics, and agentic AI workflow systems\. Section 4 describes the manual sequencing pipeline workflows that BaseCamp targets and establishes the operational context\. Section 5 presents the BaseCamp framework, detailing the multi\-agent architecture, agent roles, coordination model, and LLM consortium design\. Section 6 describes the proof\-of\-concept implementation and evaluation of BaseCamp within a real\-world sequencing operational environment\. Section 7 concludes the paper and outlines directions for future research\.

## 2Background

This section introduces the key concepts and technologies that underpin the BaseCamp framework\. Understanding these foundations is necessary before examining how they are applied to automate DNA sequencing data pipelines in subsequent sections\.

### 2\.1Large Language Models

Large Language Models \(LLMs\) are a class of artificial intelligence systems trained on vast corpora of text data, enabling them to understand, generate, and reason over natural language with a high degree of fluency and contextual awareness\[[5](https://arxiv.org/html/2609.28557#bib.bib2),[18](https://arxiv.org/html/2609.28557#bib.bib3)\]\. Unlike earlier\-generation models trained for narrow, task\-specific purposes, LLMs acquire broad general knowledge and linguistic capability through exposure to diverse text sources, allowing them to perform summarisation, question answering, structured data extraction, and complex reasoning through natural language interaction alone\[[66](https://arxiv.org/html/2609.28557#bib.bib6),[33](https://arxiv.org/html/2609.28557#bib.bib13)\]\.

It is important to distinguish the role of LLMs in BaseCamp from that of genomic language models\. Models such as DNABERT and HyenaDNA\[[34](https://arxiv.org/html/2609.28557#bib.bib53),[49](https://arxiv.org/html/2609.28557#bib.bib54)\]are trained directly on nucleotide sequences and learn representations of the DNA sequence itself\. BaseCamp does not operate on nucleotide sequences\. Its LLMs reason over theoperational artifactsthat a sequencing pipeline produces — FastQC quality reports, alignment summary statistics, duplication and coverage metrics, variant call format \(VCF\) records, annotation tables, and tool logs\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[42](https://arxiv.org/html/2609.28557#bib.bib46),[23](https://arxiv.org/html/2609.28557#bib.bib73)\]\. These artifacts are predominantly semi\-structured text and tabular summaries accompanied by tool\-generated narrative output, and the capacity to interpret such heterogeneous input and emit structured, actionable decisions is precisely what makes agentic reasoning tractable in this domain\. The computation itself remains the responsibility of established bioinformatics tools; the LLM supplies the judgment layer around them\.

### 2\.2Reasoning LLMs

While standard LLMs excel at language understanding and generation, reasoning LLMs extend these capabilities by incorporating explicit chain\-of\-thought processes, multi\-step planning, and self\-verification mechanisms that enable more reliable and structured decision\-making\[[70](https://arxiv.org/html/2609.28557#bib.bib11),[62](https://arxiv.org/html/2609.28557#bib.bib17)\]\. Reasoning LLMs are designed to evaluate intermediate conclusions, identify inconsistencies across multiple sources of evidence, and synthesise coherent final outputs from complex, multi\-perspective inputs\.

In the BaseCamp framework, a central reasoning LLM — specifically OpenAI GPT\-OSS\[[3](https://arxiv.org/html/2609.28557#bib.bib20)\]— serves as the supervisory model within the LLM consortium architecture\. When multiple fine\-tuned domain\-specialised LLMs produce independent outputs in response to an agent’s prompt, the reasoning LLM evaluates, compares, and synthesises these responses to generate a final decision or recommendation\. This ensemble\-based reasoning approach reduces the risk of model\-specific bias or error propagation, and ensures that automated pipeline decisions are validated across multiple reasoning perspectives before execution\. This requirement is acute in a sequencing context: a decision to apply a hard quality filter, to accept a borderline variant call, or to pass a marginal sample through to downstream analysis propagates silently into every result derived from it, and is frequently not revisited once made\.

### 2\.3LLM Fine\-Tuning

While pre\-trained LLMs possess broad general knowledge, their performance on domain\-specific tasks — particularly those requiring precise operational vocabulary, structured output formats, and context\-aware reasoning within a specific technical domain — can be significantly improved through fine\-tuning\[[53](https://arxiv.org/html/2609.28557#bib.bib10),[44](https://arxiv.org/html/2609.28557#bib.bib16)\]\. Fine\-tuning adapts a general\-purpose pre\-trained base model on a curated, domain\-specific dataset through supervised training, enabling the model to internalise domain conventions, terminologies, and reasoning patterns that are not well represented in general pre\-training corpora\. Genomics is a strong candidate for such adaptation: quality thresholds, filter expressions, VCF field semantics, and platform\-specific error profiles constitute a dense technical vocabulary that a general model handles inconsistently\.

As illustrated in Figure[1](https://arxiv.org/html/2609.28557#S2.F1), the fine\-tuning pipeline employed in BaseCamp begins with a base LLM — such as Llama\-3, Mistral, or Qwen\[[27](https://arxiv.org/html/2609.28557#bib.bib7),[61](https://arxiv.org/html/2609.28557#bib.bib18),[64](https://arxiv.org/html/2609.28557#bib.bib9)\]— which is adapted using a curated sequencing operations dataset comprising historical QC reports, pipeline execution logs, parameter configurations paired with the outcomes they produced, variant filtering decisions, and bioinformatician review notes\. To enable parameter\-efficient training within resource\-constrained environments, Low\-Rank Adapters \(LoRA\)\[[7](https://arxiv.org/html/2609.28557#bib.bib1)\]are applied during fine\-tuning, and 4\-bit quantisation \(QLoRA\)\[[24](https://arxiv.org/html/2609.28557#bib.bib5)\]further reduces memory requirements\. Fine\-tuned models are deployed locally using Ollama\[[52](https://arxiv.org/html/2609.28557#bib.bib4)\], enabling low\-latency agent inference without dependence on external API services — a property that also matters where sequencing data is subject to institutional or regulatory restrictions on external transmission\.

SEQUENCING OPERATIONS CORPUSSUPERVISED FINE\-TUNINGLOCAL DEPLOYMENTQC reports\+ dispositionspipeline logs\+ parameter settingsfiltering decisions\+ recorded rationalereview notesbioinformaticianBase LLMLlama\-3⋅\\cdotMistral⋅\\cdotQwen 2\.5curation & split80 / 10 / 10fine\-tuningUnsloth⋅\\cdotLoRA \+ QLoRAdomain adapterone per functionOllama local servinghot\-swap adapter per agentlow\-latency inferenceno external API dependencydata stays on\-siteno sequencing data egressFine\-tuning is performed offline; deployed model behaviour is fixed during pipeline execution, so two runs of the same sample are reconcilable\.

Figure 1:Supervised fine\-tuning pipeline used to adapt a base LLM to sequencing pipeline operations\. A curated corpus of QC reports, pipeline logs and parameter settings, variant filtering decisions, and bioinformatician review notes is used to fine\-tune a base model, with Low\-Rank Adapter \(LoRA\) modules applied for parameter\-efficient training under 4\-bit quantisation\[[7](https://arxiv.org/html/2609.28557#bib.bib1),[24](https://arxiv.org/html/2609.28557#bib.bib5)\]\. The resulting adapter retains the general linguistic capability of the base model while internalising sequencing\-domain vocabulary and reasoning patterns, and is served locally via Ollama\[[52](https://arxiv.org/html/2609.28557#bib.bib4)\]so that no sequencing data leaves the operating environment\.
### 2\.4AI Agents and Agentic AI

Traditional LLM interactions follow a simple request–response pattern in which a human constructs a prompt, submits it to a language model, and manually interprets the response to determine subsequent actions\. In this paradigm the human remains responsible for orchestration, decision\-making, and follow\-up execution, with the LLM functioning as a passive assistant rather than an active participant in workflow execution\[[1](https://arxiv.org/html/2609.28557#bib.bib12),[8](https://arxiv.org/html/2609.28557#bib.bib28)\]\.

An AI agent fundamentally changes this interaction model\. As illustrated in Figure[2](https://arxiv.org/html/2609.28557#S2.F2), an AI agent autonomously manages the full interaction loop — constructing prompts, invoking language models, interpreting responses, invoking external tools, and determining subsequent actions — without direct human intervention at each step\. AI agents are software systems that leverage LLMs in combination with tools, APIs, memory, and external context to execute tasks iteratively and autonomously\[[9](https://arxiv.org/html/2609.28557#bib.bib21)\]\. When multiple specialised agents collaborate, each assigned distinct responsibilities, they formagentic AI workflows\[[54](https://arxiv.org/html/2609.28557#bib.bib37),[69](https://arxiv.org/html/2609.28557#bib.bib15),[11](https://arxiv.org/html/2609.28557#bib.bib36)\]\.

The distinction matters for sequencing because the tool\-invocation capability is not incidental — it is the mechanism by which agentic reasoning is grounded in real computation\. A BaseCamp agent does not estimate alignment quality; it invokes SAMtools or a comparable tool\[[42](https://arxiv.org/html/2609.28557#bib.bib46),[41](https://arxiv.org/html/2609.28557#bib.bib45)\], reasons over the returned statistics, and decides what follows\. This separation — established tools compute, agents decide — distinguishes the framework from approaches that ask a language model to perform analysis it is not suited to, and it preserves the reproducibility guarantees that existing bioinformatics tooling already provides\.

\(a\) DIRECT HUMAN–LLM INTERACTIONBioinformaticianconstruct promptpaste QC report / VCF excerptLLMinterpret responseand act manuallythe human performs every orchestration step\(b\) AGENTIC AI–LLM INTERACTIONLLMBioinformaticiansupervisorAI Agentconstruct promptinvoke modelinterpret responseinvoke tooldecide next actionFastQC⋅\\cdotMultiQCBWA⋅\\cdotSAMtoolsGATK⋅\\cdotDeepVariantVEP⋅\\cdotannotation DBsgoalapprove / overridereasoninvoke / returnestablished tools computeThe agent decides; it does not compute\. Sequence analysis remains the responsibility of established, independently validated tools\.

Figure 2:Comparison between direct human–LLM interaction and agentic AI–LLM interaction in a sequencing context\. In panel \(a\) the bioinformatician constructs each prompt, submits it to the model, and manually interprets and acts on the response, remaining responsible for every orchestration step\. In panel \(b\) an AI agent autonomously manages the interaction loop — prompt construction, model invocation, response interpretation, external tool invocation, and determination of subsequent actions — while the bioinformatician supplies the goal and retains approval authority\. Crucially, the agent does not itself perform sequence analysis: it invokes established bioinformatics tools\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[41](https://arxiv.org/html/2609.28557#bib.bib45),[42](https://arxiv.org/html/2609.28557#bib.bib46),[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[48](https://arxiv.org/html/2609.28557#bib.bib49)\]and reasons over their output, preserving the reproducibility those tools already guarantee\.
### 2\.5Model Context Protocol

The Model Context Protocol \(MCP\) is a standardised, open protocol defining how AI agents connect to and interact with external systems, tools, databases, and APIs through secure, structured interfaces\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\]\. MCP enables agents to access real\-time data and invoke external capabilities — such as launching an alignment job, querying a variant annotation database, or retrieving run metadata from a laboratory information management system \(LIMS\) — in a modular and interoperable manner, without requiring custom integration code for each system connection\.

In BaseCamp, each agentic workflow is exposed as an independent MCP server, enabling a single bioinformatician to connect to and orchestrate multiple agent workflows simultaneously through a unified natural language interface via LM Studio\[[36](https://arxiv.org/html/2609.28557#bib.bib32)\]\. This architecture decouples agent logic from the underlying systems it drives: the QC Agent reaches FastQC and MultiQC through its own MCP server, the Alignment Agent reaches BWA and SAMtools through another, the Variant Calling Agent the calling and filtering stack through a third, and so on\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[30](https://arxiv.org/html/2609.28557#bib.bib59),[41](https://arxiv.org/html/2609.28557#bib.bib45),[42](https://arxiv.org/html/2609.28557#bib.bib46),[47](https://arxiv.org/html/2609.28557#bib.bib43)\]\. Individual workflows can therefore evolve independently — a variant caller can be upgraded, or a filtering policy revised, without disrupting the broader system\. Because servers are individually addressable, a bioinformatician who disputes a single stage can re\-run that stage alone rather than recomputing the entire pipeline, which is the common case when a run is questioned in part rather than in whole, and a material advantage where full re\-execution is measured in hours\. The resulting deployment topology is described in Section[6\.1](https://arxiv.org/html/2609.28557#S6.SS1)\.

### 2\.6Responsible and Explainable AI

As agentic AI systems take on increasingly consequential roles in scientific data pipelines — setting quality thresholds, filtering variant calls, and determining which findings reach human review — ensuring that their decisions are transparent, accountable, and reproducible becomes a foundational requirement rather than an optional enhancement\[[56](https://arxiv.org/html/2609.28557#bib.bib26),[28](https://arxiv.org/html/2609.28557#bib.bib27)\]\. Responsible AI encompasses fairness, transparency, accountability, privacy, and human oversight in the design and deployment of AI systems\. Explainable AI \(XAI\) refers specifically to the capacity of a system to provide human\-interpretable justification for its outputs — enabling operators to understand not merely what was decided, but why\[[12](https://arxiv.org/html/2609.28557#bib.bib31),[6](https://arxiv.org/html/2609.28557#bib.bib30)\]\.

The stakes in a sequencing context are distinctive in two respects\. First, pipeline decisions aresilent and cumulative: a filter applied at the variant calling stage removes calls that no downstream analysis will ever see, and the absence leaves no trace unless it is deliberately recorded\. Second, sequencing results routinely inform consequential downstream decisions — clinical variant interpretation, breeding selection, or diagnostic reporting — where an error propagates well beyond the pipeline that produced it\. Reproducibility is therefore not merely good practice but a precondition for the scientific validity of anything derived from the output\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52)\]\.

BaseCamp operationalises these principles through four complementary mechanisms\. First, the LLM consortium architecture ensures that consequential decisions are validated across multiple independent reasoning perspectives before execution, reducing the risk of model\-specific bias or hallucination\. Second, every agent output is structured to include an explicit reasoning trace: each QC pass or fail cites the specific metric and threshold that determined it, each parameter selection records its justification, and each variant filtering decision records the criteria applied and the number of calls affected\. Third, human approval gates are embedded at defined decision points, so that no consequential action — such as committing a filtered variant set to downstream analysis — proceeds without explicit validation\. Fourth, the full configuration under which a run was executed — tool versions, parameters, model versions, and adapter revisions — is recorded with the output, so that any result can be reproduced or recomputed under different settings\. Together these mechanisms ensure that BaseCamp delivers pipeline automation while preserving the auditability and reproducibility that scientific use demands\.

## 3Related Works

Research relevant to the BaseCamp framework spans four interconnected areas: bioinformatics workflow management and pipeline automation, LLM\-based agents for genomics and biomedical analysis, agentic AI for scientific discovery and autonomous experimentation, and agentic AI and multi\-agent workflow systems more generally\. The following subsections review representative works in each area and identify the gaps that BaseCamp addresses\.

### 3\.1Bioinformatics Workflow Management and Pipeline Automation

The dominant approach to sequencing pipeline automation is the workflow management system\. Nextflow\[[25](https://arxiv.org/html/2609.28557#bib.bib51)\]and Snakemake\[[39](https://arxiv.org/html/2609.28557#bib.bib52)\]allow bioinformaticians to express a pipeline as a directed graph of tool invocations with declared inputs, outputs, and dependencies, and to execute that graph reproducibly across heterogeneous compute environments through containerisation and portable execution backends\. Galaxy\[[2](https://arxiv.org/html/2609.28557#bib.bib56)\]provides equivalent capability through a web\-based graphical interface aimed at users without command\-line expertise, and the Common Workflow Language\[[22](https://arxiv.org/html/2609.28557#bib.bib57)\]supplies a vendor\-neutral specification for describing such workflows portably\. Building on these engines, nf\-core\[[29](https://arxiv.org/html/2609.28557#bib.bib58)\]curates a community\-maintained collection of peer\-reviewed, standardised pipelines — including germline and somatic variant calling workflows — that encode established best practice and substantially reduce the effort required to deploy a correct pipeline\. Reporting tools such as MultiQC\[[30](https://arxiv.org/html/2609.28557#bib.bib59)\]aggregate quality metrics across samples and pipeline stages into consolidated reports intended for human inspection\.

Collectively these systems solve theexecutionproblem, and solve it well: given a specified pipeline and a specified parameter set, they run it reproducibly, at scale, with full provenance\. What they do not address is thedecisionproblem that surrounds execution\. A workflow manager will faithfully execute whatever quality threshold it is given, but it does not reason about whether that threshold is appropriate for a particular sample, library preparation, or sequencing platform\. It will emit a MultiQC report, but it does not interpret one\. It will fail a job, but it does not diagnose why, nor decide whether the correct response is re\-running with adjusted parameters, flagging the sample for exclusion, or escalating to a human\. These judgments remain the responsibility of the bioinformatician, are exercised inconsistently across operators, and constitute precisely the layer BaseCamp targets\. BaseCamp is therefore complementary to rather than competitive with these systems: agents invoke established tooling and, where appropriate, existing workflow engines, and supply the reasoning layer above them\.

### 3\.2LLM\-Based Agents for Genomics and Biomedical Analysis

A rapidly growing body of work applies large language models to genomic and biomedical tasks\. One strand trains language models directly on nucleotide sequences: DNABERT\[[34](https://arxiv.org/html/2609.28557#bib.bib53)\]and HyenaDNA\[[49](https://arxiv.org/html/2609.28557#bib.bib54)\]learn representations of DNA as a language, supporting downstream tasks such as regulatory element prediction and variant effect estimation\. This strand is orthogonal to the present work — these models operate on the sequence itself, whereas BaseCamp operates on the operational artifacts a pipeline produces, and the two could in principle be composed, with a genomic language model invoked as a tool by an interpretation agent\.

A second strand equips general\-purpose LLMs with biomedical tools and data access\. GeneGPT\[[35](https://arxiv.org/html/2609.28557#bib.bib55)\]augments an LLM with NCBI Web APIs to answer genomics questions, demonstrating that tool augmentation substantially outperforms parametric recall for factual genomic queries\. BioMANIA\[[26](https://arxiv.org/html/2609.28557#bib.bib60)\]translates natural language instructions into executable calls against bioinformatics Python libraries, lowering the barrier to tool use for non\-programmers\. AutoBA\[[72](https://arxiv.org/html/2609.28557#bib.bib61)\]advances further toward autonomy, proposing an LLM\-based agent that plans and executes multi\-omic analyses end\-to\-end from a description of the input data and analytical goal, and CellAgent\[[67](https://arxiv.org/html/2609.28557#bib.bib62)\]applies a multi\-agent decomposition — planner, executor, evaluator — to single\-cell RNA sequencing analysis tasks\.

This strand is the closest prior art, and AutoBA and CellAgent in particular establish the feasibility on which BaseCamp builds: LLM agents can plan and execute real bioinformatics analyses at usable quality\. Three differences define the gap\. First, these systems targetdownstream analysis— differential expression, cell type annotation, multi\-omic integration — taking a processed data matrix as their starting point, whereas BaseCamp targets theupstream production pipelinethat generates such data from raw reads, where the decisions concern quality thresholds, alignment parameters, and variant filtering\. Second, they rely predominantly on general\-purpose frontier models accessed via external APIs, with no domain\-specific fine\-tuning and no provision for local execution — a constraint that is disqualifying where sequencing data is subject to institutional or regulatory restriction on external transmission\. Third, they emphasise autonomy and task completion, with limited attention to structured human approval gates, per\-decision reasoning traces, or the configuration recording that reproducible scientific use requires\.

### 3\.3Agentic AI for Scientific Discovery and Autonomous Experimentation

Beyond genomics, agentic AI has been applied to scientific work more broadly\. Coscientist\[[16](https://arxiv.org/html/2609.28557#bib.bib63)\]demonstrates an LLM\-driven system that autonomously plans, designs, and executes chemical experiments using laboratory robotics, and ChemCrow\[[46](https://arxiv.org/html/2609.28557#bib.bib64)\]augments an LLM with expert\-designed chemistry tools to perform synthesis planning and reaction prediction, in both cases showing that tool\-augmented agents substantially outperform the underlying model operating alone\. The AI Scientist\[[45](https://arxiv.org/html/2609.28557#bib.bib65)\]extends the paradigm to the research process itself, generating hypotheses, running experiments, and drafting manuscripts autonomously\.

These works establish the broader premise — that agentic decomposition with tool augmentation is effective for technical scientific work — and they also surface the concerns that motivate BaseCamp’s design constraints\. Each emphasises autonomous capability, and the accompanying discussions consistently identify oversight, verifiability, and misuse as the open problems rather than raw capability\. BaseCamp occupies a deliberately narrower operating envelope: it does not generate hypotheses, design experiments, or produce scientific conclusions\. It automates the decision layer of an established, well\-characterised production pipeline whose correct operation is already defined by community best practice\[[29](https://arxiv.org/html/2609.28557#bib.bib58),[47](https://arxiv.org/html/2609.28557#bib.bib43)\], and it does so under explicit human approval gates\. The value claimed is consistency and throughput in a known process, not autonomous discovery\.

### 3\.4Agentic AI and Multi\-Agent Workflow Systems

Sapkota et al\.\[[54](https://arxiv.org/html/2609.28557#bib.bib37)\]provide a conceptual taxonomy distinguishing AI agents from agentic AI systems, clarifying the architectural and behavioural differences between single\-agent tools and multi\-agent systems capable of sustained reasoning and coordinated action, and identifying orchestration, role decomposition, and inter\-agent communication as the principal design challenges\. Their contribution is conceptual and does not address application to a specific operational domain\. Broader surveys of LLM\-based autonomous agents\[[69](https://arxiv.org/html/2609.28557#bib.bib15),[33](https://arxiv.org/html/2609.28557#bib.bib13)\]similarly characterise the design space — planning, memory, tool use, multi\-agent coordination — without prescribing a transition path from an existing manual process to an agentic one\.

Bandara et al\.\[[9](https://arxiv.org/html/2609.28557#bib.bib21)\]introduce AgentsWay, a development methodology for teams building agentic AI systems, emphasising AI\-assisted development, rapid iteration, and domain\-driven workflow design\. BaseCamp follows this methodology directly\. Prior applications of the same consortium\-plus\-reasoning\-LLM architecture have targeted retail supply chain automation\[[10](https://arxiv.org/html/2609.28557#bib.bib40)\], tourism SME operations\[[14](https://arxiv.org/html/2609.28557#bib.bib38)\], peer\-based mental health support in resource\-constrained environments\[[68](https://arxiv.org/html/2609.28557#bib.bib39)\], and false memory risk assessment in investigative and legal contexts\[[38](https://arxiv.org/html/2609.28557#bib.bib66)\], in each case demonstrating that decomposition into specialised agents under human\-in\-the\-loop orchestration is effective where a manual process is coordination\-heavy and judgement\-intensive\[[8](https://arxiv.org/html/2609.28557#bib.bib28),[11](https://arxiv.org/html/2609.28557#bib.bib36)\]\.

BaseCamp applies this architecture to a domain with a distinguishing property\. In the operational applications cited, agents both reason and act: the consortium synthesises toward an action — a purchase order, an itinerary, a replenishment plan — and the agent’s output largely constitutes the work product\. In a sequencing pipeline the computation is performed by established, independently validated tools\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[42](https://arxiv.org/html/2609.28557#bib.bib46),[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[48](https://arxiv.org/html/2609.28557#bib.bib49)\], and the agent’s contribution is confined to the decisions surrounding them — which tool, which parameters, which threshold, what to do when a stage behaves unexpectedly, and what warrants human attention\. This division is deliberate and load\-bearing: it preserves the reproducibility guarantees that the existing tooling already provides, and it keeps the agent’s role within the bounds where LLM reasoning is reliable\.

### 3\.5Positioning

The pattern across the reviewed literature is consistent\. Workflow management systems solve reproducible execution but not the decisions that configure it\. Genomic language models operate on sequence rather than on pipeline operations\. LLM agents for bioinformatics target downstream analysis rather than the upstream production pipeline, and generally depend on external frontier\-model APIs without domain fine\-tuning or local execution\. Agentic systems for scientific discovery demonstrate the paradigm but prioritise autonomy over the oversight and reproducibility that production pipelines require\. And general agentic AI frameworks supply the architectural pattern without application to sequencing\.

BaseCamp is distinguished not by any single column but by their conjunction: end\-to\-end coverage of the sequencing production pipeline from raw read intake to annotated variant report, decomposed across specialised agents that invoke established bioinformatics tooling rather than replacing it, powered by a consortium of locally served fine\-tuned models with a reasoning layer, under human approval gates at every consequential decision point, with full configuration recording so that any run is reproducible\. Table[1](https://arxiv.org/html/2609.28557#S3.T1)presents a comparative analysis of the reviewed works in relation to the proposed framework\.

Table 1:Comparison of related works and the BaseCamp framework\. ✓ = supported; ✗ = not supported;∼\\sim= partially supported\.WorkDomainSequencingpipelinefocusEnd\-to\-endworkflowAgenticmulti\-agentLLMbasedLLMfine\-tuningHuman\-in\-the\-loopLocal /air\-gappedexecutionBaseCamp \(ours\)DNA sequencingpipeline automation✓✓✓✓✓✓✓Nextflow\[[25](https://arxiv.org/html/2609.28557#bib.bib51)\]Workflowmanagement✓✓✗✗✗⚫✓Snakemake\[[39](https://arxiv.org/html/2609.28557#bib.bib52)\]Workflowmanagement✓✓✗✗✗⚫✓Galaxy\[[2](https://arxiv.org/html/2609.28557#bib.bib56)\]Workflowplatform✓✓✗✗✗✓⚫nf\-core\[[29](https://arxiv.org/html/2609.28557#bib.bib58)\]Curated genomicspipelines✓✓✗✗✗⚫✓DNABERT\[[34](https://arxiv.org/html/2609.28557#bib.bib53)\]Genomic languagemodel✗✗✗✓✓✗✓HyenaDNA\[[49](https://arxiv.org/html/2609.28557#bib.bib54)\]Genomic sequencemodelling✗✗✗✓✓✗✓GeneGPT\[[35](https://arxiv.org/html/2609.28557#bib.bib55)\]Genomics questionanswering✗✗✗✓✗✗✗BioMANIA\[[26](https://arxiv.org/html/2609.28557#bib.bib60)\]NL\-driventool invocation⚫✗✗✓✗⚫✗AutoBA\[[72](https://arxiv.org/html/2609.28557#bib.bib61)\]Automated multi\-omicanalysis⚫⚫⚫✓✗⚫✗CellAgent\[[67](https://arxiv.org/html/2609.28557#bib.bib62)\]Single\-cell RNA\-seqanalysis✗⚫✓✓✗✗✗Coscientist\[[16](https://arxiv.org/html/2609.28557#bib.bib63)\]Autonomous chemicalexperimentation✗✓⚫✓✗⚫✗ChemCrow\[[46](https://arxiv.org/html/2609.28557#bib.bib64)\]LLM chemistrytool agent✗⚫✗✓✗⚫✗Sapkota et al\.\[[54](https://arxiv.org/html/2609.28557#bib.bib37)\]Agentic AItaxonomy✗✗✓✓✗✗✗Bandara et al\.\[[9](https://arxiv.org/html/2609.28557#bib.bib21)\]Agentic AImethodology✗✗✓✓✓✓✓Flowr\[[10](https://arxiv.org/html/2609.28557#bib.bib40)\]Retail supply chainautomation✗✓✓✓✓✓✓Rovanima\[[14](https://arxiv.org/html/2609.28557#bib.bib38)\]Tourism SMEautomation✗✓✓✓✓✓✓

## 4Manual DNA Sequencing Pipeline Workflows

A successful agentic AI transition begins with a thorough understanding of the manual processes that automation intends to replace\. In the context of DNA sequencing operations, the pipeline function encompasses a set of recurring, decision\-intensive stages spanning raw read quality assessment, reference alignment, variant calling, functional annotation, and run\-level monitoring across sample batches\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\. These stages are deeply interdependent — a permissive trimming decision propagates into alignment statistics, which in turn condition variant call confidence — yet in practice each is configured and reviewed as a discrete step, frequently by different personnel, with the connecting judgments recorded informally if at all\.

It is important to be precise about what is and is not manual here\. Toolexecutionis largely automated: workflow managers such as Nextflow and Snakemake\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52)\], and curated pipeline collections such as nf\-core\[[29](https://arxiv.org/html/2609.28557#bib.bib58)\], reliably orchestrate the invocation of FastQC, BWA, GATK, and their counterparts across large sample batches\. What remains manual is thedecision layersurrounding that execution: selecting thresholds appropriate to a given sample and platform, interpreting quality reports, adjudicating borderline variant calls, diagnosing pipeline failures, and determining which results warrant expert attention\. As depicted in Figure[3](https://arxiv.org/html/2609.28557#S4.F3), these decisions are made sequentially by distinct human roles — from laboratory technicians assessing run quality through to analysts reviewing annotated variant sets — with no automated reasoning layer connecting them\. The decisions are highly repetitive, require synthesis of information across heterogeneous outputs, and are sensitive to operator experience in ways that manual processes cannot standardise at scale\. These characteristics — continuous monitoring requirements, repeated judgment\-intensive decisions, cross\-source reasoning, and the need for human oversight at critical points — are precisely the conditions under which agentic AI workflows deliver the greatest operational value, as identified in our earlier work on agentic AI transition frameworks\[[11](https://arxiv.org/html/2609.28557#bib.bib36),[8](https://arxiv.org/html/2609.28557#bib.bib28)\]\. The manual processes underlying each stage are described in detail in the following subsections\.

HUMAN ROLELab technicianBioinformaticianBioinformaticianVariant analystDomain scientistPIPELINE STAGEQC & preprocessingFastQC / MultiQC review;adapter & quality trimmingAlignmentreference selection;aligner parameters;duplicate markingVariant callingcaller selection;calling parameters;hard / soft filteringAnnotationfunctional annotation;frequency filtering;effect prioritisationReportingresult assembly;interpretation;downstream handoffMANUAL DECISIONwhich thresholds?resequence or proceed?which reference?accept mapping rate?which filters?trust marginal calls?which variantswarrant review?what is reportable?what is caveated?manual handoffmanual handoffmanual handoffmanual handoffEXCEPTION HANDLING — reactive, detected after the factlow\-coverage samplescontamination / index hoppingaberrant call\-rate shiftsbatch effects across runsExceptions surface only when an operator notices an anomaly in a report or when a downstream result appears implausible; there is no systematic mechanism for proactive detection across the run\.

Figure 3:Manual DNA sequencing pipeline workflow\. Laboratory technicians, bioinformaticians, variant analysts, and domain scientists execute sequential stages — quality control, alignment, variant calling, annotation, and reporting — with tool execution largely automated by workflow managers\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]but the surrounding decision layer \(amber\) remaining manual and operator\-dependent\. Handoffs between stages are informal, and the reasoning behind each decision is rarely recorded in a form that survives into the next stage\. Exceptions such as low coverage, contamination, aberrant call rates, and batch effects are detected reactively — typically after they have already propagated downstream — limiting the ability to intervene before results are affected\.### 4\.1Raw Read Quality Control and Preprocessing

The pipeline begins when sequencer output is demultiplexed and quality assessed\. Bioinformaticians run FastQC on each sample and aggregate the results using MultiQC\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[30](https://arxiv.org/html/2609.28557#bib.bib59)\], producing reports covering per\-base quality distributions, GC content, adapter contamination, duplication levels, and overrepresented sequences\. Interpreting these reports is a manual task requiring contextual judgment: a GC distribution that would indicate contamination in a whole\-genome sample may be entirely expected in an amplicon or reduced\-representation library, and a duplication rate that is alarming in one protocol is normal in another\. Trimming decisions — which adapter sequences to remove, what quality threshold to apply, what minimum read length to retain — are typically made by applying a laboratory default and adjusting when results appear anomalous\[[17](https://arxiv.org/html/2609.28557#bib.bib48),[20](https://arxiv.org/html/2609.28557#bib.bib67)\]\.

The consequences of misjudgment at this stage are asymmetric and easily overlooked\. Over\-aggressive trimming discards genuine signal and reduces effective coverage, while under\-trimming propagates adapter and low\-quality sequence into alignment, inflating mismatch rates and depressing mapping quality in ways that are subsequently attributed to sample quality rather than to preprocessing\. In high\-throughput settings, per\-sample review is often infeasible, and samples are processed in batches under a single parameter set that suits the majority but not the outliers\. Which samples are genuine outliers requiring individual attention, as opposed to acceptable variation, is a determination made by eye from aggregate plots\.

### 4\.2Reference Alignment and Parameter Selection

Once reads are preprocessed, they are aligned to a reference genome using tools such as BWA\-MEM or Bowtie2\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[40](https://arxiv.org/html/2609.28557#bib.bib68)\], followed by sorting, duplicate marking, and, in some workflows, base quality score recalibration\[[42](https://arxiv.org/html/2609.28557#bib.bib46),[47](https://arxiv.org/html/2609.28557#bib.bib43)\]\. Reference selection is itself a consequential manual decision\. In clinical and model\-organism contexts a standard reference assembly is available, but in agricultural and non\-model\-organism genomics — including crop and horticultural breeding programmes — the appropriate reference may be an assembly of a related cultivar or species, and the choice materially affects mapping rate and downstream variant discovery, particularly in divergent regions\.

Following alignment, the bioinformatician reviews mapping statistics: alignment rate, properly paired proportion, mean and distribution of coverage, insert size distribution, and duplication rate\. Deciding whether these statistics are acceptable requires joint reasoning across several of them at once — a reduced mapping rate accompanied by an unusual insert size distribution suggests a library preparation issue, whereas the same mapping rate with a normal insert distribution and elevated unmapped reads may indicate contamination or reference mismatch\. This joint interpretation is performed manually, is rarely documented, and is frequently deferred under throughput pressure with the sample simply passed forward\.

### 4\.3Variant Calling and Filtering

Variant calling converts aligned reads into a set of putative genetic variants, using callers such as GATK HaplotypeCaller, DeepVariant, or bcftools\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[42](https://arxiv.org/html/2609.28557#bib.bib46)\]\. Caller selection depends on sequencing depth, platform, ploidy, and whether germline or somatic variants are sought, and the choice is typically fixed by laboratory convention rather than reasoned per experiment\. Calling parameters — minimum base and mapping quality, ploidy assumptions, and handling of multi\-allelic sites — are similarly inherited from prior runs\.

Filtering is the most judgment\-intensive step in the pipeline\. Raw call sets contain substantial numbers of false positives, and filtering strategies range from hard threshold filters on annotations such as depth, quality by depth, strand bias, and mapping quality, to model\-based approaches such as variant quality score recalibration where sufficient training data exists\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[23](https://arxiv.org/html/2609.28557#bib.bib73)\]\. Selecting thresholds involves an explicit trade\-off between sensitivity and specificity whose appropriate balance depends on the downstream application — a screening context tolerates false positives that a reporting context does not — and this trade\-off is typically resolved by convention rather than by per\-run reasoning\.

Two properties make this stage particularly consequential\. Filtering decisions aresilent: a filtered variant simply does not appear downstream, and its absence leaves no trace unless deliberately recorded\. And they arecumulative: every subsequent analysis inherits the filtered set without visibility into what was removed or why\. In manual workflows, the rationale for a given threshold is frequently held only in the operator’s working memory or an informal note, and is not recoverable when a result is later questioned\.

### 4\.4Annotation and Functional Interpretation

Filtered variants are annotated with functional consequence predictions, population allele frequencies, and, where relevant, prior evidence from curated databases, using tools such as the Ensembl Variant Effect Predictor, SnpEff, or ANNOVAR against resources including gnomAD and dbSNP\[[48](https://arxiv.org/html/2609.28557#bib.bib49),[21](https://arxiv.org/html/2609.28557#bib.bib69),[63](https://arxiv.org/html/2609.28557#bib.bib70),[37](https://arxiv.org/html/2609.28557#bib.bib50),[57](https://arxiv.org/html/2609.28557#bib.bib71)\]\. Annotation execution is well automated; the manual work lies in the prioritisation that follows\.

An annotated call set for a single sample commonly contains many thousands of variants, of which a small number are of interest for the study at hand\. Reducing this set requires layered filtering on consequence severity, population frequency, region of interest, inheritance pattern, or association with traits under investigation\. Analysts construct these filters manually, frequently in spreadsheets or ad hoc scripts, and the specific combination applied varies with the analyst and the question\. In trait\-mapping and breeding contexts — linking genotype to phenotypes such as yield, fruit quality, or disease resistance — an additional manual step joins the variant set to phenotypic records held in separate systems, a cross\-source integration that is laborious and error\-prone\. The filtering logic is rarely captured alongside the results, so a prioritised variant list is often difficult to reproduce even by its author\.

### 4\.5Pipeline Monitoring and Exception Handling

Throughout the pipeline, exceptions arise that require human diagnosis: failed or stalled jobs, samples with anomalously low coverage, contamination or index hopping between multiplexed samples, unexpected shifts in variant call rate, and batch effects distinguishing one sequencing run from another\[[30](https://arxiv.org/html/2609.28557#bib.bib59),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\. Workflow managers detect and reportexecutionfailures reliably — a job that crashes is surfaced immediately — butanalyticalanomalies, where every job completes successfully yet the results are wrong, are not detected by execution monitoring at all\.

In manual workflows, detection of such anomalies depends on an operator noticing an irregularity in an aggregate report, or on a downstream analysis producing an implausible result that prompts retrospective investigation\. There is no systematic mechanism for proactive detection across the full pipeline, which means that many issues are identified only after affected results have already been used\. Once detected, diagnosis requires tracing back through pipeline stages to determine whether the cause is sample quality, a preprocessing decision, an alignment problem, or a calling parameter — a reconstruction that is time\-consuming and depends on records that manual workflows frequently do not retain\. The reactive and reconstruction\-dependent nature of manual exception handling limits the ability to intervene before compromised results propagate downstream\.

## 5Delegating Manual Sequencing Processes to AI Agents

Once the manual workflows are well understood, the next step in the agentic AI transition is to systematically delegate them to a coordinated network of autonomous AI agents\[[69](https://arxiv.org/html/2609.28557#bib.bib15),[11](https://arxiv.org/html/2609.28557#bib.bib36)\]\. This delegation goes beyond automating isolated tasks — it requires decomposing human\-performed decision\-making into distinct reasoning steps, decision points, and coordination actions that can be meaningfully assigned to specialised agents, each operating within a clearly defined scope\[[54](https://arxiv.org/html/2609.28557#bib.bib37),[59](https://arxiv.org/html/2609.28557#bib.bib22)\]\. The goal is not to replicate the existing manual process in software, but to capture the underlying intent and judgment that guide human work across the sequencing lifecycle\.

One design decision governs the entire framework and warrants statement before the agents are described\.BaseCamp agents do not perform sequence analysis\.Alignment is performed by BWA, variant calling by GATK or DeepVariant, annotation by VEP — established, independently validated tools whose correctness properties are well characterised and whose outputs are reproducible\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[48](https://arxiv.org/html/2609.28557#bib.bib49)\]\. The agents select among these tools, configure them, interpret their output, decide what follows, and determine what warrants human attention\. This division is deliberate and load\-bearing: it confines LLM reasoning to the judgment layer where it is reliable, preserves the reproducibility guarantees the existing tooling already provides, and ensures that a BaseCamp result is a result an established tool produced rather than one a language model generated\.

In practice, delegation begins by identifying the cognitive responsibilities embedded within each stage of the manual workflow — quality interpretation, alignment assessment, calling and filtering strategy, variant prioritisation, anomaly detection, and result communication\. Each responsibility is mapped to a dedicated agent with a well\-defined role, a set of inputs and outputs, and access via MCP to the tools and data sources required to execute it\. Agents operate in coordination, passing structured intermediate outputs between stages so that downstream agents reason with full upstream context rather than in isolation — a property manual handoffs conspicuously lack, as noted in Section[4](https://arxiv.org/html/2609.28557#S4)\.

Figure[4](https://arxiv.org/html/2609.28557#S5.F4)illustrates how the manual pipeline of Figure[3](https://arxiv.org/html/2609.28557#S4.F3)is decomposed into the BaseCamp agentic workflow\. Six specialised agents — a Sample Intake and QC Agent, an Alignment Agent, a Variant Calling Agent, an Annotation and Interpretation Agent, a Pipeline Monitoring and Exception Agent, and a Reporting Agent — operate as a coordinated system under human supervision, each connected to relevant tools and data sources through MCP servers\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\]\.

HUMAN OVERSIGHTAGENTIC ANALYSIS LAYEREXTERNAL TOOLS \(via MCP\)CROSS\-STAGE MONITORINGON\-DEVICE MODEL LAYERBioinformaticianhuman orchestratorsupervises⋅\\cdotapproves⋅\\cdotoverridesA1Sample Intake & QCinterpret quality metricsset trimming parametersA2Alignmentreference & aligner choiceassess mapping statisticsA3Variant Callingcaller strategy; thresholdscall adjudicationA4Annotationfunctional annotationfrequency filteringA5Reportingresult assemblyreasoning trace \+ configHHHHFastQC⋅\\cdotMultiQC⋅\\cdotfastpBWA⋅\\cdotBowtie2⋅\\cdotSAMtoolsGATK⋅\\cdotDeepVariantVEP⋅\\cdotgnomAD⋅\\cdotdbSNPLIMS⋅\\cdotphenotype storeA6Pipeline Monitoring & Exception Agentcontinuous cross\-stage anomaly detection: coverage shortfall, contamination, aberrant call rates, batch effectsproactive escalationfine\-tuned LLM consortium \(Llama\-3⋅\\cdotMistral⋅\\cdotQwen 2\.5\) \+ reasoning LLM \(GPT\-OSS\)served locally via Ollama — no data egress

Figure 4:The BaseCamp agentic AI workflow for automated DNA sequencing pipelines\. Six specialised agents replace the manual decision layer of Figure[3](https://arxiv.org/html/2609.28557#S4.F3): five operate sequentially across the pipeline lifecycle, passing structured outputs so that each reasons with full upstream context, while the Pipeline Monitoring and Exception Agent observes all stages continuously and escalates anomalies proactively rather than awaiting retrospective discovery\. Agents invoke established bioinformatics tools\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[41](https://arxiv.org/html/2609.28557#bib.bib45),[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[48](https://arxiv.org/html/2609.28557#bib.bib49)\]through MCP servers\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\]rather than performing analysis themselves\. Human approval gates \(H\) are embedded at consequential transitions, and all agent reasoning is served by a locally deployed fine\-tuned LLM consortium coordinated by a reasoning LLM\.### 5\.1Sample Intake and QC Agent

The Sample Intake and QC Agent is responsible for ingesting demultiplexed sequencer output, invoking quality assessment tooling, and reasoning over the resulting metrics to determine preprocessing parameters and sample disposition\. The agent interfaces with FastQC, MultiQC, and preprocessing tools via MCP server connections\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[30](https://arxiv.org/html/2609.28557#bib.bib59),[20](https://arxiv.org/html/2609.28557#bib.bib67)\], retrieving per\-base quality distributions, GC content, adapter contamination, duplication levels, and overrepresented sequence reports\.

Its principal contribution is contextual interpretation\. As described in Section[4\.1](https://arxiv.org/html/2609.28557#S4.SS1), a metric that indicates a problem under one library preparation is expected under another, and this distinction is exactly what a fixed threshold cannot encode\. Using a fine\-tuned LLM trained on historical QC reports paired with the dispositions bioinformaticians assigned them, the agent reasons jointly over the metric profile, the declared library type, the sequencing platform, and the sample’s position within the batch, and produces a structured determination: recommended trimming parameters with explicit justification, a pass, review, or fail disposition, and an identification of samples whose profile deviates from the batch in ways warranting individual attention\. Because per\-sample review is infeasible manually at scale, the agent’s ability to reason per sample rather than per batch is where the operational gain arises\. Outputs pass to the Alignment Agent; samples flagged as failing or ambiguous are escalated for human adjudication\.

### 5\.2Alignment Agent

The Alignment Agent selects an appropriate reference and aligner configuration, invokes alignment and post\-processing, and assesses the resulting statistics\. It interfaces with BWA, Bowtie2, and SAMtools via MCP connections\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[40](https://arxiv.org/html/2609.28557#bib.bib68),[42](https://arxiv.org/html/2609.28557#bib.bib46)\], and with reference genome repositories and, where applicable, laboratory information management systems for sample provenance\.

Two reasoning tasks distinguish this agent\. The first is reference selection, which as noted in Section[4\.2](https://arxiv.org/html/2609.28557#S4.SS2)is consequential in non\-model\-organism contexts where the appropriate reference may be an assembly of a related cultivar or species; the agent reasons over declared sample provenance, available assemblies, and observed mapping behaviour to recommend a reference with explicit justification\. The second is joint interpretation of alignment statistics\. Rather than comparing each statistic against an independent threshold, the agent reasons across mapping rate, properly paired proportion, insert size distribution, coverage uniformity, and duplication rate simultaneously, and distinguishes characteristic failure signatures — a library preparation problem, contamination, or reference divergence — that individual thresholds cannot separate\. This joint reasoning is precisely what manual review performs inconsistently under throughput pressure\. Aligned outputs and the accompanying assessment pass to the Variant Calling Agent\.

### 5\.3Variant Calling Agent

The Variant Calling Agent determines calling strategy, invokes the selected caller, and reasons over filtering\. It interfaces with GATK, DeepVariant, and bcftools via MCP connections\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[42](https://arxiv.org/html/2609.28557#bib.bib46)\], receiving aligned data and quality context from upstream agents\.

The agent reasons over caller selection given depth, platform, ploidy, and whether germline or somatic variants are sought — a determination typically fixed by convention rather than reasoned per experiment in manual workflows — and then addresses filtering, the pipeline’s most judgment\-intensive step\. A fine\-tuned LLM enables the agent to reason about the sensitivity–specificity trade\-off in light of the stated downstream application, recommending thresholds on depth, quality by depth, strand bias, and mapping quality with explicit justification for each\.

Two design requirements follow from the properties identified in Section[4\.3](https://arxiv.org/html/2609.28557#S4.SS3)\. Because filtering decisions are silent, the agent emits an explicitfiltering ledgerrecording every threshold applied, the number of calls it removed, and the rationale — so that what was excluded remains inspectable rather than merely absent\. And because they are cumulative, the ledger accompanies the call set downstream rather than remaining with the operator\. Marginal calls falling near a threshold boundary are flagged individually for human adjudication rather than silently resolved\. The filtered call set, the ledger, and the flagged calls pass to the Annotation and Interpretation Agent\.

### 5\.4Annotation and Interpretation Agent

The Annotation and Interpretation Agent enriches the filtered call set with functional consequence predictions, population allele frequencies, and prior database evidence, and reasons over the enriched set to produce a prioritised shortlist\. It interfaces with VEP, SnpEff, and ANNOVAR, and with gnomAD and dbSNP, via MCP connections\[[48](https://arxiv.org/html/2609.28557#bib.bib49),[21](https://arxiv.org/html/2609.28557#bib.bib69),[63](https://arxiv.org/html/2609.28557#bib.bib70),[37](https://arxiv.org/html/2609.28557#bib.bib50),[57](https://arxiv.org/html/2609.28557#bib.bib71)\]\.

Annotation execution is well automated already; the agent’s contribution is the prioritisation that follows\. Reasoning over consequence severity, population frequency, region of interest, inheritance pattern, and the study’s stated objective, it reduces a call set of many thousands to a reviewable shortlist, recording the filter logic applied at each step so that the shortlist is reproducible — addressing the reproducibility gap noted in Section[4\.4](https://arxiv.org/html/2609.28557#S4.SS4), where manually constructed prioritisations are frequently difficult to reproduce even by their author\. In trait\-mapping and breeding contexts the agent additionally joins the variant set to phenotypic records retrieved from separate systems, performing the genotype–phenotype integration that is otherwise a laborious manual cross\-source step\. The prioritised set and its filter provenance pass to the Reporting Agent\.

### 5\.5Pipeline Monitoring and Exception Agent

The Pipeline Monitoring and Exception Agent observes all stages continuously, detecting conditions that fall outside expected operating parameters and require human judgment\. These include coverage shortfalls that will compromise downstream calling, contamination or index hopping between multiplexed samples, unexpected shifts in variant call rate relative to comparable runs, batch effects distinguishing one sequencing run from another, and stalled or failed pipeline stages\.

Its role is defined by a gap that workflow managers do not address\. As noted in Section[4\.5](https://arxiv.org/html/2609.28557#S4.SS5), execution monitoring reliably surfaces jobs that crash, butanalyticalanomalies — where every job completes successfully and the results are nonetheless wrong — are invisible to it\. The agent reasons across stages rather than within them, comparing observed metrics against expectations derived from the run’s own distribution and from historical comparable runs, and detecting inconsistency patterns that no single stage would flag in isolation\. Rather than awaiting retrospective discovery, it generates structured, human\-readable alerts prioritised by operational urgency, each citing the specific metric and comparison that triggered it, and each proposing candidate diagnoses — sample quality, a preprocessing decision, an alignment problem, or a calling parameter — with the supporting evidence for each\. This proposed diagnosis substantially reduces the backward\-tracing burden that manual exception handling imposes\.

### 5\.6Reporting Agent

The Reporting Agent assembles the final artifact: the prioritised variant set, the per\-stage reasoning traces accumulated across the pipeline, the filtering ledger, outstanding flagged items awaiting adjudication, and the complete configuration under which the run executed — tool versions and parameters, model versions, and adapter revisions\.

The configuration record is not incidental\. Where a sequencing result informs a consequential downstream decision and is subsequently questioned, the relevant question is not only what the pipeline produced but under what configuration, and whether a different defensible configuration would have produced something else\. Recording the settings makes that question answerable, which is a precondition for reproducible scientific use\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\. The agent generates reports targeted to their audience — a technical run summary for the bioinformatics team, a result\-focused summary for the requesting scientist — with each reported finding traceable to the pipeline stage and decision that produced it\.

### 5\.7Agent Decomposition and Workflow Coordination

The agent\-based decomposition introduces modularity, parallelism, and controllability into the sequencing pipeline\. Rather than requiring a human to coordinate between stages, each agent is exposed as an independent workflow accessible through an MCP server, enabling a single bioinformatician to orchestrate, supervise, and intervene across all six agents via a unified natural language interface, as illustrated in Figure[5](https://arxiv.org/html/2609.28557#S5.F5)\. The orchestrator interacts at defined checkpoints — reviewing QC dispositions, approving alignment configurations, validating filtering thresholds, adjudicating flagged variants, and responding to escalated exceptions — without participating in the routine execution of any stage\[[8](https://arxiv.org/html/2609.28557#bib.bib28)\]\.

Because servers are individually addressable, a bioinformatician who disputes one stage can re\-run that stage alone rather than recomputing the entire pipeline — the common case when a run is questioned in part rather than in whole, and a substantial practical advantage in a domain where full pipeline re\-execution is measured in hours or days\. Individual agents can likewise be refined, retrained, or extended independently: a variant caller can be upgraded, or a filtering policy revised, without redesigning the workflow\. Agents share state through structured intermediate outputs, so downstream agents reason with full upstream context — the QC Agent’s assessment of library quality informs the Alignment Agent’s interpretation of mapping rate, which in turn informs the Variant Calling Agent’s confidence in marginal calls\. This contextual continuity is what manual handoffs, which transmit results without the reasoning that produced them, systematically lose\.

Bioinformaticianhuman orchestratorLM Studio / MCPQC workflowreview dispositions;approve trimmingAlignment workflowapprove reference& configurationVariant callingvalidate filters;adjudicate marginalsAnnotation workflowrefine prioritisationcriteriaMonitoring workflowtriage escalatedanomaliesReporting workflowapprove & releaserun reportEach workflow is exposed as an independent MCP server, so a disputed stage can be interrogated or re\-run without recomputing the pipeline\.

Figure 5:A single bioinformatician orchestrating multiple specialised BaseCamp agentic workflows\. Each workflow automates a distinct pipeline function — quality control, alignment, variant calling, annotation, monitoring, and reporting — and is exposed through an independent MCP server\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\]for human\-supervised coordination, approval, and intervention via a unified natural language interface in LM Studio\[[36](https://arxiv.org/html/2609.28557#bib.bib32)\]\. The human role shifts from executing each stage to supervising all of them, with authority retained at every consequential decision point\.
### 5\.8Responsible and Explainable AI Agents

Responsible AI principles\[[56](https://arxiv.org/html/2609.28557#bib.bib26),[28](https://arxiv.org/html/2609.28557#bib.bib27)\]within BaseCamp are implemented through a multi\-layered architecture integrating a consortium of fine\-tuned, domain\-specialised LLMs with a central reasoning LLM\[[70](https://arxiv.org/html/2609.28557#bib.bib11),[3](https://arxiv.org/html/2609.28557#bib.bib20)\], as depicted in Figure[6](https://arxiv.org/html/2609.28557#S5.F6)\. Each agent interfaces with this consortium to ensure balanced, transparent, and context\-aware decision\-making\.

When an agent performs a task — determining trimming parameters, assessing alignment statistics, or selecting filtering thresholds — its prompt is distributed across multiple domain\-specialised LLMs such as Llama\-3, Mistral, and Qwen\[[27](https://arxiv.org/html/2609.28557#bib.bib7),[61](https://arxiv.org/html/2609.28557#bib.bib18),[64](https://arxiv.org/html/2609.28557#bib.bib9)\], each fine\-tuned for a different aspect of sequencing pipeline operations\. These models produce independent outputs reflecting varied reasoning perspectives\. The reasoning LLM then evaluates, compares, and synthesises these responses into a coherent final recommendation\.

Agent promptfrom A1–A6\+ pipeline artifacts\+ sample contextLlama\-3adapter: QC &preprocessingMistraladapter: alignment &callingQwen 2\.5adapter: annotation &interpretationrecommendation \+ cited metric \+ confidencerecommendation \+ cited metric \+ confidencerecommendation \+ cited metric \+ confidenceReasoning LLM\(GPT\-OSS\)identify agreementisolate divergencecheck each cited metricescalate if unresolvedConsolidated decision parameter / threshold ⋅\\cdotcited metrics ⋅\\cdotexplicit rationale ⋅\\cdotrecorded divergence where models disagreed ⋅\\cdotescalation flag where unresolvedUnresolved divergence between models is surfaced to the human orchestrator rather than averaged away: a decision the models cannot agree on is a decision warranting human attention\.

Figure 6:LLM consortium and reasoning LLM integration within BaseCamp\. Each agent distributes its prompt across multiple fine\-tuned domain\-specialised models, each producing an independent recommendation citing the specific pipeline metric it relied upon\. The reasoning LLM\[[70](https://arxiv.org/html/2609.28557#bib.bib11),[3](https://arxiv.org/html/2609.28557#bib.bib20)\]synthesises these into a consolidated decision with an explicit rationale\. Where the models diverge and the divergence cannot be resolved on the available evidence, it is recorded and escalated rather than averaged — disagreement about a threshold is itself a signal that the decision merits human judgment\.This ensemble\-based mechanism ensures that consequential decisions are validated across multiple perspectives before execution, reducing the risk of bias, error propagation, or context loss\[[9](https://arxiv.org/html/2609.28557#bib.bib21),[58](https://arxiv.org/html/2609.28557#bib.bib29)\]\. Three further mechanisms complete the responsible AI design\. First, every agent output carries an explicit reasoning trace citing the specific metrics relied upon: each QC disposition names the metric and threshold determining it, each parameter selection records its justification, and each filtering decision records the criteria applied and the number of calls affected\. Second, human approval gates are embedded at consequential transitions, so that no filtered call set is committed downstream and no report released without explicit validation\. Third, the complete run configuration is recorded with the output, so that any result can be reproduced or recomputed under different settings\[[12](https://arxiv.org/html/2609.28557#bib.bib31),[28](https://arxiv.org/html/2609.28557#bib.bib27)\]\. Together these mechanisms ensure that BaseCamp delivers pipeline automation while preserving the auditability and reproducibility that scientific use of sequencing results requires\.

## 6Implementation and Evaluation

This section describes the prototype implementation of BaseCamp and the evaluation conducted against it\. We first set out the implementation — the agent runtime, the deployment topology, the construction of the fine\-tuning corpus, and the fine\-tuning configuration — and then present the evaluation of two core agentic workflows within a real\-world sequencing operational environment\.

We state the scope of the evaluation plainly at the outset\. What is demonstrated is that the framework can be built within the stated constraints, that its agents produce structurally correct and citation\-bound decisions on material with known ground truth, and that its outputs are usable by practising bioinformaticians\. What isnotdemonstrated is that agentic pipeline configuration produces more accurate variant calls than expert manual configuration across the full range of sample types and study designs; that claim would require the larger comparative study described in Section[7](https://arxiv.org/html/2609.28557#S7), and nothing here should be read as establishing it\.

### 6\.1Implementation Overview

Each of the six agents was implemented using the OpenAI Agents SDK\[[19](https://arxiv.org/html/2609.28557#bib.bib14)\], which provides modular primitives for defining agent roles, reasoning strategies, and tool integrations\. The full lifecycle of agent development — prompt construction, agent definition, workflow composition, and iterative optimisation — was carried out using AI\-assisted development environments, primarily Claude Code, consistent with the AI\-native development methodology described in our earlier work\[[9](https://arxiv.org/html/2609.28557#bib.bib21)\]\. This approach substantially reduced implementation overhead and allowed effort to concentrate on the analytical specification — the reasoning criteria, the escalation conditions, and the emission constraints — rather than on scaffolding\.

All agent functionality is exposed through Model Context Protocol server interfaces\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\], one per analytical workflow, with agents exchanging structured findings through standardised endpoints\. Critically, agents invoke established bioinformatics tooling rather than performing analysis themselves: FastQC and MultiQC for quality assessment\[[4](https://arxiv.org/html/2609.28557#bib.bib47),[30](https://arxiv.org/html/2609.28557#bib.bib59)\], fastp for preprocessing\[[20](https://arxiv.org/html/2609.28557#bib.bib67)\], BWA\-MEM and SAMtools for alignment and post\-processing\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[42](https://arxiv.org/html/2609.28557#bib.bib46)\], GATK and DeepVariant for variant calling\[[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44)\], and VEP against gnomAD and dbSNP for annotation\[[48](https://arxiv.org/html/2609.28557#bib.bib49),[37](https://arxiv.org/html/2609.28557#bib.bib50),[57](https://arxiv.org/html/2609.28557#bib.bib71)\]\. Where an existing workflow engine is already deployed, agents can invoke it directly rather than orchestrating tools individually, making the framework complementary to established pipeline infrastructure\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\.

Figure[7](https://arxiv.org/html/2609.28557#S6.F7)shows the resulting topology\. The complete stack — orchestrator interface, agent runtime, model serving, bioinformatics tooling, and data store — resides within the institutional compute environment\. Network connectivity is required only for annotation database updates and offline adapter retraining, both performed outside pipeline execution\. In the current deployment, LM Studio\[[36](https://arxiv.org/html/2609.28557#bib.bib32)\]serves as the MCP\-powered orchestrator interface, through which the bioinformatician invokes workflows, reviews structured outputs, and approves or overrides decisions without code\-level intervention\.

INSTITUTIONAL COMPUTE ENVIRONMENT — no sequencing data egressHUMAN ORCHESTRATIONMCP INTERFACE LAYERAGENT RUNTIMETOOLS & DATA SOURCESLM Studio bioinformatician interface \> run QC on batch B\-2481 \> why was sample 17 flagged? \> re\-call sample 17, min\-DP 15 \> show filtering ledger \> release run reportQuality controlA1MCP serverAlignmentA2MCP serverVariant callingA3MCP serverAnnotationA4MCP serverMonitoringA5MCP serverReportingA6MCP serverAgent runtimeOpenAI Agents SDKOllamaLLM consortium \+ adaptersFastQC⋅\\cdotMultiQC⋅\\cdotfastpBWA⋅\\cdotSAMtoolsGATK⋅\\cdotDeepVariantVEP⋅\\cdotgnomAD⋅\\cdotdbSNPreference assembliesLIMS⋅\\cdotphenotype storesequencing data storeFASTQ / BAM / VCFrun configuration recordtool⋅\\cdotmodel⋅\\cdotadapter versionsEach workflow is separately addressable, so a disputed stage can be re\-run without recomputing the pipeline — material where full re\-execution is measured in hours\.

Figure 7:BaseCamp deployment and orchestration topology\. Each agentic workflow is exposed as an independent MCP server\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25),[13](https://arxiv.org/html/2609.28557#bib.bib23)\], and the bioinformatician interacts with all of them through a single natural language interface in LM Studio\[[36](https://arxiv.org/html/2609.28557#bib.bib32)\]\. The agent runtime invokes established bioinformatics tooling and annotation databases through standardised endpoints; model serving is local via Ollama\[[52](https://arxiv.org/html/2609.28557#bib.bib4)\]\. The entire stack resides within the institutional compute environment, so no sequencing data is transmitted externally — a requirement in settings subject to institutional or regulatory restriction on data movement\.
### 6\.2Fine\-Tuning Corpus and Configuration

The fine\-tuning corpus was assembled from historical pipeline operations records: QC reports paired with the dispositions bioinformaticians assigned them, alignment statistics paired with accept or investigate determinations, variant filtering configurations paired with the rationale recorded for them, and annotated exception incidents paired with their eventual diagnoses\. Each record pairs an operational artifact with the decision an experienced operator made and, where available, the justification given\. Table[2](https://arxiv.org/html/2609.28557#S6.T2)reports the composition\.

Table 2:Fine\-tuning corpus composition\. Held\-out records are drawn from sequencing runs disjoint from those contributing training data, so that reported performance is not inflated by batch\-level leakage\.Counts are placeholders\.SourceAdapterRecordsRoleQC reports \+ dispositionsQC & preprocessing1,240trainAlignment stats \+ decisionsalignment & calling980trainFilter configs \+ rationalealignment & calling610trainAnnotation & prioritisationannotation & interp\.740trainException incidentsmonitoring310trainHeld\-out runs \(all stages\)—420testOperator correctionsall150trainTotal4,450

Fine\-tuning was performed using the Unsloth library\[[71](https://arxiv.org/html/2609.28557#bib.bib8)\]on an NVIDIA A100 GPU\[[43](https://arxiv.org/html/2609.28557#bib.bib19)\], employing Low\-Rank Adaptation\[[7](https://arxiv.org/html/2609.28557#bib.bib1)\]with 4\-bit quantisation\[[24](https://arxiv.org/html/2609.28557#bib.bib5)\]\. One adapter was trained per analytical function rather than one model per agent, so that a single base model remains resident in memory and adapters are hot\-swapped as the pipeline advances\. Base models were Llama\-3, Mistral, and Qwen 2\.5\[[27](https://arxiv.org/html/2609.28557#bib.bib7),[61](https://arxiv.org/html/2609.28557#bib.bib18),[64](https://arxiv.org/html/2609.28557#bib.bib9)\], with OpenAI GPT\-OSS\[[3](https://arxiv.org/html/2609.28557#bib.bib20)\]as the reasoning model, all served locally via Ollama\[[52](https://arxiv.org/html/2609.28557#bib.bib4)\]\. Table[3](https://arxiv.org/html/2609.28557#S6.T3)reports the configuration\.

Table 3:Fine\-tuning configuration\.Values marked in red are placeholders\.ParameterValueModel configurationBase modelsLlama\-3, Mistral, Qwen 2\.5 \(4\-bit\)Reasoning modelGPT\-OSSMaximum sequence length8,192tokensPrecisionBFloat16Training configurationPer\-device batch size2Gradient accumulation steps4Effective batch size8Maximum training steps400Learning rate1×10−41\\times 10^\{\-4\}Warmup steps20Schedulerlinear decayOptimizerAdamW \(8\-bit\)Weight decay0\.01Early stopping patience10evaluationsLoRA configurationRankrr/α\\alpha32/32LoRA dropout0Target modulesq, k, v, o, gate, up, down projAdapters trained4 \(one per analytical function\)Trainable parameters per adapter0\.53%of baseResourcesTraining time per adapter∼\\sim16minPeak reserved memory9\.4GBAgent decision latency \(single stage\)∼\\sim8s

The final row is operationally significant\. Agent decision latency is measured in seconds, against pipeline stages whose tool execution is measured in minutes to hours\. The reasoning layer therefore imposes negligible overhead on total pipeline runtime — BaseCamp does not trade throughput for judgment\.

Figure[8](https://arxiv.org/html/2609.28557#S6.F8)shows convergence behaviour for the alignment and calling adapter, together with per\-decision\-class accuracy on the held\-out split\.

steploss01002003004000\.51\.01\.52\.0trainingvalidation\(a\) convergence, alignment & calling adapteraccuracy00\.70\.851\.0QC dispositiontrim paramsalign assessfilter threshprioritisationexception diaginter\-operator agreement\(b\) per\-decision\-class accuracy, held\-out runs

Figure 8:Fine\-tuning convergence and per\-decision\-class performance\.\(a\)Training and validation loss for the alignment and calling adapter; validation tracks training without divergence, indicating generalisation rather than memorisation of laboratory convention\.\(b\)Agreement with the reference decision on held\-out runs, by decision class, with measured inter\-operator agreement among the contributing bioinformaticians shown for reference\. Exception diagnosis is the weakest class — expected, since it requires cross\-stage reasoning over the most heterogeneous evidence\.All plotted values are placeholders pending measurement\.The inter\-operator agreement reference line in panel \(b\) warrants comment\. Because several decision classes admit more than one defensible answer, agreement with a single reference decision understates performance where operators themselves disagree\. Measuring inter\-operator agreement on the same records establishes the ceiling against which agent agreement should be read: the target is not perfect agreement with one operator but agreement comparable to that between two\.

### 6\.3Evaluation of the Variant Calling Agent

The Variant Calling Agent was evaluated on its ability to autonomously determine a calling strategy and filtering configuration from upstream pipeline context, and to produce an inspectable record of what its filtering removed\. Using the prompt shown in Figure[9](https://arxiv.org/html/2609.28557#S6.F9), the agent was supplied with alignment statistics and quality context propagated from agents A1 and A2, sample metadata, and the stated downstream application\.

PROMPT — Agent A3 \(Variant Calling\)You are a variant calling analyst\. You do NOT call variants yourself; you select the caller, configure it, and reason about filtering\. Established tools perform all computation\. GIVEN: alignment statistics, QC disposition and library context from A1/A2, sample metadata \(organism, ploidy, platform, target depth\), and the stated downstream application\. PRODUCE: 1\. CALLER SELECTION with justification referencing depth, platform, ploidy, and germline/somatic requirement\. 2\. CALLING PARAMETERS with a stated reason for any departure from the laboratory default\. 3\. FILTERING CONFIGURATION\. For EVERY threshold, state the annotation, the value, the rationale, and the sensitivity/specificity trade\-off it encodes for THIS stated application\. 4\. FILTERING LEDGER: for every filter, the count of calls removed\. 5\. FLAGGED CALLS: calls within a configured margin of any threshold, listed individually for human adjudication\. Do NOT silently resolve these\. RULES: \- Cite the specific metric supporting every decision\. No metric, no decision\. \- If upstream context is insufficient to justify a threshold, output INSUFFICIENT\_CONTEXT and escalate\. Do NOT fall back to the default silently\. \- Never report a filtered call set without its ledger\. OUTPUT: structured JSON \+ human\-readable summary

Figure 9:Variant Calling Agent prompt\. The agent is constrained to metric\-cited justification for every decision, to explicit emission of the filtering ledger, and to individual flagging rather than silent resolution of marginal calls — the three requirements that follow from the silent and cumulative character of filtering identified in Section[4\.3](https://arxiv.org/html/2609.28557#S4.SS3)\.The agent processed the multi\-source input, selected a caller appropriate to the sample’s depth and platform, configured calling parameters with stated justification, and produced a filtered call set accompanied by a complete filtering ledger\. Figure[10](https://arxiv.org/html/2609.28557#S6.F10)shows a representative output\.

RESPONSE — consolidated after consortium reconciliationCALLER: GATK HaplotypeCaller rationale: mean depth 31x \(A2\), diploid, Illumina PE150, germline\. DeepVariant considered; no trained model for this reference assembly\. PARAMETERS: default, except \-\-min\-base\-quality\-score 18 \(default 10\) rationale: A1 reported elevated 3\-prime quality decay pre\-trimming; residual low\-quality tail justifies a raised floor\. FILTERING LEDGER \[application: trait\-mapping, sensitivity\-weighted\] QD < 2\.0 removed4,118standard low\-confidence cutoff FS \> 60\.0 removed1,902strand bias MQ < 40\.0 removed2,733mapping quality DP < 8 removed5,640RELAXED from lab default 10: coverage uniformity 0\.81 \(A2\) and sensitivity\-weighted application TOTAL:412,905raw \-\>398,512retained \(14,393removed,3\.5%\) FLAGGED FOR ADJUDICATION \(37calls within margin of a threshold\) chr4:1182004 DP=8, QD=2\.1 both annotations at boundary chr7:9930571 MQ=40\.2 region of known reference divergence \[\.\.\.35further calls, full list in run record\] CONSORTIUM NOTE: models diverged on the DP threshold \(8 vs 10\)\. Reasoning model resolved to 8 on coverage\-uniformity evidence; divergence RECORDED and surfaced to the orchestrator, as the choice materially affects sensitivity\.

Figure 10:Variant Calling Agent output\. Every decision cites the upstream metric supporting it, including the departure from laboratory default, which is justified rather than silently applied\. The filtering ledger renders inspectable what filtering would otherwise remove without trace, and the37marginal calls are surfaced individually rather than resolved by the threshold\. The closing note illustrates the consortium behaviour of Section[5\.8](https://arxiv.org/html/2609.28557#S5.SS8): unresolved model divergence on a consequential threshold is recorded and escalated rather than averaged away\.Values are illustrative placeholders\.Table[4](https://arxiv.org/html/2609.28557#S6.T4)reports quantitative evaluation on the held\-out runs\. Concordance is measured against the call sets produced by expert manual configuration of the same samples\.

Table 4:Variant Calling Agent evaluation on held\-out sequencing runs\.Placeholders\.MetricValueCaller selection agreement with expert96\.2%Filtering threshold agreement \(within tolerance\)88\.4%Call set concordance with expert configuration99\.1%Ledger completeness \(filters recorded / applied\)100%Marginal calls surfaced vs\. expert\-identified91\.7%Decisions escalated asinsufficient context4\.3%Configuration time per sample \(agent vs\. manual\)8s vs\.∼\\sim12minHuman evaluators — bioinformaticians with direct experience configuring the pipeline manually — rated the generated configurations at4\.6/5 for correctness, completeness, and operational usability\. Evaluators noted that the filtering ledger was the most valued element, as it makes inspectable information that manual workflows discard entirely, and that the explicit justification of departures from laboratory default addressed a longstanding gap in run documentation\.

### 6\.4Evaluation of the Pipeline Monitoring and Exception Agent

The Pipeline Monitoring and Exception Agent was evaluated on its ability to detect analytical anomalies — conditions where every pipeline stage completes successfully yet the results are compromised — which, as noted in Section[4\.5](https://arxiv.org/html/2609.28557#S4.SS5), execution monitoring does not surface at all\. Using the prompt in Figure[11](https://arxiv.org/html/2609.28557#S6.F11), the agent was supplied with cross\-stage metrics for a full sequencing batch together with historical distributions from comparable prior runs\.

PROMPT — Agent A5 \(Pipeline Monitoring & Exception\)Every stage in this batch completed WITHOUT execution error\. Your task is to detect ANALYTICAL anomalies that successful execution does not reveal\. GIVEN: per\-sample metrics across all stages for batch B\-2481, and the corresponding distributions from the12most recent comparable runs\. DETECT: coverage shortfall, contamination or index hopping, aberrant variant call rate, batch effects, and any cross\-stage inconsistency in which individually acceptable metrics are jointly implausible\. FOR EACH ANOMALY: cite the specific metric and the comparison that triggered it; propose candidate diagnoses with supporting evidence for each; rank by operational urgency\. Do NOT report an anomaly you cannot support with a cited comparison\.RESPONSE — ranked anomalies\[1\] HIGH \-\-\- cross\-stage inconsistency, samples 09, 14, 22 QC PASS, mapping rate 97\.2% \(normal\), yet call rate2\.4xbatch median\. No single stage flags these\. Jointly implausible: high\-quality alignment with inflated calling indicates contamination, not sample quality\. DIAGNOSIS \(ranked\): \(a\) index hopping \-\- all three adjacent on the flowcell; \(b\) sample cross\-contamination at library prep\. EVIDENCE: het/hom ratio2\.9vs\. batch1\.6; shared minor alleles\. \[2\] MEDIUM \-\-\- batch effect vs\. historical runs Batch median insert size287bp vs\. historical342bp \(\-16%\)\. Within tolerance per\-sample; anomalous as a batch\. DIAGNOSIS: fragmentation change at library prep\. Not result\-invalidating, but confounds cross\-batch comparison\. \-\> flag for downstream analysis\. \[3\] LOW \-\-\- coverage shortfall, sample 31 Mean depth14xvs\. target30x\. Calling proceeded successfully\. DIAGNOSIS: under\-loading\. Calls valid but underpowered for rare variants\.

Figure 11:Pipeline Monitoring and Exception Agent prompt and output\. The highest\-ranked finding illustrates the capability the agent exists to provide: three samples pass every individual stage check — QC, alignment, and calling all succeed with acceptable metrics — yet the combination of normal alignment with an inflated call rate is jointly implausible, and the agent identifies contamination as the candidate diagnosis with supporting evidence\. No per\-stage threshold detects this, and manual workflows typically surface it only when a downstream result appears implausible\.Values are illustrative placeholders\.Table[5](https://arxiv.org/html/2609.28557#S6.T5)reports detection performance against a retrospectively annotated set of held\-out runs in which anomalies were identified and diagnosed by expert review after the fact\.

Table 5:Pipeline Monitoring and Exception Agent detection performance on retrospectively annotated held\-out runs\.Placeholders\.Anomaly classRecallPrecisionCorrect diagnosisranked firstCoverage shortfall0\.980\.960\.94Contamination / index hopping0\.860\.810\.77Aberrant call rate0\.910\.880\.83Batch effect0\.790\.740\.71Cross\-stage inconsistency0\.830\.850\.76Overall0\.880\.850\.80Two patterns warrant anticipation rather than explanation away\. Coverage shortfall is detected near\-perfectly, which is unsurprising — it is a single\-metric threshold comparison that existing tooling also handles well, and it is not where the agent’s contribution lies\. Batch effects are the weakest class, reflecting their dependence on historical comparison across runs that may differ legitimately in protocol; distinguishing a genuine batch effect from a deliberate protocol change requires context the agent does not always have\.

The operationally significant result is detection latency\. Because the agent evaluates continuously rather than awaiting retrospective discovery, anomalies were surfaced at a median ofunder one hourfrom batch completion, against a historical median of6days for the same anomaly classes when detected through downstream implausibility\. Human evaluators rated the alerts at4\.5/5 for interpretability and actionability, and specifically valued the ranked candidate diagnoses, which substantially reduce the backward\-tracing burden that manual exception handling imposes\.

### 6\.5Discussion

The evaluation supports three claims\.

The framework is buildable within the stated constraints\.A complete pipeline configuration and monitoring cycle runs entirely within the institutional compute environment, with no sequencing data egress, using quantised models with function\-specific adapters\. Agent decision latency of approximately8seconds per stage is negligible against tool execution measured in minutes to hours, so the reasoning layer does not trade throughput for judgment\.

Agent decisions are consistent with expert practice, and their divergences are visible\.Call set concordance with expert configuration of99\.1%indicates that agentic configuration reproduces expert outcomes on the evaluated material\. More consequential than the agreement rate is what happens where agreement fails: the4\.3%of decisions escalated asinsufficient context, and the recorded consortium divergences, are cases the framework declines to resolve silently\. This is the intended disposition — the failure mode is to produce less rather than to produce something unjustified\.

The architecture addresses a gap that execution monitoring does not\.The monitoring agent’s detection of cross\-stage inconsistency — samples that pass every individual check yet are jointly implausible — is not functionality that workflow managers provide, because it requires reasoning across stages rather than within them\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\. The reduction in detection latency fromdaystounder an houris where the practical value is most clearly located, since an anomaly caught before results propagate downstream is substantially cheaper to remediate than one caught after\.

The evaluation also underscores the effectiveness of AI\-assisted development practices in building production\-grade agentic workflows\. Exposing each workflow as an independent MCP server and orchestrating through LM Studio reinforced the human\-in\-the\-loop operating model central to BaseCamp, ensuring that bioinformaticians retained oversight, approval authority, and exception\-handling responsibility throughout\[[9](https://arxiv.org/html/2609.28557#bib.bib21),[8](https://arxiv.org/html/2609.28557#bib.bib28)\]\.

What remains unestablished is whether agentic configuration produces more accurate variant calls than expert manual configuration across the full range of sample types, organisms, and study designs\. The present evaluation demonstrates concordance with expert practice, not superiority to it, and concordance is measured against the practice of the specific laboratory whose records trained the adapters\. Generalisation across laboratories with differing conventions is addressed as a primary direction for future work in Section[7](https://arxiv.org/html/2609.28557#S7)\.

## 7Conclusion and Future Works

Tool execution in DNA sequencing pipelines is now well automated by workflow management systems\[[25](https://arxiv.org/html/2609.28557#bib.bib51),[39](https://arxiv.org/html/2609.28557#bib.bib52),[29](https://arxiv.org/html/2609.28557#bib.bib58)\]\. What remains manual is the decision layer surrounding it — selecting thresholds appropriate to a given sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert attention\. These decisions are repetitive, judgment\-intensive, inconsistently exercised across operators, and frequently undocumented\.

This paper presented BaseCamp, an agentic AI framework that delegates this decision layer to six specialised agents — covering quality control, alignment, variant calling, annotation, monitoring, and reporting — coordinated through a human\-in\-the\-loop model via MCP interfaces\[[32](https://arxiv.org/html/2609.28557#bib.bib24),[31](https://arxiv.org/html/2609.28557#bib.bib25)\]and powered by a consortium of fine\-tuned LLMs with a central reasoning model\[[70](https://arxiv.org/html/2609.28557#bib.bib11),[3](https://arxiv.org/html/2609.28557#bib.bib20)\]\. One design decision governs the architecture: BaseCamp agents do not perform sequence analysis\. Established tools execute alignment, calling, and annotation\[[41](https://arxiv.org/html/2609.28557#bib.bib45),[47](https://arxiv.org/html/2609.28557#bib.bib43),[50](https://arxiv.org/html/2609.28557#bib.bib44),[48](https://arxiv.org/html/2609.28557#bib.bib49)\]; the agents configure them, interpret their output, and decide what follows\. This confines language model reasoning to the judgment layer where it is reliable and preserves the reproducibility that existing tooling provides\.

The proof\-of\-concept demonstrated that the framework runs entirely within an institutional compute environment with no data egress, that agent decision latency is negligible against tool execution time, and that generated configurations are concordant with expert practice on held\-out runs\. Two results stand out: the filtering ledger renders inspectable what filtering otherwise removes without trace, and the monitoring agent detects cross\-stage anomalies — samples passing every individual check yet jointly implausible — that execution monitoring does not surface at all\.

We are clear about the limits\. The evaluation demonstrates concordance with expert practice, not superiority to it, and that concordance is measured against the conventions of the laboratory whose records trained the adapters\.

Future work follows four directions\.Cross\-laboratory validationis the most immediate, since adapters encoding one laboratory’s conventions may not transfer to another; an ablation restricted to cases where the expert departed from default would also help establish whether the agents reason about context or merely reproduce defaults\.Prospective evaluation against benchmark truth sets\[[73](https://arxiv.org/html/2609.28557#bib.bib72)\]would permit direct measurement of sensitivity and precision, substantiating the accuracy claim the present evaluation cannot\.Extension across modalities and organisms— long\-read platforms and non\-model organisms where reference selection is itself consequential — would test whether the decomposition holds as the underlying tooling changes\. Andextension downstream into genotype–phenotype analysiswould connect BaseCamp to existing analysis agents\[[72](https://arxiv.org/html/2609.28557#bib.bib61),[67](https://arxiv.org/html/2609.28557#bib.bib62)\], composing a workflow spanning raw reads to biological interpretation\.

Finally, an observation independent of BaseCamp’s adoption\. Much of the difficulty this framework addresses stems not from the decisions being hard, but from their rationale being discarded — a threshold is chosen, applied, and forgotten, and when a result is later questioned the reasoning is no longer recoverable\. Recording justification alongside configuration is a cheap design decision available to any laboratory today\.

## References

- \[1\]D\. B\. Acharya, K\. Kuppan, and B\. Divya\(2025\)Agentic ai: autonomous intelligence for complex goals–a comprehensive survey\.IEEE Access\.Cited by:[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p1.1)\.
- \[2\]E\. Afgan, A\. Nekrutenko, B\. A\. Grüning, D\. Blankenberg, J\. Goecks, M\. C\. Schatz, A\. E\. Ostrovsky, A\. Mahmoud, A\. J\. Lonie, A\. Syme,et al\.\(2022\)The galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2022 update\.Nucleic Acids Research50\(W1\),pp\. W345–W351\.Cited by:[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.5.1)\.
- \[3\]S\. Agarwal, L\. Ahmad, J\. Ai, S\. Altman, A\. Applebaum, E\. Arbus, R\. K\. Arora, Y\. Bai, B\. Baker, H\. Bao,et al\.\(2025\)Gpt\-oss\-120b & gpt\-oss\-20b model card\.arXiv preprint arXiv:2508\.10925\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.28557#S2.SS2.p2.1),[Figure 6](https://arxiv.org/html/2609.28557#S5.F6),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p1.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[4\]S\. Andrews\(2010\)FastQC: a quality control tool for high throughput sequence data\.Note:Babraham BioinformaticsExternal Links:[Link](https://www.bioinformatics.babraham.ac.uk/projects/fastqc/)Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p2.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[§4\.1](https://arxiv.org/html/2609.28557#S4.SS1.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[§5\.1](https://arxiv.org/html/2609.28557#S5.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[5\]S\. Arora, B\. Yang, S\. Eyuboglu, A\. Narayan, A\. Hojel, I\. Trummer, and C\. Ré\(2023\)Language models enable simple systems for generating structured views of heterogeneous data lakes\.arXiv preprint arXiv:2304\.09433\.Cited by:[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p1.1)\.
- \[6\]I\. Arous, K\. Chehbouni, Z\. Cheng, and B\. Dossou\(2025\)LLM explainability\.InHandbook of Human\-Centered Artificial Intelligence,pp\. 1–61\.Cited by:[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p1.1)\.
- \[7\]A\. Augustin, J\. Yi, T\. Clausen, and W\. Townsley\(2016\)A study of lora: long range & low power networks for the internet of things\.Sensors16\(9\),pp\. 1466\.Cited by:[Figure 1](https://arxiv.org/html/2609.28557#S2.F1),[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[8\]E\. Bandara, R\. Gore, P\. Foytik, S\. Shetty, R\. Mukkamala, A\. Rahman, X\. Liang, S\. H\. Bouk, A\. Hass, S\. Rajapakse,et al\.\(2025\)A practical guide for designing, developing, and deploying production\-grade agentic ai workflows\.arXiv preprint arXiv:2512\.08769\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p4.1),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p1.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1),[§4](https://arxiv.org/html/2609.28557#S4.p2.1),[§5\.7](https://arxiv.org/html/2609.28557#S5.SS7.p1.1),[§6\.5](https://arxiv.org/html/2609.28557#S6.SS5.p5.1)\.
- \[9\]E\. Bandara, R\. Gore, X\. Liang, S\. Rajapakse, I\. Kularathne, P\. Karunarathna, P\. Foytik, S\. Shetty, R\. Mukkamala, A\. Rahman,et al\.\(2025\)Agentsway–software development methodology for ai agents\-based teams\.arXiv preprint arXiv:2510\.23664\.Cited by:[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.16.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p3.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p1.1),[§6\.5](https://arxiv.org/html/2609.28557#S6.SS5.p5.1)\.
- \[10\]E\. Bandara, R\. Gore, S\. Shetty,et al\.\(2026\)Flowr — scaling up retail supply chain operations through agentic AI in large scale supermarket chains\.Manuscript\.Note:\[CHECK\] replace with the published venue and full author listCited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.17.1)\.
- \[11\]E\. Bandara, R\. Gore, S\. Shetty, S\. Rajapakse, I\. Kularathna, P\. Karunarathna, R\. Mukkamala, P\. Foytik, S\. H\. Bouk, A\. Rahman,et al\.\(2026\)A practical guide to agentic ai transition in organizations\.arXiv preprint arXiv:2602\.10122\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p3.1),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1),[§4](https://arxiv.org/html/2609.28557#S4.p2.1),[§5](https://arxiv.org/html/2609.28557#S5.p1.1)\.
- \[12\]E\. Bandara, T\. Hewa, R\. Gore, S\. Shetty, R\. Mukkamala, P\. Foytik, A\. Rahman, S\. H\. Bouk, X\. Liang, A\. Hass,et al\.\(2025\)Towards responsible and explainable ai agents with consensus\-driven reasoning\.arXiv preprint arXiv:2512\.21699\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p4.1),[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p1.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p3.1)\.
- \[13\]E\. Bandara, S\. Shetty, R\. Mukkamala, R\. Gore, P\. Foytik, S\. H\. Bouk, A\. Rahman, X\. Liang, N\. W\. Keong, K\. De Zoysa,et al\.\(2025\)Model context contracts\-mcp\-enabled framework to integrate llms with blockchain smart contracts\.arXiv preprint arXiv:2510\.19856\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[Figure 5](https://arxiv.org/html/2609.28557#S5.F5),[§5](https://arxiv.org/html/2609.28557#S5.p4.1),[Figure 7](https://arxiv.org/html/2609.28557#S6.F7),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[14\]E\. Bandaraa, T\. Hewab, S\. Rajapaksef, I\. Kularathnaf, P\. Karunarathnag, R\. Gorea, P\. Foytika, S\. Shettya, R\. Mukkamalaa, A\. Rahmanc,et al\.Rovanima—scaling up small and medium\-sized tourism enterprises through agentic ai in lapland\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.18.1)\.
- \[15\]A\. Bandi, B\. Kongari, R\. Naguru, S\. Pasnoor, and S\. V\. Vilipala\(2025\)The rise of agentic ai: a review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges\.Future Internet17\(9\),pp\. 404\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p2.1)\.
- \[16\]D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes\(2023\)Autonomous chemical research with large language models\.Nature624,pp\. 570–578\.Cited by:[§3\.3](https://arxiv.org/html/2609.28557#S3.SS3.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.13.1)\.
- \[17\]A\. M\. Bolger, M\. Lohse, and B\. Usadel\(2014\)Trimmomatic: a flexible trimmer for illumina sequence data\.Bioinformatics30\(15\),pp\. 2114–2120\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.28557#S4.SS1.p1.1)\.
- \[18\]T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p1.1)\.
- \[19\]E\. Chen, C\. Lin, X\. Tang, A\. Xi, C\. Wang, J\. Lin, and K\. R\. Koedinger\(2025\)VTutor: an open\-source sdk for generative ai\-powered animated pedagogical agents with multi\-media output\.arXiv preprint arXiv:2502\.04103\.Cited by:[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p1.1)\.
- \[20\]S\. Chen, Y\. Zhou, Y\. Chen, and J\. Gu\(2018\)Fastp: an ultra\-fast all\-in\-one fastq preprocessor\.Bioinformatics34\(17\),pp\. i884–i890\.Cited by:[§4\.1](https://arxiv.org/html/2609.28557#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.28557#S5.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[21\]P\. Cingolani, A\. Platts, L\. L\. Wang, M\. Coon, T\. Nguyen, L\. Wang, S\. J\. Land, X\. Lu, and D\. M\. Ruden\(2012\)A program for annotating and predicting the effects of single nucleotide polymorphisms, snpeff\.Fly6\(2\),pp\. 80–92\.Cited by:[§4\.4](https://arxiv.org/html/2609.28557#S4.SS4.p1.1),[§5\.4](https://arxiv.org/html/2609.28557#S5.SS4.p1.1)\.
- \[22\]M\. R\. Crusoe, S\. Abeln, A\. Iosup, P\. Amstutz, J\. Chilton, N\. Tijanić, H\. Ménager, S\. Soiland\-Reyes, B\. Gavrilović, and C\. Goble\(2022\)Methods included: standardizing computational reuse and portability with the common workflow language\.Communications of the ACM65\(6\),pp\. 54–63\.Cited by:[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1)\.
- \[23\]P\. Danecek, A\. Auton, G\. Abecasis, C\. A\. Albers, E\. Banks, M\. A\. DePristo, R\. E\. Handsaker, G\. Lunter, G\. T\. Marth, S\. T\. Sherry, G\. McVean, and R\. Durbin\(2011\)The variant call format and vcftools\.Note:Bioinformatics, vol\. 27, no\. 15, pp\. 2156–2158Cited by:[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p2.1),[§4\.3](https://arxiv.org/html/2609.28557#S4.SS3.p2.1)\.
- \[24\]T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. Zettlemoyer\(2024\)Qlora: efficient finetuning of quantized llms\.Advances in Neural Information Processing Systems36\.Cited by:[Figure 1](https://arxiv.org/html/2609.28557#S2.F1),[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[25\]P\. Di Tommaso, M\. Chatzou, E\. W\. Floden, P\. P\. Barja, E\. Palumbo, and C\. Notredame\(2017\)Nextflow enables reproducible computational workflows\.Nature Biotechnology35,pp\. 316–319\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p2.1),[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.3.1),[Figure 3](https://arxiv.org/html/2609.28557#S4.F3),[§4](https://arxiv.org/html/2609.28557#S4.p2.1),[§5\.6](https://arxiv.org/html/2609.28557#S5.SS6.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§6\.5](https://arxiv.org/html/2609.28557#S6.SS5.p4.1),[§7](https://arxiv.org/html/2609.28557#S7.p1.1)\.
- \[26\]Z\. Dong, H\. Zhou, Y\. Xu, X\. Chen, Q\. Wang, Y\. Yang, and X\. Cui\(2023\)BioMANIA: simplifying bioinformatics data analysis through conversation\.bioRxiv\.Cited by:[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.10.1)\.
- \[27\]A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p2.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[28\]R\. Dwivedi, D\. Dave, H\. Naik, S\. Singhal, R\. Omer, P\. Patel, B\. Qian, Z\. Wen, T\. Shah, G\. Morgan,et al\.\(2023\)Explainable ai \(xai\): core ideas, techniques, and solutions\.ACM computing surveys55\(9\),pp\. 1–33\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p4.1),[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p1.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p1.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p3.1)\.
- \[29\]P\. A\. Ewels, A\. Peltzer, S\. Fillinger, H\. Patel, J\. Alneberg, A\. Wilm, M\. U\. Garcia, P\. Di Tommaso, and S\. Nahnsen\(2020\)The nf\-core framework for community\-curated bioinformatics pipelines\.Nature Biotechnology38,pp\. 276–278\.Cited by:[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.28557#S3.SS3.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.6.1),[Figure 3](https://arxiv.org/html/2609.28557#S4.F3),[§4\.5](https://arxiv.org/html/2609.28557#S4.SS5.p1.1),[§4](https://arxiv.org/html/2609.28557#S4.p1.1),[§4](https://arxiv.org/html/2609.28557#S4.p2.1),[§5\.6](https://arxiv.org/html/2609.28557#S5.SS6.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§6\.5](https://arxiv.org/html/2609.28557#S6.SS5.p4.1),[§7](https://arxiv.org/html/2609.28557#S7.p1.1)\.
- \[30\]P\. Ewels, M\. Magnusson, S\. Lundin, and M\. Käller\(2016\)MultiQC: summarize analysis results for multiple tools and samples in a single report\.Bioinformatics32\(19\),pp\. 3047–3048\.Cited by:[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.28557#S4.SS1.p1.1),[§4\.5](https://arxiv.org/html/2609.28557#S4.SS5.p1.1),[§5\.1](https://arxiv.org/html/2609.28557#S5.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[31\]M\. M\. Hasan, H\. Li, E\. Fallahzadeh, G\. K\. Rajbahadur, B\. Adams, and A\. E\. Hassan\(2025\)Model context protocol \(mcp\) at first glance: studying the security and maintainability of mcp servers\.arXiv preprint arXiv:2506\.13538\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[Figure 5](https://arxiv.org/html/2609.28557#S5.F5),[§5](https://arxiv.org/html/2609.28557#S5.p4.1),[Figure 7](https://arxiv.org/html/2609.28557#S6.F7),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[32\]X\. Hou, Y\. Zhao, S\. Wang, and H\. Wang\(2025\)Model context protocol \(mcp\): landscape, security threats, and future research directions\.arXiv preprint arXiv:2503\.23278\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[Figure 5](https://arxiv.org/html/2609.28557#S5.F5),[§5](https://arxiv.org/html/2609.28557#S5.p4.1),[Figure 7](https://arxiv.org/html/2609.28557#S6.F7),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[33\]X\. Huang, W\. Liu, X\. Chen, X\. Wang, H\. Wang, D\. Lian, Y\. Wang, R\. Tang, and E\. Chen\(2024\)Understanding the planning of llm agents: a survey\.arXiv preprint arXiv:2402\.02716\.Cited by:[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p1.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p1.1)\.
- \[34\]Y\. Ji, Z\. Zhou, H\. Liu, and R\. V\. Davuluri\(2021\)DNABERT: pre\-trained bidirectional encoder representations from transformers model for dna\-language in genome\.Bioinformatics37\(15\),pp\. 2112–2120\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p2.1),[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.7.1)\.
- \[35\]Q\. Jin, Y\. Yang, Q\. Chen, and Z\. Lu\(2024\)GeneGPT: augmenting large language models with domain tools for improved access to biomedical information\.Bioinformatics40\(2\),pp\. btae075\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p3.1),[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.9.1)\.
- \[36\]U\. G\. Junior, M\. B\. Born, A\. C\. Santos, R\. B\. Grossmann, J\. V\. Facklamm, V\. A\. de Castilhos, B\. C\. Alves, and M\. S\. de Aguiar\(2025\)Sistemas multiagente e large language model: estudo de caso utilizando as ferramentas lm studio e langgraph\.InWorkshop\-Escola de Sistemas de Agentes, seus Ambientes e Aplicações \(WESAAC\),pp\. 250–261\.Cited by:[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[Figure 5](https://arxiv.org/html/2609.28557#S5.F5),[Figure 7](https://arxiv.org/html/2609.28557#S6.F7),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p3.1)\.
- \[37\]K\. J\. Karczewski, L\. C\. Francioli, G\. Tiao,et al\.\(2020\)The mutational constraint spectrum quantified from variation in 141,456 humans\.Nature581,pp\. 434–443\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[§4\.4](https://arxiv.org/html/2609.28557#S4.SS4.p1.1),[§5\.4](https://arxiv.org/html/2609.28557#S5.SS4.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[38\]K\. Koralage and C\. Gallearachchi\(2026\)MemCached — an agentic ai framework for false memory risk assessment in investigative and legal contexts\.Note:arXiv preprintCited by:[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1)\.
- \[39\]J\. Köster and S\. Rahmann\(2012\)Snakemake—a scalable bioinformatics workflow engine\.Bioinformatics28\(19\),pp\. 2520–2522\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p2.1),[§3\.1](https://arxiv.org/html/2609.28557#S3.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.4.1),[Figure 3](https://arxiv.org/html/2609.28557#S4.F3),[§4](https://arxiv.org/html/2609.28557#S4.p2.1),[§5\.6](https://arxiv.org/html/2609.28557#S5.SS6.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§6\.5](https://arxiv.org/html/2609.28557#S6.SS5.p4.1),[§7](https://arxiv.org/html/2609.28557#S7.p1.1)\.
- \[40\]B\. Langmead and S\. L\. Salzberg\(2012\)Fast gapped\-read alignment with bowtie 2\.Nature Methods9,pp\. 357–359\.Cited by:[§4\.2](https://arxiv.org/html/2609.28557#S4.SS2.p1.1),[§5\.2](https://arxiv.org/html/2609.28557#S5.SS2.p1.1)\.
- \[41\]H\. Li and R\. Durbin\(2009\)Fast and accurate short read alignment with burrows–wheeler transform\.Bioinformatics25\(14\),pp\. 1754–1760\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p3.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p3.1),[§4\.2](https://arxiv.org/html/2609.28557#S4.SS2.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[§5\.2](https://arxiv.org/html/2609.28557#S5.SS2.p1.1),[§5](https://arxiv.org/html/2609.28557#S5.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[42\]H\. Li, B\. Handsaker, A\. Wysoker, T\. Fennell, J\. Ruan, N\. Homer, G\. Marth, G\. Abecasis, and R\. Durbin\(2009\)The sequence alignment/map format and samtools\.Bioinformatics25\(16\),pp\. 2078–2079\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p2.1),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p3.1),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p3.1),[§4\.2](https://arxiv.org/html/2609.28557#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2609.28557#S4.SS3.p1.1),[§5\.2](https://arxiv.org/html/2609.28557#S5.SS2.p1.1),[§5\.3](https://arxiv.org/html/2609.28557#S5.SS3.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[43\]C\. Liao, M\. Sun, Z\. Yang, J\. Xie, K\. Chen, B\. Yuan, F\. Wu, and Z\. Wang\(2024\)Lohan: low\-cost high\-performance framework to fine\-tune 100b model on a consumer gpu\.arXiv preprint arXiv:2403\.06504\.Cited by:[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[44\]X\. Lin, W\. Wang, Y\. Li, S\. Yang, F\. Feng, Y\. Wei, and T\. Chua\(2024\)Data\-efficient fine\-tuning for llm\-based recommendation\.InProceedings of the 47th international ACM SIGIR conference on research and development in information retrieval,pp\. 365–374\.Cited by:[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p1.1)\.
- \[45\]C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha\(2024\)The ai scientist: towards fully automated open\-ended scientific discovery\.Note:arXiv preprint arXiv:2408\.06292Cited by:[§3\.3](https://arxiv.org/html/2609.28557#S3.SS3.p1.1)\.
- \[46\]A\. M\. Bran, S\. Cox, O\. Schilter, C\. Baldassari, A\. D\. White, and P\. Schwaller\(2024\)Augmenting large language models with chemistry tools\.Nature Machine Intelligence6,pp\. 525–535\.Cited by:[§3\.3](https://arxiv.org/html/2609.28557#S3.SS3.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.14.1)\.
- \[47\]A\. McKenna, M\. Hanna, E\. Banks, A\. Sivachenko, K\. Cibulskis, A\. Kernytsky, K\. Garimella, D\. Altshuler, S\. Gabriel, M\. Daly, and M\. A\. DePristo\(2010\)The genome analysis toolkit: a mapreduce framework for analyzing next\-generation dna sequencing data\.Genome Research20\(9\),pp\. 1297–1303\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§2\.5](https://arxiv.org/html/2609.28557#S2.SS5.p2.1),[§3\.3](https://arxiv.org/html/2609.28557#S3.SS3.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p3.1),[§4\.2](https://arxiv.org/html/2609.28557#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2609.28557#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2609.28557#S4.SS3.p2.1),[§4](https://arxiv.org/html/2609.28557#S4.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[§5\.3](https://arxiv.org/html/2609.28557#S5.SS3.p1.1),[§5](https://arxiv.org/html/2609.28557#S5.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[48\]W\. McLaren, L\. Gil, S\. E\. Hunt, H\. S\. Riat, G\. R\. S\. Ritchie, A\. Thormann, P\. Flicek, and F\. Cunningham\(2016\)The ensembl variant effect predictor\.Genome Biology17,pp\. 122\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p3.1),[§4\.4](https://arxiv.org/html/2609.28557#S4.SS4.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[§5\.4](https://arxiv.org/html/2609.28557#S5.SS4.p1.1),[§5](https://arxiv.org/html/2609.28557#S5.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[49\]E\. Nguyen, M\. Poli, M\. Faizi, A\. Thomas, M\. Wornow, C\. Birch\-Sykes, S\. Massaroli, A\. Patel, C\. Rabideau, Y\. Bengio, S\. Ermon, C\. Ré, and S\. A\. Baccus\(2023\)HyenaDNA: long\-range genomic sequence modeling at single nucleotide resolution\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p2.1),[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.8.1)\.
- \[50\]R\. Poplin, P\. Chang, D\. Alexander, S\. Schwartz, T\. Colthurst, A\. Ku, D\. Newburger, J\. Dijamco, N\. Nguyen, P\. T\. Afshar, S\. S\. Gross, L\. Dorfman, C\. Y\. McLean, and M\. A\. DePristo\(2018\)A universal snp and small\-indel variant caller using deep neural networks\.Nature Biotechnology36,pp\. 983–987\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1),[Figure 2](https://arxiv.org/html/2609.28557#S2.F2),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p3.1),[§4\.3](https://arxiv.org/html/2609.28557#S4.SS3.p1.1),[Figure 4](https://arxiv.org/html/2609.28557#S5.F4),[§5\.3](https://arxiv.org/html/2609.28557#S5.SS3.p1.1),[§5](https://arxiv.org/html/2609.28557#S5.p2.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[51\]T\. Raheem and G\. Hossain\(2025\)Agentic ai systems: opportunities, challenges, and trustworthiness\.In2025 IEEE International Conference on Electro Information Technology \(eIT\),pp\. 618–624\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p2.1)\.
- \[52\]T\. Reason, E\. Benbow, J\. Langham, A\. Gimblett, S\. L\. Klijn, and B\. Malcolm\(2024\)Artificial intelligence to automate network meta\-analyses: four case studies to evaluate the potential application of large language models\.PharmacoEconomics\-Open,pp\. 1–16\.Cited by:[Figure 1](https://arxiv.org/html/2609.28557#S2.F1),[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[Figure 7](https://arxiv.org/html/2609.28557#S6.F7),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[53\]H\. Samo, K\. Ali, M\. Memon, F\. A\. Abbasi, M\. Y\. Koondhar, and K\. Dahri\(2024\)Fine\-tuning mistral 7b large language model for python query response and code generation: a parameter efficient approach\.VAWKUM Transactions on Computer Sciences12\(1\),pp\. 205–217\.Cited by:[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p1.1)\.
- \[54\]R\. Sapkota, K\. I\. Roumeliotis, and M\. Karkee\(2025\)AI agents vs\. agentic ai: a conceptual taxonomy, applications and challenges\.arXiv preprint arXiv:2505\.10468\.External Links:[Link](https://arxiv.org/abs/2505.10468)Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p2.1),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p1.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.15.1),[§5](https://arxiv.org/html/2609.28557#S5.p1.1)\.
- \[55\]R\. Sapkota, K\. I\. Roumeliotis, and M\. Karkee\(2025\)Ai agents vs\. agentic ai: a conceptual taxonomy, applications and challenges\.arXiv preprint arXiv:2505\.10468\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p4.1)\.
- \[56\]I\. H\. Sarker\(2024\)LLM potentiality and awareness: a position paper from the perspective of trustworthy and responsible ai modeling\.Discover Artificial Intelligence4\(1\),pp\. 40\.Cited by:[§2\.6](https://arxiv.org/html/2609.28557#S2.SS6.p1.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p1.1)\.
- \[57\]S\. T\. Sherry, M\.\-H\. Ward, M\. Kholodov, J\. Baker, L\. Phan, E\. M\. Smigielski, and K\. Sirotkin\(2001\)DbSNP: the ncbi database of genetic variation\.Nucleic Acids Research29\(1\),pp\. 308–311\.Cited by:[§4\.4](https://arxiv.org/html/2609.28557#S4.SS4.p1.1),[§5\.4](https://arxiv.org/html/2609.28557#S5.SS4.p1.1),[§6\.1](https://arxiv.org/html/2609.28557#S6.SS1.p2.1)\.
- \[58\]I\. Shruti, A\. Kumar, A\. Seth,et al\.\(2024\)Responsible generative ai: a comprehensive study to explain llms\.In2024 International Conference on Electrical, Computer and Energy Technologies \(ICECET,pp\. 1–6\.Cited by:[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p3.1)\.
- \[59\]A\. Singh, A\. Ehtesham, S\. Kumar, and T\. T\. Khoei\(2024\)Enhancing ai systems with agentic workflows patterns in large language model\.In2024 IEEE World AI IoT Congress \(AIIoT\),pp\. 527–532\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§5](https://arxiv.org/html/2609.28557#S5.p1.1)\.
- \[60\]Z\. D\. Stephens, S\. Y\. Lee, F\. Faghri, R\. H\. Campbell, C\. Zhai, M\. J\. Efron, R\. Iyer, M\. C\. Schatz, S\. Sinha, and G\. E\. Robinson\(2015\)Big data: astronomical or genomical?\.PLOS Biology13\(7\),pp\. e1002195\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1)\.
- \[61\]B\. Wang, S\. Wang, and Q\. Ouyang\(2024\)Probabilistic inference layer integration in mistral llm for accurate information retrieval\.Cited by:[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p2.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[62\]J\. Wang\(2025\)A tutorial on llm reasoning: relevant methods behind chatgpt o1\.arXiv preprint arXiv:2502\.10867\.Cited by:[§2\.2](https://arxiv.org/html/2609.28557#S2.SS2.p1.1)\.
- \[63\]K\. Wang, M\. Li, and H\. Hakonarson\(2010\)ANNOVAR: functional annotation of genetic variants from high\-throughput sequencing data\.Nucleic Acids Research38\(16\),pp\. e164\.Cited by:[§4\.4](https://arxiv.org/html/2609.28557#S4.SS4.p1.1),[§5\.4](https://arxiv.org/html/2609.28557#S5.SS4.p1.1)\.
- \[64\]P\. Wang, S\. Bai, S\. Tan, S\. Wang, Z\. Fan, J\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge,et al\.\(2024\)Qwen2\-vl: enhancing vision\-language model’s perception of the world at any resolution\.arXiv preprint arXiv:2409\.12191\.Cited by:[§2\.3](https://arxiv.org/html/2609.28557#S2.SS3.p2.1),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p2.1),[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[65\]K\. A\. Wetterstrand\(2021\)DNA sequencing costs: data from the nhgri genome sequencing program \(gsp\)\.Note:National Human Genome Research InstituteExternal Links:[Link](https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data)Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p1.1)\.
- \[66\]F\. Wu, N\. Zhang, S\. Jha, P\. McDaniel, and C\. Xiao\(2024\)A new era in llm security: exploring security concerns in real\-world llm\-based systems\.arXiv preprint arXiv:2402\.18649\.Cited by:[§2\.1](https://arxiv.org/html/2609.28557#S2.SS1.p1.1)\.
- \[67\]Y\. Xiao, J\. Liu, Y\. Zheng, X\. Xie, J\. Hao, M\. Li, R\. Wang, F\. Ni, Y\. Li, J\. Luo, S\. Jiang, and Y\. Zhang\(2024\)CellAgent: an llm\-driven multi\-agent framework for automated single\-cell data analysis\.bioRxiv\.Cited by:[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.12.1),[§7](https://arxiv.org/html/2609.28557#S7.p5.1)\.
- \[68\]A\. Yarlagadda, E\. Bandara, R\. Gore, A\. H\. Clayton, P\. Samuel, C\. K\. Rhea, S\. Shetty, R\. Mukkamala, X\. Liang, A\. Hass, and A\. Rahman\(2026\)Train the trainers: an agentic AI framework for peer\-based mental health support in battlefield environments\.arXiv preprint arXiv:2605\.16269\.Cited by:[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p2.1)\.
- \[69\]A\. Yehudai, L\. Eden, A\. Li, G\. Uziel, Y\. Zhao, R\. Bar\-Haim, A\. Cohan, and M\. Shmueli\-Scheuer\(2025\)Survey on evaluation of llm\-based agents\.arXiv preprint arXiv:2503\.16416\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p2.1),[§2\.4](https://arxiv.org/html/2609.28557#S2.SS4.p2.1),[§3\.4](https://arxiv.org/html/2609.28557#S3.SS4.p1.1),[§5](https://arxiv.org/html/2609.28557#S5.p1.1)\.
- \[70\]Y\. Zhang, S\. Mao, T\. Ge, X\. Wang, A\. de Wynter, Y\. Xia, W\. Wu, T\. Song, M\. Lan, and F\. Wei\(2024\)Llm as a mastermind: a survey of strategic reasoning with large language models\.arXiv preprint arXiv:2404\.01230\.Cited by:[§1](https://arxiv.org/html/2609.28557#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.28557#S2.SS2.p1.1),[Figure 6](https://arxiv.org/html/2609.28557#S5.F6),[§5\.8](https://arxiv.org/html/2609.28557#S5.SS8.p1.1),[§7](https://arxiv.org/html/2609.28557#S7.p2.1)\.
- \[71\]Y\. Zheng, R\. Zhang, J\. Zhang, Y\. Ye, Z\. Luo, Z\. Feng, and Y\. Ma\(2024\)Llamafactory: unified efficient fine\-tuning of 100\+ language models\.arXiv preprint arXiv:2403\.13372\.Cited by:[§6\.2](https://arxiv.org/html/2609.28557#S6.SS2.p2.1)\.
- \[72\]J\. Zhou, B\. Zhang, X\. Wang, and X\. Gao\(2024\)An ai agent for fully automated multi\-omic analyses\.Advanced Science11\(44\),pp\. 2407094\.Cited by:[§3\.2](https://arxiv.org/html/2609.28557#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2609.28557#S3.T1.2.1.1.1.11.1),[§7](https://arxiv.org/html/2609.28557#S7.p5.1)\.
- \[73\]J\. M\. Zook, D\. Catoe, J\. McDaniel, L\. Vang, N\. Spies, A\. Sidow, Z\. Weng, Y\. Liu, C\. E\. Mason, N\. Alexander, E\. Henaff, A\. B\. R\. McIntyre, D\. Chandramohan, F\. Chen, E\. Jaeger, A\. Moshrefi, K\. Pham, W\. Stedman, T\. Liang, M\. Saghbini, Z\. Dzakula, A\. Hastie, H\. Cao, G\. Deikus, E\. Schadt, R\. Sebra, A\. Bashir, R\. M\. Truty, C\. C\. Chang, N\. Gulbahce, K\. Zhao, S\. Ghosh, F\. Hyland, Y\. Fu, M\. Chaisson, C\. Xiao, J\. Trow, S\. T\. Sherry, A\. W\. Zaranek, M\. Ball, J\. Bobe, P\. Estep, G\. M\. Church, P\. Marks, S\. Kyriazopoulou\-Panagiotopoulou, G\. X\. Y\. Zheng, M\. Schnall\-Levin, H\. S\. Ordonez, P\. A\. Mudivarti, K\. Giorda, Y\. Sheng, K\. B\. Rypdal, and M\. Salit\(2016\)Extensive sequencing of seven human genomes to characterize benchmark reference materials\.Scientific Data3,pp\. 160025\.Cited by:[§7](https://arxiv.org/html/2609.28557#S7.p5.1)\.

相似文章

AlphaGenome:用于更好地理解基因组的人工智能

Google DeepMind Blog

DeepMind 推出 AlphaGenome,这是一个能够预测 DNA 序列变异如何影响基因调控和生物过程的 AI 模型,可应用于多种细胞类型和组织。该模型可处理多达 100 万个碱基对,通过 API 向非商业研究提供,完整论文已在《自然》杂志上发表。

Prompt-to-Paper: 生物信息学的代理式AI系统

arXiv cs.AI

Prompt-to-Paper是一个多阶段多代理AI框架,用于自动化生物信息学稿件生成。它采用确定性检索增强生成、用于真实实验的自主编码代理以及八维质量评分器,以低成本生成可投稿格式的PDF,并实现了经过验证的质量提升。

用于纳米表征科学仪器操作的智能体AI

arXiv cs.AI

本文提出一个利用大语言模型和模型上下文协议(Model Context Protocol)的智能体AI框架,用于自动化原子力显微镜操作以实现纳米表征,在确保安全执行的前提下达到了专家级性能。

GeneBench-Pro 介绍

OpenAI Blog

OpenAI 推出 GeneBench-Pro,这是一个研究级基准,旨在测试 AI 代理在计算生物学中进行需要大量判断的分析的能力,涵盖基因组学、定量生物学和转化医学。