PATHFinder Agent for Tailored Prenatal Care

arXiv cs.AI Papers

Summary

This paper presents PATHFinder Agent, an end-to-end conversational AI system that generates personalized prenatal care plans following ACOG's PATH guidelines, integrating patient intake, dynamic dialogue, plan synthesis, and clinician oversight. Evaluation of frontier LLMs, including GPT-5.2, shows promising but incomplete performance, highlighting gaps in antenatal testing recommendations.

arXiv:2607.24768v1 Announce Type: new Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211. The system features a four-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight. We evaluate frontier large language models (LLMs) on expert-curated rubrics across five clinical dimensions, finding that GPT-5.2 achieves the highest average score (77.6\%) while identifying key gaps in antenatal testing recommendations. We discuss future validation through human participant studies and randomized controlled trials.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:51 AM

# PATHFinder Agent for Tailored Prenatal Care
Source: [https://arxiv.org/html/2607.24768](https://arxiv.org/html/2607.24768)
,Carissa SamuelUniversity of MichiganAnn ArborUSA,Samia AbdelnabiUniversity of MichiganAnn ArborUSA,Alex PeahlUniversity of MichiganAnn ArborUSAandElizabeth Bondi\-KellyUniversity of MichiganAnn ArborUSA[ecbk@umich\.edu](https://arxiv.org/html/2607.24768v1/mailto:[email protected])

\(2026\)

###### Abstract\.

Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals\. The American College of Obstetricians and Gynecologists \(ACOG\) recently introduced guidelines advocating tailored prenatal care, called PATH \(Plan for Tailored Healthcare\)\. We presentPATHFinder Agent\(Planner for Appropriate Tailored Healthcare\), an end\-to\-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211\. The system features a four\-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight\. We evaluate frontier large language models \(LLMs\) on expert\-curated rubrics across five clinical dimensions, finding that GPT\-5\.2 achieves the highest average score \(77\.6%\) while identifying key gaps in antenatal testing recommendations\. We discuss future validation through human participant studies and randomized controlled trials\.

Large Language Models, Large Language Model Agents, Health, Reproductive Health, Maternal Health, Healthcare Agents, LLM\-in\-the\-loop

††copyright:acmlicensed††journalyear:2026††copyright:cc††conference:Interactive Health Conference; July 05–08, 2026; Porto, Portugal††booktitle:Interactive Health Conference \(IH ’26\), July 05–08, 2026, Porto, Portugal††doi:10\.1145/3786579\.3804996††isbn:979\-8\-4007\-2422\-0/2026/07††ccs:Computing methodologies Natural language processing††ccs:Human\-centered computing Natural language interfaces††ccs:Applied computing Consumer health## 1\.Introduction

Prenatal care is a crucial preventive service that improves pregnancy outcomes for mothers and their children\(Peahlet al\.,[2020c](https://arxiv.org/html/2607.24768#bib.bib13)\), with nearly four million pregnant patients each year in the United States receiving prenatal care\. Prenatal care is multifaceted, involving planning and providing medical care, screening tests, answering questions, and connecting people to appropriate social and community resources\.

The American College of Obstetricians and Gynecologists \(ACOG\), a professional association of physicians specializing in obstetrics and gynecology in the United States, recently proposed a transformation to existing guidelines, with the aim to begin“carefully tailoring prenatal care”\(Peahlet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib1)\)\. These guidelines, called PATH \(Plan for Appropriate Tailored Healthcare\), broadly tackle\(a\) addressing unmet social needsand\(b\) incorporating alternative care modalitiesto help tailor the plan to the patient and provide prenatal care to many more birthing people\. PATH was carefully designed after conducting interviews with 110 patients, clinicians, and policy makers representing 25 organizations and more than 75 clinics\.

Adopting PATH is a complex task for both the patients and clinicians \(Figure[1\(a\)](https://arxiv.org/html/2607.24768#S1.F1.sf1)\)\. For patients with unmet social needs, additional prenatal visits with a maternity care professional are unlikely to address underlying needs and may create additional burden\(Peahlet al\.,[2021b](https://arxiv.org/html/2607.24768#bib.bib7)\)\. Furthermore, increased demands on physicians and other health care professionals to address patients’ unmet social needs with insufficient resources have been associated with burnout\(Tabata\-Kellyet al\.,[2024](https://arxiv.org/html/2607.24768#bib.bib14)\)\.

We propose to help mitigate these challenges by creating a conversational AI agent\-based interface,PATHFinder Agent, with abilities to\(a\) follow upwith questions and clarifications requiring medical and conversational knowledge,\(b\)perform actions via tool\-calls\(searching for resources, generating a report, etc\.\), and\(c\)follow instructionsto provide a comprehensive plan\. We further carefully work towards avoiding unintended behaviors, ensuring oversight and grounding, and testing for deployment, as state\-of\-the\-art medical systems like g\-AMIE have called for\(Vedadiet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib20)\)\. We anticipatePATHFinder Agenthas the potential to support broad deployment of the PATH guidelines, thereby improving health and well\-being for pregnant and birthing individuals\.

![Refer to caption](https://arxiv.org/html/2607.24768v1/x1.png)\(a\)Current practices involve complex, disjoint conversations with patients to provide prenatal care\.
![Refer to caption](https://arxiv.org/html/2607.24768v1/x2.png)\(b\)Centralized, scalable conversation withPATHFinder Agentequipped with the latest guidelines and grounded resources\.

Figure 1\.Illustrations of current practices \(left\) and our proposed solution \(right\)
## 2\.Background

We identify three key areas from the literature that have particularly helped us shape the system\.

#### Prenatal Care

Over 80% of adverse pregnancy outcomes are preventable through essential prenatal care and the management of unmet social needs\(Trostet al\.,[2022](https://arxiv.org/html/2607.24768#bib.bib2)\)\. However, the traditional 12–14 visit in\-person model\(Kilpatricket al\.,[2017](https://arxiv.org/html/2607.24768#bib.bib3)\)is often inaccessible due to social drivers of health—such as housing instability and inflexible work schedules\(Semegaet al\.,[2021](https://arxiv.org/html/2607.24768#bib.bib4); National Academies of Sciences, Engineering, and Medicine,[2019](https://arxiv.org/html/2607.24768#bib.bib5)\)—leading many patients to be underserved and even feel unheard\(Belleroseet al\.,[2022](https://arxiv.org/html/2607.24768#bib.bib6); Betronet al\.,[2018](https://arxiv.org/html/2607.24768#bib.bib33); Mohamoudet al\.,[2023](https://arxiv.org/html/2607.24768#bib.bib34)\)\. In response, ACOG developed the Plan for Appropriate Tailored Healthcare in pregnancy \(PATH\)\(Peahlet al\.,[2021b](https://arxiv.org/html/2607.24768#bib.bib7)\), which utilizes shared decision\-making to tailor care through flexible visit frequencies\(Peahlet al\.,[2020a](https://arxiv.org/html/2607.24768#bib.bib35); Turrentine,[2023](https://arxiv.org/html/2607.24768#bib.bib36); Balket al\.,[2023a](https://arxiv.org/html/2607.24768#bib.bib37)\)and telehealth modalities\(Balket al\.,[2023b](https://arxiv.org/html/2607.24768#bib.bib38)\), for example\. Despite its potential, gaps remain in implementing PATH effectively within complex healthcare systems\(Nijagalet al\.,[2021](https://arxiv.org/html/2607.24768#bib.bib39)\), and taking patient preferences into account will be vital\(Peahlet al\.,[2021a](https://arxiv.org/html/2607.24768#bib.bib40),[2020b](https://arxiv.org/html/2607.24768#bib.bib41)\)\.

#### Medical Agents

Gemini, AMIE\(Saabet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib21); Palepuet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib22); Vedadiet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib20); Johriet al\.,[2024](https://arxiv.org/html/2607.24768#bib.bib28)\)and MedAgentBench\(Jianget al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib27)\)focus on electronic health records, and MedAgentGym focuses on programming\-centric agentic tasks in medicine\(Xuet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib19)\)\. These systems do not focus on integrating both domain\-specific needs, like prenatal care, and domain\-agnostic needs, like safety measures\.

#### Tool\-use and Conversational Systems

To utilize such medical agents to support patients and providers implementing PATH, we build on work where agents must ask clarifying questions and follow instructions to achieve the user’s goal, as well as work in which agents need to call external tools\. Research works like\(Qinet al\.,[2023](https://arxiv.org/html/2607.24768#bib.bib23); Patilet al\.,[2023](https://arxiv.org/html/2607.24768#bib.bib29)\)investigated LLMs’ capabilities to invoke specific “tools” to solve tasks in mathematical reasoning, program synthesis, and general tasks\. LLMs can decide to invoke a tool by generating tokens in an expected format, which enables them to navigate autonomously by recursively invoking tools\. Instruction\-tuned LLMs paired with strategies like ReACT \(Reasoning and Acting\)\(Yaoet al\.,[2022](https://arxiv.org/html/2607.24768#bib.bib30)\)have demonstrated near\-perfect capabilities to accurately invoke tools\. On the other hand, researchers have also studied task\-oriented conversations\(Budzianowskiet al\.,[2018](https://arxiv.org/html/2607.24768#bib.bib25); Chenet al\.,[2021](https://arxiv.org/html/2607.24768#bib.bib26)\)between humans and automated evaluation of conversational models\(Güret al\.,[2018](https://arxiv.org/html/2607.24768#bib.bib24)\)\. Subsequent works\(Yaoet al\.,[2024](https://arxiv.org/html/2607.24768#bib.bib15); Luet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib18); Barreset al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib16); Patilet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib17)\)take a step closer to real\-world applications with multi\-turn interactions between the human and agents\.

## 3\.PATHFinder Agent: System Design

![Refer to caption](https://arxiv.org/html/2607.24768v1/x3.png)Figure 2\.PATHFinder interfaces\.PATHFinder Agenthas a four\-stage overflow: 1\) Patient information intake \(standardized form\), 2\) LLM\-generated UI for structured, open\-ended dialogue, 3\) Draft report review and clarification for the patient and 4\) Clinician review during provision of care\.### 3\.1\.Problem Statement

Given a patient’s medical characteristics and social context,PATHFinder Agentmust \(i\) gather relevant information through structured, open\-ended dialogue, \(ii\) curate an individualized prenatal care plan aligned with the specified guidelines, and \(iii\) surface Michigan 211 community resources matched to identified social needs—while enabling ongoing clinician review and oversight \(Figure[3](https://arxiv.org/html/2607.24768#S3.F3)illustrates the workflow\)\.

### 3\.2\.Objectives and Requirements

Core design requirements include interfaces and question flows that mirror existing clinical intake processes, and targeted, non\-redundant follow\-up questions with an interaction mode to prevent patient fatigue in long sessions while capturing data efficiently\.

### 3\.3\.Design

#### Team\.

PATHFinderwas co\-designed by a board\-certified Obstetrician and Gynecologist and Certified Nurse Midwife \(CNM\) alongside two computer scientists, ensuring jointly validated clinical and technical requirements\.

Table 1\.Agent tool suite \(13 tools across 4 categories\)\.
#### Tools and Orchestration\.

The agent is equipped with 13 tools across four categories \(Table[1](https://arxiv.org/html/2607.24768#S3.T1)\)\. Medical tools include a TOLAC \(trial of labor after cesarean\) calculator and a structured referral\-to\-clinician action\. Resource tools implement a hierarchical Michigan 211 query interface\. A personalization tool consults a secondary LLM to generate context\-aware follow\-up questions\. Report tools incrementally build the patient and clinician summaries\. The system instructions consists of∼\\sim14,000 tokens, with domain knowledge \(from ACOG guidelines\(Peahlet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib1)\)\), tool use instructions, workflow and safety policies, where the agent retrieves the knowledge and instructions to orchestrate the conversation\.

#### Social needs resources data\.

We use Michigan 211111[https://mi211\.org/](https://mi211.org/)data organized by category \(food, housing, transportation, utilities, clothing\) and can be queried by subcategory and ZIP code via dedicated tool calls\.

#### Interface design\.

PATHFinder Agentsupports interaction viaForm\+Chat Interface, where the patient first completes a standardized intake form mirroring current clinical processes; the agent then transitions to an adaptive conversational phase to ask individualized follow\-up questions, where another LLM reviews the questions and generates a form\-based interface for the user to interact which supports buttons, check boxes, and switches to reduce the text entered by the user \(see Figure[4\(b\)](https://arxiv.org/html/2607.24768#S3.F4.sf2)\)\. Based on all the available user information and guidelines\(Peahlet al\.,[2025](https://arxiv.org/html/2607.24768#bib.bib1)\), appropriate reports and timelines are generated for the clinicians to review and patients to look at and clarify the content\. Similarly, the agent is available for the clinician to make edits to the report if needed\.

![Refer to caption](https://arxiv.org/html/2607.24768v1/figures/workflow_ih.png)Figure 3\.System architecture\.React frontend communicates with a FastAPI backend routing to thePATHFinder Agentagent \(LLM \+ tool executor\), oversight classifier, FHIR service, and report generator\. Conversation state is persisted in a relational database\.
#### Workflow\.

The end\-to\-end workflow has four stages\.Stage 1 \(Intake\):the patient fills a standardized form capturing demographics, medical history, gestational age, and social factors \(Figure[4\(a\)](https://arxiv.org/html/2607.24768#S3.F4.sf1)\)\.Stage 2 \(Dynamic interaction\):conditioned on responses and guidelines, the agent asks personalized follow\-up questions to elicit unmet social needs, preferences, and barriers and uses tools in Table[1](https://arxiv.org/html/2607.24768#S3.T1)to compose a plan \(Figure[4\(b\)](https://arxiv.org/html/2607.24768#S3.F4.sf2)\)\.Stage 3 \(Plan synthesis\):the agent invokes report tools to produce a patient\-facing summary, a clinical summary, a visit\-schedule recommendation, and curated Michigan 211 resources \(Figure[4\(c\)](https://arxiv.org/html/2607.24768#S3.F4.sf3)\)\.Stage 4 \(Clinician review\):the clinician reviews the plan with capabilities to approve or edit \(Figure[4\(d\)](https://arxiv.org/html/2607.24768#S3.F4.sf4)\)\.

![Refer to caption](https://arxiv.org/html/2607.24768v1/figures/intake.png)\(a\)Patient intake form \(Stage1\)\.
![Refer to caption](https://arxiv.org/html/2607.24768v1/figures/dynamic_interaction.png)\(b\)Dynamic interaction with LLM generating intermediate UI\.
![Refer to caption](https://arxiv.org/html/2607.24768v1/figures/patient_report.png)\(c\)Patient draft report review and clarification \(Stage 3\)\.
![Refer to caption](https://arxiv.org/html/2607.24768v1/figures/clinician_review.png)\(d\)Clinician oversight and review \(Stage4\)\.

Figure 4\.PATHFinder Agentinterfaces\.\(a\)Structured patient intake form; \(b\)PATHFinder Agentinterviews the user to understand additional medical history details, and social determinants of health factors; \(c\) Patient reviews the draft report for transparency and clarification; \(d\)Clinician oversight dashboard showing conversation stage, risk flags, and escalation controls\.

### 3\.4\.Clinician Oversight

Once the draft plan is generated, the clinician reviews the patient summary, report, and timeline to validate it and change it based on additional conversations and concerns\. We highlight that this workflow potentially allows the clinician to better address concerns and provide care to patients while being involved in the prenatal care planning and validation of the plan\.

### 3\.5\.Evaluation

We evaluatePATHFinder Agenton synthetic patient profiles spanning diverse medical histories \(e\.g\., high\-risk pregnancies, TOLAC candidates\)\. Each profile pairs with expert\-curated rubrics specifying the expectations along the following dimensions: 1\) the right visit frequency, 2\) the right services \(testing, recommendations, etc\.\), 3\) timing for antenatal testing, 4\) timing for growth ultrasound, and 5\) modality of care to the patient \(mandatory in\-person, mix of in\-person and group health, etc\.\)\.

We adopt LLM\-as\-judge scoring to measure the performance of frontier models’ recommendations across the five dimensions with the final score per condition normalized to 1\. Table[5](https://arxiv.org/html/2607.24768#S3.F5)shows aggregate rubric scores across state\-of\-the\-art LLMs from OpenAI \(GPT\-5\.2, GPT\-4o\) and Google \(Gemini 2\.5 pro and flash\)\. Furthermore, Figure[5](https://arxiv.org/html/2607.24768#S3.F5)breaks down scores by dimension, with visit frequency being the easiest across all models and recommending antenatal testing and other services being the most difficult\. These results suggest that robust oversight measures, both LLM\- and human\-driven, must be integrated into these systems with appropriate communication to ensure that deployed models do not make mistakes\.

Table 2\.Average rubric\-based scores \(0\-100%\) across frontier models\. Higher is better\.![Refer to caption](https://arxiv.org/html/2607.24768v1/x4.png)Figure 5\.Per\-dimension rubric scores\.Antenatal testing and services recommendation show the widest performance gap across models \(1 is lowest, 5 is highest\)\.

## 4\.Future Work

Future research will focus on establishing formal accuracy guarantees forPATHFinder Agentto strengthen the system’s technical reliability and performance standards\. We will concurrently conduct a series of human participant experiments\.

## Acknowledgements

This work was partially supported by funding from Google and the University of Michigan \(including the Raoul Wallenberg Institute, E\-Health and Artificial Intelligence, and the Center for Academic Innovation\)\. We also thank David Stutz for helpful discussions during the project\.

## References

- E\. Balk, V\. Danilack, M\. Bhuma,et al\.\(2023a\)Reduced compared with traditional schedules for routine antenatal visits: a systematic review\.Obstetrics & Gynecology142\(1\),pp\. 8–18\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Balk, V\. Danilack, W\. Cao,et al\.\(2023b\)Televisits compared with in\-person visits for routine antenatal care: a systematic review\.Obstetrics & Gynecology142\(1\),pp\. 19–29\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- V\. Barres, H\. Dong, S\. Ray, X\. Si, and K\. Narasimhan \(2025\)τ2\\tau^\{2\}\-Bench: evaluating conversational agents in a dual\-control environment\.External Links:2506\.07982,[Link](https://arxiv.org/abs/2506.07982)Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Bellerose, M\. Rodriguez, and P\. Vivier \(2022\)A systematic review of the qualitative literature on barriers to high\-quality prenatal and postpartum care among low\-income women\.Health Services Research57\(4\),pp\. 775–785\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Betron, T\. McClair, S\. Currie, and J\. Banerjee \(2018\)Expanding the agenda for addressing mistreatment in maternity care: a mapping review and gender analysis\.Reproductive Health15\(1\),pp\. 143\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- P\. Budzianowski, T\. Wen, B\. Tseng, I\. Casanueva, S\. Ultes, O\. Ramadan, and M\. Gašić \(2018\)Multiwoz–a large\-scale multi\-domain wizard\-of\-oz dataset for task\-oriented dialogue modelling\.arXiv preprint arXiv:1810\.00278\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Chen, H\. Chen, Y\. Yang, A\. Lin, and Z\. Yu \(2021\)Action\-based conversations dataset: a corpus for building more in\-depth task\-oriented dialogue systems\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 3002–3017\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- I\. Gür, D\. Hakkani\-Tür, G\. Tür, and P\. Shah \(2018\)User modeling for task oriented dialogues\.In2018 IEEE Spoken Language Technology Workshop \(SLT\),Vol\.,pp\. 900–906\.External Links:[Document](https://dx.doi.org/10.1109/SLT.2018.8639652)Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Jiang, K\. C\. Black, G\. Geng, D\. Park, J\. Zou, A\. Y\. Ng, and J\. H\. Chen \(2025\)MedAgentBench: a virtual ehr environment to benchmark medical llm agents\.NEJM AI2\(9\),pp\. AIdbp2500144\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Johri, J\. Jeong, B\. A\. Tran, D\. I\. Schlessinger, S\. Wongvibulsin, Z\. R\. Cai, R\. Daneshjou, and P\. Rajpurkar \(2024\)CRAFT\-md: a conversational evaluation framework for comprehensive assessment of clinical llms\.InAAAI 2024 Spring Symposium on Clinical Foundation Models,Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Kilpatrick, L\. Papile, and G\. Macones \(2017\)Guidelines for perinatal care\.8th edition,American Academy of Pediatrics/The American College of Obstetricians and Gynecologists,Elk Grove Village, IL/Washington, D\.C\.\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Lu, T\. Holleis, Y\. Zhang, B\. Aumayer, F\. Nan, H\. Bai, S\. Ma, S\. Ma, M\. Li, G\. Yin,et al\.\(2025\)Toolsandbox: a stateful, conversational, interactive evaluation benchmark for llm tool use capabilities\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1160–1183\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Mohamoud, E\. Cassidy, E\. Fuchs,et al\.\(2023\)Vital signs: maternity care experiences \- United States, April 2023\.MMWR Morbidity and Mortality Weekly Report72\(35\),pp\. 961–967\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- National Academies of Sciences, Engineering, and Medicine \(2019\)Integrating social care into the delivery of health care: moving upstream to improve the nation’s health\.National Academies Press,Washington, D\.C\.\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Nijagal, D\. Patel, C\. Lyles,et al\.\(2021\)Using human centered design to identify opportunities for reducing inequities in perinatal care\.BMC Health Services Research21\(1\),pp\. 714\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Palepu, V\. Liévin, W\. Weng, K\. Saab, D\. Stutz, Y\. Cheng, K\. Kulkarni, S\. S\. Mahdavi, J\. Barral, D\. R\. Webster,et al\.\(2025\)Towards conversational ai for disease management\.arXiv preprint arXiv:2503\.06074\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- S\. G\. Patil, T\. Zhang, X\. Wang, and J\. E\. Gonzalez \(2023\)Gorilla: large language model connected with massive apis, 2023\.URL https://arxiv\. org/abs/2305\.15334\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- S\. G\. Patil, H\. Mao, C\. Cheng\-Jie Ji, F\. Yan, V\. Suresh, I\. Stoica, and J\. E\. Gonzalez \(2025\)The berkeley function calling leaderboard \(bfcl\): from tool use to agentic evaluation of large language models\.InForty\-second International Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- A\. F\. Peahl, R\. Gourevitch, E\. Luo,et al\.\(2020a\)Right\-sizing prenatal care to meet patients’ needs and improve maternity care value\.Obstetrics & Gynecology135\(5\),pp\. 1027–1037\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- A\. F\. Peahl, A\. Novara, M\. Heisler, V\. K\. Dalton, M\. H\. Moniz, and R\. D\. Smith \(2020b\)Patient preferences for prenatal and postpartum care delivery: a survey of postpartum women\.Obstetrics & Gynecology135\(5\),pp\. 1038–1046\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- A\. F\. Peahl, A\. Powell, H\. Berlin,et al\.\(2021a\)Patient and provider perspectives of a new prenatal care model introduced in response to the coronavirus disease 2019 pandemic\.American Journal of Obstetrics and Gynecology224\(4\),pp\. 384\.e1–384\.e11\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- A\. F\. Peahl, C\. Zahn, M\. Turrentine,et al\.\(2021b\)The Michigan plan for appropriate tailored health care in pregnancy prenatal care recommendations\.Obstetrics & Gynecology138\(4\),pp\. 593–602\.Cited by:[§1](https://arxiv.org/html/2607.24768#S1.p3.1),[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- A\. F\. Peahl, R\. D\. Smith, and M\. H\. Moniz \(2020c\)Prenatal care redesign: creating flexible maternity care models through virtual care\.American journal of obstetrics and gynecology223\(3\),pp\. 389–e1\.Cited by:[§1](https://arxiv.org/html/2607.24768#S1.p1.1)\.
- A\. Peahl, J\. C\. Phillippi, and M\. A\. Turrentine \(2025\)Tailored prenatal care delivery for pregnant individuals\.OBSTETRICS AND GYNECOLOGY145\(5\),pp\. 565–577\.Cited by:[§1](https://arxiv.org/html/2607.24768#S1.p2.1),[§3\.3](https://arxiv.org/html/2607.24768#S3.SS3.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2607.24768#S3.SS3.SSS0.Px4.p1.1)\.
- Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian,et al\.\(2023\)Toolllm: facilitating large language models to master 16000\+ real\-world apis\.arXiv preprint arXiv:2307\.16789\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- K\. Saab, J\. Freyberg, C\. Park, T\. Strother, Y\. Cheng, W\. Weng, D\. G\. Barrett, D\. Stutz, N\. Tomasev, A\. Palepu,et al\.\(2025\)Advancing conversational diagnostic ai with multimodal reasoning\.arXiv preprint arXiv:2505\.04653\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Semega, M\. Kollar, J\. Creamer, and A\. Mohanty \(2021\)Income and poverty in the united states: 2018\.Technical reportU\.S\. Census Bureau\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Tabata\-Kelly, X\. Hu, M\. J\. Dill, P\. M\. Alberti, K\. Bullock, W\. Crown, M\. Fair, P\. May, P\. Ortega, and J\. Perloff \(2024\)Physician engagement in addressing health\-related social needs and burnout\.JAMA network open7\(12\),pp\. e2452152–e2452152\.Cited by:[§1](https://arxiv.org/html/2607.24768#S1.p3.1)\.
- S\. Trost, J\. Beauregard, G\. Chandra,et al\.\(2022\)Pregnancy\-related deaths: data from maternal mortality review committees in 36 US states, 2017–2019\.Note:[https://www\.cdc\.gov/reproductivehealth/maternalmortality/erase\-mm/data\-mmrc\.html](https://www.cdc.gov/reproductivehealth/maternalmortality/erase-mm/data-mmrc.html)Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Turrentine \(2023\)Prenatal care visit frequency: how much is too much, and how little is too little?\.Obstetrics & Gynecology142\(1\),pp\. 6–7\.Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Vedadi, D\. Barrett, N\. Harris, E\. Wulczyn, S\. Reddy, R\. Ruparel, M\. Schaekermann, T\. Strother, R\. Tanno, Y\. Sharma,et al\.\(2025\)Towards physician\-centered oversight of conversational diagnostic ai\.arXiv preprint arXiv:2507\.15743\.Cited by:[§1](https://arxiv.org/html/2607.24768#S1.p4.1),[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Xu, Y\. Zhuang, Y\. Zhong, Y\. Yu, X\. Tang, H\. Wu, M\. D\. Wang, P\. Ruan, D\. Yang, T\. Wang,et al\.\(2025\)Medagentgym: training llm agents for code\-based medical reasoning at scale\.InThe Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance,Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2024\)τ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.External Links:2406\.12045,[Link](https://arxiv.org/abs/2406.12045)Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. Cao \(2022\)React: synergizing reasoning and acting in language models\.InThe eleventh international conference on learning representations,Cited by:[§2](https://arxiv.org/html/2607.24768#S2.SS0.SSS0.Px3.p1.1)\.

Similar Articles

Pioneering an AI clinical copilot with Penda Health

OpenAI Blog

OpenAI partnered with Penda Health in Kenya to study an LLM-powered clinical copilot called AI Consult, which demonstrated a 16% relative reduction in diagnostic errors and 13% reduction in treatment errors across 39,849 patient visits. The study highlights successful real-world implementation of AI in primary care and provides a template for safe, effective deployment of LLMs to support clinicians.