Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study

arXiv cs.CL Papers

Summary

A study evaluating zero-shot GPT-4o-mini for predicting child stunting from Bangladesh Demographic and Health Survey data, comparing against a random forest baseline and assessing fairness across demographic groups and temporal robustness. Results show comparable balanced accuracy but notable fairness disparities across residence and wealth categories.

arXiv:2607.29082v1 Announce Type: new Abstract: Child malnutrition remains a major public health challenge in low- and middle-income countries, particularly in South Asia, where early identification of vulnerable children is critical for timely intervention and resource allocation. This study aims to evaluate the feasibility, fairness, and temporal robustness of using a pretrained large language model (LLM) in a zero-shot setting for child stunting prediction using population health survey data. Using Bangladesh Demographic and Health Survey (BDHS) data collected between 2007 and 2022, we transformed maternal, child, healthcare, and household characteristics into semantically interpretable prompt-based representations and evaluated GPT-4o-mini for zero-shot stunting prediction, comparing its performance against a random forest baseline and assessing fairness across demographic and socioeconomic groups as well as temporal robustness across survey waves. The results demonstrate that zero-shot inference using GPT-4o-mini achieved comparable balanced accuracy to the supervised baseline while exhibiting substantially higher sensitivity for identifying stunting cases, relatively consistent performance across child sex groups, and stable predictive behaviour across BDHS waves; however, important fairness disparities were observed across residence and household wealth categories, highlighting the need for further investigation before deployment of foundation models in public health prediction settings.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:35 AM

# Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study
Source: [https://arxiv.org/html/2607.29082](https://arxiv.org/html/2607.29082)
\\copyrightclause

Copyright for this paper by its authors\. Use permitted under Creative Commons License Attribution 4\.0 International \(CC BY 4\.0\)\.

\\conference

Joint Proceedings of the AIME 2026 Workshops: 1st International Workshop on Multicentric and Privacy\-preserving Learning in Healthcare, Foundation Models for Public Health and Epidemiology: From Promise to Practice, and First International Workshop on Knowledge Graphs for Health, July 10, 2026, Ottawa, Canada

\[orcid=0000\-0002\-6798\-6535, email=akabir@csu\.edu\.au, \]\\cormark\[1\]

\[orcid=0000\-0003\-3452\-8367, email=mdhaque@csu\.edu\.au \]

\\cortext

\[1\]Corresponding author\.

Md Ahshanul HaqueSchool of Computing, Mathematics and Engineering, Charles Sturt University, Bathurst, NSW 2795, Australia

\(2026\)

###### Abstract

Child malnutrition remains a major public health challenge in low\- and middle\-income countries, particularly in South Asia, where early identification of vulnerable children is critical for timely intervention and resource allocation\. This study aims to evaluate the feasibility, fairness, and temporal robustness of using a pretrained large language model \(LLM\) in a zero\-shot setting for child stunting prediction using population health survey data\. Using Bangladesh Demographic and Health Survey \(BDHS\) data collected between 2007 and 2022, we transformed maternal, child, healthcare, and household characteristics into semantically interpretable prompt\-based representations and evaluated GPT\-4o\-mini for zero\-shot stunting prediction, comparing its performance against a random forest baseline and assessing fairness across demographic and socioeconomic groups as well as temporal robustness across survey waves\. The results demonstrate that zero\-shot inference using GPT\-4o\-mini achieved comparable balanced accuracy to the supervised baseline while exhibiting substantially higher sensitivity for identifying stunting cases, relatively consistent performance across child sex groups, and stable predictive behaviour across BDHS waves; however, important fairness disparities were observed across residence and household wealth categories, highlighting the need for further investigation before deployment of foundation models in public health prediction settings\.

###### keywords:

LLM\\sepZero\-shot\\sepChild malnutrition\\sepStunting\\sepFairness\\sepRobustness

## 1Introduction

Child malnutrition remains a major global public health challenge, particularly in low\- and middle\-income countries, where undernutrition during early childhood is associated with impaired growth, cognitive development, increased disease susceptibility, and elevated mortality risk\[victora2008maternal,black2013maternal\]\. Despite improvements in maternal and child healthcare, South Asian countries, including Bangladesh, continue to experience substantial burdens of childhood stunting, wasting, and underweight\[unicef2023malnutrition,bdhs2022\]\. Early identification of vulnerable children is therefore important for timely intervention and evidence\-based public health planning\.

Maternal, child, healthcare, and household socioeconomic characteristics have consistently been identified as important determinants of child malnutrition\[black2013maternal,victora2008maternal\]\. Motivated by the increasing availability of Bangladesh Demographic and Health Survey \(BDHS\) datasets, recent studies have applied conventional machine learning \(ML\) approaches such as logistic regression, random forests, and gradient boosting to predict childhood malnutrition outcomes\[talukder2020machine,khan2021model,islam2024prediction\]\. However, these studies primarily rely on task\-specific supervised learning and provide limited investigation into the applicability of large language models \(LLMs\) for public health prediction tasks\.

Recent advances in LLMs have demonstrated strong zero\-shot reasoning capabilities across various biomedical and healthcare applications\[bommasani2021opportunities,singhal2023large,thirunavukarasu2023large\]\. Nevertheless, the use of zero\-shot LLMs for structured population health survey data remains largely unexplored, particularly regarding fairness across demographic and socioeconomic groups and temporal robustness\.

In this study, we evaluate a pretrained LLM, GPT\-4o\-mini, in a zero\-shot setting for child stunting prediction using BDHS data collected between 2007 and 2022\. Using prompt\-based feature–value representations of maternal, child, healthcare, and household characteristics, we assess predictive performance, fairness across child sex, place of residence, and household wealth categories, and temporal robustness across BDHS waves\. Our findings contribute to the emerging discussion on the practical applicability and limitations of LLMs for population\-level public health prediction tasks\.

## 2Methodology

### 2\.1Dataset

This study utilised child\-level BDHS data collected between 2007 and 2022\[bdhs2022,dhsprogram\]\. BDHS is a nationally representative cross\-sectional survey containing extensive maternal, child, healthcare, demographic, and household socioeconomic information, and has been widely used in population health and machine learning studies\[talukder2020machine,khan2021model,islam2024prediction\]\. Children under five years of age with complete anthropometric and covariate information were included in the analysis\. The target outcome was childhood stunting, defined according to the World Health Organization \(WHO\) child growth standards as height\-for\-age z\-score \(HAZ\) below−2\-2standard deviations from the WHO reference population\[who2006child\]\. Features considered in this study were identified as significant determinants of child malnutrition in prior literature\[black2013maternal,victora2008maternal,talukder2020machine,islam2024prediction\]and are summarised in Table[1](https://arxiv.org/html/2607.29082#S2.T1)\. Samples with missing values for any selected feature were excluded during preprocessing, resulting in a final analytical dataset consisting of 17106 children, including 5623 stunted and 11483 non\-stunted cases\.

Table 1:Summary of the selected features used in this study stratified by child stunting status\. Values are presented as mean±\\pmstandard deviation or frequency \(%\)\. For binary variables, only one category \(e\.g\., Yes or No\) is reported for brevity\. Thepp\-values indicate statistical differences between stunted and non\-stunted groups\.Feature nameFeature valueStunting\(n=5623, %\)Non\-stunting\(n=11483, %\)p\-valueMother’s age \(years\)Mean±\\pmSD25\.37±\\pm6\.0725\.16±\\pm5\.67p=0\.032p=0\.032Mother’s education levelSecondary2449 \(43\.6\)5674 \(49\.4\)p<0\.001p<0\.001Primary1838 \(32\.7\)2787 \(24\.3\)Tertiary384 \(6\.8\)1970 \(17\.2\)No education952 \(16\.9\)1052 \(9\.2\)Mother’s BMIMean±\\pmSD20\.82±\\pm3\.5921\.96±\\pm3\.91p<0\.001p<0\.001Worked in last 12 monthsNo4290 \(76\.3\)8868 \(77\.2\)p=0\.176p=0\.176Number of living childrenMean±\\pmSD2\.23±\\pm1\.352\.01±\\pm1\.16p<0\.001p<0\.001Mother’s age at first birthMean±\\pmSD17\.98±\\pm3\.0918\.69±\\pm3\.45p<0\.001p<0\.001Involved in decisionsYes2558 \(45\.5\)5461 \(47\.6\)p=0\.011p=0\.011Believes wife beatingNo3997 \(71\.1\)8818 \(76\.8\)p<0\.001p<0\.001Exposed to mass mediaYes3165 \(56\.3\)7693 \(67\.0\)p<0\.001p<0\.001Experienced death of childNo4848 \(86\.2\)10404 \(90\.6\)p<0\.001p<0\.001Father’s education levelSecondary1598 \(28\.4\)3914 \(34\.1\)p<0\.001p<0\.001Primary1972 \(35\.1\)3169 \(27\.6\)No education1546 \(27\.5\)1940 \(16\.9\)Tertiary507 \(9\.0\)2460 \(21\.4\)Father’s occupationWorker2628 \(46\.7\)5157 \(44\.9\)p<0\.001p<0\.001Business or prof1322 \(23\.5\)3708 \(32\.3\)Agriculture1563 \(27\.8\)2375 \(20\.7\)Not working110 \(2\.0\)243 \(2\.1\)Child age \(months\)Mean±\\pmSD24\.46±\\pm13\.6919\.37±\\pm14\.45p<0\.001p<0\.001Child sexMale3014 \(53\.6\)5901 \(51\.4\)p=0\.007p=0\.007Birth order of the childMean±\\pmSD2\.41±\\pm1\.552\.12±\\pm1\.30p<0\.001p<0\.001Birth interval≥\\geq24 months3238 \(57\.6\)6360 \(55\.4\)p<0\.001p<0\.001No prior birth1904 \(33\.9\)4457 \(38\.8\)<<24 months481 \(8\.6\)666 \(5\.8\)Breastfeeding started earlyYes4212 \(74\.9\)8201 \(71\.4\)p<0\.001p<0\.001Child is being breastfedYes4205 \(74\.8\)9028 \(78\.6\)p<0\.001p<0\.001Child received vitamin A supplYes3614 \(64\.3\)6829 \(59\.5\)p<0\.001p<0\.001Skilled birth attendantNo3602 \(64\.1\)5389 \(46\.9\)p<0\.001p<0\.001Delivered in a health facilityNo3817 \(67\.9\)5909 \(51\.5\)p<0\.001p<0\.001Delivery by caesarean sectionNo4623 \(82\.2\)7924 \(69\.0\)p<0\.001p<0\.001Received at least 4 ANC visitsNo4284 \(76\.2\)7321 \(63\.8\)p<0\.001p<0\.001Drinking water sourceSafe4943 \(87\.9\)9889 \(86\.1\)p=0\.001p=0\.001Type of toilet facilityHygienic2883 \(51\.3\)7030 \(61\.2\)p<0\.001p<0\.001Place of residence was ruralYes4103 \(73\.0\)7575 \(66\.0\)p<0\.001p<0\.001Household wealth indexPoorest1621 \(28\.8\)1947 \(17\.0\)p<0\.001p<0\.001Richest679 \(12\.1\)2787 \(24\.3\)Richer961 \(17\.1\)2465 \(21\.5\)Poorer1261 \(22\.4\)2086 \(18\.2\)Middle1101 \(19\.6\)2198 \(19\.1\)Number of householdMean±\\pmSD6\.01±\\pm2\.686\.12±\\pm2\.71p=0\.011p=0\.011
### 2\.2Prompt Construction and Zero\-Shot Inference

Each child’s record was represented by serializing feature–value pairs into a structured list\-based prompt format \(Table[3](https://arxiv.org/html/2607.29082#S2.T3)\)\. Prior studies have shown that list\-style serialization preserves tabular semantics and reduces ambiguity when applying LLMs to structured data inference tasks\[hegselmann2023tabllm\]\. Following instruction\-tuning paradigms used in modern LLMs, the prompt explicitly defined the task objective, binary output space, and serialized input record\. The prompt template was iteratively refined through empirical validation following prior instruction\-based prompt optimization approaches\.

Table 3:Example of the prompt used for zero\-shot inference\.You are a clinical classification model\. Based on the provided characteristics, predict whether the child is likely to be stunted\. Output must be a single token: 0 = not stunted, 1 = stunted\.Characteristics: Mother’s current age \(years\): 24, Child sex: Female, …Answer with exactly one token \(0 or 1\)\.GPT\-4o\-mini was used as a pretrained LLM for zero\-shot inference due to its strong instruction\-following capability, cost efficiency, and support for token\-level log\-probabilities, enabling probabilistic classification through likelihood comparison of predefined output tokens\. Inference through the OpenAI API was performed using deterministic decoding with a temperature set to 0\. For baseline comparison, a random forest classifier withclass\_weight=‘balanced’andrandom\_state=42was implemented to account for class imbalance\.

## 3Results and Discussion

Table[4](https://arxiv.org/html/2607.29082#S3.T4)presents the fairness and temporal robustness performance of the zero\-shot GPT\-4o\-mini model for child stunting prediction\. Overall, GPT\-4o\-mini achieved a balanced accuracy of 58%, and AUROC of 0\.632, which were reasonably comparable to the random forest baseline model evaluated using 5\-fold cross\-validation \(balanced accuracy: 57%, AUROC: 0\.685\)\. Notably, GPT\-4o\-mini demonstrated substantially higher sensitivity \(77\.5%\) than the random forest \(23\.4%\), indicating greater capability in identifying stunting cases\. However, this was accompanied by considerably lower specificity \(38\.6% versus 90\.7%\), suggesting that the zero\-shot LLM tended to over\-predict positive stunting outcomes relative to the supervised machine learning baseline\.

Table 4:Fairness and temporal robustness performance of the zero\-shot LLM for child stunting prediction\. CV: cross\-validation\.CategoryValueTestBalancedAUROCSensitivitySpecificitysampleaccuracy \(%\)\(%\)\(%\)Random Forest5\-fold CV57\.00\.68523\.490\.7GPT\-4o\-mini1710658\.00\.63277\.538\.6Fairness of GPT\-4o\-miniChild sexMale891558\.10\.63176\.739\.4Female819158\.00\.63378\.337\.7ResidenceRural1167855\.40\.62086\.024\.9Urban542859\.80\.64554\.465\.2WealthPoorest356850\.00\.581100\.00\.0indexPoorer334750\.90\.56098\.73\.0Middle329953\.10\.56373\.033\.1Richer342655\.00\.57854\.155\.9Richest346652\.90\.58324\.781\.1Temporal robustness of GPT\-4o\-miniBDHS200721958\.20\.61147\.668\.9round2011258756\.50\.62283\.829\.22014166156\.60\.62177\.135\.92018167657\.90\.61773\.642\.3202290957\.90\.61871\.744\.1The fairness evaluation showed relatively consistent performance across child sex groups, with nearly identical balanced accuracy and AUROC values for male and female children\. In contrast, larger disparities were observed across residence and wealth categories\. GPT\-4o\-mini exhibited substantially higher sensitivity for rural children \(86%\) but lower specificity \(24\.9%\), indicating a tendency to infer stunting more frequently among rural populations\. Similarly, the model achieved 100% sensitivity and 0% specificity for the poorest wealth category, suggesting potential over\-classification of stunting among socioeconomically disadvantaged children\. To further investigate whether this behaviour was driven primarily by the wealth index variable itself, we re\-evaluated the 3568 samples from the poorest category after excluding the wealth index feature from the prompt\. The model still demonstrated very high sensitivity \(96%\) and low specificity \(6%\), suggesting that the observed prediction pattern may not solely originate from explicit wealth information, but potentially from correlated maternal, household, or other characteristics\. These findings warrant further investigation regarding implicit socioeconomic associations learned by pretrained LLMs\.

The temporal robustness analysis demonstrated relatively stable performance across BDHS survey waves between 2007 and 2022\. Balanced accuracy remained within a narrow range \(56\.5%–58\.2%\), while AUROC values varied only modestly between 0\.611 and 0\.622\. These findings suggest that the zero\-shot LLM maintained reasonably consistent predictive behaviour despite evolving demographic, socioeconomic, and healthcare distributions across survey years\. Nevertheless, sensitivity and specificity varied substantially across BDHS rounds\. Earlier survey waves, particularly 2007, showed lower sensitivity and higher specificity, whereas later waves exhibited the opposite trend\. This shift may reflect temporal changes in population characteristics and healthcare access patterns\.

## 4Conclusions and Future Work

This study evaluated the feasibility of using a pretrained LLM, GPT\-4o\-mini, in a zero\-shot setting to predict child stunting\. The findings based on BDHS dataset demonstrate that zero\-shot inference using a pretrained LLM can achieve comparable balanced accuracy to a conventional random forest baseline while exhibiting substantially higher sensitivity for identifying stunting cases\. The fairness analysis revealed relatively consistent performance across child sex but highlighted important disparities across residence and household wealth categories, indicating the need for further investigation on potential socioeconomic biases in LLM\-based public health prediction\. The temporal robustness evaluation further showed relatively stable predictive performance across BDHS waves despite changing population distributions over time\. Overall, the study highlights both the potential and limitations of LLMs for population\-level health prediction tasks without task\-specific training\. Future work should further investigate fairness implications, explore few\-shot prompting and prompt optimization strategies, and compare different LLMs with robust ML approaches to improve the reliability and generalizability of foundation models for public health applications\.

## Declaration on Generative AI

During the preparation of this work, the authors used ChatGPT\-5\.4 for grammar and spelling checks\. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the publication’s content\.

## References

Similar Articles

Zero-shot World Models Are Developmentally Efficient Learners [R]

Reddit r/MachineLearning

Researchers introduce Zero-shot World Models (ZWM), an approach that achieves visual competence comparable to state-of-the-art models while trained on minimal data (single child's visual experience) without task-specific training. This work demonstrates a path toward more data-efficient AI systems that match human developmental learning efficiency.