FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

arXiv cs.AI Papers

Summary

Introduces FAM-Bench, a multimodal benchmark with 2500 expert-verified instances across 13 diet-related health conditions, designed to evaluate AI models' ability to assess dish suitability for specific health conditions, moving beyond basic food recognition to condition-aware reasoning.

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health condition. Existing food AI benchmarks primarily evaluate dish recognition, recipe understanding, nutrient estimation, or general nutrition question answering, leaving this health-aware decision layer largely untested. We introduce FAM-Bench, a multi-modal Food-as-Medicine benchmark with 2500 nutrition-expert-verified instances across 13 diet-related health conditions. The benchmark contains two complementary tasks: dish-level suitability assessment, where models judge whether a dish is suitable for a condition from its image and ingredient list, and comparative dish analysis, where models rank four candidate dishes by condition-specific suitability. Both tasks require integrating ingredient evidence, visual preparation cues, and clinical nutrition constraints, providing a standardized testbed for grounded health-aware reasoning in language and vision-language models.
Original Article
View Cached Full Text

Cached at: 06/01/26, 09:27 AM

# FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning
Source: [https://arxiv.org/html/2605.31410](https://arxiv.org/html/2605.31410)
Mingyang Mao1,\*,Bhargav Rishi Medisetti2,\*,Utkarsh Grover1,\*,Tanvir Ibrahim2, Wenyan Li3,Tingting Zhang2,Xiaomin Lin1,†

1Department of Electrical Engineering, University of South Florida 2Muma College of Business, University of South Florida 3Computer Science, University of Copenhagen

\*Equal contribution\. †Corresponding author:xlin2@usf\.edu

###### Abstract

Food\-as\-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health condition\. Existing food AI benchmarks primarily evaluate dish recognition, recipe understanding, nutrient estimation, or general nutrition question answering, leaving this health\-aware decision layer largely untested\. We introduce FAM\-Bench, a multi\-modal Food\-as\-Medicine benchmark with 2500 nutrition\-expert\-verified instances across 13 diet\-related health conditions\. The benchmark contains two complementary tasks: dish\-level suitability assessment, where models judge whether a dish is suitable for a condition from its image and ingredient list, and comparative dish analysis, where models rank four candidate dishes by condition\-specific suitability\. Both tasks require integrating ingredient evidence, visual preparation cues, and clinical nutrition constraints, providing a standardized testbed for grounded health\-aware reasoning in language and vision\-language models\.

FAM\-Bench: A Multimodal Benchmark for Condition\-Aware Food\-as\-Medicine Reasoning

Mingyang Mao1,\*, Bhargav Rishi Medisetti2,\*, Utkarsh Grover1,\*, Tanvir Ibrahim2,Wenyan Li3,Tingting Zhang2,Xiaomin Lin1,†1Department of Electrical Engineering, University of South Florida2Muma College of Business, University of South Florida3Computer Science, University of Copenhagen\*Equal contribution\.†Corresponding author:xlin2@usf\.edu

## 1Introduction

*“Let food be thy medicine\.”*– Hippocrates

Diet is a major modifiable factor in chronic disease prevention and management\. Dietary patterns are strongly associated with cardiometabolic and gastrointestinal conditions, including cardiovascular disease, diabetes, obesity, hypertension, and related disorders\(Willett,[1994](https://arxiv.org/html/2605.31410#bib.bib10); Mozaffarian,[2016](https://arxiv.org/html/2605.31410#bib.bib9)\)\. Chronic diseases also impose substantial clinical and economic burden, accounting for nearly 90% of the United States’ $4\.9 trillion annual health\-care expenditure\(Centers for Disease Control and Prevention,[2025](https://arxiv.org/html/2605.31410#bib.bib12)\)\. These pressures have renewed interest in*Food Is Medicine*, which integrates clinically appropriate food resources into health care to prevent, manage, or treat disease\(Volppet al\.,[2023](https://arxiv.org/html/2605.31410#bib.bib13)\)\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x1.png)Figure 1:From food understanding to Food\-as\-Medicine reasoning\.Prior benchmarks ask*what is this dish?*,*what does it contain?*, or*is it appropriate?*on text\-only triples\.FAM\-Benchadds the missing decision layer: given a dish image, its ingredient list, and a target condition, the model must produce a suitability verdict grounded in the offending ingredients\.The core challenge is decision\-oriented\. A dish is not universally healthy or unhealthy: its suitability depends on the target condition, ingredients, preparation method, and nutritional implications\. The same meal may be acceptable for one health context but inappropriate for another\. A dish with hidden sodium may be risky for hypertension; added sugar or refined carbohydrates may conflict with type 2 diabetes; high potassium, phosphorus, or protein load may matter for chronic kidney disease; and trigger ingredients may affect GERD\.

Current food AI benchmarks do not directly test this capability\. Existing resources primarily evaluate dish recognition, recipe understanding, ingredient extraction, nutrient estimation, or general nutrition question answering\(Bossardet al\.,[2014](https://arxiv.org/html/2605.31410#bib.bib14); Marınet al\.,[2021](https://arxiv.org/html/2605.31410#bib.bib15); Yagciogluet al\.,[2018](https://arxiv.org/html/2605.31410#bib.bib16); Wróblewskaet al\.,[2022](https://arxiv.org/html/2605.31410#bib.bib17); Thameset al\.,[2021](https://arxiv.org/html/2605.31410#bib.bib18); Huaet al\.,[2024](https://arxiv.org/html/2605.31410#bib.bib19); Zhanget al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib26)\)\. These tasks provide important foundations, but they remain largely descriptive\. They do not systematically evaluate whether models can convert multimodal food evidence into condition\-specific dietary decisions\. Figure[1](https://arxiv.org/html/2605.31410#S1.F1)illustrates how FAM\-Bench adds the missing detection layer\.

We introduceFAM\-Bench,111Code, data, and evaluation scripts are released at an anonymized repository for review:[https://github\.com/anonymous\-research\-artifact123/Food\-as\-medicine](https://github.com/anonymous-research-artifact123/Food-as-medicine)\. The repository will be de\-anonymized upon acceptance\.a multimodal Food\-as\-Medicine benchmark for this missing decision layer, which requires models to answer not only*what is this dish?*or*what does it contain?*, but*is this dish appropriate for this health condition?*FAM\-Bench contains 2,500 nutrition\-expert\-verified instances derived from 3,859 unique recipes and spanning 13 diet\-related health conditions\.

We evaluate FAM\-Bench on five multimodal models spanning frontier closed\-source systems \(GPT\-5\.4\(OpenAI,[2025](https://arxiv.org/html/2605.31410#bib.bib56)\), Claude Sonnet 4\.6\(Anthropic,[2026](https://arxiv.org/html/2605.31410#bib.bib57)\), Gemini 2\.5 Pro\(Comaniciet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib58)\)\) and open\-weight VLMs \(Qwen3\-VL\-8B\(Baiet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib59)\), Gemma\-3\-12B\(Kamathet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib60)\)\), under baseline, chain\-of\-thought\(COT\)\(Weiet al\.,[2022](https://arxiv.org/html/2605.31410#bib.bib51)\), knowledge injection\(KI\), and CoT\+KI prompting\.

Our contributions can be summarized as:

- •Problem formulation\.We frame Food\-as\-Medicine as a multimodal, condition\-aware decision problem over dish images, ingredient lists, and health\-condition prompts\.
- •Benchmark construction\.We introduce FAM\-Bench, with 2,500 nutrition\-expert\-verified instances from 3,859 recipes across 13 diet\-related health conditions\.
- •Evaluation protocol\.We define two tasks dish\-level suitability assessment and comparative dish analysis with metrics for accuracy, grounding, ranking, and cross\-task consistency\.
- •Empirical findings\.We evaluate five VLMs under four prompting modes and show that verdicts remain easier than rationale grounding or condition\-aware ranking\.

## 2Related Work

#### Food understanding and nutrition benchmarks\.

Food AI benchmarks have primarily evaluated descriptive food understanding\. Food\-101 studies dish classificationBossardet al\.\([2014](https://arxiv.org/html/2605.31410#bib.bib14)\), Recipe1M\+ aligns recipes with images for cross\-modal retrievalMarınet al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib15)\), RecipeQA evaluates procedural recipe comprehensionYagciogluet al\.\([2018](https://arxiv.org/html/2605.31410#bib.bib16)\), and TASTEset extracts structured recipe entities such as ingredients, quantities, and cooking processesWróblewskaet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib17)\)\. Generation\-oriented corpora such as RecipeNLG further support large\-scale recipe synthesisBieńet al\.\([2020](https://arxiv.org/html/2605.31410#bib.bib37)\)\. Nutrition benchmarks move from food identity to food composition: Nutrition5k provides visual and nutritional annotations for real dishesThameset al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib18)\), NutriBench evaluates LLMs on calorie and macronutrient estimationHuaet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib19)\), and recent multimodal benchmarks such as the January Food Benchmark and DiningBench broaden evaluation to ingredient reasoning, nutrition estimation, and dietary\-domain VQAHosseinianet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib20)\); Jinet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib21)\)\. These resources provide the visual, linguistic, and nutritional substrate for food AI, but their targets remain descriptive: what a dish is, what it contains, or how it is prepared\.

#### Personalized and health\-aware dietary reasoning\.

A related line of work moves from description toward dietary guidance\. ChatDiet, NutriGen, HealthGenie, and NutriVision use LLMs, structured knowledge, user preferences, or vision\-language inputs for nutrition recommendation and meal planningYanget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib22)\); Khamesianet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib39)\); Gaoet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib23)\); Veeramreddyet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib40)\)\. Food\-specialized models such as LLaVA\-Chef and FoodSky adapt multimodal and language models to recipe generation, culinary reasoning, and dietetic knowledgeMohbat and Zaki \([2024](https://arxiv.org/html/2605.31410#bib.bib30)\); Zhouet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib31)\)\. The closest benchmark line is health\-aware nutritional reasoning: NGQA formulates nutrition reasoning as graph question answering over users, foods, nutrients, and medical conditionsZhanget al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib26)\), with related work on constrained food\-graph QA, clinical health\-aware generation, food\-safety knowledge graphs, and medicine food homologyChenet al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib27)\); Fenget al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib28)\); Anet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib29)\); Gonget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib24)\); Shaet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib44)\)\. These works connect diet to health context, but they primarily evaluate graph QA, knowledge recall, or system\-specific recommendation rather than multimodal dish\-level decisions over real recipes\.

#### Safety, grounding, and decision\-oriented evaluation\.

Medical\-LLM evaluation has increasingly shifted from accuracy alone toward harm, factuality, hallucination, and reliabilitySinghalet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib45)\); Bediet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib46)\); Palet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib47)\); Minet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib48)\)\. Reasoning and retrieval methods, including Chain\-of\-Thought prompting, Medprompt, RAG, and MedRAG/MIRAGE, provide common mechanisms for grounding knowledge\-intensive medical generationWeiet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib51)\); Kojimaet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib52)\); Noriet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib55)\); Lewiset al\.\([2020](https://arxiv.org/html/2605.31410#bib.bib53)\); Xionget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib54)\)\. Food\-safety benchmarks study contamination, storage, unsafe preparation, chemical exposure, and adversarial food\-safety promptsJacxsenset al\.\([2010](https://arxiv.org/html/2605.31410#bib.bib32)\); Le Vallée and Charlebois \([2015](https://arxiv.org/html/2605.31410#bib.bib33)\); Bryanet al\.\([1992](https://arxiv.org/html/2605.31410#bib.bib34)\); Munckeet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib35)\); Pekmezciet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib36)\); Luoet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib50)\)\. Yet chronic\-condition\-aware dietary suitability remains underrepresented: existing benchmarks rarely test whether models can ground a dietary decision in a dish image, ingredient evidence, and a target health condition\.

#### Position of this work\.

FAM\-Bench operationalizes the missing decision layer in food AI: condition\-specific dietary judgment grounded in dish images and ingredient evidence\. Unlike prior benchmarks that describe foods, estimate nutrients, or test health\-knowledge recall, FAM\-Bench evaluates whether models can assess dish\-level suitability and rank alternatives under explicit health conditions\. This shifts evaluation from food understanding to grounded Food\-as\-Medicine decision making\. We provide a broader discussion of food benchmarks, personalized nutrition systems, food\-domain LLMs, medical\-LLM safety evaluation, and knowledge augmented health reasoning in Appendix[A](https://arxiv.org/html/2605.31410#A1)\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x2.png)Figure 2:Overview of FAM\-Bench\. Recipes are aggregated from health\-information and general\-food publication sources \(§[3\.1](https://arxiv.org/html/2605.31410#S3.SS1)\), normalized into structured dish records, and annotated against a curated knowledge base of 13 diet\-sensitive conditions through a rule\-grounded, expert\-verified pipeline \(§[3\.3](https://arxiv.org/html/2605.31410#S3.SS3)\)\. The resulting 2,500 instances instantiate two complementary tasks: dish\-level suitability assessment and comparative dish analysis \(§[3\.2](https://arxiv.org/html/2605.31410#S3.SS2)\)\.

## 3Benchmark Construction

FAM\-Bench contains 2,500 expert\-verified instances spanning 13 diet\-sensitive health conditions, split into a 1,500\-instance*dish\-level suitability*task and a 1,000\-instance*comparative ranking*task\. Instances are derived from a corpus of recipes curated by medical, clinical\-nutrition, and dietitian sources, and validated by registered nutritionists\. Figure[2](https://arxiv.org/html/2605.31410#S2.F2)overviews the construction pipeline\.

### 3\.1Recipe Collection

#### Sources\.

The corpus aggregates 3,859 recipes from 54 web domains across two complementary categories: a*health\-information*tier \(recipe portals maintained by medical societies, clinical nutrition programs, and public\-health agencies\) and a*general\-food\-publication*tier \(food publications and dietitian\-curated cooking sites\)\. Source composition and the full domain inventory are reported in Appendix[B](https://arxiv.org/html/2605.31410#A2)\. Figure[3](https://arxiv.org/html/2605.31410#S3.F3)summarizes the geographic coverage\.

#### Extraction and filtering\.

We extract recipe metadata, ingredient lines, preparation instructions, and the published nutrition table, then deduplicate and discard entries with unparseable ingredients or unreachable images\. Each retained recipe is stored as a normalized structured record \(Appendix[C](https://arxiv.org/html/2605.31410#A3)\)\.

### 3\.2Task Formulation

Let a dish bed=\(I,G\)d=\(I,G\), whereIIis the dish image andG=\{g1,…,gm\}G=\\\{g\_\{1\},\\ldots,g\_\{m\}\\\}is its ingredient list, and leth∈ℋh\\in\\mathcal\{H\}denote a target health condition\.ℋ\\mathcal\{H\}spans 13 diet\-sensitive conditions covering cardiometabolic, gastrointestinal, hepatic, and renal disorders; the full list is reported in Appendix[G](https://arxiv.org/html/2605.31410#A7)\(Table[5](https://arxiv.org/html/2605.31410#A7.T5)\)\. On this representation we define two complementary tasks that probe absolute and relative dietary judgment, respectively\. Example question layouts for the two tasks are shown in Appendix[K](https://arxiv.org/html/2605.31410#A11)\.

#### Dish\-Level Suitability Assessment\.

An instance isx=\(d,h\)x=\(d,h\)with label

y∈𝒴=\{suitable,not suitable\}\.y\\in\\mathcal\{Y\}=\\\{\\textsc\{suitable\},\\textsc\{not suitable\}\\\}\.A dish issuitablewhen its ingredients and visible preparation cues are compatible withhh, andnot suitablewhen the evidence indicates a condition\-relevant conflict \(e\.g\., added sugar for type 2 diabetes, high\-sodium content for hypertension, fried or high\-fat preparation for GERD or NAFLD\)\. The output is a condition\-specific judgment, not a calorie or macronutrient estimate\.

#### Comparative Dish Analysis\.

An instance isx=\(𝒟,h\)x=\(\\mathcal\{D\},h\)with𝒟=\{d1,…,d4\}\\mathcal\{D\}=\\\{d\_\{1\},\\ldots,d\_\{4\}\\\}, wherehhmay be a single condition or a small set of co\-occurring conditions \(e\.g\., hypertension \+ type 2 diabetes\); the reference output is a permutation

π⋆=\(d\(1\),d\(2\),d\(3\),d\(4\)\)\\pi^\{\\star\}=\(d\_\{\(1\)\},d\_\{\(2\)\},d\_\{\(3\)\},d\_\{\(4\)\}\)ordered from most to least suitable forhh\. This setting reflects real\-world recommendation, in which users choose among alternatives rather than judge a dish in isolation, and forces the model to reason about relative trade\-offs across ingredients, preparation, and condition\-specific risks\. Candidate sets are constructed so that the ranking cannot be solved by dish recognition alone or by a single obvious ingredient\.

### 3\.3Dataset Creation

The 3,859 dish records from Section[3\.1](https://arxiv.org/html/2605.31410#S3.SS1)are converted into the 2,500 benchmark instances of the two tasks defined in Section[3\.2](https://arxiv.org/html/2605.31410#S3.SS2)through a rule\-grounded, expert\-verified annotation pipeline\.

#### Knowledge base\.

Annotation is anchored to a curated knowledge base𝒦\\mathcal\{K\}that, for each conditionhh, specifies the beneficial, limit, and avoid food categories together with the supporting clinical references\.𝒦\\mathcal\{K\}is compiled from established dietary guidelines \(AHA, NIH, Harvard Nutrition Source, American Liver Foundation\) and reviewed by nutrition experts; the full schema is given in Appendix[D](https://arxiv.org/html/2605.31410#A4)\.

#### Canonical ingredient vocabulary\.

Ingredient lines are normalized against a canonical vocabulary𝒱\\mathcal\{V\}aligned with the food and concept entries of𝒦\\mathcal\{K\}\(e\.g\., “skim milk”→\\to“low fat milk”; “granulated sugar”→\\to“sugar”\), so that ingredient citations produced during annotation can be compared directly against the condition\-level rule sets\.

#### Structured annotation\.

For each dishdd, we produce a structured annotation

a​\(d\)=\(Rec​\(d\),NotRec​\(d\)\),a\(d\)=\(\\mathrm\{Rec\}\(d\),\\ \\mathrm\{NotRec\}\(d\)\),whereRec​\(d\)\\mathrm\{Rec\}\(d\)andNotRec​\(d\)\\mathrm\{NotRec\}\(d\)each contain a list of entries of the form\(h,Eh\)\(h,E\_\{h\}\): a health conditionhhtogether with the canonical ingredientsEhE\_\{h\}\(drawn from𝒱\\mathcal\{V\}\) that drive the decision forhh\. A condition cannot appear in bothRec​\(d\)\\mathrm\{Rec\}\(d\)andNotRec​\(d\)\\mathrm\{NotRec\}\(d\); conditions for which the dish is neutral are simply omitted\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x3.png)Figure 3:Geographical Distribution of collected recipes
#### Validation and expert verification\.

Each candidate annotation is cross\-checked against𝒦\\mathcal\{K\}through three rule families ingredient concept consistency, condition\-specific nutrient\-threshold vetoes, and ingredient\-fallback rules; annotations on which the rule\-based check and the LLM proposal disagree on theRec\\mathrm\{Rec\}/NotRec\\mathrm\{NotRec\}side of anyhhare dropped\. Nutrition experts then review each surviving candidate and confirm or correct every\(h,Eh\)\(h,E\_\{h\}\)entry; the retained annotations form the poolAA\. Detailed rules and exclusion criteria are listed in Appendices[E](https://arxiv.org/html/2605.31410#A5)and[F](https://arxiv.org/html/2605.31410#A6)\.

#### Benchmark instantiation\.

Dish\-level task \(1,500 instances\)\.Each instance is sampled from a recipe witha​\(d\)∈Aa\(d\)\\in Aby drawing a focal conditionh∈Rec​\(d\)∪NotRec​\(d\)h\\in\\mathrm\{Rec\}\(d\)\\cup\\mathrm\{NotRec\}\(d\)and setting the labelyyaccording to which sidehhfalls on; the ingredient evidenceEhE\_\{h\}is retained as an internal annotation but is not exposed to the model at evaluation time\. The split is approximately label\-balanced\.

Comparative task \(1,000 instances\)\.Each instance pairs one to three target conditions with four candidate dishes drawn fromAA, spanning four distractor types: condition discrimination, eligible\-dish ranking, ingredient\-coverage discrimination, and combined\-constraint discrimination\. The gold record provides the full rankingπ⋆\\pi^\{\\star\}, per\-option suitability labels, and avoid\-list violation flags; answer positions are balanced across\{A,B,C,D\}\\\{A,B,C,D\\\}\.

Both splits are balanced across the 13 conditions; per\-condition coverage and split statistics \(label balance, distractor distribution, answer\-position balance\) are reported in Appendix[G](https://arxiv.org/html/2605.31410#A7)\(Tables[4](https://arxiv.org/html/2605.31410#A7.T4)and[5](https://arxiv.org/html/2605.31410#A7.T5)\)\.

## 4Experiments

### 4\.1Models

We evaluate five instruction\-tuned vision language models spanning two deployment tiers\. The*closed\-source frontier*tier consists of GPT\-5\.4\(OpenAI,[2025](https://arxiv.org/html/2605.31410#bib.bib56)\), Claude Sonnet 4\.6\(Anthropic,[2026](https://arxiv.org/html/2605.31410#bib.bib57)\), and Gemini 2\.5 Pro\(Comaniciet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib58)\), queried through their hosted APIs \(parameter counts undisclosed\)\. The*open\-weight*tier consists of Qwen3\-VL\-8B\-Instruct\(Baiet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib59)\)and Gemma\-3\-12B\-IT\(Kamathet al\.,[2025](https://arxiv.org/html/2605.31410#bib.bib60)\), served locally via vLLM; the full serving stack and decoding settings are in Appendix[J](https://arxiv.org/html/2605.31410#A10)\. All five accept interleaved image and text inputs and emit decisions under the same structured output contract\.

### 4\.2Prompting Methods

We evaluate each model under four prompting modes that share the same input schema\(I,G,h\)\(I,G,h\)and the same decision\-bearing outputs: a binary recommendation and a list of condition\-relevant ingredients\. CoT modes additionally emit an intermediate reasoning trace, ignored by the scorer\. Full prompt templates and a worked knowledge injection reference are in Appendix[L](https://arxiv.org/html/2605.31410#A12)\.

#### Baseline\.

The model is given\(I,G,h\)\(I,G,h\)and emits the decision directly, isolating its zero\-shot multimodal dietary prior\.

#### Chain\-of\-Thought \(CoT\)\.

The system prompt requires the model to first emit a short array of intermediate reasoning steps condition identification, ingredient screening, evidence aggregation before committing to the decision\. The three\-step scaffold is held fixed across models and is more constrained than free\-form CoT\(Weiet al\.,[2022](https://arxiv.org/html/2605.31410#bib.bib51)\)\.

#### Knowledge Injection \(KI\)\.

For each target conditionhh, we append a fixed condition\-level referenceℛh\\mathcal\{R\}\_\{h\}from the curated knowledge base𝒦\\mathcal\{K\}of Section[3\.3](https://arxiv.org/html/2605.31410#S3.SS3)\. The reference lists beneficial and avoid food categories forhh\. We select the reference deterministically from the condition named in the benchmark prompt\.ℛh\\mathcal\{R\}\_\{h\}carries no per\-dish labels or rationales, so the model must still apply the rules to the observed ingredients\.

#### CoT \+ KI\.

Combines both: the model readsℛh\\mathcal\{R\}\_\{h\}, emits a reasoning trace that compares the condition reference with the observed ingredients, then decides\. This tests whether reasoning supervision and a fixed condition\-level reference are complementary\.

### 4\.3Evaluation Metrics

#### Task 1: Dish\-level suitability\.

*Decision accuracy*is the fraction of instances whose predicted binary label matches the gold label\.*Rationale macro\-F1*and*micro\-F1*measure overlap between the model\-cited condition\-relevant ingredients and the expert\-verified gold rationale for the focal condition: macro\-F1 averages per\-condition F1 with equal weight across the 13 conditions; micro\-F1 pools true positives, false positives, and false negatives across all instances\.

#### Task 2: Comparative ranking\.

*Top\-1 accuracy*is the fraction of queries whose top\-ranked dish matches the gold most\-suitable dish inπ⋆\\pi^\{\\star\}\.*Mean Reciprocal Rank*\(MRR\) is

MRR=1\|Q\|​∑i=1\|Q\|1ranki,\\mathrm\{MRR\}=\\frac\{1\}\{\|Q\|\}\\sum\_\{i=1\}^\{\|Q\|\}\\frac\{1\}\{\\mathrm\{rank\}\_\{i\}\},\(1\)whereranki\\mathrm\{rank\}\_\{i\}is the position of the gold dish in the returned ordering, taken as∞\\inftywhen absent\. MRR awards partial credit for near\-miss rankings\.

#### Cross\-task consistency \(CTC\)\.

For each query, we elicit pairwise\(rec,not\_rec\)\(\\textsc\{rec\},\\textsc\{not\\\_rec\}\)orderings from the model’s own Task 1 per\-option decisions and check whether its Task 2 ranking respects them\. CTC is the micro\-averaged fraction of such pairs in agreement, with chance=0\.5=0\.5\. As both inputs come from the same model, CTC measures internal coherence independent of ground truth\.

## 5Results

### 5\.1Main Results

Table 1:Main results across five vision\-language models and four prompting modes\.Cross Task Consistency \(CTC\)is the intersection\-pool micro\-averaged fraction of discriminating\(rec,not\_rec\)\(\\textsc\{rec\},\\textsc\{not\\\_rec\}\)pairs elicited from the model’s own per\-option decisions that are ordered correctly in the model’s own Task 2 ranking; chance=0\.50=0\.50, ceiling=1\.00=1\.00\. For each model, the four modes share a common pool of questions on which all four modes emitted valid decisions and rankings\. Best value per column within each model block inbold\.Table[1](https://arxiv.org/html/2605.31410#S5.T1)reports both tasks under all four prompting modes for the five models\. We read the results through the benchmark design: whether models can decide suitability, ground that decision in ingredient evidence, compare alternatives, and remain consistent across the two settings\.

#### Verdicts are easier than rationales\.

The clearest signal in FAM\-Bench is a persistent gap between suitability verdicts and ingredient\-level grounding\. The best Task 1 accuracy reaches82\.80%82\.80\\%\(Gemini 2\.5 Pro with KI\), yet the best rationale macro\-F1 is only0\.26140\.2614\(GPT\-5\.4 with CoT\+KI\)\. This gap matters because the rationale targets are not post\-hoc explanations: they are the expert\-verified ingredients that drive the condition label in Section[3\.3](https://arxiv.org/html/2605.31410#S3.SS3)\. Current VLMs can often answer whether a dish is suitable, but they much less reliably identify the ingredients that make it suitable or unsafe\.

#### CoT and knowledge injection address different bottlenecks\.

CoT and KI improve different parts of the task\. CoT mainly helps models cite condition\-relevant ingredients, while KI mainly improves the binary suitability decision by supplying condition rules\. Combining them is therefore the strongest overall setting, but the combined prompt still narrows rather than closes the verdict–rationale gap\.

#### Comparative ranking is the harder, more realistic setting\.

Task 2 asks models to choose among four candidate dishes, matching the recommendation setting introduced in Section[3\.2](https://arxiv.org/html/2605.31410#S3.SS2.SSS0.Px2)\. Its top\-1 accuracy ranges from 29\.30% to 41\.50% across all cells, above the 25% chance level but still modest, with MRR in \[0\.5335, 0\.6258\]\. Thus, even models that perform well on binary suitability struggle to rank alternatives by condition\-specific utility\. This is especially visible for the open\-weight VLMs, whose Task 2 gains remain small even when condition knowledge is injected\.

#### Consistency reveals failures not captured by accuracy\.

CTC measures whether a model’s Task 2 ranking respects its own Task 1 per\-option judgments\. Most cells are high, but Claude Sonnet 4\.6 baseline is the exception: its Task 2 accuracy is near chance and its CTC is only0\.62810\.6281\. CoT raises the same model’s CTC to0\.93060\.9306, suggesting that the reasoning scaffold does not merely change isolated answers; it makes the model’s absolute and comparative judgments more coherent\.

#### Non\-expert human evaluation\.

To contextualize task difficulty, we collect a small non\-expert human baseline on 100 held\-out items\. Participants achieve63\.0%63\.0\\%accuracy on dish\-level suitability and37\.0%37\.0\\%top\-1 accuracy on comparative ranking, above the respective chance levels of50%50\\%and25%25\\%\. As shown in Figure[4](https://arxiv.org/html/2605.31410#S5.F4), the gap between Task 1 and Task 2 also appears in human judgments: deciding whether a single dish is suitable is easier than ranking alternatives under the same health constraint\. The VLM comparison shows a similar pattern, with recipe text contributing more than image evidence and multimodal inputs offering only limited gains over text alone\. This supports the benchmark design: comparative dish analysis is the more realistic and more difficult Food\-as\-Medicine setting\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x4.png)Figure 4:Human evaluation and VLM modality comparison\.Accuracy for chance baselines, human test participants, VLM with recipe text only, VLM with image only, and VLM with text & image\. VLM values are averaged across the five evaluated models and four prompting modes\. Chance is50%50\\%for Task 1 binary suitability and25%25\\%for Task 2 four\-way ranking\.Task 1 \(Acc\.\)Task 2 \(Acc\.\)ModeT\+IT\-onlyI\-only𝚫img\\boldsymbol\{\\Delta\_\{\\text\{img\}\}\}𝚫txt\\boldsymbol\{\\Delta\_\{\\text\{txt\}\}\}T\+IT\-onlyI\-only𝚫img\\boldsymbol\{\\Delta\_\{\\text\{img\}\}\}𝚫txt\\boldsymbol\{\\Delta\_\{\\text\{txt\}\}\}Qwen3\-VL\-8Bbaseline73\.0074\.4764\.67−1\.47\-1\.47\+8\.3331\.3031\.0024\.80\+0\.30\+6\.50cot70\.0770\.2065\.00−0\.13\-0\.13\+5\.0729\.3031\.3024\.90−2\.00\-2\.00\+4\.40ki75\.2075\.0066\.80\+0\.20\+8\.4031\.8032\.3023\.60−0\.50\-0\.50\+8\.20cot\+ki75\.7375\.4066\.47\+0\.33\+9\.2630\.8031\.3023\.80−0\.50\-0\.50\+7\.00Gemma\-3\-12Bbaseline67\.0767\.6064\.07−0\.53\-0\.53\+3\.0030\.6032\.0024\.10−1\.40\-1\.40\+6\.50cot68\.8065\.8762\.60\+2\.93\+6\.2031\.3031\.9024\.20−0\.60\-0\.60\+7\.10ki69\.0772\.6766\.33−3\.60\-3\.60\+2\.7432\.3031\.6023\.30\+0\.70\+9\.00cot\+ki71\.8069\.2065\.20\+2\.60\+6\.6032\.3032\.1025\.50\+0\.20\+6\.80GPT\-5\.4baseline77\.6778\.3369\.33−0\.66\-0\.66\+8\.3432\.1035\.3029\.90−3\.20\-3\.20\+2\.20cot78\.6779\.0761\.93−0\.40\-0\.40\+16\.7433\.1034\.8030\.00−1\.70\-1\.70\+3\.10ki80\.4782\.0770\.13−1\.60\-1\.60\+10\.3436\.9035\.9030\.50\+1\.00\+6\.40cot\+ki82\.4082\.4767\.53−0\.07\-0\.07\+14\.8737\.4037\.0031\.20\+0\.40\+6\.20Claude Sonnet 4\.6baseline77\.5377\.8068\.00−0\.27\-0\.27\+9\.5330\.9035\.0029\.20−4\.10\-4\.10\+1\.70cot77\.1377\.4767\.07−0\.34\-0\.34\+10\.0635\.7037\.8032\.80−2\.10\-2\.10\+2\.90ki80\.8081\.0768\.40−0\.27\-0\.27\+12\.4035\.1038\.4030\.70−3\.30\-3\.30\+4\.40cot\+ki81\.2080\.6768\.67\+0\.53\+12\.5341\.5039\.3032\.50\+2\.20\+9\.00Gemini 2\.5 Probaseline80\.8081\.6768\.47−0\.87\-0\.87\+12\.3332\.7034\.7031\.60−2\.00\-2\.00\+1\.10cot80\.2079\.7364\.20\+0\.47\+16\.0034\.6035\.2032\.10−0\.60\-0\.60\+2\.50ki82\.8083\.1368\.93−0\.33\-0\.33\+13\.8739\.8038\.5033\.40\+1\.30\+6\.40cot\+ki82\.6081\.2066\.53\+1\.40\+16\.0741\.3039\.7032\.90\+1\.60\+8\.40Table 2:Input\-modality ablation across Tasks 1 and 2\.Decision accuracy \(%\)\.T\+I: full multimodal input \(from Table[1](https://arxiv.org/html/2605.31410#S5.T1)\);T\-only: recipe text only;I\-only: image only\.𝚫img=\\boldsymbol\{\\Delta\_\{\\text\{img\}\}\}=T\+I−\-T\-only \(marginal contribution of the image; negative means the image*hurts*\)\.𝚫txt=\\boldsymbol\{\\Delta\_\{\\text\{txt\}\}\}=T\+I−\-I\-only \(marginal contribution of the recipe text\)\.Best per\-model column inboldfor T\+I, T\-only, and I\-only\.
#### Per\-condition variation\.

Figure[5](https://arxiv.org/html/2605.31410#S5.F5)decomposes baseline Task 1 accuracy across the 13 health conditions: per\-condition accuracy varies substantially within each model, and the five models share a broadly similar condition profile\. CoT, RAG, and CoT\+RAG breakdowns are in Appendix[H](https://arxiv.org/html/2605.31410#A8)\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x5.png)Figure 5:Task 1 per\-condition decision accuracy across the five models under theBaselineprompting mode \(text \+ image, Q2\-v3 1500\-sample evaluation\)\. See Appendix[H](https://arxiv.org/html/2605.31410#A8)for the corresponding CoT, KI, and CoT\+KI plots\.

### 5\.2Input\-Modality Ablation

To probe how each input modality contributes to performance, we re\-run the four prompting modes on GPT\-5\.4 and Qwen3\-VL\-8B\-Instruct under two modality\-stripped conditions:*text\-only*\(dish image dropped\) and*image\-only*\(recipe text dropped\)\. Table[2](https://arxiv.org/html/2605.31410#S5.T2)shows that recipe text carries most of the usable evidence\. Dropping the image preserves nearly all text\+image accuracy, and sometimes slightly improves it; dropping the recipe text causes a much larger drop\. This pattern is consistent with the annotation design: most gold rationales are ingredient\-level constraints, while images mainly provide preparation cues such as frying, creaminess, or visible portion composition\. Current VLMs do not yet extract enough additional visual evidence to offset noise from the image channel\. Under image\-only inputs, CoT can amplify unsupported inferences, while knowledge injection partially recovers performance by supplying condition knowledge that the image itself does not contain\.

### 5\.3Error Analysis

The Claude Sonnet 4\.6 baseline posts the lowest rationale macro\-F1 in Table[1](https://arxiv.org/html/2605.31410#S5.T1)\(0\.05830\.0583at77\.53%77\.53\\%accuracy; cross\-task consistency0\.62810\.6281, near the chance floor of0\.50\.5\), exposing a pattern that recurs across the baseline runs: high accept/reject accuracy can co\-exist with near\-random ingredient\-level rationales\. We inspect this error pool and identify three failure modes; full case walkthroughs are in Appendix[I](https://arxiv.org/html/2605.31410#A9)\.

#### Three failure modes\.

\(i\)Single\-axis risk checkingingredients are evaluated along the sodium axis only; a “no salt added” or “low sodium” cue is read as clearance even when the condition also restricts potassium, phosphorus, or saturated fat \(Cases 1\-2\)\. \(ii\)Rationale undercoveragethe rationale fills with plant\-derived items and skips the animal ingredients that actually drive the gold label, even when the final verdict is nonetheless correct \(Cases 1–2\)\. \(iii\)Missing rule knowledgea clearly visible ingredient is not flagged because the model’s working knowledge of the condition lacks the relevant rule, e\.g\., the model fails to flag alcohol despite its known adverse effect on bone health in osteoporosis \(Case 3\)\. This mode is distinct from \(ii\): in \(ii\) the model knows the rule but fails to surface the relevant ingredient, whereas here the rule itself is missing from the model’s working knowledge\.

## 6Conclusion and Future Work

FAM\-Bench reframes food AI evaluation as condition\-aware dietary decision making\. It provides 2,500 nutrition\-expert\-verified instances across 13 conditions and tests two capabilities: dish\-level suitability assessment and comparative dish analysis\. Across five vision\-language models and four prompting modes, models predict suitability more reliably than they ground those verdicts in the ingredients that justify them; ranking condition\-specific alternatives remains harder still\.

These findings motivate health\-oriented food assistants that move beyond dish recognition and nutrient estimation\. Future systems should integrate food evidence with portion size, longitudinal diet, medical history, medications, laboratory values, allergies, and clinician\-prescribed constraints, while supporting rather than replacing clinicians and registered dietitians\. FAM\-Bench offers a first step toward evaluating such grounded, clinically supervised dietary reasoning\.

## Limitations

FAM\-Bench provides a controlled testbed for condition\-aware dietary reasoning, but its design has several scope limits\.

#### Hidden ingredients and preparation ambiguity\.

FAM\-Bench relies on dish images, ingredient lists, and available recipe metadata\. These inputs capture many condition\-relevant cues, such as added sugar, high\-sodium ingredients, refined carbohydrates, saturated\-fat sources, and trigger ingredients\. However, recipes may omit hidden sodium or sugar in sauces, unspecified cooking fats, brand\-specific packaged ingredients, or preparation details that affect suitability\. Images also cannot reliably reveal hidden ingredients or cooking methods\. We flag or exclude cases with insufficient evidence, but the benchmark inherits the limits of real recipe data\.

#### Portion size and quantitative nutrition\.

Dietary suitability is dose\-dependent\. A small portion of a rich dish may be acceptable under some conditions, while a larger portion of the same dish may exceed relevant limits for sodium, saturated fat, added sugar, carbohydrates, or calories\. For example, one slice of cake or a small serving of butter chicken may raise a different dietary judgment than consuming the full dish\. Although we use available nutrition tables and nutrient\-threshold checks during validation, many recipe sources provide incomplete or inconsistent serving\-size information\. FAM\-Bench therefore evaluates ingredient\- and preparation\-grounded suitability, not precise portion\-level or dose\-dependent nutrition assessment\.

#### Geographic and cultural coverage\.

The corpus draws from multiple web domains but remains U\.S\.\-heavy\. This may bias the benchmark toward U\.S\. ingredient terminology, meal formats, dietary patterns, and guideline assumptions\. Performance on FAM\-Bench should therefore not be read as uniform competence across cuisines, cultures, or regional dietary practices\. Future versions would expand multilingual sources, non\-U\.S\. guidelines, and culturally specific food\-health relationships\.

## Ethics Statements

The authors used AI assistants for language polishing, wording suggestions, LaTeX editing, and code/debugging support during manuscript preparation\. All scientific claims, analyses, and final text were reviewed and approved by the authors, who retain full responsibility for the paper\.

The human comparison in this work was completed by student annotators following the benchmark task format\. These judgments are used only as a non\-expert human baseline and should not be interpreted as professional medical, clinical nutrition, or dietetic advice\.

FAM\-Bench is intended solely for research on condition\-aware food reasoning\. The dataset, labels, and model outputs must not be used for real medical advice, diagnosis, treatment, or personalized dietary recommendations without review by qualified health professionals\.

## References

- Knowledge graph and large language model synergy for food safety: approaches and perspectives\.Trends in Food Science & Technology,pp\. 105592\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- Anthropic \(2026\)Introducing Claude Sonnet 4\.6\.Note:Anthropic Release Announcement[https://www\.anthropic\.com/news/claude\-sonnet\-4\-6](https://www.anthropic.com/news/claude-sonnet-4-6)Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§4\.1](https://arxiv.org/html/2605.31410#S4.SS1.p1.1)\.
- S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, Q\. Huang, J\. Huang, F\. Huang, B\. Hui, S\. Jiang, Z\. Li, M\. Li, M\. Li, K\. Li, Z\. Lin, J\. Lin, X\. Liu, J\. Liu, C\. Liu, Y\. Liu, D\. Liu, S\. Liu, D\. Lu, R\. Luo, C\. Lv, R\. Men, L\. Meng, X\. Ren, X\. Ren, S\. Song, Y\. Sun, J\. Tang, J\. Tu, J\. Wan, P\. Wang, P\. Wang, Q\. Wang, Y\. Wang, T\. Xie, Y\. Xu, H\. Xu, J\. Xu, Z\. Yang, M\. Yang, J\. Yang, A\. Yang, B\. Yu, F\. Zhang, H\. Zhang, X\. Zhang, B\. Zheng, H\. Zhong, J\. Zhou, F\. Zhou, J\. Zhou, Y\. Zhu, and K\. Zhu \(2025\)Qwen3\-VL technical report\.arXiv preprint arXiv:2511\.21631\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§4\.1](https://arxiv.org/html/2605.31410#S4.SS1.p1.1)\.
- S\. Bedi, Y\. Liu, L\. Orr\-Ewing, D\. Dash, S\. Koyejo, A\. Callahan, J\. A\. Fries, M\. Wornow, A\. Swaminathan, L\. S\. Lehmann,et al\.\(2025\)Testing and evaluation of health care applications of large language models: a systematic review\.Jama333\(4\),pp\. 319–328\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Bień, M\. Gilski, M\. Maciejewska, W\. Taisner, D\. Wiśniewski, and A\. Lawrynowicz \(2020\)RecipeNLG: a cooking recipes dataset for semi\-structured text generation\.InProceedings of the 13th international conference on natural language generation,pp\. 22–28\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Bossard, M\. Guillaumin, and L\. Van Gool \(2014\)Food\-101–mining discriminative components with random forests\.InEuropean conference on computer vision,pp\. 446–461\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- F\. L\. Bryan, W\. H\. Organization,et al\.\(1992\)Hazard analysis critical control point evaluations: a guide to identifying hazards and assessing risks associated with food preparation and storage\.World Health Organization\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- Centers for Disease Control and Prevention \(2025\)Fast facts: health and economic costs of chronic conditions\.Note:[https://www\.cdc\.gov/chronic\-disease/data\-research/facts\-stats/index\.html](https://www.cdc.gov/chronic-disease/data-research/facts-stats/index.html)Accessed: 2026\-05\-11Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p2.1)\.
- Y\. Chen, A\. Subburathinam, C\. Chen, and M\. J\. Zaki \(2021\)Personalized food recommendation as constrained question answering over a large\-scale food knowledge graph\.InProceedings of the 14th ACM international conference on web search and data mining,pp\. 544–552\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§4\.1](https://arxiv.org/html/2605.31410#S4.SS1.p1.1)\.
- S\. Y\. Feng, V\. Khetan, B\. Sacaleanu, A\. Gershman, and E\. Hovy \(2023\)CHARD: clinical health\-aware reasoning across dimensions for text generation models\.InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics,pp\. 313–327\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Gao, X\. Zhao, D\. Xia, Z\. Zhou, R\. Yang, J\. Lu, H\. Jiang, C\. Park, and I\. Li \(2025\)Healthgenie: empowering users with healthy dietary guidance through knowledge graph and large language models\.arXiv preprint arXiv:2504\.14594\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Gong, H\. Sha, R\. Liu, T\. Wu, B\. Liu, and H\. Wang \(2024\)Integrating tcm’s “one root of medicine and food” principle into dietary recommendations with retrieval\-augmented llms\.InChina Health Information Processing Conference,pp\. 251–267\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Hosseinian, A\. D\. Zahedani, U\. Mansoor, N\. Hashemi, and M\. Woodward \(2025\)January food benchmark \(jfb\): a public benchmark dataset and evaluation suite for multimodal food analysis\.arXiv preprint arXiv:2508\.09966\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Hua, M\. P\. Dhaliwal, L\. Pullela, R\. Burke, and Y\. Qin \(2024\)Nutribench: a dataset for evaluating large language models on nutrition estimation from meal descriptions\.arXiv preprint arXiv:2407\.12843\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Jacxsens, M\. Uyttendaele, F\. Devlieghere, J\. Rovira, S\. O\. Gomez, and P\. Luning \(2010\)Food safety performance indicators to benchmark food safety output of food safety management systems\.International Journal of Food Microbiology141,pp\. S180–S187\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Jin, J\. Zhang, X\. Zhang, Z\. Tian, F\. Jiang, G\. Yin, W\. Lin, Y\. Liu, and R\. Yan \(2026\)DiningBench: a hierarchical multi\-view benchmark for perception and reasoning in the dietary domain\.arXiv preprint arXiv:2604\.10425\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard,et al\.\(2025\)Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§4\.1](https://arxiv.org/html/2605.31410#S4.SS1.p1.1)\.
- S\. Khamesian, A\. Arefeen, S\. M\. Carpenter, and H\. Ghasemzadeh \(2025\)NutriGen: personalized meal plan generator leveraging large language models to enhance dietary and nutritional adherence\.In2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society \(EMBC\),pp\. 1–7\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Kim, R\. Venkataramanan, and A\. Sheth \(2024\)A survey on food ingredient substitutions\.arXiv preprint arXiv:2501\.01958\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1)\.
- T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa \(2022\)Large language models are zero\-shot reasoners\.Vol\.35,pp\. 22199–22213\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Le Vallée and S\. Charlebois \(2015\)Benchmarking global food safety performances: the era of risk intelligence\.Journal of food protection78\(10\),pp\. 1896–1913\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Vol\.33,pp\. 9459–9474\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- K\. J\. Li, S\. Balloccu, O\. Dušek, and E\. Reiter \(2025\)When llms can’t help: real\-world evaluation of llms in nutrition\.InProceedings of the 18th International Natural Language Generation Conference,pp\. 753–779\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1)\.
- W\. Li, C\. Zhang, J\. Li, Q\. Peng, R\. Tang, L\. Zhou, W\. Zhang, G\. Hu, Y\. Yuan, A\. Søgaard,et al\.\(2024\)FoodieQA: a multimodal dataset for fine\-grained understanding of chinese food culture\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 19077–19095\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px2.p1.1)\.
- J\. Liu, P\. Zhou, Y\. Hua, D\. Chong, Z\. Tian, A\. Liu, H\. Wang, C\. You, Z\. Guo, L\. Zhu,et al\.\(2023\)Benchmarking large language models on cmexam\-a comprehensive chinese medical exam dataset\.Advances in Neural Information Processing Systems36,pp\. 52430–52452\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1)\.
- W\. Luo, X\. Wen, T\. Huang, H\. Wang, Z\. Xiang, C\. Xiao, K\. Gligorić, and M\. Chen \(2026\)Cooking up risks: benchmarking and reducing food safety risks in large language models\.arXiv preprint arXiv:2604\.01444\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Marın, A\. Biswas, F\. Ofli, N\. Hynes, A\. Salvador, Y\. Aytar, I\. Weber, and A\. Torralba \(2021\)Recipe1m\+: a dataset for learning cross\-modal embeddings for cooking recipes and food images\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(1\),pp\. 187–203\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. Hajishirzi \(2023\)Factscore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12076–12100\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- F\. Mohbat and M\. J\. Zaki \(2024\)Llava\-chef: a multi\-modal generative model for food recipes\.InProceedings of the 33rd ACM International Conference on Information and Knowledge Management,pp\. 1711–1721\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Mozaffarian \(2016\)Dietary and policy priorities for cardiovascular disease, diabetes, and obesity: a comprehensive review\.Circulation133\(2\),pp\. 187–225\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p2.1)\.
- J\. Muncke, M\. Touvier, L\. Trasande, and M\. Scheringer \(2025\)Health impacts of exposure to synthetic chemicals in food\.Nature medicine31\(5\),pp\. 1431–1443\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- H\. Nori, Y\. T\. Lee, S\. Zhang, D\. Carignan, R\. Edgar, N\. Fusi, N\. King, J\. Larson, Y\. Li, W\. Liu,et al\.\(2023\)Can generalist foundation models outcompete special\-purpose tuning? case study in medicine\.arXiv preprint arXiv:2311\.16452\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- OpenAI \(2025\)Introducing GPT\-5\.4\.Note:OpenAI Release Announcement[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§4\.1](https://arxiv.org/html/2605.31410#S4.SS1.p1.1)\.
- A\. Pal, L\. K\. Umapathi, and M\. Sankarasubbu \(2023\)Med\-halt: medical domain hallucination test for large language models\.InProceedings of the 27th Conference on Computational Natural Language Learning \(CoNLL\),pp\. 314–334\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- H\. Pekmezci, S\. Sipahi, and B\. Başaran \(2025\)Health risk assessment of dietary chemical exposures: a comprehensive review\.Foods14\(23\),pp\. 4133\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- H\. Sha, F\. Gong, B\. Liu, R\. Liu, H\. Wang, and T\. Wu \(2025\)Leveraging retrieval\-augmented large language models for dietary recommendations with traditional chinese medicine’s medicine food homology: algorithm development and validation\.JMIR Medical Informatics13\(1\),pp\. e75279\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.\(2023\)Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. Thames, A\. Karpur, W\. Norris, F\. Xia, L\. Panait, T\. Weyand, and J\. Sim \(2021\)Nutrition5k: towards automatic nutritional understanding of generic food\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 8903–8911\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Veeramreddy, A\. K\. Pradhan, S\. Ghanta, L\. Rachakonda, and S\. P\. Mohanty \(2024\)NUTRIVISION: a system for automatic diet management in smart healthcare\.arXiv preprint arXiv:2409\.20508\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- K\. G\. Volpp, S\. A\. Berkowitz, S\. V\. Sharma, C\. A\. Anderson, L\. C\. Brewer, M\. S\. Elkind, C\. D\. Gardner, J\. E\. Gervis, R\. A\. Harrington, M\. Herrero,et al\.\(2023\)Food is medicine: a presidential advisory from the american heart association\.Circulation148\(18\),pp\. 1417–1439\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p2.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Vol\.35,pp\. 24824–24837\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p6.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2605.31410#S4.SS2.SSS0.Px2.p1.1)\.
- W\. C\. Willett \(1994\)Diet and health: what should we eat?\.Science264\(5158\),pp\. 532–537\.Cited by:[§1](https://arxiv.org/html/2605.31410#S1.p2.1)\.
- A\. Wróblewska, A\. Kaliska, M\. Pawłowski, D\. Wiśniewski, W\. Sosnowski, and A\. Ławrynowicz \(2022\)TASTEset–recipe dataset and food entities recognition benchmark\.arXiv preprint arXiv:2204\.07775\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Xiong, Q\. Jin, Z\. Lu, and A\. Zhang \(2024\)Benchmarking retrieval\-augmented generation for medicine\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 6233–6251\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px6.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Yagcioglu, A\. Erdem, E\. Erdem, and N\. Ikizler\-Cinbis \(2018\)Recipeqa: a challenge dataset for multimodal comprehension of cooking recipes\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 1358–1368\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Yang, E\. Khatibi, N\. Nagesh, M\. Abbasian, I\. Azimi, R\. Jain, and A\. M\. Rahmani \(2024\)ChatDiet: empowering personalized nutrition\-oriented food recommender chatbots through an llm\-augmented framework\.Smart Health32,pp\. 100465\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- W\. Yue, X\. Wang, W\. Zhu, M\. Guan, H\. Zheng, P\. Wang, C\. Sun, and X\. Ma \(2024\)Tcmbench: a comprehensive benchmark for evaluating large language models in traditional chinese medicine\.arXiv preprint arXiv:2406\.01126\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1)\.
- Y\. Zhang \(2021\)Diet according to traditional chinese medicine for health and longevity\.InNutrition, Food and Diet in Ageing and Longevity,pp\. 331–356\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1)\.
- Z\. Zhang, Y\. Li, N\. H\. L\. Le, Z\. Wang, T\. Ma, V\. Galassi, K\. Murugesan, N\. Moniz, W\. Geyer, N\. V\. Chawla,et al\.\(2025\)Ngqa: a nutritional graph question answering benchmark for personalized health\-aware nutritional reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5934–5966\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2605.31410#S1.p4.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.
- P\. Zhou, W\. Min, C\. Fu, Y\. Jin, M\. Huang, X\. Li, S\. Mei, and S\. Jiang \(2025\)FoodSky: a food\-oriented large language model that can pass the chef and dietetic examinations\.Patterns6\(5\)\.Cited by:[Appendix A](https://arxiv.org/html/2605.31410#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2605.31410#S2.SS0.SSS0.Px2.p1.1)\.

## Appendix

## Appendix ARelated Work

#### Food recognition and recipe understanding\.

Food AI benchmarks have historically focused on recognizing dishes and structuring recipe content\. Food\-101 established large\-scale visual dish classificationBossardet al\.\([2014](https://arxiv.org/html/2605.31410#bib.bib14)\), while Recipe1M\+ paired recipes with food images for cross\-modal retrieval and representation learningMarınet al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib15)\)\. RecipeQA extended food understanding to procedural reasoning over recipe text and imagesYagciogluet al\.\([2018](https://arxiv.org/html/2605.31410#bib.bib16)\), and TASTEset introduced structured extraction of ingredients, quantities, cooking processes, and ingredient properties from recipesWróblewskaet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib17)\)\. Generation\-oriented corpora such as RecipeNLG further enabled controlled recipe synthesis at scaleBieńet al\.\([2020](https://arxiv.org/html/2605.31410#bib.bib37)\)\. These resources provide the visual and linguistic foundations for food understanding, but they remain descriptive: they identify what a dish is, what it contains, or how it is prepared, without evaluating whether the dish is appropriate for a specific health condition\.

#### Nutrition estimation and dietary\-domain benchmarks\.

A second wave of benchmarks shifts from food recognition toward nutrient estimation and broader dietary reasoning\. Nutrition5k provides visual and nutritional annotations for real dishesThameset al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib18)\), and NutriBench evaluates large language models on calorie and macronutrient estimation from natural\-language meal descriptions, including a downstream blood\-glucose simulation for type\-1 diabetesHuaet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib19)\)\. Recent multimodal benchmarks such as the January Food Benchmark and DiningBench broaden the setting to ingredient reasoning, nutrition estimation, and dietary\-domain VQAHosseinianet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib20)\); Jinet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib21)\), while FoodieQA curates multimodal Chinese food\-culture QA but explicitly excludes nutrition and clinical contentLiet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib38)\)\. These benchmarks make food AI more nutritionally grounded, yet their targets remain primarily descriptive calories, macronutrients, ingredients, or food attributes rather than evaluating whether those food properties imply condition\-specific dietary suitability\.

#### Personalized nutrition and food\-domain LLMs\.

A parallel line of work delivers system contributions: LLM\-based pipelines that issue dietary advice\. ChatDiet couples user data with an LLM front\-end for personalized nutrition counselingYanget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib22)\); NutriGen generates calorie\-bounded meal plans under user preferencesKhamesianet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib39)\); HealthGenie grafts a knowledge graph onto an LLM for chronic\-condition\-aware recommendationGaoet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib23)\); and NutriVision integrates vision with conversational AI for smart\-healthcare diet managementVeeramreddyet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib40)\)\. Constraint\-aware approaches further survey ingredient substitution for chronic conditionsKimet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib41)\)\. Food\-specialized models adapt LLMs and vision–language models to culinary reasoning and dietetic knowledge: LLaVA\-Chef targets multimodal recipe generationMohbat and Zaki \([2024](https://arxiv.org/html/2605.31410#bib.bib30)\), and FoodSky reports strong performance on chef and dietetic examinationsZhouet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib31)\)\. These contributions highlight the practical value of LLMs for nutrition support, but each is evaluated on its own task suite, making cross\-system comparison and safety claims difficult to verify against a common standard\.

#### Health\-aware nutritional reasoning\.

The closest benchmark direction is personalized health\-aware nutritional reasoning\. NGQA formulates nutrition reasoning as graph question answering over users, foods, nutrients, and medical conditions, explicitly connecting dietary choices to health profilesZhanget al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib26)\)\. Related lines of work study constrained QA over food knowledge graphsChenet al\.\([2021](https://arxiv.org/html/2605.31410#bib.bib27)\), clinical health\-aware text generationFenget al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib28)\), and knowledge\-graph grounding for food\-safety reasoningAnet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib29)\)\. In the Chinese\-medicine setting, CMExam and TCMBench probe traditional\-medicine knowledge through exam\-style MCQsLiuet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib42)\); Yueet al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib43)\), while two recent retrieval\-augmented studies target*medicine–food homology*: Gong et al\.Gonget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib24)\)integrate the TCM “one root of medicine and food” principle into dietary recommendations, and Sha et al\.Shaet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib44)\)report a validation study on TCM\-styled cohorts\. Earlier work surveys TCM diet from a clinical\-nutrition standpointZhang \([2021](https://arxiv.org/html/2605.31410#bib.bib25)\)\. Across this stream, evaluation predominantly takes the form of structured graph QA or exam\-style knowledge recall rather than multimodal dish\-level decision making on real recipes\.

#### Safety\- and risk\-aware evaluation of medical LLMs\.

A growing body of work shifts medical\-LLM evaluation from accuracy alone toward safety and harm\-awareness\. Med\-PaLM and the MultiMedQA suite introduced multi\-axis human evaluation including harm, factuality, and biasSinghalet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib45)\), and MedHELM aggregates 35 benchmarks across 121 clinical tasks under an LLM\-jury whose scores correlate with clinician judgmentBediet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib46)\)\. Hallucination benchmarks such as Med\-HALTPalet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib47)\)and atomic\-fact methods like FActScoreMinet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib48)\)target factuality failures, while the recent RCT of Li et al\.Liet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib49)\)documents real\-world failures of LLM\-augmented nutrition chatbots and explicitly calls for risk\-aware intrinsic evaluation\. Adjacent to safety, “Cooking Up Risks” stress\-tests LLMs against adversarial food\-safety prompts, targeting misuse rather than well\-intentioned but clinically risky adviceLuoet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib50)\)\. Existing food\-safety benchmarks themselves remain centered on contamination, storage, and harmful preparationJacxsenset al\.\([2010](https://arxiv.org/html/2605.31410#bib.bib32)\); Le Vallée and Charlebois \([2015](https://arxiv.org/html/2605.31410#bib.bib33)\); Bryanet al\.\([1992](https://arxiv.org/html/2605.31410#bib.bib34)\); Munckeet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib35)\); Pekmezciet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib36)\)\. Despite this momentum, nutrition and chronic\-condition\-aware dietary advice remain underrepresented in safety\-oriented medical\-LLM evaluation: the existing literature offers strong methodological precedent but limited coverage of the dietary domain\.

#### Reasoning and retrieval for health\-aware generation\.

Two prompting strategies dominate recent work on knowledge\-intensive medical generation\. Chain\-of\-Thought promptingWeiet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib51)\); Kojimaet al\.\([2022](https://arxiv.org/html/2605.31410#bib.bib52)\)improves multi\-step reasoning, and Medprompt demonstrates large gains on clinical MCQA through CoT\-style prompting and ensemblingNoriet al\.\([2023](https://arxiv.org/html/2605.31410#bib.bib55)\)\. Retrieval\-Augmented Generation grounds outputs in external knowledgeLewiset al\.\([2020](https://arxiv.org/html/2605.31410#bib.bib53)\), and medical extensions such as MedRAG / MIRAGE show that retrieval over curated biomedical corpora substantially reduces hallucinationXionget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib54)\)\. In the food domain, retrieval has been applied to TCM medicine–food homologyGonget al\.\([2024](https://arxiv.org/html/2605.31410#bib.bib24)\); Shaet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib44)\)and to knowledge\-graph grounding for dietary guidanceGaoet al\.\([2025](https://arxiv.org/html/2605.31410#bib.bib23)\); Anet al\.\([2026](https://arxiv.org/html/2605.31410#bib.bib29)\)\. While individual systems incorporate CoT or RAG, a systematic comparison of these strategies on dietary recommendation across heterogeneous health conditions has not, to our knowledge, been reported\.

#### Position of this work\.

We introduce two Food\-as\-Medicine benchmarks for evaluating the decision layer of food AI\. The first benchmark tests multimodal dish\-level safety assessment: given a dish image, recipe ingredients, and a health condition, a model must determine whether the dish is suitable for that condition\. The second benchmark tests comparative dietary reasoning: given a health condition and four candidate dishes, a model must rank the dishes by condition\-specific suitability\. Together, these tasks shift evaluation from recognizing food and estimating nutrients to making health\-aware dietary decisions grounded in multimodal food evidence\.

## Appendix BRecipe Source Inventory

Table[3](https://arxiv.org/html/2605.31410#A2.T3)lists every web domain from which recipes were collected, together with the number of unique recipes contributed by that domain and the source category into which it is grouped in Section[3\.1](https://arxiv.org/html/2605.31410#S3.SS1)\. The*health\-information*category \(H\) comprises medical societies, clinical nutrition programs, academic medical centers, and public\-health agency portals; the*general\-food\-publication*category \(G\) comprises general\-audience food publications and dietitian\-curated cooking sites\. Percentages are computed against the corpus total of 3,859 unique recipes\.

Cat\.Domain\#%Hmds\.culinarymedicine\.org1,31234\.00Hdiabetesfoodhub\.org53113\.76Hrecipes\.heart\.org41110\.65Hheartandstroke\.ca2356\.09Harthritis\.org1042\.69Hkidney\.org681\.76Hucdincommon\.com471\.22Hmayoclinic\.org461\.19Hfruitsandveggies\.org310\.80Hcurearthritis\.org160\.41Hcookforyourlife\.org150\.39Hemersonhealth\.org90\.23Hhw\.qld\.gov\.au50\.13Hdietitiansaustralia\.org\.au50\.13Hheartfoundation\.org\.au40\.10Hcrohnscolitisfoundation\.org10\.03Hmasseycancercenter\.org10\.03Hnsw\.gov\.au10\.03Health\-information subtotal2,84273\.65Gbbcgoodfood\.com3338\.63G101cookbooks\.com2025\.23Geatingwell\.com1985\.13Gmamaknowsglutenfree\.com1824\.72Golivemagazine\.com280\.73Gbbc\.co\.uk150\.39Grealsimple\.com130\.34Glive2thrive\.org100\.26Gtaste\.com\.au40\.10Gthehealthychef\.com20\.05Gtheguthealthdoctor\.com20\.05Gsuperchargedfood\.com20\.05Gwomensweeklyfood\.com\.au20\.05Glivelighter\.com\.au20\.05Gjamieoliver\.com10\.03Gepicurious\.com10\.03Gdeliciousmagazine\.co\.uk10\.03Griverford\.co\.uk10\.03Gminimalistbaker\.com10\.03Gwellplated\.com10\.03Gcookieandkate\.com10\.03Gjoythebaker\.com10\.03Gwellnessmama\.com10\.03Ghealthymummy\.com10\.03Gtheprettybee\.com10\.03Gthehealthyfoodie\.com10\.03Gdizzybusyandhungry\.com10\.03Giowagirleats\.com10\.03Gkillingthyme\.net10\.03Gbromabakery\.com10\.03Gwillcookforfriends\.com10\.03Gwendypolisi\.com10\.03Gtrafficlightcook\.com10\.03Gqcwacountrykitchens\.com\.au10\.03Gdish\.co\.nz10\.03Gpunchfork\.com10\.03General\-food\-publication subtotal1,01726\.35Total3,859100\.00Table 3:Recipe source inventory across the benchmark\. Counts denote unique recipes per domain after URL\-level deduplication\. Percentages are computed against the corpus total of 3,859 recipes\.
## Appendix CNormalized Recipe Record Example

Each recipe retained after extraction and filtering \(Section[3\.1](https://arxiv.org/html/2605.31410#S3.SS1)\) is stored as a normalized JSON record\. Listing[1](https://arxiv.org/html/2605.31410#LST1)shows one such record\. Note how condition\-relevant modifiers \(“unsweetened,” “low\-fat,” “low sodium”\) are preserved verbatim in theingredientsanddietary\_tagsfields, and how thenutritiontable exposes the per\-serving micronutrient detail used by the validation rules in Appendix[E](https://arxiv.org/html/2605.31410#A5)\.

Listing 1:A normalized recipe record drawn from the FAM\-Bench corpus\. Long URLs and notes are abbreviated with “…” for readability; nutrition fields not used by downstream validation are omitted\.\{

"title":"FrozenBerrySmoothie",

"url":"https://mds\.culinarymedicine\.org/recipes/frozen\-berry\-smoothie/",

"image\_url":"https://mds\.culinarymedicine\.org/\.\.\./strawberries\.jpg",

"meal\_type":\["breakfast","beverage"\],

"dietary\_tags":\["lowsodium"\],

"cuisine":\["American"\],

"servings":2,

"serving\_size":"2cups",

"prep\_time\_minutes":5,

"cook\_time\_minutes":15,

"ingredients":\[

"2cupblueberries,frozen,unsweetened\(mayalsousefrozenstrawberries,raspberries,orotherberries\)",

"1cuporangejuice",

"1cupplainlow\-fatyogurt",

"1largebanana\(frozen\)"

\],

"instructions":\[

"Placethefrozenblueberries,orangejuice,yogurt,andfrozenbananainablenderorfoodprocessor\.",

"Blendtheingredientsuntilsmooth\.",

"Addwaterasneededandcontinueblendinguntilthedesiredconsistencyisreached\.",

"Serveimmediately\."

\],

"nutrition":\{

"calories":"273",

"total\_fat":"3g",

"saturated\_fat":"1g",

"sodium":"89mg",

"total\_carbohydrate":"55g",

"dietary\_fiber":"6g",

"total\_sugars":"39g",

"protein":"8g",

"potassium":"862mg"

\},

"notes":"Thisisalowsodiumrecipe\.Frozenberriescanbesubstitutedwithotherfrozenfruits\.\.\.",

"last\_updated":"2024\-06\-02",

"image\_valid":true

\}

## Appendix DKnowledge Base Schema

The condition\-level dietary knowledge base𝒦\\mathcal\{K\}used in Section[3\.3](https://arxiv.org/html/2605.31410#S3.SS3)has the form

𝒦=\{\(ℱh\+,ℱh−,ℱh×,𝒞h,𝒯h,ℛh\)\}h∈ℋ,\\mathcal\{K\}=\\big\\\{\(\\mathcal\{F\}\_\{h\}^\{\+\},\\mathcal\{F\}\_\{h\}^\{\-\},\\mathcal\{F\}\_\{h\}^\{\\times\},\\mathcal\{C\}\_\{h\},\\mathcal\{T\}\_\{h\},\\mathcal\{R\}\_\{h\}\)\\big\\\}\_\{h\\in\\mathcal\{H\}\},covering the 13 target conditions\. For eachhh,ℱh\+\\mathcal\{F\}\_\{h\}^\{\+\},ℱh−\\mathcal\{F\}\_\{h\}^\{\-\}, andℱh×\\mathcal\{F\}\_\{h\}^\{\\times\}list the beneficial foods, foods to limit, and foods to avoid;𝒞h\\mathcal\{C\}\_\{h\}is their canonical concept\-level abstraction \(e\.g\.,whole\_grains,processed\_meats,sugar\_sweetened\_beverages,high\_sodium\_foods,refined\_grains,healthy\_unsaturated\_fats\);𝒯h\\mathcal\{T\}\_\{h\}contains dietary\-tag hints aligned with common recipe metadata \(e\.g\., “low\-sodium,” “DASH,” “high\-fiber,” “diabetic\-friendly,” “heart\-healthy”\); andℛh\\mathcal\{R\}\_\{h\}records the supporting clinical references\.

## Appendix EValidation Rule Details

The rule\-grounded validation step described in Section[3\.3](https://arxiv.org/html/2605.31410#S3.SS3)applies three families of checks\.Ingredient–concept consistency:\(h,Eh\)\(h,E\_\{h\}\)pairs whose ingredients map toℱh×\\mathcal\{F\}\_\{h\}^\{\\times\}but are placed inRec\\mathrm\{Rec\}, or toℱh\+\\mathcal\{F\}\_\{h\}^\{\+\}but placed inNotRec\\mathrm\{NotRec\}, are flagged\.Nutrient threshold vetoes: condition\-specific nutrient ceilings are applied when the published nutrition table allows, such as sodium above the hypertension limit, saturated fat above the cardiovascular\-disease limit, and added sugar above the type\-2\-diabetes limit\.Ingredient\-fallback rules: emptyEhE\_\{h\}entries are backfilled from𝒱\\mathcal\{V\}patterns when the recipe contains canonical instances of the relevant rule set\. Annotations violating the first two are returned for a single corrective pass with the specific conflicts surfaced; persistent violations are dropped from the candidate pool\.

## Appendix FExclusion Criteria

During expert verification, candidate annotations are excluded when the available evidence cannot support a reliable judgment\. Common cases include missing ingredients, unobservable hidden ingredients, unclear preparation methods, and cases in which portion size is decisive but unavailable\.

## Appendix GBenchmark Split Statistics and Per\-Condition Coverage

Table[4](https://arxiv.org/html/2605.31410#A7.T4)summarizes the split\-level statistics referenced in Section[3\.3](https://arxiv.org/html/2605.31410#S3.SS3): label balance for the dish\-level task, the distractor\-type distribution for the comparative task, and the answer\-position balance across\{A,B,C,D\}\\\{A,B,C,D\\\}\.

Table 4:Split statistics for FAM\-Bench: label balance for the dish\-level task, distractor\-type distribution for the comparative task, and answer\-position balance\.Table[5](https://arxiv.org/html/2605.31410#A7.T5)reports the per\-condition coverage of the two benchmark splits\. The dish\-level column counts focal conditions in dish\-level suitability instances; the comparative\-ranking column decomposes multi\-condition comparative prompts and counts each condition mention\.

Table 5:Condition mention distributions for the two benchmark splits\. The dish\-level column counts focal conditions in dish\-level suitability instances; the comparative\-ranking column decomposes multi\-condition comparative prompts and counts each condition mention\. IBS: irritable bowel syndrome; NAFLD: nonalcoholic fatty liver disease; GERD: gastroesophageal reflux disease; CKD: chronic kidney disease\.
## Appendix HPer\-Condition Radar Plots: CoT, KI, and CoT\+KI

Figures[6](https://arxiv.org/html/2605.31410#A8.F6),[7](https://arxiv.org/html/2605.31410#A8.F7), and[8](https://arxiv.org/html/2605.31410#A8.F8)extend Figure[5](https://arxiv.org/html/2605.31410#S5.F5)from the main text by reporting Task 1 per\-condition decision accuracy under the remaining three prompting modes\. Axes, model set, and 13\-condition scope match the baseline plot; only the prompting mode varies\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/x6.png)Figure 6:Per\-condition decision accuracy underCoTprompting\.![Refer to caption](https://arxiv.org/html/2605.31410v1/x7.png)Figure 7:Per\-condition decision accuracy underKnowledge Injectionprompting\.![Refer to caption](https://arxiv.org/html/2605.31410v1/x8.png)Figure 8:Per\-condition decision accuracy underCoT \+ KIprompting\.
## Appendix IError Analysis Case Studies

The three failure modes named in Section[5\.3](https://arxiv.org/html/2605.31410#S5.SS3)are illustrated below with one missed\-risk false negative each, drawn from the Claude Sonnet 4\.6 baseline cell\. Recipe titles and ingredient lists are reproduced verbatim from the benchmark; the model’s predicted rationale is taken verbatim from its JSON output\.

#### Case 1\. Creamy Leek Soup for CKD\.

The recipe contains*white beans*\(high potassium\) and*semisoft goat cheese*\(high phosphorus\) alongside olive oil, leeks, “no salt added” vegetable stock, salt, and black pepper\. Asked whether the dish is suitable for someone managing chronic kidney disease, the model flagged the plant ingredients and the two “no salt added” items as condition\-relevant and predictedrecommend\. The gold label isnot\_recommend; the offending ingredients are the*bean*and the*cheese*\.

Two failure modes account for the error: \(i\)single\-axis risk checking— ingredients are evaluated along the sodium axis only, so the “no salt added” qualifier is treated as sufficient to clear the beans and the potassium and phosphorus axes are never audited, even though CKD restricts all three; \(ii\)rationale undercoverage— the goat cheese, the single highest\-phosphorus item in the recipe, does not appear in the rationale at all, consistent with a reader that attends to plant\-derived items and skips dairy\. A third candidate explanation — anchoring on the dish name “Creamy Leek Soup” — would require a name\-ablation experiment to confirm and is not pursued here\.

#### Case 2\. Lamb Sauce with Olives for stroke\.

The recipe pairs ground lamb \(1010g saturated fat per serving\) and grated Parmesan with onion, garlic, leeks, olive oil, “no salt added” tomato paste, green olives, herbs, and whole wheat fettuccine; the recipe’s own dietary tags include “low sodium”\. Asked whether the dish is suitable for someone managing stroke, the model marked eleven plant\-derived items and the no\-salt\-added tomato paste as condition\-relevant and predictedrecommend\. The gold label isnot\_recommend; the offending ingredient is the*cheese*, with the dish’s saturated\-fat load as the underlying rationale — low sodium does not offset it\.

The same two modes recur on a new condition: sodium cues drive the verdict \(the “low sodium” tag and “no salt added” qualifier are read as clearance\), and the rationale undercovers the two animal ingredients — the eponymous lamb and the Parmesan — in favour of every plant item in the recipe\.

#### Case 3\. Pistachio Crusted Pork for osteoporosis\.

The recipe contains boneless pork chops, pistachios, fresh basil, spray olive oil, lemon juice, honey, unsalted butter, and six tablespoons of*white wine*\. Asked whether the dish is suitable for someone managing osteoporosis, the model marked the pork, pistachios, lemon juice, and butter as condition\-relevant and predictedrecommend\. The gold label isnot\_recommend; the offending ingredient is the*wine*\(alcohol limits bone health and the recipe contributes no meaningful calcium or vitamin D to offset it\)\.

This case exposes a third mode beyond the CKD and stroke pattern:missing dietary rule\. The wine is not undercovered in the sense of selective reading — the model could easily see it in the ingredient list\. It is unflagged because the model’s working knowledge of osteoporosis does not include alcohol as a restricted item\. This is a knowledge gap, not an attention gap, and predicts exactly the asymmetry observed in Table[1](https://arxiv.org/html/2605.31410#S5.T1): Knowledge Injection lifts the binary verdict \(the retrieved rule fires\) but barely moves rationale macro\-F1 \(the selection over already\-visible evidence is not the bottleneck here\)\.

## Appendix JInference Settings

This appendix records the exact decoding and serving configuration used to produce every cell of Table[1](https://arxiv.org/html/2605.31410#S5.T1), extending the one\-line summary in Section[4\.1](https://arxiv.org/html/2605.31410#S4.SS1)\.

#### Decoding\.

All five models use greedy decoding with temperatureT=0T=0and a40964096\-token output budget\. JSON mode is enforced wherever supported by passingresponse\_format=\{"type":"json\_object"\}on the OpenAI\-compatible chat completion call; this covers the OpenAI, Anthropic\-via\-compat, Gemini, Qwen, and locally served vLLM endpoints\. Three per\-provider fallbacks are applied automatically when the API rejects a parameter: \(i\) for newer OpenAI reasoning models that disallow a customtemperature, the field is dropped and the model’s default sampling temperature is used; \(ii\) for Gemini models served through the OpenAI\-compatible shim that rejectmax\_tokens, the field is dropped; \(iii\) for providers that reject the JSONresponse\_format, the field is dropped and the JSON contract is enforced through the system prompt alone\. None of the five models in the main results table fell back beyond \(i\)\.

#### Concurrency and hardware\.

Closed\-source APIs \(GPT\-5\.4, Claude Sonnet 4\.6, Gemini 2\.5 Pro\) are queried at concurrency88from a single client host\. Open\-weight models \(Qwen3\-VL\-8B\-Instruct, Gemma\-3\-12B\-IT\) are served locally via vLLM on a single NVIDIA RTX 5090 \(32 GB\) at concurrency44; the concurrency limit reflects the GPU’s KV\-cache budget at40964096\-token output for an88–1212B vision–language backbone, not an API rate limit\.

#### Image handling\.

Dish imagesIIare submitted at their native source resolution\. Closed\-source APIs perform any further server\-side resizing internally; vLLM hands the raw image tensor to the model’s own vision pre\-processor\. No client\-side resizing, cropping, or re\-encoding is applied\.

#### Failure handling\.

A request is considered persistently failed only after three retries with exponential backoff\. Across the5×4×2=405\\times 4\\times 2=40model\-mode\-task cells, two persistent failures were observed \(one on each of two distinct closed\-source cells, both due to upstream5​x​x5xxresponses\) and the corresponding instances are dropped from the per\-model pool for that cell only; all reported metrics are computed over the surviving instances\. No model\-mode\-task cell loses more than0\.07%0\.07\\%of its instances\.

## Appendix KExample Questions

Figure[9](https://arxiv.org/html/2605.31410#A11.F9)shows an example of the dish\-level suitability assessment task, and Figure[10](https://arxiv.org/html/2605.31410#A11.F10)shows an example of the comparative dish\-ranking task\.

![Refer to caption](https://arxiv.org/html/2605.31410v1/Images/task1.png)Figure 9:Example question for Task 1: dish\-level suitability assessment\.![Refer to caption](https://arxiv.org/html/2605.31410v1/x9.png)Figure 10:Example question for Task 2: comparative dish ranking\.
## Appendix LPrompting Configurations

This appendix gives the verbatim system prompts used in Section[4\.2](https://arxiv.org/html/2605.31410#S4.SS2)and a worked example of the per\-condition knowledge injection referenceℛh\\mathcal\{R\}\_\{h\}\. All four modes share the*baseline*JSON output contract; CoT modes prepend the*CoT*system prompt, and KI modes append the*KI addendum*\. The user message in every mode carries the recipe payload\(I,G,h\)\(I,G,h\); in KI modes it additionally carries adisease\_food\_referencefield built fromℛh\\mathcal\{R\}\_\{h\}\.

#### Baseline system prompt\.

Youareevaluatingwhetheradishissuitableforahealth\-managementquestion\.

Task:

\-Inferwhetherthedishshouldberecommendedornotrecommendedfortheaskedcondition\(s\)\.

\-Explainyourdecisionbylistingcondition\-specificsupportingingredients\.

\-Useonlyinformationfromtheprovidedrecipe/questioncontext\.

ReturnONLYvalidJSONinthisexactshape:

\{

"decision":"recommend\|notrecommend",

"rationale\_ingredients":\[

\{

"condition":"<conditionname\>",

"ingredients":\["<ingredientshortname\>","\.\.\."\]

\}

\]

\}

Rules:

\-decisionmustbeexactly"recommend"or"notrecommend"\.

\-rationale\_ingredientscanbeemptyonlyifnovalidingredientevidenceisavailable\.

\-Keepconditionnamesconciseandcloselyalignedwiththeaskedquestion\.

\-Ingredientsmustcomefromtheprovidedrecipeingredientlist\.

#### Chain\-of\-Thought \(CoT\) system prompt\.

Replaces the baseline prompt in CoT and CoT\+KI modes\.

Youareevaluatingwhetheradishissuitableforahealth\-managementquestion\.

Thinkstepbystepbeforeanswering\.ProduceyourreasoningasaJSONarray

ofshortstrings\(onethoughtperstep\),thencommittoafinaldecision\.

Reasoningchecklist:

1\.Identifyeachhealthconditionimpliedbythequestion\.

2\.Foreachcondition,scantherecipeingredientsandnutritionfor

factorsthatclearlysupportorconflictwiththatcondition\.

3\.Weighsupportingvs\.conflictingevidence\.

4\.Decide"recommend"onlywhentheoverallevidencesupportssuitability;

otherwise"notrecommend"\.

Useonlyinformationfromtheprovidedrecipe/questioncontext\.

ReturnONLYvalidJSONinthisexactshape:

\{

"reasoning\_steps":\["<step1\>","<step2\>","\.\.\."\],

"decision":"recommend\|notrecommend",

"rationale\_ingredients":\[

\{

"condition":"<conditionname\>",

"ingredients":\["<ingredientshortname\>","\.\.\."\]

\}

\]

\}

Rules:

\-reasoning\_stepsmustcontain2\-6concisesteps\.

\-decisionmustbeexactly"recommend"or"notrecommend"\.

\-rationale\_ingredientscanbeemptyonlyifnovalidingredientevidenceisavailable\.

\-Keepconditionnamesconciseandcloselyalignedwiththeaskedquestion\.

Hardformattingrules\(theoutputisscoredbyexactstringmatch\):

\-EACHconditionmustappearasitsownentryinrationale\_ingredients\.

Nevermergemultipleconditionsintoone"condition"string\.DoNOTuse

"and","&","/","\+",orcommastojoinconditionnames\.

\-Eachingredientmustbethebarefoodnameasitwouldappearona

cleanshoppinglist\.Stripquantities,units,preparation/state

qualifiers,brandandorigindescriptors,parentheticalsandtrailing

notes\.Uselowercase,singular\-form\-as\-listed\.

\-Onlyincludeingredientsthatgenuinelydrivethe\(non\-\)recommendation

forthatspecificcondition\.

#### KI addendum\.

Appended to the baseline or CoT system prompt in KI and CoT\+KI modes\.

Additionalreferencerule:

\-Theuserpayloadmayincludedisease\_food\_referencefrom

merged\_disease\_food\_recommendations\.json\.

\-Usethatreferenceasbackgroundguidanceaboutfoodsthatmay

supportorconflictwitheachcondition\.

\-Thefinaldecisionmuststillbebasedontheprovided

recipe/questioncontext\.

\-Ingredientsinrationale\_ingredientsmuststillcomefromtherecipe

ingredientlist,notfromthereferencealone\.

#### Worked knowledge injection referenceℛh\\mathcal\{R\}\_\{h\}\(hypertension\)\.

For knowledge injection modes, the user payload carries adisease\_food\_referenceblock populated with the focal condition’s entry from𝒦\\mathcal\{K\}\. Listing[2](https://arxiv.org/html/2605.31410#LST2)shows the hypertension entry as it is injected \(truncated for space; the full 13\-condition reference is released with the benchmark\)\.

Listing 2:Workedℛh\\mathcal\{R\}\_\{h\}for hypertension\. The “recommend” and “not\_recommend” food lists are the rules the model must apply to the observed ingredient list; no dish\-specific labels appear\."hypertension":\{

"recommend":\[

"fruits","vegetables","grains",

"low\-fatdairyproducts","low\-fatorfat\-freemilk",

"reduced\-fatcheeses","bread","plaincereal",

"brownorwhiterice","pasta",

"cannedfruitinjuiceorwater",

"frozenvegetableswithoutaddedbutterorsauces",

"cannedvegetablesorvegetablesoupslowinsodium",

"lentils","blackbeans","chickpeas",

"whitemeatskinlesschickenandturkey","unbreadedfish",

"porktenderloin","extra\-leangroundbeef",

"herbs","spices","flavoredvinegars",

"fruits\(4\-5servings/day\)",

"vegetables\(includingcolorful,legumes\)",

"low\-fatdairy\(2\-3servings/day\)",

"wholegrains\(7\-8servings/day\)",

"nuts\(4\-5servings/week\)","beans\(\>3servings/week\)"

\],

"not\_recommend":\[

"high\-sodiumfoods","processedfoodshighinsodium",

"soups","broths","cannedvegetables",

"condiments","saladdressings","sauces","dips",

"ketchup","mustard","relish","pickles","olives",

"morethanamoderateamountofalcohol",

"morethan2cupsofcoffeeaday",

"fattymeats","full\-fatdairy",

"sugar\-sweetenedbeverages","sweets",

"butter","cheese","pizza",

"delimeatsandwiches","coldcuts","curedmeats",

"burgers","burritos","tacos",

"savorysnacks\(chips,crackers,popcorn\)"

\]

\}

Similar Articles

Introducing HealthBench

OpenAI Blog

OpenAI introduces HealthBench, a new benchmark for evaluating AI systems in healthcare contexts, created with 262 physicians across 60 countries. The benchmark includes 5,000 realistic health conversations with physician-written rubrics to assess model performance on meaningful, trustworthy, and improvable metrics.

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

arXiv cs.AI

Introduces CliniCARE-Bench, a deployment-oriented benchmark for evaluating AI agents on clinical audit tasks over longitudinal EHR data, with 25 clinician-validated scenarios and 750 patient cases. It assesses verdict accuracy, evidence grounding, policy adherence, and calibrated abstention, finding that raw accuracy overstates investigation quality.

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

arXiv cs.CL

This paper introduces BehaviorBench, a comprehensive benchmark for evaluating foundation models on behavioral science tasks including behavior prediction, strategic decision-making, subject-trait inference, and behavioral knowledge application. It also presents Be.FM-1.5, a fine-tuned model that achieves strong distributional alignment, highlighting the gap between general-purpose and behaviorally adapted models.