A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
Summary
This paper presents a human-centred benchmarking approach for evaluating large language models in parenting advice, assessing 15 LLMs across 100 scenarios in English and Chinese using expert-derived rubrics.
View Cached Full Text
Cached at: 08/18/26, 09:53 AM
# A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
Source: [https://arxiv.org/html/2608.14622](https://arxiv.org/html/2608.14622)
Yunke Zhao\\equalcontrib1, Isobel Voysey\\equalcontrib1, Alastair van Heerden2, Rob Hughes2, Jun Zhao1
###### Abstract
People are increasingly using large language models \(LLMs\) to seek advice, including for parenting\. Parenting is a critical and socially sensitive domain\. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses\. With a multi\-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages \(English and Chinese\), using an LLM\-as\-a\-judge method\. Results show that aggregate scores can hide rubric item\-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses\. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM\-generated advice in domains like parenting\. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user\-facing parenting advice applications\.
## 1Introduction
Parents routinely make complex decisions about how best to care for their children\. To support these decisions, parents have long relied on advice from their family, friends, and professionals, and a large proportion seek information and advice through online sources, including social media, parenting forums, and traditional search engines\(Mertenset al\.[2024](https://arxiv.org/html/2608.14622#bib.bib36); Bakeret al\.[2017](https://arxiv.org/html/2608.14622#bib.bib6)\)\. More and more, large language models \(LLMs\) such as ChatGPT are becoming part of this informational ecosystem, with early evidence suggesting that some parents are beginning to use LLMs to seek parenting guidance and support\(Sharmin and Afrin[2026](https://arxiv.org/html/2608.14622#bib.bib46)\)111Further evidence includes popular media depictions of parents using LLMs for advice, such as a recent subplot inAmandaland, a British comedy series about parenting\(amandaland\)\.\.
Existing research evaluating LLMs for parents or in parenting\-adjacent domains has largely focused on health information tasks, where outputs can often be assessed against relatively clear standards of factual correctness, safety, or completeness\(McFaydenet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib35); Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10); Sezginet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib45)\)\. However, the quality of parenting advice extends beyond factual accuracy alone\. Parenting is a socially and culturally situated practice shaped by values, norms, beliefs about child development, and differing theories of parent\-child relationships\(Linet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib30); Lansford[2022](https://arxiv.org/html/2608.14622#bib.bib26)\)\. In many parenting situations, there may be no single objectively correct answer\. Instead, advice quality may be affected by how actionable it is, the tone taken, and how relatable it is to specific cultural contexts or long\-term challenges facing a family\.
These characteristics create challenges for evaluating LLM\-generated parenting advice\. Traditional benchmark approaches centred on accuracy or matching human preferences\(Beanet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib9)\)may fail to capture the relational and value\-laden nature of parenting support\. Evaluating parenting advice therefore requires morehuman\-centred approachesthat consider not only whether advice is safe and accurate, but also social qualities of the support provided by the model, what parenting style the model implicitly reinforces, and how advice differs across cultural and linguistic contexts\.
In this paper, we present a multilingual evaluation pipeline for assessing LLM\-generated parenting advice\. The pipeline combines scenario generation and refinement, large\-scale response generation, rubric\-based evaluation using LLM judges, and parenting\-style analysis grounded in established parenting theory\. Using 100 parenting scenarios in English and Mandarin \(Chinese\), we compare responses from 15 LLMs across multiple dimensions of advice quality and parenting style\.
Our work addresses the following questions:
1. 1\.How does rubric\-based evaluation reveal different forms of user\-facing advice quality across models and languages?
2. 2\.How implicit parenting styles are reflected in LLM\-generated advice, and how does this vary across models and languages?
Through our work we make three primary contributions:
- •A scalable multilingual pipeline for evaluating LLM\-generated parenting advice: The pipeline encompasses scenario generation, response generation, LLM\-based judging, and automated analysis\. The pipeline also includes various failure repair strategies, an auditable output trace for each score, and the option to conduct the same evaluation in another language\.
- •A human\-centered evaluation framework for parenting advice: The expert\-informed rubric moves beyond factual correctness to incorporate relational and behavioural qualities of LLM\-generated advice, which is further extended through classification of implicit parenting style\.
- •Empirical comparison of LLMs across advice quality and parenting styles: Our findings show a wide range in advice quality across models and model families and models displayed different capability strengths across the eight evaluation rubrics, confirming a multi\-dimensionality of LLMs for supporting user\-facing scenarios\. Furthermore, when responding in English, models lean more authoritative in parenting style than they do when responding in Chinese, where they lean more authoritarian\.
## 2Background
### 2\.1LLMs and Parenting
With the global rise of nuclear families, many parents go online to seek information and social support for parenting\(Mertenset al\.[2024](https://arxiv.org/html/2608.14622#bib.bib36); Bakeret al\.[2017](https://arxiv.org/html/2608.14622#bib.bib6); Dworkinet al\.[2013](https://arxiv.org/html/2608.14622#bib.bib11)\), marking a shift from previous generations that predominately relied on families, close friends, and childcare professionals \(e\.g\., teachers, paediatricians\)\. Developments in LLMs have introduced a new opportunities for parents, especially those from in Western countries, who are in dire need of tools that provide information, advice, and social support\(Sharmin and Afrin[2026](https://arxiv.org/html/2608.14622#bib.bib46)\), necessitating the evaluation of these systems\.
Research has not yet established the real\-world prevalence of parents using LLMs, but patterns of parents’ existing online behaviour combined with high usage among the general public makes LLMs for parenting support worth study\. For example,Yun and Bickmore \([2025](https://arxiv.org/html/2608.14622#bib.bib56)\)found that 21% of adult survey participants were using LLMs to seek health information, motivating researchers to evaluate LLMs as a tool for parents seeking health information about their child\(McFaydenet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib35); Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10); Sezginet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib45); Leslie\-Milleret al\.[2024](https://arxiv.org/html/2608.14622#bib.bib28)\)\. All of these works indicate the potential for LLMs in these situations, but the authors note shortcomings, like missing or inaccurate references\(Sezginet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib45); McFaydenet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib35)\)or the failure to provide appropriate, actionable recommendations\(Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10); McFaydenet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib35)\)\.
ChatGPT has been the primary LLM under investigation in these cases\. Researchers have looked into the effect of different base models, e\.g\., GPT\-3\.5 vs GPT\-4\(Kimet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib22); Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10)\), and different data sources, e\.g, training data only vs access to internet vs access to parenting\-specific content repository\(Kimet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib22)\)on parenting advice quality, but no work so far has compared a broad set of models from several different families for parenting advice\. To address this gap, we present the first comparison of LLM\-generated advice for general\-purpose parenting advice from multiple model families\.
### 2\.2Human\-Centered Evaluation of LLMs
Evaluating LLMs remains a challenge\. Traditional approaches have relied on automated benchmarks capturing performance on a narrowly specified task, but this falls short in tasks where good performance is highly subjective\(Xiaoet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib54)\)\. Researchers have therefore attempted to assess how aligned AI is with human values and preferences, but this presents its own challenges\. Many forms of human\-in\-the\-loop evaluation are time\-consuming and resource\-intensive, limiting the extent of evaluation\. This is reflected in the literature on LLMs in the parenting domain, where evaluations have typically been based on only a small number of questions or vignettes, ranging from as few as 3\(Leslie\-Milleret al\.[2024](https://arxiv.org/html/2608.14622#bib.bib28)\)to 22\(Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10)\)\. When it comes to optimising for a set of preferences, differences may be substantial across groups, leading researchers to prompt careful consideration of whose preferences are optimised for in what tasks\(Kirket al\.[2024](https://arxiv.org/html/2608.14622#bib.bib23)\)\.
Parenting advice is often subjective, making it challenge to evaluate\. Previous work on LLMs for parents seeking information was largely restricted to the medical domain, where experts largely concur on appropriate and inappropriate courses of action, meaning these works focused mainly on accuracy of information\.Bushuvenet al\.\([2023](https://arxiv.org/html/2608.14622#bib.bib10)\)analysed LLM responses to vignettes about paediatric emergencies, coding the diagnosis, advice to call medical professionals, and advice given to first aiders as correct or incorrect\.McFaydenet al\.\([2024](https://arxiv.org/html/2608.14622#bib.bib35)\)scored ChatGPT responses to questions about autism on three scales, correctness, clarity, and conciseness\. However, providing a set of ‘correct’ answers to compare an LLM response to is not a viable option for more general parenting questions, like suitable discipline practices, particularly given the aforementioned individual and cultural variations in parenting\.
Furthermore, the finding bySharmin and Afrin \([2026](https://arxiv.org/html/2608.14622#bib.bib46)\)that some mothers are turning to LLMs for advice to avoid social judgement suggests that relational dimensions of advice, like empathy, are particularly important\. These have largely been neglected in existing work in the parenting domain\.McFaydenet al\.\([2024](https://arxiv.org/html/2608.14622#bib.bib35)\)considered how understandable and actionable advice was, as well as whether the language used was more medical or neurodiversity\-affirming\. Response characteristics like this may affect parents’ perceptions of LLMs;Leslie\-Milleret al\.\([2024](https://arxiv.org/html/2608.14622#bib.bib28)\)found no difference between the trustworthiness and perceived expertise of responses from ChatGPT and paediatricians when parents were blind to the source\.
Evaluating parenting advicerequires consideration of subjective and relational dimensionsthat extend beyond simple correctness\. LLM\-as\-a\-judge approaches enable scalable assessment of such complex responses, but prior work has identified potential biases, including preferences for longer answers and models in the same family\(Stureborget al\.[2024](https://arxiv.org/html/2608.14622#bib.bib47); Kooet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib24)\)\. We therefore develop a pipeline that combines multiple LLM judges with an expert\-informed multidimensional rubric designed to assess both informational quality and the relational characteristics of advice\.
### 2\.3Culturally\-Sensitive Parenting Advice
Evaluating LLMs for parenting requires considering not only the relational aspects of the interaction between the LLM and the parent, but also the advised interaction between the parent and the child\. Parenting is a socially and culturally situated practice\(Linet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib30); Lansford[2022](https://arxiv.org/html/2608.14622#bib.bib26)\), which means that parenting advice tends to go beyond purely informational content to promote certain patterns of behaviour, attitudes, and approaches—normative guidance about how parents should react and relate to their children—commonly referred to as parenting styles\. What is considered by the majority to be an appropriate or ideal parenting style varies across cultures\(Linet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib30); Grusecet al\.[2017](https://arxiv.org/html/2608.14622#bib.bib14)\)and within them, as an individual’s parenting decisions are also shaped by their personal experiences\(Rutiglianoet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib43)\)\. Furthermore, research suggests the parenting style associated with positive child outcomes is also dependent on culture; authoritarian parenting may result in negative outcomes for children in individualistic societies where autonomy is prioritised but work well in collectivist societies where group needs are paramount\(Rudy and Grusec[2001](https://arxiv.org/html/2608.14622#bib.bib42)\)\. As such, professionals advocate for understanding and integrating culture in support for parents\(Harkness and Super[2020](https://arxiv.org/html/2608.14622#bib.bib16); Schillinget al\.[2021](https://arxiv.org/html/2608.14622#bib.bib44); Baumannet al\.[2015](https://arxiv.org/html/2608.14622#bib.bib7)\)\.
Existing work indicates that generic LLMs are poor at responding in culturally sensitive ways, i\.e\., providing advice that considers the user’s beliefs and attitudes, which are shaped by the values, norms, and assumptions prevalent in their community\(Aldaweeshet al\.[2026](https://arxiv.org/html/2608.14622#bib.bib1)\), largely representing WEIRD contexts\(qiu2025evaluating\)\. Previous evaluations of LLMs for parenting advice have not brought in culture or context as a relevant factor, usually considering a small set of scenarios with no acknowledgement of or variation in context \(e\.g\., “How do I set boundaries and discipline my child?”\(Kimet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib22)\)\)\. Similarly, LLM performance and response style can vary across languages\(chevi2025individual\), which has not yet been analysed in the parenting domain\.
In this work, we extend the existing evaluations through the development of a pipeline that evaluates the LLM across a set of context\-specified scenarios and classifies parenting style according to Baumrind, Maccoby, and Martin’s responsiveness\-demandingness theory\(Baumrind[1971](https://arxiv.org/html/2608.14622#bib.bib8); Maccoby and Martin[1983](https://arxiv.org/html/2608.14622#bib.bib33)\)in order to indicate models’ normative stance on how parents should react to and relate to children\. This evaluation is carried out in English and Chinese to further investigate cross\-language variation in advice quality and characteristics\.
## 3Methods
### 3\.1Overview of Evaluation Pipeline
We developed a multi\-stage evaluation pipeline to assess the quality and characteristics of LLM parenting advice across multiple models and languages222https://github\.com/UndeadCZ13/parenting\-advice\-llm\-evals\. The pipeline consisted of four stages, which are shown in Figure[1](https://arxiv.org/html/2608.14622#S3.F1):
1. 1\.Expert\-guided parenting scenario generation
2. 2\.Response generation
3. 3\.LLM\-based judging
4. 4\.Automated quantitative analysis and reporting of benchmarking results
The evaluation was conducted over 15 models using 100 parenting scenarios presented in both English and Chinese\. Responses generated by LLMs were assessed using rubric\-based quality scoring and parenting\-style classification by two different LLM judges\.
Scenario GenerationParenting scenario creationDevelop diverse parenting scenarios in English\.Expert review and refinementExperts review scenarios for clarity, realism, and appropriateness\.Scenario translationTranslate scenarios into Chinese\.Response GenerationModel samplingSample responses from multiple LLMs under a shared protocol\.Response validation and repairCheck for issues and regenerate flagged responses when needed\.LLM\-Based JudgingRubric\-based quality judgingLLM judges score responses using a rubric and provide brief comments\.Parenting style classificationLLM judges classify responses into parenting styles\.Automated ReportingComparison of modelsEvaluate and compare performance across LLM models\.Comparison of judgesAssess agreement and consistency between LLM judges\.Comparison of languagesCompare results between English and Chinese scenarios\.scenariopromptsN×LN\{\\times\}Lresponses\+ metadataN×M×LN\{\\times\}M\{\\times\}LjudgedrecordsN×M×L×JN\{\\times\}M\{\\times\}L\{\\times\}JScale reference:NN= number of scenarios;MM= number of models;LL= number of languages;JJ= number of judges\.
Figure 1:Simplified evaluation workflow
### 3\.2Scenario Generation
#### Parenting Scenario Creation
An initial set of 100 parenting scenarios was generated using ChatGPT\. The scenarios were designed to represent a broad range of common parenting situations encountered across developmental stages and family contexts\. The generated scenarios covered multiple topics, including health, nutrition, safety, development, discipline, education, sleep, parent wellbeing, social\-emotional support, special needs, and low\-resource contexts\. Each scenario consisted of a single\-turn prompt describing a parenting situation and requesting advice\. No follow\-up clarification or conversational context was provided, allowing all models to respond under identical single\-shot conditions\.
#### Expert Review and Refinement
The initial scenarios were manually reviewed and refined by an expert group with domain expertise in parenting and child development\. The review process aimed to improve realism, diversity, and clarity\.
#### Scenario Translation
All finalised scenarios were translated from English into Chinese to enable cross\-language comparison\. This translation was carried out using GPT\-5 under prompt constraints\. The prompt gave explicit instructions to preserve semantic content, retain the original level of risk and tone, avoid adding or removing key facts, and avoid introducing extra cultural assumptions \(see Appendix\)\. A subset of scenarios were reviewed by an author fluent in both English and Chinese to verify the accuracy of the translation\.
### 3\.3Response Generation
#### Model Sampling
Parenting advice responses were generated using 15 LLMs across 100 scenarios in 2 languages \(English and Chinese\)\. The models evaluated spannedseveral model families and deployment types, including frontier systems, compact variants, and smaller open or service\-based models\.
The set included four GPT\-family models \(GPT\-5\.2, GPT\-5 Nano, GPT\-OSS 20B, and GPT\-4o Mini\), two DeepSeek models \(DeepSeek V3\.1 and DeepSeek R1\), two Qwen models \(Qwen3 32B and Qwen3 8B\), two Llama models \(Llama 3\.3 70B and Llama 3\.1 8B\), two GLM models \(GLM\-4\.6 and GLM\-4 9B\), and three additional models: Kimi K2 Thinking, MiniMax M2, and Ministral 3 14B333Claude models were not included because access was not available through our institution at the time of evaluation, which we acknowledge as a limitation of the work\. Extending the comparison to Claude models is an important avenue for future work\.\.
Each model received the full scenario text, including the context, tags, and parent question\. Thecontextgives relevant information about the child, family, or setting\. Thequestiondefines the advice\-seeking task\. Thetagslocate the scenario within broader parenting topics\. Prompts were language\-aligned: English scenarios elicited English responses, and Chinese scenarios elicited Chinese responses\. The task was framed as giving practical advice to a parent in the described situation\.
The prompt used a final\-answer\-only design\. Models were instructed to provide the parent\-facing answer without intermediate reasoning or wrapper text\. This reduces format variation across model families and makes the judged response unit more consistent\.
#### Response Validation and Repair
To ensure reliable assessment, we implemented a response validation, repair, and cleaning stage to identify incomplete, malformed, or failed generations that could distort later scores for reasons unrelated to advice quality\.
Problematic responses were automatically flagged using generation metadata and rule\-based heuristics, including API errors, truncation indicators, missing terminal punctuation, cut\-off endings, and unusually short final fragments\.
Flagged cases were then manually reviewed before rerun, since simple surface rules can over\-flag complete responses, especially across languages and punctuation systems\.
Confirmed defective response were regenerated for a maximum of three attempts\. Repaired responses were merged back into the canonical answer files, with repair manifests and rerun logs recording which rows were changed\. This preserved auditability while reducing the risk that execution failures contaminated model comparison\.
A further cleaning step was applied after generation because different model families often produce different output shapes\. Reasoning\-like segments, wrapper formats such as`<think\>`blocks, and explicit final\-answer markers were stripped from the raw output to remove this as a confounding factor from judging\. The original model output was stored, alongside the cleaned output and further metadata recording if and how many characters were removed\.
To further enhance auditability, generation metadata was recorded alongside each response\. This included finish reason, retry use, suspected truncation flags, and API\-related errors\. These signals helped distinguish complete but weak answers from outputs affected by backend failure, empty generation, or length\-limited stopping\.
### 3\.4LLM\-Based Judging
Generated parenting advice was evaluated using an LLM\-as\-a\-judge framework employing two independent judge models: GPT\-5\.2 as the primary judge and DeepSeek V3\.1 as the secondary judge\. The secondary judge provided a robustness check from a different model family\. During the judge process, each LLM\-generated response wasassessed across eight predefined quality rubrics, measuring dimensions of parenting advice quality, and then further wereclassified by parenting style\. As a result, the judging stage produced quantitative rubric scores, parenting style classifications, and qualitative annotations for scoring decisions\.
#### Rubric\-Based Judging
The pipeline used LLM\-as\-a\-judge as the main scoring method because parenting advice is open\-ended and cannot be evaluated through a single reference answer\. The judging stage used eight rubrics for advice quality inspired by HealthBench\(Aroraet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib2)\)and informed by parenting experts\. They include the follows:
1. 1\.Accuracy: factual correctness and evidence basis
2. 2\.Safety: harm avoidance and risk awareness
3. 3\.Helpfulness: actionable, practical steps
4. 4\.Empathy: supportive, non\-judgmental tone
5. 5\.Completeness: covers key aspects without major omissions
6. 6\.Bias avoidance: avoids stereotypes and harmful assumptions
7. 7\.Limitation awareness: recognises limits and refers to professionals when appropriate
8. 8\.Communication: presents clear structure and asks clarifying questions when needed
These rubrics allow response quality to be analysed across multiple dimensions\. They make it possible to distinguish, for example, between advice that is practical but weak in safety and advice that is safe but too vague to be useful\.
The judge prompt presented the scenario, the model response, and the rubric definitions\. The judge assigned a score from 0 to 100 for each rubric and provided one short comment explaining the judgement\. The output was constrained to a fixed JSON schema so that scores and comments could be parsed and merged consistently\. The primary LLM judge evaluated each response three times, and the final score for each rubric item was computed as the mean across repeats\. The overall score was the mean across the eight rubric items\. This process was repeated for the secondary judge\.
#### Parenting Style Classification
Rubric scores describe response quality, but they do not fully capture the nature of the advice\. For example, two responses may receive similar quality scores while taking different interpersonal stances\. In parenting advice, this matters because the answer may imply a particular balance of warmth, structure, reassurance, and boundary\-setting\.
The parenting style judging was based on the Responsiveness\-Demandingness framework fromBaumrind \([1971](https://arxiv.org/html/2608.14622#bib.bib8)\)andMaccoby and Martin \([1983](https://arxiv.org/html/2608.14622#bib.bib33)\)\.Responsivenesscaptures warmth, emotional support, and sensitivity to the child’s situation\.Demandingnesscaptures structure, expectations, behavioural guidance, and boundary\-setting\. Together, these dimensions definefour broad parenting styles: authoritative, authoritarian, permissive, and uninvolved\.Authoritativeadvice is warm and provides structure to the child\.Authoritarianadvice is firmer and less attuned\.Permissiveadvice is warm but loosely structured\.Neglectfuladvice is weak in both warmth and structure\.
The style was determined through a dedicated judging prompt\. The prompt defined Responsiveness and Demandingness and asked the judge to return structured JSON scores\. Each answer receives a Responsiveness score \(RR\) and a Demandingness score \(DD\) from 0 to 1\. These are then converted into a four\-style probability distribution: authoritative =R×DR\\times D, authoritarian =\(1−R\)×D\(1\-R\)\\times D, permissive =R×\(1−D\)R\\times\(1\-D\), and neglectful =\(1−R\)×\(1−D\)\(1\-R\)\\times\(1\-D\), with the four probabilities normalised to sum to one\. This two\-step calculation is used because the four styles derive from the two dimensions\. Direct classification would force mixed responses into one label, while separateRRandDDscores preserve gradation for later analysis\.
Keeping the style branch separate from rubric scoring preserves a clear distinction between advice quality and advice style, as shown in Section[4\.2](https://arxiv.org/html/2608.14622#S4.SS2)\. These scores were used to analyse how parenting style varies across models and languages\.
### 3\.5Automated Reporting
After generation, quality control, and judging, the rubric outputs were exported into structured score tables and merged into a single analysis\-ready matrix\. The core analysis unit was a scored response indexed by scenario, model, language, and judge\. Each row retained the generated answer, repeated judge scores, aggregated rubric means, one judge comment, and parenting style classification probabilities\. This organisation supported the main comparisons used later in the paper: model comparison within a language \(English vs\. Chinese\), language comparison within a model, and robustness comparison across judges \(GPT vs\. Deepseek\)\. Summary statistics and figures visualising these comparisons were also generated, a selection of which are included in the following section\.
## 4Results
Here we present our key findings, regarding overall model performances, in addition to rubric\-, scenario\-, and language\-specific performance comparisons \(Section[4\.1](https://arxiv.org/html/2608.14622#S4.SS1)\), a reflection of parenting style by models \(Section[4\.2](https://arxiv.org/html/2608.14622#S4.SS2)\), and judge robustness \(Section[4\.3](https://arxiv.org/html/2608.14622#S4.SS3)\)\.
### 4\.1Advice Quality
#### 4\.1\.1 Overall Score
Figure 2:Overall score distribution across the scenarios by model\. Primary judge = GPT\-5\.2The benchmark shows clear overall performance differences across the 15 evaluated models\. Figure[2](https://arxiv.org/html/2608.14622#S4.F2)reports the distribution of overall scores under the primary judge\. GPT\-5\.2 occupies the strongest position, followed by GPT\-5 Nano and Kimi K2 Thinking\. GLM\-4 9B and Llama 3\.1 8B are lower across most of the distribution\.
Figure 3:Heatmap of pairwise win\-rate of overall rubric score for each scenario\. Primary judge = GPT\-5\.2Pairwise win\-rate tests whether model differences remain stable across scenarios\. As shown in Figure[3](https://arxiv.org/html/2608.14622#S4.F3), GPT\-5\.2 dominates almost all pairwise comparisons, while GPT\-5 Nano and Kimi K2 Thinking also retain strong positions\. Lower\-ranked models accumulate broad losses, suggesting that the benchmark captures a stable comparative structure rather than isolated favourable cases\.
#### 4\.1\.2 Rubric\-Level Scores
Figure[2](https://arxiv.org/html/2608.14622#S4.F2)establishes the broad performance tiers of the benchmark\. However, overall score does not explain what kind of advice quality produces each model’s position\. A model may give practical guidance while remaining weak in risk framing, or remain cautious while giving insufficient next steps\. The rest of the section therefore looks inside the score profile rather than treating the leaderboard as the endpoint\.
##### Model Profiles
Figure 4:Radar plot showing mean rubric\-item score across scenarios for selected models\. Primary judge = GPT\-5\.2Figure[4](https://arxiv.org/html/2608.14622#S4.F4)compares rubric profiles for six representative models: GPT\-5\.2, GPT\-5 Nano, Kimi K2 Thinking, DeepSeek V3\.1, Ministral 3 14B, and GPT\-4o Mini\. The radar plot shows how overall performance is distributed across the eight rubrics\.
The main observation is that strong models do not share a single profile shape\. GPT\-5\.2 and GPT\-5 Nano are strong and balanced across rubrics\. Kimi K2 Thinking is especially strong on Helpfulness, Empathy, and Communication, while DeepSeek V3\.1 is relatively stronger on Accuracy, Safety, and Bias Avoidance\. GPT\-4o Mini shows larger deficits, especially on Completeness and Limitation Awareness\.
##### Rubric\-Item Comparison
While aggregate rubric scores provide a useful overall comparison between models, examining individual rubric dimensions reveals more specific differences in the kinds of advice models produce\. Full rubric\-level results are available in the appendix\. We focus here on Helpfulness and Safety as representative dimensions because effective parenting advice must balance practical utility with appropriate risk management\. Helpfulness captures whether a response provides clear and actionable guidance, while Safety reflects whether the response appropriately manages risk, avoids harmful suggestions, and sets suitable boundaries\.
Figure[5](https://arxiv.org/html/2608.14622#S4.F5)compares each model’s mean helpfulness and safety scores\. GPT\-5\.2 and GPT\-5 Nano remain strong and relatively balanced on both dimensions\. Kimi K2 Thinking also performs strongly overall, although its gap between helpfulness and safety is more visible than in the top GPT models\. DeepSeek V3\.1 and Qwen3 32B remain reasonably balanced, but their scores are lower than the leading models\. GPT\-4o Mini and Llama 3\.1 8B show broader limitations, with lower scores on both dimensions\.
In summary, these results show that model differences appear both in overall score and in key user\-facing rubrics\.
Figure 5:Dumbbell plot of mean helpfulness and safety score for each model\. Primary judge = GPT\-5\.2
#### 4\.1\.3 Comparison Across Languages
##### Overall Score Differences
Figure[6](https://arxiv.org/html/2608.14622#S4.F6)shows model\-level cross\-language differences in overall score under the primary judge\. The main pattern is heterogeneity\. Language\-conditioned differences are present, but models separate in different directions and by different magnitudes in the English\-versus\-Chinese scatter\.
GLM\-4 9B shows the strongest positive difference in Chinese, with DeepSeek V3\.1 and Qwen3 32B also moving upward\. By contrast, Llama 3\.1 8B and GLM\-4\.6 show the clearest negative differences, with Llama 3\.3 70B, Kimi K2 Thinking, and GPT\-OSS 20B also moving downward\.
These results show a cross\-language difference in advice quality that is strongly model\-dependent\. However, it does not yet explain which parts of the answer changed\.
Figure 6:Scatter plot of overall score by language and model\. Positive delta values and positioning above the dashed line in the scatter plot indicate higher scores in Chinese than English
##### Rubric\-Level Differences
Figure[7](https://arxiv.org/html/2608.14622#S4.F7)shows rubric\-level cross\-language differences under the primary judge\. The clearest pattern is that Chinese answers often increase in Completeness and Limitation Awareness, while declining in Accuracy and Empathy\.
In parenting scenarios, higher Completeness suggests that Chinese responses cover more subpoints or provide more extended explanation, while higher Limitation Awareness suggests more explicit acknowledgement of uncertainty, boundaries, or professional support\. Lower Empathy suggests that this expanded coverage may come with weaker reassurance or emotional attunement with the user\.
Figure 7:Heatmap of rubric\-level cross\-language difference by model\. Primary judge = GPT\-5\.2\. Positive values and red shading indicate higher scores in Chinese than English
##### Scenario\-Level Differences
Figure 8:Bar plot for the 12 scenarios with the highest difference in overall rubric score between English and Chinese, averaged over modelsModel\-level averages show that cross\-language difference exists, but they can hide where the difference occurs\. Some scenarios produce little movement across most models, while others show huge language\-conditioned change\. Figure[8](https://arxiv.org/html/2608.14622#S4.F8)identifies the highest\-difference scenarios under the primary judge\.
The high\-difference scenarios are concentrated in several task types\. Low\-resource safety and health tasks recur, including open\-fire safety, river safety, emergency first aid, and possible hearing loss with limited medical access\. The direction of change is not consistent; in three of these scenarios the Chinese advice is higher quality and in others the English advice is higher quality\.
### 4\.2Parenting Style
#### 4\.2\.1 Responsiveness and Demandingness
Figure[9](https://arxiv.org/html/2608.14622#S4.F9)reports the distribution of Responsiveness and Demandingness scores under the primary judge\. In the figure, the models are ordered top to bottom by their overall rubric score\. The patterns indicate a correlation between overall score and Demandingness, i\.e\., high\-performing models are more likely to give advice providing high levels of structure for the child compared to lower performing models\. The results for Responsiveness are more mixed across models\.
Figure 9:Overall distribution of Responsiveness and Demandingness scores across the scenarios by model\. Primary judge = GPT\-5\.2
#### 4\.2\.2 Style Profiles
Figure 10:Stacked bar plot showing the parenting style profiles of different models across four parenting types \(authoritative, authoritarian, permissive, and neglectful\) and two languages \(English and Chinese\)\. Positive values and red shading in the heatmap indicate that parenting style was more prevalent in Chinese than English for the given modelFigure[10](https://arxiv.org/html/2608.14622#S4.F10)shows the parenting\-style profiles\. The English style space is centred on authoritative advice\. GPT\-5\.2 and Kimi K2 Thinking have the strongest authoritative profiles, while GPT\-5 Nano carries a larger authoritarian component\. Ministral 3 14B has a relatively larger permissive component among the stronger models\. At the lower end, GPT\-4o Mini and GLM\-4 9B show weaker authoritative cores and larger neglectful shares\. Reflecting the patterns in Demandingness across models, Figure[10](https://arxiv.org/html/2608.14622#S4.F10)also indicates a correlation between parenting style and overall score, i\.e\., high\-performing models are more likely to take an authoritative stance compared to lower performing models\.
#### 4\.2\.3 Comparison Across Languages
The model profiles when responding in Chinese shows a different distribution to when responding in English, indicating that language can also affect the advisory stance of the response\. \(Figure[10](https://arxiv.org/html/2608.14622#S4.F10)\)\. GPT\-5\.2 and Kimi K2 Thinking are still the most likely to adopt an authoritative parenting style, but both provide a greater proportion of authoritarian\-style advice\. The columns of the heatmap show an overall pattern in Chinese responses compared to English responses: the parenting stance adopted becomes less authoritative or permissive and more authoritarian or neglectful, suggesting advice becomes lower in warmth \(Responsiveness\) and somewhat higher in imposed structure \(Demandingness\)\. This is particularly the case for GPT\-5 Nano, Qwen3 8B, Ministral 3 14B, and GPT\-OSS 20B\.
### 4\.3Judge Robustness
#### 4\.3\.1 Rubric Judge
\(a\)Overall rubric score
\(b\)Responsiveness
\(c\)Demandingness
Figure 11:Scatter plots of primary and secondary judge scores for each of the 1,500 model\-scenario combinations across English and ChineseThe main results above use GPT\-5\.2 as the primary judge\. This section tests whether the same broad structure remains visible under the secondary judge, DeepSeek V3\.1\. The comparison is made on overlapping evaluation units defined by scenario, answer model, and language\. Figure[11\(a\)](https://arxiv.org/html/2608.14622#S4.F11.sf1)visualises judge agreement on overall mean score\. Across 3,000 overlapping units, the two judges reach a Pearson correlation of 0\.80 and a mean absolute error of 8\.2 points\. This indicates a substantial shared signal, but not identical calibration\. GPT\-5\.2 generally assigns lower overall scores than DeepSeek V3\.1, with the average difference varying by answer model\.
##### Judge\-Sensitive Rubrics
Table[1](https://arxiv.org/html/2608.14622#S4.T1)provides the comparison on the rubric\-item level\. Agreement between judges is strongest for Completeness, Empathy, Helpfulness, and Communication\. Limitation Awareness remains relatively aligned but with a larger spread\. Accuracy, Safety, and Bias Avoidance show weaker agreement and should be read as more judge\-sensitive dimensions\.
Table 1:Rubric\-item level metrics comparing agreement between GPT\-5\.2 and DeepSeek V3\.1 as judgesRubric ItemPearson’srrMAEAccuracy0\.6510\.5Safety0\.618\.1Helpfulness0\.867\.0Empathy0\.8712\.0Completeness0\.899\.1Bias Avoidance0\.593\.6Limitation Awareness0\.8414\.3Communication0\.816\.1
#### 4\.3\.2 Style Judge
Figures[11\(b\)](https://arxiv.org/html/2608.14622#S4.F11.sf2)and[11\(c\)](https://arxiv.org/html/2608.14622#S4.F11.sf3)visualise judge agreement on Responsiveness and Demandingness scores over the 3,000 scenario\-model\-language combinations\. For Responsiveness the judges reach a Pearson correlation of 0\.61 and a mean absolute error of 0\.16 points, and for Demandingness the Pearson correlation is 0\.61 and mean absolute error is 0\.14 points\. Again, this indicates a shared signal, but not identical calibration\. GPT\-5\.2 generally assigns higher Demandingness and lower Responsiveness scores than DeepSeek V3\.1\.
## 5Discussion
### 5\.1Model and Language Variations in Parenting Advice Quality and Style
Our results show that LLM\-generated parenting advice varies substantially across models, languages, and evaluation dimensions, indicating that parenting advice cannot be evaluated adequately through a single aggregate quality score\. While the overall ranking identifies broad performance tiers, the rubric\-level analysis shows that models reach similar total scores through different response profiles\. Some models are stronger in practical guidance and communication, while others are stronger in safety, accuracy, or bias avoidance \(see Section[4\.1](https://arxiv.org/html/2608.14622#S4.SS1)\)\. The identification of mixed profiles is important because parents may have different needs at different moments of advice\-seeking: safety in high\-risk situations, emotional support in times of stress, or specificity when deciding on a course of action\. Aggregate scores may oversimplify individual strength and weaknesses, potentially directly affecting user experience, particularly when the opacity of LLMs is further compounded in build user\-facing applications\.
Our results also suggest that supporting broader parenting advice using LLMs requires a more multi\-dimensional approach\. For example, a given response can be factually acceptable but poorly calibrated emotionally, or appropriately cautious but too vague to be useful\. Existing evaluations of LLMs for parenting and child\-related information have mainly evaluated models in health contexts, where responses can be judged against clear standards of correctness and completeness\(Bushuvenet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib10); McFaydenet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib35); Sezginet al\.[2023](https://arxiv.org/html/2608.14622#bib.bib45)\), and our work extended this\. Our finding aligns with arguments that LLM performance in open\-ended, and thus, socially situated tasks should be assessed in relation to user\-facing qualities rather than only task success or benchmark accuracy\(Xiaoet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib54)\)\.
The cross\-language results further show that advice quality is not stable across language conditions\. Some models improved in Chinese, while others worsened, and these differences were not evenly distributed across rubrics or scenarios\. This highlights that multilingual evaluation should not be treated as a simple translation exercise\. Even when scenario content is fixed, model behaviour may shift in ways that affect both practical advice quality and relational aspects of the response\. These patterns may reflect uneven multilingual training and alignment data\(Grattafioriet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib13)\), English\-centric representations\(Wendleret al\.[2024](https://arxiv.org/html/2608.14622#bib.bib53)\), translation\-mediated data\(Thompsonet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib50); Guoet al\.[2024](https://arxiv.org/html/2608.14622#bib.bib15)\), or language\-specific differences in judging\(Fu and Liu[2025](https://arxiv.org/html/2608.14622#bib.bib12)\)\.
The parenting style analysis adds a further layer here\. In English, higher\-performing models were more likely to produce advice classified as authoritative, which combined warmth towards the child with structure and expectations for them\. In Chinese, many models shifted toward more authoritarian or neglectful profiles, suggesting lower responsiveness and stronger imposed structure\. The origin of this difference is undetermined\. It may reflect cultural patterns in parenting norms, linguistic differences in how advice is conventionally expressed, or model\-specific training and alignment effects\. These results do not indicate that one language condition is better or worse, but they show that language might change the normative stance of the advice, including the implied relationship between parent and child\.
Our findings have implications for the evaluation and deployment of LLMs for parenting advice\. A model that performs well in English cannot be assumed to provide equivalent support in another language\. Similarly, a model that scores well overall may still adopt a style of advice that does not match the values, needs, or context of a particular family\. For parenting advice systems,model selection should therefore consider not only overall quality, but also rubric\-specific strengths and implicit parenting style\. Inconsistent models may be especially concerning: a model that performs well in many scenarios but fails sharply in safety\- or health\-related cases may create greater risk than a consistently limited model whose weaknesses are easier to anticipate\.
### 5\.2Parenting Advice Systems as Relational Technologies
Parenting advice systems should be understood not only as information\-retrieval or question\-answering tools, but also as relational technologies\(Hertlein[2018](https://arxiv.org/html/2608.14622#bib.bib17)\)\. The advice from an LLM may shape how a parent interprets and responds to their child’s behaviour, thereby influencing the relationship between the parent and a child\.
This raises a difficult question: what counts as good parenting advice from an LLM? Advice that aligns with evidence\-based child development knowledge, or advice that is culturally sensitive and responsive to the parents’ values and circumstances? Both are important goals, but may be in tension\. Some culturally familiar or traditional practices may be harmful, while some evidence\-based recommendations may be unachievable in certain settings or communicated in a way that feels alienating or insensitive\. For example, advice about discipline, boundaries, or independence may be received differently depending on cultural expectations, family structure, and beliefs about child\-rearing\. Therefore, a question such as “How do I set boundaries and discipline my child?”\(Kimet al\.[2025](https://arxiv.org/html/2608.14622#bib.bib22)\)cannot be evaluated as if there were a single universally accepted ideal response\.
Our parenting\-style analysis provides one way to make these normative dimensions visible\. We do not claim that authoritative advice is always best or that any arbitrary descriptor of ‘parenting style’ can fully capture the quality of a response\. Rather, it makes explicit that LLMs do not provide value\-neutral parenting guidance—there is no such thing\. The models implicitly recommend particular forms of parent\-child interaction, including different balances of warmth and structure\. This becomes especially relevant when considering alignment efforts, as parenting presents another example of a domain where values are plural, contested, and culturally situated\. A one\-size\-fits\-all model of ‘good parenting advice’ risks privileging one set of norms while presenting them as general guidance\.
When it comes to designing parenting advice systems, they need tosupport pluralistic values and context\-sensitive interaction\. This does not mean systems should simply align with any user preference; safety remains essential, and advice should not normalise violations of children’s rights\. However, within safe boundaries, systems might need to adapt to different family contexts, make assumptions explicit, seeking further information where necessary, and offer a set of options rather than a single prescriptive answer\.
### 5\.3Robust, Auditable Evaluation Pipelines
Through our work we highlight the importance of robust and auditable evaluation pipelines\. Parenting advice is a high\-stakes and socially sensitive domain, and small implementation failures can distort conclusions about model quality\. In this study, response validation and repair were necessary because generation failures, truncation, wrapper text, and reasoning\-like segments could otherwise be mistaken for poor advice quality\. Backend metadata was useful but insufficient on its own, especially across languages—a complete Chinese answer ending with Chinese punctuation may be falsely flagged by English\-style punctuation rules\. This supports the need for evaluation pipelines that record not only final scores, but also the generation, cleaning, repair, and judging steps that produced them\.
Auditability is also important because LLM\-as\-a\-judge evaluation is useful but not definitive\. The two judges showed substantial agreement on overall scores, suggesting that the benchmark captures a meaningful signal\. However, agreement varied by rubric\. Accuracy, Safety, and Bias Avoidance were more judge\-sensitive than Completeness, Empathy, Helpfulness, and Communication\. This pattern is important because the more judge\-sensitive rubrics include dimensions that are central to harm prevention\. A benchmark that reports only aggregate agreement may overstate robustness\. Future evaluations should consider rubric\-level judge agreements when applicable and treat judge comments as useful qualitative evidence rather than merely intermediate outputs\.
## 6Limitations and Future Work
This study has several limitations\. First, although the rubric and scenarios were expert\-informed, the final scoring relied on LLM judges rather than human annotators\. The use of two judges provides a robustness check, but judge agreement varied by rubric, and dimensions such as safety, accuracy, and bias avoidance should be interpreted with particular caution\. Future work should validate the rubric scores and parenting\-style classifications against judgements from parenting experts, parents, and culturally diverse annotators, taking into account that what is regarded as a good answer may differ across families and cultural contexts\.
Second, the evaluation used translated versions of the same 100 scenarios in English and Chinese\. This design supports controlled cross\-language comparison, but it does not fully capture parenting situations that originate within different cultural and linguistic contexts\. Future work should therefore develop language\-specific and culture\-specific scenario sets rather than relying only on translation\.
Third, the pipeline could be improved and scaled by automating remaining manual steps, including review of flagged generations, and refining rubrics and judging prompts through further iterative testing, particularly for culturally sensitive cases and high\-risk scenarios where small differences in framing may have substantial implications for the advice given\.
## 7Conclusion
Overall, this work shows that evaluating LLM parenting advice requires attention to multiple forms of variation: across models, languages, rubric dimensions,parenting styles, and judges\. We developed a multilingual evaluation pipeline that combines expert\-informed scenarios, response generation across 15 models with validation and repair, rubric\-based LLM judging, parenting\-style classification, and automated analysis, all with auditable traces\. The results show that aggregate scores can obscure important differences in the kinds of advice models produce, including their relative strengths and weaknesses in safety, helpfulness, empathy, limitation awareness, and other user\-facing dimensions\. They also show that parenting advice is not value\-neutral: models implicitly recommend different balances of warmth and structure, and these patterns shift across English and Chinese responses\. More broadly, this paper highlights the need for human\-centred benchmarks that capture relational and behavioural aspects of response and make normative assumptions visible\. For socially sensitive advice domains, responsible evaluation requires moving beyond uni\-dimensional leaderboards toward multi\-dimensional and social and cultural context\-aware assessment approaches\.
## References
- ”It Hasn’t Lived in Our Society”: Investigating Cultural Sensitivity in LLM Chatbots for Emotional Support\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,CHI ’26,New York, NY, USA,pp\. 1–20\.External Links:[Document](https://dx.doi.org/10.1145/3772318.3791060),ISBN 979\-8\-4007\-2278\-3Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p2.1)\.
- R\. K\. Arora, J\. Wei, R\. S\. Hicks, P\. Bowman, J\. Quiñonero\-Candela, F\. Tsimpourlas, M\. Sharman, M\. Shah, A\. Vallone, A\. Beutel, J\. Heidecke, and K\. Singhal \(2025\)HealthBench: Evaluating Large Language Models Towards Improved Human Health\.arXiv\.External Links:2505\.08775,[Document](https://dx.doi.org/10.48550/arXiv.2505.08775)Cited by:[§3\.4](https://arxiv.org/html/2608.14622#S3.SS4.SSSx1.p1.1)\.
- S\. Baker, M\. R\. Sanders, and A\. Morawska \(2017\)Who Uses Online Parenting Support? A Cross\-Sectional Survey Exploring Australian Parents’ Internet Use for Parenting\.Journal of Child and Family Studies26\(3\),pp\. 916–927\.External Links:ISSN 1573\-2843,[Document](https://dx.doi.org/10.1007/s10826-016-0608-1)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p1.1)\.
- A\. A\. Baumann, B\. J\. Powell, P\. L\. Kohl, R\. G\. Tabak, V\. Penalba, E\. K\. Proctor, M\. M\. Domenech\-Rodriguez, and L\. J\. Cabassa \(2015\)Cultural adaptation and implementation of evidence\-based parent\-training: A systematic review and critique of guiding evidence\.Children and Youth Services Review53,pp\. 113–120\.External Links:ISSN 0190\-7409,[Document](https://dx.doi.org/10.1016/j.childyouth.2015.03.025)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- D\. Baumrind \(1971\)Current patterns of parental authority\.Developmental Psychology4\(1, Pt\.2\),pp\. 1–103\.External Links:ISSN 1939\-0599,[Document](https://dx.doi.org/10.1037/h0030372)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p3.1),[§3\.4](https://arxiv.org/html/2608.14622#S3.SS4.SSSx2.p2.1)\.
- A\. M\. Bean, R\. O\. Kearns, A\. Romanou, F\. S\. Hafner, H\. Mayne, J\. Batzner, N\. Foroutan, C\. Schmitz, K\. Korgul, H\. Batra, O\. Deb, E\. Beharry, C\. Emde, T\. Foster, A\. Gausen, M\. Grandury, S\. Han, V\. Hofmann, L\. Ibrahim, H\. Kim, H\. R\. Kirk, F\. Lin, G\. K\. Liu, L\. Luettgau, J\. Magomere, J\. Rystrøm, A\. Sotnikova, Y\. Yang, Y\. Zhao, A\. Bibi, A\. Bosselut, R\. Clark, A\. Cohan, J\. Foerster, Y\. Gal, S\. A\. Hale, I\. D\. Raji, C\. Summerfield, P\. H\. S\. Torr, C\. Ududec, L\. Rocher, and A\. Mahdi \(2025\)Measuring what Matters: Construct Validity in Large Language Model Benchmarks\.arXiv\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2511.04703)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p3.1)\.
- S\. Bushuven, M\. Bentele, S\. Bentele, B\. Gerber, J\. Bansbach, J\. Ganter, M\. Trifunovic\-Koenig, and R\. Ranisch \(2023\)”ChatGPT, Can You Help Me Save My Child’s Life?” \- Diagnostic Accuracy and Supportive Capabilities to Lay Rescuers by ChatGPT in Prehospital Basic Life Support and Paediatric Advanced Life Support Cases \- An In\-silico Analysis\.Journal of Medical Systems47\(1\),pp\. 123\.External Links:ISSN 1573\-689X,[Document](https://dx.doi.org/10.1007/s10916-023-02019-x)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p2.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p2.1),[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p2.1)\.
- J\. Dworkin, J\. Connell, and J\. Doty \(2013\)A literature review of parents’ online behavior\.Cyberpsychology: Journal of Psychosocial Research on Cyberspace7\(2\)\.External Links:ISSN 1802\-7962,[Document](https://dx.doi.org/10.5817/CP2013-2-2)Cited by:[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p1.1)\.
- X\. Fu and W\. Liu \(2025\)How Reliable is Multilingual LLM\-as\-a\-Judge?\.Note:https://arxiv\.org/abs/2505\.12201v1Cited by:[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p3.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. De Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The Llama 3 Herd of Models\.Note:https://arxiv\.org/abs/2407\.21783v3Cited by:[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p3.1)\.
- J\. E\. Grusec, T\. Danyliuk, H\. Kil, and D\. O’Neill \(2017\)Perspectives on parent discipline and child outcomes\.International Journal of Behavioral Development41\(4\),pp\. 465–471\.External Links:ISSN 0165\-0254,[Document](https://dx.doi.org/10.1177/0165025416681538)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- Y\. Guo, S\. Conia, Z\. Zhou, M\. Li, S\. Potdar, and H\. Xiao \(2024\)Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs\.Note:https://arxiv\.org/abs/2410\.15956v3Cited by:[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p3.1)\.
- S\. Harkness and C\. M\. Super \(2020\)Why understanding culture is essential for supporting children and families\.Applied Developmental Science25\(1\),pp\. 14–25\.External Links:ISSN 1088\-8691,[Document](https://dx.doi.org/10.1080/10888691.2020.1789354)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- K\. M\. Hertlein \(2018\)Technology in Relational Systems: Roles, Rules, and Boundaries\.InFamilies and Technology,J\. Van Hook, S\. M\. McHale, and V\. King \(Eds\.\),pp\. 89–102\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-95540-7%5F5),ISBN 978\-3\-319\-95540\-7Cited by:[§5\.2](https://arxiv.org/html/2608.14622#S5.SS2.p1.1)\.
- Y\. Kim, S\. L\. Vilches, S\. Shapiro, and A\. Clarkson \(2025\)Testing the capability of generative artificial intelligence for parent and caregiver information seeking\.Family Relations74\(3\),pp\. 1266–1284\.External Links:ISSN 1741\-3729,[Document](https://dx.doi.org/10.1111/fare.13167)Cited by:[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p3.1),[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p2.1),[§5\.2](https://arxiv.org/html/2608.14622#S5.SS2.p2.1)\.
- H\. R\. Kirk, A\. Whitefield, P\. Röttger, A\. Bean, K\. Margatina, J\. Ciro, R\. Mosquera, M\. Bartolo, A\. Williams, H\. He, B\. Vidgen, and S\. Hale \(2024\)The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models\.Advances in Neural Information Processing Systems37,pp\. 105236–105344\.External Links:[Document](https://dx.doi.org/10.52202/079017-3342)Cited by:[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p1.1)\.
- R\. Koo, M\. Lee, V\. Raheja, J\. I\. Park, Z\. M\. Kim, and D\. Kang \(2024\)Benchmarking Cognitive Biases in Large Language Models as Evaluators\.arXiv\.External Links:2309\.17012,[Document](https://dx.doi.org/10.48550/arXiv.2309.17012)Cited by:[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p4.1)\.
- J\. E\. Lansford \(2022\)Annual Research Review: Cross\-cultural similarities and differences in parenting\.Journal of Child Psychology and Psychiatry63\(4\),pp\. 466–479\.External Links:ISSN 1469\-7610,[Document](https://dx.doi.org/10.1111/jcpp.13539)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- C\. J\. Leslie\-Miller, S\. L\. Simon, K\. Dean, N\. Mokhallati, and C\. C\. Cushing \(2024\)The critical need for expert oversight of ChatGPT: Prompt engineering for safeguarding child healthcare information\.Journal of Pediatric Psychology49\(11\),pp\. 812–817\.External Links:ISSN 0146\-8693, 1465\-735X,[Document](https://dx.doi.org/10.1093/jpepsy/jsae075)Cited by:[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p3.1)\.
- G\. Lin, M\. Mikolajczak, H\. Keller, E\. Akgun, G\. Arikan, K\. Aunola, E\. Barham, E\. Besson, M\. A\. Blanchard, E\. Boujut, M\. E\. Brianda, A\. Brytek\-Matera, F\. César, B\. Chen, G\. Dorard, L\. C\. dos Santos Elias, S\. Dunsmuir, N\. Egorova, M\. J\. Escobar, N\. Favez, A\. M\. Fontaine, H\. Foran, K\. Furutani, M\. Gannagé, M\. Gaspar, L\. Godbout, A\. Goldenberg, J\. J\. Gross, M\. A\. Gurza, O\. Hatta, A\. Heeren, M\. Helmy, M\. Huynh, E\. Kaneza, T\. Kawamoto, N\. Kellou, B\. L\. Kpassagou, L\. Lazarevic, S\. Le Vigouroux, A\. Lebert\-Charron, V\. Leme, C\. MacCann, D\. Manrique\-Millones, O\. Medjahdi, R\. B\. Millones Rivalles, M\. I\. Miranda Orrego, M\. Miscioscia, S\. F\. Mousavi, B\. Moutassem\-Mimouni, H\. Murphy, A\. Ndayizigiye, T\. J\. Ngnombouowo, S\. Olderbak, S\. Ornawka, D\. O\. Cádiz, P\. A\. Pérez\-Díaz, K\. Petrides, A\. Prikhidko, F\. Salinas\-Quiroz, M\. Santelices, C\. Schrooyen, P\. Silva, A\. Simonelli, M\. Sorkkila, E\. Stănculescu, E\. Starchenkova, D\. Szczygieł, J\. Tapia, M\. Tremblay, T\. M\. T\. Tri, A\. M\. Üstündağ\-Budak, M\. Valdés Pacheco, H\. van Bakel, L\. Verhofstadt, J\. Wendland, S\. Yotanyamaneewong, and I\. Roskam \(2023\)Parenting Culture\(s\): Ideal\-Parent Beliefs Across 37 Countries\.Journal of Cross\-Cultural Psychology54\(1\),pp\. 4–24\.External Links:ISSN 0022\-0221,[Document](https://dx.doi.org/10.1177/00220221221123043)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- E\. E\. Maccoby and J\. A\. Martin \(1983\)Socialization in the context of the family: Parent\-child interaction\.InHandbook of Child Psychology: Vol\. 4: Socialization, Personality and Social Development ; E\. Mavis Hetherington, Volume Editor,P\. H\. Mussen \(Series Ed\.\) and E\. M\. Hetherington \(Vol\. Ed\.\) \(Eds\.\),pp\. 1–101\.Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p3.1),[§3\.4](https://arxiv.org/html/2608.14622#S3.SS4.SSSx2.p2.1)\.
- T\. C\. McFayden, S\. Bristol, O\. Putnam, and C\. Harrop \(2024\)ChatGPT: Artificial Intelligence as a Potential Tool for Parents Seeking Information About Autism\.Cyberpsychology, Behavior, and Social Networking27\(2\),pp\. 135–148\.External Links:ISSN 2152\-2715, 2152\-2723,[Document](https://dx.doi.org/10.1089/cyber.2023.0202)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p3.1),[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p2.1)\.
- E\. Mertens, G\. Ye, E\. Beuckels, and L\. Hudders \(2024\)Parenting Information on Social Media: Systematic Literature Review\.JMIR Pediatrics and Parenting7\(1\),pp\. e55372\.External Links:[Document](https://dx.doi.org/10.2196/55372)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p1.1)\.
- D\. Rudy and J\. E\. Grusec \(2001\)Correlates of Authoritarian Parenting in Individualist and Collectivist Cultures and Implications for Understanding the Transmission of Values\.Journal of Cross\-Cultural Psychology32\(2\),pp\. 202–212\.External Links:ISSN 0022\-0221,[Document](https://dx.doi.org/10.1177/0022022101032002007)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- B\. E\. Rutigliano, A\. L\. Randolph, and C\. N\. Park \(2023\)Understanding Parents’ Self\-Awareness of Their Parenting Style\(s\) and Its Influences on Their Parenting Choices \- A Grounded Theory Study\.The Family Journal31\(3\),pp\. 385–391\.External Links:ISSN 1066\-4807,[Document](https://dx.doi.org/10.1177/10664807231163268)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- S\. Schilling, A\. Mebane, and K\. M\. Perreira \(2021\)Cultural Adaptation of Group Parenting Programs: Review of the Literature and Recommendations for Best Practices\.Family Process60\(4\),pp\. 1134–1151\.External Links:ISSN 1545\-5300,[Document](https://dx.doi.org/10.1111/famp.12658)Cited by:[§2\.3](https://arxiv.org/html/2608.14622#S2.SS3.p1.1)\.
- E\. Sezgin, F\. Chekeni, J\. Lee, and S\. Keim \(2023\)Clinical Accuracy of Large Language Models and Google Search Responses to Postpartum Depression Questions: Cross\-Sectional Study\.Journal of Medical Internet Research25,pp\. e49240\.External Links:ISSN 1438\-8871,[Document](https://dx.doi.org/10.2196/49240)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p2.1)\.
- S\. Sharmin and S\. Afrin \(2026\)Avoiding Social Judgment, Seeking Privacy: Investigating why Mothers Shift from Facebook Groups to Large Language Models\.arXiv\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2602.13941)Cited by:[§1](https://arxiv.org/html/2608.14622#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p3.1)\.
- R\. Stureborg, D\. Alikaniotis, and Y\. Suhara \(2024\)Large Language Models are Inconsistent and Biased Evaluators\.arXiv\.External Links:2405\.01724,[Document](https://dx.doi.org/10.48550/arXiv.2405.01724)Cited by:[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p4.1)\.
- B\. Thompson, M\. P\. Dhaliwal, P\. Frisch, T\. Domhan, and M\. Federico \(2024\)A Shocking Amount of the Web is Machine Translated: Insights from Multi\-Way Parallelism\.Note:https://arxiv\.org/abs/2401\.05749v2Cited by:[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p3.1)\.
- C\. Wendler, V\. Veselovsky, G\. Monea, and R\. West \(2024\)Do Llamas Work in English? On the Latent Language of Multilingual Transformers\.Note:https://arxiv\.org/abs/2402\.10588v4Cited by:[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p3.1)\.
- Z\. Xiao, W\. H\. Deng, M\. S\. Lam, M\. Eslami, J\. Kim, M\. Lee, and Q\. V\. Liao \(2024\)Human\-Centered Evaluation and Auditing of Language Models\.InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems,CHI EA ’24,New York, NY, USA,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1145/3613905.3636302),ISBN 979\-8\-4007\-0331\-7Cited by:[§2\.2](https://arxiv.org/html/2608.14622#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2608.14622#S5.SS1.p2.1)\.
- H\. S\. Yun and T\. Bickmore \(2025\)Online Health Information–Seeking in the Era of Large Language Models: Cross\-Sectional Web\-Based Survey Study\.Journal of Medical Internet Research27\(1\),pp\. e68560\.External Links:[Document](https://dx.doi.org/10.2196/68560)Cited by:[§2\.1](https://arxiv.org/html/2608.14622#S2.SS1.p2.1)\.
## Appendix AExtended Results
### A\.1Rubric\-Level Means
Table 2:Rubric\-item means by modelModelAccuracySafetyHelpfulnessEmpathyComplete\-nessBiasAvoidanceLimitationAwarenessCommuni\-cationGPT\-5\.291\.691\.492\.587\.390\.090\.187\.493\.1GPT\-5 Nano87\.988\.990\.282\.987\.489\.283\.789\.9Kimi K2 Thinking86\.686\.790\.487\.984\.487\.877\.790\.7DeepSeek V3\.186\.287\.485\.684\.377\.289\.471\.388\.4MiniMax M284\.085\.486\.481\.279\.787\.773\.988\.0Ministral 3 14B81\.482\.085\.785\.379\.986\.673\.587\.5Qwen3 32B82\.984\.882\.079\.874\.888\.470\.686\.2GLM\-4\.681\.884\.581\.380\.873\.487\.866\.085\.2GPT\-OSS 20B80\.480\.485\.979\.379\.787\.070\.788\.0DeepSeek R180\.683\.080\.383\.573\.287\.167\.985\.0Qwen3 8B79\.680\.380\.174\.772\.086\.767\.883\.8Llama 3\.3 70B79\.482\.075\.172\.267\.286\.964\.482\.1GPT\-4o Mini79\.582\.872\.969\.763\.586\.656\.280\.9Llama 3\.1 8B72\.775\.769\.571\.162\.284\.059\.478\.6GLM\-4 9B73\.878\.364\.970\.357\.984\.053\.575\.8Similar Articles
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
This paper presents HumRightsBench, the first expert-validated, scenario-based benchmark for evaluating large language models' legal reasoning about international human rights law. Pilot results on frontier models show overall accuracy between 0.339 and 0.577, highlighting notable gaps in detecting obligation violations.
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.
Benchmarking LLMs
A study or report on benchmarking large language models, likely comparing performance across various tasks.
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
DLawBench is a new benchmark for evaluating large language models in multi-turn legal consultation, covering Chinese and US law with four client types. Experiments show significant room for improvement, with the best model achieving only 0.562 on legal reasoning.
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
CulturALL introduces a 2,610-sample benchmark across 14 languages and 51 regions to evaluate LLMs on real-world, culturally grounded tasks; top model scores only 44.48%, highlighting large room for improvement.