What happens when frontier LLMs are deployed in rural Rwanda? Lessons on usefulness, language gaps, and incorrect answers [D]
Summary
GiveDirectly's pilot in rural Rwanda paired unconditional cash transfers with a general-purpose AI chatbot, revealing both value as an always-available advisor and critical limitations including language gaps, irrelevant responses, and confidently incorrect answers, raising questions about evaluating models beyond benchmarks.
Similar Articles
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
This paper examines the deployment of an LLM-based machine translation system for text-to-911 emergency services, highlighting common misconceptions and providing recommendations for stakeholders to ensure safe and effective use of AI in critical contexts.
Using LLMs
The article reflects on the limited use and understanding of LLMs such as ChatGPT among personal acquaintances, raising questions about widespread AI tool adoption.
Evaluated a RAG chatbot and the most expensive model was the worst performer. Notes on what actually moved the needle.
A detailed evaluation of a RAG customer support chatbot reveals that retrieval issues often masquerade as LLM problems, heuristic evaluators are misleading, deduplication improves quality, stricter grounding trades helpfulness for accuracy, and model sweeping can dramatically reduce cost while improving performance.
Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication
This benchmark study evaluates 46 large language models against human experts for coding qualitative humanitarian data, finding that LLMs can achieve comparable reliability with structured prompts and reasoning, but require careful oversight for nuanced themes.
Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration
This paper proposes the RP-RCAF prompting strategy to generate culturally sensitive mental health advice in low-resource languages, and introduces the G-REFS evaluation framework, showing significant improvement over conventional prompting across multiple LLMs.