Tag
Luth-2 releases two French small language models (0.8B and 2B) that achieve state-of-the-art results on French benchmarks for their size, with open-weights and data on Hugging Face.
OntoBook converts medical ontology graphs into synthetic textbook prose using LLMs, then uses the resulting 1.3M French textbooks to pretrain ModernCamemBERT, achieving significant gains on medical coding benchmarks.
OpenLLM-France releases Luciole-23B-Instruct-1.1, an open-source multilingual instruction-tuned language model under Apache 2.0, with smaller 8B and 1B variants also available.
User shares their experience learning French in 33 days using Claude AI, and provides 6 prompts for others to try.
This paper introduces a French OSCE dialogue dataset of 240 interactions and a controllable LLM-based pipeline for generating synthetic OSCE dialogues, enabling realistic virtual patient simulations for medical training with automatic feedback.
This paper compares the geometric structures induced by deep learning vector embeddings (CamemBERT) and lexical co-occurrence graph models on the French 'Great National Debate' corpus, finding similar local topology but distinct global organization, highlighting complementarity between the two approaches.
Describes the conversion of French verb Lexicon-Grammar tables into the LMF format, enhancing interoperability and standardization for NLP dictionaries.