Tag
This paper introduces ConRub-Med, a reinforcement learning approach that uses consensus rubrics from multiple language models to reward open-ended medical question answering, achieving state-of-the-art results on several benchmarks including HealthBench-Hard.
Small open-weight 4B LLMs achieve up to 87% accuracy on Swedish medical licensing exam questions, approaching o3-level performance with reasoning enabled and post-training techniques.
Hybrid-IR introduces a dual-path retrieval framework combining graph-based and dense retrieval with iterative reasoning to improve complex medical QA, addressing limitations in existing RAG methods. Experiments on three benchmarks show effectiveness.