Medical model: Reasoning-Medical-27B (Qwen3.6-27B finetune)
Summary
Reasoning-Medical-27B is a fine-tuned Qwen3.6-27B model for advanced medical reasoning, trained on 370k Q&A examples with Chain-of-Thought reasoning using GRPO and Unsloth optimization.
Similar Articles
Reasoning-Medical0.1-27B (Qwen3.5-27B medical finetune, claims to surpass MedGemma)
EpistemeAI released Reasoning-Medical0.1-27B, a fine-tuned version of Qwen3.5-27B for medical reasoning, claiming to surpass MedGemma on several medical benchmarks by incorporating chain-of-thought reasoning on a curated dataset of 100,000 records.
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
MedGuideX transforms clinical practice guidelines into executable decision logic to generate factual and counterfactual QA data for training medical LLMs, achieving a 10.28% relative improvement in average accuracy across clinical reasoning benchmarks.
Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering
This paper introduces a multi-agent peer-reviewed reasoning method where multiple LLMs independently generate chain-of-thought reasoning and then evaluate each other's outputs to select the best answer. The method outperforms single-model reasoning and majority voting on medical QA benchmarks.
Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6
The article discusses Qwen 3.8 27B, a 27B parameter model that uses extensive reasoning tokens to compete with larger models, emphasizing trade-offs in token usage and benefits for local deployment.
1.7B model leading strict-7 formal reasoning above Qwen3-8B and Gemma-4-26B - specialists eating generalist territory?
TwIL-LM2, a specialized 1.7B model fine-tuned for formal logic translation, outperforms larger generalist models like Qwen3-8B and Gemma-4-26B on strict scoring benchmarks, highlighting the potential of narrow AI specialists for efficient reasoning.