Tag
Arcee AI, Loka, AWS, and Prime Intellect post-trained an open model using reinforcement learning to improve scientific tool use and biological reasoning, achieving notable gains on drug tool and Gene Ontology benchmarks.
This paper investigates how post-training stages such as continued pre-training, supervised fine-tuning, and reinforcement learning affect generalization in biological reasoning models, finding that these stages have distinct impacts on in-domain and out-of-domain performance.