arXiv · 2604.19060
Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports
Abstract
Accurate disease classification from radiology reports is essential for many applications. While supervised fine-tuning (SFT) of lightweight LLMs improves accuracy, it can degrade reasoning. We propose a two-stage approach: SFT on disease labels followed by Group Relative Policy Optimization (GRPO) to refine predictions by optimizing accuracy and format without reasoning supervision. Across three radiologist-annotated datasets, SFT outperformed baselines and GRPO further improved classification and enhanced reasoning recall and comprehensiveness.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yishu Wei, Yi Lin, Adam Flanders, George Shih, Yifan Peng. 2026-04-21. Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports. https://arxiv.org/abs/2604.19060
Cite the original work for its findings. Save a collection to share your selection of sources.