arXiv ScienceSearch

arXiv subjects

Shan Yang

Publications and source records attributed to Shan Yang.

1 recordsLinked to original sources

Physics-R1: An Audited Olympiad Corpus and Released Verifiers for Visual Physics Reasoning

Trackable improvement in multimodal physics reasoning rests on a training-and-evaluation system that is itself rarely verified: the corpora a model trains on, the reward it is optimized against, and the benchmarks and judges that score it. We audit this system end to end and find that standard construction practices systematically distort measurement: contamination slips past n-gram deduplication, translation degrades problems, saturated multiple-choice formats overstate capability, partial-credit training rewards are easier to exploit than to earn, and open-ended grading silently depends on the choice of judge. Left unverified, these distortions inflate reported progress and leak test knowledge into training. We answer with a released verifier system: a three-stage contamination audit that certifies the train/test boundary behind an audited multimodal training corpus and a held-out olympiad benchmark; a binary answer verifier that supplies the reinforcement-learning training reward; and an answer-judging harness that brackets every open-ended score between a deterministic strict layer and a large-language-model liberal layer. All per-record verdicts are released and cross-checked against an independent open-weight judge, whose substitution shifts absolute scores but preserves the sign of every base-to-trained lift. Training against the binary answer verifier confirms the certified corpus supports training: a reference recipe lifts an 8B open-source base by 18.3 points on the held-out benchmark across three seeds. The simple binary reward also beats a dense partial-credit variant on three of four open-ended benchmarks, tying the fourth. All verifiers, verdicts, and datasets are public.

cs.CL