When2Think: Learning When and How Much to Reason
Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly control whether}to reason and how much computation to allocate within reasoning. We call the resulting difficulty-dependent loss in accuracy under computation reduction the efficiency tax. We propose When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation. Its core mechanism, Instance-level Difficulty-Aware Control (IDAC), uses cached reference statistics of success and token cost to modulate a correctness-gated efficiency bonus based on generated token count. Importance sampling supports exploration of Think and NoThink, while Batch-Wise Standardization constructs standardized advantages for critic-free optimization. The framework requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates. On AIME24, When2Think improves Pass@3 by 10.0 percentage points while reducing token usage by 27.9% relative to the backbone.