An integer programming-based approach to construct exact two-sample binomial tests with maximum power
Comparing binomial proportions is a foundational task in clinical trials; however, practitioners often use a liberal likelihood-based test or the conservative Fisher's exact test. To increase power without sacrificing validity, we use a linear integer program to construct finite-sample exact binomial tests with maximum power. Our proposed Average Power Knapsack (APK) test finds a decision boundary that maximizes average power, guaranteeing it cannot uniformly be improved in terms of power, while enforcing type I error rate control across the null hypothesis parameter space using Lipschitz continuity. Our numerical evaluation reveals consistent pointwise power gains of up to 35% for the APK test over Fisher's exact test, as well as systematic power improvements over state-of-the-art exact Berger and Boos implementations of the mid-p-value and Z-pooled tests. When compared against non-exact procedures exhibiting moderate type I error rate inflation, the exact APK test maintains comparable power. For sample sizes where APK construction becomes computationally intensive, Fisher's mid-p-value test emerges as the best non-exact alternative. To support practical application, we developed an interactive R Shiny tool that enables seamless implementation of our optimized exact tests for binomial proportions.