arXiv Science⌕ Search

arXiv · 2609.38693

Batting Average as the Product of Two Rates: Skill, Luck, and the Disappearance of the .400 Hitter

Abstract

Batting average factors exactly as BA = c times f, where c = (AB - SO)/AB is the rate of avoiding a strikeout and f = H/(AB - SO) is the rate at which non-strikeout at-bats become hits. Using the 256 major-league hitters with at least 300 at-bats in 2025, we show that the two factors behave very differently. Strikeout avoidance is highly repeatable, with a median year-to-year correlation of 0.86 over 19 consecutive-season pairs from 2004 to 2025. The finishing rate is not: its median correlation is 0.44, and batting average itself (0.44) is no more repeatable than its noisier factor. A bivariate logistic-normal random-effects model fit to the 2025 season estimates the correlation between the two talents at -0.68 (95\% profile interval -0.85 to -0.50), far stronger than the raw correlation of -0.44, and implies single-season reliabilities of 0.91 for c and 0.37 for f. The model also yields a closed-form bivariate shrinkage estimator in which a hitter's strikeout rate informs the estimate of his finishing rate. Applied to decades of American and National League data, the decomposition revisits Gould's explanation for the disappearance of the .400 hitter. Since the dead-ball era, the talent variance of strikeout avoidance has grown roughly 4.6-fold while that of finishing has halved, and the correlation between them has moved from near zero to about -0.53. That emerging trade-off, rather than a general narrowing of talent, accounts for the reduced spread of modern batting averages.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jim Albert. 2026-09-30. Batting Average as the Product of Two Rates: Skill, Luck, and the Disappearance of the .400 Hitter. https://arxiv.org/abs/2609.38693

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A General Design-Based Framework and Estimator for Randomized Experiments

We describe a widely applicable design-based framework for drawing causal inference in randomized experiments. Causal effects are defined as linear functionals evaluated at unit-level potential outcome functions. Assumptions about the potential outcome functions are encoded as function spaces. This makes the framework expressive, allowing experimenters to formulate and investigate a wide range of causal questions that previously could not be investigated with design-based methods. The framework is particularly suited for complex, non-discrete interventions and causal interference. We describe a class of estimators for estimands defined using the framework and investigate their properties. We provide necessary and sufficient conditions for unbiasedness and consistency. We also describe a class of conservative variance estimators, which facilitate the construction of confidence intervals. In order to demonstrate the value of our approach in practice, we provide several illustrative examples of causal investigations which can be handled within our framework, but that could not be addressed using conventional design-based methods.

stat.ME↗

Bayesian Neural-Net-Assisted Multi-Treatment Mixture Cure Survival Model with Application in Pediatric Oncology

Estimating covariate-conditional treatment effects in multi-arm oncology studies is complicated when treatment arms have common distributional features and a non-negligible fraction of patients achieve long-term remission. We propose a joint mixture cure model with covariate-dependent mixtures of log-normal kernels with treatment-specific inclusions. Both linear and neural-network-assisted non-linear covariate links are proposed. Specifically, the susceptible survival distributions use a common finite dictionary of log-normal components, and then a binary inclusion matrix determines which components are active in each treatment arm. All parameters, including the hidden bases of the neural network, are learned jointly, while the output coefficients remain treatment- or component-specific. Posterior inference is performed using gradient-based MCMC, and treatment effects are summarized by covariate-conditional differences in restricted mean survival time (RMST). Variable importance is assessed using thresholded marginal best linear projections with data partitioning. Across two simulation settings, the proposed method demonstrates good finite-sample performance, with lower RMST-based estimation error than flexsurvcure. Compared with pairwise grf fits, the proposed method yields lower RMST-contrast MSE in most comparisons while ensuring mutually coherent multi-treatment contrasts. Finally, the application to the AALL0434 trial reveals covariate-dependent patterns in RMST posterior across methotrexate-based regimens and provides new insights into how these differences vary with patient covariates, highlighting the method's practical utility for studying heterogeneous treatment effects in pediatric oncology trials.

stat.ME↗

Slice Monte Carlo Integration

Numerical integration involving expensive target functions is a common bottleneck in Bayesian inference and simulation. When a cheap surrogate is available, standard approaches such as reweighting or importance sampling often suffer from high variance and inefficient use of function evaluations. We introduce Slice Monte Carlo integration (S$\ell$MC), a method that leverages a Nested Sampling-like procedure on the surrogate to partition the space into informative strata, or slices, while generating samples in the parameter space drawn from the prior within each slice. This enables stratified Monte Carlo integration of the expensive target function over the surrogate-induced partition, yielding an efficient estimate of the target integral. The surrogate level sets therefore make the induced ordering of the parameter space the relevant information for S$\ell$MC, rather than the pointwise target-to-surrogate ratios governing importance-sampling efficiency. Another key advantage of S$\ell$MC is the decoupling of slice volume estimation from the evaluation of the target, allowing the refinement of the slice-volume estimates without additional target evaluations. We investigate the properties of S$\ell$MC, demonstrate how to efficiently generate posterior samples, and assess its robustness in controlled Gaussian benchmark families with continuously tunable surrogate-target mismatch.

stat.ME↗