arXiv ScienceSearch

arXiv subjects

Shibshankar Dey

Publications and source records attributed to Shibshankar Dey.

5 recordsLinked to original sources

Robust Chance-Constrained Optimization using a Continuous Parameter Space Wasserstein-2 Ambiguity Set of Gaussian Mixtures

We study distributionally robust linear chance-constrained problems in which uncertainty is modeled by a Gaussian mixture model (GMM). Finite-support distributionally robust (FDR) formulations, widely used in data-driven robust optimization, robustify over empirical mixture support points and therefore primarily stress-test the fitted nominal mixture. This can be insufficient when service reliability depends on structural misspecification of the nominal mixture-support parameters. To address this limitation, we describe the ambiguity set of distributions by developing a novel formulation of a Wasserstein-2 metric that uses the Bures-Wasserstein (BW) metric over probability measures with finite second moments. Unlike FDR, which generally sets finitely many empirical support points a priori, the proposed ambiguity set allows the worst-case distribution to endogenously determine both how many mixture components receive mass and where their means and covariances lie within a continuous support. For the resulting ambiguity set, under mild regularity conditions, we prove strong duality for the inner worst-case chance-constraint problem and derive its semi-infinite reformulation. We then develop an adaptive cutting-surface algorithm, which endogenously determines the locations of mixture components receiving mass, and the mean and covariances of the Gaussian distributions at these locations. The algorithm attains any prescribed optimality gap in finitely many iterations, while a block-alternating local search identifies new components. A case study using the electric-vehicle charging-station energy-allocation problem demonstrates the framework's practical value in achieving any reliability targets. CDR also induces structural changes in energy allocations, unlike FDR, whose allocations remain close to the nominal solution.

math.OC

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce initialization memory: the dependence of the validation-selected predictor on the scale of the random initialization. We perform controlled CIFAR-10 experiments on ResNets where initialization memory already sharply separates training regimes. Low-learning-rate SGD can interpolate while still remembering its initialization: on ResNet-9 with batch size $b=128$, test accuracy varies by $26.5$ percentage points across initialization scales despite $\ge99.5\%$ training accuracy. This is not undertraining: extending the same low-learning-rate regime to $5{,}000$ epochs leaves the spread essentially unchanged. In contrast, Adam-family methods largely erase the dependence. SGD can also be made to forget when larger learning rates are paired with explicit $L_2$ norm control. We interpret these findings in terms of the time scale of forgetting: gradient-flow-like dynamics can preserve initialization memory, whereas stochastic finite-step effects, explicit norm decay, and adaptive preconditioning erase it on scales governed by the size of explicit or implicit regularization. The practical inductive bias of a trained network is therefore not the architectural prior alone, but the architectural prior after being filtered by the forgetting dynamics of the training pipeline; and the same regularizers that improve generalization are precisely those that erase memory of initialization.

cs.LG

On Solving Chance-Constrained Models with Gaussian Mixture Distribution

We study linear chance-constrained problems where the coefficients follow a Gaussian mixture distribution. We provide mixed-binary quadratic programs that give inner and outer approximations of the chance constraint based on piecewise linear approximations of the standard normal cumulative density function. We show that $O\left(\sqrt{\ln(1/τ)/τ} \right)$ pieces are sufficient to attain $τ$-accuracy in the chance constraint. We also show that any desired optimality gap can be achieved under a constraint qualification condition by controlling the approximation accuracy. Extensive computations using a commercial solver show that problems with up to one thousand random coefficients specified with up to fifteen Gaussian mixture components, generated under diverse settings, can be solved to near optimality within 18 hours, while satisfying chance constraint satisfaction probabilities of up to $0.999$. The solution times are significantly lower for problems with fewer random coefficients and mixture terms. For example, problems with one hundred random coefficients, ten mixture terms, and a constraint satisfaction probability of $0.999$ can be solved in a minute or less. Sample average approximations fail to provide meaningful solutions even for the smaller problems.

math.OC

A Cutting-plane and Benders' Decomposition Algorithm for Two-Stage Distributionally Robust Convex programs

We present a finitely convergent cutting-plane algorithm for solving a general mixed-integer convex program given an oracle for solving a general convex program. This method is extended to solve a family of two-stage mixed-integer convex programs using cutting planes, with applications to solving distributionally-robust two-stage stochastic mixed-integer convex programs. Analysis is also given for the case where convex programming oracle provides an $epsilon$-optimal solution. We combine the cut generation with a branch-and-union scheme to develop a more practical algorithm. Computational results on generated test problems show the practicality of our algorithm. Specifically, results show that in the tested problems our algorithm achieves < 5% optimality gap in 12 hours. This gap is >17% with a commercial solver.

math.OC

Optimization Modeling for Pandemic Vaccine Supply Chain Management: A Review and Future Research Opportunities

During various stages of the COVID-19 pandemic, countries implemented diverse vaccine management approaches, influenced by variations in infrastructure and socio-economic conditions. This article provides a comprehensive overview of optimization models developed by the research community throughout the COVID-19 era, aimed at enhancing vaccine distribution and establishing a standardized framework for future pandemic preparedness. These models address critical issues such as site selection, inventory management, allocation strategies, distribution logistics, and route optimization encountered during the COVID-19 crisis. A unified framework is employed to describe the models, emphasizing their integration with epidemiological models to facilitate a holistic understanding. This article also summarizes evolving nature of literature, relevant research gaps, and authors' perspectives for model selection. Finally, future research scopes are detailed both in the context of modeling and solutions approaches.

math.OC