arXiv Science⌕ Search

arXiv subjects

Chung-Piaw Teo

Publications and source records attributed to Chung-Piaw Teo.

9 recordsLinked to original sources

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

Large language models (LLMs) can translate natural-language problem descriptions into optimization code, but the code is prone to silent failures: it executes and returns a solver-feasible solution while encoding a semantically incorrect formulation. On compositional problems, the resulting feasibility-correctness gap reaches 90 percentage points. We introduce ReLoop, which combines two mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify) to reduce formulation errors during generation. Behavioral verification detects the errors that remain by testing whether the formulation responds correctly to solver-based parameter perturbation, a signal that comes from the solver rather than from LLM self-review and requires no ground truth. The two mechanisms address different error structures: structured generation gives the largest gain on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), and behavioral verification gives its largest gain on localized defects (+4.4pp on MAMO-ComplexLP). With diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6, and relative to direct generation it raises or preserves every reported metric of the three chat-tuned foundation models on all three benchmarks. For the narrowly fine-tuned SFT model we test, the chain-of-thought prompt conflicts with its learned output format and lowers its accuracy on MAMO-ComplexLP; we document and analyze this interaction. We release RetailOpt-190, 190 compositional retail optimization scenarios in which several constraints interact.

cs.SE↗

Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth answer to decide with. Existing learners admit trajectories by matching known optima or labels, and label-free substitutes such as execution success or agreement at one instance can admit wrong models. We introduce ADMITOR, a label-free admission gate. It generates models from three model families, runs each on the stated problem and on instances with resampled parameters, keeps the largest group of models whose optimal values agree on every instance across families, and applies a threshold fitted on solver-verified problems to accept, abstain, or escalate, with a finite-sample bound on the false-discovery rate among accepted values. Inside a state-of-the-art skill learner, ADMITOR raises candidate-level admission precision to 0.927, against 0.871 for majority vote over the host's own samples and 0.726 for execution success, and its library, the smallest of the four, reaches the highest macro accuracy over five public benchmarks, 58.4 against 54.8 for majority vote. An ablation on the same records shows that the gain comes from the accepted value being external to the learner and unanimous across families; on this stream, resampling never changed an accepted value and only reduced coverage. The false-discovery bound holds on the calibration set but not on the benchmark stream: an audit of every false certificate traces most of them to benchmark texts that omit or round the numbers needed to reproduce the labeled answer, and a label-free check of the extracted numbers against the text flags most of these cases.

cs.AI↗

Optimizer as Detector: Stochastic Gradient Descent for Latent Mixture Models

Pooling latent subpopulations can obscure relationships and yield misleading regression conclusions, including Simpson's paradox (SP). We propose a detector based on the steady-state dynamics of constant-step stochastic gradient descent (SGD). Unlike likelihood-based mixture tests and confounder-search methods, it requires neither a normal-mixture specification nor observed candidate confounders. Two linear pieces compete under a winner-take-all squared-error loss, and their normalized terminal separation forms the test statistic. Using diffusion approximations, we derive its asymptotic null distribution for general centered scalar covariates and for Gaussian multivariate covariates under symmetric noise. The distribution-dependent null center reduces to the dimension-free constant $4/π$ when both covariates and noise are Gaussian. We establish asymptotic size control and consistency against fixed mixture alternatives under regularity conditions. We extend the method to intercept and partial-mixture heterogeneity and study endogeneity, heteroskedasticity, and nonlinear misspecification. Simulations examine calibration, power, and robustness. Finally, a three-stage Detect--Screen--Verify toolkit separates evidence of heterogeneity from its substantive explanation. Across eight public datasets, it recovers four established SP benchmarks and identifies four cases that, to our knowledge, have not been documented previously. The detector requires neither latent-group labels nor a prespecified number of mixture components.

stat.ME↗

Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision pipelines is challenging because a requirement may admit multiple intervention paths with different downstream effects. We formulate requirement-driven adaptation as the joint selection of an intervention route and an admissible module-level change, and propose a graph-constrained agentic framework in which domain agents expose admissible reformulation interfaces and a central processor searches over bounded intervention paths. Candidates are validated and compared using downstream KPIs. In collaboration with a large retail partner, we evaluate 100 warehouse requirements elicited from practitioner interviews, with GPT, Qwen, and DeepSeek as base LLMs. Relative to direct LLM reformulation, our framework improves correctness and end-to-end success across all three models, raising end-to-end success from 72--76% to 79--83%.

cs.AI↗

Extreme-Case Distorted Utility under Moment Ambiguity

Many operations decisions under distributional ambiguity, from pricing and inventory to capacity and contracting, evaluate an action through a tail-sensitive distorted utility of an uncertain payoff and hedge against the least favorable distribution consistent with a few known moments; the resulting worst-case evaluation is the inner problem of a moment-based distributionally robust decision. We study this inner problem, the extreme-case distorted utility under moment constraints, for a locally Lipschitz utility that may be nonsmooth and neither convex nor concave together with a general, possibly atomic, distortion. Recasting the problem in the quantile domain, we develop a unified method that yields exact first-order optimality conditions and closed-form extremal values and distributions for both the worst and best cases, drawing on nonsmooth variational analysis. A central step treats the monotonicity constraint by isotonic projection onto the monotone cone, turning an abstract infinite-dimensional restriction into an inexpensive inner solve that scales linearly in the discretization. The method recovers and extends classical moment bounds through three examples: a range value-at-risk extension of the Scarf bound, GlueVaR distortions with a reward--penalty utility, and a capped incentive contract under conditional value-at-risk. As the inner oracle of a robust min-max decision, the characterization embeds directly in outer robust optimization, illustrated on a real capacity-provisioning problem for generative artificial intelligence inference where accounting for moment ambiguity lowers required capacity while preserving service compliance.

math.OC↗

Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow

Large-scale optimization is a key backbone of modern business decision-making. However, building these models is often labor-intensive and time-consuming. We address this by proposing LEAN-LLM-OPT, a LightwEight AgeNtic workflow construction framework for LLM-assisted large-scale OPTimization auto-formulation. LEAN-LLM-OPT takes as input a problem description together with associated datasets and orchestrates a team of LLM agents to produce an optimization formulation. Specifically, upon receiving a query, two upstream LLM agents dynamically construct a workflow that specifies, step-by-step, how optimization models for similar problems can be formulated. A downstream LLM agent then follows this workflow to generate the final output. The agentic workflow leverages common modeling practices to structure the modeling process into a sequence of sub-tasks, offloading mechanical data-handling operations to auxiliary tools. This reduces the LLM's burden in planning and data handling, allowing us to exploit its flexibility to address unstructured components. Extensive simulations show that LEAN-LLM-OPT, instantiated with GPT-4.1 and the open source gpt-oss-20B, achieves strong performance on large-scale optimization modeling tasks and is competitive with state-of-the-art approaches. In addition, in a Singapore Airlines choice-based revenue management use case, LEAN-LLM-OPT demonstrates practical value by achieving leading performance across a range of scenarios. Along the way, we introduce Large-Scale-OR and Air-NRM, the first comprehensive benchmarks for large-scale optimization auto-formulation. The code and data of this work is available at https://github.com/CoraLiang01/lean-llm-opt.

cs.AI↗

Enhancing binary classification: A new stacking method via leveraging computational geometry

Stacking, a potent ensemble learning method, leverages a meta-model to harness the strengths of multiple base models, thereby enhancing prediction accuracy. Traditional stacking techniques typically utilize established learning models, such as logistic regression, as the meta-model. This paper introduces a novel approach that integrates computational geometry techniques, specifically solving the maximum weighted rectangle problem, to develop a new meta-model for binary classification. Our method is evaluated on multiple open datasets, with statistical analysis showing its stability and demonstrating improvements in accuracy compared to current state-of-the-art stacking methods with out-of-fold predictions. This new stacking method also boasts two significant advantages: enhanced interpretability and the elimination of hyperparameter tuning for the meta-model, thus increasing its practicality. These merits make our method highly applicable not only in stacking ensemble learning but also in various real-world applications, such as hospital health evaluation scoring and bank credit scoring systems, offering a fresh evaluation perspective.

cs.LG↗

Limousine Service Management: Capacity Planning with Predictive Analytics and Optimization

The limousine service in luxury hotels is an integral component of the whole customer journey in the hospitality industry. One of the largest hotels in Singapore manages a fleet of both in-house and outsourced vehicles around the clock, serving 9000 trips per month on average. The need for vehicles may scale up rapidly, especially during special events and festive periods in the country. The excess demand is met by having additional outsourced vehicles on standby, incurring millions of dollars of additional expenses per year for the hotel. Determining the required number of limousines by hour of the day is a challenging service capacity planning problem. In this paper, a recent transformational journey to manage this problem in the hotel is introduced, driving up to S\$3.2 million of savings per year with improved service level. The approach builds on widely available open-source statistical and spreadsheet optimization tools, along with robotic process automation, to optimize the schedule of its fleet of limousines and drivers, and to support decision-making for planners/controllers to drive sustained business value.

stat.AP↗

On Policies for Single-leg Revenue Management with Limited Demand Information

In this paper we study the single-item revenue management problem, with no information given about the demand trajectory over time. When the item is sold through accepting/rejecting different fare classes, Ball and Queyranne (2009) have established the tight competitive ratio for this problem using booking limit policies, which raise the acceptance threshold as the remaining inventory dwindles. However, when the item is sold through dynamic pricing instead, there is the additional challenge that offering a low price may entice high-paying customers to substitute down. We show that despite this challenge, the same competitive ratio can still be achieved using a randomized dynamic pricing policy. Our policy incorporates the price-skimming technique from Eren and Maglaras (2010), but importantly we show how the randomized price distribution should be stochastically-increased as the remaining inventory dwindles. A key technical ingredient in our policy is a new "valuation tracking" subroutine, which tracks the possible values for the optimum, and follows the most "inventory-conservative" control which maintains the desired competitive ratio. Finally, we demonstrate the empirical effectiveness of our policy in simulations, where its average-case performance surpasses all naive modifications of the existing policies.

cs.DS↗