arXiv · 2609.14299
Learning Source Acquisition Policies by Offline Planning
Abstract
Predicting under an acquisition budget requires choosing feature groups whose value can depend on later queries. O-MPAC transfers finite-horizon risk-cost targets from complete training records into a shared source-action scorer. At inference time, the scorer uses partial observations and source metadata, re-scores after each query, and applies a hard cost mask. We analyze how tied teacher targets and the remaining planning horizon affect the learned decisions. Uniform supervision over tied minima preserves the target distribution under source relabeling. In a five-seed routing experiment, it achieves 0.965 accuracy under both original and context-last orders. On six real tasks, validation selects H1 without action cross-entropy in all thirty splits. O-MPAC has the highest mean budget-integrated accuracy on five tasks against source-adapted GDFS, DIME, AACO+NN and a static policy.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ziqi Zhao, Run Xu, Qingjian Ni. 2026-09-13. Learning Source Acquisition Policies by Offline Planning. https://arxiv.org/abs/2609.14299
Cite the original work for its findings. Save a collection to share your selection of sources.