arXiv Science⌕ Search

arXiv · 2609.34887

Conformal Prediction and Conditional Coverage for Tabular Foundation Models

Abstract

Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sungwoo Park, Sunghee Park, Won Chang. 2026-09-28. Conformal Prediction and Conditional Coverage for Tabular Foundation Models. https://arxiv.org/abs/2609.34887

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Certified Adaptive Refresh: Anytime-Valid Monitoring for Federated Conformal RAG

Question-answering services built on retrieval-augmented generation (RAG), in which a language model answers from retrieved documents, are inspected continuously and upgraded repeatedly, so their reliability guarantee must survive both. We study federated conformal RAG: nodes holding private corpora score candidate answers with a shared language model and send compressed scores to a hub that returns an answer set; a miss omits the true answer. We formulate monitoring as a sequential test: alarm when misses exceed a certified bound on the miss rate (the allowance). But a conformal certificate covers one frozen configuration at one look fixed in advance, so it gives no anytime-valid alarm; a threshold calibrated for one model certifies nothing about its replacement; and re-certifying each upgrade at full error level compounds failures. We propose Anytime-FC-RAG with three auditable rules: pick each configuration before its fresh calibration, charge every certificate attempt to one trajectory-wide budget $δ_{\rm cal}$, and fix threshold, allowance and bet before each query. One betting wealth $E_t$, growing in expectation only when misses exceed the allowance, never resets across deployments. Our certified adaptive-refresh theorem shows that, with evidence budget $δ_e$, alarming when $E_t$ reaches $1/δ_e$ has false-alarm probability at most $δ_e + δ_{\rm cal}$ under a random number of history-selected, evidence-driven refreshes of model, retriever, corpus, score or threshold; $δ_{\rm cal}$ pays for wrong certificates. With an exact order-statistic certificate the bound is distribution-free, architecture-agnostic and about misses, not per-input coverage or distribution shift as such. On Qwen2.5/MMLU-Pro the alarm separated audited-harmful from benign shifts completely at a 5% budget; in a model swap, only fresh calibration kept a downgrade's miss rate in check.

stat.ML↗

Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction

Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection with its owner, or node, which scores candidate answers from its own documents; a central hub combines the scores. Some nodes, called Byzantine, may be compromised, faulty, or misled by instructions hidden in documents, and report arbitrary scores. Conformal prediction returns a set containing the correct answer with a chosen probability, using a cutoff set in a calibration step on questions with known answers. An unknown group of nodes, no larger than a declared bound, may misreport both in this step and at query time. Existing methods assume every node is honest or protect only the calibration step. We observe that the honest nodes are the same in both steps. The hub therefore has all nodes score the same calibration questions, and keeps a candidate only if some plausible group of honest nodes, using its own scores in both steps, would keep it. We prove that the resulting sets contain the correct answer with the chosen probability in finite samples, whatever the Byzantine nodes report. No method using the same information can return smaller sets without risking the loss of an answer the honest nodes support. If nodes fail at random, the guarantee weakens only by the probability that more nodes fail than declared. In simulations, on real question-answering tasks including medical exams, and with language models as nodes, some hijacked, our sets reached the target whenever no more nodes misbehaved than declared, while plain averaging could miss it. They were also clearly smaller than those of simpler methods with the same protection, most of all when the declared bound was generous, so a cautious bound costs little.

stat.ML↗

Hierarchical Utility Calibration for Structured Multiclass Decisions

In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.

stat.ML↗