arXiv ScienceSearch

arXiv subjects

Zhaozhao Ma

Publications and source records attributed to Zhaozhao Ma.

3 recordsLinked to original sources

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in safety-critical clinical decision-making. Most existing explainability methods in healthcare operate in a post-hoc manner and are predominantly designed for unimodal data, which limits their applicability in increasingly prevalent multimodal diagnostic settings. This paper addresses the problem of self-explainable multimodal diagnosis by formulating it within the information bottleneck (IB) framework. We propose a unified learning paradigm that jointly optimizes predictive performance and modality-specific explainability by identifying the most informative elements inside each modality that contribute to diagnostic decisions. To enable tractable and stable optimization, we employ a matrix-based Renyi's $α$-order entropy functional under the assumption of sufficiently expressive encoders. Extensive experiments on representative medical datasets spanning heterogeneous modalities demonstrate that the proposed method consistently achieves strong diagnostic performance, including an absolute accuracy improvement of 9.1 percentage points on the iCTCF dataset. Moreover, the learned explanations provide transparent and modality-aware insights into feature relevance, thereby improving both the explainability and generalization.

cs.CV

MM-ARC: Multimodal Adaptive Routing of Capital with Robustness-Audited Strategy Pools

Financial trading systems must convert multimodal market history into executable positions while limiting overfitting from repeated strategy search. We introduce MM-ARC (MultiModal Adaptive Routing of Capital), which routes capital across trend, reversal, breakout, and exposure-control experts using aligned chart, numerical, and technical-text views. Within each market, regime-conditioned strategy pools are shared with bounded asset-specific adjustments. Robustness-Audited Bayesian Optimization (RABO) filters candidates proposed by Bayesian optimization on purged validation blocks using after-cost benchmark exceedance, lower-tail performance, stability, and turnover; a common portfolio layer then produces market-feasible orders. We evaluate 62 instruments across five asset classes using five training seeds and a frozen July 2025--June 2026 trading holdout. Under an all-in one-way cost of 10 basis points per unit of executed turnover, MM-ARC attains an equal-market Sharpe ratio of 1.33 and maximum drawdown of -13.7, versus 0.53 and -18.3 for the LLMoE-style routing baseline. The global learned-static control reaches 1.12 and -15.3, respectively. Paired block-bootstrap intervals favor the prespecified contrasts, while ablation point estimates are consistent with contributions from visual inputs, adaptive routing, exposure control, and robustness-audited admission. Family-level data-snooping tests also reject their prespecified nulls (SPA p= .039; Reality Check p= .021); we therefore interpret the evidence as benchmark-relative support within the evaluated candidate family and holdout, not as universal or future-regime superiority.

q-fin.TR

Explainable Multimodal Regression via Information Decomposition

Multimodal regression aims to predict a continuous target from heterogeneous input sources and typically relies on fusion strategies such as early or late fusion. However, existing methods lack principled tools to disentangle and quantify the individual contributions of each modality and their interactions, limiting the interpretability of multimodal fusion. We propose a novel multimodal regression framework grounded in Partial Information Decomposition (PID), which decomposes modality-specific representations into unique, redundant, and synergistic components. The basic PID framework is inherently underdetermined. To resolve this, we introduce inductive bias by enforcing Gaussianity in the joint distribution of latent representations and the transformed response variable (after inverse normal transformation), thereby enabling analytical computation of the PID terms. Additionally, we derive a closed-form conditional independence regularizer to promote the isolation of unique information within each modality. Experiments on six real-world datasets, including a case study on large-scale brain age prediction from multimodal neuroimaging data, demonstrate that our framework outperforms state-of-the-art methods in both predictive accuracy and interpretability, while also enabling informed modality selection for efficient inference. Implementation is available at https://github.com/zhaozhaoma/PIDReg.

cs.LG