arXiv ScienceSearch

arXiv subjects

Zihao Wu

Publications and source records attributed to Zihao Wu.

At least 19 recordsLinked to original sources

JADES: Dynamical Measurements of Dark Matter Halo Masses at $z \approx 7$ Using Galaxy Pairs

We present dynamical measurements of dark matter halo masses at $z\approx7$ using galaxy pairs. Galaxy pairs in the early Universe are usually in their first infall, before substantial orbital evolution or phase mixing. Their relative velocities closely trace the halo mass, especially when the pair separation is near the halo virial radius. In the TNG100 simulation, we find that dynamical mass estimators can recover halo masses with an intrinsic scatter of $\sim$0.2 dex. We apply this method to 10 pair systems at z=6-8 with [O III] spectroscopy from the JWST Advanced Deep Extragalactic Survey (JADES). We infer the stellar-to-halo mass relation while marginalizing over projection effects, measurement uncertainties, and intrinsic scatter. We find a mean halo mass of $\log\,(M_{200}/M_\odot)=11.06\pm0.21$ at $\log\,(M_{*}/M_\odot)=9$. These galaxies do not appear to inhabit unusually massive halos. Our results instead indicate a high stellar-to-halo mass ratio, corresponding to an integrated star formation efficiency of $5^{+3}_{-2}\%$, about twice the TNG100 prediction, although the current statistical significance is limited. Future observations with larger samples will tighten these constraints and directly examine whether enhanced star formation efficiency drives the overabundance of luminous galaxies at cosmic dawn.

astro-ph.GA

The Roman eXtreme Deep Field (RXDF)

The Roman eXtreme Deep Field (RXDF) program is one of the five General Astrophysics Survey (GAS) programs approved for observing time with the Nancy Grace Roman Space Telescope in Cycles 1 and 2. It has been allocated 386.41 hours to carry out an imaging survey to AB = 30 mag (5-sigma) over ~140x larger area than the Hubble eXtreme Deep Field (HXDF) full-depth area (ACS+WFC3/IR). The RXDF will cover the full Roman wavelength range with 7 bands, reaching AB = 30 mag in RZYJH, 29 mag in F, and 28 mag in K, over a full-depth area of 678.75 arcmin^2 embedded in a total area of 1,243 arcmin^2, and far exceeding the depths of the Roman Core Community Surveys (CCS). The RXDF is within the Euclid Ultra Deep Field (EUDF) near the North Ecliptic Pole (NEP), a strategic long-term field for generational space facilities, with a wealth of multi-wavelength data including extensive coverage from the James Webb Space Telescope (JWST) NEXUS Treasury program. The observations will cover 3 epochs at a 1-year cadence, each epoch divided into 3 sub-epochs ~10 days apart, enabling time-domain studies on time baselines from ~10 days to over ~2 years. The RXDF is uniquely positioned to address critical questions in reionization, large scale structure (LSS), growth of supermassive black holes (SMBHs), little red dots (LRDs), and high-z supernovae (SNe); the volumes probed by HST+JWST are too small at these extreme depths, and even the deepest CCS tiers are too shallow. In addition to our key objectives, a wealth of additional science will be enabled by engaging the community with our rapidly released datasets, revolutionizing a wide range of science for a lasting legacy. This short document, which is converted from the approved RXDF proposal, aims to provide the community with a summary of the program.

astro-ph.GA

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.

cs.LG

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often endure lesion noises, modality misalignment, hallucination, and missed contextual grounding in complex clinical cases. Moreover, prevailing agent systems usually depend on static and non-adaptable pipelines and lack the versatility necessary for complex medical reasoning. To resolve these difficulties, we present BioMed-Agent-RL, a unified medical agent that incorporates adaptive orchestration, policy, and reward-based reinforcement learning (RL) models for biomedical applications. To ensure reliability, it invokes clinical context-aware preference optimization (CPO), direct preference optimization (DPO), and group relative policy optimization (GRPO) with dynamic entropy regulation. This pipeline utilizes a multimodal meta-learning approach that operates as a field-specific expert and human judgment synthesizer. The agent adaptively utilizes a set of model-level expertise, such as clinical grounding and reasoner, lesion segmenter, and field-specific synthesizer, across various clinical modalities (e.g., X-ray) by utilizing an iterative and adaptive RL approach. The agent learns to seriously synthesize misleading, conflicting vision cues and trust in inherent reasoning, while specialist advice is faulty. An intensive ablation study is conducted across multiple benchmarks, and the agent significantly outperforms existing state of the art models, such as GPT-5, attaining up to ~73% accuracy (gain of ~5%) over contemporary baselines. As a result, the framework suggests a new standard for building factual, reliable, robust, and expert-like intelligent agent systems for independent clinical reasoning.

cs.LG

Low Ly$α$ Visibility in Galaxy Overdensities: Reionization Topology and Neutral-Fraction Ceilings from DIVER over $4.8<z<11$

Ly-alpha emission is widely used to trace cosmic reionization, but its interpretation depends on how Ly-alpha visibility varies with galaxy environment. We use deep JWST/NIRSpec observations from Deep Insights into UV Spectroscopy at the Epoch of Reionization (DIVER) in GOODS-N to measure Ly-alpha visibility for 250 galaxies at 4.8 25 A. We combine these measurements with H-alpha and [O III] emitters from JWST/NIRCam wide-field slitless spectroscopy to map the density field around each DIVER galaxy. Galaxies with high Ly-alpha equivalent widths (W_Lyalpha>25 A) or high effective Ly-alpha escape fractions (f_esc,Lyalpha^eff>0.05) tend to lie farther from nearby H-alpha and [O III] emitters than galaxies with lower Ly-alpha visibility. The clearest signal occurs near the prominent GOODS-N overdensity at z~5.2, where fewer than 15% of galaxies show strong Ly-alpha emission. This trend is opposite to the simplest inside-out reionization expectation that overdensities produce larger ionized regions and enhance Ly-alpha visibility. Possible explanations include circumgalactic and local intergalactic opacity, dense absorbers, and gas kinematics. We also derive an empirical upper envelope for f_esc,Lyalpha^eff and calibrate it with reionization simulations. Interpreting this envelope as a limiting IGM-attenuation signal gives neutral-fraction ceilings of _max=0.36, 0.76, 0.74, 0.84, and 1.0 at z~5.2, 5.8, 6.7, 7.7, and 9.8, respectively. The z~8 ceiling disfavors an almost completely neutral IGM at this epoch. These results support patchy reionization already underway by z~8 and show that galaxy Ly-alpha visibility encodes both large-scale ionization topology and near-source gas structure.

astro-ph.GA

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

cs.CV

Revolutionizing Finance with LLMs: An Overview of Applications and Insights

In recent years, Large Language Models (LLMs) like ChatGPT have seen considerable advancements and have been applied in diverse fields. Built on the Transformer architecture, these models are trained on extensive datasets, enabling them to understand and generate human language effectively. In the financial domain, the deployment of LLMs is gaining momentum. These models are being utilized for automating financial report generation, forecasting market trends, analyzing investor sentiment, and offering personalized financial advice. Leveraging their natural language processing capabilities, LLMs can distill key insights from vast financial data, aiding institutions in making informed investment choices and enhancing both operational efficiency and customer satisfaction. In this study, we provide a comprehensive overview of the emerging integration of LLMs into various financial tasks. Additionally, we conducted holistic tests on multiple financial tasks through the combination of natural language instructions. Our findings show that GPT-4 effectively follow prompt instructions across various financial tasks. This survey and evaluation of LLMs in the financial domain aim to deepen the understanding of LLMs' current role in finance for both financial practitioners and LLM researchers, identify new research and application prospects, and highlight how these technologies can be leveraged to solve practical challenges in the finance industry.

cs.CL

PhotoIFU: NIRCam as a Photometric Integral Field Unit for Mapping Feedback in Galaxies

We present PhotoIFU, a workflow that uses deep multi-band imaging as a low-resolution photometric integral field unit. Applied to PSF-matched JWST/NIRCam imaging, PhotoIFU treats each spatial pixel as a coarse SED element and fits the pixel SEDs with Prospector to map resolved stellar-population and ISM-related properties. We apply this approach to three galaxies at $z=1.3$--3.7 in JADES: two systems with extended ionized line emission and one post-starburst galaxy with an exceptionally strong neutral outflow. Pixel-by-pixel SED fitting gives maps of stellar-mass surface density, specific star formation rate, dust attenuation, gas-phase metallicity, and recent star-formation history. We find that regions selected from the extended-emission or outflow geometry occupy distinct parts of the resolved SED-property distribution compared with the full host. In the systems with extended ionized emission, these regions are generally less dusty, consistent with ionized emission being observed along dust-poor, low-column-density pathways through the host. In the neutral-outflow system, the selected regions show enhanced recent star formation, suggesting that compact rejuvenation may mark the aftermath of an earlier energetic phase. These results show that galactic outflows and extended emission-line structures can be spatially associated with measurable differences in resolved host-galaxy stellar populations and ISM-related properties. PhotoIFU provides an imaging-based method for resolved SED mapping of feedback-related structures in larger galaxy samples where full spectroscopic integral-field mapping is unavailable.

astro-ph.GA

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.

cs.CL

Beyond the Dot: an LRD-like nucleus at the Heart of an IR-Bright Galaxy and its implications for high-redshift LRDs

Little Red Dots (LRDs) are compact, red sources discovered by JWST at high redshift ($z \gtrsim 4$), marked by distinctive 'V-shaped' spectral energy distributions (SEDs) and often interpreted as rapidly accreting Active Galactic Nuclei (AGNs). Their true nature remains unclear, however, and their evolutionary connection to their lower-redshift counterparts is still poorly constrained. Thus, we present WISEA J123635.56+621424.2, here dubbed {\it the Saguaro}, a $z=2.0145$ galaxy in GOODS-North, as a possible analog of high-redshift LRDs and a potential missing link in their evolutionary path toward lower-redshift systems. It features a compact LRD-like nucleus surrounded by a face-on spiral host. Its connection to LRDs includes that: (1) its nuclear spectrum shows a clear `V-shaped'' SED; and (2) when redshifted to $z=7$, surface-brightness dimming makes the host undetectable, thus mimicking an LRD. This suggests that high-redshift LRDs may be embedded in extended hosts. To test this, we stack rest-frame UV images of 99 photometrically selected LRDs, revealing faint, diffuse emission. Stacking in redshift bins reveals mild radial growth, consistent with the expected galaxy size evolution. A simple analytic model confirms that surface-brightness dimming alone can explain their compact appearance. Lastly, we show that {\it the Saguaro} is not unique by describing similar objects from the literature at $z\lesssim3.5$. Taken together, our results support a scenario in which LRDs may not be a distinct population, but could instead be the visible nuclei of galaxies undergoing a short-lived, perhaps AGN-dominated, evolutionary phase, with their compact, red appearance driven largely by observational biases.

astro-ph.GA

JADES: the mass-metallicity relation at $z=1-10$. New calibrations, extremely metal-poor galaxies, and chemical diversity

We present gas-phase metallicities of star-forming galaxies at $z=1$-10 with deep JWST/NIRSpec spectra from the JADES full data release, Dark Horse, and OASIS programmes. We stack $\sim$1500 medium-resolution spectra, yielding detections of the [OIII]$λ$4363 auroral line down to $12+\log(\mathrm{O/H})=7.0$ to derive stack-based strong-line calibrations over the metallicity range $12+\log(\mathrm{O/H})=7.0$-8.7. At a fixed metallicity, our stacks exhibit [OIII]$λ$5007/H$β$ and [OIII]$λ$5007/[OII]$λλ$3726,3729 values generally lower than calibrations based on high-$z$ individual auroral-line emitters, suggesting an observational bias towards higher excitation introduced when requiring auroral line detections in individual spectra. Based on our new calibrations, we obtain canonical mass-metallicity relations (MZRs) at z$=$1-10, identifying a decrease in metallicities from $z\sim0$ to z$\sim$4-10, without significant change in slope. Moreover, we identify 50 promising candidates of extremely metal-poor galaxies (EMPGs) with $12+\log(\mathrm{O/H})=6.7$-7.3 (1-4\% solar metallicity) at $z=1.2$-9.1. The MZRs of EMPGs are characterised by a large scatter, with those having lower metallicities generally exhibiting lower sSFRs, opposite of what expected from the local Fundamental Metallicity Relation. These results support a stochastic star-formation history involving gas consumption/ejection and metal-poor inflow, strongly affecting metallicities of low-mass galaxies. Furthermore, we identify two Little Red Dots in our EMPG candidates, both exhibiting broad H$α$ and prominent Ly$α$, offering insights into the early black-hole growth in extremely metal-poor environments.

astro-ph.GA

JADES: A Prominent Galaxy Overdensity Candidate within the First 500 Myr

We report a galaxy overdensity candidate at $z\approx10.5$ in the JWST Advanced Deep Extragalactic Survey (JADES). This overdensity contains 18 galaxies with consistent photometric redshifts within 8 comoving Mpc in projection. The galaxy number density is four times higher than the field expectation, accounting for one-third of comparably bright galaxies and nearly 50% of the total star formation rate at $10<z_\mathrm{phot}<12$ in the GOODS-S field. Galaxies in the overdensity more frequently have close companions or substructure, with one-third showing such features within 1 kpc at consistent photometric redshifts, implying enhanced interactions. Most galaxies have stellar masses of 0.6-3$\times10^8 M_\odot$, half-light radii of $\sim200$ pc, and star formation rates (SFRs) of $\sim5 M_\odot \mathrm{yr^{-1}}$. Their stellar masses and SFRs are slightly higher than those of field galaxies, but remain broadly consistent with typical high-redshift scaling relations. Two compact objects show possible Balmer breaks, suggestive of evolved stellar populations or little red dots (LRDs). We find tentative evidence for a spatially varying Ly$α$ transmission inferred photometrically, consistent with an emerging ionized bubble. This overdensity provides a rare opportunity for probing the environmental impact on galaxy evolution and the onset of cosmic reionization within the first 500 Myr.

astro-ph.GA

On the spanning cuts consistency problem in the IBP reductions of Feynman integrals

The spanning cuts method is a powerful approach to reduce the cost of IBP reduction while computing Feynman integrals. However, its usage is limited due to the so-called consistency problem. It was unclear why the IBP reduction coefficients can be inconsistent with each other between different cuts. In this paper, we report a mechanism behind this inconsistency. We found that the IBP relations can be violated under the cuts, if we blindly erase the hidden terms that are proportional to the ``vanishing'' Feynman prescription parameters in the relations. In some cases, the cut introduces pinch singularities, which cancel the vanishing Feynman prescription parameters, making the hidden terms finite. In various cases, the error comes from omitting such finite hidden terms. We also claimed that the pinch singularity under the cuts are related to some hidden relations between the propagators. In this paper, we provide an algorithm and its implementation to find the linear hidden relations.

hep-ph

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artificial general intelligence, enabling agents to predict, plan, and reason within learned representations. Despite rapid progress across reinforcement learning, robotics, autonomous driving, and video generation, the field lacks a unified framework integrating its diverse architectural choices, training methods, reasoning mechanisms, and application settings. This survey addresses that gap with a multi-axis taxonomy organized along four dimensions: (i) architecture, encompassing representation format, dynamics formulation, input modality, learning paradigm, and downstream application; (ii) methodological family, including state-space and recurrent approaches, transformer-based models, diffusion-based generators, physics-informed networks, and language-augmented multimodal systems; (iii) reasoning strategy, covering imagination-based planning, latent policy learning, counterfactual reasoning, and planning under uncertainty; and (iv) application domain, spanning robotics, autonomous driving, video prediction, multimodal agents, reinforcement learning, scientific modeling, medical imaging, educational measurement, and business and finance. Tracing the field from early cognitive-science foundations to milestone systems such as PlaNet, the Dreamer family, MuZero, Sora, Cosmos, and Genie, we examine how these dimensions interact and highlight the recent convergence of chain-of-thought reasoning with world-model imagination. We review evaluation protocols and benchmarks, identify persistent challenges such as compounding prediction errors, sim-to-real transfer, and fragmented evaluation, and outline future directions toward unified multimodal world models, foundation-scale interactive simulators, and safe deployment in safety-critical domains.

cs.LG

JWST Advanced Deep Extragalactic Survey (JADES) Data Release 5: stellar population catalogue for galaxies in GOODS-N and GOODS-S

We present the galaxy stellar population catalogue from the JWST Advanced Deep Extragalactic Survey (JADES) Data Release 5 (DR5), providing homogeneous Bayesian inference of physical galaxy properties in GOODS-N and GOODS-S. Using deep JWST/NIRCam and MIRI imaging combined with ancillary multi-wavelength data, we model the spectral energy distributions of ~500,000 sources with the Prospector framework. Our modelling incorporates flexible non-parametric star-formation histories (SFHs), nebular emission, dust attenuation, metallicities, and mid-infrared AGN and dust emission. We adopt an evolving star-forming main sequence (SFMS) prior for modelling the SFHs, which provides a physically-motivated long-term shape of SFHs while retaining non-parametric flexibility. The prior links stellar mass growth and SFR through the observed redshift-dependent SFMS, shaping the global behaviour of the inferred SFHs but allowing substantial deviations and scatters wherever supported by the data. We derive posterior distributions for stellar masses, SFRs, SFHs, dust attenuation, metallicities, and AGN contributions. The depth and wavelength coverage of JADES enable robust stellar mass measurements down to low-mass limits, as well as improved constraints on recent star-formation activity for ~350,000 galaxies at z = 1 - 9. The adoption of a physically motivated prior mitigates unphysical solutions and reduces degeneracies between redshift, age, dust, and metallicity, particularly for faint sources. We validate the catalogue through consistency checks and comparison to spectroscopic redshifts where available. The resulting value-added catalogue provides a uniform set of stellar population parameters suitable for statistical studies of galaxy growth, quenching, and the build-up of stellar mass across cosmic time. The full catalogue and posterior summaries are publicly released as part of JADES DR5.

astro-ph.GA

JWST Advanced Deep Extragalactic Survey (JADES) Data Release 5: Photometrically Selected Galaxy Candidates at z > 8

We present a sample of 2081 sources selected at photometric redshift $z_{\mathrm{phot}} > 8$ across the JADES DR5 data release in GOODS-S and GOODS-N over a total area of 469 square arcmin. These sources range from $M_{\mathrm{UV}} = -22$ to $M_{\mathrm{UV}} = -16$, with 19 objects at $z_{\mathrm{phot}} > 14$. We estimate the UV slopes for the full sample from fits to the photometry and find evidence for a steepening of the relationship between the UV continuum slope and $M_{\mathrm{UV}}$ to higher redshifts, a result that differs from prior analyses of brighter samples in the literature. We provide evidence that over one quarter of our sources have evidence for being morphologically extended, with many galaxies showing multiple bright knots or clumps even out to $z \sim 13 - 14$, an indication of how galaxies at Cosmic Dawn are growing and evolving. We discuss JADES-GN+189.15982+62.28899, a GOODS-N F200W dropout galaxy at $z_{\mathrm{phot}} \sim 15 - 18$ which has been observed spectroscopically with JWST/NIRSpec in prism mode, resulting in a very low signal-to-noise spectrum that is consistent with the photometry and rules out a number of low-redshift solutions for the source. Finally, we use a subsample of 123 objects in our sample with spectroscopic redshifts to explore the usage of alternate fitting templates and a prescription for Ly-$α$ damping wing absorption, finding that both produce significant improvements to the estimated photometric redshifts.

astro-ph.GA

A World Model of Radiologist Reading for Medical Image Representation Learning

Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.

cs.CV

Same Brain, Different Prediction: How Preprocessing Choices Undermine EEG Decoding Reliability

Electroencephalography (EEG) is a cornerstone of brain-computer interfaces and clinical neuroscience, yet deep learning models are typically trained and evaluated under a single, unreported preprocessing pipeline. We formalize preprocessing choices as a counterfactual intervention space and show that EEG predictions are surprisingly unstable under this space: across six datasets spanning four paradigms, up to 42% of trial-level predictions flip when only the preprocessing changes, a variability that standard uncertainty methods do not explicitly quantify because they condition on a fixed preprocessing pipeline. We provide three tools to make this instability measurable, decomposable, and reducible. First, a Walsh-Hadamard decomposition of the 2^7 pipeline space reveals that sensitivity is near-additive in practice under the binary intervention design, enabling efficient step-by-step optimization. Second, we introduce Preprocessing Uncertainty (PU), a per-trial diagnostic that captures a dimension of instability complementary to model-based confidence. Third, we study Normalized Adaptive PGI (NA-PGI), a graph-structured regularizer that exploits the compositional structure of preprocessing interventions as one mitigation strategy with clear scope conditions.

cs.LG