arXiv ScienceSearch

arXiv subjects

Jia Huang

Publications and source records attributed to Jia Huang.

At least 19 recordsLinked to original sources

Controllable Dysarthric Speech Synthesis with Patient-Specific Conditioning for Speaker-Diverse ASR Augmentation

Dysarthric speech recognition is limited by high speaker variability and scarce labeled data. Existing synthesis methods often couple speaker identity with dysarthric articulation, reducing control over generated speech. We propose a controllable dysarthric speech synthesis framework for ASR augmentation with separate prompt-derived timbre prefixes and learnable patient-specific pathology prefixes. Built on a pre-trained neural codec language model adapted using LoRA, the framework combines both prefixes through additive conditioning. A dual-classifier objective with gradient reversal and voice-conversion-based counterfactual augmentation promotes factor separation, allowing learned pathology conditions to be paired with different target speakers, including healthy speakers. Experiments on TORGO show that generated speech can partially replace real dysarthric training data and provides effective speaker-diverse augmentation when combined with real data. Objective ASR, perceptual, factor-separation, and phoneme-level analyses indicate preservation of target-speaker timbre and pathology-dependent patterns consistent with real dysarthric speech.

cs.SD

Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets

Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.

cs.LG

What Makes a Redundant Representation Remember? Lineage Isolation, Not Masking

Memory-based evolutionary algorithms for dynamic optimization often carry a redundant second copy of the genotype and expose only one copy to the objective, on the assumption that the shielded copy accumulates information about past optima. We show this assumption is false as usually implemented, and identify the structural property that actually determines whether the shielded copy retains information. We formalize such methods as a gated dual-copy representation with two independent design axes: a gating rule deciding which copy is evaluated, and an inheritance rule deciding whether the two copies mix across generations. A ablation shows retained information is governed almost entirely by the inheritance rule (21.4 vs. 1.3 bits) and is nearly invariant to the gating rule. Per-locus independent inheritance reshuffles cross-locus structure every generation, so shielding preserves the variance of the hidden copy while destroying the pattern that constitutes a memory. Under isolated inheritance the memory effect is real: against a single-copy baseline matched for representation budget, the method gains +0.010 AUC when optima recur periodically and loses 0.078 when they drift unidirectionally---a 0.089 separation under otherwise identical settings, which excludes explanations based on added capacity. We show the readout rate is also the corruption rate, predicting and confirming an interior optimum replicated across two implementations. We report one negative result with a mechanism: dual-copy representations lower the mutational error threshold, because gated expression is a selector rather than a joint decoder and therefore provides no coding gain. Finally, we document a benchmarking hazard: on dynamic benchmarks the choice of recombination operator alone shifted our baseline by 0.062 AUC, six times the effect size under study.

cs.NE

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $ϕ$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $ρ> 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $ρ\in [0.32, 0.52]$).

cs.AI

Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration

Restoring 3D scenes with low-light conditions is challenging, and most existing methods depend on precomputed camera poses and scene-specific optimization, which greatly restricts their application to real-world scenarios. To overcome these limitations, we propose Lumos3D, a pose-free single-forward framework for 3D low-light scene restoration. First, we develop a cross-illumination distillation scheme, where a frozen teacher network takes normal-light ground truth images as input to distill accurate geometric information to the student model. Second, we define a Lumos loss to improve the restoration quality of the reconstructed 3D Gaussian space. Trained on a single dataset, Lumos3D performs inference in a purely feed-forward manner, directly restoring illumination and structure from unposed, low-light multi-view images without any per-scene training or optimization. Experiments on real-world datasets demonstrate that Lumos3D achieves competitive restoration results compared to scene-specific methods. Our codes are available at https://github.com/HanzhouLiu/Lumos3D.

cs.CV

Wave Activity at MHD-ion Scales Associated with Switchbacks

Magnetic switchbacks (SB) -- the localized magnetic structures with magnetic field direction inclined at an angle $θ$ relative to the background $B_0$ -- in the young solar wind have been associated with enhanced ion-scale wave activity and local plasma heating. It remains debated whether the apparent wave-power increase is intrinsic or mainly caused by sampling geometry. In this work, we analyze magnetic and electric field fluctuations measured by Parker Solar Probe, focusing on the 0.1--3~\(f_{cp}\) frequency band that spans the transition from the MHD inertial range to ion-kinetic scales. By decomposing magnetic fluctuations into field-aligned and transverse components and comparing SB and non-SB intervals at the same local magnetic field angle, we test whether SBs sample an anisotropic cascade from different viewing angles or host intrinsically amplified wave activity. We find that the transverse magnetic power $δB_{\perp}$ is systematically enhanced inside switchbacks across a wide range of magnetic field rotation angles $θ$. The enhancement persists even at small and intermediate deflections, where geometric projection alone predicts weak power, indicating an intrinsic origin beyond sampling geometry. The inertial-range spectral indices also remain similar between SB and non-SB intervals despite the enhanced wave power inside SBs, suggesting that the underlying turbulence cascade is largely preserved. This excess $δB_{\perp}$ coincides with elevated proton temperatures and enhanced electric-field fluctuations, supporting the interpretation that SBs act as localized sites of cross-scale energy transfer and ion-scale dissipation in the near-Sun solar wind.

physics.space-ph

GO: The Great Outdoors Multimodal Dataset

The Great Outdoors (GO) dataset is a multi-modal annotated data resource aimed at advancing ground robotics research in unstructured environments. Existing off-road datasets often lack sensor diversity and exclude vital modalities like thermal and radar that are critical for operation in degraded conditions (e.g., low visibility or adverse weather). To address these gaps, we introduce a large-scale multimodal off-road dataset with six complementary sensor modalities, along with semantic annotations and GPS traces, to support tasks such as semantic segmentation, object detection, and SLAM. The diverse environmental conditions represented in the dataset present significant real-world challenges, which provide opportunities to develop more robust solutions to support the continued advancement of field robotics, autonomous exploration, and perception systems in natural environments. The dataset can be downloaded at: https://www.unmannedlab.org/the-great-outdoors-dataset/

cs.RO

Critical groups for Hopf algebra modules

This paper considers an invariant of modules over a finite-dimensional Hopf algebra, called the critical group. This generalizes the critical groups of complex finite group representations studied by Benkart, Klivans, Reiner and Gaetz. A formula is given for the cardinality of the critical group generally, and the critical group for the regular representation is described completely. A key role in the formulas is played by the greatest common divisor of the dimensions of the indecomposable projective representations.

math.CO

A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus on execution topology -- how data flows -- while cognitive science surveys focus on cognitive function -- what the agent does. Neither axis alone disambiguates architecturally distinct systems: the same Orchestrator-Workers topology can implement Plan-and-Execute, Hierarchical Delegation, or Adversarial Verification -- three patterns with fundamentally different failure modes and design trade-offs. We propose a two-dimensional classification that combines (1) a Cognitive Function axis with seven categories (Perception, Memory, Reasoning, Action, Reflection, Collaboration, Governance) and (2) an Execution Topology axis with six structural archetypes (Chain, Route, Parallel, Orchestrate, Loop, Hierarchy). The resulting 7x6 matrix identifies 28 named patterns, 15 with original names. We demonstrate orthogonality through systematic cross-axis analysis, define eight representative patterns in detail, and validate descriptive coverage across four real-world domains (financial lending, legal due diligence, network operations, healthcare triage). Cross-domain analysis yields five empirical laws of pattern selection governing the relationship between environmental constraints (time pressure, action authority, failure cost asymmetry, volume) and architectural choices. The framework provides a principled, framework-neutral, and model-agnostic vocabulary for AI agent architecture design.

cs.AI

Moments for generalizations of a coin flip game

We derive a recursive formula for the moments of the number of flips using a possibly biased coin to produce a prescribed finite binary string $S$ when $S$ is either a run of heads or a run of heads followed by a tails. Our recursive formula involve certain sums, which we simplify by using a one-parameter extension of the well-studied Eulerian number, which belongs to the two-parameter family of numbers introduced by Graham, Knuth, and Patashnik. We also use the Goulden--Jackson cluster method and Faà di Bruno's formula to establish a closed formula for the moments in a more general situation where a die having an arbitrary number of faces with possibly different probabilities is rolled repeatedly until a prescribed finite word occurs.

math.CO

Timing Jitter Induced by Stochastic Baseline Fluctuations in High-Count-Rate Superconducting Nanowire Single-Photon Detectors

Superconducting nanowire single-photon detectors (SNSPDs) have demonstrated timing jitter in the few-picosecond regime, yet their timing resolution deteriorates substantially under high-count-rate operation. Existing interpretations mainly attribute this degradation to deterministic waveform distortions, such as multiphoton responses and pulse pile-up, yet the experimentally observed jitter broadening at high count rates cannot be fully accounted for within this picture. Here, we show that stochastic baseline fluctuations arising from finite-memory readout dynamics constitute an intrinsic source of the count-rate-dependent timing jitter in SNSPD systems. For stochastically arriving photons, overlapping recovery responses accumulate in the readout chain and generate statistically fluctuating baselines, which are converted into timing uncertainty through threshold-based timing extraction. We develop a stochastic-process framework that quantitatively connects photon statistics, readout dynamics, and timing jitter. The framework predicts characteristic scaling behaviors, including a nonmonotonic dependence of baseline fluctuations under pulsed excitation with a maximum near half of the repetition frequency. These predictions are quantitatively verified through systematic variations of count rate, circuit time constant, and detector dynamical properties. Our results identify stochastic baseline dynamics as a fundamental mechanism limiting timing resolution in high-count-rate SNSPD operation and provide a general framework for optimizing finite-memory high-speed photon-counting systems.

physics.app-ph

Fortress: A Case Study in Stabilizing Search Recommendations via Temporal Data Augmentation and Feature Pruning

In search and recommendation systems, predictive models often suffer from temporal instability when certain input features introduce volatility in output scores. This instability can degrade model reliability and user experience especially in multi-stage systems where consistent predictions are critical for downstream decision making. We introduce Fortress, a general framework for enhancing model stability and accuracy by identifying and pruning features that contribute to inconsistent prediction scores over time. Fortress leverages historical snapshots temporally partitioned datasets capturing score fluctuations for the same entity across periods and follows a four-step process: (1) collect historical snapshots, (2) identify samples with unstable predictions, (3) isolate and remove instability-inducing features, and (4) retrain models using only stable features. While semantic features from LLMs and BERT-based models improve generalization, they often lack full query or entity coverage. Engagement-based features offer strong predictive power but tend to introduce temporal instability. Fortress mitigates this trade-off by suppressing the volatility of engagement signals while retaining their predictive value leading to more stable and accurate models. We validate Fortress on a query-to-app relevance model in a large-scale app marketplace. Offline experiments demonstrate notable improvements in prediction stability (measured by Coefficient of Variation) and classification performance (measured by PR-AUC).

cs.IR

MHD modeling of magnetic flux evolution around solar maximum by the coronal model COCONUT

In this paper, we simulate the magnetic flux evolution at different heliocentric distances during two solar-maximum Carrington rotations (CRs) using the time-evolving coronal magnetohydrodynamic (MHD) model COCONUT to investigate the ``open flux problem". The simulated open magnetic flux (OMF) near the solar surface is comparable to that derived from \textit{in situ} observations by PSP and WIND satellites, and is about 5 times larger than that derived from SDO coronal hole (CH) observations, and the variation in the simulated radial solar wind speed is consistent with the evolution of the OMF evaluated around the corresponding solar disk center. We find that the OMF is reduced by up to $45\%$ from 1.01~$R_s$ to 0.1~AU and increases with a higher-resolution mesh. The OMF decreases mainly within 3~$R_s$, where the closed magnetic flux drops more rapidly, from about $60\%$ of the total magnetic flux at 1.01~$R_s$ to about $4\%$ at 3~$R_s$. Moderate adjustment of the heating source term can effectively regulate the simulated OMF. Preprocessing the photospheric magnetograms with a potential field solver that removes many high-order spherical harmonic components reduces the OMF in the low corona, while having little impact beyond 3~$R_s$. Additionally, the ratio of the maximum to the minimum OMF can reach 1.4 during a single solar maximum CR. These findings highlight the necessity of considering higher grid resolution, more realistic heating mechanisms, and the time-evolving regime of coronal MHD modeling when further addressing the ``open flux problem".

astro-ph.SR

Parker Solar Probe Observations of Compound Reconnection Exhaust Boundaries and Mirror-Mode Structures in the Near-Sun Heliospheric Current Sheet

Magnetic reconnection is a fundamental physical process that can drive rapid conversion of magnetic energy into plasma bulk flows, thermal heating, and particle acceleration in space and astrophysical plasmas. Classical reconnection theory predicts that the Alfvenic reconnection exhausts are bounded by pairs of slow-mode shocks. However, identifying and characterizing these shocks through in situ spacecraft observations remains a challenge. Here we report Parker Solar Probe (PSP) observations of a reconnection exhaust embedded in the heliospheric current sheet (HCS) at a heliocentric distance of 12.2 R_O. The reconnection exhaust is bounded on both boundaries by compound magnetic structures rather than a pair of pure slow shocks. Each boundary consists of a rapidly evolving, steep inner slow shock, whose Mach numbers and shock-normal angles change significantly within several minutes, and an outer, gradual compound structure which comprises a slow shock and a rotational discontinuity. These slow shocks are quasi-perpendicular and are accompanied by enhanced proton perpendicular heating. Deep within the reconnection exhaust, high perpendicular temperature together with large plasma beta trigger mirror instability and generate mirror-mode structures. These observations provide new insights into the structure of reconnection exhaust boundaries and their role in energy conversion in the near-Sun plasma.

astro-ph.SR

Stylos: Multi-View 3D Stylization with Single-Forward Gaussian Splatting

We present Stylos, a single-forward 3D Gaussian framework for 3D style transfer that operates on unposed content, from a single image to a multi-view collection, conditioned on a separate reference style image. Stylos synthesizes a stylized 3D Gaussian scene without per-scene optimization or precomputed poses, achieving geometry-aware, view-consistent stylization that generalizes to unseen categories, scenes, and styles. At its core, Stylos adopts a Transformer backbone with two pathways: geometry predictions retain self-attention to preserve geometric fidelity, while style is injected via global cross-attention to enforce visual consistency across views. With the addition of a voxel-based 3D style loss that aligns aggregated scene features to style statistics, Stylos enforces view-consistent stylization while preserving geometry. Experiments across multiple datasets demonstrate that Stylos delivers high-quality zero-shot stylization, highlighting the effectiveness of global style-content coupling, the proposed 3D style loss, and the scalability of our framework from single view to large-scale multi-view settings. Our codes are available at https://github.com/HanzhouLiu/Stylos.

cs.CV

Machine learning prediction of plasma behavior from discharge configurations on WEST

Accurately predicting plasma behavior based on discharge configurations is essential for the safe and efficient operation of tokamak experiments. While physics-based integrated modeling codes provide valuable insights, their high computational cost limits their applicability for fast scenario design and control optimization. In this study, we propose a transformer-based machine learning model to predict key global plasma parameters on the WEST tokamak, including the normalized beta ($β_{n}$), toroidal beta ($β_{t}$), poloidal beta ($β_{p}$), plasma stored energy ($W_{\mathrm{mhd}}$), safety factor at the magnetic axis ($q_{0}$), and safety factor at the 95% flux surface ($q_{95}$). The model uses only signals that can be defined before the discharge, such as magnetic coil currents, auxiliary heating power, plasma current reference, and line-averaged plasma density. Trained on 550 discharges from the WEST campaigns, the model demonstrates an average mean square error (MSE) loss of 0.026, an average coefficient of determination $R^{2}$ of 0.94, and achieves inference times on the order of 0.1 seconds. These results highlight the potential of data-driven surrogate models for assisting in discharge planning, scenario evaluation, and real-time control of tokamak plasmas.

physics.plasm-ph

Solar Wind Heating Near the Sun: A Radial Evolution Approach

Characterizing the plasma state in the near-Sun environment is essential to constrain the mechanisms that heat and accelerate the solar wind. In this study, we use Parker Solar Probe (PSP) observations from Encounters 1 through 24 to investigate the radial evolution of solar wind plasma and magnetic field properties in this region. Using intervals with high field-of-view ($>85\%$) coverage, we derive the radial profiles of magnetic field strength ($|B|$), proton density ($N$), bulk speed ($V$), total proton temperature ($T$), parallel ($T_\parallel$) and perpendicular ($T_\perp$) temperatures, temperature anisotropy ($T_\perp/T_\parallel$), plasma beta ($β$), Alfvén Mach number ($M_A$), and magnetic field fluctuations ($δB/B$) for sub and super-Alfvénic regions. In super-Alfvénic regions, power-law of $|B|$, $N$, $V$, and $T$ as a function of heliocentric distance are broadly consistent with previous \textit{Helios} results at $>0.3$ AU. The radial evolution of the components of the temperature tensor reveals distinct behavior: $T_\perp$ decreases monotonically with distance, whereas $T_\parallel$ exhibits a non-monotonic trend -- decreasing in the sub-Alfvénic region, increasing just beyond the Alfvén surface. We interpret the increase in $T_\parallel$ as a proxy for proton beam occurrence. We further examine the evolution of magnetic field fluctuations, finding decreasing radial/parallel fluctuations but enhanced tangential/normal/perpendicular fluctuations in sunward direction. These fluctuations may provide free energy for beam generation and particle heating via wave-particle interactions.

astro-ph.SR

Evidence of energy conversion in weakly collisional plasma during an interplanetary coronal mass ejection

Intervals of enhanced turbulent fluctuations are typically less frequent within the magnetic cloud region of an interplanetary coronal mass ejection (ICME). We investigate two such intervals inside an ICME observed by the \textit{Wind} spacecraft on 8--9 June 2000 and characterize their associated wave populations. We focus on spectral analysis and plasma instability analysis, using ion-scale normalized magnetic helicity and polarization properties with respect to the background magnetic field $B_0$. In the first interval, the ion-scale normalized magnetic helicity shows a left-handed circularly polarized signature. In the second interval, the left-handed signature persists and an additional high-frequency right-handed population appears. The propagation is approximately parallel to $B_0$. The left-handed fluctuations are compatible with Alfvén ion-cyclotron (AIC) waves, while the right-handed fluctuations are consistent with fast magnetosonic/whistler (FM/W) waves. The ICME plasma accesses resonance conditions that support multiple ion-scale wave modes. Evolving anisotropies in the plasma and the approach to marginal stability allow the coexistence of AIC-like and fast-magnetosonic/whistler-like fluctuations, with enhanced electron heating favoring the growth of the FM/W contribution and strengthening the density--magnetic-field magnitude correlation.

astro-ph.SR