arXiv ScienceSearch

arXiv subjects

Rui Huang

Publications and source records attributed to Rui Huang.

At least 19 recordsLinked to original sources

Development, Evaluation, and Multicenter Clinical-Trial Application of an Artificial Intelligence-Assisted MRI Method for Quantitative Knee Cartilage Morphometry

Objective: To develop and evaluate an AI-assisted MRI method for quantitative knee cartilage morphometry in a multicenter phase III knee osteoarthritis trial. Methods: AI pre-segmentation used 3D full-resolution nnU-Net. Version 1.0 used separate femorotibial- and patellar-cartilage models, whereas version 2.0 used a unified three-class model trained on gold-standard annotations. Trial images then underwent two-reader correction and third-reader adjudication. Adjudicated masks were partitioned into medial/lateral femoral and tibial cartilage plus patellar cartilage. Cartilage volume was measured in physical coordinates, mean thickness by 3D ray tracing (3D-RT), and surface area with local thickness <1.5 mm by a 3D ray-based area method (3D-RBA). Evaluation included 1,189 phase III MRI examinations, reader agreement, 20 synthetic thinning models, and a 69-participant longitudinal comparison with 3D-PMA and three comparator thickness methods. Results: Overall pre-segmentation Dice was 0.964 +/- 0.030 (median 0.970), with 78.7% achieving Dice >=0.95. Inter-reader ICCs for cartilage volume were 0.959-0.995. In the 69-participant subset, total cartilage volume increased from 14,184.366 mm^3 at V0 to 15,359.345 mm^3 at V8; 3D-RBA and 3D-PMA decreased by 4.70% and 6.88%, and all four thickness measures were highest at V8. In 20 geometric experiments, MAPE was 5.73%, CCC 0.822, and Dice 0.956. The workflow was applied to 1,188 MRI examinations from 416 participants. From V0 to V8, the treatment group showed +3.45% total cartilage volume, +2.46% mean thickness, and -4.54% 3D-RBA, versus -2.08%, -1.32%, and +0.16% in controls. Conclusion: This workflow provided a reproducible MRI cartilage assessment framework for a multicenter KOA trial. Cross-method agreement and geometric validation supported 3D-RT and 3D-RBA for therapeutic efficacy evaluation.

eess.IV

Transfer Learning with Heterogeneous Feature Spaces in Linear Regression

Transfer learning improves target-task performance by leveraging related source data. Most methods assume shared feature spaces, yet in many applications, each source observes only a subset of target covariates. Classical imputation fails here due to block missingness, and standard imputation matrices are not optimized for target parameter estimation. We study low- and high-dimensional linear regression and propose Heterogeneous Importance Weighting (HIW). Our method aligns feature spaces via projection-based imputation and transfers information through sample-selected importance weighting. This framework accommodates diverse projection matrices to construct target-oriented imputation. We develop a classification-based procedure with pseudo-responses to estimate conditional error densities for the weights. We establish entry-wise and global convergence rates for the estimator, with numerical and real-data studies demonstrating its effectiveness.

stat.ME

Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature caching is widely used to accelerate diffusion transformers; however, as the number of steps increases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an online regression model to dynamically correct this deviation, but it faces the issue of label data being unavailable during the acceleration process. This paper presents a statistical observation that the residuals between the features of full computation steps using caching methods and reference full-compute trajectory locally exhibit a zero-mean Gaussian distribution. By treating the features of full computation steps as noisy observations of reference features, the data acquisition problem is resolved. Based on this observation, a plug-and-play GP-Refiner correction framework is proposed. This method utilizes Gaussian Process Regression for correction and, leveraging the properties of GPR, introduces an uncertainty-adaptive computation strategy that triggers necessary full-computation calibration by monitoring the posterior variance in real time. Experiments demonstrate significant improvements across different models when combined with various state-of-the-art methods. Integrating the proposed framework with TaylorSeer reduces the computational load by 19.3% while improving PSNR by 0.9 dB and reducing LPIPS from 0.46 to 0.29. Code is available in https://github.com/Aredstone/GP-Refiner.

cs.CV

LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus

Weakly supervised 3D occupancy prediction reduces the reliance on costly 3D annotations by learning from 2D pseudo-labels generated by vision foundation models. However, existing methods typically use these imperfect pseudo-labels directly as supervision, making occupancy learning vulnerable to erroneous geometric and semantic targets. We observe that agreement across repeated observations provides an inexpensive and reliable cue for assessing pseudo-label reliability. Based on this observation, we propose \textbf{LetOccVote}, a weakly supervised Gaussian-based occupancy framework that leverages cross-frame voting to improve both geometric and semantic supervision. For geometry, Depth Vote exploits cross-frame geometric agreement to refine supported pseudo depth and reject contradictory estimates before volumetric lifting and depth supervision. For semantics, Semantic Vote aggregates pseudo-semantic observations in a shared 3D space to identify reliable and contested evidence, strengthening reliable semantic supervision while filtering unreliable pseudo-label segments. The entire framework is trained solely with 2D pseudo-label supervision without requiring 3D occupancy annotations. On Occ3D-nuScenes, LetOccVote achieves 53.27 IoU and 20.39 mIoU, establishing state-of-the-art performance among methods with 2D pseudo-label supervision.

cs.CV

RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation

Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion- and flow-matching robot policies lack tractable likelihoods, limiting their use in likelihood-based offline RL post-training. AR-NFs offer both expressive action modeling and exact likelihood evaluation, but their sequential sampling incurs substantial sampling overhead during policy optimization and deployment. We present RoMAN-Flow (Robotic Manipulation with Autoregressive Normalizing Flows), an offline reinforcement learning framework that makes AR-NF policies practical for robotic manipulation by addressing this sampling bottleneck in both stages. During policy optimization, RoMAN-Flow employs a sampling-free, advantage-weighted likelihood objective that assigns higher likelihood to high-advantage actions from the offline dataset without sampling from the autoregressive policy. For efficient deployment, it distills the optimized autoregressive policy into a one-step action generator, enabling low-latency action prediction. Experiments across multiple simulated manipulation benchmarks and real-world robotic platforms demonstrate that RoMAN-Flow achieves competitive policy performance while substantially reducing inference latency. Code is available at https://github.com/konnyaku28/RoMAN-Flow.

cs.CV

LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse labware and instruments and execute long-horizon, state-dependent experimental procedures. Yet existing benchmarks do not jointly capture dexterous hand use, real-world laboratory interactions, and multi-stage experimental procedures, limiting systematic training and evaluation. To bridge this gap, we introduce LabDex, a large-scale real-world dataset and benchmark for dexterous manipulation in chemistry laboratories, organized around a hierarchical task taxonomy spanning atomic skills, compositional tasks, and long-horizon experiments. First, LabDex is cross-platform and, for the first time, unifies real-world and simulation platforms under a common framework, providing standardized task definitions, demonstrations, and evaluation protocols. Second, LabDex is large-scale and systematically organizes chemistry laboratory operations into three interconnected levels: Atomic Skills, which characterize fundamental dexterous manipulation capabilities; Compositional Skills; and Long-Horizon Laboratory Workflows. This hierarchical design not only supports the evaluation of end-task performance, but also enables the analysis of how fundamental dexterous skills compose and influence more complex laboratory operations. We conduct cross-level evaluations of representative robot learning methods in both real-world and simulation environments. The experimental results validate the effectiveness of the LabDex task design and demonstration data, and show that the benchmark supports the training and systematic evaluation of existing robotic policies across laboratory dexterous manipulation tasks at different levels, providing a foundation for further research and development of autonomous laboratory robots.

cs.RO

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.0, an efficient and powerful embodied foundation model specifically designed for embodied agents operating in the physical world. To cultivate such capabilities from the pre-training stage onward, we define an action-centric capability taxonomy comprising three progressive dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Guided by this taxonomy, we develop a systematic data pipeline and curate data mixtures spanning both pre-training and post-training. To deliver strong physical-world understanding and interaction capabilities while supporting latency-sensitive deployment, we build our model on the Hy3-A3B language backbone and the Hy-ViT2 vision encoder. Its efficient Mixture-of-Experts architecture combines strong model capacity with high inference efficiency. We evaluate Hy-Embodied-VLM-1.0 on a comprehensive suite of 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning. The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3. Compared with the previous-generation Hy-Embodied-0.5 MoT-2B, Hy-Embodied-VLM-1.0 improves average performance by 8.4%. Despite activating only 3B parameters, it achieves performance close to that of the previous-generation model with 32B activated parameters. Beyond static benchmark evaluation, Hy-Embodied-VLM-1.0 also demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.

cs.CV

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attention compute and KVCache storage significantly compared to Full Attention, realizing these gains in production requires substantial engineering effort. We systematically optimize the KVCache system with layerwise prefetch, SWA-aware prefix cache trees, and specialized placement strategies, achieving strict $O(W)$ SWA storage and high cache hit rates. We further build GCache, a high-performance distributed cache infrastructure with RDMA-optimized networking, and develop a KVCache-affinity router to reduce computation while preserving load balancing. We also optimize for multimodal inputs, including GPU image preprocessing, parallel video decoding, and multimodal cache sharing. Together, these optimizations constitute the first large-scale LLM serving system in production that efficiently covers the Hybrid SWA + MoE + multimodal composite architecture.

cs.AR

Correlated Plasmonic Excitation in Twisted Nematic Plasmonic Superlattices

Superlattices with twisted configurations, such as moire lattices, have recently been extensively exploited for their unique electronic, magnetic, and optical properties. One remarkable feature of nanoscale twisted superlattices is the distinct lattice symmetries and the continuous phase transitions between periodic or aperiodic phases, representing a unique opportunity to study many emerging physical phenomena. Here, we report a correlated light and matter interaction between the collective polarization effect of nematic plasmonic superstructures and the plasmonic excitation of individual constituent nanorods in reconfigurable twisted plasmonic superlattices. Using hybrid Fe3O4 and Au nanorods as building blocks, we assembled plasmonic nematic liquid crystals with unidirectionally aligned nanorods, which could be further assembled into moire plasmonic lattices through a vertical stacking assembly method. A twist angle dependent plasmonic excitation is recognized in the twisted bilayer of two plasmonic superlattices, featuring enhanced transverse and longitudinal plasmonic excitation at a twisting angle of 0 degree and 90 degree, respectively. Such correlated plasmonic excitation in twisted plasmonic superstructures is induced by the correlation between the collective polarization effect of the liquid crystal phases and the anisotropic plasmonic excitation of individual nanorods. The magnetic orientation control allows for precise alignment of hybrid Fe3O4 and Au nanorods in polymer substrates and enables the coding of nematic domains and plasmonic patterns in each sublattice. The correlated plasmonic excitation and light polarization create reconfigurable photonic moire superlattices with well-defined domain colors, feature sizes, periodicities, symmetries, and dimensions determined by twist angles and displacements in the twisted plasmonic lattices.

cond-mat.mtrl-sci

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search environments, and reward signals independently, causing synthesized structural metadata to be discarded, environments to rely on irreproducible external engines, and RL rewards to remain sparse at the trajectory level. We present \textbf{SearchEyes}, which uses a typed knowledge graph as the backbone of a \emph{simulated search world} that unifies all three components. We propose \textbf{Perception-Knowledge Chains (PKC)} to sample constrained multi-hop paths over the visual-knowledge intersection of Wikidata5M, retaining hop-level entity metadata that simultaneously defines a self-contained search world and step-level reward anchors. We further propose \textbf{Hop-Anchored Policy Optimization (HaPO)}, which reuses these anchors for step-level credit assignment without a separately trained process reward model. Experiments on six multimodal knowledge-intensive benchmarks show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.%

cs.AI

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead. Visual token pruning offers a natural solution, yet existing methods are imperfect: attention-based criteria tend to retain redundant tokens, while diversity-based criteria are often agnostic to user instructions. Even methods that combine multiple criteria still lack a principled formulation of the intrinsic objective of token pruning. In this paper, we revisit visual token pruning from a first-principles perspective and formulate it as constructing Token Optimal Preservation Sets. Through a top-down information-theoretic analysis, we identify three fundamental principles for effective token selection: Task Relevance, Information Coverage, and Semantic Diversity. Based on these principles, we propose TOPS, a training-free and model-agnostic pruning module that can be applied to various MLLMs. Extensive experiments on 7 MLLM backbones and 14 benchmarks demonstrate that TOPS outperforms prior methods under diverse pruning settings. Notably, on LLaVA-NeXT, TOPS removes 77.8% of visual tokens while preserving 100.0% and 100.6% performance on its 7B and 13B models, respectively, suggesting that pruning redundant visual tokens can sometimes mitigate hallucination and inspire future lightweight MLLM design.

cs.AI

Denoising Implicit Feedback for Cold-start Recommendation

Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g., clickbait, position bias). Meanwhile, recommenders inevitably face the item cold-start problem due to the continuous influx of new items. We identify that cold items are more prone to noisy samples due to the aforementioned factors, and researchers often overlook the significance of denoising implicit feedback for cold items. Previous denoising studies usually identify noisy samples based on heuristic patterns, such as higher loss values, and mitigate noise through sample selection or re-weighting. However, these methods have limited adaptability and are ineffective in cold-start scenarios. To achieve denoising implicit feedback for cold-start recommendation, we propose a model-agnostic denoising method called DIF. First, user preferences for content remain stable, which allows us to infer pseudo-labels indicating whether a user is interested in a cold item through content-similar warm items. Furthermore, to improve pseudo-label accuracy, we model the confidence of pseudo-labels based on the content similarity between the cold item and warm items, and then aggregate multiple pseudo-labels for each sample. Finally, we explicitly estimate the uncertainty of the noisy sample label by considering its relative entropy and the cold-start status of the item, which adaptively guides the role of pseudo-labels to correct the noisy labels at the sample level. DIF's superiority is supported by both theoretical justification and extensive experiments on real-world datasets. The method has been deployed on a billion-user scale short video application Kuaishou and has significantly improved various commercial metrics within cold-start scenarios.

cs.AI

ChargeBD: Character-Aware Heterogeneous Agent Reasoning for Guided Engineering in Battery Development

Redox-flow battery (RFB) research spans molecular design, electrolyte optimization, electrode and membrane materials, stack operation, system management, and safety analysis, making it a constrained, multi-scale, and multi-objective energy-storage R&D problem. Although large language models (LLMs) can support scientific knowledge integration and proposal generation, generic LLM reasoning remains insufficiently adaptive across innovation-oriented exploration, rule-based execution, mechanistic modeling, and system-level trade-offs. Here we introduce ChargeBD, a character-aware heterogeneous-agent reasoning framework for guided engineering in battery development. Starting from a 50-question RFB-specific task set, we construct a 500-question ESS-LLM Benchmark and define MBTI-inspired persona agents as structured cognitive-bias templates rather than psychometric instruments or representations of real personalities. DeepSeek-V3-Plus is selected as the shared base model, and 16 MBTI-inspired persona agents are evaluated to construct a persona capability matrix and a cognitive advantage matrix.

stat.AP

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving

Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for interpolation. When extended to future extrapolation, they suffer from ghosting artifacts under large displacements and are constrained by simplified motion assumptions or strict future priors. To overcome these challenges, we propose Envision4D, a fully self-supervised feed-forward framework for pose-free future extrapolation. Specifically, we introduce a Future Pose Prediction module that infers future camera parameters via an iterative denoising process. Furthermore, to capture non-linear dynamics, we propose In-layer Temporal Attention and employ Conditioned Motion Lifting, which transforms the highly uncertain extrapolation process into robust relational mappings. Finally, a Progressive Training Strategy is utilized to stabilize unsupervised motion learning against error accumulation. Extensive experiments demonstrate that Envision4D achieves state-of-the-art performance, significantly outperforming existing methods in future view synthesis.

cs.CV

DIffuse X-ray Explorer (DIXE): Sky Survey Strategy and Collimator Response Demodulation

DIffuse X-ray Explorer (DIXE) is a proposed high-resolution X-ray spectroscopic surveyor aimed at studying large structures of hot gas in the Milky Way. Its payload is designed to have a field of view (FoV) of $10^\circ$ (half-power diameter) and an energy resolution of better than 6 eV, covering an energy range of 0.1-10 keV. It will be mounted on the China Space Station (CSS) and follow the CSS orbit to conduct the survey with fixed zenith pointing in order to optimize the coverage of key science targets. The payload will avoid the Sun passively via an operable sunshade, where a minimum $25^\circ$ angular separation between the pointing axis and the direction of the Sun is required. Two Sun-avoidance strategies are considered: one focusing on minimizing mechanical risk and the other on maximizing exposure time. The one-year exposure maps indicate that DIXE will cover approximately $72.5\%$ of the sky, with typical exposure times of 26 ks and 68 ks for the two strategies, respectively. Although mechanically collimated, the imaging performance of the payload can be enhanced with a demodulation method based on Markov Chain Monte Carlo sampling using the collimator response. Through simulation, we found that the method could achieve a localization accuracy of $1^\circ$ for point-like sources and a spatial resolution of $3^\circ$ for the extended sources of complex surface brightness distribution, both of which are significantly smaller than the FoV.

astro-ph.IM

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation

Recent progress in speech dialogue systems requires Text-to-Speech (TTS) models to be faster and more responsive. Modern speech dialogue systems impose two primary requirements on TTS models: low latency and support for streaming inputs and outputs. However, most existing single-codebook LLM-based TTS methods rely on multi-stage pipelines that lack native streaming capabilities. These systems typically suffer from high end-to-end latency due to slow autoregressive prediction and multi-step flow matching. To address these limitations, we propose FlashTTS, an open-source and low-latency streaming TTS framework. FlashTTS introduces a lagged multi-track architecture that natively processes streaming text and speech inputs, thereby eliminating the need for sentence-level buffering. To accelerate acoustic generation, we integrate parallel Multi-Token Prediction (MTP) with an X-pred mean flow matching decoder. This configuration achieves high-fidelity token-to-mel generation in exactly two function evaluations (2-NFE). By jointly optimizing input processing and decoding efficiency, FlashTTS offers a practical foundation for real-time speech dialogue systems. Experiments show that FlashTTS substantially reduces First-Packet Latency to 325ms compared to robust streaming baselines, all while preserving strong zero-shot voice cloning and cross-lingual intelligibility. Speech samples are available. The model code and checkpoints will be released as open source.

eess.AS

Energy Barriers for Reversible Chain Scission and Healing under Tension with Displacement Control

Polymer chain scission is a key mechanism for fracture of soft materials. It is well known from single-molecule force spectroscopy experiments that the critical condition for chain scission depends on the loading rate and other environmental effects (e.g., temperature and solvent). Common approaches to describing the kinetics of chain scission often assume force-controlled conditions, that is, when a polymer chain is stretched by a prescribed force. As a result of this assumption, chain scission is irreversible, excluding the possibility of healing. In many soft materials, however, self-healing has been observed after fracture, suggesting possibly reversible chain scission. Here, we show that reversible chain scission is possible under displacement-controlled conditions, that is, when a polymer chain is stretched with a prescribed end-to-end distance. We present a breakable freely-jointed chain model, assuming that a polymer chain breaks when one of its links breaks while the other links remain nearly rigid. At a prescribed end-to-end distance, the free energy of the chain has two local minima and a local maximum (the transition state), giving rise to energy barriers for chain scission and healing. As the prescribed displacement increases, the energy barrier decreases for scission but increases for healing, depending on the chain length (number of links) and the potential energy of the link. With the energy barriers, we adopt a kinetic approach to predict the statistics and kinetics of a single polymer chain under tension, first by integrating the rate equation and then by kinetic Monte Carlo simulations. Notably, the present model predicts rate-dependent chain scission, with a lower bound for the rupture force that could be several orders of magnitude lower than the upper bound (which is close to the theoretical strength of the covalent bonds).

cond-mat.soft

Detector Development for HUBS I: Initial Testing of Small-Area TES Microcalorimeters

We report progress on the ongoing development of microcalorimeter detector technology for the Hot Universe Baryon Surveyor (HUBS) mission. We show the results from testing and characterizing selected pixels in a 10$\times$10 microcalorimeter array. The microcalorimeter is based on a Mo/Cu transition-edge sensor (TES) coupled to an Au absorber. To better understand the properties of the devices, we have first measured the energy resolution of a selected pixel in a TES array of the same design with a pulsed laser system that produces 3 eV photons, and found that individual photon peaks are easily resolved with the TES, indicating good performance. We have then exposed the microcalorimeter array to radiation from a $^{55}$Fe source, and found that the pixels tested show energy resolutions as good as 3.7$\pm$0.1 eV at 5.9 keV. The energy resolution is found to vary monotonically with the bias point for all the devices, showing little evidence for the presence of the so-called excess noise. This is consistent with the results from modeling the measured noise spectrum. The effects of thermal crosstalk are evident, leading to the degradation of energy resolution.

astro-ph.IM