arXiv ScienceSearch

arXiv subjects

Jianing Yang

Publications and source records attributed to Jianing Yang.

At least 19 recordsLinked to original sources

Existence of Weak Solutions to a Power-Law Model for Compressible Non-Newtonian Fluids on 1D Unbounded Domain

This paper is concerned with the analysis of a one-dimensional power-law model for compressible fluid dynamics on $\mathbb{R}$, in which the shear stress takes the form $\mu |\partial_{x}u|^{p-2}\partial_{x}u$, where $\mu$ is the viscosity coefficient and $u$ is the velocity. We prove that, in the singular limit $p\rightarrow\infty$, the solutions converge to functions $(\rho,u)$ satisfying $|\partial_{x}u|\leq 1$, $\tau = \pi \partial_{x}u$, $\pi \geq 0$, and $\pi (1 - |\partial_{x}u|) = 0$ a.e. on $\mathbb{R}$. Moreover, we rigorously justify the existence of weak solutions to the limiting equation. The convergence as $p \to \infty$ is obtained via domain truncation and compactness arguments, of which the key challenge is to show that the density remains bounded away from zero and infinity on any compact subset. This extends the recent result of Bresch, Burtea, and Szlenk [Nonlinearity 26 (2026), no. 5, Paper No. 055010.] from one-dimensional periodic domain to the whole real line.

math.AP

Aperture-aware Dispersion 5-D Light-field Imaging Spectrometer

Enhancing perceptual dimensions while miniaturizing imaging systems presents significant challenges for high-dimensional visual sensing. Conventionally, the acquisition of the 5D (x,y,u,v,{\lambda}) spectral light field (5D-SLF) data cube relies on bulky and expensive camera arrays, which are impractical for widespread application. Existing single-detector systems are fundamentally limited by a trade-off between the resolutions of different dimensions owing to insufficient coding capabilities. Here we introduce an Aperture-aware Dispersion Light-field Imaging Spectrometer (ADLIS), that targets a synergy between compactness and resolution through aperture-multiplexed modulation, leveraging the inherent spectral-filtering properties of birefringent material. Using only a manufacturing-friendly and cost-effective phase plate made of birefringent quartz crystal, the aperture of the proposed ADLIS enables compact angular-spectral encoding that is highly sensitive to both the incident angle and spectrum of incoming light. In contrast to the viewpoint-separation approach of microlens arrays, ADLIS employs aperture encoding to superimpose all viewpoints onto each sensor pixel. This shifts the design paradigm from spatial division to encoding integration, aiming to achieve full-resolution light field recovery. Thus, we develop the Aperture-aware Dispersion Light-field Imaging (ADLI) framework, which optimizes the aperture design and 5D-SLF reconstruction in an end-to-end (E2E) manner. Trained by simulation data and validated through real-world experiments, our system achieves robust high-performance 5D-SLF imaging while maintaining full spatial resolution.

cs.CV

Observation of a tripartite quantum phase for coexisting extended, localized, and critical states

The disordered quantum world hosts three fundamental types of states: extended, localized, and critical, of which the critical states are confined to fine-tuned critical points or mobility edges in randomly disordered systems. The tripartite phase, with all three types of states coexisting over finite spectral windows, represents a hallmark distinction between quasiperiodic and truly random systems in the localization physics. Here, we report the realization of this exotic phase in a quasi-periodically driven orbital optical lattice with ultracold atoms. The optical lattice with a quasiperiodic Floquet modulation coupling s and p orbitals is realized in experiment and shown to host the tripartite phase from exact theory. We develop a two-stage protocol to precisely prepare and detect the three types of quantum states. The characteristic exponents of these states are determined from expansion dynamics, showing their distinct universal transport properties. Our study marks a significant advancement in exploring unconventional critical phenomena and localization physics with ultracold atoms.

cond-mat.quant-gas

DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization

Spoken dialog systems with cascaded ASR-LLM-TTS modules retain strong LLM intelligence, but VAD segmentation often forces half-duplex turns and brittle control. On the other hand, VAD-free end-to-end model support full-duplex interaction but is hard to maintain conversational intelligence. In this paper, we present DuplexCascade, a VAD-free cascaded streaming pipeline for full-duplex speech-to-speech dialogue. Our key idea is to convert conventional utterance-wise long turns into chunk-wise micro-turn interactions, enabling rapid bidirectional exchange while preserving the strengths of a capable text LLM. To reliably coordinate turn-taking and response timing, we introduce a set of conversational special control tokens that steer the LLM's behavior under streaming constraints. On Full-DuplexBench and VoiceBench, DuplexCascade delivers state-of-the-art full-duplex turn-taking and strong conversational intelligence among open-source speech-to-speech dialogue systems.

cs.CL

RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily occluded. Recent vision-language-action (VLA) models have begun to incorporate memory mechanisms; however, their evaluations remain confined to narrow, non-standardized settings. This limits systematic understanding, comparison, and progress measurement. To address these challenges, we introduce RoboMME: a large-scale standardized benchmark for evaluating and advancing VLA models in long-horizon, history-dependent scenarios. Our benchmark comprises 16 manipulation tasks constructed under a carefully designed taxonomy that evaluates temporal, spatial, object, and procedural memory. We further develop a suite of 14 memory-augmented VLA variants built on the {\pi}0.5 backbone to systematically explore different memory representations across multiple integration strategies. Experimental results show that the effectiveness of memory representations is highly task-dependent, with each design offering distinct advantages and limitations across different tasks. Videos and code can be found at our website https://robomme.github.io.

cs.RO

Global well-posedness for one-dimensional compressible Navier--Stokes system in dynamic combustion with small $BV\cap L^1$ initial data

We establish the global well-posedness theory of small BV weak solutions to a one-dimensional compressible Navier--Stokes model for reacting gas mixtures in dynamic combustion. The unknowns of the PDE system consist of the specific volume, velocity, temperature, and mass fraction of the reactant. For initial data that are small perturbations around the constant equilibrium state $(1, 0, 1, 0)$ in the $L^1(\mathbb{R}) \cap {\rm BV}(\mathbb{R})$-norm, we establish the local-in-time existence of weak solutions via an iterative scheme, show the stability and uniqueness of local weak solutions, and prove the global-in-time existence of solutions for initial data with small BV-norm via an analysis of the Green's function of the linearised system. The large-time behaviour of the global BV weak solutions is also characterised. This work is motivated by and extends the recent global well-posedness theory for BV weak solutions to the one-dimensional isentropic Navier--Stokes and Navier--Stokes--Fourier systems developed in [T.-P. Liu, S.-H. Yu, Commun. Pure Appl. Math. 75 (2022), 223--348] and [H. Wang, S.-H. Yu, X. Zhang, Arch. Ration. Mech. Anal. 245 (2022), 375--477].

math.AP

DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction

With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However, during fine-tuning, these SSL-based MOS prediction models often suffer from catastrophic forgetting of the pretrained knowledge and tend to overfit the training set, resulting in poor generalization performance. In this study, we propose DistilMOS, a novel method that learns to predict not only MOS but also token IDs obtained by clustering the hidden representations of each layer in the pretrained SSL model. These layer-wise token targets serve as self-distillation signals that enables the MOS prediction model to extract rich internal knowledge from SSL models, enhancing both prediction accuracy and generalization capability. Experimental evaluations demonstrate that our method significantly outperforms standard SSL-based MOS prediction models on both in-domain and out-of-domain evaluations, verifying the effectiveness and practicality of the proposed method.

cs.SD

SAM 3D: 3Dfy Anything in Images

We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual recognition cues from context play a larger role. We achieve this with a human- and model-in-the-loop pipeline for annotating object shape, texture, and pose, providing visually grounded 3D reconstruction data at unprecedented scale. We learn from this data in a modern, multi-stage training framework that combines synthetic pretraining with real-world alignment, breaking the 3D "data barrier". We obtain significant gains over recent work, with at least a 5:1 win rate in human preference tests on real-world objects and scenes. We will release our code and model weights, an online demo, and a new challenging benchmark for in-the-wild 3D object reconstruction.

cs.CV

Error-Driven Scene Editing for 3D Grounding in Large Language Models

Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. This limitation stems in part from training data that focuses on language reasoning rather than spatial understanding due to scarce 3D resources, leaving inherent grounding biases unresolved. To address this, we propose 3D scene editing as a key mechanism to generate visual counterfactuals that mitigate these biases through fine-grained spatial manipulation, without requiring costly scene reconstruction or large-scale 3D data collection. Furthermore, to make these edits targeted and directly address the specific weaknesses of the model, we introduce DEER-3D, an error-driven framework that diagnoses grounding failures and generates targeted counterfactual training supervision via a structured "Decompose, Diagnose, Edit, and Retrain" loop. Specifically, given a grounding failure, DEER-3D first identifies the predicate-level error (e.g., attribute or spatial relation). It then performs minimal predicate-aligned scene edits, such as recoloring or repositioning, and constructs aligned question-answer pairs that explicitly target the failed predicate, forming targeted counterfactual training examples. We evaluate our editing pipeline across multiple benchmarks for 3D grounding and scene understanding tasks, consistently demonstrating improvements across all grounding datasets through iterative refinement (4-6% gains). DEER-3D underscores the effectiveness of targeted, error-driven scene editing in bridging linguistic reasoning with spatial grounding in 3D LLMs.

cs.CV

Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel emotional TTS method that enables fine-grained phoneme-level emotion embedding prediction while disentangling intrinsic attributes of the reference speech. The proposed method employs a style disentanglement method to guide two feature extractors, reducing mutual information between timbre and emotion features, and effectively separating distinct style components from the reference speech. Experimental results demonstrate that our method outperforms baseline TTS systems in generating natural and emotionally rich speech. This work highlights the potential of disentangled and fine-grained representations in advancing the quality and flexibility of emotional TTS systems.

cs.SD

Quench spectroscopy for Lieb-Liniger bosons in the presence of harmonic trap

Quench spectroscopy has emerged as a novel and powerful technique for probing the energy spectrum of various quantum phases for quantum systems from out-of-equilibrium dynamics. While its efficacy has been demonstrated in the homogeneous systems theoretically, most experimental setups feature a confining potential, such as a harmonic trap, which complicates the practical implementations. In this work, we experimentally probe the quench spectroscopy for one-dimensional bosons in optical lattices with the presence of a harmonic trap, and comparing our results with the density matrix renormalization group simulation. For the Mott insulator phase, although a gap is still observed, the band signal is broadened along the frequency space and cut at the half Brillouin zone, which can be explained by the nearest-neighbor tunneling excitations under harmonic confinement. Comparing with the superfluid spectrum, we can see a clear distinction between the two phases and find the inverse quench with larger amplitude yields the clearest spectrum. Our work offers pivotal insights into conducting quench spectroscopy effectively in practical systems.

cond-mat.quant-gas

4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction model that takes input from unconstrained views and timestamps and renders arbitrary novel view-time combinations. Unlike prior 4D approaches, e.g., optimization-based, geometry-based, or generative, that struggle with efficiency, generalization, or faithfulness, 4D-LRM learns a unified space-time representation and directly predicts per-pixel 4D Gaussian primitives from posed image tokens across time, enabling fast, high-quality rendering at, in principle, infinite frame rate. Our results demonstrate that scaling spatiotemporal pretraining enables accurate and efficient 4D reconstruction. We show that 4D-LRM generalizes to novel objects, interpolates across time, and handles diverse camera setups. It reconstructs 24-frame sequences in one forward pass with less than 1.5 seconds on a single A100 GPU.

cs.CV

SAB3R: Semantic-Augmented Backbone in 3D Reconstruction

We introduce a new task, Map and Locate, which unifies the traditionally distinct objectives of open-vocabulary segmentation - detecting and segmenting object instances based on natural language queries - and 3D reconstruction, the process of estimating a scene's 3D structure from visual inputs. Specifically, Map and Locate involves generating a point cloud from an unposed video and segmenting object instances based on open-vocabulary queries. This task serves as a critical step toward real-world embodied AI applications and introduces a practical task that bridges reconstruction, recognition and reorganization. To tackle this task, we introduce a simple yet effective baseline, which we denote as SAB3R. Our approach builds upon MASt3R, a recent breakthrough in 3D computer vision, and incorporates a lightweight distillation strategy. This method transfers dense, per-pixel semantic features from 2D vision backbones (eg, CLIP and DINOv2) to enhance MASt3R's capabilities. Without introducing any auxiliary frozen networks, our model generates per-pixel semantic features and constructs cohesive point maps in a single forward pass. Compared to separately deploying MASt3R and CLIP, our unified model, SAB3R, achieves superior performance on the Map and Locate benchmark. Furthermore, we evaluate SAB3R on both 2D semantic segmentation and 3D tasks to comprehensively validate its effectiveness.

cs.CV

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-order-of-magnitude gap that severely limits performance. We introduce $\textbf{LIFT-GS}$, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This render-supervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with $25.7\%$ mAP on open-vocabulary instance segmentation (vs. $20.2\%$ prior SOTA) and consistent $10-30\%$ improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies fine-tuning datasets by 2X, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project page: https://liftgs.github.io

cs.CV

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R's Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy.

cs.CV

On global existence and large-time behaviour of weak solutions to the compressible barotropic Navier--Stokes Equations on $\mathbb{T}^2$ with density-dependent bulk viscosity: beyond the Va\u{\i}gant--Kazhikhov regime

We are concerned with the compressible barotropic Navier--Stokes equations for a $\gamma$-law gas with density-dependent bulk viscosity coefficient $\lambda=\lambda(\rho)=\rho^\beta$ on the two-dimensional periodic domain $\mathbb{T}^2$. The global existence of weak solutions with initial density bounded away from zero and infinity for $\beta>3$, $\gamma>1$ has been established by Va\u{\i}gant--Kazhikhov [Sib. Math. J. 36 (1995), 1283--1316]. When $\gamma=\beta>3$, the large-time behaviour of the weak solutions and, in particular, the absence of formation of vacuum and concentration of density as $t \to \infty$, has been proved by Perepelitsa [\textit{SIAM J. Math. Anal.} 39 (2007/08), 1344--1365]. Huang--Li [J. Math. Pures Appl. 106 (2016), 123--154] extended these results by establishing the global existence of weak solutions and large-time behaviour under the assumptions $\beta >3/2$, $1< \gamma<4\beta-3$, and that the initial density stays away from infinity (but may contain vacuum). Improving upon the works listed above, we prove that in the regime of parameters as in Huang--Li, namely that $\beta >3/2$ and $1< \gamma<4\beta-3$, if the density has no vacuum or concentration at $t=0$, then it stays away from zero and infinity at all later time $t \in ]0,\infty[$. Moreover, assuming $\beta>1$, $\gamma>1$ and a technical condition, we establish the global existence of weak solutions on $\mathbb{T}^2$. One of the key ingredients of our proof is a novel application --- motivated by the recent work due to Danchin--Mucha [Comm. Pure Appl. Math. 76 (2023), 3437--3492] --- of Desjardins' logarithmic interpolation inequality.

math.AP

Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use

In real-world scenarios, it is desirable for embodied agents to have the ability to leverage human language to gain explicit or implicit knowledge for learning tasks. Despite recent progress, most previous approaches adopt simple low-level instructions as language inputs, which may not reflect natural human communication. It's not clear how to incorporate rich language use to facilitate task learning. To address this question, this paper studies different types of language inputs in facilitating reinforcement learning (RL) embodied agents. More specifically, we examine how different levels of language informativeness (i.e., feedback on past behaviors and future guidance) and diversity (i.e., variation of language expressions) impact agent learning and inference. Our empirical results based on four RL benchmarks demonstrate that agents trained with diverse and informative language feedback can achieve enhanced generalization and fast adaptation to new tasks. These findings highlight the pivotal role of language use in teaching embodied agents new tasks in an open world. Project website: https://github.com/sled-group/Teachable_RL

cs.CL

Multi-Object Hallucination in Vision-Language Models

Large vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images. While current benchmarks for object hallucination primarily concentrate on the presence of a single object class rather than individual entities, this work systematically investigates multi-object hallucination, examining how models misperceive (e.g., invent nonexistent objects or become distracted) when tasked with focusing on multiple objects simultaneously. We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate ambiguity. With comprehensive empirical studies and analysis of potential factors leading to multi-object hallucination, we found that (1). LVLMs suffer more hallucinations when focusing on multiple objects compared to a single object. (2). The tested object class distribution affects hallucination behaviors, indicating that LVLMs may follow shortcuts and spurious correlations. (3). Hallucinatory behaviors are influenced by data-specific factors, salience and frequency, and model intrinsic behaviors. We hope to enable LVLMs to recognize and reason about multiple objects that often occur in realistic visual scenes, provide insights, and quantify our progress towards mitigating the issues.

cs.CV