arXiv ScienceSearch

arXiv subjects

Yadong Liu

Publications and source records attributed to Yadong Liu.

At least 19 recordsLinked to original sources

Higher Order Convergence for the Sharp Interface Limit of 3D Navier--Stokes/Allen--Cahn Systems

We show convergence of solutions to a Navier--Stokes/Allen--Cahn system as the interfacial thickness $\varepsilon>0$ tends to zero for well-prepared initial data as long as the limit system possesses a sufficiently smooth solution. The limit system consists of a two-phase Navier--Stokes system separated by a sharp interface in the presence of surface tension coupled to a convective mean curvature flow equation. In comparison to previous results we obtain improved convergence estimates for higher-order norms. These enable us to prove convergence in the case of three space dimensions and non-constant viscosity, which was unknown before. The convergence results relies crucially on uniform higher-order estimates for the associated linearized Navier--Stokes/Allen--Cahn system in suitably weighted $L^2$-Sobolev spaces. Here a novel problem-adapted weight proportional to the sum of $\varepsilon$ and the distance to the sharp interface of the limit, which gives improved and sharp estimates, is an important new ingredient. This approach can be potentially adapted to other sharp interface limits as well.

math.AP

ConFi-GS Confidence-Guided High-Frequency Injection for 3D Gaussian Splatting Super-Resolution

Reconstructing high-quality 3D scenes from low-resolution multi-view images remains challenging for 3D Gaussian Splatting (3DGS), because insufficient high-frequency observations often lead to blurred textures, weak boundaries, and view-inconsistent details. Existing approaches either apply super-resolution guidance uniformly or localize enhancement regions based mainly on geometric sampling. However, they typically do not distinguish between two fundamentally different questions: where additional detail is needed, and whether the corresponding candidate high-frequency content is reliable enough to be internalized into a multi-view consistent 3D representation. In this paper, we propose a reliability-aware frequency modeling framework for low-resolution 3DGS reconstruction. The framework first estimates a geometry-guided detail-demand prior to locate regions that are likely under-detailed under low-resolution supervision. It then computes a frequency-aware reliability map to determine whether candidate high-frequency details are structurally supported, spectrally unresolved, and cross-view stable. Combining these signals yields a detail-injection map that guides where super-resolved details should be introduced during optimization. Based on this map, we design a unified optimization scheme comprising spatially selective supervision, coarse-to-fine frequency regularization, and reliability-aware Gaussian densification. This scheme controls where reliable details are injected, when high-frequency supervision is activated, and how unresolved yet reliable details are internalized into the Gaussian representation. Experiments on multiple benchmarks show improved fidelity and perceptual quality while suppressing unstable or view-inconsistent details.

cs.CV

Global weak solutions to a diffuse-interface model for quasi-incompressible two-phase flows with unmatched densities and singular potential

We study a thermodynamically consistent diffuse-interface model that describes the motion of two macroscopically immiscible, incompressible, and viscous Newtonian fluids with unmatched densities. This model is compatible with continuum mixture theory. It adopts a mass-averaged (barycentric) velocity so that the two-phase flow is quasi-incompressible: the velocity is no longer divergence-free, and the pressure enters the equation of the chemical potential. For the initial-boundary value problem in $\mathbb{T}^3$ with a class of physically relevant singular free energy densities, we prove the existence of global-in-time weak solutions. The proof relies on a suitable reduction of the original system to a Korteweg-type fluid model combined with a two-layer approximation, together with delicate estimates for the mass density and the phase-field variable inspired by the celebrated Bresch-Desjardins entropy. A key observation is that capillarity at the free interface provides a damping effect on the density evolution. For the limiting procedure, we derive delicate tail estimates to exclude possible concentrations of the singular potential, since no integrability of the pressure is available \textit{a priori}. This work appears to be the first existence result for the Navier-Stokes/Cahn-Hilliard type system with unmatched densities and mass-averaged velocity without spatial regularization.

math.AP

A Framework for Deploying Learning-based Quadruped Loco-Manipulation

Quadruped mobile manipulators offer strong potential for agile loco-manipulation but remain difficult to control and transfer reliably from simulation to reality. Reinforcement learning (RL) shows promise for whole-body control, yet most frameworks are proprietary and hard to reproduce on real hardware. We present an open pipeline for training, benchmarking, and deploying RL-based controllers on the Unitree B1 quadruped with a Z1 arm. The framework unifies sim-to-sim and sim-to-real transfer through ROS, re-implementing a policy trained in Isaac Gym, extending it to MuJoCo via a hardware abstraction layer, and deploying the same controller on physical hardware. Sim-to-sim experiments expose discrepancies between Isaac Gym and MuJoCo contact models that influence policy behavior, while real-world teleoperated object-picking trials show that coordinated whole-body control extends reach and improves manipulation over floating-base baselines. The pipeline provides a transparent, reproducible foundation for developing and analyzing RL-based loco-manipulation controllers and will be released open source to support future research.

cs.RO

FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation

Multimodal sentiment analysis remains a challenging task due to the inherent heterogeneity across modalities. Such heterogeneity often manifests as asynchronous signals, imbalanced information between modalities, and interference from task-irrelevant noise, hindering the learning of robust and accurate sentiment representations. To address these issues, we propose a factorized multimodal fusion framework that first disentangles each modality into shared and unique representations, and then suppresses task-irrelevant noise within both to retain only sentiment-critical representations. This fine-grained decomposition improves representation quality by reducing redundancy, prompting cross-modal complementarity, and isolating task-relevant sentiment cues. Rather than manipulating the feature space directly, we adopt a mutual information-based optimization strategy to guide the factorization process in a more stable and principled manner. To further support feature extraction and long-term temporal modeling, we introduce two auxiliary modules: a Mixture of Q-Formers, placed before factorization, which precedes the factorization and uses learnable queries to extract fine-grained affective features from multiple modalities, and a Dynamic Contrastive Queue, placed after factorization, which stores latest high-level representations for contrastive learning, enabling the model to capture long-range discriminative patterns and improve class-level separability. Extensive experiments on multiple public datasets demonstrate that our method consistently outperforms existing approaches, validating the effectiveness and robustness of the proposed framework.

cs.MM

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses despite safety alignment. Existing defenses typically improve safety through external filtering, auxiliary guardrails, or decoding-time control. However, these interventions often reduce practical deployability because they may require additional model access, introduce extra inference cost, or affect benign-task utility. In this paper, we propose Safety-Aware Intent Defense (SAID), a training-free jailbreak defense framework based on intent-level safety probing. SAID first distills potentially obfuscated user inputs into concise core intents using the target model itself. It then applies a validated safety prefix to probe each distilled intent and elicit the model's safety-aware response. Finally, a conservative aggregation rule rejects the original request if any distilled intent is identified as unsafe. This design enables black-box-compatible defense without updating model parameters or modifying the decoding process. Experiments on four open-source LLMs under six representative jailbreak attacks show that SAID achieves state-of-the-art defense performance in reducing harmful responses while maintaining competitive utility on benign tasks. Further analyses on prefix variants, hierarchical distillation, and inference efficiency demonstrate that SAID provides a practical safety-utility trade-off for securing LLMs against jailbreak threats.

cs.CR

Weak solutions and incompressible limit of a quasi-incompressible Navier--Stokes/Cahn--Hilliard model for viscous two-phase flows

We study a quasi-incompressible Navier--Stokes/Cahn--Hilliard coupled system which describes the motion of two macroscopically immiscible incompressible viscous fluids with partial mixing in a small interfacial region and long-range interactions. The case of unmatched densities with mass-averaged velocity is considered so that the velocity field is no longer divergence-free, and the pressure enters the equation of the chemical potential. We first prove the existence of global weak solutions to the model in a three-dimensional periodic domain, for which the implicit time discretization together with a fixed-point argument to the approximate system is employed. In particular, we obtain a new regularity estimate of the order parameter by exploiting the partial damping effect of the capillary force. Then utilizing the relative entropy method, we establish the incompressible limit -- the quasi-incompressible two-phase model converges to model H as the density difference tends to zero. Crucial to the passage of the incompressible limit, due to the lack of regularity of the pressure, are some non-standard uniform-in-density difference controls of the pressure, which are derived from the structure of the momentum equations and the improved regularity of the order parameter.

math.AP

A Thermodynamically Consistent Free Boundary Model for Two-Phase Flows in an Evolving Domain with Bulk-Surface Interaction

We derive a thermodynamically consistent model, which describes the time evolution of a two-phase flow in an evolving domain. The movement of the free boundary of the domain is driven by the velocity field of the mixture in the bulk, which is determined by a Navier--Stokes equation. In order to take interactions between bulk and boundary into account, we further consider two materials on the boundary, which may be the same or different materials as those in the bulk. The bulk and the surface materials are represented by respective phase-fields, whose time evolution is described by a bulk-surface convective Cahn--Hilliard equation. This approach allows for a transfer of material between bulk and surface as well as variable contact angles between the diffuse interface in the bulk and the boundary of the domain. To provide a more accurate description of the corresponding contact line motion, we include a generalized Navier slip boundary condition on the velocity field. Based on local mass balance laws, we derive our model from scratch in two different ways: by the Lagrange Multiplier Approach and (in the case of matched densities and no mass flux between bulk and surface) by the Energetic Variational Approach. We further show that our model generalizes previous models from the literature, which can be recovered from our system by either dropping the dynamic boundary conditions or assuming a static boundary of the domain.

math.AP

Local-in-time existence of strong solutions to a quasi-incompressible Cahn--Hilliard--Navier--Stokes system

We analyze a quasi-incompressible Cahn--Hilliard--Navier--Stokes system (qCHNS) for two-phase flows with unmatched densities. The order parameter is the volume fraction difference of the two fluids, while mass-averaged velocity is adopted. This leads to a quasi-incompressible model where the pressure also enters the equation of the chemical potential. We establish local existence and uniqueness of strong solutions by the Banach fixed point theorem and the maximal regularity theory.

math.AP

Comparative Analysis of Extrinsic Factors for NER in French

Named entity recognition (NER) is a crucial task that aims to identify structured information, which is often replete with complex, technical terms and a high degree of variability. Accurate and reliable NER can facilitate the extraction and analysis of important information. However, NER for other than English is challenging due to limited data availability, as the high expertise, time, and expenses are required to annotate its data. In this paper, by using the limited data, we explore various factors including model structure, corpus annotation scheme and data augmentation techniques to improve the performance of a NER model for French. Our experiments demonstrate that these approaches can significantly improve the model's F1 score from original CRF score of 62.41 to 79.39. Our findings suggest that considering different extrinsic factors and combining these techniques is a promising approach for improving NER performance where the size of data is limited.

cs.CL

Integrating Posture Control in Speech Motor Models: A Parallel-Structured Simulation Approach

Posture is an essential aspect of motor behavior, necessitating continuous muscle activation to counteract gravity. It remains stable under perturbation, aiding in maintaining bodily balance and enabling movement execution. Similarities have been observed between gross body postures and speech postures, such as those involving the jaw, tongue, and lips, which also exhibit resilience to perturbations and assist in equilibrium and movement. Although postural control is a recognized element of human movement and balance, particularly in broader motor skills, it has not been adequately incorporated into existing speech motor control models, which typically concentrate on the gestures or motor commands associated with specific speech movements, overlooking the influence of postural control and gravity. Here we introduce a model that aligns speech posture and movement, using simulations to explore whether speech posture within this framework mirrors the principles of bodily postural control. Our findings indicate that, akin to body posture, speech posture is also robust to perturbation and plays a significant role in maintaining local segment balance and enhancing speech production.

eess.AS

Weak solutions and singular limits for a compressible fluid-structure interaction problem with slip boundary conditions

We study a system describing the compressible barotropic fluids interacting with (visco) elastic solid shell/plate. In particular, the elastic structure is part of the moving boundary of the fluid, and the Navier-slip type boundary condition is taken into account. Depending on the reference geometry (flat or not), we show the existence of weak solutions to the coupled system provided the adiabatic exponent satisfies $\gamma > \frac{12}{7}$ without damping and $\gamma > \frac{3}{2}$ with structure damping, utilizing the domain extension and regularization approximation. Moreover, via a modified relative entropy method in time-dependent domains, we give a rigorous justification of the incompressible inviscid limit of the compressible fluid-structure interaction problem with a flat reference geometry, in the regime of low Mach number, high Reynolds number, and well-prepared initial data. As a byproduct, with a fixed Reynolds number, we derive the incompressible limit without extra assumption. To the best of our knowledge, this is the first result concerning the singular limit problem for compressible fluids interacting with elastic structures.

math.AP

High-coherence parallelization in integrated photonics

Coherent optics has profoundly impacted diverse applications ranging from communications, LiDAR to quantum computations. However, building coherent systems in integrated photonics previously came at great expense in hardware integration and energy efficiency: the lack of a power-efficient way to generate highly coherent light necessitates bulky lasers and amplifiers, while frequency and phase recovery schemes require huge digital signal processing resources. In this work, we demonstrate a high-coherence parallelization strategy that facilitates advanced integrated coherent systems at a minimum price. Using a self-injection locked microcomb to injection lock a distributed feedback laser array, we boost the microcomb power by a record high gain of up to 60 dB on chip with no degradation in coherence. This strategy enables tens of highly coherent channels with an intrinsic linewidth down to the 10 Hz level and power of more than 20 dBm. The overall electrical to optical wall-plug efficiency reaches 19%, comparable with that of the state-of-the-art semiconductor lasers. Driven by this parallel source, we demonstrate a silicon photonic communication link with an unprecedented data rate beyond 60 Tbit/s. Importantly, the high coherence we achieve reduces the coherent-related DSP consumption by 99.999% compared with the traditional III-V laser pump scheme. This work paves a way to realizing scalable, high-performance coherent integrated photonic systems, potentially benefiting numerous applications.

physics.optics

Algorithms for Object Detection in Substations

Inspection of high-voltage power equipment is an effective way to ensure power supply reliability. Object recognition, one of the key technologies in automatic power equipment inspection, attracts attention of many researchers and engineers. Although quite a few existing models have some their own advantages, object relationship between equipment which is very important in this task is scarcely considered. This paper combining object relationship modeling and Transformer Model proposes a Relation Transformer Model. It has four parts -- backbone, encoder, decoder and prediction heads. With this structure, the proposed method shows in experiments a much better performance than other three commonly used models in object recognition in substation, largely promoting the development of automatic power equipment inspection.

cs.CV

On a diffuse interface model for incompressible viscoelastic two-phase flows

This paper concerns a diffuse interface model for the flow of two incompressible viscoelastic fluids in a bounded domain. More specifically, the fluids are assumed to be macroscopically immiscible, but with a small transition region, where the two components are partially mixed. Considering the elasticity of both components, one ends up with a coupled Oldroyd-B/Cahn--Hilliard type system, which describes the behavior of two-phase viscoelastic fluids. We prove the existence of weak solutions to the system in two dimensions for general (unmatched) mass densities, variable viscosities, different shear moduli, and a class of physically relevant and singular free energy densities that guarantee that the order parameter stays in the physically reasonable interval. The proof relies on a combination of a regularization of the original system and a new hybrid implicit time discretization for the regularized system together with the analysis of an Oldroyd-B type equation.

math.AP

Bootstrapping meaning through listening: Unsupervised learning of spoken sentence embeddings

Inducing semantic representations directly from speech signals is a highly challenging task but has many useful applications in speech mining and spoken language understanding. This study tackles the unsupervised learning of semantic representations for spoken utterances. Through converting speech signals into hidden units generated from acoustic unit discovery, we propose WavEmbed, a multimodal sequential autoencoder that predicts hidden units from a dense representation of speech. Secondly, we also propose S-HuBERT to induce meaning through knowledge distillation, in which a sentence embedding model is first trained on hidden units and passes its knowledge to a speech encoder through contrastive learning. The best performing model achieves a moderate correlation (0.5~0.6) with human judgments, without relying on any labels or transcriptions. Furthermore, these models can also be easily extended to leverage textual transcriptions of speech to learn much better speech embeddings that are strongly correlated with human annotations. Our proposed methods are applicable to the development of purely data-driven systems for speech mining, indexing and search.

cs.CL

Convolutional Neural Networks with A Topographic Representation Module for EEG-Based Brain-Computer Interfaces

Objective: Convolutional Neural Networks (CNNs) have shown great potential in the field of Brain-Computer Interfaces (BCIs). The raw Electroencephalogram (EEG) signal is usually represented as 2-Dimensional (2-D) matrix composed of channels and time points, which ignores the spatial topological information. Our goal is to make the CNN with the raw EEG signal as input have the ability to learn EEG spatial topological features, and improve its performance while essentially maintaining its original structure. Methods:We propose an EEG Topographic Representation Module (TRM). This module consists of (1) a mapping block from the raw EEG signal to a 3-D topographic map and (2) a convolution block from the topographic map to an output of the same size as input. According to the size of the kernel used in the convolution block, we design 2 types of TRMs, namely TRM-(5,5) and TRM-(3,3). We embed the TRM into 3 widely used CNNs, and tested them on 2 publicly available datasets (Emergency Braking During Simulated Driving Dataset (EBDSDD), and High Gamma Dataset (HGD)). Results: The results show that the classification accuracies of all 3 CNNs are improved on both datasets after using the TRM. With TRM-(5,5), the average accuracies of DeepConvNet, EEGNet and ShallowConvNet are improved by 6.54%, 1.72% and 2.07% on EBDSDD, and by 6.05%, 3.02% and 5.14% on HGD, respectively; with TRM-(3,3), they are improved by 7.76%, 1.71% and 2.17% on EBDSDD, and by 7.61%, 5.06% and 6.28% on HGD, respectively. Significance: We improve the classification performance of 3 CNNs on 2 datasets by the use of TRM, indicating that it has the capability to mine the EEG spatial topological information. In addition, since the output of TRM has the same size as the input, CNNs with the raw EEG signal as input can use this module without changing their original structures.

eess.SP