arXiv ScienceSearch

arXiv subjects

Jianbo Yu

Publications and source records attributed to Jianbo Yu.

13 recordsLinked to original sources

Critical Morrey Rigidity and Removable Singularities for Five-Dimensional Stationary Navier-Stokes Flows

We prove a critical Morrey rigidity theorem for the five-dimensional stationary Navier--Stokes equations. More precisely, every smooth solution on $\mathbb R^5\setminus\{0\}$ satisfying \[ \sup_{R>0}R^{-2}\int_{B_R}|u|^3\,dx<\infty \] is identically zero, up to an additive constant in the pressure. This replaces the pointwise Type-I control in the known higher-dimensional rigidity theory by a velocity-only, scale-invariant averaged condition that allows spatial concentration. The proof develops a weak head-pressure mechanism that does not rely on pointwise pressure estimates or classical normal traces. We reconstruct a canonical pressure from the velocity, derive a renormalized inequality for the positive head pressure, and introduce two monotone radial fluxes. Annular energy estimates, suitable-weak compactness, and blow-up and blow-down limits are then used to identify the endpoint fluxes and force rigidity. As an application, we obtain a removable-singularity criterion in dimension five: if a suitable weak solution is smooth away from one point and either its scale-invariant Dirichlet energy or its cubic velocity Morrey quantity remains bounded near that point, then the singularity is removable. Thus, within the isolated-singularity class, the smallness assumption in the classical stationary regularity criterion is replaced by boundedness. We also prove the corresponding velocity-only cubic Morrey rigidity theorem in dimension four by a different finite-energy argument.

math.AP

Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning

Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open challenge. While E(3)-equivariant neural networks excel at point estimates, they lack rigorous confidence measures. We focus on symmetric rank-2 tensor prediction, where the target has six Kelvin--Mandel coordinates and full uncertainty is represented by a $6\times6$ covariance matrix. We introduce a framework for E(3)-equivariant UQ, modeling the full predictive distribution where both mean and covariance preserve rotational symmetry. Our approach decomposes the covariance into irreducible representations $\mathrm{Sym}^2(\rho_c) \cong 2\times(l=0) \oplus 2\times(l=2) \oplus 1\times(l=4)$. By mapping from the flat Lie algebra $\mathfrak{sym}(6)$ to the curved SPD manifold via matrix exponentiation, we strictly ensure positive-definite covariances while maintaining exact equivariance. Furthermore, we formulate a Log-Euclidean Equivariant Scoring Objective (LE-ESO)---a robust surrogate loss based on the Multivariate Laplace distribution---providing robustness to heavy-tailed errors and stable optimization. Validation on ModelNet40 inertia tensors and Materials Project dielectric tensors demonstrates that our method achieves competitive performance and provides physically consistent, symmetry-preserving uncertainty estimates with useful risk and OOD sensitivity.

cs.LG

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation

In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable efficient long-horizon manipulation, we propose GTP-FA (Grasp-Then-Plan with Failure Attribution), a task-oriented two-stage grasp-then-plan framework that generates grasp candidates and performs downstream motion planning conditioned on the selected grasp. Given a failed manipulation trajectory, we learn a failure attribution model that generalizes to unseen grasps and produces a stable distribution over failure modes for diagnosis-guided optimization. Based on these attribution results, we then optimize both modules in a diagnosis-driven manner: on the grasping side, we inject task-level priors and risk penalties into grasp candidate scoring and optimization to suppress unstable or task-incompatible grasps; on the planning side, we target high-risk initial states through data collection and fine-tuning to address genuine planning bottlenecks. We evaluate the proposed framework in both simulation and real-robot experiments, and show that GTP-FA improves the corresponding base learners across RL, IL, diffusion-policy, and VLA-based settings, achieving substantially higher overall task success rates.

cs.RO

Existence and partial regularity of suitable weak solutions to the 3D Navier-Stokes-Vlasov-Fokker-Planck equations

In this paper, we investigate the incompressible Navier-Stokes equations coupled with the Vlasov-Fokker-Planck equation, which describes a two-phase mixture of the viscous incompressible fluid with particles or bubbles through a frictional force term. In the three-dimensional whole space, we construct a new class of suitable weak solutions to the Navier-Stokes-Vlasov-Fokker-Planck system satisfying energy estimates and three local or global energy inequalities of different forms. These obtained local energy inequalities play an important role in characterizing the measure of the singularity set of weak solutions. The main difficulties in deriving these inequalities lie in establishing the convergence of the density function $f$ in bounded or unbounded domains and dealing with the convergence of the non-local frictional force term. The strong convergence of both $f$ and $f \log f$ weighted by $|v|^k$ is proved by exploring some new a priori quantities of the velocity with the help of Tao's $L^p$ decomposition and the DiPerna-Lions compactness method. Moreover, as an immediate consequence of the existence result, we are able to describe the Hausdorff dimension of set of singular points of the fluid velocity $u$ and also establish the $\alpha$-H\"{o}lder continuity of $f$ at the regular points of $u$.

math.AP

IPEC: Test-Time Incremental Prototype Enhancement Classifier for Few-Shot Learning

Metric-based few-shot approaches have gained significant popularity due to their relatively straightforward implementation, high interpret ability, and computational efficiency. However, stemming from the batch-independence assumption during testing, which prevents the model from leveraging valuable knowledge accumulated from previous batches. To address these challenges, we propose a novel test-time method called Incremental Prototype Enhancement Classifier (IPEC), a test-time method that optimizes prototype estimation by leveraging information from previous query samples. IPEC maintains a dynamic auxiliary set by selectively incorporating query samples that are classified with high confidence. To ensure sample quality, we design a robust dual-filtering mechanism that assesses each query sample based on both global prediction confidence and local discriminative ability. By aggregating this auxiliary set with the support set in subsequent tasks, IPEC builds progressively more stable and representative prototypes, effectively reducing its reliance on the initial support set. We ground this approach in a Bayesian interpretation, conceptualizing the support set as a prior and the auxiliary set as a data-driven posterior, which in turn motivates the design of a practical "warm-up and test" two-stage inference protocol. Extensive empirical results validate the superior performance of our proposed method across multiple few-shot classification tasks.

cs.LG

InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery

Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for real-world, open-world applications. However, existing methods often rely on pseudo-labeling, or two-stage clustering, which lack a principled mechanism to explicitly disentangle essential, category-defining signals from instance-specific noise. In this paper, we address this fundamental limitation by re-framing GCD from an information-theoretic perspective, grounded in the Information Bottleneck (IB) principle. We introduce InfoSculpt, a novel framework that systematically sculpts the representation space by minimizing a dual Conditional Mutual Information (CMI) objective. InfoSculpt uniquely combines a Category-Level CMI on labeled data to learn compact and discriminative representations for known classes, and a complementary Instance-Level CMI on all data to distill invariant features by compressing augmentation-induced noise. These two objectives work synergistically at different scales to produce a disentangled and robust latent space where categorical information is preserved while noisy, instance-specific details are discarded. Extensive experiments on 8 benchmarks demonstrate that InfoSculpt validating the effectiveness of our information-theoretic approach.

cs.CV

Enhancing Visual In-Context Learning by Multi-Faceted Fusion

Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single best visual prompt, a practice that often discards valuable contextual information from other suitable candidates. While recent work has explored fusing the top-K prompts into a single, enhanced representation, this still simply collapses multiple rich signals into one, limiting the model's reasoning capability. We argue that a more multi-faceted, collaborative fusion is required to unlock the full potential of these diverse contexts. To address this limitation, we introduce a novel framework that moves beyond single-prompt fusion towards an multi-combination collaborative fusion. Instead of collapsing multiple prompts into one, our method generates three contextual representation branches, each formed by integrating information from different combinations of top-quality prompts. These complementary guidance signals are then fed into proposed MULTI-VQGAN architecture, which is designed to jointly interpret and utilize collaborative information from multiple sources. Extensive experiments on diverse tasks, including foreground segmentation, single-object detection, and image colorization, highlight its strong cross-task generalization, effective contextual fusion, and ability to produce more robust and accurate predictions than existing methods.

cs.CV

Beyond Single Prompts: Synergistic Fusion and Arrangement for VICL

Vision In-Context Learning (VICL) enables inpainting models to quickly adapt to new visual tasks from only a few prompts. However, existing methods suffer from two key issues: (1) selecting only the most similar prompt discards complementary cues from other high-quality prompts; and (2) failing to exploit the structured information implied by different prompt arrangements. We propose an end-to-end VICL framework to overcome these limitations. Firstly, an adaptive Fusion Module aggregates critical patterns and annotations from multiple prompts to form more precise contextual prompts. Secondly, we introduce arrangement-specific lightweight MLPs to decouple layout priors from the core model, while minimally affecting the overall model. In addition, an bidirectional fine-tuning mechanism swaps the roles of query and prompt, encouraging the model to reconstruct the original prompt from fused context and thus enhancing collaboration between the fusion module and the inpainting model. Experiments on foreground segmentation, single-object detection, and image colorization demonstrate superior results and strong cross-task generalization of our method.

cs.CV

EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers

Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful representational capacity. However, fine-tuning such large models demands extensive GPU memory and prolonged training time, making them impractical for many real-world low-resource scenarios. To bridge this gap, we propose EfficientFSL, a query-only fine-tuning framework tailored specifically for few-shot classification with ViT, which achieves competitive performance while significantly reducing computational overhead. EfficientFSL fully leverages the knowledge embedded in the pre-trained model and its strong comprehension ability, achieving high classification accuracy with an extremely small number of tunable parameters. Specifically, we introduce a lightweight trainable Forward Block to synthesize task-specific queries that extract informative features from the intermediate representations of the pre-trained model in a query-only manner. We further propose a Combine Block to fuse multi-layer outputs, enhancing the depth and robustness of feature representations. Finally, a Support-Query Attention Block mitigates distribution shift by adjusting prototypes to align with the query set distribution. With minimal trainable parameters, EfficientFSL achieves state-of-the-art performance on four in-domain few-shot datasets and six cross-domain datasets, demonstrating its effectiveness in real-world applications.

cs.CV

MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context Learning

Log parsing converts semi-structured logs into structured templates, forming a critical foundation for downstream analysis. Traditional syntax and semantic-based parsers often struggle with semantic variations in evolving logs and data scarcity stemming from their limited domain coverage. Recent large language model (LLM)-based parsers leverage in-context learning (ICL) to extract semantics from examples, demonstrating superior accuracy. However, LLM-based parsers face two main challenges: 1) underutilization of ICL capabilities, particularly in dynamic example selection and cross-domain generalization, leading to inconsistent performance; 2) time-consuming and costly LLM querying. To address these challenges, we present MicLog, the first progressive meta in-context learning (ProgMeta-ICL) log parsing framework that combines meta-learning with ICL on small open-source LLMs (i.e., Qwen-2.5-3B). Specifically, MicLog: i) enhances LLMs' ICL capability through a zero-shot to k-shot ProgMeta-ICL paradigm, employing weighted DBSCAN candidate sampling and enhanced BM25 demonstration selection; ii) accelerates parsing via a multi-level pre-query cache that dynamically matches and refines recently parsed templates. Evaluated on Loghub-2.0, MicLog achieves 10.3% higher parsing accuracy than the state-of-the-art parser while reducing parsing time by 42.4%.

cs.SE

Resolution and Robustness Bounds for Reconstructive Spectrometers

Reconstructive spectrometers are a promising emerging class of devices that combine complex light scattering with inference to enable compact, high-resolution spectrometry. Thus far, the physical determinants of these devices' performance remain under-explored. We show that under a broad range of conditions, the noise-induced error for spectral reconstruction is governed by the Fisher information. We then use random matrix theory to derive a closed-form relation linking the variance bound to a set of key physical parameters: the spectral correlation length, the mean transmittance, and the number of frequency and measurement channels. The analysis reveals certain fundamental trade-offs between these physical parameters, and establishes the conditions for a spectrometer to achieve ``super-resolution'' below the limit set by the spectral correlation length. Our theory is confirmed using numerical validations with a random matrix model as well as full-wave simulations. These results establish a physically-grounded framework for designing and analyzing performant and noise-robust reconstructive spectrometers.

physics.optics

Wavelength-scale noise-resistant on-chip spectrometer

Performant on-chip spectrometers are important for advancing sensing technologies, from environmental monitoring to biomedical diagnostics. As device footprints approach the scale of the operating wavelength, previously strategies, including those relying on multiple scattering in diffusive media, face fundamental accuracy constraints tied to limited optical path lengths. Here, we demonstrate a wavelength-scale, CMOS-compatible on-chip spectrometer that overcomes this challenge by exploiting inverse-designed quasinormal modes in a complex photonic resonator. These modes extend the effective optical path length beyond the physical device dimensions, producing highly de-correlated spectral responses. We show that this strategy is theoretically optimal for minimizing spectral reconstruction error in the presence of measurement noise. The fabricated spectrometer occupies a lateral footprint of only 3.5 times the free-space operating wavelength, with a spectral resolution of 10 nm across the 3.59-3.76 micrometer mid-infrared band, which is suitable for molecular sensing. The design of this miniaturized noise-resistant spectrometer is readily extensible to other portions of the electromagnetic spectrum, paving the way for lab-on-a-chip devices, chemical sensors, and other applications.

physics.optics

Liouville type theorems for the fractional Navier-Stokes equations without the integrability condition of velocity in $\mathbb{R}^3$

Motivated by the classification of solutions of harmonic functions, we investigate Liouville type theorems for the fractional Navier-Stokes equations in $\mathbb{R}^3$ under some conditions on the boundedness of fractional derivatives. We prove that the smooth solution must be a trivial solution provided that it uniformly converges to a nonzero constant vector at infinity by applying Lizorkin's multiplier theorem to establish \(L^p\) estimates for the fractional linear Oseen system and Coifman-McIntosh-Meyer type commutator estimates for the dissipation term. It is noteworthy that the integrability of velocity is not required here.

math.AP