arXiv ScienceSearch

arXiv subjects

Qing Yu

Publications and source records attributed to Qing Yu.

At least 37 records · Page 2Linked to original sources

Granular jamming and rheology in microgravity

Understanding how granular materials behave in low gravity is crucial for planetary science and space exploration. It can also help us understand granular phenomena usually hidden by gravity. On Earth, gravity dominates granular behavior, but disentangling its role from intrinsic particle interactions is challenging. We present a series of compression and shear experiments conducted in microgravity using the Center of Applied Space Technology and Microgravity (ZARM) drop tower and GraviTower Bremen (GTB). Our in-house developed experimental setup enables precise measurement of packing density and in-situ shear stress via a Taylor-Couette rheometer. We find that the jamming transition occurs at lower packing density in microgravity than on Earth, confirming that gravity promotes densification. Rheological measurements further reveal that in microgravity, the lack of a secondary force field and predominance of cohesive interparticle forces increase the stress needed for granular media to flow. These findings highlight gravity's dual role in enhancing both compaction and flow, and demonstrate the need for tailored granular models, valid in low- and microgravity environments.

cond-mat.soft

Who You Are Matters: Bridging Topics and Social Roles via LLM-Enhanced Logical Recommendation

Recommender systems filter contents/items valuable to users by inferring preferences from user features and historical behaviors. Mainstream approaches follow the learning-to-rank paradigm, which focus on discovering and modeling item topics (e.g., categories), and capturing user preferences on these topics based on historical interactions. However, this paradigm often neglects the modeling of user characteristics and their social roles, which are logical confounders influencing the correlated interest and user preference transition. To bridge this gap, we introduce the user role identification task and the behavioral logic modeling task that aim to explicitly model user roles and learn the logical relations between item topics and user social roles. We show that it is possible to explicitly solve these tasks through an efficient integration framework of Large Language Model (LLM) and recommendation systems, for which we propose TagCF. On the one hand, TagCF exploits the (Multi-modal) LLM's world knowledge and logic inference ability to extract realistic tag-based virtual logic graphs that reveal dynamic and expressive knowledge of users, refining our understanding of user behaviors. On the other hand, TagCF presents empirically effective integration modules that take advantage of the extracted tag-logic information, augmenting the recommendation performance. We conduct both online experiments and offline experiments with industrial and public datasets as verification of TagCF's effectiveness, and we empirically show that the user role modeling strategy is potentially a better choice than the modeling of item topics. Additionally, we provide evidence that the extracted logic graphs are empirically a general and transferable knowledge that can benefit a wide range of recommendation tasks. Our code is available in https://github.com/Code2Q/TagCF.

cs.IR

Nonlinear Stability of Large-Period Traveling Waves Bifurcating from the Heteroclinic Loop in the FitzHugh-Nagumo Equation

A wave front and a wave back that spontaneously connect two hyperbolic equilibria, known as a heteroclinic wave loop, give rise to periodic waves with arbitrarily large spatial periods through the heteroclinic bifurcation. The nonlinear stability of these periodic waves is established in the setting of the FitzHugh-Nagumo equation, which is a well-known reaction-diffusion model with degenerate diffusion. First, for general systems, we give the expressions of spectra with small modulus for linearized operators about these periodic waves via the Lyapunov-Schmidt reduction and the Lin-Sandstede method. Second, applying these spectral results to the FitzHugh-Nagumo equation, we establish their diffusive spectral stability. Finally, we consider the nonlinear stability of these periodic waves against localized perturbations. We introduce a spatiotemporal phase modulation $\varphi$, and couple it with the associated modulated perturbation $\mathbf{V}$ along with the unmodulated perturbation $\mathbf{\widetilde{V}}$ to close a nonlinear iteration argument.

math.AP

A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models

Out-of-distribution (OOD) detection is a task that detects OOD samples during inference to ensure the safety of deployed models. However, conventional benchmarks have reached performance saturation, making it difficult to compare recent OOD detection methods. To address this challenge, we introduce three novel OOD detection benchmarks that enable a deeper understanding of method characteristics and reflect real-world conditions. First, we present ImageNet-X, designed to evaluate performance under challenging semantic shifts. Second, we propose ImageNet-FS-X for full-spectrum OOD detection, assessing robustness to covariate shifts (feature distribution shifts). Finally, we propose Wilds-FS-X, which extends these evaluations to real-world datasets, offering a more comprehensive testbed. Our experiments reveal that recent CLIP-based OOD detection methods struggle to varying degrees across the three proposed benchmarks, and none of them consistently outperforms the others. We hope the community goes beyond specific benchmarks and includes more challenging conditions reflecting real-world scenarios. The code is https://github.com/hoshi23/OOD-X-Benchmarks.

cs.CV

Determination of $\alpha_s(M_Z)$ via a high-precision effective coupling $\alpha^{g_1}_s(Q)$

We propose a novel method to determine the strong coupling of quantum chromodynamics (QCD) and fix its running behavior at all scales by using the Bjorken sum rules (BSR). The BSR defines an effective coupling $\alpha^{g_1}_s(Q)$ which includes the nonperturbative high-twist corrections and perturbative QCD (pQCD) corrections to the leading-twist part. For the leading-twist part of $\alpha^{g_1}_s(Q)$, we adopt the infinite-order scale-setting procedure of the principle of maximum conformality ($\rm{PMC}_\infty$) to deal with its pQCD corrections, which reveals the intrinsic conformality of series and eliminates conventional renormalization scheme-and-scale ambiguities. Using the $\rm{PMC}_\infty$ approach, we not only eliminate \textit{the first kind of residual scale dependence} due to uncalculated higher-order terms, but also resolve the previous ``self-consistence problem". The holographic light-front QCD model is used for $\alpha^{g_1}_s(Q)$ in the infrared region, which also reveals a conformal behavior at $Q\to 0$. As a combination, we obtain a precise $\alpha^{g_1}_s(Q)$ at all scales, which matches well with the known experimental data with $p$-value $\sim99\%$, we determine the strong coupling constant at the critical scale $M_Z$, $\alpha_s(M_Z)=0.1191\pm{0.0012}\mp0.0006$, where the first error comes from $\Delta\kappa$ of LFHQCD model and the second error is from \textit{the second kind of residual scale dependence} that is negligible.

hep-ph

Global Estimation of Building-Integrated Facade and Rooftop Photovoltaic Potential by Integrating 3D Building Footprint and Spatio-Temporal Datasets

This research tackles the challenges of estimating Building-Integrated Photovoltaics (BIPV) potential across various temporal and spatial scales, accounting for different geographical climates and urban morphology. We introduce a holistic methodology for evaluating BIPV potential, integrating 3D building footprint models with diverse meteorological data sources to account for dynamic shadow effects. The approach enables the assessment of PV potential on facades and rooftops at different levels-individual buildings, urban blocks, and cities globally. Through an analysis of 120 typical cities, we highlight the importance of 3D building forms, cityscape morphology, and geographic positioning in measuring BIPV potential at various levels. In particular, our simulation study reveals that among cities with optimal facade PV performance, the average ratio of facade PV potential to rooftop PV potential is approximately 68.2%. Additionally, approximately 17.5% of the analyzed samples demonstrate even higher facade PV potentials compared to rooftop installations. This finding underscores the strategic value of incorporating facade PV applications into urban sustainable energy systems.

cs.IR

Multiscale spatiotemporal heterogeneity analysis of bike-sharing system's self-loop phenomenon: Evidence from Shanghai

Bike-sharing is an environmentally friendly shared mobility mode, but its self-loop phenomenon, where bikes are returned to the same station after several time usage, significantly impacts equity in accessing its services. Therefore, this study conducts a multiscale analysis with a spatial autoregressive model and double machine learning framework to assess socioeconomic features and geospatial location's impact on the self-loop phenomenon at metro stations and street scales. The results reveal that bike-sharing self-loop intensity exhibits significant spatial lag effect at street scale and is positively associated with residential land use. Marginal treatment effects of residential land use is higher on streets with middle-aged residents, high fixed employment, and low car ownership. The multimodal public transit condition reveals significant positive marginal treatment effects at both scales. To enhance bike-sharing cooperation, we advocate augmenting bicycle availability in areas with high metro usage and low bus coverage, alongside implementing adaptable redistribution strategies.

cs.LG

Updated Determination of Ellis-Jaffe Sum Rules up to $\rm N^{3}LO$ QCD corrections

In this paper, we explore the properties of the Ellis-Jaffe Sum Rule (EJSR) by employing the Principle of Maximum Conformality (PMC) approach to address its perturbative part up to next-to-next-to-next-to-leading order ($\rm N^{3}LO$) QCD contributions. By applying the PMC, we achieve a precise perturbative QCD prediction for the EJSR, free from conventional ambiguities associated with the renormalization scale choices. Considering the presence of the $\alpha_s$ Landau pole near the asymptotic scale, we incorporate the low-energy $\alpha_s$ model based on analytic perturbation theory (APT) to refine the EJSR behavior in the infrared region. By combining the PMC approach with the low-energy APT model, the agreement between theoretical calculations and experimental measurements of EJSR is significantly improved, as evidenced by the reduced discrepancy from $\chi^{2}/d.o. f|_{\rm Conv.}=1.86$ to $\chi^{2}/d.o. f|_{\rm PMC}=1.19$, thereby validating the effectiveness of our approach.

hep-ph

Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Detecting out-of-distribution (OOD) samples is crucial for ensuring the safety of machine learning systems and has shaped the field of OOD detection. Meanwhile, several other problems are closely related to OOD detection, including anomaly detection (AD), novelty detection (ND), open set recognition (OSR), and outlier detection (OD). To unify these problems, a generalized OOD detection framework was proposed, taxonomically categorizing these five problems. However, Vision Language Models (VLMs) such as CLIP have significantly changed the paradigm and blurred the boundaries between these fields, again confusing researchers. In this survey, we first present a generalized OOD detection v2, encapsulating the evolution of these fields in the VLM era. Our framework reveals that, with some field inactivity and integration, the demanding challenges have become OOD detection and AD. Then, we highlight the significant shift in the definition, problem settings, and benchmarks; we thus feature a comprehensive review of the methodology for OOD detection and related tasks to clarify their relationship to OOD detection. Finally, we explore the advancements in the emerging Large Vision Language Model (LVLM) era, such as GPT-4V. We conclude with open challenges and future directions. The resource is available at https://github.com/AtsuMiyai/Awesome-OOD-VLM.

cs.CV

Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models

With the release of large-scale motion datasets with textual annotations, the task of establishing a robust latent space for language and 3D human motion has recently witnessed a surge of interest. Methods have been proposed to convert human motion and texts into features to achieve accurate correspondence between them. Despite these efforts to align language and motion representations, we claim that the temporal element is often overlooked, especially for compound actions, resulting in chronological inaccuracies. To shed light on the temporal alignment in motion-language latent spaces, we propose Chronologically Accurate Retrieval (CAR) to evaluate the chronological understanding of the models. We decompose textual descriptions into events, and prepare negative text samples by shuffling the order of events in compound action descriptions. We then design a simple task for motion-language models to retrieve the more likely text from the ground truth and its chronologically shuffled version. CAR reveals many cases where current motion-language models fail to distinguish the event chronology of human motion, despite their impressive performance in terms of conventional evaluation metrics. To achieve better temporal alignment between text and motion, we further propose to use these texts with shuffled sequence of events as negative samples during training to reinforce the motion-language models. We conduct experiments on text-motion retrieval and text-to-motion generation using the reinforced motion-language models, which demonstrate improved performance over conventional approaches, indicating the necessity to consider temporal elements in motion-language alignment.

cs.CV

Unfolding a Hopf bifurcation in a linear reaction-diffusion equation with strongly localized impurity existence of breathing pulses

This paper presents a general framework to derive the weakly nonlinear stability near a Hopf bifurcation in a special class of multi-scale reaction-diffusion equations. The main focus is on how the linearity and nonlinearity of the fast variables in system influence the emergence of the breathing pulses when the slow variables are linear and the bifurcation parameter is around the Hopf bifurcation point. By applying the matching principle to the fast and slow changing quantities and using the relevant theory of singular perturbation, we obtain explicit expressions for the stationary pulses. Then, the normal form theory and the center manifold theory are applied to give Hopf normal form expressions. Finally, one of these expressions is verified by the numerical simulation.

math.DS

Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches

To build a cross-modal latent space between 3D human motion and language, acquiring large-scale and high-quality human motion data is crucial. However, unlike the abundance of image data, the scarcity of motion data has limited the performance of existing motion-language models. To counter this, we introduce "motion patches", a new representation of motion sequences, and propose using Vision Transformers (ViT) as motion encoders via transfer learning, aiming to extract useful knowledge from the image domain and apply it to the motion domain. These motion patches, created by dividing and sorting skeleton joints based on body parts in motion sequences, are robust to varying skeleton structures, and can be regarded as color image patches in ViT. We find that transfer learning with pre-trained weights of ViT obtained through training with 2D image data can boost the performance of motion analysis, presenting a promising direction for addressing the issue of limited motion data. Our extensive experiments show that the proposed motion patches, used jointly with ViT, achieve state-of-the-art performance in the benchmarks of text-to-motion retrieval, and other novel challenging tasks, such as cross-skeleton recognition, zero-shot motion classification, and human interaction recognition, which are currently impeded by the lack of data.

cs.CV

Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. The code is available at https://github.com/AtsuMiyai/UPD.

cs.CV

Chiral perturbation theory and Bose-Einstein condensation in QCD

We present recent results in three-flavor chiral perturbation theory at finite isospin $\mu_I$ and strangeness $\mu_s$ chemical potentials at zero temperature. The phase diagram to ${\cal O}(p^2)$ in the $\mu_I$--$\mu_S$ plane is mapped out with and without electromagnetic effects. The phase diagram consists of a vacuum phase and three Bose-condensed phases with condensates of $\pi^{\pm}$, $K^{\pm}$, and $K^{0}/\bar{K}^0$, respectively. Including electromagnetic interactions, the Bose-condensed phases become Higgs phases via the Higgs mechanism. The tree-level spectrum for the mesons and gauge bosons is also derived. We calculate the pressure, energy density, isospin density, and speed of sound in the pion-condensed phase to ${\cal O}(p^4)$ for three-flavor $\chi$PT. The results are compared with recent lattice simulations and the agreement is very good for isospin chemical potentials up to approximately 200 MeV. Moreover, by integrating out the $s$-quark, we show that the thermodynamic quantities can be mapped onto their two-flavor counterparts with renormalized parameters. %to ${\cal O}(p^6)$ for two-flavor $\chi$PT in the chiral limit. We also consider the nonrelativistic limit. It is shown that the energy density can be matched onto the classic result by Lee, Huang and Yang (LHY) for a dilute Bose, with an $s$-wave scattering length that includes radiative corrections. The breaking of the $U(1)$ symmetry in the Bose-condensed phases gives rise to a Goldstone bosons, whose dispersion is linear for momenta $p\ll\mu_I$. In this regime, we use Son's prescription to construct an effective theory for the Goldstone field which is valid in this regime. It is shown that its damping rate is of order $p^5$. This result is in agreement with Beliav's for a dilute Bose gas.

hep-ph

Can Pre-trained Networks Detect Familiar Out-of-Distribution Data?

Out-of-distribution (OOD) detection is critical for safety-sensitive machine learning applications and has been extensively studied, yielding a plethora of methods developed in the literature. However, most studies for OOD detection did not use pre-trained models and trained a backbone from scratch. In recent years, transferring knowledge from large pre-trained models to downstream tasks by lightweight tuning has become mainstream for training in-distribution (ID) classifiers. To bridge the gap between the practice of OOD detection and current classifiers, the unique and crucial problem is that the samples whose information networks know often come as OOD input. We consider that such data may significantly affect the performance of large pre-trained networks because the discriminability of these OOD data depends on the pre-training algorithm. Here, we define such OOD data as PT-OOD (Pre-Trained OOD) data. In this paper, we aim to reveal the effect of PT-OOD on the OOD detection performance of pre-trained networks from the perspective of pre-training algorithms. To achieve this, we explore the PT-OOD detection performance of supervised and self-supervised pre-training algorithms with linear-probing tuning, the most common efficient tuning method. Through our experiments and analysis, we find that the low linear separability of PT-OOD in the feature space heavily degrades the PT-OOD detection performance, and self-supervised models are more vulnerable to PT-OOD than supervised pre-trained models, even with state-of-the-art detection methods. To solve this vulnerability, we further propose a unique solution to large-scale pre-trained models: Leveraging powerful instance-by-instance discriminative representations of pre-trained models and detecting OOD in the feature space independent of the ID decision boundaries. The code will be available via https://github.com/AtsuMiyai/PT-OOD.

cs.CV

CPR-Coach: Recognizing Composite Error Actions based on Single-class Training

The fine-grained medical action analysis task has received considerable attention from pattern recognition communities recently, but it faces the problems of data and algorithm shortage. Cardiopulmonary Resuscitation (CPR) is an essential skill in emergency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this paper constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compression and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper thoroughly investigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable Single-class Training & Multi-class Testing problem, we propose a humancognition-inspired framework named ImagineNet to improve the model's multierror recognition performance under restricted supervision. Extensive experiments verify the effectiveness of the framework. We hope this work could advance research toward fine-grained medical action analysis and skill assessment. The CPR-Coach dataset and the code of ImagineNet are publicly available on Github.

cs.CV

Open-Set Domain Adaptation with Visual-Language Foundation Models

Unsupervised domain adaptation (UDA) has proven to be very effective in transferring knowledge obtained from a source domain with labeled data to a target domain with unlabeled data. Owing to the lack of labeled data in the target domain and the possible presence of unknown classes, open-set domain adaptation (ODA) has emerged as a potential solution to identify these classes during the training phase. Although existing ODA approaches aim to solve the distribution shifts between the source and target domains, most methods fine-tuned ImageNet pre-trained models on the source domain with the adaptation on the target domain. Recent visual-language foundation models (VLFM), such as Contrastive Language-Image Pre-Training (CLIP), are robust to many distribution shifts and, therefore, should substantially improve the performance of ODA. In this work, we explore generic ways to adopt CLIP, a popular VLFM, for ODA. We investigate the performance of zero-shot prediction using CLIP, and then propose an entropy optimization strategy to assist the ODA models with the outputs of CLIP. The proposed approach achieves state-of-the-art results on various benchmarks, demonstrating its effectiveness in addressing the ODA problem.

cs.CV

Pion condensation in dense QCD, the dilute Bose gas, and speedy Goldstone bosons

We consider pion condensation in QCD at finite isospin density $\mu_I$ and zero temperature using two-flavor chiral perturbation theory ($\chi$PT). The pressure is calculated to next-to-leading order (NLO) in the low-energy expansion. In the nonrelativistic limit, we recover the classic result by Lee, Huang, and Yang for the energy density of a dilute Bose gas with an $s$-wave scattering length that includes loop corrections from $\chi$PT. In the chiral limit, higher-order calculations are tractable. We calculate the pressure to next-to-next-to-leading order (NNLO) in the low-energy expansion, which is an expansion in powers of $\mu_I^2/(4\pi)^2f^2$, where $f$ is the (bare) pion decay constant. The spontaneous breakdown of the global internal symmetry $U(1)_{I_3}$ gives rise to a massless Goldstone boson or phonon. We discuss the properties of the low-energy effective theory describing this mode. Finally, we compare our results for the pressure and the speed of sound with recent lattice simulations with 2+1 flavors. The agreement is very good for isospin chemical potentials up to 180-200 MeV, depending on the physical quantity.

hep-ph