arXiv ScienceSearch

arXiv subjects

Yuzhuo Wang

Publications and source records attributed to Yuzhuo Wang.

At least 19 recordsLinked to original sources

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

Measuring the novelty of scientific papers is a central concern in research evaluation and scientometrics. From a recombination perspective, prior studies have largely focused on the co-occurrence of knowledge units to assess the novelty of scientific papers. However, these studies often overlook other relationships between knowledge units. This narrow view may result in inaccurate or incomplete evaluations of novelty for scientific papers. To fill this gap, this study introduces a comprehensive novelty measurement that incorporates three types of relationships between knowledge units: network, semantic, and hierarchical. These relationships are used to quantify the latent distances among knowledge units. Using a dataset of 142,036 articles published in PLoS ONE and a validation dataset from the H1 Connect platform, our results demonstrate that (1) each relationship type captures distinct latent distances between MeSH terms; (2) compared to the widely used indicators proposed by Uzzi et al. (2013), our measures show stronger alignment with peer judgements; and (3) combining all three distance metrics yields more effective identification of novel papers than using any single perspective alone.

cs.DL

Revealing the Technology Development of Natural Language Processing: A Scientific Entity-Centric Perspective

Most studies on technology development have been conducted from a thematic perspective, but the topics are coarse-grained and insufficient to accurately represent technology. The development of automatic entity recognition techniques makes it possible to extract technology-related entities on a large scale. Thus, we perform a more accurate analysis of technology development from an entity-centric perspective. To begin with, we extract technology-related entities such as methods, datasets, metrics, and tools in articles on Natural Language Processing (NLP), and we apply a semi-automatic approach to normalize the entities. Subsequently, we calculate the z-scores of entities based on their co-occurrence networks to measure their impact. We then analyze the development trends of new technologies in the NLP domain since the beginning of the 21st century. The findings of this paper include three aspects: Firstly, the continued increase in the average number of entities per paper implies a growing burden on researchers to acquire relevant technical background knowledge. However, the emergence of pre-trained language models has injected new vitality into the technological innovation of the NLP domain. Secondly, Methods dominate among the 179 high-impact entities. An analysis of the z-score trend about the top 10 entities reveals that pre-trained language models, exemplified by BERT and Transformer, have become mainstream in recent years. Unlike the trend of the other eight method entities, the impact of Wikipedia dataset and BLEU metric has continued to rise in the long term. Thirdly, in recent years, there has been a remarkable surge in popularity for new high-impact technologies than ever before, and their acceptance by researchers has accelerated at an unprecedented speed. Our study provides a new perspective on analyzing technology development in a specific domain.

cs.CL

Exploring Motivations for Algorithm Mention in the Domain of Natural Language Processing: A Deep Learning Approach

With the rise of data-intensive science, algorithms have become central to scientific research. In academic papers, algorithms are mentioned for different purposes, such as describing, using, comparing, or improving methods for specific research tasks. Identifying these purposes can reveal relationships among algorithms and help assess their roles and value. Taking natural language processing (NLP) as an example, this study proposes a sentence-level framework for identifying, analyzing, and tracing the evolution of motivations for mentioning algorithms. We first identify algorithm entities and algorithm-related sentences from full-text papers through manual annotation and machine learning. We then classify mention motivations using pretrained models and data augmentation, and analyze their distribution and temporal evolution. The results show that deep learning models trained with augmented data outperform traditional machine learning models in motivation classification. In NLP papers, more than half of algorithm-related sentences express direct use, whereas improvement is the least frequent motivation. The diversity of motivations has increased over time. For specific algorithm categories, grammar-based algorithms are more often mentioned for description, while machine learning algorithms are more often mentioned for use. Over time, use motivations have gradually replaced description motivations across different algorithms, and the number of motivation types associated with individual algorithms has declined significantly. This study reveals how authors mention algorithm entities in academic writing and provides a basis for future research on algorithm relationship identification and algorithm impact evaluation.

cs.CL

Unveiling Novelty Evolution in the field of Library and Information Science in China

This study analyzes the novelty distribution of scholarly papers in the field of Library and Information Science (LIS) in China, with a focus on differences across journals, research topics, and time periods. Articles published in Chinese LIS journals indexed by the Chinese Social Sciences Citation Index (CSSCI) from 2000 to 2022 were collected as the research sample. BERTopic was applied to paper abstracts to identify research topics, and novelty scores were calculated based on the combinatorial innovation theory of reference pairs cited by focal papers. The study then examined the novelty of papers under different topics and further analyzed author collaboration patterns to explain how collaboration may be associated with paper novelty. The results show that archival research topics generally have lower novelty, whereas topics related to journal evaluation and patent technology display higher novelty in Chinese LIS research. Overall, the novelty of papers in this field has gradually increased over time. Papers with different topics and novelty levels also show distinct collaboration patterns: low-novelty topics are more often associated with solo authorship, while high-novelty topics tend to involve a higher proportion of inter-institutional collaboration. This study reveals the topic-level characteristics and temporal trends of novelty in Chinese LIS research and provides a new perspective for understanding how research topics and collaboration patterns influence scholarly innovation.

cs.DL

Do more heads imply better performance? An empirical study of team thought leaders' impact on scientific team performance

Thought leadership plays a crucial role in boosting team performance; thus, teams with more thought leaders may perform better. However, the impact of the number of thought leaders on team performance in a scientific context remains understudied. In this study, we consider the authors of a publication as a scientific team and define authors responsible for conceptual tasks, such as conceived and designed the experiments in the PLOS contribution statement classification system, as thought leaders. Leveraging more than 140,000 papers from PLOS journals, we examine the relationship between the number of thought leaders and two aspects of team performance, namely team impact and team disruptiveness, from both correlational and causal perspectives. The results show that (1) an inverted U-shaped relationship exists between the number of thought leaders and team impact, and (2) teams with more thought leaders tend to produce less disruptive ideas. We also explore how international collaboration, team size, and gender diversity interact with the number of thought leaders in shaping team performance, and find that (3) international collaboration improves team impact but lowers the disruptiveness of team outputs. This study advances scholarly understanding of thought leadership in scientific teams and provides valuable insights for policymakers and team managers.

cs.CY

Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers

Algorithms have become central to scientific research in the era of artificial intelligence (AI). Although algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation and pay limited attention to the collective influence formed through their interconnections. This study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigates algorithm influence from a network perspective. Using deep learning models, we extract algorithm entities and build overall, cumulative, and annual co-occurrence networks. We analyze their structural characteristics and apply multiple centrality measures to assess the group influence of algorithms across the whole field and over time. The results show that algorithm networks display typical features of complex networks, with increasingly dense connections developing over approximately two decades. Classic, high-performing algorithms and those located at the intersections of different research periods tend to have high popularity, control, centrality, and balanced influence. When the influence of an algorithm declines, it usually loses its core network position first, followed by weaker associations with other algorithms. This study is the first large-scale analysis of algorithm co-occurrence networks. Covering more than four decades of academic publications, it provides a temporal and structural view of algorithm influence and offers a foundation for future research on networks linking algorithms, scholars, and tasks.

cs.AI

Transit Noise in Spin Squeezing Experiments with Coated Rubidium Vapor Cell

Spin squeezing can suppress quantum projection noise via interparticle entanglement, therefore enabling measurement sensitivities beyond the standard quantum limit. In practice, however, the Gaussian and finite intensity profiles of the optical probe beam induce spatially inhomogeneous atom-light interactions. As polarized atoms move within a vapor cell, they experience position-dependent optical intensities, generating transit noise that limits spin squeezing performance. Here, we investigate the transit noise in a coated rubidium vapor cell through combined theoretical analysis and experimental measurements. By varying the probe beam diameter, we quantify the dependence of transit noise on beam size and atomic Larmor frequency. Our results show that, for a vapor cell with fixed dimensions, the transit noise increases as the probe beam spot area decreases. Moreover, when the Larmor frequency is below the characteristic linewidth of the transit noise, the noise contribution becomes larger. We further calculated and measured spin squeezing for different beam sizes and found an experimental difference of $2.7 \pm 0.2$ dB between 2~mm and 0.6~mm, similar to the theoretical prediction of $3.0 \pm 0.3$ dB. Theoretical analysis under conditions of stronger squeezing shows that transit noise becomes an even more critical limiting factor. These results provide practical guidance for optimizing probe beam parameters and suppressing transit noise in spin squeezing experiments.

quant-ph

Beyond Single-Dimension Novelty: How Combinations of Theory, Method, and Results-based Novelty Shape Scientific Impact

Scientific novelty drives advances at the research frontier, yet it is also associated with heightened uncertainty and potential resistance from incumbent paradigms, leading to complex patterns of scientific impact. Prior studies have primarily ex-amined the relationship between a single dimension of novelty -- such as theoreti-cal, methodological, or results-based novelty -- and scientific impact. However, because scientific novelty is inherently multidimensional, focusing on isolated dimensions may obscure how different types of novelty jointly shape impact. Consequently, we know little about how combinations of novelty types influence scientific impact. To this end, we draw on a dataset of 15,322 articles published in Nature Communications. Using the DeepSeek-V3 model, we classify articles into three novelty dimensions based on the content of their Introduction sections: theoretical novelty, methodological novelty, and results-based novelty. These dimensions may coexist within the same article, forming distinct novelty configura-tions. Scientific impact is measured using five-year citation counts and indicators of whether an article belongs to the top 1% or top 10% highly cited papers. Descriptive results indicate that results-based novelty alone and the simultaneous presence of all three novelty types are the dominant configurations in the sample. Regression results further show that articles with results-based novelty only re-ceive significantly more citations and are more likely to rank among the top 1% and top 10% highly cited papers than articles exhibiting all three novelty types. These findings advance our understanding of how multidimensional novelty configurations shape knowledge diffusion.

cs.DL

NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing pressure on human reviewers. While large language models (LLMs), including those fine-tuned on peer review data, have shown promise in generating review comments, the absence of a dedicated benchmark has limited systematic evaluation of their ability to assess research novelty. To address this gap, we introduce NovBench, the first large-scale benchmark designed to evaluate LLMs' capability to generate novelty evaluations in support of human peer review. NovBench comprises 1,684 paper-review pairs from a leading NLP conference, including novelty descriptions extracted from paper introductions and corresponding expert-written novelty evaluations. We focus on both sources because the introduction provides a standardized and explicit articulation of novelty claims, while expert-written novelty evaluations constitute one of the current gold standards of human judgment. Furthermore, we propose a four-dimensional evaluation framework (including Relevance, Correctness, Coverage, and Clarity) to assess the quality of LLM-generated novelty evaluations. Extensive experiments on both general and specialized LLMs under different prompting strategies reveal that current models exhibit limited understanding of scientific novelty, and that fine--tuned models often suffer from instruction-following deficiencies. These findings underscore the need for targeted fine-tuning strategies that jointly improve novelty comprehension and instruction adherence.

cs.CL

The Effect of Gender Diversity on Scientific Team Impact: A Team Roles Perspective

The influence of gender diversity on the success of scientific teams is of great interest to academia. However, prior findings remain inconsistent, and most studies operationalize diversity in aggregate terms, overlooking internal role differentiation. This limitation obscures a more nuanced understanding of how gender diversity shapes team impact. In particular, the effect of gender diversity across different team roles remains poorly understood. To this end, we define a scientific team as all coauthors of a paper and measure team impact through five-year citation counts. Using author contribution statements, we classified members into leadership and support roles. Drawing on more than 130,000 papers from PLOS journals, most of which are in biomedical-related disciplines, we employed multivariable regression to examine the association between gender diversity in these roles and team impact. Furthermore, we apply a threshold regression model to investigate how team size moderates this relationship. The results show that (1) the relationship between gender diversity and team impact follows an inverted U-shape for both leadership and support groups; (2) teams with an all-female leadership group and an all-male support group achieve higher impact than other team types. Interestingly, (3) the effect of leadership-group gender diversity is significantly negative for small teams but becomes positive and statistically insignificant in large teams. In contrast, the estimates for support-group gender diversity remain significant and positive, regardless of team size.

cs.CL

A Liquid-Nitrogen-Cooled Ca+ Ion Optical Clock with a Systematic Uncertainty of 4.4E-19

We report a single-ion optical clock based on the 4S_1/2-3D_5/2 transition of the 40Ca+ ion, operated in a liquid nitrogen cryogenic environment,achieving a total systematic uncertainty of 4.4E-19. We employ a refined temperature evaluation scheme to reduce the frequency uncertainty due to blackbody radiation (BBR), and the 3D sideband cooling has been implemented to minimize the second-order Doppler shift. We have precisely determined the average Zeeman coefficient of the 40Ca+ clock transition to be 14.345(40) Hz/mT^2, thereby significantly reducing the quadratic Zeeman shift uncertainty. Moreover, the cryogenic environment enables the lowest reported heating rate due to ambient electric field noise in trapped-ion optical clocks.

physics.atom-ph

Time delay of mean field interaction in thermal Rydberg atomic gas

Mean field theory is commonly employed to study nonequilibrium dynamics in hot Rydberg atomic ensembles, but the fundamental mechanism behind the generation of the mean-field interactions remains poorly understood. In this work, we experimentally observe a time-delay effect in the buildup of mean-field interaction, which reveals the key role of collision ionization. We analyze the relevant collision channels and propose a microscopic mechanism that quantitatively explains the hysteresis window observed in optical bistability. Then, using square-wave modulation spectroscopy (SMS) to monitor the growth of the mean-field interaction, we experimentally demonstrate a delay in its dynamical buildup following the initial Rydberg excitation. Finally, we demonstrate how this delay effect may help understand the recently observed self-sustained oscillations in thermal Rydberg gases. Our findings provide compelling evidence for the contribution of ionization processes in the nonequilibrium dynamics of thermal Rydberg gas, a system of growing interest for quantum sensing and quantum information science.

physics.atom-ph

Uncertainty Evaluation of the Caesium Fountain Primary Frequency Standard NIM6

A new caesium (Cs) fountain clock NIM6 has been developed at the National Institute of Metrology (NIM) in China, for which a comprehensive uncertainty evaluation is presented. A three-dimensional magneto-optical trap (3D MOT) loading optical molasses is employed to obtain more cold atoms rapidly and efficiently with a tunable, uniform density distribution. A heat pipe surrounding the flight tube maintains a consistent and stable temperature within the interrogation region. Additionally, a Ramsey cavity with four azimuthally distribution feeds is utilized to mitigate distributed cavity phase shifts. The Cs fountain clock NIM6 achieves a short-term stability of 1.0x10-13 {\tau}-1/2 at high atomic density, and a typical overall fractional type-B uncertainty is estimated to be 2.3x10-16. Comparisons of frequency between the Cs fountain NIM6 and other Cs fountain Primary Frequency Standards (PFSs) through Coordinated Universal Time (UTC) have demonstrated an agreement within the stated uncertainties.

quant-ph

Complex-valued 3D atomic spectroscopy with Gaussian-assisted inline holography

When a laser-cooled atomic sample is optically excited, the envelope of coherent forward scattering can often be decomposed into a few complex Gaussian profiles. The convenience of Gaussian propagation helps addressing key challenges in digital holography. In this work, we develop a Gaussian-decomposition-assisted approach to inline holography, for single-shot, simultaneous measurements of absorption and phase-shift profiles of small atomic samples sparsely distributed in 3D. The samples' axial positions are resolved with micrometer resolution, and their spectroscopy are extracted from complex-valued images recorded at various probe frequencies. The phase-angle readout is not only robust against transition saturation, but also insensitive to atom-number and optical-pumping-induced interaction-strength fluctuations. Benefiting from such features, we achieve hundred-kHz-level single-shot resolution to the transition frequency of a $^{87}$Rb D2 line, with merely hundreds of atoms. We further demonstrate single-shot 3D field sensing by measuring local light shifts to the atomic array with micrometer spatial resolution.

physics.atom-ph

How do software citation formats evolve over time? A longitudinal analysis of R programming language packages

Under the data-driven research paradigm, research software has come to play crucial roles in nearly every stage of scientific inquiry. Scholars are advocating for the formal citation of software in academic publications, treating it on par with traditional research outputs. However, software is hardly consistently cited: one software entity can be cited as different objects, and the citations can change over time. These issues, however, are largely overlooked in existing empirical research on software citation. To fill the above gaps, the present study compares and analyzes a longitudinal dataset of citation formats of all R packages collected in 2021 and 2022, in order to understand the citation formats of R-language packages, important members in the open-source software family, and how the citations evolve over time. In particular, we investigate the different document types underlying the citations and what metadata elements in the citation formats changed over time. Furthermore, we offer an in-depth analysis of the disciplinarity of journal articles cited as software (software papers). By undertaking this research, we aim to contribute to a better understanding of the complexities associated with software citation, shedding light on future software citation policies and infrastructure.

cs.DL

Absolute frequency measurements with a robust, transportable ^{40}Ca^{+} optical clock

We constructed a transportable 40Ca+ optical clock (with an estimated minimum systematic shift uncertainty of 1.3*10^(-17) and a stability of 5*10^(-15)/sqrt{tau} ) that can operate outside the laboratory. We transported it from the Innovation Academy for Precision Measurement Science and Technology, Chinese Academy of Sciences, Wuhan to the National Institute of Metrology, Beijing. The absolute frequency of the 729 nm clock transition was measured for up to 35 days by tracing its frequency to the second of International System of Units. Some improvements were implemented in the measurement process, such as the increased effective up-time of 91.3 % of the 40Ca+ optical clock over a 35-day-period, the reduced statistical uncertainty of the comparison between the optical clock and hydrogen maser, and the use of longer measurement times to reduce the uncertainty of the frequency traceability link. The absolute frequency measurement of the 40Ca+ optical clock yielded a value of 411042129776400.26 (13) Hz with an uncertainty of 3.2*10^(-16), which is reduced by a factor of 1.7 compared with our previous results. As a result of the increase in the operating rate of the optical clock, the accuracy of 35 days of absolute frequency measurement can be comparable to the best results of different institutions in the world based on different optical frequency measurements.

physics.atom-ph

Automatic Recognition and Classification of Future Work Sentences from Academic Articles in a Specific Domain

Future work sentences (FWS) are the particular sentences in academic papers that contain the author's description of their proposed follow-up research direction. This paper presents methods to automatically extract FWS from academic papers and classify them according to the different future directions embodied in the paper's content. FWS recognition methods will enable subsequent researchers to locate future work sentences more accurately and quickly and reduce the time and cost of acquiring the corpus. The current work on automatic identification of future work sentences is relatively small, and the existing research cannot accurately identify FWS from academic papers, and thus cannot conduct data mining on a large scale. Furthermore, there are many aspects to the content of future work, and the subdivision of the content is conducive to the analysis of specific development directions. In this paper, Nature Language Processing (NLP) is used as a case study, and FWS are extracted from academic papers and classified into different types. We manually build an annotated corpus with six different types of FWS. Then, automatic recognition and classification of FWS are implemented using machine learning models, and the performance of these models is compared based on the evaluation metrics. The results show that the Bernoulli Bayesian model has the best performance in the automatic recognition task, with the Macro F1 reaching 90.73%, and the SCIBERT model has the best performance in the automatic classification task, with the weighted average F1 reaching 72.63%. Finally, we extract keywords from FWS and gain a deep understanding of the key content described in FWS, and we also demonstrate that content determination in FWS will be reflected in the subsequent research work by measuring the similarity between future work sentences and the abstracts.

cs.CL

Spectroscopic localization of atomic sample plane for precise digital holography

In digital holography, the coherent scattered light fields can be reconstructed volumetrically. By refocusing the fields to the sample planes, absorption and phase-shift profiles of sparsely distributed samples can be simultaneously inferred in 3D. This holographic advantage is highly useful for spectroscopic imaging of cold atomic samples. However, unlike (e.g., biological samples or solid particles), the quasi-thermal atomic gases under laser-cooling are typically featureless without sharp boundaries, invalidating a class of standard numerical refocusing methods. Here, we extend the refocusing protocol based on the Gouy phase anomaly for small phase objects to free atomic samples. With a prior knowledge on a coherent spectral phase angle relation for cold atoms that is robust against probe condition variations, an ``out-of-phase'' response of the atomic sample can be reliably identified, which flips the sign during the numeric back-propagation across the sample plane to serve as the refocus criterion. Experimentally, we determine the sample plane of a laser-cooled $^{39}$K gas released from a microscopic dipole trap, with a $\delta z\approx 1~{\rm \mu m}$$\ll 2\lambda_p/{\rm NA}^2$ axial resolution, with a NA=0.3 holographic microscope at $\lambda_p=770~$nm probe wavelength.

physics.optics