arXiv ScienceSearch

arXiv · 2407.04199

Top Research Performance Over Three Decades: A Multidimensional Micro-Data Approach

Abstract

In this research, the contributions of a highly productive minority of scientists to the national Polish research output over the past three decades (1992-2021) is explored. In almost all previous research, the approaches to high research productivity are missing the time component. Cross-sectional studies were not complemented by longitudinal studies: Scientists comprising the classes of top performers have not been tracked over time. Three classes of top performers (the upper 1%, 5%, and 10%) are examined, and a surprising temporal stability of productivity patterns is found. The 1/10 and 10/50 rules consistently apply across the three decades: The upper 1% of scientists, on average, account for 10% of the national output, and the upper 10% account for almost 50% of total output, with significant disciplinary variations. The Relative Presence Index (RPI) we constructed shows that men are overrepresented and women underrepresented in all top performers classes. Top performers are studied longitudinally through their detailed publishing histories, with micro-data coming from the raw Scopus dataset. Econometric models identify the three most important predictors that change the odds ratio estimates of membership in the top performance classes: gender, academic age, and research collaboration. The downward trend in fixed effects over successive six-year periods indicates increasing competition in academia. A large population of all internationally visible Polish scientists (N=152,043) with their 587,558 articles is studied.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marek Kwiek, Wojciech Roszka. 2024-10-17. Top Research Performance Over Three Decades: A Multidimensional Micro-Data Approach. https://arxiv.org/abs/2407.04199

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Funding to Findings (FIND): An Open Database of NSF Awards and Research Outputs

Public funding plays a central role in driving scientific discovery. To better understand the link between research inputs and outputs, we introduce FIND (Funding-Impact NSF Database), an open-access dataset that systematically links NSF grant proposals to their downstream research outputs, including publication metadata and abstracts. The primary contribution of this project is the creation of a large-scale, structured dataset that enables transparency, impact evaluation, and metascience research on the returns to public funding. To illustrate the potential of FIND, we present two proof-of-concept NLP applications. First, we analyze whether the language of grant proposals can predict the subsequent citation impact of funded research. Second, we leverage large language models to extract scientific claims from both proposals and resulting publications, allowing us to measure the extent to which funded projects deliver on their stated goals. Together, these applications highlight the utility of FIND for advancing metascience, informing funding policy, and enabling novel AI-driven analyses of the scientific process.

cs.DL

Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions

We introduce document embedding geometry as a quantitative observable of conceptual reorganization and develop a counterfactual ablation framework for measuring how individual concepts influence the organization of scientific knowledge, providing a quantitative framework for detecting scientific revolutions. The observable is defined by the geometric perturbation induced when removing documents associated with a candidate concept from the embedding space before and after its historical emergence. Statistical validation is performed using five historical case studies spanning physics, mathematics, and machine learning: special relativity, Gödel's incompleteness theorems, the Higgs mechanism, deep learning, and the attention mechanism underlying transformer architectures. Across the historical case studies, the framework identifies measurable geometric signatures associated with conceptual reorganization, while the validation studies expose important limitations arising from document assignment and sparse historical data. These results establish embedding geometry as a medium for quantifying conceptual reorganization, providing a new approach for studying how scientific fields restructure over time.

cs.DL

Toward non-textual representation of social anthropology: Modeling cultures as knowledge graphs

The study examines the emerging field of knowledge representation in the context of the semantic web and linked data, with a focus on knowledge produced within the social sciences - particularly in sociocultural anthropology. It starts from the premise that natural language, especially its textualized form, has long been the primary vehicle for producing and communicating anthropological research. Informed by theoretical approaches from information science, the study explores how computational methods may offer alternative modes of structuring and representing anthropological knowledge. It challenges the dominance of text as the sole representational medium and highlights the potential of semantic modeling to open new epistemological pathways. At the same time, it acknowledges the conceptual and methodological challenges involved in such a transition. This approach shifts emphasis away from metrics and programming, foregrounding processes of conceptualization, semantics, meaning, and reasoning as key to engaging with anthropological knowledge in digital environments.

cs.DL