arXiv Science⌕ Search

arXiv · 2610.00528

Best practices in software citation

Abstract

Software is both a foundational tool and a primary output of modern computational research, yet citation practices for software remain inconsistent, incomplete, and rarely machine-actionable. Existing infrastructure designed for paper and data citation does not adequately serve the distinct needs of software citation, leaving a gap that impedes reproducibility, misattributes scholarly credit, and obscures the labor embedded in research pipelines. Drawing on a NASA-funded community workshop held in April 2026, we present an analysis of four interconnected themes: (I)~the cultural barriers to consistent citation practice; (II)~the need for clearer community norms and conventions; (III)~gaps in existing technical infrastructure and workflow; and (IV)~the emerging challenges posed by AI-assisted research. For each theme we identify targeted interventions and assign responsibility across stakeholder groups. We conclude that meaningful progress requires simultaneous action on technical and cultural fronts. Journal editors and publishers represent the single highest-leverage point for accelerating this change, and correct citation must become the path of least resistance within researchers' existing workflows.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Phil R. Van-Lane, Floor S. Broekgaarden, Daniel S. Katz, Bhavesh Patel, Pengyin Shan, Jonathan Starr, Samantha Teplitzky, Peter K. G. Williams, Alice Allen, Lucas M. de Sá, Andrew Fullard, Sandra Gesing, Tom Wagg, Andrea Zonca. 2026-09-30. Best practices in software citation. https://arxiv.org/abs/2610.00528

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Augmenting Italian Cultural Heritage with Virtual and Digital Technologies: the final outcomes of Project CHANGES' Spoke 4

CHANGES (Cultural Heritage Innovation for Next-Gen Sustainable Society) was a project coordinated by the CHANGES Foundation, bringing together complementary disciplines and expertise across the entire cultural heritage lifecycle. This article focuses on the project's research area dedicated to applying virtual technologies to museums and art collections. This research enabled us to experiment with different types of museums and art collections across Italy, designing case studies and best practices that institutions and similar contexts can adapt and reuse. The meta-analysis of the case studies showed that museum types play a decisive role in shaping not only technological choices but also organisational strategies, training models, and pathways to sustainability. This work also allowed us to identify technological and organisational gaps that go beyond single cases and reflect shared challenges across museum contexts. Finally, we documented the success of the engagement, exploitation, and dissemination strategy proposed in the project.

cs.DL↗

Foreign-trained faculty and the collaborative organization of high-impact U.S. science

Internationally mobile scientists are central to national research and innovation systems. We link faculty rosters from the Academic Analytics Research Center to OpenAlex publication records for 2011-2020, yielding more than 12 million faculty-publication observations for 236,394 tenure-system faculty at more than 300 major U.S. universities. Foreign-trained faculty, defined by a terminal degree awarded outside the United States, constitute about 11% of the observed faculty workforce but account for 13-14% of publications and 14-16% of top-1% cited elite output. We find that this elevated representation in elite output is closely related to their organizational embeddedness: when faculty are compared within the same institution, scientific domain, rank, and year, the difference in elite output narrow substantially while overall productivity and collaboration differences persist within those settings. In addition, foreign-trained faculty enter each publication year with larger and broader prior collaboration networks, and collaborator reach is more strongly associated with subsequent elite output. By contrast, although raw topic breadth is greater among foreign-trained faculty, after accounting for prior publication volume, we find that their topic breadth is slightly narrower and more cognitively concentrated. These findings recast international training in the lens of scientific capacity and organizational integration in that internationally accumulated scientific capabilities become embedded in institutions and relationships through which research is organized and produced.

cs.DL↗

Which reported inputs govern molecular docking reproducibility? A benchmark from reporting audit to independent re-execution

Computational docking results can be re-executed only when the molecular system and protocol are specified, yet reporting checklists do not quantify how strongly individual method fields affect the reported score. We combined a span-verified audit of 50 open-access papers with controlled one-factor perturbations, a Vinardo scoring-function check and a Vina cross-docking extension covering 12 targets in 7 protein families, and re-execution of 37 published claims on local and independent cloud infrastructure. Search-related fields were sparsely reported, but perturbing box centre, box size, exhaustiveness or random seed produced median absolute score changes of no more than 0.08 kcal mol$^{-1}$; ligand protonation and identity produced median changes of 0.19 and 0.29 kcal mol$^{-1}$, respectively, whereas receptor structure produced the largest change (1.02 kcal mol$^{-1}$; n = 85; P < 0.001), and 22% of 116 unique reported ligand strings remained unresolved or ambiguous after deterministic name normalization. These measurements yielded empirical field weights for the NewScience Evidence score and identified exact receptor structure, machine-resolvable ligand identity and protonation state as primary reporting elements; 29 of 37 local and 27 of 37 cloud re-executions were within 2.0 kcal mol$^{-1}$ of the reported score, but these descriptive rates do not estimate experimental replication or a calibrated probability of reproduction.

cs.DL↗