arXiv · 2602.12206
Making the complete OpenAIRE citation graph easily accessible through compact data representation
Abstract
The OpenAIRE graph contains a large citation graph dataset, with over 200 million publications and over 2 billion citations. The current graph is available as a dump with metadata which, when uncompressed, totals $\sim$2.5 TB. This makes it hard to process on conventional computers. To make this network more accessible for the community, we provide a processed OpenAIRE graph which is downscaled to 16 GB RAM, while preserving the full graph structure. Apart from this we offer the processed data in a very simple format, which allows for further straightforward manipulation. We also provide (1) a Python pipeline, which can be used to process the next releases of the OpenAIRE graph, and (2) a larger version of the dataset including more publication fields such as, the title, list of authors.
Explore related subjects
Keep this discovery
Joakim Skarding, Pavel Sanda. 2026-02-12. Making the complete OpenAIRE citation graph easily accessible through compact data representation. https://doi.org/10.5334/johd.520
Cite the original work for its findings. Save a collection to share your selection of sources.