arXiv · cs/0702012
Plagiarism Detection in arXiv
Abstract
We describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different research disciplines. The methodology efficiently detects a variety of problematic author behaviors, and heuristics are developed to reduce the number of false positives. The methods are also efficient enough to implement as a real-time submission screen for a collection many times larger.
Explore related subjects
Keep this discovery
Daria Sorokina, Johannes Gehrke, Simeon Warner, Paul Ginsparg. 2007-02-01. Plagiarism Detection in arXiv. https://doi.org/10.1109/icdm.2006.126
Cite the original work for its findings. Save a collection to share your selection of sources.