arXiv · 2407.10440
A novel multi-threaded web crawling model
Abstract
This paper proposes a novel model for web crawling suitable for large-scale web data acquisition. This model first divides web data into several sub-data, with each sub-data corresponding to a thread task. In each thread task, web crawling tasks are concurrently executed, and the crawled data are stored in a buffer queue, awaiting further parsing. The parsing process is also divided into several threads. By establishing the model and continuously conducting crawler tests, it is found that this model is significantly optimized compared to single-threaded approaches.
Explore related subjects
Keep this discovery
Weijie. Jiang. 2024-05-09. A novel multi-threaded web crawling model. https://arxiv.org/abs/2407.10440
Cite the original work for its findings. Save a collection to share your selection of sources.