arXiv Science⌕ Search

arXiv · 2609.32978

Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) Arrays

Abstract

As networks expand and become an ever more critical infrastructure to modern society the need to analyze these networks with the highest regard for privacy is essential to ensure their proper function. Depending on the level of the network layer to be analyzed, sources and destinations can be any combination of physical, logical, or persona/agentic endpoints, which requires the ability to handle diverse data. Invaluable to these analyses are mathematical tools that enable sophisticated mathematical algorithms to be expressed succinctly while achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations of hardware) performance. Associative (token) array mathematics and corresponding libraries is one approach that can meet these requirements. Accelerating these libraries with GPUs enables the analysis of the largest networks. The MIT/IEEE/Amazon Anonymized Network Sensing Graph Challenge provides a venue for highlighting the applicability of accelerated associative arrays for these types of problems. The D4M associative library has been implemented in a number of languages. This work benchmarks a prototype Matlab D4M GPU accelerated implementation of the Anonymized Network Sensing challenge across a wide range of CPU and GPU hardware. Scalable performance is demonstrated within and across CPU cores, CPU nodes, and GPU nodes. Horizontal scaling across multiple nodes was linear. Running on hundreds of GPU nodes simultaneously achieved a sustained processing rate sufficient to potentially analyze a 10 Petabit/s network.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jeremy Kepner, Hayden Jananthan, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Chansup Byun, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Piotr Luszczek, Peter Michaleas, Lauren Milechin, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Alex Pentland. 2026-09-26. Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) Arrays. https://arxiv.org/abs/2609.32978

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cell-Free Massive MIMO Under Mobility: A Fairness-Differentiated Handover Scheme

While cell-free massive MIMO (CF-mMIMO) offers high and uniform network-wide throughput in static networks, its performance in mobile networks is not yet fully addressed. In this paper, we evaluate the throughput performance of urban mobile CF-mMIMO networks under a comprehensive throughput model and show that it suffers from large performance degradation due to the combined effect of channel aging and handover overheads. To restore the uniformly good performance of CF-mMIMO under mobility, we formulate a novel optimization problem to maximize the nett throughput that considers both channel aging and handover cost. We derive a near-optimal solution nearOpt for our transformed and relaxed optimization problem with Newton's method. We then design a heuristic handover algorithm, FairDiff, to differentiate prioritized and optional handovers using a policy threshold based on Jain's fairness index, in order to achieve uniform throughput over the network. Our extensive evaluation of the mobile throughput performance of our handover schemes in realistic urban mobile networks shows that, unlike the existing literature benchmarks that obtain very low throughput under mobility, our FairDiff scheme consistently achieves the near-optimal throughput comparable to nearOpt and highest network-wide throughput with the lowest computational complexity among all considered schemes. We thus for the first time propose a handover scheme that delivers the promise of uniformly good throughput for mobile CF-mMIMO, making it a feasible architecture for practical mobile networks.

cs.NI↗

Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence

Embodied artificial intelligence (AI) couples perception and learned decision making to actions that change the physical world. This coupling distinguishes an embodied agent from a conventional connected controller: the agent maintains task state and uncertainty, reasons about the consequences of actions, and adapts from subsequent observations. Wireless networking becomes relevant when perception, inference, or coordination is distributed, but it should not replace local safety control. This article develops a tutorial perception--communication--action (PCA) architecture that exposes task state, action deadlines, uncertainty, agent intent, and safety envelopes to a 6G orchestration plane. It separates capabilities already addressed by 5G and 5G-Advanced from functions that motivate 6G, including task-state interfaces, semantic freshness, predictive digital twins, and safety-aware coordination across agents. A multi-robot simulation study is retained to illustrate joint sensing, communication, and computation control. The results show where network orchestration improves task utility and where local autonomy remains essential.

cs.NI↗

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.

cs.NI↗