arXiv Science⌕ Search

arXiv · 2610.03457

Cross-Facility LLM Pre-training on HPC: Elastic Aggregation, Data Leasing, and Queue-Aware Placement

Abstract

Academic compute is fragmented: allocations are granted per facility, and facilities differ in accelerators and software stacks, schedule jobs independently, and share neither a network nor a filesystem. We present a system that pools such allocations to pre-train a single language model across three supercomputers on two continents, up to 7,400km apart: Snellius (NVIDIA H100), LUMI and Frontier (both AMD MI250X). It combines (i) DiLoCo-style two-loop training with an elastic, token-weighted Nesterov outer step for which zero, one or many live sites are all normal states; (ii) DARL, a data-leasing protocol whose heartbeat-backed leases guarantee that, within an epoch, no sample is trained twice or lost under crashes, late joins and work stealing; and (iii) queue-aware placement driven by unprivileged sbatch --test-only probes. Training Qwen3-0.6B on C4 for 20,000 optimizer steps, the three-site run reaches a held-out perplexity of 34.7, against 28.2 for a centralised baseline. Per-round overhead (weight exchange and checkpointing) stays near 110 s regardless of the number of local steps H, so its share of wall-clock time falls from 32% at H=100 to 5.7% at H=1,000 and 3.1% at H=2,000. In a 23.8 h three-site run with four site departures, 1.2% of granted data blocks were reclaimed and none was duplicated or lost. In an idealized queue-model projection, queue-aware placement shortens time-to-target by 18-43% compared with waiting for all sites to be allocated. Cross-site pre-training is thus operational rather than competitive: it turns fragmented allocations into one training run at a measured cost.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zarè Palanciyan, Thomas van Osch, Douwe van der Wal, Olivera Kotevska, Tim Kok. 2026-10-02. Cross-Facility LLM Pre-training on HPC: Elastic Aggregation, Data Leasing, and Queue-Aware Placement. https://arxiv.org/abs/2610.03457

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PoCQ: Defending Decentralised Federated Learning Against Model Poisoning and Collusion with Verifiable Evidence

Decentralised federated learning removes the central coordinator but requires participants to establish model update integrity autonomously. Existing frameworks either use inexpensive yet unverifiable proxies, including dataset size and epoch count, or validate updates through retraining, which is computationally costly. This paper introduces Proof of Contribution Quality (PoCQ), a framework in which trust is based on published evidence rather than on private judgement. Updates are committed before disclosure and assessed by small assigned peer committees against the updates each committee already holds. No common validation dataset is required, no validator retrains the evaluated model, and no participant requires a view of all updates in a round. Every vote is published with the values that produced it, so any peer can recompute it and contradict a dishonest validator. Accountability is graded by the strength of that evidence, with permanent removal reserved for provable misconduct and bounded suspensions for statistical anomalies, which honest nodes can produce under non-IID data. Across five model poisoning attacks on PathMNIST, PoCQ improves the average update level detection precision of the best current framework by 49% under non-IID data and by 206% under IID, and improves the precision of malicious node detection by 106% under non-IID data. This is achieved while reducing the time required to validate by half against the fastest existing framework and by 92% against retraining based validation. Colluding bad-mouthing validators are removed by recomputing their published evidence, and free riders who replay or copy an update exactly are identified directly from the ledger.

cs.DC↗

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediation outputs across BGL, HDFS, Thunderbird, and OpenStack. The primary evaluation compares Qwen2.5-14B and Mistral-Small-22B, served through vLLM on one NVIDIA RTX PRO 6000 GPU (96 GB), under zero-shot, few-shot, and retrieval-augmented generation (RAG) prompting across three data-sampling seeds. We report F1, bootstrap confidence intervals, predicted-positive rates, throughput, and memory use, with DeepLog as a held-out classical baseline. Mistral-Small attains a higher macro-averaged F1 than Qwen2.5-14B (0.644 versus 0.560), whereas Qwen provides approximately twice the throughput. A separate log-probability evaluation compares raw decisions with Contextual Calibration (CC) and Batch Calibration (BC): neither correction consistently improves prediction-balance diagnostics across prompting strategies. Supplementary single-run comparisons extend evaluation to 4-bit Llama-3.1-70B via local Ollama and Claude Haiku 4.5 via Anthropic's cloud API; RAG improves F1 on all four datasets for both models. Claude's reported aggregate F1 increases from 0.566 to 0.695, while the estimated API cost rises from \$0.94 to \$2.15 per 1,000 incidents. Local deployment ablations show approximately 41-fold throughput scaling with batching and 20% lower latency with 4-bit quantization on the tested workload. These findings support retrieval as useful context for anomaly decisions, while the supplementary protocols, unvalidated explanation quality, and prediction-balance diagnostics limit broader claims about RCA accuracy and probabilistic calibration.

cs.DC↗

OACM: Optimistic Asynchronous Communication Model for Large-Scale SNN Simulation

Spiking Neural Network (SNN) simulation serves as a crucial tool for understanding brain dynamics and advancing neuromorphic computing, but faces significant scalability challenges in large-scale distributed implementations. The primary bottlenecks arise from frequent synchronization overhead in time-driven simulators and extensive secondary rollbacks in optimistic PDES approaches, compounded by inefficient communication patterns that underutilize network bandwidth. In this paper, we propose OACM, an Optimistic Asynchronous Communication Model that addresses these challenges through three key innovations. First, we design Optimistic Hybrid SNN Simulation that combines the implementation simplicity of time-driven approaches with reduced synchronization frequency through strategic rollback mechanisms within synchronization windows. Second, we implement asynchronous one-sided communication using the UNR library, eliminating handshake latency and achieving complete computation-communication overlap. Third, we develop adaptive message aggregation and routing strategies within a 2D-HyperX virtual topology to optimize bandwidth utilization for small message traffic. Experimental evaluation on a high-performance computing cluster demonstrates that OACM achieves up to 1.4x speedup over the original CORTEX simulator and over 22.6x speedup compared to NEST when simulating a multi-area marmoset brain model at scales of up to 174 compute nodes.

cs.DC↗