arXiv Science⌕ Search

arXiv · 2610.10386

Efficient Heuristics and Machine Learning Approach for Fault Characterization in Distributed Self-Stabilizing Programs

Abstract

Modern large-scale systems rely on distributed protocols to maximize efficiency while preserving the correctness guarantees of single-process execution. However, designing such protocols is non-trivial: adding resources to a system inherently increases its complexity, which in turn introduces faults that must be addressed during design. One such fault class, arising in systems that use distributed shared memory (e.g., replicated databases), is consistency violating fault (cvf), a fault in which a process accessing shared memory reads stale data previously written by another process. Cvfs are inherent to systems that prioritize availability over strict consistency. Prior work has shown that self-stabilizing programs remain correct under high availability, provided such faults stay within a tolerable bound. Characterizing the behavior of self-stabilizing programs in the presence of cvfs can therefore help system designers build systems that are both correct and efficient. In this paper, we study self-stabilizing programs under cvfs to understand when and where such faults occur, enabling runtime measures that improve computational efficiency. Since exhaustive state-space exploration does not scale, it suffers from state-space explosion; our study investigates two complementary approaches: a conflict-based state-space exploration technique for monotonically stabilizing $(Δ + 1)$-Graph Coloring programs, and a machine-learning-based approach for a Maximal Independent Set program. We propose a conflicting-edges-based state-space exploration algorithm alongside a rigorously validated ML model, evaluated on an out-of-distribution dataset through structural consistency checks against known graph invariants.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amit Garu, Duong Nguyen. 2026-10-07. Efficient Heuristics and Machine Learning Approach for Fault Characterization in Distributed Self-Stabilizing Programs. https://arxiv.org/abs/2610.10386

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning

The well-known Asynchronous Verifiable Information Dispersal (AVID) problem lets a sender disperse a message across $N=3F+1$ nodes such that it remains recoverable despite up to $F$ Byzantine failures and unbounded message delays. An optimal AVID scheme requires $3\times$ the message size in communication and storage: up to $F$ nodes may be delayed in responding, and among the remaining $2F+1$ respondents, up to $F$ may be Byzantine. This paper presents eAVID, an elastic AVID scheme that provides two key improvements. First, after dissemination, nodes can prune up to half of their stored information without requiring anyone to reconstruct or recode the message. Second, pruning is enabled by an extremely simple protocol. Nodes collect acknowledgments from one another confirming receipt of their coded pieces, after which each node locally prunes its stored information. eAVID achieves this $2\times$ reduction while making two tradeoffs: the sender generates $2N$ fragments rather than $N$, and pruning may necessitate contacting $2F+1$ nodes for message retrieval rather than $F+1$. eAVID was implemented in DispersedSimplex, a simple-to-understand BFT consensus protocol that erasure-codes its blocks. The implementation demonstrates two features. First, eAVID achieves these improvements using a flat erasure-coding scheme that requires no metadata or bookkeeping at the nodes during reconstruction. Second, it delivers the storage savings without a performance cost: with all nodes responsive, pruning fires on over $99\%$ of blocks and steady-state per-node storage falls by $42$-$46\%$ for committees of $10$ to $22$ nodes. Throughput and latency match the unmodified protocol showing that we can achieve these storage savings without any performance costs.

cs.DC↗

IAPRepair: In-Network Aggregation Enhanced Proactive Repair for Erasure-Coded Storage System

Erasure-coded storage provides fault tolerance with substantially lower storage overhead than full replication, but repairing lost or at-risk blocks requires intensive cross-node data transfer. Existing reactive repair schemes start only after a failure, while proactive schemes can move data before failure but commonly treat migration, reconstruction, and network aggregation as loosely coupled operations. As a result, receiver?side bottlenecks, heterogeneous available bandwidth, and limited programmable-switch state continue to constrain repair paral?lelism. This paper presents IAPRepair, an in-network aggre?gation enhanced proactive repair framework for erasure-coded storage. IAPRepair jointly constructs each repair batch, assigns reconstruction providers and replacement nodes according to normalized transmission loads, schedules migration around the remaining receive capacity, and selectively enables in-network aggregation for reconstruction blocks that would otherwise over?load healthy nodes. The selective design reduces receiver-side traffic while retaining migration parallelism and respecting a configurable switch-resource budget. We implement IAPRepair with a Tofino programmable switch and 16 storage nodes, and evaluate it using both a prototype testbed and large-scale simu?lations. Across coding parameters, block sizes, node populations, and multiple STF-node scenarios, IAPRepair reduces repair time by at least 47.83% compared with the evaluated state-of-the-art methods. The results demonstrate that coordinating proactive repair decisions with in-network processing is an effective way to improve repair efficiency under bandwidth heterogeneity.

cs.DC↗

Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.

cs.DC↗