arXiv Science⌕ Search

arXiv · 2609.34048

Schedule Repair for DAG Workflows under Link Disruptions

Abstract

Schedules for directed acyclic graph (DAG) workflows in networked IoT systems are typically computed assuming a static or generally stable network. In contested and adversarial environments, this assumption is not valid. Links degrade and fail due to mobility, interference, and jamming. We study schedule repair: when a link disruption invalidates part of a schedule, how much of it should be rescheduled? We introduce a spectrum of repair policies that vary in repair scope, how much of the pending schedule each may move: wait out the disruption, reroute data around it, reschedule only the affected tasks locally, or reschedule all pending tasks globally. We evaluate each against an oracle and charge every repair a decision latency proportional to the extent to which it moves. Across 100 workload instances spanning synthetic task graphs, RIoTBench pipelines, and WfCommons scientific workflows, each run at five communication-to-computation ratios (CCRs) and disrupted by processes with deliberately different correlation structure, we find that no single scope wins: rerouting nearly erases isolated failures that cost waiting 30%, global repair comes within 4% of the oracle under jamming blackouts, waiting is favored under memoryless link flapping for larger and communication-heavy workloads (the scheduling analog of route-flap damping), self-healing mobility outages reward patience over reaction, and accounting for repair latency erodes large scopes first. We conclude that the scope of the repair should be adapted to the disruption process and the repair cost, rather than fixed by the scheduler.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohammadali Khodabandehlou, Jared Coleman, Bhaskar Krishnamachari, Kevin Chan. 2026-09-28. Schedule Repair for DAG Workflows under Link Disruptions. https://arxiv.org/abs/2609.34048

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment

This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays. The server aims to recover Kc linear combinations of K intermediate outcomes, where each intermediate outcome is a separable function of one dataset. We consider a more general setting with arbitrary heterogeneous data assignment across users, where ''arbitrary'' means that the data assignment is given in advance (which can be in any form) and ''heterogeneous'' means that the users may hold different numbers of datasets. Under this assignment, each user computes the intermediate outcomes of its assigned datasets and sends masked messages to its associated relay. The relays subsequently process and forward the received messages to the server. We impose two security constraints: (i) security against server, requiring the server to learn only the desired task function without gaining any additional information about users' inputs; and (ii) security against relays, ensuring each relay learns nothing about users' inputs. Moreover, the server or any relay may collude with a subset of users. For Kc=1, the underlying computation reduces to distributed gradient coding. We propose a secure scheme tolerating user dropouts and user collusion, achieving the optimal two-layer communication rates in one regime and order-optimal communication rates within a factor of 2 in the other regime. For Kc>1, we extend the proposed construction to multi-dimensional linearly separable tasks under the no-dropout setting.

cs.DC↗

ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching

Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.

cs.DC↗