arXiv · 2606.18425
Complexity and Scale in AI-Assisted Workflow Management: A Federated Learning Case Study
Abstract
Federated learning over medical images is a demanding workflow application. Each round fans out across parallel client jobs and converges on an aggregation step that feeds the next round. At scale this yields 101 sub-workflows and 2,679 jobs on GPUs at four sites, which takes an expert months to build, mostly on workflow mechanics rather than science. We ask how far AI assistance can automate such workflows. An LLM agent, grounded in a released plugin of Pegasus-specific skills, first produces a reviewable specification of checkable constraints and acceptance criteria, then generates the executable workflow. A validation loop repairs runtime failures, checks code against those constraints, and regenerates the implementation from the specification alone. We evaluate three LLM agents, report end-to-end runs on the FABRIC testbed, and show how conformance checking against the specification caught three silent errors that failure-driven debugging missed, including one that trained 1,700 jobs on random tensors.
Explore related subjects
Keep this discovery
Komal Thareja, Hamza Safri, Rajiv Mayani, Anirban Mandal, Ewa Deelman. 2026-06-16. Complexity and Scale in AI-Assisted Workflow Management: A Federated Learning Case Study. https://arxiv.org/abs/2606.18425
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.