arXiv Science⌕ Search

arXiv · 2609.33315

AgentLoop: Runtime Control of Slot-closed Execution Loops for Tool-augmented LLM Agents

Abstract

Tool-augmented large language model (LLM) agents are becoming an important execution unit in service computing, but existing agent loops still lack explicit runtime signals for assessing task completion. The challenge lies in the fact that an agent may continue reasoning or invoking services even after the runtime context has stopped changing, while evidence already collected remains unsynthesized into a complete answer, which leads to inefficiency in resource usage. To address these challenges, this paper presents AgentLoop, which provides runtime control of slot-closed execution loops for tool-augmented agents. Slot closure means that the information slots required by a request have been covered by sufficient runtime evidence, and that unresolved slots are explicitly identified before the loop stops. AgentLoop converts open-ended agent iteration into state-driven execution control: it maintains a compact runtime state, uses model-assisted structured verification to check answer completeness and missing evidence, and applies bounded stability and low-gain signals over neighboring LLM/tool rounds before selecting one of three actions: Continue Invocation, Answer Synthesis, or Terminate Iteration. Experiments show that AgentLoop reduces redundant execution and context growth, with total token cost reduced by up to 88.44% and average service invocations reduced by up to 76.85% against baselines. The ablation study further shows that the slot-centered control path plays a central role, since disabling it increases execution depth and substantially reduces accuracy. Overall, the results suggest that efficient tool-augmented agents can benefit from explicit runtime signals for deciding when further LLM/tool iterations no longer add useful context or supported evidence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wanyi Zheng, Minxian Xu, Kan Hu, Kejiang Ye, Chengzhong Xu. 2026-09-27. AgentLoop: Runtime Control of Slot-closed Execution Loops for Tool-augmented LLM Agents. https://arxiv.org/abs/2609.33315

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment

This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays. The server aims to recover Kc linear combinations of K intermediate outcomes, where each intermediate outcome is a separable function of one dataset. We consider a more general setting with arbitrary heterogeneous data assignment across users, where ''arbitrary'' means that the data assignment is given in advance (which can be in any form) and ''heterogeneous'' means that the users may hold different numbers of datasets. Under this assignment, each user computes the intermediate outcomes of its assigned datasets and sends masked messages to its associated relay. The relays subsequently process and forward the received messages to the server. We impose two security constraints: (i) security against server, requiring the server to learn only the desired task function without gaining any additional information about users' inputs; and (ii) security against relays, ensuring each relay learns nothing about users' inputs. Moreover, the server or any relay may collude with a subset of users. For Kc=1, the underlying computation reduces to distributed gradient coding. We propose a secure scheme tolerating user dropouts and user collusion, achieving the optimal two-layer communication rates in one regime and order-optimal communication rates within a factor of 2 in the other regime. For Kc>1, we extend the proposed construction to multi-dimensional linearly separable tasks under the no-dropout setting.

cs.DC↗

ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching

Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.

cs.DC↗