arXiv Science⌕ Search

arXiv · 2609.40141

LLTA: A Simplicity-Oriented Open-Source WCET Analyser

Abstract

Deriving a safe upper bound on the worst-case execution time (WCET) of a real-time task is essential for hard real-time systems. Many WCET analysers exist, but they are either (i) closed source or (ii) do not provide a WCET for an existing microcontroller. Hence, it is difficult to obtain a trustworthy WCET bound for the binary that is flashed to an available embedded hardware platform and to observe whether the bound is safe. To solve this, we present LLTA, an open-source WCET analyser for Commercial-Off-The-Shelf (COTS) microcontrollers. LLTA is designed based on the following goals: it is simple (G1), runs on real COTS hardware (G2), reports results that are unsound due to programming or device constraints (G3), and is available open-source without licensing complications (G4). LLTA generates WCETs for the ESP32-C6 and the MSP430(FR5994), which are broadly available. This paper systematically lays out our design decisions for LLTA with those goals in mind. To demonstrate its ease of use, we propose scenarios where LLTA can be used for lab exercises of real-time systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nils Hölscher, Kay Heider, Jian-Jia Chen. 2026-09-30. LLTA: A Simplicity-Oriented Open-Source WCET Analyser. https://arxiv.org/abs/2609.40141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based autotuner, improving the efficiency of the tuning process by reducing tuning overheads and circumventing subpar evaluations. By leveraging knowledge from related tasks, we are able to effectively exploit the transfer relationship to access high-performing configurations in fewer samples than traditional techniques that rely upon itera- tive refinement. Our framework can achieve similar performance improvements as state-of-the-art autotuning techniques with up to 61.67% fewer evaluations, averaging 27.85% fewer evaluations across various HPC benchmarks.

cs.PF↗

Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection

In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and the number of KV cache slots at compile time. We define such an environment as a compile-time-static serving substrate and analyze the execution-time cost that tool waiting and re-arrival incur in it. On a single LLM instance, we run synthetic workloads following a measured tool waiting time distribution and compare, on the same inputs, a baseline configuration with settings {1, 2, 4, 8}, 8, and 8 against configurations that change some of them. Because re-arrival times differ across configurations, we build a simulator that replays request processing in time order, select the candidate with the lowest predicted cost among 2,077 configurations, and validate it on new inputs. We identify three mechanisms: discrete batch alignment, KV cache survival, and prefill interference. At a concurrency of 6, absent from the bucket set, tool waiting lowered the padding ratio (0.235 to 0.120) yet increased decode execution time 1.51-fold, so padding alone did not indicate cost. On new inputs at a concurrency of 8, where the baseline reused KV in 9 of 24 re-arrivals, enlarging the maximum batch size alone cut execution cost by 8.25%, and the selected configuration, which also adjusted the bucket set, by 9.72%. Where 17 of 18 re-arrivals were already reused, the effect was 0.59%. Compile-time configurations should thus be selected by diagnosing KV reuse loss and the resulting change in execution.

cs.PF↗

PASCAL: A Progress Divergence-Aware Shared-Cache Model

In modern AI accelerators and GPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query (Q) tiles share the same key and value (K/V) blocks, GEMM, where every tile in a row reads the same slice, and many other operators. We name this pattern shared cyclic scan. Due to a significant amount of data reuse in this pattern, the cache is expected to capture as much data reuse as possible and largely reduce requests sent to the main memory for both performance and energy consumption concerns. However, in reality, because of the intrinsic asynchrony of multi-cores, the actual cache miss rate and DRAM traffic can be much higher compared to ideal cases. In this paper, we propose PASCAL, a shared-cache model for shared cyclic scans. It calibrates finite-run traffic, which reflects progress divergence, on reference configurations and interpolates it along static program structure such as occupancy to predict the miss rate before execution. PASCAL supports software configuration exploration without target traces or counters at scales where cycle-accurate simulation is impractical, while its analysis gives a sharp $2σ-1$ sufficient capacity condition for preserving LRU sharing. Across held-out scans and GEMM on GB10 and Thor, PASCAL reaches 12.04% balanced fill-equivalent miss-rate MAPE. Replacing TileSight's cache component with PASCAL lowers GB10 GEMM latency MAPE from 18.17% to 13.14%.

cs.PF↗