arXiv Science⌕ Search

arXiv subjects

Vlad-Petru Nitu

Publications and source records attributed to Vlad-Petru Nitu.

3 recordsLinked to original sources

Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Second, workloads such as serverless functions frequently invoke OS services. At the same time, the OS is a complex codebase spanning many subsystems (e.g., memory management, networking), making it hard to localize the specific code path responsible for a slowdown. Existing profilers expose measurements that require interpretation(e.g., perf and Intel VTune) or can perturb short operations when extensively instrumented (e.g., ftrace). Diagnosing OS bottlenecks can therefore require repeated kernel instrumentation and manual interpretation. We introduce Argus, an agentic LLM-based profiler that produces instrumentation code and autonomously reasons over potential OS-level bottlenecks. Argus integrates two key mechanisms: (i) a calibration methodology that involves collecting a measurement from an idle system and using it as a reference point to discover potential bottlenecks, and (ii) a tree-based data structure that represents the different OS execution paths, improving the agent's bottleneck localization accuracy. Argus aims to identify a specific kernel code path rather than stop at a subsystem-level diagnosis. In two case studies, we employ Argus to autonomously discover bottlenecks present in the memory management subsystem caused by (i) a THP aggressor co-running with other applications, and (ii) applications that incur different types of page faults. Argus produces 19 times fewer incorrect deep-path diagnoses than the strongest evaluated LLM-based baseline, which lacks reference calibration, while preserving low time-to-diagnosis (approximately 31 s)

cs.OS↗

Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation

Physical memory allocation establishes virtual-to-physical mappings on demand. In current systems, each minor page fault traps into the kernel and triggers pipeline flushes, stalls, and a long sequence of allocation steps that can cost tens of thousands of cycles. These overheads are increasingly significant for short-lived workloads such as serverless functions and microservices, where minor faults can account for up to 54% of runtime and up to 40% of system energy. Prior hardware allocation proposals avoid traps and context switches, but either sacrifice useful placement optimizations or rely on fixed-function logic that cannot adapt to new policies or changing hardware conditions. We present Valinor, a hardware-OS cooperative memory allocation substrate that combines software flexibility with hardware-class performance. Valinor introduces a programmable hardware allocation engine that executes compact OS-supplied allocation libraries at close to fixed-hardware speed. It supports diverse policies, including short-lived object allocators, integrity mechanisms, and hardware-telemetry-guided placement. We implement Valinor on a BOOM RISC-V soft core running Linux and in a full-system simulator. On real hardware, Valinor accelerates allocation by 17x, improves end-to-end performance by 16%, and reduces energy consumption by up to 8%. Full-system simulation further evaluates the programmable allocation engine and six allocation libraries, showing that Valinor provides hardware-class performance without sacrificing programmability.

cs.AR↗

Revelator: Rapid Data Fetching via System-Software-Guided Hash-based Speculative Address Translation

Address translation is a major performance bottleneck in modern computing systems. Predicting the physical address (PA) of requested data before address translation completes can hide this latency, but accurate virtual address (VA)-to-PA prediction is difficult because conventional operating systems make VA-to-PA mappings unpredictable. Prior work improves predictability but relies on large pages or VA-to-PA contiguity, or stores speculation metadata in costly hardware structures. We introduce Revelator, a hardware-OS cooperative technique that uses hashing to enable accurate speculative address translation with small system modifications. Revelator employs a tiered hash-based memory allocation policy for both program data and last-level page table entries (PTEs), creating predictable VA-to-PA and VA-to-PTE mappings. After an L2 TLB miss, a lightweight hardware speculation engine uses the OS hash functions to predict these mappings and prefetch the corresponding cache blocks before translation completes, hiding address translation latency and accelerating page table walks (PTWs). Revelator does not rely on large pages or VA-to-PA contiguity and requires only small OS and hardware changes. Across 11 data-intensive workloads, Revelator improves performance by 15.3% on average over the state-of-the-art speculative address translation technique under high memory fragmentation. In virtualized environments, it predicts both guest and host physical addresses, providing a 13.6% average speedup over Nested Paging. In 16-core systems, Revelator achieves 1.40x (1.50x) speedup over Transparent Huge Pages across 30 server workload mixes from Google under medium (high) memory fragmentation. RTL synthesis shows only 0.02% area and 0.03% power overheads on a high-end server-grade CPU. Revelator is freely available at \href{https://github.com/CMU-SAFARI/Virtuoso}{github.com/CMU-SAFARI/Virtuoso}.

cs.AR↗