arXiv Science⌕ Search

arXiv · 2610.05554

RetainZ: Reclaim-Time-Aware Placement for AI Checkpoints on Zoned SSDs

Abstract

Solid State Drives (SSDs) are becoming the dominant medium for performance-critical storage. Zoned Namespace (ZNS) SSDs are getting more and more attractive because sequential writes and explicit zone resets reduce address-mapping, over-provisioning, and internal garbage-collection costs. However, reclaiming space requires resetting an entire zone, so deleting one file does not free its space while other files there must be kept. This erase-before-write constraint becomes a bottleneck for write-intensive workloads, particularly AI checkpointing during model training, where frequent saves of model and optimizer state compete for space with long-retained checkpoints. We observe that retention policies inherently encode when data can be retired, avoiding the need for a learned lifetime predictor. We present RetainZ, a storage backend that translates these policies into explicit reclaim epochs. Its allocator isolates long-retained data and groups objects with nearby reclaim epochs, balancing delayed space reclamation against unused space at zone ends. Evaluated on FEMU with multi-tenant streams and checkpoint manifests up to Pythia-1B, RetainZ reduces space occupied by deleted data and unused zone tails by 41.7% and mean per-input checkpoint p99 latency by 47.5% at 84% target load compared with the allocator of ZenFS, an established lifetime-aware ZNS backend, and lowers host write amplification from 1.165 to 1.003.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Minxing Chu, Ruoxi Yang. 2026-10-04. RetainZ: Reclaim-Time-Aware Placement for AI Checkpoints on Zoned SSDs. https://arxiv.org/abs/2610.05554

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

FDP: The Data Placement Promise of Modern NVMe SSDs

NVMe SSDs are now widely deployed as the storage tier in data centers. As SSDs have evolved over the past decade, the commu- nity has continued to debate the interfaces they expose and how operating systems and storage systems should exploit them. The NVMe Flexible Data Placement (FDP) proposal is the latest point in this design space. FDP introduces an interface based on Reclaim Units that enables explicit data placement to reduce device write amplification without the software engineering costs of sequential- write constraints and host garbage collection. FDP-enabled SSDs are emerging in commercial products and early data center de- ployments. Their compatibility with conventional block I/O allows existing applications to run unchanged, allowing a frictionless adop- tion in industry. This paper presents an experimental evaluation of FDP SSDs to characterize their data placement guarantees over the raw device interface. We then revisit two widely deployed and distinct open- source storage systems, MySQL and RocksDB, and examine whether lifetime-based data separation and distinct write patterns built into their architectures can be mapped onto FDP SSDs without inva- sive changes. Our evaluation shows end-to-end WAF reductions at higher device utilization, along with QoS and throughput improve- ments under synthetic and real-world workloads. These results demonstrate that FDP provides a practical and deployable cross- layer mechanism for data placement with open-source ecosystem support on Linux. They also highlight why FDP SSDs are gaining traction in industry.

cs.OS↗

No More Translation at Runtime: LLM-Empowered Static Binary Translation

While AArch64 CPUs are becoming strong market contenders, their software ecosystem lags behind the mature x86-64 environment, hindering the adoption of the new architectures and impacting user experience. Binary translation bridges this divide by converting binary code from one architecture (e.g., x86-64) to run on another (e.g., AArch64), allowing legacy software to benefit from modern hardware's performance and energy efficiency advantages. Current translation methods are typically either dynamic, which adds significant runtime overhead, or static, which struggles with reliability due to the inherent complexities of binary analysis. This paper introduces a new static, assembly-to-assembly translation paradigm that transforms binary code ahead of execution, generating portable, efficient native-like binaries that run on AArch64 devices without runtime frameworks. Benefiting from recent breakthroughs in large language models (LLMs), we provide a practical and automated translation engine that produces high-quality code with minimal human intervention. To ensure correctness, we introduce a crucial verification step, where we split the assembly code into simplified snippets, enabling efficient and scalable semantic verification. Our evaluation shows that this approach significantly outperforms existing open-source solutions with a large margin, producing binaries with near-native performance. Furthermore, it shows substantial improvements over the leading industrial translator, ExaGear, illuminating a promising new direction for cross-architecture binary translation research.

cs.OS↗

ARMS: Adaptive and Robust Memory Tiering System

Memory tiering systems seek cost-effective memory scaling by adding multiple tiers of memory. For maximum performance, frequently accessed (hot) data must be placed close to the host in faster tiers and infrequently accessed (cold) data can be placed in farther slower memory tiers. Existing tiering solutions such as HeMem, Memtis, and TPP use rigid policies with pre-configured thresholds to make data placement and migration decisions. These thresholds make the systems brittle - they fail to perform well in all scenarios. Our analysis of existing systems revealed that incorrect tiering parameters lead to: inaccurate hot page identification, delayed response to hot set changes, and wasteful migrations. Based on this study, we designed ARMS that replaces sensitive parameters with robust policies and mechanisms. We develop a novel hot/cold page identification mechanism that uses relative scoring rather than threshold comparison, a hot set change detector to adapt to workload distribution changes, an adaptive migration policy based on cost/benefit analysis, and a bandwidth-aware batched migration scheduler. Combined, these approaches provide an out-of-the-box performance that matches the best tuned performance of prior systems, while being 1.22-1.85x better than prior systems without tuning.

cs.OS↗