arXiv · 2506.17255
UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
Abstract
Large language models (LLMs) require larger GPU memory size these days, necessitating efficient and extreme weight compression methods. Existing compression methods are either theoretically limited by 1 bit per weight or face severe performance degradation and inefficiency. To deploy LLMs in resource-constrained scenarios, we introduce UltraSketchLLM, compressing LLMs with data sketch. It reduces peak GPU memory footprint with a high compression rate down to 0.5 bit per weight. Combined with hardware-friendly implementation, UltraSketchLLM keeps tolerable performance degradation and extremely low latency overhead with 14.9x speedup compared to naive sketch solution.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sunan Zou, Xueting Sun, Ziyun Zhang, Guojie Luo. 2025-06-08. UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators. https://doi.org/10.1145/3770743.3804314
Cite the original work for its findings. Save a collection to share your selection of sources.