HyQuant proposes keeping critical attention states in high precision to reduce LLM memory costs
A new arXiv preprint describes HyQuant, a hybrid-precision method that compresses most LLM attention data into low-bit formats while preserving selected tokens and local context in higher precision.