Researchers introduced OptR, a method for ultra-low-bit INT2 quantization of key-value caches in large language models. It optimizes rotations to reduce errors after attention and the output projection step. Existing rotation-based INT2 approaches mainly target cache statistics or proxy errors before full attention readout, which overlooks how quantization noise propagates through attention operations and subsequent projections that determine final model outputs. OptR decomposes errors into key- and value-driven components, learns per-head orthogonal corrections, and reparameterizes keys without altering softmax distributions, delivering better reasoning, coding, and long-context retrieval benchmarks with minimal inference overhead.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.