Runtime-Regularized KV Cache Management
Algorithm-hardware framework for end-to-end acceleration of generative models.
This project designs an algorithm-hardware framework that compresses the KV cache context in generative models by keeping the same number of tokens across inputs, heads, and layers.
The method is paired with a memory management unit for efficient memory access and a scheduling mechanism for batch processing. This work appears in JETCAS 2025.