Runtime-Regularized KV Cache Management

Algorithm-hardware framework for end-to-end acceleration of generative models.

This project designs an algorithm-hardware framework that compresses the KV cache context in generative models by keeping the same number of tokens across inputs, heads, and layers.

The method is paired with a memory management unit for efficient memory access and a scheduling mechanism for batch processing. This work appears in JETCAS 2025.