Hybrid Attention Accelerator for LLMs
Fabricated charge-based in-memory computing accelerator for analog and digital Transformer attention.
This project designs and implements a 9T-1C SRAM bitcell with analog in-memory computing capability inside a 16Kb CIM cache. The chip also includes a digital core of roughly 150K gates for attention-related computations in large language models.
The 2.00 mm by 1.60 mm 65nm taped-out die achieves 1.65 TOPS/W SoC energy efficiency and supports operation up to 1.1 GHz. This work appears in ESSERC 2024 and JSSC 2025.