Accelerator Design for Ultra-Compressed Transformers
Neural engine design for accelerating ultra-compressed Transformer blocks from algorithm to fabrication.
This ongoing project designs and fabricates a neural engine for ultra-compressed Transformer blocks. The compression stack combines tensor decomposition, mixed-precision quantization, and token pruning.
The goal is to connect algorithm-level compression with hardware support that can sustain efficient Transformer inference.