Research Adaptive Optimal Rank Selection for KV Cache Compression Soft-thresholded low-rank KV cache compression with custom Triton kernels for faster LLM generation. Hybrid Attention Accelerator for LLMs Fabricated charge-based in-memory computing accelerator for analog and digital Transformer attention. Sparse Attention Acceleration with In-Memory Pruning ReRAM-based sparse attention accelerator that reduces off-chip data movement and recomputes on chip. Token-Adaptive Low-Rank KV Cache Compression PIM-enabled low-rank KV cache compression with token-adaptive rank selection for LLM inference. Accelerator Design for Ultra-Compressed Transformers Neural engine design for accelerating ultra-compressed Transformer blocks from algorithm to fabrication. Runtime-Regularized KV Cache Management Algorithm-hardware framework for end-to-end acceleration of generative models. Reconfigurable Transformer Accelerator Encoder and decoder accelerator with sparsity-aware data mapping and task-based precision control. Hyperdimensional In-Memory Computing for Mass Spectrometry NVM-based in-memory computing system for full-stack mass spectrometry analysis. Coursework GAN for Medical Image Synthesis DCGAN-based Chest X-Ray augmentation for improving pneumonia classification. Reinforcement Learning for Safe Driving Q-learning agent for a Markov decision process in an autonomous driving scenario. Branch Predictor Design Custom TAGE branch predictor implemented in C for computer architecture workloads.