GPU performance and ML systems. MS in Data Science, University of Maryland (2026).
I write and benchmark GPU code, and I study how inference workloads use the memory hierarchy. Before graduate school I spent three years at Oracle building data pipelines and ML-ready feature processing at scale.
- CUDA kernels and roofline analysis (memory-bound vs. compute-bound behavior)
- LLM inference: KV-cache management, batching, quantization
- Open-source work on ROCm support in vLLM (in progress)

