β¨ π§ β β β‘ β¨
π§’
π¦π» ββββΆ β¨οΈ ββββΆ π₯οΈ ββββΆ π€
"just one more kernel..." β
Senior Software Engineer Β· custom operators Β· kernel optimization Β· LLM/VLM bring-up on hardware accelerators
I build hardware-aware ML systems, from cycle-level kernel tuning to deploying full models with high efficiency for inference on custom accelerators. 4+ years across DSPs, runtimes and inference servers.
| π°οΈ Radar-SDK porting | FFT, CA-CFAR, OS-CFAR and DML target detection on a custom DSP platform, integrated with the SDK framework |
| βοΈ SIMD / VECC kernels | Optimized kernels for SensPro DSP workloads, targeting cycle-level performance and hardware utilization |
| π DMA & memory optimization | Single/double-buffered transfers; theoretical vs. achieved cycle analysis |
| π§© Custom device operators | Native convolution, deformable convolution and arithmetic kernels in C++/Python |
| π ONNX custom operators | Custom ops in ONNX Runtime (x86) plus a symbolic-mapping bridge between PyTorch and ONNX |
| π§ LLM / VLM bring-up | Text and multimodal models integrated into vLLM/Inference Servers |
Also: SIMD/VECC, DSP acceleration, MLA attention, Decoupled-RoPE, MTP, model conversion, profiling and debugging, AI agentic workflows.
- Fusion vs. scheduling overhead: distributed inference is often limited by kernel launch granularity and memory copies, not raw compute.
- DMA double-buffering: it hides transfer latency only when pipeline depth matches the accelerator's command queue depth. Profile, don't guess.
- PyTorch β ONNX: numerical precision silently degrades at symbolic-mapping boundaries. Validate at every one.
- On-device LLMs: memory-constrained first, compute-bound second. Design for quantization from day one.
Always interested in challenging ML systems work and hardware-accelerated inference. See my portfolio, or reach me by email or on LinkedIn.