Skip to content
View Niklesh99's full-sized avatar

Block or report Niklesh99

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Niklesh99/README.md
        ✨        🧠  βˆ‘  βˆ‡  ⚑        ✨
   🧒
  πŸ‘¦πŸ’»  ───▢  ⌨️  ───▢  πŸ–₯️  ───▢  πŸ€–
  "just one more kernel..."   β˜•

Hi, I'm Nikilesh πŸ‘‹

Senior Software Engineer Β· custom operators Β· kernel optimization Β· LLM/VLM bring-up on hardware accelerators

Email LinkedIn Portfolio


🧭 About

I build hardware-aware ML systems, from cycle-level kernel tuning to deploying full models with high efficiency for inference on custom accelerators. 4+ years across DSPs, runtimes and inference servers.

πŸ”§ What I work on

πŸ›°οΈ Radar-SDK porting FFT, CA-CFAR, OS-CFAR and DML target detection on a custom DSP platform, integrated with the SDK framework
βš™οΈ SIMD / VECC kernels Optimized kernels for SensPro DSP workloads, targeting cycle-level performance and hardware utilization
🚚 DMA & memory optimization Single/double-buffered transfers; theoretical vs. achieved cycle analysis
🧩 Custom device operators Native convolution, deformable convolution and arithmetic kernels in C++/Python
πŸ” ONNX custom operators Custom ops in ONNX Runtime (x86) plus a symbolic-mapping bridge between PyTorch and ONNX
🧠 LLM / VLM bring-up Text and multimodal models integrated into vLLM/Inference Servers

🧰 Tech

C/C++ Python PyTorch ONNX vLLM SGLang Docker Kubernetes GCP Bash

Also: SIMD/VECC, DSP acceleration, MLA attention, Decoupled-RoPE, MTP, model conversion, profiling and debugging, AI agentic workflows.

πŸ“ Engineering notes

  • Fusion vs. scheduling overhead: distributed inference is often limited by kernel launch granularity and memory copies, not raw compute.
  • DMA double-buffering: it hides transfer latency only when pipeline depth matches the accelerator's command queue depth. Profile, don't guess.
  • PyTorch ↔ ONNX: numerical precision silently degrades at symbolic-mapping boundaries. Validate at every one.
  • On-device LLMs: memory-constrained first, compute-bound second. Design for quantization from day one.

🀝 Let's talk

Always interested in challenging ML systems work and hardware-accelerated inference. See my portfolio, or reach me by email or on LinkedIn.

Popular repositories Loading

  1. cameraCalibration_And_perspectiveTransformation cameraCalibration_And_perspectiveTransformation Public

    Extracting Intrinsic and Extrinsic Parameters using Zhang Method, extracting cameraMatrix, Rotation and Translational Vectors to calibrate the camera using CheckersBoard Pattern

    Python 3

  2. Cats-vs-Dogs-using-CNN-and-keras- Cats-vs-Dogs-using-CNN-and-keras- Public

    This Project is based on Neural Network to classify between Dogs and Cats.

    Python 1

  3. Hand-Written-Digit-Classification-using-Keras-and-Neural-Network Hand-Written-Digit-Classification-using-Keras-and-Neural-Network Public

    Jupyter Notebook 1

  4. Sentiment-Stock-Analysis-Prediction-using-RandomForestClassifier-and-MultinominalNB Sentiment-Stock-Analysis-Prediction-using-RandomForestClassifier-and-MultinominalNB Public

    Jupyter Notebook 1

  5. IMAGE_STITCHING_using_opencv_python IMAGE_STITCHING_using_opencv_python Public

    Python 1

  6. plotting-data-readme-data-science-intro-000 plotting-data-readme-data-science-intro-000 Public

    Forked from melanieshi0120/plotting-data-readme-data-science-intro-000

    Jupyter Notebook 1