I implement generative AI papers from scratch and document what breaks.
Machine Learning Engineer working on applied ML problems in MedTech, mostly navigating ambiguity and vague ideas. Outside of that I rebuild the models behind modern generative AI, from VAEs and discrete latents to CLIP, GPT-style LLMs, and autoregressive text-to-image, so I understand the mechanism instead of the API. Currently moving into speech and talking avatars.
Learn - Build - Feedback - Iterate
-
speech-to-speech-latency-profiling - How much of a cascade's latency is the cascade itself. Three ASR to LLM to TTS arrangements, one process and one GPU per stage, built to be compared against Moshi, a full-duplex spoken dialogue model that does the whole job in one network. Overlapping the LLM with the TTS is worth 4.4x on time to first audio, 11.9s down to 2.7s. Running the ASR during speech on top of that saves 106ms at the median and gives it back at p95. Profiling CosyVoice2 for this turned up its speech-token LM and flow decoder blocking each other on the Python interpreter lock, worth 40% of the LM's time. Next: Moshi.
-
generative-ai-from-scratch - VAE, VQ-VAE, VQ-VAE-2, CLIP, and DALL·E-1-style autoregressive text-to-image, each implemented by hand and pushed through ablations. Perceptual loss took VQ-VAE-2 from 0.44 to 0.097 LPIPS. The text-to-image write-up tracks a generation bug through two wrong hypotheses before pinning it on data scale.
-
talking_avatar_journey - Learning the building blocks behind audio-driven talking avatars through hands-on implementation. Covers audio features for lip-sync (mel spectrograms, MFCCs, DeepSpeech/wav2vec2 extraction), a from-scratch pix2pix conditional GAN for image-to-image translation, Wav2Lip audio-driven lip-sync inference, and a Spatial Transformer Network (STN) for pose warping with comparative benchmarking against a standard CNN on MNIST. Next step: DINet.
-
attention_to_llm - A GPT-2 style LLM built from tokenizer to DPO. Also holds architecture experiments with plots: pre-norm vs post-norm gradient stress test, RoPE vs absolute position embeddings across context lengths the model never saw, and MoE routing with and without a load balancing loss.
-
microscaling-formats - Benchmark of per-tensor, block-32, and microscaling (MX) INT quantization at 8, 4, and 2 bits across 7 vision-language models on VQAv2 and TextVQA. INT8 is essentially free, per-tensor INT4 collapses every model to near zero accuracy because one scale per tensor rounds ~90% of weights to zero, and block-INT4 beats MXINT4 by 3 to 9%. Blog post
-
upasak - A flexible, mindful to privacy, no-code/low-code framework for fine-tuning large language models, built around Hugging Face Transformers. Streamlit interface, multi-format dataset support, built-in PII sanitization. Tutorial
-
text_recognition - End-to-end OCR pipeline with CRAFT for detection and CRNN for recognition, trained from scratch on ~45,000 images and tracked with CometML.
I am a firm believer of figuring out as we go.
I started my career in 2024 and went deeper into AI to stop treating it as a black box. The further in I got, the more I fell for the fundamental mechanisms and principles behind deep learning architectures.
I'm broadly interested in maths, finance, tech, and ways to stop self-sabotaging. Aiming for generalist status, but my comfort zone has me in a death grip. Don't believe in fluff projects, either you learn something or you solve a problem.
Permanently buried under a backlog of papers to read.

