AI Engineer working on time-series forecasting, NLP and LLM systems.
Currently a Research Assistant at Fraunhofer IEG, where I fine-tune a time-series foundation model (Chronos-2) for electricity price and demand forecasting and run ablation studies across market, weather, calendar and ERA5 forecast covariates. Before my Master's I spent 1.5 years as a software engineer at Tata Consultancy Services.
I like building things end to end — ingestion, training, evaluation, and the part where it actually runs.
- 🔬 M.Sc. Applied Computer Science (Data Science), University of Göttingen
- ⚡ Working on: time-series foundation models, cross-modal feature fusion, LLM evaluation
- 📫 uwaishmohd05@gmail.com
Most of what I'm proudest of lives in other people's repositories.
Apogee — private in-browser AI summarizer
9 merged PRs, including most of a 7-PR series adding llama.cpp as a local inference backend alongside WebLLM, in-browser WASM and Ollama.
- Built the llama-server HTTP client from scratch — SSE stream parsing, runtime context-window detection, bearer-token auth, typed error hierarchy — with 31 unit tests
- Found and fixed a connection leak on stream cancellation (early generator return skipping the catch block, leaving the response body locked)
- Shipped live tokens/sec throughput metering across all four inference backends, computed in the background service worker so the reading survives a popup close mid-stream
TraceRoot — observability layer for AI agents (YC S25, 750+ ⭐)
3 PRs to the TypeScript/Python monorepo — CLA, maintainer review, 18-check CI.
- Fixed timestamp reasoning in the agent system prompt by injecting current UTC date at session start, so time-bounded tool calls resolve correctly
- Fixed LLM cost tracking where per-token rates had drifted from published pricing and an unmatched model alias returned null — added absolute-rate assertions the existing ratio-only tests couldn't catch
FairSample — my own, published to PyPI
pip install fairsample
Resampling library for imbalanced datasets: 14+ techniques and 40+ dataset complexity measures across feature, instance and structural overlap, so you can diagnose why a dataset is hard before you resample it. Docs
| Project | What it is |
|---|---|
| Agents Jailbreaking Agents | 3-agent adversarial framework (Jailbreaker / Victim / Judge) testing multi-turn LLM safety across 14 attack techniques and 520+ prompts. Multi-turn attacks hit 3× the success rate of single-turn. |
| GeoRAG | Hybrid RAG — vector retrieval + Neo4j knowledge graph, with query classification routing between strategies. 11-metric evaluation framework. |
| ZeroDesk | Enterprise RAG support chatbot. FastAPI + Next.js, Dockerized. |
| Universität Kompass | Matches your CV to German university programs using GPT-4 + FAISS semantic search over scraped DAAD data. |
| Meeting Summarization Testbench | Upload a HuggingFace model, get ROUGE / BLEU / BERTScore plus linguistic analysis. Django REST + Plotly. |
Languages · Python · SQL / PL-SQL · JavaScript · TypeScript · Java
ML & DL · PyTorch · TensorFlow · scikit-learn · XGBoost · Transformers · LSTM · Chronos-2 · time-series forecasting
NLP & LLMs · HuggingFace · BERT · BERTopic · spaCy · NLTK · fine-tuning · prompt engineering
RAG & Agents · LangChain · LangGraph · LlamaIndex · RAGAS · FAISS · ChromaDB · Pinecone · Neo4j
Backend & Web · FastAPI · Django · Flask · Streamlit · React · Next.js · Node.js
Data · pandas · NumPy · Plotly · Tableau · Scrapy · PostgreSQL · MySQL · Oracle · MongoDB
Infra · Docker · Git · GitHub Actions · GitLab CI · MLflow · AWS · GCP · Linux
Email · LinkedIn · Portfolio · Xing
Open to ML, AI engineering and data science roles. Based in Germany.

