Data Scientist building end-to-end machine learning systems from data pipeline to deployed API, not just model training.
I build and deploy ML systems that go beyond a notebook real APIs, real statistical rigor, real verification of my own results. My recent work includes cost-sensitive fraud detection, statistical A/B testing frameworks, credit risk modeling, and hybrid retrieval-augmented generation (RAG) systems.
Python | SQL | Machine Learning | Statistical Modeling | XGBoost | Scikit-learn | SHAP | Hypothesis Testing | A/B Testing FastAPI | Docker | LangChain | ChromaDB | BM25 | RAG | Streamlit | Git | GitHub Actions | pytest | Pandas | NumPy
π Fraud Detection System Cost-sensitive XGBoost model (0.867 PR-AUC) deployed live with FastAPI + Docker, SHAP explainability, drift monitoring. Live API β
π A/B Testing & Experimentation Framework 7-stage statistical pipeline β SRM checks, power analysis, Bonferroni-corrected guardrail metrics β on real messy data cleaned via SQL.
π³ Credit Risk Scoring Leakage-safe credit default prediction with SQL-based EDA and stratified cross-validation.
π Hybrid RAG Document Intelligence BM25 + dense vector retrieval fused with Reciprocal Rank Fusion, citation-grounded generation.
π« Reach me: LinkedIn Β· sahithibrunda@gmail.com