FraudLens is an end-to-end fraud detection system that combines machine learning, explainable AI, and retrieval-augmented generation (RAG) to assist human reviewers in making informed fraud decisions.
This system provides an intelligent fraud review workflow that:
- Detects fraudulent transactions using XGBoost with calibrated probabilities
- Explains model decisions using SHAP values
- Retrieves relevant context (similar cases, policies) using advanced RAG techniques
- Enables human reviewers to override model decisions and provide feedback
- Continuously improves through active learning and model retraining
- XGBoost-based fraud detection with 86% precision and 86% recall
- Risk bucket classification (LOW/MEDIUM/HIGH) for automated decision routing
- Optimal threshold optimization balancing precision and recall
- Probability calibration using isotonic regression or Platt scaling for reliable confidence scores
- SHAP (SHapley Additive exPlanations) for transaction-level feature importance
- Global feature importance analysis across the dataset
- Interactive visualizations showing which features drive each prediction
- Zero-value filtering for cleaner, more interpretable explanations
- Hybrid search combining semantic (vector) and keyword (BM25) retrieval
- Re-ranking using cross-encoder models for improved document relevance
- Multi-query retrieval generating query variations for better coverage
- Intelligent summarization using LangChain and local LLMs (Ollama)
- Context-aware retrieval of similar fraud cases and relevant policies
- Review queue for transactions requiring human judgment
- Override capability allowing reviewers to disagree with model predictions
- Feedback collection storing reviewer decisions and notes in SQLite database
- Analytics dashboard tracking agreement rates, escalation patterns, and review statistics
- Uncertainty-based selection prioritizing transactions where the model is least confident
- Diversity-based selection ensuring selected transactions cover diverse patterns
- Analytics integration tracking the value of active learning reviews
- Queue management for efficient review workflow
- Confidence calibration adjusting probabilities to match observed frequencies
- Uncertainty quantification measuring model confidence using entropy or margin
- Model retraining incorporating human feedback with multiple weighting strategies
- Version management tracking model iterations and performance
- MCP (Model Context Protocol) client infrastructure for external data enrichment
- Mock customer history simulation (demonstrates integration pattern)
- Mock merchant risk indicators (demonstrates external API pattern)
- Mock geographic risk assessment (demonstrates location-based enrichment)
- Note: Currently uses mock data for demo purposes. The infrastructure is in place to connect to real MCP servers or external APIs.
- XGBoost 3.1.3: Gradient boosting for fraud detection
- scikit-learn 1.8.0: Model calibration, metrics, and utilities
- SHAP 0.50.0: Model explainability
- NumPy 2.2.4, Pandas 2.3.3: Data manipulation
- ChromaDB 1.4.0: Vector database for document storage and retrieval
- sentence-transformers 5.2.0: Embeddings for semantic search
- LangChain 1.2.0: LLM orchestration and prompt management
- langchain-ollama 1.0.1: Local LLM integration via Ollama
- rank-bm25 0.2.2: Keyword-based search (BM25)
- Streamlit 1.53.0: Interactive web dashboard
- SQLite: Lightweight database for human feedback storage
- Matplotlib 3.10.8: Visualization and plotting
┌─────────────────────────────────────────────────────────────────┐
│ User Interface │
│ (Streamlit Dashboard) │
└────────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Transaction Analysis Tab │
│ • Load test data / Upload transactions │
│ • View predictions, probabilities, risk buckets │
│ • SHAP explanations │
│ • RAG context retrieval │
└────────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Core ML Pipeline │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌───────────────┐ │
│ │ XGBoost │───▶│ Calibration │───▶│ Uncertainty │ │
│ │ Model │ │ Module │ │ Quantification│ │
│ └──────────────┘ └──────────────┘ └───────────────┘ │
│ │ │ │ │
│ └───────────────────┴─────────────────────┘ │
│ │ │
│ ▼ │
│ ┌───────────────┐ │
│ │ Risk Bucket │ │
│ │ Classification│ │
│ └───────────────┘ │
└────────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Explainability Layer │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌───────────────┐ │
│ │ SHAP │───▶│ Feature │───▶│ Visualization │ │
│ │ Values │ │ Importance │ │ (Plots) │ │
│ └──────────────┘ └──────────────┘ └───────────────┘ │
└────────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ RAG System │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Hybrid │───▶│ Re-ranking │───▶│ Multi-query │ │
│ │ Search │ │ (Cross-enc) │ │ Retrieval │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │ │ │ │
│ └───────────────────┴─────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────┐ │
│ │ LangChain │ │
│ │ Summarization│ │
│ └──────────────┘ │
└────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Human Review Workflow │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Review │───▶│ Feedback │───▶│ Analytics │ │
│ │ Queue │ │ Collection │ │ Dashboard │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │ │ │ │
│ └───────────────────┴─────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────┐ │
│ │ Active │ │
│ │ Learning │ │
│ └──────────────┘ │
└────────────────────────────┬───────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Model Improvement │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Feedback │───▶│ Weighted │───▶│ Retraining │ │
│ │ Database │ │ Preparation │ │ Script │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────────┘
Transaction Input
│
▼
┌──────────────────┐
│ Feature Extract │
└────────┬─────────┘
│
▼
┌─────────────────┐ ┌──────────────┐
│ XGBoost Model │────▶│ Probability │
└────────┬────────┘ └──────┬───────┘
│ │
│ ▼
│ ┌─────────────────┐
│ │ Risk Bucket │
│ │ Classification │
│ └────────┬────────┘
│ │
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ SHAP Explain │ │ Decision Logic │
│ (Why this?) │ │ (Auto/Review) │
└────────┬────────┘ └────────┬────────┘
│ │
│ ▼
│ ┌─────────────────┐
│ │ RAG Context │
│ │ (Similar cases)│
│ └────────┬────────┘
│ │
└───────────────────────┘
│
▼
┌────────────────────┐
│ Human Reviewer │
│ (Override/Approve)│
└────────┬───────────┘
│
▼
┌─────────────────┐
│ Feedback DB │
│ (SQLite) │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Model Retrain │
│ (Incremental) │
└─────────────────┘
FinRAG/
├── app/ # Main application package
│ ├── main_streamlit.py # Streamlit dashboard (main entry point)
│ ├── main_streamlit_al_tab.py # Active Learning tab
│ ├── data_pipeline.py # Data loading and preprocessing
│ ├── model.py # Model loading and scoring
│ ├── explainability.py # SHAP explanations
│ ├── calibration.py # Probability calibration
│ ├── uncertainty.py # Uncertainty quantification
│ ├── active_learning.py # Active learning selection
│ ├── retraining.py # Model retraining utilities
│ ├── rag.py # Core RAG functionality
│ ├── rag_hybrid.py # Hybrid search (semantic + keyword)
│ ├── rag_rerank.py # Document re-ranking
│ ├── rag_multiquery.py # Multi-query retrieval
│ ├── langchain_rag.py # LangChain integration
│ ├── mcp_client.py # MCP client for external data
│ ├── human_feedback.py # Human feedback database management
│ └── config.py # Centralized configuration
│
├── scripts/ # Utility scripts
│ ├── train_model.py # Train XGBoost model
│ ├── retrain_with_feedback.py # Retrain with human feedback
│ ├── generate_synthetic_cases.py # Generate fraud cases for RAG
│ └── prepare_rag_data.py # Process documents for RAG
│
├── data/ # Data directory
│ ├── raw/ # Raw datasets (creditcard.csv)
│ ├── processed/ # Processed data
│ └── human_feedback.db # SQLite database for reviews
│
├── models/ # Trained models
│ ├── fraud_xgb.joblib # Base XGBoost model
│ ├── fraud_xgb_calibrated.joblib # Calibrated model
│ ├── optimal_threshold.json # Optimal decision threshold
│ └── calibration_info.json # Calibration metadata
│
├── rag_docs/ # RAG knowledge base
│ ├── raw/ # Original documents
│ └── processed/ # Processed and chunked documents
│ ├── *.txt # Processed text files
│ └── chroma_db/ # ChromaDB vector store
│
├── requirements.txt # Python dependencies
└──README.md # This file
-
Clone the repository
git clone https://github.com/pmr123/FraudLens.git
-
Create and activate virtual environment
python -m venv venv # On Windows: venv\Scripts\activate # On Linux/Mac: source venv/bin/activate
-
Install dependencies
pip install -r requirements.txt
-
Install and configure Ollama (for LLM features)
# Download from https://ollama.ai # Pull a small model (fits in 8GB GPU): ollama pull llama3.2:3b
-
Download dataset
- Download the Credit Card Fraud Detection dataset from Kaggle
- Place
creditcard.csvindata/raw/
-
Train the model
python -m scripts.train_model
This will:
- Train the XGBoost model
- Find optimal threshold
- Save calibrated model (if enabled)
- Save model artifacts to
models/
-
Prepare RAG documents (optional)
# Generate synthetic fraud cases python -m scripts.generate_synthetic_cases # Process documents for RAG python -m scripts.prepare_rag_data
-
Start the application
python -m streamlit run app/main_streamlit.py
Note: Always use
python -m streamlit run(not juststreamlit run) to ensure the correct Python environment is used, especially with conda environments.
The dashboard provides four main tabs:
-
Transaction Analysis
- Load test data or upload CSV files
- View model predictions with probabilities and risk buckets
- Explore SHAP explanations for individual transactions
- Retrieve RAG context (similar cases and policies)
- Generate AI-powered review reports
-
Review Queue
- View transactions requiring human review
- Approve or block transactions
- Add reviewer notes and escalate cases
- Track review history
-
Active Learning
- Generate prioritized queue of uncertain transactions
- Review transactions selected by active learning
- View uncertainty scores and statistics
- Submit feedback for model improvement
-
Analytics
- Review statistics (total reviews, agreement rates)
- Active learning analytics
- Disagreement case analysis
- Performance metrics
After collecting human feedback, retrain the model:
python -m scripts.retrain_with_feedback \
--mode incremental \
--strategy combined \
--min-reviews 50Options:
--mode:incremental(add to existing) orfull(retrain from scratch)--strategy:equal,uncertainty,al_priority,time_decay, orcombined--min-reviews: Minimum number of reviews required--only-disagreements: Only use transactions where human disagreed with model--only-al: Only use transactions selected by active learning
Key configuration options in app/config.py:
- Model: Model path, threshold settings
- RAG: Search method (semantic/keyword/hybrid), re-ranking, multi-query
- Calibration: Enable/disable, method (isotonic/platt)
- Uncertainty: Enable/disable, method (entropy/margin)
- Active Learning: Enable/disable, selection method, number of transactions
- LLM: Ollama base URL, model name, temperature
Combines semantic (vector) and keyword (BM25) search for better retrieval:
- Semantic search finds conceptually similar documents
- Keyword search finds exact term matches
- Weighted combination or Reciprocal Rank Fusion (RRF) for merging results
Improves document relevance by re-scoring retrieved documents:
- Cross-encoder: Most accurate, processes query+document together
- LLM-based: Uses Ollama to score relevance
- Feature-based: Fast metadata-based re-ranking
Generates query variations to improve coverage:
- LLM generates alternative phrasings
- Retrieves documents for each variation
- Aggregates results with deduplication
Selects most informative transactions for review:
- Entropy-based: Prioritizes high uncertainty
- Margin-based: Focuses on borderline cases
- Diverse: Combines uncertainty with feature diversity
Incorporates human feedback into model:
- Multiple weighting strategies (uncertainty, time decay, AL priority)
- Incremental or full retraining modes
- Model versioning and performance tracking
- Model Performance: 86% precision, 86% recall on test set
- Inference Speed: <10ms per transaction (XGBoost)
- RAG Retrieval: <500ms for hybrid search + re-ranking
- SHAP Computation: <1s per transaction (TreeExplainer)