Skip to content
pmr123Public

About

FraudLens is an end-to-end fraud detection system that combines machine learning, explainable AI, and retrieval-augmented generation (RAG) to assist human reviewers in making informed fraud decisions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

FraudLens

FraudLens is an end-to-end fraud detection system that combines machine learning, explainable AI, and retrieval-augmented generation (RAG) to assist human reviewers in making informed fraud decisions.

Overview

This system provides an intelligent fraud review workflow that:

  • Detects fraudulent transactions using XGBoost with calibrated probabilities
  • Explains model decisions using SHAP values
  • Retrieves relevant context (similar cases, policies) using advanced RAG techniques
  • Enables human reviewers to override model decisions and provide feedback
  • Continuously improves through active learning and model retraining

Key Features

Core Fraud Detection

  • XGBoost-based fraud detection with 86% precision and 86% recall
  • Risk bucket classification (LOW/MEDIUM/HIGH) for automated decision routing
  • Optimal threshold optimization balancing precision and recall
  • Probability calibration using isotonic regression or Platt scaling for reliable confidence scores

Explainability

  • SHAP (SHapley Additive exPlanations) for transaction-level feature importance
  • Global feature importance analysis across the dataset
  • Interactive visualizations showing which features drive each prediction
  • Zero-value filtering for cleaner, more interpretable explanations

Advanced RAG System

  • Hybrid search combining semantic (vector) and keyword (BM25) retrieval
  • Re-ranking using cross-encoder models for improved document relevance
  • Multi-query retrieval generating query variations for better coverage
  • Intelligent summarization using LangChain and local LLMs (Ollama)
  • Context-aware retrieval of similar fraud cases and relevant policies

Human-AI Collaboration

  • Review queue for transactions requiring human judgment
  • Override capability allowing reviewers to disagree with model predictions
  • Feedback collection storing reviewer decisions and notes in SQLite database
  • Analytics dashboard tracking agreement rates, escalation patterns, and review statistics

Active Learning

  • Uncertainty-based selection prioritizing transactions where the model is least confident
  • Diversity-based selection ensuring selected transactions cover diverse patterns
  • Analytics integration tracking the value of active learning reviews
  • Queue management for efficient review workflow

Model Improvement

  • Confidence calibration adjusting probabilities to match observed frequencies
  • Uncertainty quantification measuring model confidence using entropy or margin
  • Model retraining incorporating human feedback with multiple weighting strategies
  • Version management tracking model iterations and performance

External Data Integration (Mock/Demo)

  • MCP (Model Context Protocol) client infrastructure for external data enrichment
  • Mock customer history simulation (demonstrates integration pattern)
  • Mock merchant risk indicators (demonstrates external API pattern)
  • Mock geographic risk assessment (demonstrates location-based enrichment)
  • Note: Currently uses mock data for demo purposes. The infrastructure is in place to connect to real MCP servers or external APIs.

Tech Stack

Machine Learning

  • XGBoost 3.1.3: Gradient boosting for fraud detection
  • scikit-learn 1.8.0: Model calibration, metrics, and utilities
  • SHAP 0.50.0: Model explainability
  • NumPy 2.2.4, Pandas 2.3.3: Data manipulation

RAG & LLM

  • ChromaDB 1.4.0: Vector database for document storage and retrieval
  • sentence-transformers 5.2.0: Embeddings for semantic search
  • LangChain 1.2.0: LLM orchestration and prompt management
  • langchain-ollama 1.0.1: Local LLM integration via Ollama
  • rank-bm25 0.2.2: Keyword-based search (BM25)

UI & Infrastructure

  • Streamlit 1.53.0: Interactive web dashboard
  • SQLite: Lightweight database for human feedback storage
  • Matplotlib 3.10.8: Visualization and plotting

System Architecture

┌─────────────────────────────────────────────────────────────────┐
│                         User Interface                          │
│                    (Streamlit Dashboard)                        │
└────────────────────────────┬────────────────────────────────────┘
                             │
                             ▼
┌─────────────────────────────────────────────────────────────────┐
│                    Transaction Analysis Tab                     │
│  • Load test data / Upload transactions                         │
│  • View predictions, probabilities, risk buckets                │
│  • SHAP explanations                                            │
│  • RAG context retrieval                                        │
└────────────────────────────┬────────────────────────────────────┘
                             │
                             ▼
┌─────────────────────────────────────────────────────────────────┐
│                      Core ML Pipeline                           │
│                                                                 │
│  ┌──────────────┐    ┌──────────────┐     ┌───────────────┐     │
│  │   XGBoost    │───▶│ Calibration  │───▶│ Uncertainty   │     │
│  │   Model      │    │   Module     │     │ Quantification│     │
│  └──────────────┘    └──────────────┘     └───────────────┘     │
│         │                   │                     │             │
│         └───────────────────┴─────────────────────┘             │
│                            │                                    │
│                            ▼                                    │
│                   ┌───────────────┐                             │
│                   │ Risk Bucket   │                             │
│                   │ Classification│                             │
│                   └───────────────┘                             │
└────────────────────────────┬────────────────────────────────────┘
                             │
                             ▼
┌─────────────────────────────────────────────────────────────────┐
│                    Explainability Layer                         │
│                                                                 │
│  ┌──────────────┐    ┌──────────────┐     ┌───────────────┐     │
│  │   SHAP       │───▶│ Feature      │───▶│ Visualization │     │
│  │  Values      │    │ Importance   │     │   (Plots)     │     │
│  └──────────────┘    └──────────────┘     └───────────────┘     │
└────────────────────────────┬────────────────────────────────────┘
                             │
                             ▼
┌─────────────────────────────────────────────────────────────────┐
│                      RAG System                                 │
│                                                                 │
│  ┌──────────────┐     ┌──────────────┐    ┌──────────────┐      │
│  │   Hybrid     │───▶│  Re-ranking  │───▶│ Multi-query  │      │
│  │   Search     │     │  (Cross-enc) │    │ Retrieval    │      │
│  └──────────────┘     └──────────────┘    └──────────────┘      │
│         │                   │                     │             │
│         └───────────────────┴─────────────────────┘             │
│                            │                                    │
│                            ▼                                    │
│                   ┌──────────────┐                              │
│                   │  LangChain   │                              │
│                   │ Summarization│                              │
│                   └──────────────┘                              │
└────────────────────────────┬────────────────────────────────────┘
                             │
                             ▼
┌────────────────────────────────────────────────────────────────┐
│                    Human Review Workflow                       │
│                                                                │
│  ┌──────────────┐     ┌──────────────┐     ┌──────────────┐    │
│  │   Review     │───▶│   Feedback    │───▶│  Analytics   │    │
│  │   Queue      │     │   Collection │     │  Dashboard   │    │
│  └──────────────┘     └──────────────┘     └──────────────┘    │
│         │                   │                     │            │
│         └───────────────────┴─────────────────────┘            │
│                            │                                   │
│                            ▼                                   │
│                   ┌──────────────┐                             │
│                   │ Active       │                             │
│                   │ Learning     │                             │
│                   └──────────────┘                             │
└────────────────────────────┬───────────────────────────────────┘
                             │
                             ▼
┌─────────────────────────────────────────────────────────────────┐
│                    Model Improvement                            │
│                                                                 │
│  ┌──────────────┐     ┌──────────────┐    ┌──────────────┐      │
│  │   Feedback   │───▶│  Weighted    │───▶│  Retraining  │      │
│  │   Database   │     │  Preparation │    │  Script      │      │
│  └──────────────┘     └──────────────┘    └──────────────┘      │
└─────────────────────────────────────────────────────────────────┘

Data Flow

Transaction Input
    │
    ▼
┌──────────────────┐
│ Feature Extract  │
└────────┬─────────┘
         │
         ▼
┌─────────────────┐      ┌──────────────┐
│  XGBoost Model  │────▶│  Probability │
└────────┬────────┘      └──────┬───────┘
         │                     │
         │                     ▼
         │            ┌─────────────────┐
         │            │ Risk Bucket     │
         │            │ Classification  │
         │            └────────┬────────┘
         │                     │
         ▼                     ▼
┌─────────────────┐     ┌─────────────────┐
│  SHAP Explain   │     │  Decision Logic │
│  (Why this?)    │     │  (Auto/Review)  │
└────────┬────────┘     └────────┬────────┘
         │                       │
         │                       ▼
         │              ┌─────────────────┐
         │              │  RAG Context    │
         │              │  (Similar cases)│
         │              └────────┬────────┘
         │                       │
         └───────────────────────┘
                       │
                       ▼
              ┌────────────────────┐
              │  Human Reviewer    │
              │  (Override/Approve)│
              └────────┬───────────┘
                       │
                       ▼
              ┌─────────────────┐
              │  Feedback DB    │
              │  (SQLite)       │
              └────────┬────────┘
                       │
                       ▼
              ┌─────────────────┐
              │  Model Retrain  │
              │  (Incremental)  │
              └─────────────────┘

Project Structure

FinRAG/
├── app/                          # Main application package
│   ├── main_streamlit.py         # Streamlit dashboard (main entry point)
│   ├── main_streamlit_al_tab.py  # Active Learning tab
│   ├── data_pipeline.py          # Data loading and preprocessing
│   ├── model.py                  # Model loading and scoring
│   ├── explainability.py         # SHAP explanations
│   ├── calibration.py            # Probability calibration
│   ├── uncertainty.py            # Uncertainty quantification
│   ├── active_learning.py        # Active learning selection
│   ├── retraining.py             # Model retraining utilities
│   ├── rag.py                    # Core RAG functionality
│   ├── rag_hybrid.py             # Hybrid search (semantic + keyword)
│   ├── rag_rerank.py             # Document re-ranking
│   ├── rag_multiquery.py         # Multi-query retrieval
│   ├── langchain_rag.py          # LangChain integration
│   ├── mcp_client.py             # MCP client for external data
│   ├── human_feedback.py         # Human feedback database management
│   └── config.py                 # Centralized configuration
│
├── scripts/                      # Utility scripts
│   ├── train_model.py            # Train XGBoost model
│   ├── retrain_with_feedback.py  # Retrain with human feedback
│   ├── generate_synthetic_cases.py # Generate fraud cases for RAG
│   └── prepare_rag_data.py       # Process documents for RAG
│
├── data/                         # Data directory
│   ├── raw/                      # Raw datasets (creditcard.csv)
│   ├── processed/                # Processed data
│   └── human_feedback.db         # SQLite database for reviews
│
├── models/                       # Trained models
│   ├── fraud_xgb.joblib          # Base XGBoost model
│   ├── fraud_xgb_calibrated.joblib # Calibrated model
│   ├── optimal_threshold.json    # Optimal decision threshold
│   └── calibration_info.json    # Calibration metadata
│
├── rag_docs/                     # RAG knowledge base
│   ├── raw/                      # Original documents
│   └── processed/                # Processed and chunked documents
│       ├── *.txt                 # Processed text files
│       └── chroma_db/            # ChromaDB vector store
│
├── requirements.txt              # Python dependencies
└──README.md                     # This file

Installation

Setup Steps

  1. Clone the repository

    git clone https://github.com/pmr123/FraudLens.git
  2. Create and activate virtual environment

    python -m venv venv
    # On Windows:
    venv\Scripts\activate
    # On Linux/Mac:
    source venv/bin/activate
  3. Install dependencies

    pip install -r requirements.txt
  4. Install and configure Ollama (for LLM features)

    # Download from https://ollama.ai
    # Pull a small model (fits in 8GB GPU):
    ollama pull llama3.2:3b
  5. Download dataset

    • Download the Credit Card Fraud Detection dataset from Kaggle
    • Place creditcard.csv in data/raw/
  6. Train the model

    python -m scripts.train_model

    This will:

    • Train the XGBoost model
    • Find optimal threshold
    • Save calibrated model (if enabled)
    • Save model artifacts to models/
  7. Prepare RAG documents (optional)

    # Generate synthetic fraud cases
    python -m scripts.generate_synthetic_cases
    
    # Process documents for RAG
    python -m scripts.prepare_rag_data
  8. Start the application

    python -m streamlit run app/main_streamlit.py

    Note: Always use python -m streamlit run (not just streamlit run) to ensure the correct Python environment is used, especially with conda environments.

Usage

Streamlit Dashboard

The dashboard provides four main tabs:

  1. Transaction Analysis

    • Load test data or upload CSV files
    • View model predictions with probabilities and risk buckets
    • Explore SHAP explanations for individual transactions
    • Retrieve RAG context (similar cases and policies)
    • Generate AI-powered review reports
  2. Review Queue

    • View transactions requiring human review
    • Approve or block transactions
    • Add reviewer notes and escalate cases
    • Track review history
  3. Active Learning

    • Generate prioritized queue of uncertain transactions
    • Review transactions selected by active learning
    • View uncertainty scores and statistics
    • Submit feedback for model improvement
  4. Analytics

    • Review statistics (total reviews, agreement rates)
    • Active learning analytics
    • Disagreement case analysis
    • Performance metrics

Model Retraining

After collecting human feedback, retrain the model:

python -m scripts.retrain_with_feedback \
    --mode incremental \
    --strategy combined \
    --min-reviews 50

Options:

  • --mode: incremental (add to existing) or full (retrain from scratch)
  • --strategy: equal, uncertainty, al_priority, time_decay, or combined
  • --min-reviews: Minimum number of reviews required
  • --only-disagreements: Only use transactions where human disagreed with model
  • --only-al: Only use transactions selected by active learning

Configuration

Key configuration options in app/config.py:

  • Model: Model path, threshold settings
  • RAG: Search method (semantic/keyword/hybrid), re-ranking, multi-query
  • Calibration: Enable/disable, method (isotonic/platt)
  • Uncertainty: Enable/disable, method (entropy/margin)
  • Active Learning: Enable/disable, selection method, number of transactions
  • LLM: Ollama base URL, model name, temperature

Features in Detail

Hybrid Search

Combines semantic (vector) and keyword (BM25) search for better retrieval:

  • Semantic search finds conceptually similar documents
  • Keyword search finds exact term matches
  • Weighted combination or Reciprocal Rank Fusion (RRF) for merging results

Re-ranking

Improves document relevance by re-scoring retrieved documents:

  • Cross-encoder: Most accurate, processes query+document together
  • LLM-based: Uses Ollama to score relevance
  • Feature-based: Fast metadata-based re-ranking

Multi-query Retrieval

Generates query variations to improve coverage:

  • LLM generates alternative phrasings
  • Retrieves documents for each variation
  • Aggregates results with deduplication

Active Learning

Selects most informative transactions for review:

  • Entropy-based: Prioritizes high uncertainty
  • Margin-based: Focuses on borderline cases
  • Diverse: Combines uncertainty with feature diversity

Model Retraining

Incorporates human feedback into model:

  • Multiple weighting strategies (uncertainty, time decay, AL priority)
  • Incremental or full retraining modes
  • Model versioning and performance tracking

Performance

  • Model Performance: 86% precision, 86% recall on test set
  • Inference Speed: <10ms per transaction (XGBoost)
  • RAG Retrieval: <500ms for hybrid search + re-ranking
  • SHAP Computation: <1s per transaction (TreeExplainer)

References

About

FraudLens is an end-to-end fraud detection system that combines machine learning, explainable AI, and retrieval-augmented generation (RAG) to assist human reviewers in making informed fraud decisions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages