Skip to content

About

AI-powered invoice data extraction system with multi-OCR support (Mathpix, Tesseract, Claude Vision). Converts PDF invoices into structured JSON/CSV with validation in ~5 seconds. Built with Claude API.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

Invoice Processor

AI-powered invoice data extraction that converts PDF invoices into structured data in seconds.

Python Claude API License

The Problem

Accountants and bookkeepers spend 5+ minutes per invoice manually entering data into spreadsheets:

  • Tedious, repetitive work
  • Prone to typos and errors
  • Doesn't scale
  • Nobody enjoys it

The Solution

This system extracts invoice data automatically in ~5 seconds:

  • Works with any invoice format (LLM understands context)
  • Handles both text-based and scanned PDFs
  • Validates extracted data (math checks, required fields)
  • Exports to JSON, CSV, or Excel
  • Flags low-confidence items for human review

Architecture

PDF Invoice
     │
     ▼
┌─────────────────┐
│  PDF Extractor  │  Detects PDF type (text vs scanned)
│                 │  Extracts raw text or uses OCR
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│  LLM Processor  │  Claude API extracts structured data:
│                 │  vendor, date, total, line items...
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│    Validator    │  Checks: required fields, date format,
│                 │  currency, math (items sum to total)
└────────┬────────┘
         │
    ┌────┴────┐
    │         │
 Valid?    Invalid?
    │         │
    ▼         ▼
┌────────┐ ┌────────┐
│ Export │ │Flagged │
│JSON/CSV│ │ Review │
└────────┘ └────────┘

Quick Start

Option 1: Docker (Recommended)

# 1. Set API key
export ANTHROPIC_API_KEY=sk-ant-...

# 2. Run with Docker
docker-compose up

# Results in output/ folder

Option 2: Local Installation

# 1. Install dependencies
pip install -r requirements.txt

# 2. Configure API key
cp .env.example .env
# Edit .env: ANTHROPIC_API_KEY=sk-ant-...

# 3. Process invoices
python -m execution.process_all

# With cost tracking:
python -m execution.process_all --cost-report

Find Results

  • output/extracted/ - JSON file per invoice
  • output/exports/invoices.csv - Master spreadsheet
  • output/flagged/ - Items needing human review
  • output/cost_tracking.json - API usage & costs

Example Output

Input: PDF invoice from any vendor

Output: Structured JSON

{
  "vendor_name": "Acme Corporation",
  "vendor_address": "123 Business Ave, New York, NY 10001",
  "invoice_number": "INV-2024-0042",
  "invoice_date": "2024-01-15",
  "due_date": "2024-02-14",
  "currency": "USD",
  "subtotal": 1250.00,
  "tax_rate": 8,
  "tax_amount": 100.00,
  "total": 1350.00,
  "line_items": [
    {
      "description": "Consulting Services",
      "quantity": 5,
      "unit_price": 200.00,
      "total": 1000.00
    },
    {
      "description": "Software License",
      "quantity": 1,
      "unit_price": 250.00,
      "total": 250.00
    }
  ],
  "confidence": "high"
}

Project Structure

invoice-processor/
├── config/
│   └── settings.py          # API keys, paths, thresholds
├── directives/
│   ├── process_invoice.md   # Main workflow documentation
│   ├── handle_scanned.md    # OCR fallback process
│   ├── validation_rules.md  # Data quality rules
│   └── known_formats.md     # Vendor-specific notes
├── execution/
│   ├── pdf_extractor.py     # PDF text extraction
│   ├── llm_processor.py     # Claude/OpenAI integration
│   ├── validator.py         # Data validation
│   ├── exporter.py          # JSON/CSV/Excel export
│   ├── process_all.py       # Batch processing
│   └── utils/
│       ├── prompts.py       # LLM prompt templates
│       └── ocr.py           # Tesseract OCR wrapper
├── input/                   # Drop PDFs here
├── output/
│   ├── extracted/           # JSON per invoice
│   ├── exports/             # CSV/Excel files
│   └── flagged/             # Low confidence items
├── tests/
│   ├── test_extractor.py    # Unit tests
│   └── generate_samples.py  # Generate test invoices
├── .env.example
├── requirements.txt
└── README.md

Tech Stack

Component Technology
PDF Extraction pdfplumber, PyMuPDF
OCR (scanned PDFs) Tesseract, Claude Vision
LLM Processing Claude API (Anthropic)
Data Validation Custom + Pydantic
Export pandas, openpyxl
Progress Tracking tqdm, colorama
Error Handling Custom error handler
Retry Logic Exponential backoff
Deployment Docker, Docker Compose
CI/CD GitHub Actions
Testing pytest

Key Features

Core Processing

  • Intelligent PDF Detection - Auto-detects text vs scanned PDFs
  • LLM-Powered Extraction - Works with any format without custom rules
  • Math Validation - Verifies line items sum to totals
  • Confidence Scoring - LLM self-reports confidence, flags low scores
  • Batch Processing - Process folders with progress tracking

Production Features ✨

Cost Tracking

# Real-time API cost monitoring
python -m execution.utils.cost_tracker --verbose
  • Per-invoice and aggregate costs
  • Export to CSV for accounting
  • Supports all Claude & OpenAI models

Progress Bars

  • Visual progress during batch processing
  • Color-coded status (✓ OK, ✗ Failed, ⚠ Flagged)
  • Real-time file tracking

Enhanced Error Handling

  • Detailed, actionable error messages
  • 8 error categories with suggestions
  • Automatic recovery on retryable errors

Retry Logic

  • Exponential backoff on failures
  • Auto-retry on rate limits & 5xx errors
  • Configurable (default: 3 retries)

Docker Support

  • One-command deployment
  • Multi-stage build for size
  • Health checks included

CI/CD Pipeline

  • Automated testing on push
  • Security scanning
  • Multi-Python version support (3.10-3.12)

Testing

# Generate sample invoices
python tests/generate_samples.py --count 5

# Run unit tests
pytest tests/ -v

# Test PDF extraction
python -m execution.pdf_extractor input/sample_invoice_001.pdf --verbose

# Test validation
python -m execution.validator --file tests/sample_invoices/expected_001.json

Cost Estimate

Model Cost per Invoice 1000 Invoices/Month
Claude Haiku ~$0.001 ~$1
Claude Sonnet ~$0.01 ~$10
GPT-4o-mini ~$0.002 ~$2

Recommendation: Start with Haiku - it's excellent at structured extraction.

Documentation

Advanced Features (Optional)

Want more? These can be added:

  • Web UI (Streamlit) for upload and review
  • Database (PostgreSQL) for persistence
  • REST API (FastAPI) for integrations
  • Google Drive integration (watch folder)
  • QuickBooks/Xero direct integration
  • Duplicate invoice detection

See PRODUCTION_READY.md for current production features.

License

MIT License - see LICENSE file.


Production Ready ✅ | Built with Claude API | Documentation

About

AI-powered invoice data extraction system with multi-OCR support (Mathpix, Tesseract, Claude Vision). Converts PDF invoices into structured JSON/CSV with validation in ~5 seconds. Built with Claude API.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages