AI-powered invoice data extraction that converts PDF invoices into structured data in seconds.
Accountants and bookkeepers spend 5+ minutes per invoice manually entering data into spreadsheets:
- Tedious, repetitive work
- Prone to typos and errors
- Doesn't scale
- Nobody enjoys it
This system extracts invoice data automatically in ~5 seconds:
- Works with any invoice format (LLM understands context)
- Handles both text-based and scanned PDFs
- Validates extracted data (math checks, required fields)
- Exports to JSON, CSV, or Excel
- Flags low-confidence items for human review
PDF Invoice
│
▼
┌─────────────────┐
│ PDF Extractor │ Detects PDF type (text vs scanned)
│ │ Extracts raw text or uses OCR
└────────┬────────┘
│
▼
┌─────────────────┐
│ LLM Processor │ Claude API extracts structured data:
│ │ vendor, date, total, line items...
└────────┬────────┘
│
▼
┌─────────────────┐
│ Validator │ Checks: required fields, date format,
│ │ currency, math (items sum to total)
└────────┬────────┘
│
┌────┴────┐
│ │
Valid? Invalid?
│ │
▼ ▼
┌────────┐ ┌────────┐
│ Export │ │Flagged │
│JSON/CSV│ │ Review │
└────────┘ └────────┘
# 1. Set API key
export ANTHROPIC_API_KEY=sk-ant-...
# 2. Run with Docker
docker-compose up
# Results in output/ folder# 1. Install dependencies
pip install -r requirements.txt
# 2. Configure API key
cp .env.example .env
# Edit .env: ANTHROPIC_API_KEY=sk-ant-...
# 3. Process invoices
python -m execution.process_all
# With cost tracking:
python -m execution.process_all --cost-reportoutput/extracted/- JSON file per invoiceoutput/exports/invoices.csv- Master spreadsheetoutput/flagged/- Items needing human reviewoutput/cost_tracking.json- API usage & costs
Input: PDF invoice from any vendor
Output: Structured JSON
{
"vendor_name": "Acme Corporation",
"vendor_address": "123 Business Ave, New York, NY 10001",
"invoice_number": "INV-2024-0042",
"invoice_date": "2024-01-15",
"due_date": "2024-02-14",
"currency": "USD",
"subtotal": 1250.00,
"tax_rate": 8,
"tax_amount": 100.00,
"total": 1350.00,
"line_items": [
{
"description": "Consulting Services",
"quantity": 5,
"unit_price": 200.00,
"total": 1000.00
},
{
"description": "Software License",
"quantity": 1,
"unit_price": 250.00,
"total": 250.00
}
],
"confidence": "high"
}invoice-processor/
├── config/
│ └── settings.py # API keys, paths, thresholds
├── directives/
│ ├── process_invoice.md # Main workflow documentation
│ ├── handle_scanned.md # OCR fallback process
│ ├── validation_rules.md # Data quality rules
│ └── known_formats.md # Vendor-specific notes
├── execution/
│ ├── pdf_extractor.py # PDF text extraction
│ ├── llm_processor.py # Claude/OpenAI integration
│ ├── validator.py # Data validation
│ ├── exporter.py # JSON/CSV/Excel export
│ ├── process_all.py # Batch processing
│ └── utils/
│ ├── prompts.py # LLM prompt templates
│ └── ocr.py # Tesseract OCR wrapper
├── input/ # Drop PDFs here
├── output/
│ ├── extracted/ # JSON per invoice
│ ├── exports/ # CSV/Excel files
│ └── flagged/ # Low confidence items
├── tests/
│ ├── test_extractor.py # Unit tests
│ └── generate_samples.py # Generate test invoices
├── .env.example
├── requirements.txt
└── README.md
| Component | Technology |
|---|---|
| PDF Extraction | pdfplumber, PyMuPDF |
| OCR (scanned PDFs) | Tesseract, Claude Vision |
| LLM Processing | Claude API (Anthropic) |
| Data Validation | Custom + Pydantic |
| Export | pandas, openpyxl |
| Progress Tracking | tqdm, colorama |
| Error Handling | Custom error handler |
| Retry Logic | Exponential backoff |
| Deployment | Docker, Docker Compose |
| CI/CD | GitHub Actions |
| Testing | pytest |
- Intelligent PDF Detection - Auto-detects text vs scanned PDFs
- LLM-Powered Extraction - Works with any format without custom rules
- Math Validation - Verifies line items sum to totals
- Confidence Scoring - LLM self-reports confidence, flags low scores
- Batch Processing - Process folders with progress tracking
Cost Tracking
# Real-time API cost monitoring
python -m execution.utils.cost_tracker --verbose- Per-invoice and aggregate costs
- Export to CSV for accounting
- Supports all Claude & OpenAI models
Progress Bars
- Visual progress during batch processing
- Color-coded status (✓ OK, ✗ Failed, ⚠ Flagged)
- Real-time file tracking
Enhanced Error Handling
- Detailed, actionable error messages
- 8 error categories with suggestions
- Automatic recovery on retryable errors
Retry Logic
- Exponential backoff on failures
- Auto-retry on rate limits & 5xx errors
- Configurable (default: 3 retries)
Docker Support
- One-command deployment
- Multi-stage build for size
- Health checks included
CI/CD Pipeline
- Automated testing on push
- Security scanning
- Multi-Python version support (3.10-3.12)
# Generate sample invoices
python tests/generate_samples.py --count 5
# Run unit tests
pytest tests/ -v
# Test PDF extraction
python -m execution.pdf_extractor input/sample_invoice_001.pdf --verbose
# Test validation
python -m execution.validator --file tests/sample_invoices/expected_001.json| Model | Cost per Invoice | 1000 Invoices/Month |
|---|---|---|
| Claude Haiku | ~$0.001 | ~$1 |
| Claude Sonnet | ~$0.01 | ~$10 |
| GPT-4o-mini | ~$0.002 | ~$2 |
Recommendation: Start with Haiku - it's excellent at structured extraction.
- DEPLOYMENT.md - Docker deployment & production guide
- PRODUCTION_READY.md - Complete feature list
- PHASE1_COMPLETE.md - Development features
- CLAUDE.md - Architecture & agent instructions
- output/OUTPUT_README.md - Output format guide
Want more? These can be added:
- Web UI (Streamlit) for upload and review
- Database (PostgreSQL) for persistence
- REST API (FastAPI) for integrations
- Google Drive integration (watch folder)
- QuickBooks/Xero direct integration
- Duplicate invoice detection
See PRODUCTION_READY.md for current production features.
MIT License - see LICENSE file.
Production Ready ✅ | Built with Claude API | Documentation