Object-centric, uncertainty-aware financial close controller for Razorpay merchants.
Live console: https://paridhipawaiya.github.io/proofledger-ai/
The console is a static build and needs the API to be reachable. If it reports that the workspace could not load, the backend is asleep or not yet configured — see docs/deployment.md.
ProofLedger reconstructs how merchant orders become captured payments, refunds, settlements, bank credits, and accounting entries. It combines deterministic finance controls, calibrated matching, minimum-evidence human review, bounded Gemini explanations, and independently verifiable close certificates.
The product thesis is simple: finance teams need proof, not a confident-looking match.
Generic reconciliation tools flatten rows and return a match score. ProofLedger instead:
- Reconstructs an object/event graph across five financial sources.
- Imports all five of them — orders, settlements, refunds, bank statements, and ledger — as separately signed, separately approved, separately reversible batches.
- Stages real CSV exports behind size, encoding, shape, and row-validation boundaries.
- Suggests schema mappings from headers only, then requires controller confirmation.
- Atomically normalizes accepted rows and seals them in an Ed25519-signed manifest.
- Separates signing from authority through a hashed controller activation event.
- Recomputes graphs, matches, controls, reviews, journals, and certificates after activation.
- Rejects duplicate bank rows that attempt to mask a known source mismatch.
- Persists signed batches and ordered activation authority through API restarts.
- Applies exact and composite evidence before any probabilistic assistance.
- Abstains when candidate evidence is unsafe or ambiguous.
- Reports that abstention boundary through split-conformal calibration on held-out labels.
- Asks the smallest question likely to resolve that uncertainty.
- Hashes controller evidence separately, preserves the original source row, and recomputes.
- Enforces ten deterministic accounting and lifecycle controls.
- Proposes a balanced journal that remains pending human approval.
- Issues a proof-carrying settlement certificate only when critical controls pass.
- Detects if any imported or certified evidence changes—even by ₹1.
- Benchmarks safety using incorrect automatic approvals, not only aggregate accuracy.
The deterministic seed creates more than 1,200 records across merchant orders, Razorpay-shaped reconciliation, refunds, bank statements, and a general ledger. It injects missing bank evidence, a ₹1 amount difference, absent references, and a duplicate journal.
The default workspace (seed 2026, 600 orders, 12 settlement batches) is what the running
application and the deployed demo report:
| Method | Precision | Recall | F1 | Human review | Wrong auto-approvals |
|---|---|---|---|---|---|
| Exact ID baseline | 100% | 72.7% | 84.2% | 0% | 0 |
| Fuzzy narration baseline | 75% | 81.8% | 78.3% | 0% | 3 |
| ProofLedger | 100% | 100% | 100% | 16.7% | 0 |
The smaller reproducible CLI configuration (--seed 11 --orders 120 --settlement-size 20,
6 settlement batches) is harder for the baselines and is what the test suite pins:
| Method | Precision | Recall | F1 | Human review | Wrong auto-approvals |
|---|---|---|---|---|---|
| Exact ID baseline | 100% | 40% | 57.1% | 0% | 0 |
| Fuzzy narration baseline | 50% | 60% | 54.6% | 0% | 3 |
| ProofLedger | 100% | 100% | 100% | 33.3% | 0 |
The invariant across both configurations is the one that costs money: fuzzy matching produces three silent incorrect auto-approvals, and ProofLedger produces none. These numbers are from synthetic data and are not presented as production performance. The committed benchmark exposes its labels, methods, and caveats for inspection.
GET /api/v1/calibration fits a split-conformal minimum-confidence threshold on half the
labelled payouts and reports coverage on the half the fit never saw, then sizes the candidate
set behind every open review question at that threshold.
- Close command: captured volume, close readiness, exception runway, and proof graph counts.
- Evidence intake: real CSV preview, deterministic/AI-assisted header mapping, atomic import, signed manifest export, ₹1 tamper lab, audited activation, and live close recomputation.
- Settlement book: payout-by-payout evidence, controls, and balanced journal proposals.
- Evidence review: attach the missing bank row, verify its UTR, and watch the close recompute.
- Lifecycle graph: order → payment → refund → settlement → bank → ledger, laid out by stage. The default spine folds a payout's member records into stage totals; one toggle expands every record and source event.
- Safety benchmark: exact, fuzzy, and ProofLedger outcomes side by side, plus the split-conformal threshold, its held-out coverage, and the candidate-set size behind each open review question.
- Certificate lab: issue a valid close certificate, change one source by ₹1, and watch verification fail.
Gemini is optional. When configured, it may:
- explain a deterministic control result in plain language;
- phrase the next evidence question;
- suggest a source-to-canonical schema mapping using column names only.
Row values are never sent to Gemini. AI suggestions are constrained to existing source headers and canonical fields, cannot commit an import, and invalidate the controller confirmation checkbox.
Gemini cannot accept a match, pass a control, approve a journal, alter a source record, or close a settlement. Without an API key, the same application works with a deterministic explanation fallback.
Five source systems
└─> bounded CSV staging + controller-confirmed mapping
└─> Ed25519 manifest + immutable, SHA-256-hashed evidence records
└─> verified + controller-activated authority
└─> object-centric lifecycle graph
├─> exact / composite reconciliation
├─> deterministic safe abstention (split-conformal reports on it)
├─> ten deterministic finance controls
└─> minimum-evidence controller review
├─> balanced journal proposal
└─> verifiable settlement certificate
Authoritative money calculations use integer paise. Probabilistic output never enters an accounting equation.
apps/
api/proofledger/
domain/ immutable models, graph, controls, reconciliation, calibration
services/ synthetic data, benchmark, bounded AI, closing, demo workspace
api.py FastAPI product endpoints
main.py application entrypoint
web/
src/ React operator console
docs/ architecture, controls, evaluation, demo, threat model
tests/ ingestion, finance engine, certificate, benchmark, and API tests
.github/workflows continuous integration
Requirements:
- Python 3.10 or newer
- Node.js 22 or newer
Backend:
python -m pip install -e ".[dev]"
python -m uvicorn proofledger.main:app --reloadFrontend, in a second terminal:
cd apps/web
npm install
npm run devOpen http://localhost:5173. FastAPI documentation is available at http://localhost:8000/docs.
Committed import manifests, normalized records, and activation events persist in the configured
database. The default is sqlite:///./proofledger.db. PostgreSQL uses a standard SQLAlchemy URL,
for example postgresql+psycopg://user:password@host:5432/proofledger.
The demo does not require Gemini. To enable bounded explanations, copy .env.example to .env and set PROOFLEDGER_GEMINI_API_KEY.
Reproduce a workspace or benchmark directly:
proofledger summary
proofledger --seed 11 --orders 120 --settlement-size 20 benchmarkPersist Docker data across container replacement:
docker run --name proofledger -p 8000:8000 `
-e PROOFLEDGER_DATABASE_URL=sqlite:////data/proofledger.db `
-v proofledger-data:/data proofledger-ai:localpython -m ruff check .
python -m pytest --cov=proofledger --cov-report=term-missing
cd apps/web
npm audit --omit=dev
npm run lint
npm run test
npm run build| Endpoint | Purpose |
|---|---|
| GET /health | Runtime health |
| GET /api/v1/overview | Close metrics and graph summary |
| GET /api/v1/settlements | Settlement control book |
| GET /api/v1/settlements/{id} | Evidence, controls, decision, and journal |
| GET /api/v1/reviews | Minimum-evidence review queue |
| GET /api/v1/reviews/history | Append-only controller resolution trail |
| POST /api/v1/reviews/{id}/resolve | Attach hashed evidence and recompute the close |
| POST /api/v1/ingestion/preview | Stage, validate, sample, hash, and suggest CSV mappings |
| POST /api/v1/ingestion/{upload_id}/ai-map | Header-only bounded mapping suggestion |
| POST /api/v1/ingestion/commit | Atomically normalize records and sign the manifest |
| GET /api/v1/ingestion/manifests | List committed import manifests |
| GET /api/v1/ingestion/activations | List append-only activation/deactivation events |
| GET /api/v1/ingestion/demo-bank-statement | Generate evidence for the live missing-bank scenario |
| POST /api/v1/ingestion/manifests/{id}/verify | Verify signature/records or simulate tampering |
| POST /api/v1/ingestion/manifests/{id}/activate | Verify, audit, activate, and recompute the workspace |
| POST /api/v1/ingestion/manifests/{id}/deactivate | Audit removal and recompute downstream results |
| GET /api/v1/benchmark | Held-out baseline comparison |
| GET /api/v1/calibration | Split-conformal threshold, held-out coverage, and candidate-set sizes |
| GET /api/v1/graph/{id} | Lifecycle graph for one settlement |
| POST /api/v1/settlements/{id}/certificate | Issue certificate or return a blocking control |
| POST /api/v1/certificates/{id}/verify | Verify evidence; optionally simulate ₹1 tampering |
| POST /api/v1/controls/{id}/explain | Bounded Gemini or deterministic explanation |
Step-by-step instructions, including the free-tier behaviour that matters when recording a demo, are in docs/deployment.md.
- Backend: Dockerfile and render.yaml are included for Render.
- Frontend: deploy apps/web to Vercel and set VITE_API_BASE_URL to the deployed API URL followed by /api/v1.
- Backend CORS: set PROOFLEDGER_CORS_ORIGINS to the final frontend origin.
- Gemini: PROOFLEDGER_GEMINI_API_KEY is optional and must remain server-side.
- Database: set PROOFLEDGER_DATABASE_URL to managed PostgreSQL for durable deployment storage.
The built-in close scenario uses deterministic synthetic data inspired by public payment entity shapes; the evidence-intake lab also accepts local CSV exports. Activated imports participate in the same graph, reconciliation, controls, review queue, journal proposals, and certificates as the built-in evidence. Signed imports and activation history persist in SQLite/PostgreSQL; temporary previews, review resolutions, and issued certificates remain in memory. This project does not claim access to Razorpay production data, move money, execute refunds, post a real journal, or provide tax/legal advice. Real deployment requires encrypted storage, tenant authentication, key management, retention controls, maker-checker approval, observability, and merchant validation.
See docs/limitations.md and docs/threat-model.md before treating this as production software.
- docs/architecture.md — system boundaries and data flow
- docs/deployment.md — Render and Vercel walkthrough
- docs/control-catalog.md — all ten authoritative controls
- docs/evaluation.md — benchmark protocol and caveats
- docs/threat-model.md — abuse cases and mitigations
- docs/decisions.md — architecture decision records
PARIDHIPAWAIYA
Released under the MIT License.