Autonomous, full-stack AI research platform for literature analysis, experiment planning, problem ideation, gap detection, citation graph intelligence, and benchmark discovery.
π Live Platform Demo β’ π» GitHub Repository β’ π API Reference β’ π System Architecture
PaperLens AI is an autonomous, full-stack AI research orchestrator designed to transform unstructured academic literature into structured, actionable research outputs. Unlike basic single-prompt document chatbots, PaperLens AI operates as a multi-capability scientific assistant featuring a dual-pipeline RAG architecture (in-memory BM25 + FAISS hybrid search for instant single-session analysis, paired with remote Supabase pgvector persistence for cross-session synthesis), a multi-provider fallback engine, real-time Server-Sent Events (SSE) citation tracking, and an autonomous ReAct agent loop with Model Context Protocol (MCP) server support.
The system is engineered specifically to run under tight production constraints (such as Render's 500MB free-tier memory cap) through generator-based document parsing, lazy-loaded embedding models, and token-optimized prompt design.
PaperLens AI incorporates systematic LLM orchestration optimizations validated by automated test suites (benchmark_token_savings.py, test_agent_architecture.py, test_call_consolidation.py, and test_structured_outputs.py):
| Optimization Benchmark | Measured Performance Gain | Technical Implementation |
|---|---|---|
| Deterministic Fast-Path Router | 1.5s Latency Savings (~600 tokens/query) | Direct keyword pattern matcher skips the LLM router call for single-intent queries (e.g. dataset lookup, literature search). |
| LLM Call Consolidation | 54.5% API Call Reduction | Consolidated dual-pass synthesis & critique into single LLM passes; batched multi-chunk summarization. |
| Turn-Based System Prompt Compression | ~40% Context Token Reduction | Turn 1 sends full tool JSON schemas; Turns 2β6 automatically compress tools into signature representations (ACTIVE TOOLS: search_papers(domain, limit)). |
| Memory-Safe Extraction | 0MB Heap Bloat (500MB Cap Compliant) | Generator-based PyMuPDF stream parsing combined with lazy-loaded SentenceTransformer vector models. |
| Structured Output Enforcement | 0 Parsing Retries / 0 Regex Hacks | Strict Pydantic v2 schemas (ReActDecision, SynthesisAndCritiqueResult) + XML tags (<summary>, <limitations>, <future_work>). |
| 4-Stage Citation Matching Resilience | <1% Missing Citation Rate | Automatic 4-stage search fallback (DOI |
flowchart TD
subgraph ClientLayer ["π» Client Layer (React 18 + TypeScript + Vite)"]
UI["React Dashboard UI (Tailwind CSS + Framer Motion)"]
Auth["Clerk JWT Authentication"]
end
subgraph APIGateway ["β‘ Backend Gateway (FastAPI + Async Uvicorn)"]
JWKS["Clerk RSA-256 JWKS Validator"]
Limiter["Async Token-Bucket Rate Limiter"]
Routes["REST & SSE Route Handlers (/api/*)"]
end
subgraph AgentEngine ["π€ Autonomous ReAct Agent Loop (backend/app/services/agents/)"]
Router{"Fast-Path Router\n(Keyword vs LLM Intent)"}
ReactLoop["ReAct Execution Loop & Scratchpad Memory"]
TurnCompress["Turn-Based System Prompt Compressor"]
PydanticGuard["Pydantic & XML Output Guardrails"]
end
subgraph CapabilitiesSubsystem ["π Capability Engines & Scoped Tools"]
PaperAnalyzer["π Paper Analyzer (In-Memory BM25/FAISS + pgvector)"]
ExpPlanner["π§ͺ Experiment Planner (Roadmap Generator)"]
ProblemGen["π‘ Problem Generator (Ideation & Brief Expansion)"]
GapDetect["π Gap Detection (Methodological Flaw Scoring)"]
DatasetFinder["π Dataset & Benchmark Finder"]
CitationIntel["π Citation Intelligence (SSE + 4-Stage Fallback)"]
MCPProtocol["π Model Context Protocol (MCP) Server"]
end
subgraph InfrastructureLayer ["βοΈ Infrastructure & External APIs"]
Groq["Groq LLM Engine\n(llama-3.1-8b / gpt-oss-120b)"]
ModelFallback["Multi-Model Resilience & Fallback Router"]
Supabase["Supabase PostgreSQL + pgvector"]
AcademicAPIs["Semantic Scholar / Crossref / arXiv APIs"]
end
UI -->|Bearer JWT| Auth --> JWKS --> Limiter --> Routes
Routes --> Router
Router -->|Single Intent| CapabilitiesSubsystem
Router -->|Open-Ended Task| ReactLoop
ReactLoop --> TurnCompress --> PydanticGuard --> CapabilitiesSubsystem
CapabilitiesSubsystem --> ModelFallback --> Groq
CapabilitiesSubsystem --> Supabase
CapabilitiesSubsystem --> AcademicAPIs
Each capability is built as a modular domain engine accessible via REST endpoints, SSE streams, or the autonomous ReAct agent loop:
-
π Paper Analyzer: Dual RAG architecture providing instant in-memory BM25 + FAISS single-session Q&A, plus remote Supabase
pgvectorchunking (tiktoken) for persistent Map-Reduce document summarization. - π§ͺ Experiment Planner: Generates structured, step-by-step 6-phase research roadmaps with parameter recommendations, baseline configurations, and risk assessments.
- π‘ Problem Generator: Two-stage ideation engine that identifies novel research problems in a target domain and expands surface ideas into comprehensive methodology briefs.
-
π Gap Detection: Analyzes manuscripts or text inputs for methodological flaws, unstated assumptions, missing literature, and severity scores (Low/Medium/High) on a pinned lightweight LLM route (
llama-3.1-8b-instant). - π Dataset & Benchmark Finder: Matches research problems with standard datasets, evaluation benchmarks, evaluation metrics, and baseline models.
-
π Citation Intelligence: Evaluates paper bibliographies via a 4-stage fallback matcher (DOI
$\rightarrow$ Exact$\rightarrow$ Title$\rightarrow$ Loose), streams real-time matching progress over Server-Sent Events (SSE), and computes prioritized reading paths. - π€ Agent Mode: Flagship autonomous multi-agent orchestrator implementing a turn-compressed ReAct loop, task-scoped tools, deterministic fast-path routing, and native Model Context Protocol (MCP) server integration.
| Layer | Primary Technologies | Architecture & Rationale |
|---|---|---|
| Frontend UI | React 18, TypeScript, Vite, Tailwind CSS, shadcn/ui, Framer Motion | Fully typed component architecture with dynamic state drawers, interactive execution graphs, and SSE streaming handlers. |
| Backend Framework | Python 3.10+, FastAPI, Uvicorn, Pydantic v2, Asyncio | Async ASGI server delivering high concurrency, automatic OpenAPI documentation, and strict request/response data contracts. |
| Auth & Security | Clerk JWT, RSA-256 JWKS Verification, Custom Token Bucket Rate Limiter | Stateless token validation via Clerk public keys; per-user sliding-window rate limiting protecting LLM & database endpoints. |
| Database & Persistence | PostgreSQL, Supabase pgvector, SQLAlchemy 2.0 ORM, Alembic |
Relational tracking for user activity & documents combined with remote vector embeddings (paper_chunks) and Alembic migration control. |
| RAG & Vector Search | PyMuPDF (fitz), rank_bm25, faiss-cpu, all-MiniLM-L6-v2, tiktoken |
Hybrid lexical (BM25) + dense vector (FAISS/pgvector) retrieval; generator-based parsing preventing memory bloat on massive PDFs. |
| LLM & Fallback Engine | Groq Cloud API (llama-3.1-8b-instant, openai/gpt-oss-120b, llama-3.3-70b) |
Ultra-low latency inference with custom per-attempt model fallback routing for maximum API uptime and cost control. |
- Python 3.10+
- Node.js 18+ & npm 9+
- Supabase Project (with
pgvectorextension enabled) - Clerk, Groq Cloud, and Semantic Scholar API keys
Run the SQL DDL script from backend/supabase_migration.sql in your Supabase SQL Editor to establish the paper_chunks schema and match_chunks RPC vector search function.
cd backend
python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txtCreate backend/.env:
DATABASE_URL=postgresql://postgres:password@db.supabase.co:5432/postgres
SUPABASE_URL=https://your-project.supabase.co
SUPABASE_KEY=your-supabase-service-key
CLERK_SECRET_KEY=sk_test_...
GROQ_API_KEY=gsk_...
SEMANTIC_SCHOLAR_API_KEY=your-keyRun the development server:
uvicorn app.main:app --reload --port 8000cd frontend
npm installCreate frontend/.env.local:
VITE_CLERK_PUBLISHABLE_KEY=pk_test_...
VITE_API_URL=http://localhost:8000Run the client:
npm run devFor technical interviewers, hiring managers, and system contributors seeking in-depth architectural and design analysis:
- π System Architecture Guide (
docs/ARCHITECTURE.md) β Complete multi-layer architecture, sequence diagrams, dual RAG pipeline execution, and ReAct agent state machine. - π‘ Design Decisions & Engineering Trade-Offs (
docs/DESIGN_DECISIONS.md) β Rationale behind FAISS vs pgvector, ReAct agent loops vs static DAGs, Groq inference, and Pydantic/XML constraints. - π‘οΈ Security & Performance Optimization Audit (
docs/SECURITY_PERFORMANCE.md) β In-depth breakdown of Clerk JWKS authentication, token-bucket rate limiting, LLM token benchmarks, and memory profiling. - π» Detailed Tech Stack & Dependencies (
docs/TECH_STACK.md) β Detailed rationale for technology choices, framework comparisons, and complete dependency tree analysis. - π Database Schema & Vector Architecture (
docs/DATABASE.md) β Entity-relationship diagrams, PostgreSQL schemas, Supabase pgvectorpaper_chunksDDL, RPC matching algorithms, and Alembic migrations. - π OpenAPI-Grade API Reference (
docs/API_REFERENCE.md) β Complete specification of all REST endpoints, SSE streaming endpoints, MCP server protocols, request/response schemas, and status code matrix. - π Production Deployment & DevOps Guide (
docs/DEPLOYMENT.md) β Setup instructions for Render (backend 500MB RAM tier optimization), Vercel (frontend), environment variable security, and production migrations.
Distributed under the MIT License. See LICENSE for more details.
Built by Arpan Pramanik.