Skip to content

Repository files navigation

BetterCV

An AI-powered resume optimization platform. Upload your .docx resume and paste a job description — a multi-agent pipeline evaluates your fit, asks you what you actually did, rewrites weak bullets with missing JD keywords, and optionally swaps in stronger experiences from a pool.

Live: bettercv-2a86.azurewebsites.net


Features

  • Keyword gap analysis — exhaustive extraction of every JD skill missing from your resume
  • Anchored clarifying questions — before rewriting anything, the evaluator asks up to 8 questions, each tied to one specific bullet of yours: "When you migrated those services to the cloud, did you use Docker?" Your answers are treated as ground truth, so a rewrite only claims a skill you confirmed. Questions you answer "no" to are actively kept out of every later suggestion.
  • Per-bullet rewrites in two categories:
    • Missing Skills: bullets rewritten in STAR format with exact JD keywords inserted for ATS matching
    • STAR Improvements: bullets strengthened for clarity and impact without forcing keywords
  • Full-resume coverage — a second pass sweeps up any bullet the main rewrite skipped, so you get a suggestion for every weak line rather than the handful the model felt confident about
  • BetterCV Score — composite 0–100 score (40% keyword match · 35% overall quality · 25% experience relevance)
  • Experience pool swaps — add extra experiences; AI recommends 1-for-1 swaps when a pool entry scores 20+ points higher than what's on your resume
  • Chat assistant — ask for a specific change and its suggestions join the same review queue
  • In-browser review — approve or skip each suggestion with a live document preview; download the modified .docx instantly

Guardrails

Rewriting a resume with an LLM invites two failure modes, and both are handled deterministically rather than by asking the model nicely:

  • Invented experience. Every suggestion's "current text" is checked against the real resume before it reaches you. The model does occasionally invent a bullet — it has been caught lifting a line straight out of the job description and offering to rewrite it as if it were yours. Those are dropped and logged.
  • Silently dropped keywords. A JD phrase must appear in the rewrite character-for-character or it scores nothing with an ATS. When the model paraphrases one away, the pipeline re-asks, and splices the phrase in as a last resort, so a skill you confirmed never disappears without a trace.

How it works

The pipeline is deliberately split in two, with the user in the middle. Evaluation runs first and stops to ask questions; rewriting only starts once those answers come back. That pause is an ordinary HTTP round trip, which is why no graph checkpointer is needed — nothing persists between requests.

flowchart TD
    A(["User uploads .docx + pastes JD"]) --> B{Experience pool provided?}

    B -- Yes --> D["experience_optimizer_agent<br/>Scores each resume role<br/>against pool entries on JD fit"]
    D --> E(["User accepts or rejects swaps"])
    E -- Accept --> F["/apply-swaps-docx<br/>Rewrites the Word doc<br/>with pool experiences"]
    F --> C
    E -- Reject --> C
    B -- No --> C

    C["evaluate_node — POST /evaluate-resume<br/>Scores resume, extracts matching and<br/>missing skills, and writes clarifying<br/>questions anchored to real bullets"]

    C --> Q(["User answers the questions<br/>— the human-in-the-loop step —"])

    Q --> G["rate_node — POST /finalize-analysis<br/>1. Confirmed skills → one rewrite per answered bullet<br/>2. Rule A → keyword rewrite · Rule B → STAR rewrite<br/>3. Retry + splice so no confirmed skill is lost"]

    G --> S["_sweep_uncovered<br/>Second scoped call over any bullet<br/>the pass above skipped"]

    S --> V["Validation<br/>Drop suggestions whose current_text<br/>is not a real resume line"]

    V --> H(["Dashboard — BetterCV Score,<br/>skills gap, strengths and weaknesses"])
    H --> I(["Resume Preview — approve or skip<br/>each suggestion against a live doc"])
    I --> J(["Download improved .docx"])
Loading

Failed replacements are reported in the X-Replacements-Applied / X-Replacements-Failed response headers rather than being swallowed — the returned file is a valid .docx either way, so the UI warns you when a bullet could not be matched.


Tech stack

Layer Technology
Frontend React 19, TypeScript, Vite, Tailwind CSS, Radix UI
Backend Python 3.11, FastAPI, Uvicorn
AI LangGraph + LangChain agents, LiteLLM → OpenAI (REASONING_MODEL)
Documents python-docx (Word), PyMuPDF (PDF)
Package managers uv (Python), npm (Node)

Running locally

Prerequisites

Tool Version Install
Python 3.11+ python.org
Node.js 20+ nodejs.org
uv latest curl -LsSf https://astral.sh/uv/install.sh | sh
OpenAI API key — platform.openai.com

1. Clone the repo

git clone https://github.com/zxu73/resume-parser.git
cd resume-parser

2. Set up the backend

cd backend

# Create and activate a virtual environment
uv venv
.venv\Scripts\activate        # Windows
# source .venv/bin/activate   # macOS / Linux

# Install dependencies
uv pip install -r pyproject.toml

Create a backend/.env file:

OPENAI_API_KEY=sk-...          # required
REASONING_MODEL=gpt-4o-mini    # optional — any LiteLLM-compatible model

All environment variables:

Variable Required Default Purpose
OPENAI_API_KEY yes — All model calls. Must be plain ASCII — a smart quote pasted in from a doc fails much later, as an encoding error on the request header
REASONING_MODEL no gpt-4o-mini Used by both ChatOpenAI and LiteLLM
RESUME_STORE_DIR no temp dir Where uploads are kept between requests. Must be shared across instances in a multi-instance deployment
RESUME_STORE_TTL_HOURS no 24 Uploads older than this are purged at startup
JSEARCH_API_KEY no — Required by /job-search; that endpoint 500s without it
GREENHOUSE_API_KEY no — Only for /auto-apply against boards that require auth

3. Set up the frontend

cd frontend
npm install

4. Start both servers

From the project root:

make dev

This starts:

Or start them individually in separate terminals:

# Terminal 1 — backend
cd backend && uvicorn src.agent.app:app --reload --port 8000

# Terminal 2 — frontend
cd frontend && npm run dev

Then open http://localhost:5173 in your browser.


API endpoints

The two-step evaluate → finalize flow is the important part: /evaluate-resume returns clarifying_questions and nothing else actionable, and /finalize-analysis takes those answers back in to produce the rewrites.

Method Endpoint Description
POST /upload-resume Upload .docx or .pdf, returns extracted text and a doc_id / pdf_id (the UI currently offers .docx only)
POST /evaluate-resume Step 1 — scores, skills gap, and clarifying_questions. Does not rewrite
POST /finalize-analysis Step 2 — takes the answers, returns keyword_suggestions + star_suggestions
POST /analyze-experience-swaps Optimizer recommendations
POST /apply-swaps-docx Apply accepted swaps to the stored doc
GET /resume-doc/{doc_id} Serve original or swapped doc for preview
GET /resume-pdf/{pdf_id} Serve original PDF for preview
POST /download-modified-docx Apply approved rewrites, return .docx
POST /download-modified-pdf Same for PDF sources (PyMuPDF redaction)
POST /apply-suggestions Apply one rewrite to the stored doc without downloading
POST /chat Freeform assistant; its suggestions join the review queue
POST /job-search Find matching listings via JSearch — needs JSEARCH_API_KEY
POST /auto-apply Submit the stored resume to a Greenhouse job board

/job-search and /auto-apply are backend-only for now — there is no frontend consumer yet.


Project structure

resume-parser/
├── backend/
│   └── src/agent/
│       ├── app.py          # FastAPI routes, graph invocation, docx/pdf editing
│       ├── agent.py        # Prompts, LangGraph graphs, node fns, Pydantic schemas
│       ├── tools.py        # Resume extraction helpers
│       └── guidelines.md   # Bullet-rewriting rules, shared by all rating prompts
├── deploy/
│   └── azure-deploy-code.sh
└── frontend/
    └── src/
        ├── App.tsx                       # upload → analyze → dashboard state machine
        ├── components/
        │   ├── LandingPage.tsx
        │   ├── ClarifyingQuestions.tsx   # the human-in-the-loop step
        │   ├── AnalysisDashboard.tsx
        │   ├── ResumePreview.tsx         # three-pane approve/skip review
        │   ├── ExperienceManager.tsx
        │   ├── SwapReview.tsx
        │   └── ChatBot.tsx
        ├── hooks/                        # useResumeUpload · useResumeEvaluation · useExperienceSwap
        └── types/analysis.ts

Deployment

Deployed on Azure App Service via deploy/azure-deploy-code.sh. The backend serves the compiled React frontend as static files from /frontend/dist.

az login
./deploy/azure-deploy-code.sh

The script creates the resource group, App Service plan and web app if they are missing, builds the frontend, uploads the source, and configures the app — so the same command does the first deploy and every update after it. It reads OPENAI_API_KEY from backend/.env (or the environment) and sets it as an application setting; the key is never committed or baked into an artifact.

Uploaded resumes are written to RESUME_STORE_DIR (/home/data/... on App Service) rather than a temporary directory: /upload-resume and /download-modified-docx are separate requests, so the file has to survive a restart and be visible to every instance. Files older than RESUME_STORE_TTL_HOURS (default 24) are purged at startup.

A Dockerfile is also included for container hosts.

Releases

Packages

Contributors

Languages