A local, rate-limited and session-aware web scraper.
Config-driven RPA scraper, built as a Nest CLI monorepo (apps/api, apps/worker, libs/*). It drives a real Playwright browser through an explicit, queued pipeline — paginate a listing, open a record, extract a few DOM selectors, write a file — rather than an AI agent or a vendor scraping API. This file is enough to get it running end to end.
Every DOM selector, URL template, and on-disk filename lives in a JSON file at SITE_CONFIG_PATH, validated by SiteConfigService. Point that path at a different config and the same pipeline runs against a different site — no code changes.
- Browser — Playwright/Chromium with a persistent profile. If the target site needs a logged-in session,
npm run login:bootstrapopens it headed once; every later run reuses that session headlessly. Public sites that need no login skip this entirely (seeSITE_REQUIRES_LOGINbelow). - Queue — Kafka, via a local Redpanda broker. Each pipeline stage is its own topic; listing pagination is one message per page, so no handler blocks the consumer long enough to trigger a rebalance.
- Lock — a single Redis mutex; only one worker drives the browser session at a time.
- Storage — plain files on disk, outside this repo (
$LOCAL_STORAGE_DIR, default../scraped-data). - Ports & adapters —
QueuePort,LockPort, andStoragePortare the only seams into these systems, each with onelocalimplementation today, selected byQUEUE_PROVIDER/LOCK_PROVIDER/STORAGE_PROVIDER.
api and worker always run locally via npm, never in Docker — only Redpanda, Redis, and the Redpanda Console UI are containerized.
apps/
api/ HTTP controllers + producers (the only public surface)
worker/ Kafka consumers that actually drive the browser
Both apps group src/ by domain — listing/, detail/, titles/, filter/,
each with its own module, controller, services, dto
and specs side by side; shared helpers live in common/.
libs/
browser/ Playwright session management
site/ Configurable selectors/URLs (SiteConfigService) + DOM parsing (SiteService)
storage/ Append-only files, dedupe, HTML writers
queue/ Kafka topics, message shapes, producer wrapper
redis-lock/ Distributed mutex guarding the browser session
config/ Typed, validated env var access
-
Start the infrastructure (Redpanda + Redis + a Redpanda Console UI):
docker-compose up -d
-
Install dependencies and the Playwright browser:
npm install npm run playwright:install
-
Copy the env and site config examples, and point them at your target:
cp .env.example .env cp config/site.config.example.json config/site.config.json
Edit
config/site.config.jsonwith the target site's real selectors and URL templates (see the shape of the example file), and setSITE_BASE_LISTING_URLin.env. -
If the target site requires a login, log in once, headed, so the persistent browser profile has a session:
npm run login:bootstrap
With
SITE_REQUIRES_LOGIN=false(the default) this step is a no-op — the site is scraped without ever logging in. -
Run the API and worker (each in its own terminal):
npm run start:dev npm run start:worker
-
Kick off a crawl:
curl -X POST http://localhost:3000/scrape/listing/init -H 'Content-Type: application/json' -d '{}'
A ready-made Bruno collection for every endpoint lives in
bruno/.
Code, Kafka topics, message shapes, and endpoint paths use record/recordId and source/sources.
See .env.example for the full list — Kafka/Redis connection strings, browser profile location, data directory, pacing/retry limits, SITE_REQUIRES_LOGIN, and the *_PROVIDER adapter selectors. SITE_CONFIG_PATH points at the JSON file describing the target site's selectors and URLs (see config/site.config.example.json).
npm run test:unit # unit tests (vitest)
npm run test:e2e # end-to-end tests
npm run test:cov # coverage
npm run lint # oxlint
npm run format # prettier