A high-throughput, restart-safe, politeness-enforcing web crawler engine and microservice in Go.
crawler (Hermes Crawler) is a production-grade web crawling engine and microservice implemented in Go (github.com/hmza-hb/crawler). It accepts seed URLs or XML sitemaps, traverses hypermedia links under atomic resource budgets, and persists documents into PostgreSQL with HTTP ETag and Last-Modified conditional re-validation.
It operates both as an embeddable Go package and as a standalone REST API microservice daemon (cmd/crawld).
┌─────────────────────────────────────────┐
│ crawld REST API │
└──────────────────┬──────────────────────┘
│
▼
┌───────────────────────────────────────────┐
│ Crawler Worker Pool (crawl.go) │
└─────┬───────────────────────────────┬─────┘
│ │
▼ ▼
┌───────────────────────────────┐ ┌───────────────────────────────┐
│ MemoryFrontier Queue │ │ PgFrontier (SKIP LOCKED) │
└──────────────┬────────────────┘ └──────────────┬────────────────┘
│ │
└──────────────────┬──────────────────┘
│
▼
┌───────────────────────────────────────┐
│ Polite HTTP Fetcher (fetch.go) │
└──────────────────┬────────────────────┘
│
┌──────────────────────────┴──────────────────────────┐
▼ ▼
┌─────────────────────────┐ ┌─────────────────┐
│ RFC 9309 Robots Parser │ │ SSRF IP Guard │
└─────────────────────────┘ └─────────────────┘
Crawling web resources at scale while maintaining system safety and legal/politeness compliance presents non-trivial distributed engineering challenges:
- Strict Politeness Compliance: Uncontrolled crawlers risk overloading target hosts or violating site access policies.
crawlerincorporates a dedicated RFC 9309 compliantrobots.txtparser and domain-keyed token bucket rate limiting. - Re-validation & Bandwidth Efficiency: Downloading identical web pages repeatedly wastes network I/O and storage.
crawlerexecutes conditional HTTP GET requests (If-None-Match/If-Modified-Since), processing304 Not Modifiedresponses to prevent write amplification. - Restart Safety & Non-Blocking Leasing: Node crashes during long-running crawls can corrupt queue state.
crawlerprovides a PostgreSQL frontier (PgFrontier) utilizingFOR UPDATE SKIP LOCKEDfor concurrent work leasing and automatic lease recovery without external message brokers. - Zero-Trust Network Isolation: Malicious links or open redirects can trigger Server-Side Request Forgery (SSRF) against cloud infrastructure endpoints (
169.254.169.254, internal services).crawlerintercepts DNS resolutions and redirect chains, validating target IP addresses against private CIDR ranges prior to HTTP transport.
- RFC 9309 Compliant Engine: Parses wildcards (
*,$),Crawl-delay, group precedence, treats404 Not Foundas permissive, and fails closed on network errors. - Conditional GET Re-validation: Tracks HTTP
ETagandLast-Modifiedresponse metadata, avoiding body downloads when content remains unchanged. - SSRF Isolation & IP Range Guards: Re-evaluates every DNS lookup and redirect hop against private ranges (
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,169.254.169.254, loopback). - Distributed Work Leasing (
SKIP LOCKED): Multi-node URL scheduling backed by PostgreSQL row-level locks without lock contention. - Atomic Budget Enforcement: Hard bounds across total pages, retained bytes, wall-clock duration, crawl depth, and per-host page limits.
- Dual Infrastructure Modes: Embeddable Go package and standalone REST API daemon (
cmd/crawld). - Prometheus Observability: Native exporter tracking fetch latencies, queue depth, status codes, and HTTP server metrics.
The core architecture isolates concerns into four primary component abstractions:
┌─────────────────────────────────────────────────────────────────────────────────┐
│ Crawler │
│ Worker Pool Coordinator · Link Canonicalization · Budget Enforcement │
└───────┬─────────────────────────────────────────────────────────────────┬───────┘
│ │
▼ ▼
┌───────────────────────────────┐ ┌───────────────────────────────┐
│ Frontier │ │ Fetcher │
│ Memory / PostgreSQL Queueing │ │ Polite HTTP Client / SSRF │
└───────────────┬───────────────┘ └───────────────┬───────────────┘
│ │
└────────────────────────┬────────────────────────┘
▼
┌───────────────────────────────┐
│ Store │
│ Document & Metadata Storage │
└───────────────────────────────┘
Fetcher(fetch.go): Manages HTTP connection pooling, domain-keyed token bucket rate limiting (internal/platform/ratelimit), RFC 9309 parsing, and conditional GET re-validation.Frontier(frontier.go,pgfrontier.go): Manages URL traversal state (pending->leased->completed/failed).MemoryFrontierprovides an in-memory priority queue, whilePgFrontierhandles PostgreSQL persistence with row locking.Store(store.go,pgstore.go): Persistence layer storing HTTP status codes, headers, execution latency, raw/normalized SHA-256 content hashes, and extracted hypermedia links.Crawler(crawl.go): Multi-threaded worker pool coordinator executing depth traversal, link canonicalization, HTML/Sitemap parsing, and budget tracking.
→ Detailed Architecture Document
Rather than coupling the system to complex external message queues (e.g., Kafka, RabbitMQ, or Redis Streams), PgFrontier uses PostgreSQL FOR UPDATE SKIP LOCKED. Worker nodes query pending URLs concurrently without waiting for locked rows. This eliminates lock contention while reducing operational setup overhead to PostgreSQL.
When a worker claims a URL from PgFrontier, a lease timestamp (leased_until) is set. If a crawler worker process panics or dies mid-crawl, the lease automatically expires. Subsequent worker polls detect and reclaim expired leases seamlessly without manual intervention.
Rate limiting operates at the registrable domain level (via publicsuffix) rather than globally or per-IP. This ensures domain politeness (PerHostDelay) while allowing maximum parallel throughput across distinct web hosts.
To prevent Server-Side Request Forgery (SSRF), DNS resolution is performed explicitly prior to HTTP transport execution. Target IP addresses are evaluated against reserved private CIDRs, loopback addresses, and cloud provider metadata IPs (169.254.169.254). Redirect chains are re-checked at every hop.
Add the package to your go.mod:
go get github.com/hmza-hb/crawlerInitialize and run a bounded crawl:
package main
import (
"context"
"fmt"
"time"
"github.com/hmza-hb/crawler"
)
func main() {
cfg := crawler.DefaultConfig()
store := crawler.NewMemoryStore()
frontier := crawler.NewMemoryFrontier(cfg)
fetcher, err := crawler.NewHTTPFetcher(cfg, crawler.Options{Store: store})
if err != nil {
panic(err)
}
loop, err := crawler.New(fetcher, frontier, store, cfg, crawler.Options{})
if err != nil {
panic(err)
}
res, err := loop.Crawl(context.Background(), crawler.CrawlOptions{
Seeds: []string{"https://example.com"},
FollowLinks: true,
Budget: crawler.Budget{
MaxPages: 50,
MaxDuration: 2 * time.Minute,
},
})
if err != nil && err != crawler.ErrCrawlStopped {
panic(err)
}
fmt.Printf("Fetched: %d, Not Modified: %d, Stopped By: %s\n",
res.Fetched, res.NotModified, res.Stats.StoppedBy)
}make docker-upexport DATABASE_URL='postgres://postgres:postgres@localhost:5432/crawler?sslmode=disable'
export HTTP_ADDR=':8080'
go run ./cmd/crawld| Method | Endpoint | Description |
|---|---|---|
POST |
/v1/crawl |
Trigger a bounded, multi-page crawl run |
POST |
/v1/fetch |
Politely fetch a single URL with persistence |
GET |
/v1/documents/{id} |
Retrieve stored document metadata by UUID |
GET |
/v1/documents?url={url} |
Query the last stored document for a target URL |
GET |
/v1/hosts/{host}/robots |
Inspect cached robots.txt rules and crawl delay |
GET |
/v1/stats?run_id={id} |
Retrieve queue depth and run statistics |
GET |
/healthz, /readyz |
Liveness and database connectivity readiness probes |
GET |
/metrics |
Prometheus metrics endpoint |
curl -X POST http://localhost:8080/v1/fetch \
-H 'Content-Type: application/json' \
-d '{"url": "https://example.com"}'{
"outcome": "fetched",
"url": "https://example.com/",
"status": 200,
"status_text": "OK",
"content_type": "text/html; charset=UTF-8",
"content_hash": "a67f3...e3b1",
"etag": "\"314b5b701509a25b3992070e6508d519\"",
"last_modified": "2024-01-30T10:00:00Z",
"not_modified": false,
"filtered": false,
"body_truncated": false,
"bytes": 1256,
"elapsed_ms": 142
}The service and library are configured via Config or environment variables:
| Environment Variable | Description | Default |
|---|---|---|
DATABASE_URL |
PostgreSQL connection DSN | Required for crawld |
HTTP_ADDR |
Listen address for API microservice | :8080 |
CRAWLER_USER_AGENT |
HTTP User-Agent header sent upstream |
HermesCrawler/1.0 (+https://github.com/hmza-hb/crawler) |
CRAWLER_CONCURRENCY |
Concurrent worker goroutines per run | 10 |
CRAWLER_PER_HOST_DELAY |
Minimum delay between requests to same host | 250ms |
CRAWLER_MAX_BYTES |
Maximum allowed HTTP response body size | 10485760 (10MB) |
CRAWLER_MAX_RETRIES |
Retry attempts for transient failures | 2 |
CRAWLER_ALLOW_PRIVATE_HOSTS |
Override SSRF guard to permit loopback/private IPs | false |
GET /healthz— Service liveness probe returning HTTP 200.GET /readyz— Database readiness probe executing a ping against PostgreSQL.
GET /metrics exposes standardized operational metrics:
crawler_fetch_duration_seconds: Histogram measuring HTTP fetch latency.crawler_fetched_documents_total: Counter by HTTP status code and outcome.crawler_frontier_depth: Gauge of current pending URLs in queue.
make testgo test -v -race ./...make test-coverageExecute Go micro-benchmarking suite for robots parsing, URL normalization, and frontier queueing:
go test -bench=. -benchmem ./....
├── cmd/
│ └── crawld/ # REST API microservice entrypoint & CLI flags
├── internal/
│ └── platform/ # Platform infrastructure (db pool, httpx server, observe, ratelimit)
├── robotstxt/ # Custom RFC 9309 robots.txt parser engine
├── docs/
│ ├── architecture.md # Deep-dive system design document
│ └── openapi.yaml # OpenAPI 3.0 REST API specification
├── migrations/ # Embedded PostgreSQL migrations (`001_init.sql`, etc.)
├── docker-compose.yml # Local development stack (Postgres + Service)
├── Dockerfile # Multi-stage production container build
├── Makefile # Build, test, and container targets
├── config.go # Core configuration structures & validation
├── crawl.go # Crawler worker loop & execution manager
├── fetch.go # Polite HTTP fetcher & security guards
├── frontier.go # In-memory frontier queue implementation
├── pgfrontier.go # PostgreSQL SKIP LOCKED frontier implementation
├── store.go # Storage interfaces & memory storage
└── pgstore.go # PostgreSQL document & run persistence
- No JavaScript Execution:
crawlerparses static HTML and XML sitemaps. Single-page applications (SPAs) requiring client-side JS rendering are not executed. - Single-Process Memory Mode:
MemoryFrontieroperates in-memory for standalone single-process usage. Multi-node distributed deployment requiresPgFrontierbacked by PostgreSQL.
- Dynamic adaptive per-host crawl delay based on server response latency.
- Pluggable Object Storage (AWS S3 / GCS) backend for raw HTML content bodies.
- JSON-LD and Schema.org metadata extraction pipeline.
Contributions are welcome! Please follow these guidelines:
- Open an issue describing the proposed bug fix or feature.
- Ensure all changes include unit tests covering new behavior.
- Verify that
make testandmake lintpass before submitting pull requests.
MIT © 2026 Hamza & Contributors