Open-source cloud pricing API for C3X. Scrapes pricing data directly from AWS, Azure, and Google Cloud public APIs and serves it via a GraphQL endpoint.
At a glance: one Go binary, three subcommands (serve, scrape, seed),
PostgreSQL-backed, versioned schema, production-grade middleware (rate limiting,
auth, CORS, panic recovery, AST-based GraphQL validation). See
ARCHITECTURE.md for the ten-minute tour, and
docs/provider-quirks.md for the hard-won tribal
knowledge about each upstream pricing API.
# Clone and configure
git clone https://github.com/c3xdev/c3x-pricing-api.git
cd c3x-pricing-api
cp .env.example .env
# Edit .env, set POSTGRES_PASSWORD
# Start with Docker Compose
docker compose up -d
# Scrape pricing data: each vendor in its own container, with enough
# memory for AWS (3 GB by default; see deploy/scrape.sh)
deploy/scrape.sh # aws, azure and gcp
deploy/scrape.sh azure # or a subsetDon't scrape with docker compose exec api ...: that runs inside the API
container under its 1 GB limit, where an AWS scrape runs out of memory and
can take the API down with it. To scrape nightly, add to your crontab:
0 3 * * * /path/to/c3x-pricing-api/deploy/scrape.sh >> /var/log/c3x-pricing-scrape.log 2>&1.
The API will be available at http://localhost:4000/graphql.
export C3X_SELF_HOSTED=true
export C3X_PRICING_API_ENDPOINT=http://localhost:4000
c3x estimate --path /path/to/terraform| Endpoint | Method | Description |
|---|---|---|
/graphql |
POST | GraphQL pricing queries (supports batched requests) |
/status |
GET | Scrape status and product counts per vendor (JSON). Per-vendor status is one of ready, empty, failed, stale, scraping, never; see below. |
/healthz |
GET | Liveness probe, no DB dependency |
/readyz |
GET | Readiness probe, pings the DB |
/health |
GET | Backwards-compatible alias for /readyz |
/status reports one state per vendor, ordered by urgency so the most
actionable one wins when several apply:
| status | meaning |
|---|---|
ready |
fresh data from a run that actually ingested products |
empty |
the last successful run ingested nothing, so it produced no usable data despite being recorded a success |
failed |
the most recent run failed, so the served data is frozen at the previous run |
stale |
nothing has finished in 48 hours |
scraping |
a run is in flight |
never |
no scrape has ever run for this vendor |
empty and failed matter because a vendor with a revoked credential
scrapes nothing while the API keeps serving the rows it already had. The
data stays available and correct-as-of-its-last-good-scrape, but it stops
advancing, and the status is what tells you so.
See deploy/ for opinionated recipes:
deploy/compose/: local dev with a cron sidecar for scheduled scrapes.deploy/k8s/: plain-manifest Kubernetes deployment with per-vendor CronJobs.deploy/github-actions/: run the scraper on GitHub's free schedule against a hosted Postgres.
Scrapes are one-shot CLI invocations, guarded by a Postgres advisory lock so
overlapping runs are safe. Freshness is tracked per vendor in the scrape_runs
table.
| Variable | Default | Description |
|---|---|---|
DATABASE_URL |
required | PostgreSQL connection string |
PORT |
4000 |
HTTP server port |
API_KEY |
(empty) | API key for authentication. Empty = no auth. |
GCP_API_KEY |
(empty) | Google Cloud API key for scraping GCP prices |
SCRAPE_CONCURRENCY |
4 |
Global concurrent scrape workers |
SCRAPE_CONCURRENCY_AWS |
0 |
AWS-specific override (0 = inherit global) |
SCRAPE_CONCURRENCY_AZURE |
8 |
Azure-specific override |
SCRAPE_CONCURRENCY_GCP |
0 |
GCP-specific override (0 = inherit global) |
MAX_REQUEST_BODY_MB |
4 |
Maximum request body size in MB |
MAX_BATCH_SIZE |
50 |
Maximum number of queries per batch request |
MAX_PRODUCTS_PER_REQUEST |
5000 |
Products returned per HTTP request, summed over all products fields (aliases) and batch items. A field's limit is clamped to what is left; once spent, further fields error |
MAX_PRODUCT_QUERIES_PER_REQUEST |
50 |
products fields (each one a DB query) per HTTP request, across aliases and batch items |
MAX_INFLIGHT_REQUESTS |
0 |
Concurrent /graphql requests. 0 = DB pool max minus 2 (min 1), -1 = unbounded |
INFLIGHT_WAIT_MS |
250 |
How long a request waits for an in-flight slot before 503 + Retry-After |
QUERY_TIMEOUT_SECS |
30 |
Query execution timeout in seconds |
MAX_QUERY_DEPTH |
10 |
Maximum GraphQL query nesting depth |
RATE_LIMIT_PER_SEC |
100 |
Maximum requests per second per IP |
CORS_ALLOWED_ORIGINS |
(empty) | Comma-separated allowed CORS origins. * = all |
TRUSTED_PROXIES |
(empty) | Comma-separated CIDRs/IPs whose CF-Connecting-IP / X-Forwarded-For is trusted for the real client IP (used by rate limiting and metrics). Use the literal cloudflare to trust Cloudflare's published edge ranges. Behind Cloudflare, set this or rate limiting keys on the Cloudflare edge IP. |
DISABLE_INTROSPECTION |
true if ENV=production, else false |
Block GraphQL introspection queries (__typename is always allowed) |
DB_MAX_CONNS |
0 |
Max database pool connections (0 = pgx default) |
DB_MIN_CONNS |
0 |
Min database pool connections (0 = pgx default) |
METRICS_ADDR |
127.0.0.1:9090 |
Listen address for /metrics, which is never served on PORT. off disables it. Legacy METRICS_PORT=<p> still works and means :<p> |
CNY_USD_RATE |
6.2069 |
CNY/USD exchange rate for AWS China pricing |
# Scrape all vendors
./c3x-pricing-api scrape --vendor aws
./c3x-pricing-api scrape --vendor azure
./c3x-pricing-api scrape --vendor gcp
./c3x-pricing-api scrape --vendor all
# Seed from JSON files (for testing)
./c3x-pricing-api seed --file data/products.jsonAWS and Azure pricing data is publicly available without authentication. GCP requires a free API key. Get one from Google Cloud Console.
- Rate limiting per IP address
- Request body size limits
- Query execution timeouts
- Timing-safe API key comparison
- Graceful shutdown on SIGTERM
- Database health checks
- SHA-256 hashing (no MD5)
- No hardcoded credentials