A free-tier-first multi-provider LLM router with a complete agent harness. Stdlib-only core (no dependencies), 173 tests, zero production traffic.
One OpenAI-compatible request, five providers behind it: OpenRouter, freeinference.org, Groq, NVIDIA NIM, Cloudflare Workers AI. When one rate-limits, errors, or hangs, the next picks up mid-request.
export GROQ_KEY=gsk_... # one key is enough to start
python src/ai_failover.py "explain KV caches"
python src/ai_failover.py --model openai/gpt-oss-120b "write a haiku"
python src/ai_failover.py --json "list 3 colors" # machine-readable outputWhat the router does for you, in order:
- Adaptive provider ordering — providers are scored by an EWMA of success rate over latency. A provider that just failed sinks below the healthy ones; a fast one rises. Registry order is the cold-start tie-break, so first-run behavior is unchanged.
- In-provider retry with backoff — a transient 500 gets up to 3 attempts
on the same provider (exponential backoff + jitter,
Retry-Afterrespected) before the router walks. A flaky provider recovers without wasting a failover hop. - 429s are never retried in-provider — the quota ledger already owns rate-limit cooldowns; retrying only burns quota. Failover instead.
- Quota ledger — every provider has a daily request budget. When it's exhausted, the router skips that provider before hitting the 429. Escalating cooldowns on rate limits: 5 min → 15 min → 1 hour.
- Multi-key rotation — paste several keys comma-separated
(
GROQ_KEY=k1,k2,k3). Dead keys (401/403) are skipped permanently; exhausted keys (429) cool down and rotate back in. Key material is never written to disk by the rotation state — only indices and statuses. - Semantic response cache — near-duplicate prompts hit a SQLite-backed TF-IDF cache (cosine ≥ 0.92, 168h TTL) and return instantly without a provider call. Stateful tool conversations skip the cache.
- Hedged requests — fire the top two providers concurrently, take the first answer, abandon the loser. Kills tail latency. (Threading-based; burns 2x quota on slow tails.)
- Deadline budgets —
route(..., timeout_budget_s=30)stops walking providers and retrying when the budget is spent. - Full attempt trails — every provider attempt emits an event with
attempt number, status, latency; every backoff emits a
retry_waitevent. The run log shows the complete routing decision path, replayable.
python src/server.py # PORT env var, default 8080A stdlib-only HTTP server. Point any OpenAI client at it:
| Endpoint | What it does |
|---|---|
POST /v1/chat/completions |
OpenAI-shaped chat with full failover behind it |
GET /health |
liveness probe |
GET /metrics |
Prometheus text metrics (all flippy_ prefixed) |
GET /usage |
per-provider calls / errors / latency / tokens + cache savings |
GET /quota |
live per-provider free-tier headroom |
Auth: set FLIPPY_AUTH_TOKEN to require Authorization: Bearer <token>.
Binds localhost by default; set HOST=0.0.0.0 to expose (and set the token).
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "hello"}]}'python -m src.loomweaver agent "check disk space with the shell tool"
python -m src.loomweaver agent "fetch example.com and summarize" --session research
python -m src.loomweaver agent "list the src directory" --model groq/openai/gpt-oss-120bA ReAct-style loop: goal → model → tool call → observation → repeat, with:
- 9 guarded tools:
http_get,read_file,write_file,list_dir,shell,remember(session facts),sql_query(SELECT-only, read-only connection),json_transform(filter/map/limit),http_post_json(JSON-validated POST). - Security guards on every tool: SSRF protection (private IPs, cloud
metadata, DNS-rebinding blocked), path jail (project root + scratch only,
credential-like paths denied), shell blocklist (25+ dangerous patterns,
env stripped of
KEY/TOKEN/SECRET/PASSWORDvars, key-redaction on output), SQL writes impossible at two independent layers, cron command allowlist.LOOMWEAVER_SAFE_MODE=1disables shell entirely. - String-aware JSON action parser — the agent understands model output wrapped in prose, markdown fences, brace-containing string values, and stray-brace noise. Measured 12/12 on realistic free-tier outputs (the old greedy-regex parser scored 10/12).
- No-progress loop detection — two consecutive steps with no tool call
and no done terminate the run with a
no_progressreason instead of burning max_steps. A tool call resets the counter. - Budgeted observations — tool output re-entering context is capped at
2000 chars with an explicit
[truncated, N more chars]marker. - Persistent sessions — facts and messages survive across runs, trimmed to 40 messages.
python -m src.loomweaver armada "audit the security module and report findings"Four role-specialized agents on one mission, each restricted to its own toolset (enforced in dispatch code, not just prompts):
| Role | Tools | Job |
|---|---|---|
| Scout | read-only: http_get, read_file, list_dir, shell | gather facts, end with FINDINGS |
| Builder | read_file, write_file, list_dir, shell, http_get | implement, end with BUILT + TESTS |
| Verifier | read_file, list_dir, shell, http_get | adversarial QA, end with VERDICT: PASS/FAIL |
| Reporter | read-only + write | write the summary |
python -m src.loomweaver eval --suite basic # 5 cases: math, caps, JSON, counting
python -m src.loomweaver eval --suite reasoning # logic + code output
python -m src.loomweaver eval --suite extraction # dates, prices, emails → JSON
python -m src.loomweaver eval --suite tools # tool-protocol emission
python -m src.loomweaver eval --suite agent # 4 multi-step agent tasks, scored on
# required tools + genuine completion
python -m src.loomweaver eval-compare # suites × models comparison tableThe agent suite measures the loop, not the model: it detects looping or never-done models deterministically (a broken planner scores 0/4; the shipped loop scores 100). Scored from the event trail, so max-steps exhaustion doesn't count as success.
Load testing and latency:
python -m src.loomweaver loadtest --provider groq --concurrency 4 --requests 8
python -m src.loomweaver ttft # streaming time-to-first-token sweep across providerspython -m src.loomweaver usage # dashboard: calls, errors, cache hits, avg latency, tokens
python -m src.loomweaver quota # per-provider free-tier headroom + cooldown state
python -m src.loomweaver providers # what's configured right nowAlso live over HTTP at /usage and /quota.
python -m src.loomweaver cron --list
python -m src.loomweaver cron --run nightly-eval
python -m src.loomweaver cron --daemon # interval loop; jobs defined in cron_jobs.jsonJobs may only invoke loomweaver subcommands (allowlist enforced) and run as subprocesses — a poisoned job file can't escalate.
pip install -r requirements.txt
python src/aihub.py --health
python src/aihub.py --chat "hello"
python src/aihub.py --vision image.jpg "what is this?"
python src/aihub.py --rag add doc.txt # build a local vector store
python src/aihub.py --rag query "question"
python src/aihub.py --summarize file.txt
python src/aihub.py --tts "text" # edge-tts voice
python src/aihub.py --stt audio.wavChat, vision, RAG (local vector store), summarization, text-to-speech, speech-to-text — routed through the same litellm provider failover.
git clone https://github.com/Rawbeew/flippy && cd flippy
# Set ONE key to start (any of: OPENROUTER_KEY, FREEINFERENCE_KEY,
# CLOUDFLARE_TOKEN + CLOUDFLARE_ACCOUNT_ID, NVIDIA_KEY, GROQ_KEY)
export GROQ_KEY=gsk_...
python src/ai_failover.py "explain KV caches in one paragraph" # chat with failover
python -m src.loomweaver agent "check disk space using the shell tool"
pip install pytest && python -m pytest tests/ -q # 173 testscp .env.example .env # add your keys
docker compose up -d
curl http://localhost:8080/health
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "hello"}]}'See ARCHITECTURE.md for the full data flow, module map, scaling points, and trade-offs table.
Request → cache check → adaptive ordering → quota check → key rotation
→ provider call (≤3 attempts w/ backoff on 5xx) → usage record
↓
failover on failure
| Module | Purpose |
|---|---|
src/flippy_providers.py |
Canonical provider registry (single source of truth) |
src/ai_failover.py |
Standalone CLI router |
src/aihub.py |
litellm-powered multimodal hub (optional: vision, RAG, TTS, STT) |
src/server.py |
stdlib HTTP server: /v1/chat, /health, /metrics, /usage, /quota |
src/loomweaver/ |
Agent harness: routing core, adaptive router policy, agent loop, armada fleet, evals, security, quota ledger, key rotation, semantic cache, usage, cron |
Read SECURITY.md for the threat model: SSRF guards, path jail,
env stripping, key redaction, cron allowlist, and LOOMWEAVER_SAFE_MODE=1.
- Zero external users. No production traffic has hit this code.
- Benchmarks are self-reported from a single machine on residential WiFi.
- Free-tier providers only — paid overflow is untested.
- The "semantic" cache is lexical TF-IDF, not embedding-based. It matches word overlap, not meaning.
- No async/await. Threading-based hedging exists but the core is synchronous.
- Single-process. No multi-worker mode; SQLite state won't survive concurrent writers at scale.
- Agent tool_calls are JSON-protocol, not native function-calling (planned).
MIT