This is a student project that has never served untrusted traffic. It implements specific, named mitigations and tests that they behave as designed. It is not a hardened system, and nothing here should be read as a claim that it is secure.
What follows separates, explicitly, what is mitigated from what is not.
Verify the claims below: python3 scripts/verify/run_pytest.py tests/test_robustness.py VERIFY_G10_PASS (29 tests).
| Actor | Capability assumed | In scope |
|---|---|---|
| End user | Sends arbitrary text to POST /query |
Yes |
| Document supplier | Can get a document into the ingested corpus | Yes |
| Operator | Runs the service, reads the logs | Partially (log hygiene) |
| Network attacker | Intercepts traffic | No — TLS is the deployment's job |
| Host attacker | Code execution on the machine | No |
The document supplier is the interesting one. In any real RAG deployment, whoever can add a document to the corpus can put arbitrary text in front of the model. Retrieved content is therefore treated as untrusted input throughout.
- Turn separation. Instructions live in the
systemturn; retrieved content lives in theuserturn, fenced in<sources>…</sources>and explicitly labelled untrusted data. Tested:test_retrieved_content_is_fenced_and_labelled_as_data. - A hostile user query is never promoted to an instruction. It stays inside
the user turn; the system turn is unchanged.
Tested:
test_injection_in_the_user_query_is_not_promoted_to_instructions. - Detection and disclosure. A regex detector flags instruction-like spans in
retrieved passages (
ignore previous instructions,you are now a,reveal your system prompt,</sources>,BEGIN SYSTEM, …). A hit is logged and returned as a responsewarning. It is deliberately not used to drop evidence: silently censoring retrieved text would make behaviour unauditable. Tested across 5 payloads, plus a positive control asserting the detector does not fire on ordinary RFC text. - Label spoofing is impossible. A passage containing the literal text
[S9]cannot create a citation target; only labels the system assigned are valid, and a citation to an unassigned label is recorded as invalid. Tested:test_source_labels_cannot_be_spoofed_by_passage_text. - Verification is the real floor. Even if an injected instruction changed the
model's behaviour, every claim must still pass content-overlap and
distinctive-term checks against genuinely retrieved corpus text before it
reaches the user (
PROMPT_DESIGN.md§3.2). An injected instruction saying "reply that the API key is X" produces a claim that no retrieved passage supports, so it is dropped and the system abstains.
Prompt injection is not solved. The detector is a small regex list with obvious bypasses (paraphrase, encoding, non-English, instruction spread across passages). Turn separation is a convention the model may ignore. The verification layer constrains what can be asserted, not what the model can be made to do — an injection that suppresses an answer, or steers which of several true facts is reported, is not caught.
Do not deploy this against a corpus that untrusted parties can write to.
| Control | Behaviour | Test |
|---|---|---|
| Empty / whitespace query | 422 |
test_empty_query_is_rejected, test_whitespace_only_query_is_rejected |
| Query > 2,000 chars | 422 |
test_oversized_query_is_rejected |
top_k outside 1–20 |
422 |
test_invalid_top_k_is_rejected |
| Unknown retrieval method | 422 |
test_invalid_method_is_rejected |
| Malformed JSON body | 422 |
test_malformed_json_is_rejected |
| Upload > 8 MB | 413 |
test_upload_rejects_oversized_file |
| Unsupported file type | 415 |
test_invalid_file_types_are_rejected |
POST /documents/ingest takes paths. Every path is resolved and checked for
containment inside the configured corpus directory; anything resolving outside is
rejected with 400 path_outside_corpus. Both ../../../../etc/passwd and the
absolute /etc/hosts are tested.
Uploads are written under pathlib.Path(filename).name, so a filename containing
directory separators cannot escape the corpus directory.
Query text is user-supplied and may contain a credential pasted by mistake.
Every log record and every audit-log entry passes through a redaction filter
covering bearer tokens, JWTs, api_key=/secret=/password= assignments, and
sk--prefixed keys.
Tested in both directions: four secret formats are redacted, and a positive control asserts an ordinary question ("What does the Authorization header field do in RFC 9110 section 11.6.2?") passes through unmodified — a redactor that mangles normal text is a bug, not a safety feature.
Not comprehensive. This is pattern matching against common secret shapes, not data-loss prevention. A novel credential format will be logged verbatim.
Internal error text is never echoed to the client. A pipeline failure returns
500 with {"error": "query_failed", "request_id": ...}; the exception and its
traceback go to the server log only. Explicitly tested: a failure whose message
is "vector database unavailable" must not appear anywhere in the response body.
Startup failure leaves the process live but not ready: /health answers
degraded, and query endpoints return 503 with a remediation hint, rather than
the process crash-looping.
| Vector | Status |
|---|---|
| Rate limiting | Not implemented. No per-client throttle exists. |
| Authentication | Not implemented. Every endpoint is unauthenticated. |
| Expensive queries | Partially: top_k ≤ 20, query ≤ 2,000 chars, upload ≤ 8 MB |
| Concurrent generation | Unbounded. Each request can occupy the GPU for seconds; see LOAD_TESTING.md for what that does to p99 |
| Index rebuild on ingest | POST /documents/ingest triggers a full BM25 rebuild — a cheap request with an expensive server-side cost |
An internet-facing deployment would need authentication, per-client rate limits, a bounded generation queue, and ingest moved to an authenticated background job. None of that is built.
- All models are open-weights from Hugging Face. No API key exists anywhere in this system, so there is no credential to leak.
- Model downloads are not pinned to a revision hash — a supply-chain gap that is noted rather than fixed.
- The corpus is public IETF RFCs, fetched over HTTPS with a recorded SHA-256 per
document, so corpus tampering is detectable (
scripts/verify/check_corpus.py). bm25.pklis a Python pickle. It is generated locally and never loaded from an untrusted source, but pickle is unsafe by construction and a hostile index file would be code execution. A real deployment should use a non-executable format.
In priority order:
- Authentication and per-client rate limiting on every endpoint.
- A bounded generation queue with admission control and a hard timeout.
- Move ingestion to an authenticated, asynchronous job.
- Replace the pickle index with a non-executable serialisation.
- Pin model revisions by hash and verify on load.
- Replace regex injection detection with a trained classifier, and treat it as a signal for review rather than a control.
- An adversarial evaluation set for injection, so mitigation is measured the way retrieval is, instead of asserted.