Portfolio analytics for French institutional investors, with RAG over regulatory documents and a local LLM.
quarq computes risk metrics for a CAC 40 portfolio, answers questions from ECB, Banque de France and AMF documents with page-level citations, and writes the report narrative with a model that runs on your own machine.
General-purpose portfolio tools assume US markets and English documents. French institutional work needs different defaults:
- The benchmark is the CAC 40 and the risk-free rate is the OAT 10Y, not the S&P 500 and US Treasuries.
- The documents that matter are in French and come from the ECB, Banque de France and AMF. They need a multilingual embedder, not an English-only one.
- Every claim needs a source. An answer about systemic risk is useless to a risk committee without the document and page it came from.
- Holdings are sensitive. The narrative model should run locally, not in someone else's cloud.
quarq pulls market and macro data straight from public APIs, indexes regulatory PDFs into a local vector store, and splits LLM work between two agents: a fast one that turns metrics into prose, and a heavier one that answers research questions from the corpus.
| Portfolio metrics | CAGR, Sharpe, max drawdown, volatility, VaR 95, beta and alpha, against a configurable benchmark (default ^FCHI). Portfolios are TOML files with weighted sleeves and holdings. |
| Nine data providers | Equities via yfinance, plus FRED, ECB SDW, OECD, Eurostat, Banque de France, Euronext, AMF SFDR and CDP, all over direct REST with no framework. Responses are cached on disk. |
| Cited RAG | PDFs are chunked (512 tokens, 64 overlap), embedded with multilingual-e5-large and stored in a local ChromaDB. Retrieval has a similarity floor, and every answer returns source, page and snippet. |
| Two LLM agents | A reporting agent (default qwen3.5-9b) writes narrative on every report. A research agent (default qwen3.6-27b) answers questions on demand. Both run in LM Studio, and model names come from config. |
| Reports | HTML reports with Plotly charts (weight treemap, correlation heatmap, cumulative returns, drawdown, rolling Sharpe) and optional LLM narrative. PDF and JSON output too. |
| REST API | FastAPI server exposing metrics, RAG queries, corpus management, raw provider data and reports. |
| Open WebUI tools | Two uploadable tools that bring RAG answers and portfolio analysis into a local chat UI. |
Requires Python 3.11+ and, for the LLM features, LM Studio with a model loaded.
git clone https://github.com/yodablocks/quarq && cd quarq
pip install -e . # add '.[full]' for PDF report output
quarq status # providers, LM Studio and corpus health
quarq rag add ./docs/ # index a folder of PDFs
quarq query "Que dit la BCE sur la concentration du CAC 40 ?"
quarq report --portfolio ./demo/portfolio.toml --narrative --openThe demo portfolio is a three-sleeve CAC 40 book (Growth, Defensive, Financial) over the 12 months from September 2025 to August 2026.
A portfolio is a TOML file of weighted sleeves:
name = "CAC40 Multi-Sleeve Demo"
benchmark = "^FCHI"
start = "2025-09-01"
end = "2026-08-31"
currency = "EUR"
[[sleeve]]
name = "Growth"
weight = 0.40
[[sleeve.holding]]
ticker = "MC.PA"
weight = 0.40
# ...quarq report turns it into a self-contained HTML report. Alongside the metrics and return charts above, it shows how the holdings move together and how the weights break down by sleeve:
The building blocks are also usable as a library:
from datetime import date
from quarq.config import load_config
from quarq.ingest import get_provider
from quarq.ingest.fred import get_risk_free_rate
prices = get_provider("equity").fetch("MC.PA", date(2025, 1, 1), date(2025, 12, 31))
cfg = load_config()
rfr = get_risk_free_rate(cfg) # live OAT 10Y from FRED, e.g. 0.032Every provider returns the same shape: a DatetimeIndex with value, series_id and source columns. On failure it raises ProviderError rather than returning partial data.
flowchart LR
subgraph ingest["ingest/"]
P["9 providers<br/>yfinance, FRED, ECB,<br/>BdF, AMF, ..."]
C[(disk cache)]
end
subgraph rag["rag/"]
L["PDF loader<br/>512 / 64 chunks"]
E["multilingual-e5-large"]
V[(ChromaDB<br/>quarq_rag_v2)]
end
subgraph llm["llm/"]
R["reporting agent<br/>(fast)"]
Q["research agent<br/>(heavy)"]
end
PF["portfolio.py<br/>metrics"]
RP["report/<br/>Plotly + Jinja2"]
API["api/ FastAPI<br/>+ CLI"]
P <--> C
P --> PF --> R --> RP
PF --> RP
L --> E --> V --> Q
RP --> API
Q --> API
- Data path: providers fetch prices and rates,
portfolio.pycomputes the metrics, the reporting agent turns them into prose, andreport/renders the HTML. - Research path: PDFs are chunked with metadata (
source,doc_type,date,page,chunk_id), embedded and stored. A query takes the 20 nearest chunks above a 0.35 similarity floor, re-ranks the top 10 with a cross-encoder that reads the question and each passage together, and returns the 5 best distinct pages (one chunk per page). When the question names an exact date ("as of March 31, 2026"), documents whose manifest period covers that date come first. After the research agent answers, every figure in the answer is checked against the passages it was shown, and figures found nowhere are flagged. The research agent is prompted with the best 3 and told to answer only from them, and the answer carries citations for every retrieved chunk. - LLM selection: quarq uses LM Studio when it is reachable. If not, and
ANTHROPIC_API_KEYis set, it falls back to the Claude API. With neither, LLM features fail with a clear error and the metrics still work.
Documents get a doc_type from their filename: ecb_fsr, bdf_fsr, amf_sfdr, prospectus, factsheet, or macro by default. Filter queries on it with --doc-type.
A PDF's own date is when the file was created, not the period it covers: an annual report on 2024 is created in 2025. To record the facts, put a quarq_manifest.toml next to the PDFs:
[[document]]
source = "bdf_rapport_annuel_2024.pdf"
period_start = 2024-01-01 # required: the period the document is about
period_end = 2024-12-31
published = 2025-03-17 # optional: becomes the chunk's date
doc_type = "bdf_fsr" # optional: overrides the filename rulequarq rag add applies it when indexing, and quarq rag manifest <folder> writes it onto chunks already indexed. Unknown keys, unknown document types and reversed periods are rejected, with the document and field named.
| Command | What it does |
|---|---|
quarq status |
Provider, LM Studio and corpus health table |
quarq version |
Print version |
quarq query <question> |
Ask the RAG corpus. --doc-type, --k |
quarq rag add <path> |
Index a PDF or folder. Re-adding a file replaces its chunks |
quarq rag status |
Corpus statistics |
quarq rag coverage <path> |
Report PDF pages with no text layer (scans, image-only pages), without indexing |
quarq rag dedupe |
Remove chunks that repeat content already in the index. --dry-run |
quarq rag migrate |
Copy the previous collection (quarq_rag_v1) into the current one, without re-embedding |
quarq rag manifest <folder> |
Apply the corpus manifest to chunks already indexed, without re-embedding |
quarq config --set-lmstudio-url <url> |
Point at your LM Studio instance |
quarq serve |
FastAPI server on 127.0.0.1:8000. --host, --port, --reload |
quarq report --portfolio <toml> |
Generate a report. --format html|pdf|json, --output, --narrative, --open |
quarq eval |
Score retrieval against the gold set. --k, --doc-type-filter, --dataset, --out |
quarq eval-gen |
Draft candidate gold questions with the LLM, for human review. --per-doc-type, --seed, --out |
quarq eval checks, for each question in a human-reviewed gold set, whether retrieval returns the page that answers it. The run is deterministic and makes no LLM calls. Current results on 38 questions over a 23-document corpus (default settings: top 5, 0.35 floor), 25 September 2026:
| Retrieval | Hit@1 | Hit@3 | Hit@5 | MRR |
|---|---|---|---|---|
| Re-ranked + date-aware (default), all documents | 33 / 38 (87%) | 35 / 38 (92%) | 36 / 38 (95%) | 0.90 |
Re-ranked + date-aware (default), filtered to doc_type |
33 / 38 (87%) | 36 / 38 (95%) | 37 / 38 (97%) | 0.91 |
Embedding order only (rerank = false), all documents |
22 / 38 (58%) | 32 / 38 (84%) | 35 / 38 (92%) | 0.72 |
Embedding order only, filtered to doc_type |
23 / 38 (61%) | 34 / 38 (89%) | 36 / 38 (95%) | 0.75 |
Compared with embedding order only, re-ranking plus date-aware ordering moves 12 questions up and none down. The answer page went to first place for questions where a summary page, a neighbouring page or another edition used to win, including one that was missed entirely. When a question names an exact date, documents whose manifest period covers it come first: that fixed the last edition confusion (a June 2026 index composition ranking above the March 2026 factsheet for a March 31 question).
History on the first 28 questions: returning one result per page (instead of several chunks of the same page) raised Hit@5 from 23 to 25 and MRR from 0.71 to 0.75. The 10 later questions are year-sensitive (the same fact in the 2023, 2024 and 2025 editions) and harder, which is why the overall scores dip.
Hit@k is the share of questions whose answer page is in the top k. MRR averages 1 / rank of the first correct page. ChromaDB's search is approximate, and its settings only take effect when a collection is created: the first index, built with the defaults, left true neighbours out for 7 of the 28 questions. The current collection (quarq_rag_v2) is built with explicit settings and leaves none out. Every quarq eval run checks this: it compares the index with an exact search over the same chunks and reports how many questions it leaves true neighbours out for (today 0 of 38) and whether the final pages match exact search (38 of 38).
What the misses show:
- Editions are now kept apart. Without re-ranking, 2 of 16 year- or edition-sensitive questions ranked another edition first. The re-ranker fixes one (the 2023 annual report no longer beats the 2024 one for a 2024 figure) and date-aware ordering the other. The date rule only applies to exact dates; "in April 2025" or "at the end of 2024" leave the order alone, because the answer is often in a later document. Documents without a manifest period are never promoted.
- The right document, the wrong page. Without re-ranking, the most common miss on year questions: the correct edition came first, but through a summary or contents page. The re-ranker fixes most of these.
- Neighbouring pages win. In three ECB questions, nearby pages on the same topic (for example p112 for an answer on p113) ranked above the answer page. In one of them, the answer page still isn't in the top 5.
- Answers in footnotes lose to the main text. One ECB answer appears only in a footnote, and retrieval returned the main-text pages about the same April 2025 episode instead.
- The similarity floor never filters. Every question gets as many results as it asks for (only the 4-page factsheet set returns fewer), because retrieved chunks score far above 0.35 (about 0.8 to 0.9 in spot checks).
Caveat: 38 questions is small (one question is about 2.6 points), and they were reviewed by a single person. 28 were drafted by an LLM from the very chunks being searched, which tends to share wording with the page and flatter retrieval; the 10 year-sensitive ones were drafted from a corpus search, with every page stating the answer listed. Treat these numbers as a first baseline to compare changes against, not as expected accuracy. The gold set is in quarq/eval/datasets/quarq_gold_v1.jsonl.
Config lives at ~/.quarq/config.toml and is created with defaults on first run.
Supply secrets through the environment rather than storing them on disk:
export FRED_API_KEY=... # live OAT 10Y rate; overrides the config value
export ANTHROPIC_API_KEY=... # cloud fallback when LM Studio is unavailableRe-ranking is on by default and set in the [rag] section:
[rag]
rerank = true # false: embedding order only (faster)
reranker_model = "BAAI/bge-reranker-v2-m3" # multilingual, Apache-2.0, ~2.2 GB, downloaded on first use
rerank_top_n = 10 # candidates re-ranked per query
rerank_max_length = 512 # tokens per question + passage pair
date_aware = true # exact dates in a question favour documents covering themFRED_API_KEY is read at load time and never written back to config.toml. Without it, quarq uses the configured fallback risk-free rate (3%). ECB, OECD and the other providers need no key.
quarq ships two tool classes that expose RAG queries and portfolio analysis inside a local Open WebUI chat. Setup is in demo/OPEN_WEBUI_SETUP.md.
The tools call quarq over HTTP and default to host.docker.internal:8000, which is what resolves from inside the Open WebUI container. To run them from the host, set QUARQ_API_URL=http://127.0.0.1:8000.
Note: the tool files are uploaded into Open WebUI, not imported from this repo. After pulling changes to
demo/tools/*.py, re-upload both files in Admin → Tools, or your instance keeps running the old copies.
quarq is alpha. v0.1.0 is the first tagged release, and it has not been used in production. Known limitations:
- "Local" has exceptions. The narrative model runs on your machine, but tickers and date ranges go to Yahoo Finance and the other data APIs, and if the Claude fallback triggers, the prompt (metrics or retrieved document text) is sent to Anthropic. Leave
ANTHROPIC_API_KEYunset to keep LLM traffic local. - Retrieval misses the right page about one time in six on the first try. With re-ranking, the answer page ranks first for 33 of 38 questions and is in the top 5 for 36 (see Retrieval quality). The test set is still small.
- Scanned PDFs aren't searchable. quarq reads the PDF's text layer and has no OCR, so pages that are only images (scans, full-page photos, infographics) are not indexed.
quarq rag addwarns about them andquarq rag coveragelists them; in the current corpus that's 28 of 1,973 pages, and no document is scanned. Numbers inside charts are not searchable either. - Re-ranking costs time and memory. The cross-encoder adds about 2.5 seconds per query on an Apple GPU and about 400 MB of memory, and its first use downloads about 2.2 GB. Set
rerank = falsefor faster, less accurate retrieval. - Only figures are checked, not wording. The research agent sees the top 3 passages, each cut to 500 characters, and is told to answer from them. Every figure in its answer (counts, amounts, percentages, years) is then looked up in exactly what it was shown, and any figure found nowhere is flagged (
quarq querywarns, the API returnsunsupported_figures). On the gold set, the check accepted all 27 correct answers and caught 26 of 27 answers with one digit changed. It can't tell whether a figure is attached to the right claim: the miss was a corrupted "22%" that appears elsewhere on the same page. Sentences without figures are not checked. - The test suite is fully mocked. It needs no network, server or LM Studio, which also means it doesn't prove the live APIs still answer the same way. End-to-end checks against a live stack are manual.
- yfinance is unofficial. It scrapes Yahoo Finance and can break or rate-limit without notice.
- Not on PyPI. Install from source.
- Lint is advisory. CI reports ruff findings but doesn't fail on them yet, because the repo has existing lint debt.
experiments/ holds retrieval experiments that are not part of the package and change no default. Each one has its own README with the method, the pass conditions (written before each run), the raw results and the caveats. Results below use the 38-question gold set unless stated.
| Experiment | Question | Result |
|---|---|---|
jev_rerank |
Does TypeSafe's Jev re-rank as well as the local cross-encoder? | Hit@1 36/38 against 33/38 in the first run, the same across three wordings of the question. Jev's answers are not deterministic: rescoring the same pairs on the union pool (see recall) moved its 38-question Hit@1 by up to one question. The 36/38 here is a single run and was not repeated. Jev is a cloud API: the question and passage text leave your machine. |
recall |
Do BM25 candidates next to the embedding candidates put more answer pages in front of the re-ranker? | A round-robin union has the answer page in the first 10 chunks for 38/38 (embedding alone 36, reciprocal rank fusion 35). With Jev, Hit@1 is 38/38 in two of three passes and 37/38 in the third. Includes a one-question probe and a 15-question hard set. |
local_rerank |
Does that better candidate pool also help the local cross-encoder? | No lift at Hit@1 (33/38), though Hit@5 rises from 36 to 38. The local path is 4 to 5 questions behind Jev, which scored 38, 38 and 37 on the same pool in three passes. |
Treat these as leads, not measurements. The gold set is small (one question is 2.6 points), 28 of its 38 questions were drafted from the chunks being searched, the 15 hard questions were drafted by an LLM (a person checked the answers and pages of h-003 to h-011 against the PDFs; the four French rewrites were checked only through their English originals; and h-001 and h-002 are two of the original 38, the very questions that motivated the union merge, so the hard set is not independent of it), and a perfect score on a set that has been iterated on partly reflects saturation. The Retrieval quality numbers above describe the shipped defaults and are unchanged.
Where to pick this up. None of these is started in the repo.
- Scale test. Everything above holds on the 2,571-chunk corpus. A test with more documents was begun outside this repo, using EU legal acts from the Publications Office's Cellar service (about 17,500 in-force regulations and directives, roughly 250,000 chunks in all), and stopped. Embedding with
multilingual-e5-largeran at about 5 chunks per second on the Apple-GPU Mac used here, which overheats under sustained load, so only a step of about four times the current corpus is realistic on it. A real test needs other hardware or a cloud embedding service. None of the sources checked (ECB, Banque de France, AMF) offers a bulk API for the supervisory reports quarq indexes; their APIs cover statistics. - Local re-ranker. The local cross-encoder trails Jev by about 5 questions on the 38. A local LLM as a judge was attempted and stopped before producing a result, so the idea that the wording of the question ("does the passage state the answer" against "is it relevant") matters as much as the model is untested.
- Jev's variation. Three passes over the same pairs moved the 38-question Hit@1 by up to one question. More passes would tighten that, and the other question wordings and the embedding pool were not re-run.
- First-set questions. 28 of the original 38 questions were drafted from the chunks being searched. They have not been replaced.
- Cross-document answers. Questions whose answer needs several documents, with a source for each step, were the starting idea and are not covered by any experiment.
- Ingest-time tagging. The original plan also had Jev tag each document once at ingest (type, topic, entity, period), to filter before search and replace hand-written manifests. It was never run. The names "experiment 2" and "experiment 3" in the experiment READMEs were reused for recall and the local re-ranker; they are not that plan's experiments 2 and 3.
pip install -e '.[dev]'
pytest tests/ # offline: no network, server or LM Studio required
ruff check .CI (.github/workflows/ci.yml) runs the tests on Python 3.11 and 3.12, ruff in advisory mode, and a wheel build whose entry point is smoke-tested on every pull request and every push to master.
The Open WebUI tools are covered by contract tests that check request bodies against the real Pydantic models. To verify them against a live stack:
quarq serve &
QUARQ_API_URL=http://127.0.0.1:8000 python demo/smoke_test_tools.pyquarq/
ingest/ providers, BaseProvider, disk cache
rag/ loader, embedder, ChromaDB store, retriever, generator
llm/ LM Studio and Claude backends
report/ Plotly charts, Jinja2 template, renderer
api/ FastAPI app and routes
eval/ retrieval eval: gold set, metrics, runner, reports, draft generation
portfolio.py metrics and TOML portfolio loader
cli.py the quarq command
demo/ sample portfolio, Open WebUI tools and setup guide
experiments/ retrieval experiments, not part of the package (see Experiments)
tests/ offline test suite with mocked HTTP
MIT. Built by yodablocks.

