Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 74 additions & 9 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,11 +96,13 @@ Source-agnostic nodes (Pydantic v2, `extra="forbid"`). Implemented in PR 2 in
`modality`. *(No `experimental_narrative`; that field is removed.)* Raw free
text (abstracts, full descriptions) is **not** stored on the node — it belongs
to the per-source `RawRecord`; the canonical node holds only normalized,
design-describing fields.
- `DatasetNode` — `dataset_id`, `data_uri`, `assay`, `cell_count`/size. Its
parent study is a typed `EXTRACTED_FROM` **edge**, not a stored foreign-key
field. *(Decision, PR 2: containment/relationships live on edges only;
duplicating them as node fields invites drift and gives two sources of truth.)*
design-describing fields. *(Target, §5: this free-text `modality` is refined
into an EFO `assay` term plus a derived coarse `molecular_layer` enum — PR 4.)*
- `DatasetNode` — `dataset_id`, `data_uri`, `assay` (to be grounded to an EFO
term ID, see §5), `cell_count`/size. Its parent study is a typed
`EXTRACTED_FROM` **edge**, not a stored foreign-key field. *(Decision, PR 2:
containment/relationships live on edges only; duplicating them as node fields
invites drift and gives two sources of truth.)*
- `SampleNode` — `sample_id`, `data_uri`, and **design covariates**: `condition`,
`perturbation`, `timepoint`, `subject`, `organism`. (Reintroduced; the prior
schema was dataset-level only.) All covariates are optional — different
Expand All @@ -118,7 +120,66 @@ Source-agnostic nodes (Pydantic v2, `extra="forbid"`). Implemented in PR 2 in
Cross-source links are *emergent*: two studies share an edge target
(`ontology_id`) rather than any source-specific key.

## 5. Coding style & architecture choices
## 5. Ontology grounding

Every facet of an experiment is bound to **one designated ontology**, never a
free-text string. This is what makes context a real metric space — the
precondition for cross-source linking and for the downstream model's shared
context space.

### Facet → ontology registry

This registry is a constant in `parce/ontology/` and the single source of truth
for "which vocabulary annotates which field".

| Facet | Ontology | Notes |
|-------|----------|-------|
| Assay / platform | **EFO** | Cross-domain assay branch (scRNA-seq, ATAC-seq, MS proteomics, …). CELLxGENE already emits EFO assay IDs. |
| Assay (upper-level / fallback) | **OBI** | Where EFO lacks a term. |
| MS proteomics specifics | **PSI-MS CV** | Instruments / acquisition; matches SDRF-Proteomics. |
| Tissue / anatomy | **UBERON** | Sample source. |
| Disease / condition | **MONDO** | Unified; prefer over bare DOID. |
| Organism | **NCBITaxon** | Already in use. |
| Cell type | **CL** | Exists but **excluded as context** (data-inferred → leakage). |
| Chemical / drug perturbation | **ChEBI** | Compound treatments. |
| Genetic perturbation | gene ID (**Ensembl**/**HGNC**) + action vocab | No clean single ontology for knockout vs. knockdown; pair gene ID with a small controlled action term. |
| Data format | **EDAM** | FASTQ/BAM/mzML/H5AD as typed terms. |

### "Modality" is two controlled fields, not one string

Avoid a free-text `modality`. Instead store, per dataset/study:

1. **`assay`** — a precise **EFO term ID** (resolved once from free text, then
stable). Fine-grained (every 10x chemistry is its own term).
2. **`molecular_layer`** — a coarse enum `{genome, epigenome, transcriptome,
proteome, metabolome, …}` **derived deterministically by walking the EFO
term's `is-a` ancestors** to a small set of anchor classes. The lineage does
the classification; we never re-string it.

The model sees a clean cross-modality categorical (`molecular_layer`) plus a
precise term (`assay`) — both controlled, no free text on either.

### Resolution (deterministic-first; see §3)

- **OLS4** (EBI Ontology Lookup Service) — one REST API to search free text →
candidate terms and to validate an ID + fetch ancestors (used for the
`molecular_layer` lineage walk).
- **text2term** / **Zooma** — batch free-text → term mapping (Zooma is well
suited to messy GEO characteristics).
- **OxO** — cross-ontology ID mapping when sources disagree (e.g. DOID → MONDO).
- **LLM** — fallback only, for strings the deterministic resolvers can't
confidently map.

Template to follow rather than reinvent: **SDRF / MAGE-TAB** (and
**SDRF-Proteomics**) already specify per-sample, ontology-annotated experiment
description and *which ontology per column* — almost exactly the `SampleNode` +
this registry. Aligning to it also yields free structured terms from sources
that already ship SDRF.

> Specific EFO/MONDO/etc. IDs are **resolved/validated via OLS at runtime**, not
> hardcoded from memory. The registry pins *ontologies*, not term IDs.

## 6. Coding style & architecture choices

- **Language/runtime:** Python ≥ 3.11, `src`-layout, `from __future__ import
annotations` everywhere, full type hints on public APIs.
Expand All @@ -139,13 +200,17 @@ Cross-source links are *emergent*: two studies share an edge target
- **Tests:** offline unit tests by default; live/credentialed tests carry the
`integration` marker and are excluded from CI.

## 6. Open questions (track, don't silently decide)
## 7. Open questions (track, don't silently decide)

- **Sample granularity for CELLxGENE.** Census is per-cell/dataset, not
per-sample in the GEO sense. Defer mapping cxg to `SampleNode` until needed;
keep it dataset-level for now.
- **Ontology resolver dependency.** text2term vs. a thin OLS REST client —
decide when implementing `ontology/` (PR4); prefer the lighter dependency.
- **`molecular_layer` anchor set.** The exact EFO ancestor classes that define
each coarse layer need pinning during PR4 (and a default for assays whose
lineage doesn't reach an anchor).
- **Ontology resolver dependency.** OLS4 REST client (needed anyway for the
lineage walk) vs. adding text2term/Zooma — prefer the lightest combination
that covers messy GEO strings; decide in PR4.
- **Graph persistence/export format** for the modeling step (per-study context +
sample manifest + URIs). Specified in a later PR.
- **Multi-omics is core, not optional.** CELLxGENE alone cannot carry the
Expand Down
27 changes: 24 additions & 3 deletions docs/ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,9 +36,13 @@ Each PR is one branch, one focused scope, green CI, and a roadmap update.
the `NarrativeOutput` schema, and the step-2 block in `main.py`). Drop
cell-type extraction. Remove the `parce.tools.*` mypy exemption as those
modules move under `sources/`. *(Next up.)*
- [ ] **PR 4 — Ontology resolver.** Shared `ontology/` stage: free-text →
UBERON/MONDO/assay IDs (deterministic OLS/text2term + on-disk cache), LLM
fallback for hard cases. Wire into normalizers.
- [ ] **PR 4 — Ontology resolver.** Shared `ontology/` stage (see ARCHITECTURE
§5). Pin the **facet → ontology registry** as a constant (EFO, UBERON, MONDO,
NCBITaxon, ChEBI, PSI-MS, EDAM). Implement free-text → term resolution
(OLS4 REST + text2term/Zooma, on-disk cache; LLM fallback) and the
**`molecular_layer` derivation** (walk EFO `is-a` ancestors to anchor classes).
Decide the anchor set + the no-anchor default. Wire into normalizers. Resolve
IDs at runtime via OLS — do not hardcode term IDs.
- [ ] **PR 5 — GEO extraction agent (vertical slice).** GEO adapter
(E-utilities/GEOparse) + Azure extraction normalizer emitting the canonical
schema via `response_format`; extract sample covariates from
Expand Down Expand Up @@ -102,6 +106,23 @@ what the next session should know. Keep entries short and factual.
format --check, mypy (16 files), **52 unit tests** pass. No dep changes.
- **Next session:** PR 3 (source-adapter interface + rip out the narrative path).

### 2026-06-23 — Ontology grounding (docs)
- Decision: every experiment facet binds to one designated ontology (EFO assay,
UBERON tissue, MONDO disease, NCBITaxon organism, ChEBI/gene-ID perturbation,
PSI-MS for MS proteomics, EDAM data format). Registry will live in
`parce/ontology/`.
- Decision: **no free-text `modality` long-term.** Store `assay` (EFO term ID) +
a coarse `molecular_layer` enum derived by walking EFO `is-a` ancestors — both
controlled. PR 2 shipped with a `modality` field; this refinement now lands in
**PR 4** (see ARCHITECTURE §4–5).
- Decision: resolve term IDs at runtime via OLS4 (+ text2term/Zooma, LLM
fallback); never hardcode IDs. Follow SDRF/MAGE-TAB conventions for the record.
- Added ARCHITECTURE §5 (Ontology grounding; later sections renumbered) and
sharpened ROADMAP PR 4 (registry + lineage derivation + anchor-set open
question). Docs only, no code change.
- Authored on the foundations branch and merged on top of PR 2 / PR 3.
**Next up: PR 3.**

### 2026-06-23 — PR 1: Foundations & tooling
- Branch `restructure-context-metadata` off `main` (post-cxg-merge).
- Decision: the LLM is repurposed from **narrative writing** to **structured
Expand Down
Loading