Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion .scripts/ci/download_data.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,13 +9,19 @@

import argparse

_CNT = 0 # increment this when you want to rebuild the CI cache
_CNT = 1 # increment this when you want to rebuild the CI cache


def main(args: argparse.Namespace) -> None:
"""Call each loader once, so its files and anything it assembles from them land in the cache."""
import mantispy as mt

def jump_lite_all() -> None:
"""Every feature set, because tutorial 12 compares them and each is a separate file."""
for model in mt.ds.JUMP_LITE_MODELS:
mt.ds.jump_lite(model=model)
mt.ds.jump_lite_targets()

loaders = {
"bbbc021": mt.ds.bbbc021,
"rohban": mt.ds.rohban,
Expand All @@ -25,6 +31,7 @@ def main(args: argparse.Namespace) -> None:
"jump_target2": mt.ds.jump_target2,
"jump_cells": mt.ds.jump_cells,
"jump_plate": mt.ds.jump_plate,
"jump_lite": jump_lite_all,
"oasis_pilot": mt.ds.oasis_pilot,
}
if args.dry_run:
Expand Down
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,3 +22,10 @@ and this project adheres to [Semantic Versioning][].
- `mantispy.io`: `stamp`, which puts an `AnnData` built elsewhere — a published h5ad, another pipeline's output, a matrix of learned embeddings — on the mantispy API surface
- `mantispy.metrics`: `known_relationships`, the share of annotated perturbation pairs whose similarity falls in either tail of the distribution over all pairs, and `evaluate_correction(covariates=...)`, which reports what a representation spends its variance on besides the batch and the label
- `mantispy.pp`: `tvn`, typical variation normalization with per-batch CORAL, which aligns each batch's controls onto the pooled controls
- `mantispy.ds`: `jump_lite`, the same 1,536 JUMP Target-2 wells embedded by five models and measured by `cp_measure`, and `jump_lite_targets`, the gene each compound is annotated to act on

### Fixed

- `mantispy.io`: `cp_measure` column names are read as `<object>_<channel>/<aggregation>/<group><Feature>` rather than through the CellProfiler grammar, which left `var['channel']` empty and split one feature group into as many as the channels and aggregations it was written with
- `mantispy.pp`: `well_qc` says it expects cell resolution, instead of counting one row per well and failing every well on a well-level object
- `mantispy.pp`: `feature_select` warns when it selects nothing, rather than leaving an empty matrix for whatever runs next; `noise_removal`'s `stdev_cutoff` is documented as an absolute threshold on the scale `normalize` left the values on
7 changes: 5 additions & 2 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -323,6 +323,8 @@ Plotting functions return Matplotlib axes and do not modify the object.
ds.amish
ds.chroma
ds.jump_crispr
ds.jump_lite
ds.jump_lite_targets
ds.luad
ds.miami
ds.neuropainting
Expand Down Expand Up @@ -353,7 +355,7 @@ replicate; the other eight are the strongest movers among the wells that survive
`synthetic_plate` and `blobs` are generated locally. `synthetic_plate` is a single-cell profile table with injected
artifacts for quality control to find; `blobs` is a small `SpatialData` plate of images, labels and tables. The other
datasets download once, checked against a pinned sha256, into `mt.settings.cache_dir` (set `MANTISPY_CACHE_DIR` to
change it). Four of them carry the annotations the analysis functions need:
change it). Six of them carry the annotations the analysis functions need:

| dataset | download | perturbations | carries |
|---|---|---|---|
Expand All @@ -362,8 +364,9 @@ change it). Four of them carry the annotations the analysis functions need:
| `pki` | ~71 MB | 15 kinase inhibitors x 7 doses | cell counts, MOA labels, 32-64 replicates |
| `jump_target2` | ~0.7 GB | 302 compounds, one shared plate map | the same plate run at eleven sites, so any difference between them is technical |
| `jump_cells` | ~1.5 GB | 12 compounds and DMSO | single cells, 24 wells x 4 fields of view of `BR00121438` |
| `jump_lite` | ~10 MB per feature set | 302 compounds, four laboratories | the same 1,536 wells under five learned embeddings and `cp_measure`, so the feature set is the only thing that changes; `jump_lite_targets` gives the gene each compound acts on |

The other nine are further gallery accessions, normalized and feature-selected by their authors and read with the
The other ten are further gallery accessions, normalized and feature-selected by their authors and read with the
`io.read_profiles` defaults. Use them to run a method across a range of screens.

Check anything tuned on one dataset against the others. Cutoffs that looked universal on BBBC021 turned out to be
Expand Down
9 changes: 9 additions & 0 deletions docs/references.bib
Original file line number Diff line number Diff line change
Expand Up @@ -259,3 +259,12 @@ @article{Celik_2024
doi = {10.1371/journal.pcbi.1012463},
url = {http://dx.doi.org/10.1371/journal.pcbi.1012463}
}

@article{Munoz_2026,
title = {JUMP-lite: Compact, reproducible benchmarking of cell representations},
author = {Muñoz, Alán F. and Fredin Haslum, Johan and Shen, Rex and Carpenter, Anne E. and Singh, Shantanu},
journal = {arXiv},
year = {2026},
doi = {10.48550/arXiv.2608.07632},
url = {https://arxiv.org/abs/2608.07632}
}
2,922 changes: 2,922 additions & 0 deletions docs/tutorials/12_learned_embeddings.ipynb

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions docs/tutorials/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,4 +17,5 @@ reading_plates
09_scaling_and_sites
10_differential_features
11_dose_response
12_learned_embeddings
```
95 changes: 70 additions & 25 deletions src/mantispy/_core/features.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,10 +3,16 @@
This is the only module that parses feature names.
Its output populates ``adata.var``, and nothing downstream re-parses names.

The grammar handled here is::
Two grammars are handled. CellProfiler's::

[<Object>_]<Group>_<feature words>[_<channel>...][_<numeric params>][_<NofM>]

and cp_measure's, which separates its tokens with slashes, names the channel by index and glues the group to the feature in camel case::

<object>_<channel>/<aggregation>/<group><Feature words>[_<numeric params>][_<NofM>]

A name is read as cp_measure's only when it matches that shape in full, which a CellProfiler name cannot because CellProfiler never emits a ``/``.

Columns that are not measurements (object numbers, parent/child links, locations, file names, metadata) get ``is_feature = False`` so callers can route them somewhere other than ``X``.
"""

Expand Down Expand Up @@ -122,6 +128,13 @@ def canonical_channel(channel: str | None, aliases: dict[str, str] | None = None
_RADIAL_BIN_RE = re.compile(r"^\d+of\d+$")
_NUMERIC_RE = re.compile(r"^-?\d+(\.\d+)?$")

#: ``cell_0/max/textureContrast_3_03_256``: object, channel index, per-object aggregation, then the group glued to the
#: feature in camel case. CellProfiler separates every token with an underscore and never emits a ``/``, so a name has
#: to match this whole shape before it is read this way.
_CP_MEASURE_RE = re.compile(
r"^(?P<object>[a-z]+)_(?P<channel>\d+)/(?P<agg>[a-z]+)/(?P<group>[a-z_]+)(?P<feature>[A-Z].*)$"
)


def _is_numeric(token: str) -> bool:
return bool(_NUMERIC_RE.match(token))
Expand Down Expand Up @@ -155,6 +168,52 @@ def _split_object(tokens: list[str]) -> tuple[str | None, str, list[str]]:
return None, tokens[0], tokens[1:]


def _read_suffixes(row: dict, rest: list[str], group: str) -> None:
"""Move the radial bin and the numeric suffix of ``rest`` into their own columns, leaving the feature name."""
radial = [token for token in rest if _RADIAL_BIN_RE.match(token)]
if radial:
row["radial_bin"] = radial[0]
rest = [token for token in rest if token not in radial]

numeric = [token for token in rest if _is_numeric(token)]
if numeric:
# Keep the full numeric suffix: Zernike_2_0 and Zernike_2_2 must stay distinct.
row["params"] = "_".join(numeric)
rest = [token for token in rest if not _is_numeric(token)]
if group.lower() == "texture":
for key, value in zip(("scale", "angle", "gray_levels"), numeric, strict=False):
row[key] = float(value)
else:
row["scale"] = float(numeric[0])

row["feature"] = "_".join(rest) if rest else group
row["is_feature"] = True


def _parse_cp_measure(name: str) -> dict:
"""Annotate one ``cp_measure`` name, whose channel is an index rather than the stain's name.

cp_measure is handed one channel at a time and numbers them in the order it was given them, so the index is all
the name carries. It is kept as the channel rather than resolved to a stain, because the mapping lives in the
acquisition metadata and guessing it would put a wrong stain on every intensity feature in the screen.
"""
row: dict = dict.fromkeys(COLUMNS)
match = _CP_MEASURE_RE.match(name)
if match is None:
# The grammar was chosen from the file as a whole, so a name that does not fit it is a mixed or
# hand-edited file rather than a name to guess at.
row["is_feature"] = False
return row

# A group ending in "_" means the name separated it from the feature, as in "sizeshape_Solidity".
group = match["group"].rstrip("_")
row["object"] = match["object"]
row["feature_group"] = group
row["channel"] = match["channel"]
_read_suffixes(row, match["feature"].split("_"), group)
return row


def _parse_one(name: str, channels: frozenset[str]) -> dict:
row: dict = dict.fromkeys(COLUMNS)
row["is_feature"] = False
Expand All @@ -179,30 +238,15 @@ def _parse_one(name: str, channels: frozenset[str]) -> dict:
row["channel"] = "|".join(found_channels)
rest = [token for token in rest if token not in channels]

radial = [token for token in rest if _RADIAL_BIN_RE.match(token)]
if radial:
row["radial_bin"] = radial[0]
rest = [token for token in rest if token not in radial]

numeric = [token for token in rest if _is_numeric(token)]
if numeric:
# Keep the full numeric suffix: Zernike_2_0 and Zernike_2_2 must stay distinct.
row["params"] = "_".join(numeric)
rest = [token for token in rest if not _is_numeric(token)]
if group == "Texture":
for key, value in zip(("scale", "angle", "gray_levels"), numeric, strict=False):
row[key] = float(value)
else:
row["scale"] = float(numeric[0])

row["feature"] = "_".join(rest) if rest else group
row["is_feature"] = True
_read_suffixes(row, rest, group)
return row


def parse_feature_names(names: Sequence[str], channels: Sequence[str] | None = None) -> pd.DataFrame:
"""Parse ``names`` into a table of feature annotations indexed by name.

Which grammar to read is decided once, from ``names`` as a whole, because it is a property of the file that wrote them rather than of any one column.

Args:
names: Column names from a CellProfiler or cp_measure table.
channels: The channel vocabulary.
Expand All @@ -213,12 +257,13 @@ def parse_feature_names(names: Sequence[str], channels: Sequence[str] | None = N
Text columns are ``category`` dtype (so they survive an h5ad round trip even when they are entirely missing) and ``is_feature`` is ``bool``.
"""
names = list(names)
channel_set = frozenset(channels if channels is not None else _infer_channels(names))
parsed = pd.DataFrame(
[_parse_one(name, channel_set) for name in names],
index=pd.Index(names),
columns=COLUMNS,
)
if any(_CP_MEASURE_RE.match(str(name)) for name in names):
# cp_measure names their own channel by index, so there is no vocabulary to infer or to be given.
rows = [_parse_cp_measure(str(name)) for name in names]
else:
channel_set = frozenset(channels if channels is not None else _infer_channels(names))
rows = [_parse_one(str(name), channel_set) for name in names]
parsed = pd.DataFrame(rows, index=pd.Index(names), columns=COLUMNS)
for column in _TEXT_COLUMNS:
parsed[column] = parsed[column].astype("category")
for column in _FLOAT_COLUMNS:
Expand Down
6 changes: 6 additions & 0 deletions src/mantispy/ds/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,16 @@

from mantispy.ds._blobs import blobs
from mantispy.ds._datasets import (
JUMP_LITE_MODELS,
agnp,
amish,
bbbc021,
chroma,
jump_cells,
jump_crispr,
jump_export,
jump_lite,
jump_lite_targets,
jump_plate,
jump_target2,
luad,
Expand All @@ -29,7 +32,10 @@
"blobs",
"chroma",
"jump_cells",
"JUMP_LITE_MODELS",
"jump_crispr",
"jump_lite",
"jump_lite_targets",
"jump_export",
"jump_plate",
"jump_target2",
Expand Down
104 changes: 104 additions & 0 deletions src/mantispy/ds/_datasets.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
import pandas as pd
from scverse_misc.datasets import fetch, parse_registry, register_loader

from mantispy._core.features import empty_annotation
from mantispy._core.frames import as_frame
from mantispy._core.logging import get_logger, report_drop
from mantispy._core.schema import SCHEMA_VERSION, stamp
Expand Down Expand Up @@ -578,6 +579,109 @@ def miami(cache_dir: str | Path | None = None, **kwargs: Any) -> AnnData:
return _profiles("miami", cache_dir, **kwargs)


#: The feature sets JUMP-Lite publishes for one set of wells: five learned embeddings, and the
#: CellProfiler-equivalent measurements of ``cp_measure`` for comparison on the same rows.
JUMP_LITE_MODELS = ("openphenom", "dinov2", "dinov2_random", "subcell", "morphem", "cp_measure")


def jump_lite(
model: str = "openphenom", annotate: bool = True, cache_dir: str | Path | None = None, **kwargs: Any
) -> AnnData:
"""JUMP-Lite Target-2: 1,536 wells, four imaging sites, one feature set at a time.

``cpg0016-jump``, the compact benchmark of :cite:t:`Munoz_2026`. Four plates of the JUMP Target-2 plate map, one from each of ``source_3``, ``source_4``, ``source_5`` and ``source_6``, so the four batches are four different laboratories running the same 302 compounds with 64 DMSO wells each.

Every ``model`` covers the same 1,536 wells, which is what makes this a comparison rather than six datasets: the rows and the metadata are identical and only the feature block changes. Five are learned embeddings and one, ``"cp_measure"``, is the CellProfiler-equivalent measurement of the same images.

Args:
model: Which feature set to read, one of ``ds.JUMP_LITE_MODELS``. ``"dinov2_random"`` is the same architecture with untrained weights, which is the null model the benchmark scores the others against.
annotate: Join the JUMP well and compound tables, which name the compound of each well.
Downloads about 14 MB once and caches it.
cache_dir: Where to keep the download.
Defaults to :attr:`mantispy.settings.cache_dir`.
kwargs: Passed to :func:`mantispy.io.read_profiles`.

Returns:
Wells by features at well resolution, indexed by plate and well, with ``Metadata_Source``, ``Metadata_Batch``, ``Metadata_Plate``, ``Metadata_Well``, ``Metadata_CellCount`` and, when annotated, ``Metadata_JCP2022``, ``Metadata_Perturbation``, ``Metadata_InChIKey`` and ``Metadata_Control``.

Raises:
ValueError: ``model`` is not one of ``ds.JUMP_LITE_MODELS``.

Notes:
A dimension of a learned embedding is a coordinate in the model's own basis, not a measurement with a name to parse, so for every model but ``"cp_measure"`` the annotation columns of ``var`` are supplied empty. Anything that reads ``var["feature_group"]`` or ``var["channel"]``, such as the feature families :func:`~mantispy.pl.effect_sizes` colours by, has nothing to work with on those.

``"cp_measure"`` is CellProfiler-style measurements and keeps its parsed compartment, feature group and channel. Its channel is the index cp_measure numbered its inputs by rather than the name of a stain, because the name lives in the acquisition metadata and not in the feature name.

The embeddings are not normalized. They are the model's output on each well's images, so a per-plate control normalization is still the first step.

The trained embeddings here carry the cell count in their leading components, where it can account for more of the variance than either the laboratory or the imaging site. The untrained ``"dinov2_random"`` does not, and neither does ``"cp_measure"``, whose per-cell measurements are averaged over the well. Measure it with :func:`~mantispy.metrics.evaluate_correction` before correcting for anything else, and read :doc:`/tutorials/12_learned_embeddings` on why removing it is not obviously right.

References:
:cite:t:`Munoz_2026`, :cite:t:`Chandrasekaran_2023`, :cite:t:`Weisbart_2024`.
"""
if model not in JUMP_LITE_MODELS:
raise ValueError(f"model must be one of {JUMP_LITE_MODELS}, got {model!r}")

adata = _profiles("jump_lite", cache_dir, select=lambda name: name == f"{model}.parquet", **kwargs)
if model != "cp_measure":
# An embedding dimension is a coordinate in a learned basis, not a measurement with a name
# to parse. Left alone, the parser reads "openphenom_nahualX_17" as the "nahualX" feature
# group of an "openphenom" object, and the model's own tensor names become feature families.
empty = empty_annotation(adata.var_names)
adata.var[empty.columns] = empty

(counts_path,) = _files("jump_lite", cache_dir, select=lambda name: name == "cell_count.parquet")
counts = pd.read_parquet(counts_path, columns=["Metadata_id", "cell_count"])

obs = as_frame(adata.obs)
joined = obs.merge(counts, on="Metadata_id", how="left", validate="1:1")
joined.index = obs.index
if unmatched := int(joined["cell_count"].isna().sum()):
# A left join leaves the count missing rather than failing, and everything that reads it downstream,
# from cytotoxicity to the well filters, would quietly treat those wells as having no cells.
get_logger().warning("jump_lite(%s): %d well(s) have no cell count in the count table", model, unmatched)
adata.obs = joined.rename(columns={"cell_count": "Metadata_CellCount"})

if annotate:
from mantispy.pp._annotate import annotate_jump

annotate_jump(adata)
get_logger().info(
"jump_lite(%s): %d wells x %d features over %d source(s)",
model,
adata.n_obs,
adata.n_vars,
int(as_frame(adata.obs)["Metadata_Source"].nunique()),
)
return adata


def jump_lite_targets(cache_dir: str | Path | None = None) -> pd.DataFrame:
"""The gene each JUMP compound is annotated to act on, as a set per target.

RefChemDB annotations distributed with :cite:t:`Munoz_2026`, in the ``source``/``target`` shape :func:`~mantispy.metrics.known_relationships` reads, so two compounds annotated to the same gene count as a related pair.

Args:
cache_dir: Where to keep the download.
Defaults to :attr:`mantispy.settings.cache_dir`.

Returns:
A frame with ``source``, the gene symbol, and ``target``, the ``Metadata_JCP2022`` of a compound annotated to it. One row per annotated pairing, over every JUMP compound rather than only those of :func:`jump_lite`.

Notes:
The annotation is sparse against a plate map: most compounds on the JUMP-Lite plates carry none, and only the targets shared by more than one compound contribute a pair, so the recall is computed over a minority of the plate.

References:
:cite:t:`Munoz_2026`.
"""
(path,) = _files("jump_lite", cache_dir, select=lambda name: name == "refchem_annotations.parquet")
frame = pd.read_parquet(path, columns=["target", "Metadata_JCP2022"])
# Dropped before the cast, or an unannotated compound becomes the literal string "nan" and
# every one of them is then related to every other.
frame = frame.dropna().rename(columns={"target": "source", "Metadata_JCP2022": "target"}).astype(str)
return frame.drop_duplicates().reset_index(drop=True)


def jump_crispr(cache_dir: str | Path | None = None, **kwargs: Any) -> AnnData:
"""The assembled JUMP CRISPR arm, 51,185 wells of knockouts.

Expand Down
Loading
Loading