Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion .claude/skills/bioengine-maintainer/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -118,9 +118,16 @@ Local artifact development: `export BIOENGINE_LOCAL_ARTIFACT_PATH=/path/to/bioen
### Run tests

```bash
pytest tests/end_to_end/ -v
pytest tests/ # offline tests only
pytest tests/end_to_end/ -v --live # deploys to and calls a real cluster
```

Tests that act on a live cluster are deselected without `--live`, so the
end-to-end command above exits 5 ("no tests collected") if you omit it.
Whenever a credential resolves, the run prints the server, the workspace each
token is scoped to, and how many live-reaching tests are enabled — with or
without `--live`, so a finished run's log says what it could have touched.

Test organisation:
- `tests/end_to_end/` — integration tests for the core worker
- `tests/apps/` — per-app tests, one subdirectory per app (`tests/apps/cellpose/`, …)
Expand Down
1 change: 1 addition & 0 deletions pytest.ini
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ markers =
integration: marks tests that use real Ray or other subsystems
unit: marks tests that use mocked external dependencies
requires_gpu: marks tests that require a GPU runtime deployment to be running
live: marks tests that act on a real Hypha cluster; deselected unless --live is passed
testpaths = tests
python_files = test_*.py
python_classes = Test*
Expand Down
21 changes: 21 additions & 0 deletions tests/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,27 @@ pip install -r requirements-test.txt
### 3. Environment Configuration
The `.env` file in the project root contains required environment variables including `HYPHA_TOKEN`. This is automatically loaded by the test configuration.

### 4. Tests that act on a live cluster

Some tests deploy applications to, and call services on, the production workspaces at `https://hypha.aicell.io`. They are deselected unless you pass `--live`:

```bash
pytest tests/ # offline tests only
pytest tests/ --live # also runs tests that act on the live cluster
```

A test counts as live if it requests `hypha_client`, `hypha_token` or `model_runner`, or is marked `@pytest.mark.live`. Anything that reaches a cluster by some other route — a browser driven at the deployed app, a client built inline — has to carry the marker. Liveness has nothing to do with the `--ignore` lists in circulation: `tests/apps/model-runner/` and `tests/test_artifact_version.py` are live and neither lives under `tests/end_to_end/`.

**Do not rely on your working directory to keep a credential away from the suite.** There are at least three ways one arrives, and only the last is under your control:

- `load_dotenv()` above searches **upward** from `tests/`, not just the directory you started in. A worktree under `<repo>/.claude/worktrees/` has no `.env` of its own but still reaches the repo root's, so it is *not* protected. A worktree outside the repo tree (say under `/tmp`) does stop the walk.
- `tests/test_artifact_version.py` loads the repo-root `.env` by explicit path, wherever you run it from.
- `source .env`, `export HYPHA_TOKEN=…`, or a token already in your shell defeats all of the above regardless of location.

Running inside the worker image with only the repo bind-mounted also stops the upward walk, because the search cannot escape the mount.

Because none of that is reliable, the gate does not depend on it. Whenever a credential resolves, the run prints — before the first test, and again next to the wall clock at the end — which server it is, which workspace each token is scoped to, and how many live-reaching tests are enabled. That last number is printed **with or without `--live`**: the flag only prevents the accidental case, so once someone has opted in deliberately the count in the log is what makes an after-the-fact audit possible. Runtime is the corroborating signal; the same scope takes roughly 30s offline and several minutes against a cluster.

## Running Tests

### All Tests
Expand Down
5 changes: 5 additions & 0 deletions tests/apps/cellpose/test_ui_e2e.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,11 @@
import pytest
from playwright.sync_api import Page, expect

# Every test here drives a browser at the deployed app and injects HYPHA_TOKEN
# into its localStorage. The `page` fixture is not in LIVE_FIXTURES, so the
# marker is what gates them.
pytestmark = pytest.mark.live

SERVER_URL = "https://hypha.aicell.io"
APP_WORKSPACE = os.environ.get("HYPHA_TEST_WORKSPACE", "ri-scale")
APP_URL = f"{SERVER_URL}/{APP_WORKSPACE}/view/cellpose-finetuning"
Expand Down
5 changes: 4 additions & 1 deletion tests/apps/model-runner/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,10 @@
Requires the model-runner app to be deployed to bioimage-io/bioengine-worker.
Set BIOIMAGE_IO_TOKEN (or HYPHA_TOKEN) in the environment before running.

pytest tests/apps/model-runner/ -v -o "addopts="
These call the live service, so they need --live; without it they are
deselected and the command exits 5 with no tests collected.

pytest tests/apps/model-runner/ -v -o "addopts=" --live
"""

import io
Expand Down
173 changes: 171 additions & 2 deletions tests/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,14 +8,17 @@
"""

import asyncio
import base64
import json
import os
import re
import subprocess
import sys
import tempfile
import time
from datetime import datetime
from pathlib import Path
from typing import AsyncGenerator, Generator
from typing import AsyncGenerator, Dict, Generator, List, Optional, Tuple

import pytest
import pytest_asyncio
Expand All @@ -28,6 +31,172 @@
# Load environment variables from .env file
load_dotenv()

# Every Hypha connection in the suite is hard-coded to this deployment; no
# environment variable redirects it.
HYPHA_SERVER_URL = "https://hypha.aicell.io"

# Requesting one of these fixtures means the test authenticates against the
# real deployment above. Anything else that reaches a cluster has to say so
# with @pytest.mark.live.
LIVE_FIXTURES = frozenset({"hypha_client", "hypha_token", "model_runner"})

# Credentials the suite picks up on its own, in the order a reader should
# worry about them. `load_dotenv()` above supplies these from a repo-root
# .env, so whether they resolve depends on which checkout pytest was run from.
LIVE_TOKEN_VARS = ("HYPHA_TOKEN", "BIOIMAGE_IO_TOKEN")

_START_TIME = pytest.StashKey[float]()


def pytest_addoption(parser: pytest.Parser) -> None:
parser.addoption(
"--live",
action="store_true",
default=False,
help=(
"Run tests that deploy to and call a real Hypha cluster. Without "
"this flag they are deselected even when a token is available."
),
)


def _token_workspace(token: str) -> str:
"""Workspace a Hypha JWT is scoped to, or '' if it is not a readable JWT."""
try:
segment = token.split(".")[1]
claims = json.loads(
base64.urlsafe_b64decode(segment + "=" * (-len(segment) % 4))
)
scope = str(claims.get("scope", ""))
except Exception:
# A token we cannot parse must not abort the session; the caller
# reports the credential as present with an unknown workspace.
return ""
for part in scope.split():
if part.startswith("wid:"):
return part[len("wid:") :]
return ""


def resolved_live_targets(
environ: Optional[Dict[str, str]] = None,
) -> List[Tuple[str, str]]:
"""(variable, workspace) for every cluster credential visible to this run."""
environ = os.environ if environ is None else environ
return [
(var, _token_workspace(environ[var]) or "<unreadable token>")
for var in LIVE_TOKEN_VARS
if environ.get(var)
]


def live_exposure_banner(
targets: List[Tuple[str, str]], live_count: int, enabled: bool
) -> str:
"""The text printed before the first test when a cluster credential resolves.

The count is stated whether or not --live was given: the flag only stops
the accidental case, and once someone has opted in deliberately the count
in the log is the only thing that makes an after-the-fact audit possible.
"""
lines = [
"=" * 22 + " live cluster credentials resolved " + "=" * 22,
f"{live_count} collected test(s) can reach the live cluster at "
f"{HYPHA_SERVER_URL}, and a credential for it resolved:",
]
lines += [f" {var} -> workspace {workspace}" for var, workspace in targets]
lines.append(
f"THEY WILL RUN: --live was given, so {live_count} test(s) will act on "
"that workspace."
if enabled
else f"Deselected {live_count} test(s); pass --live to run them."
)
lines.append("=" * 78)
return "\n".join(lines)


def _emit(config: pytest.Config, text: str) -> None:
"""Write where the controller will actually see it.

Under pytest-xdist -- which `addopts` turns on with `--numprocesses=1` --
collection runs inside a worker whose terminalreporter output is dropped
and whose deselection stats never reach the summary. The worker's stderr
is inherited by the controller, so that is the channel that survives.
"""
worker = getattr(config, "workerinput", None)
if worker is not None:
if worker.get("workerid", "gw0") == "gw0":
sys.stderr.write(text + "\n")
sys.stderr.flush()
return

reporter = config.pluginmanager.get_plugin("terminalreporter")
if reporter is None:
print(text)
else:
reporter.write_line(text)


def _is_live(item: pytest.Item) -> bool:
if any(item.iter_markers("live")):
return True
return bool(LIVE_FIXTURES.intersection(getattr(item, "fixturenames", ())))


@pytest.hookimpl(tryfirst=True)
def pytest_collection_modifyitems(
config: pytest.Config, items: List[pytest.Item]
) -> None:
# Attaching the marker is what makes `-m live` agree with `--live`: most
# live tests are detected from their fixture closure, so without this they
# are gated but unnamed and `-m live` returns a wrong subset. It has to
# happen before pytest's own mark filtering -- which a conftest hookimpl
# already does under pluggy's reverse-registration order, so tryfirst is a
# guarantee against a plugin registering an earlier modifyitems, not a
# load-bearing fix. Switching it to trylast does break the invariant.
live, offline = [], []
for item in items:
if _is_live(item):
item.add_marker(pytest.mark.live)
live.append(item)
else:
offline.append(item)

enabled = config.getoption("--live")
if live and not enabled:
items[:] = offline
config.hook.pytest_deselected(items=live)

targets = resolved_live_targets()
if targets:
_emit(config, live_exposure_banner(targets, len(live), enabled))


def pytest_configure(config: pytest.Config) -> None:
config.stash[_START_TIME] = time.time()


def pytest_terminal_summary(
terminalreporter, exitstatus: int, config: pytest.Config
) -> None:
"""Restate the exposure next to the wall clock.

Runtime is the cheapest signal that live tests ran -- the same command
takes ~30s offline and several minutes against a cluster -- so an audit
needs the credential and the duration in one place.
"""
targets = resolved_live_targets()
if not targets:
return
elapsed = time.time() - config.stash.get(_START_TIME, time.time())
workspaces = ", ".join(workspace for _, workspace in targets)
terminalreporter.write_line(
f"live cluster exposure: credentials for {workspaces} on "
f"{HYPHA_SERVER_URL} were in scope for this {elapsed:.1f}s run "
f"({'--live GIVEN' if config.getoption('--live') else '--live not given'}).",
red=True,
)


@pytest.fixture(
scope="session",
Expand Down Expand Up @@ -192,7 +361,7 @@ def head_node_port(ray_address: str) -> int:
@pytest.fixture(scope="session")
def server_url() -> str:
"""Return Hypha server URL for test connections."""
return "https://hypha.aicell.io"
return HYPHA_SERVER_URL


@pytest.fixture(scope="session")
Expand Down
8 changes: 7 additions & 1 deletion tests/test_artifact_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,14 @@
3. Re-saving the same version → rejected (would overwrite a published release).
4. Saving an older version → rejected.

These act on ``bioimage-io/bioengine-worker`` in production: they call
``upload_app`` and delete the artifacts again. They are the most destructive
tests in the suite, hence the module-wide ``live`` marker.

Run with:
conda activate bioengine
source .env
pytest tests/test_artifact_version.py -v
pytest tests/test_artifact_version.py -v --live
"""

import os
Expand All @@ -26,6 +30,8 @@

load_dotenv(Path(__file__).parent.parent / ".env")

pytestmark = pytest.mark.live


# ── helpers ────────────────────────────────────────────────────────────────────

Expand Down
Loading
Loading