diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index afa8c51d..300f6d67 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -222,7 +222,7 @@ jobs: open4d/*.py open4d/streaming/*.py open4d/codec open4d/core open4d/io open4d/torch_ops open4d/visualization integrations/__init__.py integrations/open3d - examples/visualization scripts + examples/visualization examples/streaming_demo.py scripts - name: Check shell syntax run: bash -n scripts/*.sh - name: Check local Markdown links diff --git a/README.md b/README.md index 316f463d..0431e039 100644 --- a/README.md +++ b/README.md @@ -17,13 +17,18 @@ the extras you use: ```bash python -m pip install -e '.[player]' # interactive mesh viewer and GIF export -python -m pip install -e '.[open3d]' # RGB-D reconstruction +python -m pip install -e '.[open3d]' # RGB-D reconstruction (Open3D 0.19.x) python -m pip install -e '.[gaussians]' # read Gaussian PLY files ``` Research methods have additional setup below. Their source, native programs and model weights are not bundled in the Python wheel. +RGB-D reconstruction requires Open3D 0.19.x. The legacy TSDF integrator in +Open3D 0.20 rescales already-metric float depth and can return empty meshes; +the extra selects the supported version and reconstruction rejects an +incompatible manually installed runtime before processing frames. + ## Try a sequence No dataset is needed for this example: @@ -151,10 +156,10 @@ with receive() as frames: Then send from another: ```python -from open4d import stream +from open4d import send from open4d.demo import mesh_sequence -stream(mesh_sequence(frames=30)) +send(mesh_sequence(frames=30)) ``` This sends decoded mesh arrays over TCP, at their recorded frame timing. It is diff --git a/THIRD_PARTY.md b/THIRD_PARTY.md index 5c1c9402..bd55a776 100644 --- a/THIRD_PARTY.md +++ b/THIRD_PARTY.md @@ -51,7 +51,7 @@ distribution and still needs separate review before any redistribution. | `open4d/reconstruction/queen` | Imported research tree including MiDaS, SIBR, and rasterizers | Top-level NVIDIA non-commercial license plus multiple subtree licenses | `BLOCK`; record revisions, patches, all notices, model/data rights, and redistribution limits | | `open4d/reconstruction/gs_tools` | Consolidated Gaussian tooling tree containing SIBR viewers, GLM, and three rasterizer imports; one immutable upstream/patch manifest is absent | Component-local SIBR, GLM, and rasterizer license files exist with differing terms | `BLOCK`; inventory exact upstream revisions and patches, preserve every notice, and determine compatible source/binary distribution terms | | `open4d/reconstruction/vega` | Copied from `4DVideoStreaming` `baselines/Vega`, working tree above commit `6c2569de85ebe4592a8294c8cb0268efd3659212`, so not a clean revision: it carries three uncommitted modifications (`vega/encoder.py`, `vega/gov.py`, `orbitvega/prepare.py`) and two never-committed files (`orbitvega/scene_export.py`, `orbitvega/scene_render.py`). Local re-implementation of Vega (Kim et al., ACM MobiCom '25, `10.1145/3680207.3765267`); no upstream code release exists. Open4D modifications: import roots rerooted off `baselines.Vega`, `pytest.ini` added | No component license file; `citation.txt` records the paper only, and `vega_engine_pyproject.toml` names it a research prototype | `BLOCK`; identify the implementation's copyright holder and distribution terms, and pin an immutable source revision | -| `open4d/reconstruction/nevo` | Copied from `4DVideoStreaming` `baselines/NeVo` at its introducing commit `6c2569de85ebe4592a8294c8cb0268efd3659212` plus one uncommitted modification (`orbitnevo/train.py`). NeVo (Wu et al., ACM MobiCom '25, `10.1145/3680207.3723473`) has no released code, so the simulator is local; it vendors `https://github.com/aoliao12138/ReRF` (CVPR 2023) under `rerf/`, **cloned at depth 1 so no upstream revision is recorded**, byte-identical apart from four non-code deletions listed in `rerf/PATCHES.md`. Open4D modifications: import roots rerooted off `baselines.NeVo`, `MODULE_ROOT` repointed, `pytest.ini` and `orbitnevo/objects.py` added | `rerf/LICENSE` is GPL-3.0 carrying a research-purposes-only rider, and its code base derives from DVGO; no top-level component license | `BLOCK`; recover and record the exact ReRF revision, resolve the GPL-3.0 plus research-only boundary and the DVGO lineage, and identify terms for the local simulator | +| `open4d/reconstruction/rerf` | `rerf_stream/` is project-authored. `upstream/` is a clone of `https://github.com/aoliao12138/ReRF` (Wang et al., CVPR 2023) at `510e6073827ff15f58efda17a8729f2c47e13401` (2023-07-24), byte-identical to that tree apart from four non-code deletions enumerated in the module README (`.git/`, `ac_dc/` CMake intermediates, a README image); upstream carries no Open4D patches, all local adaptation living in `rerf_stream/env.py`. Supersedes the removed `open4d/reconstruction/nevo`, which vendored the same ReRF code at an unrecorded depth-1 revision and remains in history at `3d33655`. `upstream/ac_dc/` ships five prebuilt `.so` modules and an `ncvv_code` binary with no sources, only CMake residue, over a pybind11 checkout whose source is likewise absent | `upstream/LICENSE` is GPL-3.0 prefixed with a `THIS CODE CAN ONLY BE USED FOR RESEARCH PURPOSES` rider that contradicts the GPL grant; the code base derives from DVGO, whose kernels are `upstream/lib/cuda/`. No component license covers `rerf_stream/` beyond root MIT | `BLOCK`; resolve the GPL-3.0 against research-only conflict and the DVGO lineage, obtain reproducible sources and a dependency bill for the `ac_dc` binaries or exclude them from every artifact, and record terms for the project-authored `rerf_stream/` layer | | `integrations/unity` | Project glue, copied Eigen, prebuilt plugins, encoded archive | No integration-wide manifest; Eigen has multiple license files; binaries/data unresolved | `BLOCK`; inventory producers, licenses, build revisions, symbols, and fixture rights | | `integrations/unity/TVMCUnity/Unity Files/Plugins` | Prebuilt Android `.so` and macOS `.dylib` | Reproducible build and dependency bill absent | `BLOCK`; exclude until reproduced from reviewed source with notices | | `integrations/unity/TVMCUnity/EncodedExample/DanceSequence.zip` | Historical encoded dataset archive | Source dataset license/consent absent | `BLOCK`; identify redistribution permission or replace with a licensed synthetic fixture | diff --git a/docs/api.md b/docs/api.md new file mode 100644 index 00000000..e2e7ace2 --- /dev/null +++ b/docs/api.md @@ -0,0 +1,137 @@ +# Python API + +The public API loads, saves, unloads, and visualizes whole finite triangle-mesh +sequences independently of their storage format: + +```python +import open4d + +with open4d.load("capture.usdc") as sequence: + open4d.save(sequence, "capture.o4d") + open4d.visualize(sequence) + +# A path can go straight to the lazy viewer; it is closed when the window exits. +open4d.visualize("capture.o4d") +``` + +`.usd`, `.usda`, `.usdc`, and `.usdz` are OpenUSD interchange containers. +`.o4d` and the registered codec suffixes are codec artifacts. Both carry whole +sequences and use the same `Sequence` interface. `open4d.unload(sequence)` is an +explicit, idempotent alternative to the context manager. + +Frame folders and individual meshes remain supported as import paths; frames +are decoded on access: + +```python +from open4d.io import open_sequence + +with open_sequence("path/to/frames", fps=30.0) as sequence: + print(len(sequence), sequence.duration, sequence.fps) + mesh = sequence[0].geometry # TriangleMesh: positions, triangles +``` + +## Representations + +What a decoded frame *is* is deliberately separate from the codec that produced +it: a triangle mesh is a mesh whether it arrived as OBJ, as a Draco payload, or +out of a V-DMC bitstream. `open4d.Representation` is the axis to gate on, and +`MESH`, `POINTS`, and `GAUSSIANS` have concrete types — `TriangleMesh`, +`PointCloud`, and `GaussianCloud`. Already-rendered pixels are named in the +taxonomy but have no concrete type yet; they land with the camera model they +need in order to be comparable at a known pose. + +The containers and codecs described below are the mesh path, and the one most +completely covered. Check a specific codec before assuming it round-trips +points or Gaussians. + +## Writing sequences + +`write_sequence(sequence, "frames/", format="ply")` writes a versioned +`open4d.sequence.json` beside the frame files, so reopening the directory keeps +source frame indices, timestamps, frame/sequence metadata, and topology +declarations. Empty sequences are rejected before the destination is changed. + +Single mesh-file exports require `allow_lossy=True` because that storage cannot +preserve sequence timing, metadata, or topology declarations. Trimesh-backed +OFF/GLB/glTF color export also requires that opt-in because OFF drops vertex +color and GLB/glTF quantize canonical float colors to eight bits. + +OpenUSD is the public interchange container. `--pack-usd out.usdc` packs any +source into one compressed `.usdc` file carrying the frame rate, the key-frame +index, and per-frame streams alongside the geometry — see the +[visualization guide](../examples/visualization/README.md#the-openusd-container). + +## Streaming + +`open4d.stream` exports a file as a bundle and serves it to a browser: + +```python +import open4d + +open4d.stream("capture.usdc") +open4d.stream("capture.usdc", rungs=["draco", "draco@11"], out_dir="bundle/") +``` + +For a loaded sequence, pass browser options such as `name="capture"` and +`out_dir="bundle/"`. `open4d.send(sequence, host, port)` explicitly selects +decoded-mesh TCP transport and pairs with `open4d.receive`. Existing +`open4d.stream(sequence, host, port)` calls, and frame iterables without browser +options, retain that TCP behavior without requiring `open4d-streamer`. + +`rungs` is the quality ladder. The first is the rendition a client plays by +default and the rest are what it can switch to mid-playback, so a one-entry +list is a fixed-quality stream and says so. A spec is a frame format — +`ply` for interchange, `draco` for delivery — optionally with a position +quantisation, as in `draco@11`. Sizes are measured off disk rather than +predicted; quality is left unscored until something scores it. + +[`examples/streaming_demo.py`](../examples/streaming_demo.py) runs the whole +of it on the ten basketball frames the TVMC codec vendors: three rungs +built and scored, served over HTTP with the counters read back, then thirty +seconds simulated over a link that collapses mid-run. + +The implementation is the separate `open4d-streamer` package, imported on the +call rather than at load: it depends on `open4d`, so `open4d` must not depend +on it. Without it installed the call raises `open4d.StreamerDependencyError` +saying how to install it. For several clips in one bundle, a constrained link, +or the delivered-quality metrics, use that package directly — +`streamer.Bundle`, `streamer.Link`, `streamer.serve`. + +Separately, [`open4d/webclients`](../open4d/webclients) is the vendored +browser-client research tree that compares five delivery systems against each +other. It is not this API and shares no code with it. + +## Codecs + +Five lossless, in-process reference codecs are included: `raw`, `deflate`, +`bzip2`, `lzma`, and byte-level `rle` (`npz` remains the default DEFLATE alias). +They share a safe NumPy-array container so they compare storage strategies, not +research geometry models. + +Source checkouts register in-process adapters for `klt`, `n4mc`, `qndf`, and +`qndf-int8`; the lightweight wheel omits them until their provenance review is +complete. Open4D's separate `temporal-delta` and `temporal-pca` experiments are +not the repository's TVMC or TSMC pipelines. The V-DMC adapters do not execute +shell scripts, but they do invoke configured native encoder and decoder +processes once per sequence. Callers can also register another +`open4d.codec.Codec`. + +For an all-registered-codec attempt using `4d_files/Rafa_Approves_hd_4k`, open +[`examples/open4d_sequence_codec.ipynb`](../examples/open4d_sequence_codec.ipynb). +Set `OPEN4D_NOTEBOOK_REQUIRE_ALL=1` in a fully provisioned environment to make +any codec failure stop the notebook instead of appearing only in its result +table. + +### Device selection + +The N4MC and QNDF adapters accept `device="auto"` (CUDA, then Apple Metal/MPS, +then CPU), or an explicit `"cuda"`, `"mps"`, or `"cpu"`. QNDF-int8 can train on +CUDA or Metal, but its quantized decoder remains CPU-only. Override the notebook +selection with `OPEN4D_NOTEBOOK_DEVICE=mps` when needed. + +### Test coverage + +Normal CI runs dependency-complete CPU encode/fresh-decode contracts for KLT, +N4MC, QNDF, and QNDF-int8. The larger two-format Rafa quality/export matrix is +an additional CUDA acceptance test gated by `OPEN4D_TEST_RESEARCH_CODECS=1` and +`OPEN4D_RAFA_DATASET`; it is not presented as part of ordinary CI coverage. diff --git a/docs/assets/streaming-demo.png b/docs/assets/streaming-demo.png new file mode 100644 index 00000000..cfd0bef3 Binary files /dev/null and b/docs/assets/streaming-demo.png differ diff --git a/docs/components.md b/docs/components.md new file mode 100644 index 00000000..973316e1 --- /dev/null +++ b/docs/components.md @@ -0,0 +1,37 @@ +# Components + +Each component has its own README and may add native tools, GPU extensions, or +hardware requirements to the shared Python baseline. See +[Requirements](requirements.md) for those additions. + +## Mesh codecs + +| Codec | | +|---|---| +| [**N4MC**](../open4d/codecs/n4mc/README.md) | Neural TSDF-based mesh compression, including a newer modular codec under its `data`, `models`, `losses`, `training`, and `evaluation` packages | +| [**QNDF**](../open4d/codecs/qndf/README.md) | Quantized Neural Displacement Fields: static mesh compression using an SSP coarse mesh and an implicit displacement decoder. [`qndf_int8`](../open4d/codecs/qndf_int8/README.md) is its quantized variant | +| [**TVMC**](../open4d/codecs/tvmc/README.md) | A Python, .NET, and Draco pipeline for tracked time-varying mesh compression, with setup and resumable pipeline scripts | +| [**TSMC**](../open4d/codecs/tsmc/README.md) | Scene-mesh compression with optional SAM-based static/dynamic separation, ARAP volume tracking, deformation, displacement compression, and evaluation | +| [**KLT**](../open4d/codecs/klt/README.md) | Karhunen–Loève Transform baseline that compresses TSDF voxel blocks with a learned linear basis and quantized coefficients, reconstructing meshes via marching cubes | +| [**Draco**](../open4d/codecs/draco/README.md) | Google Draco mesh-compression baseline. Wraps the vendored `draco_encoder`/`draco_decoder` binaries into a per-frame encode/decode/eval pipeline for benchmarking against the neural codecs | +| [**MPEG V-DMC test model**](../open4d/codecs/vdmc/README.md) | The pinned MPEG reference implementation for video-based dynamic mesh coding — reference encoder, decoder, metric tools, and unit tests. Separate from Open4D's TVMC research pipeline | +| [**Faster V-DMC**](../open4d/codecs/faster_vdmc/README.md) | A pinned performance-oriented fork of the same test model, with exact-output and higher-throughput modes recorded in the [benchmark report](benchmarks/faster-vdmc.md) | + +## Reconstruction and streaming + +| Module | | +|---|---| +| [**RGB-D**](../open4d/streaming/README.md) | Synchronized multi-camera RGB-D ingestion, calibrated point-cloud fusion, CUDA TSDF mesh reconstruction, and live browser playback. Includes both the original native reconstruction code and the Python two-camera streaming pipeline | +| [**QUEEN**](../open4d/reconstruction/queen/README.md) | Quantized efficient encoding of dynamic Gaussians for streaming free-viewpoint video (NeurIPS 2024) | +| [**3DGStream**](../open4d/reconstruction/3dgstream/README.md) | On-the-fly training of 3D Gaussians for streaming photo-realistic free-viewpoint video (CVPR 2024) | +| [**Vega**](../open4d/reconstruction/vega/README.md) | An ORBIT adaptation of Vega (MobiCom 2025): mobile volumetric video streaming with 3D Gaussian splatting | +| [**ReRF**](../open4d/reconstruction/rerf/README.md) | Neural residual radiance fields (CVPR 2023) as a streamable compression method | +| [**gs-tools**](../open4d/reconstruction/gs_tools/README.md) | The one environment, rasterizers, and viewer the Gaussian methods share, plus `gs-tools view` for putting several methods under one camera | +| [**streamer**](../open4d/streamer/README.md) | Streaming and playback for 4D reconstructions, whatever their representation. Producers import it to describe and serve their output; it imports none of them | + +## Integrations + +| | | +|---|---| +| [**Open3D**](../integrations/open3d/README.md) | Converts decoded Open4D geometry into standard Open3D `TriangleMesh` or `PointCloud` objects. An adapter, not a loader | +| [**Unity**](../integrations/unity/README.md) | A Unity playback system for TVMC-encoded sequences: a C++ decoder backend and a C# front end, for XR targets | diff --git a/docs/requirements.md b/docs/requirements.md new file mode 100644 index 00000000..a34b71ea --- /dev/null +++ b/docs/requirements.md @@ -0,0 +1,140 @@ +# Requirements and installation + +## Baseline + +| | | +|---|---| +| Python | 3.10–3.13 | +| Operating system | macOS, Linux, or Windows | +| CPU | Any x86-64 or arm64; no particular core count | +| GPU | Not required. The viewers open a real OpenGL window, so a graphical session is needed even for `--save` | +| Memory | Roughly 1 MB of RAM per frame of playback | +| Disk | About 1.5 GB for a clone with submodules initialized | + +`pip install -e .` needs only NumPy, and reads `.obj` and `.ply` with no further +dependencies. Extras add optional readers and viewers. The comparison program +additionally needs SciPy, which the `[player]` extra installs, for its +nearest-neighbour search — the same `cKDTree` query TVMC's own evaluation uses. + +Open3D ships no 3.13 wheels, capping `.[open3d]` and the codecs at 3.12. + +## Installation + +Clone with submodules to obtain the pinned Draco, libigl, SAM3, and MPEG V-DMC +source: + +```bash +git clone --recurse-submodules https://github.com/open4dfoundation/Open4D.git +cd Open4D +``` + +For the lightweight core package: + +```bash +python -m venv .venv +source .venv/bin/activate # Windows: .venv\Scripts\activate +python -m pip install --upgrade pip +python -m pip install -e . +``` + +Optional local tooling is available through extras: + +```bash +python -m pip install -e ".[player]" # the example viewer (PyQt6 + pyqtgraph) +python -m pip install -e ".[usd]" # OpenUSD containers +python -m pip install -e ".[tools]" # trimesh, for extra mesh formats +python -m pip install -e ".[open3d]" # Open3D adapter; Python 3.12 or older +python -m pip install -e ".[qndf]" # QNDF/QNDF-INT8 in-process adapters +python -m pip install -e ".[temporal]" # experimental temporal-delta/PCA codecs +python -m pip install -e ".[all]" +``` + +These extras do not install the heavyweight codec environments. Use the setup +instructions inside the selected codec before running it. Research codec +implementations remain source-checkout-only and are excluded from the +lightweight wheel until their provenance review is complete. + +If an existing clone is missing Draco, initialize and build all three copies — +the Draco baseline codec's own, plus TSMC's and TVMC's — with: + +```bash +./scripts/setup_draco.sh +``` + +## One Python dependency set for the codecs + +The supported baseline for codec Python stages is described by +[`environment.yml`](../environment.yml) at the repository root: + +```bash +conda env create -f environment.yml +conda activate open4d +pip install -e . +``` + +The Python set is Python 3.12, NumPy 1.26.4, Open3D 0.19, and PyTorch 2.7.0. +Native projects use one external .NET 10 SDK. This replaces three Python +versions, two Open3D versions, two PyTorch versions, and three .NET targets. The +Python pins themselves live in +[`requirements-codecs.txt`](../requirements-codecs.txt), which `environment.yml` +installs; it lists direct dependencies only, so inside an existing Python 3.12 +environment `pip install -r requirements-codecs.txt` is equivalent. + +Codec-local setup scripts may create a convenience virtual environment, but +they must use these same Python and package pins rather than defining a second +dependency baseline. Native tools and GPU extensions remain separate. + +### The .NET SDK trap + +One trap worth naming, because its error message points the wrong way. The .NET +projects target `net10.0`, and a distribution's own `dotnet` under +`/usr/lib/dotnet` will shadow a newer SDK in `~/.dotnet` on `PATH`. The build +then fails with `NETSDK1045: The current .NET SDK does not support targeting +.NET 10.0`, which reads as a missing SDK when the SDK is usually installed and +merely second in line. Check with `dotnet --list-sdks` before installing +anything. Downgrading the projects to `net9.0` is not the fix: .NET 9 left +support in May 2026, and moving off end-of-life targets is why they are on +`net10.0`. + +## Compiled GPU extensions + +Some codecs additionally need compiled extensions that pip cannot resolve from a +version number alone, because each is built against one exact PyTorch and CUDA +build. Those are optional and separate, with install commands in +[`requirements-gpu.txt`](../requirements-gpu.txt): + +| Extra | Needed by | +|---|---| +| `cupy-cuda12x` | `n4mc`, `tsmc` | +| `torch-scatter` | `n4mc` | +| `nvdiffrast` | `n4mc` | +| `kaolin` | `n4mc`, `klt` | + +## What each module adds + +| Module | Adds | +|---|---| +| `codecs/tvmc` | .NET 10 SDK, CMake; Homebrew macOS or Ubuntu | +| `codecs/tsmc` | .NET 10 SDK, SAM3, `cupy`; Ubuntu 24.04, tested against Meta Quest 3. `convert_to_std_obj.py` runs inside Blender, which supplies `bpy` | +| `codecs/n4mc` | All four GPU extras and an NVIDIA GPU — 24 GB holds only about two training frames at resolution 256 | +| `codecs/qndf`, `codecs/qndf_int8` | An NVIDIA GPU for training. Evaluation (`mesh_errors.py`) runs on CPU. Building the `ssp_remesh` preprocessor needs CMake and Eigen (`libeigen3-dev`/`brew install eigen`), plus the pinned libigl submodule | +| `codecs/klt` | `kaolin` and an NVIDIA GPU; 24 GB is the same ceiling at resolution 128–256 | +| `codecs/draco` | A CMake build of the vendored Draco submodule. Open3D, pymeshlab, and OpenCV are for evaluation only | +| `codecs/vdmc`, `codecs/faster_vdmc` | The MPEG reference and optimized test models' own build requirements | +| `streaming` | Two hardware-synchronized RGB-D cameras, a Windows capture host, and an Ubuntu host with Python 3.10+, an NVIDIA GPU, and CUDA-enabled Open3D. Its legacy C++ pipeline additionally wants CUDA 12.x, Open3D 0.18, OpenCV, Eigen, jsoncpp, Draco, CMake, Ninja, and either the Azure Kinect SDK or the Orbbec K4A wrapper | +| `reconstruction/gs_tools`, `queen`, `3dgstream`, `vega` | The separate `open4d-gs` conda environment and five CUDA extensions built with `--no-build-isolation`, per [`gs_tools`](../open4d/reconstruction/gs_tools/README.md). Build on ext4; on an ntfs3 mount ninja deadlocks in `ntfs_file_write_iter` | +| `reconstruction/rerf` | Python 3.8, because `ac_dc/ncvv_ac_dc.cpython-38-*.so` ships without sources and cannot be rebuilt for a newer interpreter. Plus torch with CUDA, mmcv, bitarray, Pillow, NumPy — a separate environment from every other module here | +| `integrations/unity` | Unity, plus a C++ toolchain to rebuild the backend for anything other than the prebuilt macOS and Android/Quest 3 plugins | + +## RGB-D capture on Windows + +The RGB-D capture host is Windows and only encodes and forwards frames, so it +needs no NVIDIA GPU: just the camera vendor SDK (tested: Orbbec K4A Wrapper +1.10.5, SDK 1.10.28, two Femto Bolts), both cameras on separate USB 3 ports with +a sync hub, and an OpenSSH client. Close Orbbec Viewer first or the sender fails +with `Hardware MFT failed to start`. 5 synchronized pairs/s held over Wi-Fi and +VPN; 15 did not. + +Calibration layout and the step-by-step session walkthrough are in +[`open4d/streaming/README.md`](../open4d/streaming/README.md), +which covers how to run the pipeline and leaves requirements to this page. diff --git a/examples/streaming_demo.py b/examples/streaming_demo.py new file mode 100644 index 00000000..bdeeb73a --- /dev/null +++ b/examples/streaming_demo.py @@ -0,0 +1,150 @@ +#!/usr/bin/env python3 +"""Build a bundle, score its rungs, serve it, and play it over a bad link. + +The whole of `streamer` in one pass, on the ten basketball frames the TVMC +codec vendors. Runnable from anywhere: + + python examples/streaming_demo.py + +Needs three things beyond the base install: the optional `open4d-streamer` +package (``pip install -e open4d/streamer``), `open4d[draco]` for the Draco +rungs, and SciPy for the nearest-neighbour search the quality column uses. + +Output goes to ``examples/out/``, which the repository ignores. +""" + +from __future__ import annotations + +import json +import math +import sys +import urllib.request +from pathlib import Path + +import numpy as np +import open4d + +try: + import streamer + from streamer import link, playback, policy +except ImportError: # pragma: no cover - depends on the environment + sys.exit("this example needs: pip install -e open4d/streamer") + +from scipy.spatial import cKDTree + +ROOT = Path(__file__).resolve().parents[1] +SOURCE = ROOT / "open4d/codecs/tvmc/arap-volume-tracking/data/basketball_player" +OUT = ROOT / "examples/out/streaming-demo" +FPS = 10 +RUNGS = ["ply", "draco", "draco@11"] + + +def build() -> dict: + """Write the sequence as one clip at three qualities.""" + with open4d.load(SOURCE, fps=FPS) as sequence: + with streamer.Bundle(OUT, title="Basketball", source=SOURCE, fps=FPS) as clips: + clips.add(sequence, name="player", scene="basketball", rungs=RUNGS) + return json.loads((OUT / "view.json").read_text()) + + +def reference_frames() -> list[np.ndarray]: + with open4d.load(SOURCE, fps=FPS) as sequence: + return [np.asarray(sequence[i].geometry.positions, dtype=np.float64) + for i in range(len(sequence))] + + +def rung_frames(frames: list[str]) -> list[np.ndarray]: + """Vertex positions per frame, whichever format the rung is in.""" + if frames[0].endswith(".drc"): + import DracoPy + return [np.asarray(DracoPy.decode((OUT / f).read_bytes()).points, + dtype=np.float64) for f in frames] + # No `fps=`: write_sequence left a manifest here, and open4d refuses an + # override of timing a source already declares. + with open4d.load(OUT / Path(frames[0]).parent) as sequence: + return [np.asarray(sequence[i].geometry.positions, dtype=np.float64) + for i in range(len(sequence))] + + +def quality_db(reference: list[np.ndarray], decoded: list[np.ndarray]) -> float: + """RMS vertex deviation as dB below the model's diagonal. + + `streamer.metrics` scores *pixels* against a captured reference, and says + so for a mesh clip. Positions are what this bundle has, so this scores + those. Nearest-neighbour rather than index-wise: Draco merges duplicate + vertices, so the two sets are not the same length. + """ + span = float(np.linalg.norm(np.ptp(np.vstack(reference), axis=0))) + squared, count = 0.0, 0 + for ref, got in zip(reference, decoded): + distance, _ = cKDTree(ref).query(got, k=1) + squared += float(np.sum(distance ** 2)) + count += len(got) + rms = math.sqrt(squared / count) + return 20 * math.log10(span / rms) if rms else float("inf") + + +def ladder_of(clip: dict) -> list[policy.Rung]: + """Every rung, with its measured bitrate and its measured quality.""" + reference = reference_frames() + seconds = len(clip["frames"]) / FPS + renditions = [(clip["detail"]["rung"], clip["frames"])] + renditions += [(v["name"], v["frames"]) for v in clip["variants"]] + + rungs = [] + for name, frames in renditions: + size = sum((OUT / f).stat().st_size for f in frames) + score = quality_db(reference, rung_frames(frames)) + rungs.append(policy.Rung("player", name, size * 8 / seconds, + {"psnr": min(score, 99.0)})) + return sorted(rungs, key=lambda rung: rung.bits_per_second) + + +def serve_and_count(clip: dict) -> dict: + """Start the server, pull every frame over HTTP, read the counters.""" + server = streamer.serve(OUT, port=0, block=False, open_browser=False) + try: + base = f"http://127.0.0.1:{server.server_address[1]}/" + for frame in clip["frames"]: + urllib.request.urlopen(base + frame).read() + return server.monitor.snapshot() + finally: + server.shutdown() + + +def main() -> None: + clip = build()["clips"][0] + print(f"bundle {OUT.relative_to(ROOT)}/ — {len(clip['frames'])} frames, " + f"{clip['representation']}, {len(clip['variants']) + 1} rungs\n") + + ladder = ladder_of(clip) + print(f"{'rung':<10}{'Mbit/s':>9}{'quality dB':>12}") + for rung in ladder: + print(f"{rung.variant:<10}{rung.bits_per_second / 1e6:9.2f}" + f"{rung.quality['psnr']:12.1f}") + + counters = serve_and_count(clip) + print(f"\nserved {counters['requests']} requests, " + f"{counters['bytes'] / 1e6:.2f} MB over HTTP") + + print("\nwhat a budget buys:") + for budget in (2e6, 5e6, 20e6, 100e6): + chosen = policy.choose([ladder], budget=budget) + picked = (chosen.choices[0].variant if chosen.choices + else f"nothing — dropped {chosen.dropped[0]}") + print(f" {budget / 1e6:6.0f} Mbit/s -> {picked}") + + print("\n30 s over a link that collapses 30 -> 2.5 Mbit/s at t=15:") + clock = [0.0] + constrained = link.Link( + capacity=30e6, latency=0.03, clock=lambda: clock[0], + trace=link.Trace(at=(0.0, 15.0), capacity=(30e6, 2.5e6), loop=False)) + report = playback.Playback( + [ladder], constrained, fps=FPS, clock=lambda: clock[0]).run(30.0).as_dict() + for key in ("stalled_seconds", "frozen_seconds", "switches", "mean_quality"): + print(f" {key:<17}{report[key]}") + print(f" {'queueing':<17}{report['link']['queueing_fraction']:.1%} of the delay") + + +if __name__ == "__main__": + main() diff --git a/open4d/__init__.py b/open4d/__init__.py index b7396253..895cf6b7 100644 --- a/open4d/__init__.py +++ b/open4d/__init__.py @@ -8,18 +8,25 @@ NORMAL_DTYPE, POSITION_DTYPE, UV_DTYPE, + Dependency, + DependencyMode, Frame, FrameProvider, + GaussianCloud, + Geometry, MemoryFrameProvider, + PointCloud, + Representation, Sequence, SequenceView, TopologyMode, TriangleMesh, dtypes, ) -from ._api import load, save, unload, reconstruct +from ._api import load, save, unload, reconstruct, stream +from ._streamer import StreamerDependencyError from .codec import available_codecs, decode_sequence as decode, encode_sequence as encode -from .streaming import receive, send as stream +from .streaming import receive, send from .gaussians import GaussianSplats, GaussianRun, NeuralGaussianFrame, load_gaussians from .metrics import compare_meshes, compare_sequences from .visualization import visualize @@ -28,17 +35,24 @@ "ATTRIBUTE_FLOAT_DTYPE", "ATTRIBUTE_INT_DTYPE", "COLOR_DTYPE", + "Dependency", + "DependencyMode", "Frame", "GaussianSplats", "GaussianRun", "NeuralGaussianFrame", "FrameProvider", + "GaussianCloud", + "Geometry", "INDEX_DTYPE", "MemoryFrameProvider", "NORMAL_DTYPE", + "PointCloud", "POSITION_DTYPE", + "Representation", "Sequence", "SequenceView", + "StreamerDependencyError", "TopologyMode", "TriangleMesh", "UV_DTYPE", @@ -54,6 +68,7 @@ "receive", "stream", "save", + "send", "unload", "visualize", ] diff --git a/open4d/_api.py b/open4d/_api.py index 21106ad4..27d40704 100644 --- a/open4d/_api.py +++ b/open4d/_api.py @@ -163,3 +163,22 @@ def reconstruct(source, output=None, *, method="rgbd", **options): return reconstruct_gaussians(source, output, method=method, **options) raise ValueError("method must be 'rgbd', 'queen' or '3dgstream'") + + +def stream(source, *address, **options): + """Send mesh frames over TCP or export a sequence for browser playback. + + A path, or browser options such as out_dir/name/rungs, selects the optional + browser streamer. A frame iterable without browser options retains the + original TCP behavior, including positional or keyword host/port arguments. + Use send() to select TCP explicitly. + """ + browser_options = {"out_dir", "name", "title", "rungs", "fps", "open_browser", "block"} + if address or (not isinstance(source, (str, os.PathLike)) + and not browser_options.intersection(options)): + from .streaming import send + + return send(source, *address, **options) + from ._streamer import stream as stream_to_browser + + return stream_to_browser(source, **options) diff --git a/open4d/_streamer.py b/open4d/_streamer.py new file mode 100644 index 00000000..9f4e4b55 --- /dev/null +++ b/open4d/_streamer.py @@ -0,0 +1,151 @@ +"""Serve a sequence to a browser, if the optional streamer is installed. + +The dependency runs one way. `open4d-streamer` imports `open4d` -- for the +`Representation` vocabulary a bundle declares its clips in -- and Open4D must +not import it back, or the two become one package that cannot be released +separately and the reconstruction tree ends up in the base wheel. So this is a +verb in the public API whose implementation lives outside it, imported on the +call rather than at module load, the same arrangement `open4d.visualize` has +with Qt. + +What `stream` does is the two-step export-then-serve written as one line, +which is the common case: + + open4d.stream("capture.usdc") + +Everything it hides is reachable directly -- `streamer.Bundle` to collect +several clips, `streamer.serve` to serve a bundle already on disk, and +`streamer.Link` to constrain the link so the rate means something. +""" + +from __future__ import annotations + +import tempfile +from pathlib import Path +from typing import TYPE_CHECKING, Any, Sequence as TypingSequence + +if TYPE_CHECKING: # pragma: no cover - import cycle only matters to a checker + from open4d.core import Sequence + +__all__ = ["StreamerDependencyError", "stream"] + +#: Delivery, not interchange. Draco compresses this repository's mesh sequence +#: 12.9x, and a browser decodes it with the vendored WASM decoder; PLY at +#: 23 MB/s is an archive format that happens to be servable. `streamer.Rung` +#: parses the spec, so "draco@11" and a list of several also work. +DEFAULT_RUNGS: tuple[str, ...] = ("draco",) + + +class StreamerDependencyError(ImportError): + """`open4d.stream` was called without the streamer package installed.""" + + +def _require() -> Any: + """The `streamer` package, or an error saying how to get it. + + The message names the source checkout rather than an extra, because there + is no published `open4d-streamer` to resolve: the component is excluded + from distribution while its provenance review is open (see + ``THIRD_PARTY.md``). Advising ``pip install 'open4d[streamer]'`` would be + advice that fails. + """ + try: + import streamer + except ImportError as error: + raise StreamerDependencyError( + "open4d.stream needs the 'open4d-streamer' package, which is not " + "part of the base install -- it imports open4d, so open4d cannot " + "depend on it. From a source checkout:\n" + " python -m pip install -e open4d/streamer" + ) from error + return streamer + + +def stream( + source: "Sequence | Path | str", + *, + out_dir: Path | str | None = None, + name: str | None = None, + title: str | None = None, + rungs: TypingSequence[str] = DEFAULT_RUNGS, + fps: float | None = None, + host: str = "127.0.0.1", + port: int | None = None, + open_browser: bool = True, + block: bool = True, +) -> Any: + """Export ``source`` as a bundle and serve it to a browser. + + ``source`` is either a loaded `Sequence` or anything `open4d.load` reads. + ``rungs`` is the quality ladder: the first is the rendition a client plays + by default and the rest are what it can switch to, so a single-entry list + is a fixed-quality stream and says so. + + With ``out_dir`` omitted the bundle goes to a temporary directory, which is + **not** deleted afterwards. An export is minutes of encoding and the + directory is the artifact; removing it when the server stops would throw + away the expensive half of the call. The path is on the returned server as + ``server.bundle_dir``. + + Returns the running server, whose ``monitor`` counts what actually went + over the wire. With ``block=True`` it runs until interrupted. + """ + streamer = _require() + loaded = _is_sequence(source) + + if loaded and name is None: + raise ValueError( + "streaming a Sequence needs a name: it has no path to take one " + "from, and the name is what labels the pane in the viewer and " + "what the frame directory is called" + ) + + temporary = out_dir is None + if temporary: + out_dir = Path(tempfile.mkdtemp(prefix="open4d-stream-")) + out_dir = Path(out_dir).expanduser().resolve() + if temporary: + # Printed rather than only returned, because `serve(block=True)` does + # not return until it is interrupted -- and a caller who never named + # the directory would have no way to find it in the meantime. `serve` + # prints its URL for the same reason. + print(f"bundle: {out_dir}") + + if title is None: + title = name or (None if loaded else Path(str(source)).stem) or out_dir.name + + session = streamer.Bundle( + out_dir, + title=title, + source=None if loaded else source, + fps=int(fps) if fps else 30, + ) + with session: + if loaded: + session.add(source, name=name, rungs=rungs) + else: + session.add_source(source, name=name, fps=fps, rungs=rungs) + + server = streamer.serve( + out_dir, + host=host, + open_browser=open_browser, + block=block, + **({} if port is None else {"port": port}), + ) + # Recorded on the server for the same reason `serve` records the monitor + # there: with a temporary `out_dir` the caller never named the directory, + # and the export is the expensive half of this call. + server.bundle_dir = out_dir + return server + + +def _is_sequence(source: object) -> bool: + """Whether ``source`` is an in-memory sequence rather than a path to one. + + Structural rather than an `isinstance` against `open4d.core.Sequence`: + `SequenceView` and anything else satisfying the same protocol should be + streamable, and the two things this has to tell apart -- a sequence and a + path -- differ clearly enough that a length and an index is the whole test. + """ + return not isinstance(source, (str, Path)) and hasattr(source, "__len__") diff --git a/open4d/core/__init__.py b/open4d/core/__init__.py index cbcd1f43..9b6b806b 100644 --- a/open4d/core/__init__.py +++ b/open4d/core/__init__.py @@ -11,20 +11,38 @@ UV_DTYPE, ) from .frame import Frame -from .geometry import TriangleMesh -from .provider import FrameProvider, MemoryFrameProvider, TopologyMode +from .geometry import ( + GaussianCloud, + Geometry, + PointCloud, + Representation, + TriangleMesh, +) +from .provider import ( + Dependency, + DependencyMode, + FrameProvider, + MemoryFrameProvider, + TopologyMode, +) from .sequence import Sequence, SequenceView __all__ = [ "ATTRIBUTE_FLOAT_DTYPE", "ATTRIBUTE_INT_DTYPE", "COLOR_DTYPE", + "Dependency", + "DependencyMode", "Frame", "FrameProvider", + "GaussianCloud", + "Geometry", "INDEX_DTYPE", "MemoryFrameProvider", "NORMAL_DTYPE", + "PointCloud", "POSITION_DTYPE", + "Representation", "Sequence", "SequenceView", "TopologyMode", diff --git a/open4d/core/frame.py b/open4d/core/frame.py index 3362cdad..8ac44a4b 100644 --- a/open4d/core/frame.py +++ b/open4d/core/frame.py @@ -8,7 +8,7 @@ from types import MappingProxyType from typing import Any, Mapping -from .geometry import TriangleMesh +from .geometry import Geometry, Representation @dataclass(frozen=True, eq=False) @@ -17,7 +17,7 @@ class Frame: frame_index: int timestamp: float - geometry: TriangleMesh + geometry: Geometry metadata: Mapping[str, Any] = field(default_factory=dict) def __post_init__(self) -> None: @@ -32,8 +32,18 @@ def __post_init__(self) -> None: timestamp = float(self.timestamp) if not math.isfinite(timestamp): raise ValueError("timestamp must be finite") - if not isinstance(self.geometry, TriangleMesh): - raise TypeError("geometry must be a TriangleMesh") + # Any representation, not just triangles: see `open4d.core.geometry`. + # The check is against the protocol rather than a fixed tuple of types so + # that a new representation -- in this package or a third party's -- needs + # no edit here. + if not isinstance(self.geometry, Geometry): + raise TypeError( + "geometry must implement open4d.core.Geometry (a " + "`representation` property); got " + f"{type(self.geometry).__name__}" + ) + if not isinstance(self.geometry.representation, Representation): + raise TypeError("geometry.representation must be a Representation") if not isinstance(self.metadata, Mapping): raise TypeError("metadata must be a mapping") diff --git a/open4d/core/geometry.py b/open4d/core/geometry.py index f38393c5..dee6d307 100644 --- a/open4d/core/geometry.py +++ b/open4d/core/geometry.py @@ -3,8 +3,9 @@ from __future__ import annotations from dataclasses import dataclass, field +from enum import Enum from types import MappingProxyType -from typing import Mapping +from typing import Mapping, Protocol, runtime_checkable import numpy as np from numpy.typing import ArrayLike, NDArray @@ -32,6 +33,74 @@ def _finite(array: NDArray, name: str) -> None: raise ValueError(f"{name} must contain only finite values") +class Representation(str, Enum): + """What a decoded frame *is*, independent of how it was transported. + + This is the axis that decides which renderer can draw a frame and whether a + free camera is meaningful at all, and it is deliberately separate from the + codec that produced the frame: a triangle mesh is a mesh whether it arrived + as OBJ, as a Draco payload, or out of a V-DMC bitstream. Conflating the two + is what makes a viewer need a new branch per format instead of per + representation. + """ + + MESH = "mesh" + POINTS = "points" + GAUSSIANS = "gaussians" + #: Already-rendered pixels. Some representations cannot be decoded in the + #: consumer's process at all -- ReRF's entropy coder ships only as a + #: CPython 3.8 binary with no sources -- so server-rendered images are the + #: honest form for them rather than a fallback. The concrete type lands with + #: the camera model it needs in order to be comparable at a known pose; only + #: the taxonomy entry and :attr:`has_geometry` exist here, which is what a + #: consumer needs to gate on it today. + PIXELS = "pixels" + + @property + def has_geometry(self) -> bool: + """Whether a free camera can be aimed at this representation. + + False only for :attr:`PIXELS`, which is fixed to whichever camera + rendered it. Gate free-camera views on this rather than on a concrete + type, so that a mesh and a Gaussian cloud are treated alike without + either being named. + """ + return self is not Representation.PIXELS + + +@runtime_checkable +class Geometry(Protocol): + """What :class:`open4d.core.Frame` accepts as a frame's payload. + + One property, so a new representation can be added without editing `Frame`. + The concrete types in this module implement it; so may a third party's. + """ + + @property + def representation(self) -> Representation: + """Which :class:`Representation` this value is.""" + + +def _validate_point_attributes( + values: Mapping[str, ArrayLike], count: int +) -> dict[str, NDArray]: + """Validate attributes for a representation whose only alignment is per-point.""" + if not isinstance(values, Mapping): + raise TypeError("attributes must be a mapping") + result: dict[str, NDArray] = {} + for name, value in values.items(): + if not isinstance(name, str) or not name: + raise ValueError("attribute names must be non-empty strings") + if name in _RESERVED_ATTRIBUTES: + raise ValueError(f"{name!r} is a reserved attribute name") + attribute = _array(value, f"attribute {name!r}") + if attribute.ndim == 0 or attribute.shape[0] != count: + raise ValueError(f"attribute {name!r} must be point-aligned") + _finite(attribute, f"attribute {name!r}") + result[name] = dtypes.as_attribute(attribute, f"attribute {name!r}") + return result + + @dataclass(frozen=True, eq=False) class TriangleMesh: """A validated triangle mesh held in Open4D's canonical dtypes. @@ -91,6 +160,10 @@ def __post_init__(self) -> None: object.__setattr__(self, "texture_coordinates", texture_coordinates) object.__setattr__(self, "attributes", MappingProxyType(attributes)) + @property + def representation(self) -> Representation: + return Representation.MESH + @staticmethod def _validate_colors(value: ArrayLike | None, count: int) -> NDArray | None: if value is None: @@ -155,3 +228,153 @@ def _validate_attributes( _finite(attribute, f"attribute {name!r}") result[name] = dtypes.as_attribute(attribute, f"attribute {name!r}") return result + + +@dataclass(frozen=True, eq=False) +class PointCloud: + """Positions without connectivity, in the same canonical dtypes as a mesh. + + A distinct representation rather than a mesh with zero triangles: the RGB-D + fusion path and the point-cloud codecs in this repository produce exactly + this, and a consumer that wants to render points should not have to discover + the absence of connectivity by inspecting an empty array. + """ + + positions: NDArray + colors: NDArray | None = None + normals: NDArray | None = None + attributes: Mapping[str, NDArray] = field(default_factory=dict) + + def __post_init__(self) -> None: + positions = _array(self.positions, "positions") + if positions.ndim != 2 or positions.shape[1:] != (3,): + raise ValueError( + f"positions must have shape (N, 3); got {positions.shape}" + ) + _finite(positions, "positions") + positions = dtypes.as_positions(positions) + count = len(positions) + + object.__setattr__(self, "positions", positions) + object.__setattr__( + self, "colors", TriangleMesh._validate_colors(self.colors, count) + ) + object.__setattr__( + self, "normals", TriangleMesh._validate_normals(self.normals, count) + ) + object.__setattr__( + self, + "attributes", + MappingProxyType(_validate_point_attributes(self.attributes, count)), + ) + + @property + def representation(self) -> Representation: + return Representation.POINTS + + +@dataclass(frozen=True, eq=False) +class GaussianCloud: + """A frame of 3D Gaussians, stored activated and renderer-ready. + + Activated rather than raw -- world-space standard deviations, unit + quaternions, opacity already in [0, 1] -- because the raw parameterisation is + a training detail that differs between implementations. 3DGS stores + log-scale and pre-sigmoid opacity; a consumer should not have to know which + activation to apply, nor risk applying it twice. This is the same argument + `open4d.core.dtypes` makes for dtypes, one level up. + + ``rotations`` are ``(w, x, y, z)``, matching the order 3DGS writes into a PLY, + and are normalised on construction. + + ``colors`` is the view-*independent* term only. View-dependent appearance is + not representable here and is deliberately not faked: a method whose colour + is a hash grid or a higher-order SH expansion must either bake it for one + direction and declare that it did, or keep its own renderer. Storing a + degree-0 term and calling it the appearance is how a comparison quietly + stops being one. + """ + + positions: NDArray + scales: NDArray + rotations: NDArray + opacities: NDArray + colors: NDArray | None = None + attributes: Mapping[str, NDArray] = field(default_factory=dict) + + def __post_init__(self) -> None: + positions = _array(self.positions, "positions") + if positions.ndim != 2 or positions.shape[1:] != (3,): + raise ValueError( + f"positions must have shape (N, 3); got {positions.shape}" + ) + _finite(positions, "positions") + positions = dtypes.as_positions(positions) + count = len(positions) + + scales = _array(self.scales, "scales") + if scales.shape != (count, 3): + raise ValueError(f"scales must have shape ({count}, 3); got {scales.shape}") + _finite(scales, "scales") + if scales.size and scales.min() < 0: + raise ValueError( + "scales must be nonnegative: Open4D stores activated " + "world-space standard deviations, not log-scales" + ) + scales = dtypes.as_positions(scales, "scales") + + rotations = _array(self.rotations, "rotations") + if rotations.shape != (count, 4): + raise ValueError( + f"rotations must have shape ({count}, 4) as (w, x, y, z); " + f"got {rotations.shape}" + ) + _finite(rotations, "rotations") + rotations = dtypes.as_positions(rotations, "rotations") + norms = np.linalg.norm(rotations, axis=1, keepdims=True) + if rotations.size and float(norms.min()) == 0.0: + raise ValueError("rotations must not contain a zero quaternion") + # Normalised here so every consumer can skip it. A renderer that builds a + # covariance from a non-unit quaternion silently scales the Gaussian. + rotations = (rotations / norms).astype(rotations.dtype, copy=False) + + opacities = _array(self.opacities, "opacities") + if opacities.shape == (count, 1): + opacities = opacities.reshape(count) + if opacities.shape != (count,): + raise ValueError( + f"opacities must have shape ({count},) or ({count}, 1); " + f"got {opacities.shape}" + ) + _finite(opacities, "opacities") + if opacities.size and (opacities.min() < 0.0 or opacities.max() > 1.0): + raise ValueError( + "opacities must lie in [0, 1]: Open4D stores activated opacity, " + "not the pre-sigmoid parameter" + ) + opacities = dtypes.as_positions(opacities, "opacities") + + colors = self.colors + if colors is not None: + colors = _array(colors, "colors") + if colors.shape != (count, 3): + raise ValueError( + f"colors must have shape ({count}, 3); got {colors.shape}" + ) + _finite(colors, "colors") + colors = dtypes.as_colors(colors) + + object.__setattr__(self, "positions", positions) + object.__setattr__(self, "scales", scales) + object.__setattr__(self, "rotations", rotations) + object.__setattr__(self, "opacities", opacities) + object.__setattr__(self, "colors", colors) + object.__setattr__( + self, + "attributes", + MappingProxyType(_validate_point_attributes(self.attributes, count)), + ) + + @property + def representation(self) -> Representation: + return Representation.GAUSSIANS diff --git a/open4d/core/provider.py b/open4d/core/provider.py index e00c5cab..9c93bf26 100644 --- a/open4d/core/provider.py +++ b/open4d/core/provider.py @@ -3,7 +3,9 @@ from __future__ import annotations from collections.abc import Sequence as CollectionSequence +from dataclasses import dataclass from enum import Enum +from numbers import Integral import operator from types import MappingProxyType from typing import Any, Mapping, Protocol, runtime_checkable @@ -19,12 +21,119 @@ class TopologyMode(str, Enum): UNKNOWN = "unknown" +class DependencyMode(str, Enum): + """Whether a frame can be decoded on its own. + + :class:`Sequence` is lazy and random-access, which is the right contract for + a directory of files but assumes a frame is always reachable in one step. + Every real 4D codec here breaks that assumption in one of two ways, and a + consumer that wants to seek has to know which. + """ + + #: Every frame stands alone. A directory of per-frame OBJ or PLY files. + INDEPENDENT = "independent" + #: Frames depend on the most recent key frame and on each other in order, + #: as in Vega's group-of-volumes structure: a residual frame is undecodable + #: without its key. Seeking backwards is allowed, but only to a key frame. + GOP = "gop" + #: The decode stream is pulled in order and cannot be rewound at all, so + #: reaching an earlier frame means replaying from the start. ReRF's decoder + #: is this: `gs_tools.methods._rerf_rig_render` exists because of it. + SEQUENTIAL = "sequential" + + +@dataclass(frozen=True) +class Dependency: + """How frames in a sequence depend on one another. + + Declared by a provider so a consumer can plan a seek instead of discovering + mid-playback that frame 12 renders as garbage without frame 8. The default is + :attr:`DependencyMode.INDEPENDENT`, which is what a directory of files is and + what every existing provider gets without changing. + """ + + mode: DependencyMode = DependencyMode.INDEPENDENT + #: Ordinals that can be decoded without prior state. Required for + #: :attr:`DependencyMode.GOP` and meaningless otherwise. + key_frames: tuple[int, ...] = () + + def __post_init__(self) -> None: + if not isinstance(self.mode, DependencyMode): + raise TypeError("mode must be a DependencyMode") + keys = tuple(self.key_frames) + for key in keys: + if not isinstance(key, Integral) or isinstance(key, bool): + raise TypeError("key_frames must contain integers") + if key < 0: + raise ValueError("key_frames must be nonnegative") + keys = tuple(sorted({int(key) for key in keys})) + if self.mode is DependencyMode.GOP: + if not keys: + raise ValueError("GOP dependency requires at least one key frame") + if keys[0] != 0: + raise ValueError( + "GOP dependency requires frame 0 to be a key frame; a " + "sequence whose first frame cannot be decoded has no entry " + "point" + ) + elif keys: + raise ValueError(f"key_frames is meaningless for {self.mode.value}") + object.__setattr__(self, "key_frames", keys) + + def key_for(self, index: int) -> int | None: + """The key frame ``index`` decodes from, or None when it needs no key.""" + ordinal = operator.index(index) + if self.mode is not DependencyMode.GOP: + return None + candidates = [key for key in self.key_frames if key <= ordinal] + if not candidates: + raise IndexError(f"no key frame at or before {ordinal}") + return candidates[-1] + + def chain(self, index: int, *, decoded: int | None = None) -> tuple[int, ...]: + """Frames to decode, in order, to make ``index`` available. + + ``decoded`` is the ordinal the consumer's decoder is currently positioned + at, if any -- passing it lets a forward seek reuse that state instead of + restarting. Under :attr:`DependencyMode.SEQUENTIAL` a *backward* seek + cannot reuse it, because the stream does not rewind, so the chain comes + back as a full replay from zero. That asymmetry is the point of + declaring the mode at all. + + :attr:`DependencyMode.INDEPENDENT` always returns just ``index``: it + carries no decoder state, so whether the frame is already in a consumer's + cache is the consumer's business, not this model's. + """ + ordinal = operator.index(index) + if ordinal < 0: + raise ValueError("index must be nonnegative") + if decoded is not None: + decoded = operator.index(decoded) + if decoded < 0: + raise ValueError("decoded must be nonnegative") + + if self.mode is DependencyMode.INDEPENDENT: + return (ordinal,) + + if decoded == ordinal: + return () + + if self.mode is DependencyMode.GOP: + start = self.key_for(ordinal) + if decoded is not None and start <= decoded < ordinal: + start = decoded + 1 + else: # SEQUENTIAL + start = decoded + 1 if decoded is not None and decoded < ordinal else 0 + + return tuple(range(start, ordinal + 1)) + + @runtime_checkable class FrameProvider(Protocol): """Minimum random-access contract used by :class:`Sequence`. Providers may additionally expose ``metadata``, ``timestamps``, - ``topology``, ``has_constant_vertex_count``, + ``topology``, ``dependency``, ``has_constant_vertex_count``, ``has_vertex_correspondence``, and ``close``. Sequence consumes these declarations when present without requiring them from every provider. """ diff --git a/open4d/core/sequence.py b/open4d/core/sequence.py index 741d5154..92e72e41 100644 --- a/open4d/core/sequence.py +++ b/open4d/core/sequence.py @@ -10,7 +10,7 @@ from typing import Any, Mapping, overload from .frame import Frame -from .provider import FrameProvider, TopologyMode +from .provider import Dependency, DependencyMode, FrameProvider, TopologyMode class Sequence: @@ -37,6 +37,11 @@ def __init__(self, provider: FrameProvider) -> None: raise TypeError("provider topology must be a TopologyMode") self._topology = topology + dependency = getattr(provider, "dependency", None) or Dependency() + if not isinstance(dependency, Dependency): + raise TypeError("provider dependency must be a Dependency") + self._dependency = dependency + def __len__(self) -> int: return self._frame_count @@ -136,6 +141,32 @@ def fps(self) -> float | None: def topology(self) -> TopologyMode: return self._topology + @property + def dependency(self) -> Dependency: + """How frames in this sequence depend on one another. + + Advisory, not a restriction: :meth:`__getitem__` stays random-access, and + a provider whose codec needs prior state is expected to replay + internally to honour that. What this declares is the *cost and ordering* + of doing so -- which is exactly what a consumer needs to prefetch + sensibly, or to know that seeking backwards in a ReRF stream is a full + replay rather than a step. + """ + return self._dependency + + def decode_chain( + self, index: int, *, decoded: int | None = None + ) -> tuple[int, ...]: + """Frames to decode, in order, to reach ``index``. + + Convenience for ``sequence.dependency.chain(...)``; see + :meth:`open4d.core.Dependency.chain`. + """ + self._ensure_open() + if not 0 <= operator.index(index) < len(self): + raise IndexError("frame index out of range") + return self._dependency.chain(index, decoded=decoded) + @property def has_constant_topology(self) -> bool | None: if self.topology is TopologyMode.FIXED: @@ -193,6 +224,15 @@ def __init__(self, parent: Sequence, indices: range) -> None: if isinstance(fps, Real) and not isinstance(fps, bool) and math.isfinite(fps) and fps > 0: self.metadata = {**self.metadata, "fps": float(fps) / abs(indices.step)} self.topology = parent.topology + # Key-frame ordinals are the parent's, and a slice may start mid-group or + # skip frames, so they cannot be rebased onto this view in general. + # Declaring SEQUENTIAL is pessimistic but never wrong: a consumer walks + # the view in order, which is what it would do anyway. + self.dependency = ( + Dependency() + if parent.dependency.mode is DependencyMode.INDEPENDENT + else Dependency(mode=DependencyMode.SEQUENTIAL) + ) self.has_constant_vertex_count = parent.has_constant_vertex_count self.has_vertex_correspondence = parent.has_vertex_correspondence self.allow_nonmonotonic_timestamps = True diff --git a/open4d/core/tests/test_dependency.py b/open4d/core/tests/test_dependency.py new file mode 100644 index 00000000..60b7c598 --- /dev/null +++ b/open4d/core/tests/test_dependency.py @@ -0,0 +1,217 @@ +"""Contract tests for the dependency axis: what a seek actually costs.""" + +from __future__ import annotations + +import numpy as np +import pytest + +from open4d import ( + Dependency, + DependencyMode, + Frame, + Sequence, + TriangleMesh, +) + +pytestmark = pytest.mark.cpu + + +def mesh() -> TriangleMesh: + return TriangleMesh( + np.asarray([[0, 0, 0], [1, 0, 0], [0, 1, 0]], dtype=np.float32), + np.asarray([[0, 1, 2]], dtype=np.uint32), + ) + + +class StubProvider: + """A provider that declares a dependency. Frames themselves are irrelevant here. + + Deliberately not `MemoryFrameProvider`: that holds already-decoded frames, so + declaring anything but INDEPENDENT on it would be a lie. + """ + + def __init__(self, count: int, dependency: Dependency | None = None) -> None: + self._count = count + self.dependency = dependency + + @property + def frame_count(self) -> int: + return self._count + + def get_frame(self, index: int) -> Frame: + return Frame(index, float(index), mesh()) + + +# -------------------------------------------------------------- validation --- + + +def test_default_is_independent(): + assert Dependency().mode is DependencyMode.INDEPENDENT + assert Dependency().key_frames == () + + +def test_gop_requires_key_frames(): + with pytest.raises(ValueError, match="at least one key frame"): + Dependency(mode=DependencyMode.GOP) + + +def test_gop_requires_frame_zero_to_be_a_key(): + with pytest.raises(ValueError, match="no entry point"): + Dependency(mode=DependencyMode.GOP, key_frames=(4, 8)) + + +def test_key_frames_are_sorted_and_deduplicated(): + assert Dependency( + mode=DependencyMode.GOP, key_frames=(8, 0, 4, 4) + ).key_frames == (0, 4, 8) + + +@pytest.mark.parametrize( + "mode", [DependencyMode.INDEPENDENT, DependencyMode.SEQUENTIAL] +) +def test_key_frames_are_rejected_where_they_mean_nothing(mode): + with pytest.raises(ValueError, match="meaningless"): + Dependency(mode=mode, key_frames=(0,)) + + +def test_mode_must_be_a_dependency_mode(): + with pytest.raises(TypeError, match="DependencyMode"): + Dependency(mode="gop") + + +def test_key_frames_reject_negatives_and_bools(): + with pytest.raises(ValueError, match="nonnegative"): + Dependency(mode=DependencyMode.GOP, key_frames=(0, -1)) + with pytest.raises(TypeError, match="integers"): + Dependency(mode=DependencyMode.GOP, key_frames=(0, True)) + + +# ------------------------------------------------------------------ key_for --- + + +def test_key_for_finds_the_group_a_frame_belongs_to(): + dependency = Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8)) + assert [dependency.key_for(i) for i in range(10)] == [ + 0, 0, 0, 0, 4, 4, 4, 4, 8, 8 + ] + + +def test_key_for_is_none_when_no_key_is_involved(): + assert Dependency().key_for(3) is None + assert Dependency(mode=DependencyMode.SEQUENTIAL).key_for(3) is None + + +# -------------------------------------------------------------------- chain --- + + +def test_independent_frames_decode_alone(): + assert Dependency().chain(7) == (7,) + + +def test_independent_ignores_decoder_position(): + """It carries no state, so caching is the consumer's business, not the model's.""" + assert Dependency().chain(7, decoded=7) == (7,) + assert Dependency().chain(7, decoded=6) == (7,) + + +def test_gop_decodes_from_its_key_frame(): + dependency = Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8)) + assert dependency.chain(6) == (4, 5, 6) + assert dependency.chain(4) == (4,) + assert dependency.chain(0) == (0,) + + +def test_gop_reuses_decoder_state_on_a_forward_seek_inside_the_group(): + dependency = Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8)) + assert dependency.chain(6, decoded=5) == (6,) + assert dependency.chain(7, decoded=4) == (5, 6, 7) + + +def test_gop_restarts_at_the_key_when_seeking_backwards(): + dependency = Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8)) + assert dependency.chain(5, decoded=7) == (4, 5) + + +def test_gop_ignores_state_from_another_group(): + dependency = Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8)) + assert dependency.chain(9, decoded=2) == (8, 9) + + +def test_sequential_replays_from_zero_without_state(): + dependency = Dependency(mode=DependencyMode.SEQUENTIAL) + assert dependency.chain(3) == (0, 1, 2, 3) + + +def test_sequential_reuses_state_going_forward(): + dependency = Dependency(mode=DependencyMode.SEQUENTIAL) + assert dependency.chain(5, decoded=3) == (4, 5) + + +def test_sequential_backward_seek_is_a_full_replay(): + """The asymmetry that justifies the whole model: ReRF's stream cannot rewind.""" + dependency = Dependency(mode=DependencyMode.SEQUENTIAL) + assert dependency.chain(2, decoded=9) == (0, 1, 2) + + +@pytest.mark.parametrize("mode", [DependencyMode.GOP, DependencyMode.SEQUENTIAL]) +def test_a_frame_already_decoded_needs_no_work(mode): + keys = (0,) if mode is DependencyMode.GOP else () + assert Dependency(mode=mode, key_frames=keys).chain(5, decoded=5) == () + + +def test_chain_rejects_negative_ordinals(): + with pytest.raises(ValueError, match="nonnegative"): + Dependency().chain(-1) + with pytest.raises(ValueError, match="nonnegative"): + Dependency().chain(1, decoded=-1) + + +# ----------------------------------------------------------------- Sequence --- + + +def test_sequence_defaults_to_independent_for_existing_providers(): + """Every provider that predates this axis keeps working, unchanged.""" + assert Sequence(StubProvider(3)).dependency.mode is DependencyMode.INDEPENDENT + + +def test_sequence_surfaces_a_declared_dependency(): + declared = Dependency(mode=DependencyMode.GOP, key_frames=(0, 2)) + assert Sequence(StubProvider(4, declared)).dependency == declared + + +def test_sequence_rejects_a_non_dependency_declaration(): + with pytest.raises(TypeError, match="must be a Dependency"): + Sequence(StubProvider(2, "gop")) + + +def test_decode_chain_delegates_and_bounds_check(): + sequence = Sequence( + StubProvider(8, Dependency(mode=DependencyMode.GOP, key_frames=(0, 4))) + ) + assert sequence.decode_chain(6) == (4, 5, 6) + assert sequence.decode_chain(6, decoded=5) == (6,) + with pytest.raises(IndexError): + sequence.decode_chain(8) + + +def test_random_access_still_works_on_a_dependent_sequence(): + """`dependency` is advisory: the provider replays internally to honour indexing.""" + sequence = Sequence( + StubProvider(6, Dependency(mode=DependencyMode.SEQUENTIAL)) + ) + assert sequence[4].frame_index == 4 + + +def test_a_view_of_a_dependent_sequence_declares_the_safe_mode(): + """Key ordinals are the parent's and cannot be rebased onto a slice.""" + sequence = Sequence( + StubProvider(9, Dependency(mode=DependencyMode.GOP, key_frames=(0, 4, 8))) + ) + view = sequence[2:7] + assert view.dependency.mode is DependencyMode.SEQUENTIAL + assert view.dependency.key_frames == () + + +def test_a_view_of_an_independent_sequence_stays_independent(): + view = Sequence(StubProvider(5))[1:4] + assert view.dependency.mode is DependencyMode.INDEPENDENT diff --git a/open4d/core/tests/test_gaussians.py b/open4d/core/tests/test_gaussians.py index b9d23cd6..ae1fd985 100644 --- a/open4d/core/tests/test_gaussians.py +++ b/open4d/core/tests/test_gaussians.py @@ -171,6 +171,7 @@ def test_reconstruction_refuses_existing_output_before_launch(tmp_path, cli_runt def test_native_render_commands_do_not_call_destructive_extractor(monkeypatch, tmp_path): + monkeypatch.setitem(sys.modules, "streamer", None) root = Path(__file__).resolve().parents[2] / "reconstruction" / "gs_tools" monkeypatch.syspath_prepend(str(root)) try: diff --git a/open4d/core/tests/test_representation.py b/open4d/core/tests/test_representation.py new file mode 100644 index 00000000..f2801eef --- /dev/null +++ b/open4d/core/tests/test_representation.py @@ -0,0 +1,216 @@ +"""Contract tests for the representation axis: geometry kinds beyond triangles.""" + +from __future__ import annotations + +import numpy as np +import pytest + +from open4d import ( + Frame, + GaussianCloud, + Geometry, + PointCloud, + Representation, + TriangleMesh, +) + +pytestmark = pytest.mark.cpu + + +def points(count: int = 3) -> PointCloud: + return PointCloud(np.arange(count * 3, dtype=np.float32).reshape(count, 3)) + + +def gaussians(count: int = 2) -> GaussianCloud: + return GaussianCloud( + positions=np.zeros((count, 3), dtype=np.float32), + scales=np.ones((count, 3), dtype=np.float32), + rotations=np.tile(np.asarray([1, 0, 0, 0], dtype=np.float32), (count, 1)), + opacities=np.full(count, 0.5, dtype=np.float32), + ) + + +def mesh() -> TriangleMesh: + return TriangleMesh( + np.asarray([[0, 0, 0], [1, 0, 0], [0, 1, 0]], dtype=np.float32), + np.asarray([[0, 1, 2]], dtype=np.uint32), + ) + + +# ---------------------------------------------------------------- taxonomy --- + + +def test_each_type_declares_its_representation(): + assert mesh().representation is Representation.MESH + assert points().representation is Representation.POINTS + assert gaussians().representation is Representation.GAUSSIANS + + +def test_only_pixels_lacks_geometry(): + """The Explore/free-camera gate. Everything decodable to 3D admits a camera.""" + assert Representation.PIXELS.has_geometry is False + for value in ( + Representation.MESH, + Representation.POINTS, + Representation.GAUSSIANS, + ): + assert value.has_geometry is True + + +def test_representation_values_are_stable_strings(): + """These land in `view.json` and in URLs, so the wire values are contract.""" + assert [member.value for member in Representation] == [ + "mesh", + "points", + "gaussians", + "pixels", + ] + + +# ------------------------------------------------------------------- Frame --- + + +@pytest.mark.parametrize("geometry", [mesh(), points(), gaussians()]) +def test_frame_accepts_every_representation(geometry): + frame = Frame(0, 0.0, geometry) + assert frame.geometry is geometry + + +def test_frame_rejects_a_value_that_is_not_geometry(): + with pytest.raises(TypeError, match="open4d.core.Geometry"): + Frame(0, 0.0, object()) + + +def test_frame_rejects_a_bogus_representation(): + class Fake: + representation = "gaussians" # a string, not the enum + + assert isinstance(Fake(), Geometry) # satisfies the protocol structurally + with pytest.raises(TypeError, match="must be a Representation"): + Frame(0, 0.0, Fake()) + + +def test_a_third_party_type_needs_no_edit_to_core(): + class Voxels: + @property + def representation(self) -> Representation: + return Representation.POINTS + + assert Frame(0, 0.0, Voxels()).geometry.representation is Representation.POINTS + + +# -------------------------------------------------------------- PointCloud --- + + +def test_point_cloud_is_not_a_mesh_with_no_triangles(): + assert not hasattr(points(), "triangles") + + +def test_point_cloud_coerces_to_the_canonical_dtypes(): + cloud = PointCloud( + np.zeros((2, 3), dtype=np.float64), + colors=np.asarray([[255, 0, 0], [0, 255, 0]], dtype=np.uint8), + ) + assert cloud.positions.dtype == np.float32 + assert cloud.colors.dtype == np.float32 + assert cloud.colors.max() <= 1.0 + + +def test_point_cloud_rejects_misaligned_attributes(): + with pytest.raises(ValueError, match="point-aligned"): + PointCloud(np.zeros((3, 3), dtype=np.float32), attributes={"weight": [1.0, 2.0]}) + + +# ------------------------------------------------------------ GaussianCloud --- + + +def test_gaussian_rotations_are_normalised_on_construction(): + """A renderer building a covariance from a non-unit quaternion silently scales.""" + cloud = GaussianCloud( + positions=np.zeros((1, 3), dtype=np.float32), + scales=np.ones((1, 3), dtype=np.float32), + rotations=np.asarray([[0, 3, 4, 0]], dtype=np.float32), + opacities=np.asarray([1.0], dtype=np.float32), + ) + assert np.allclose(np.linalg.norm(cloud.rotations, axis=1), 1.0) + assert np.allclose(cloud.rotations, [[0, 0.6, 0.8, 0]]) + + +def test_gaussian_cloud_rejects_a_zero_quaternion(): + with pytest.raises(ValueError, match="zero quaternion"): + GaussianCloud( + positions=np.zeros((1, 3), dtype=np.float32), + scales=np.ones((1, 3), dtype=np.float32), + rotations=np.zeros((1, 4), dtype=np.float32), + opacities=np.asarray([1.0], dtype=np.float32), + ) + + +def test_gaussian_cloud_rejects_raw_log_scales(): + """Negative scale means the caller passed 3DGS's raw parameter, not a stddev.""" + with pytest.raises(ValueError, match="nonnegative"): + GaussianCloud( + positions=np.zeros((1, 3), dtype=np.float32), + scales=np.asarray([[-2.0, -2.0, -2.0]], dtype=np.float32), + rotations=np.asarray([[1, 0, 0, 0]], dtype=np.float32), + opacities=np.asarray([1.0], dtype=np.float32), + ) + + +@pytest.mark.parametrize("value", [-0.5, 1.5]) +def test_gaussian_cloud_rejects_unactivated_opacity(value): + with pytest.raises(ValueError, match=r"\[0, 1\]"): + GaussianCloud( + positions=np.zeros((1, 3), dtype=np.float32), + scales=np.ones((1, 3), dtype=np.float32), + rotations=np.asarray([[1, 0, 0, 0]], dtype=np.float32), + opacities=np.asarray([value], dtype=np.float32), + ) + + +def test_gaussian_cloud_accepts_a_column_of_opacities(): + """3DGS stores opacity as (N, 1); accepting it avoids a reshape at every call.""" + cloud = GaussianCloud( + positions=np.zeros((2, 3), dtype=np.float32), + scales=np.ones((2, 3), dtype=np.float32), + rotations=np.tile(np.asarray([1, 0, 0, 0], dtype=np.float32), (2, 1)), + opacities=np.full((2, 1), 0.25, dtype=np.float32), + ) + assert cloud.opacities.shape == (2,) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("scales", np.ones((3, 3), dtype=np.float32)), + ("rotations", np.tile(np.asarray([1, 0, 0, 0], dtype=np.float32), (3, 1))), + ("opacities", np.ones(3, dtype=np.float32)), + ], +) +def test_gaussian_cloud_requires_every_field_to_match_the_point_count(field, value): + kwargs = { + "positions": np.zeros((2, 3), dtype=np.float32), + "scales": np.ones((2, 3), dtype=np.float32), + "rotations": np.tile(np.asarray([1, 0, 0, 0], dtype=np.float32), (2, 1)), + "opacities": np.full(2, 0.5, dtype=np.float32), + field: value, + } + with pytest.raises(ValueError, match="must have shape"): + GaussianCloud(**kwargs) + + +def test_gaussian_cloud_is_structurally_immutable(): + cloud = gaussians() + with pytest.raises(Exception): + cloud.positions = np.zeros((2, 3), dtype=np.float32) + + +def test_empty_gaussian_cloud_is_valid(): + cloud = GaussianCloud( + positions=np.zeros((0, 3), dtype=np.float32), + scales=np.zeros((0, 3), dtype=np.float32), + rotations=np.zeros((0, 4), dtype=np.float32), + opacities=np.zeros(0, dtype=np.float32), + ) + assert cloud.representation is Representation.GAUSSIANS + assert len(cloud.positions) == 0 diff --git a/open4d/core/tests/test_temporal.py b/open4d/core/tests/test_temporal.py index 3bdc5c5e..1685b7bf 100644 --- a/open4d/core/tests/test_temporal.py +++ b/open4d/core/tests/test_temporal.py @@ -56,7 +56,7 @@ def test_frame_rejects_nonfinite_timestamps(value): def test_frame_requires_geometry_and_mapping_metadata(): - with pytest.raises(TypeError, match="TriangleMesh"): + with pytest.raises(TypeError, match="open4d.core.Geometry"): Frame(0, 0.0, object()) with pytest.raises(TypeError, match="mapping"): Frame(0, 0.0, mesh(), []) diff --git a/open4d/io/_api.py b/open4d/io/_api.py index 73ef8d2d..a0b3c297 100644 --- a/open4d/io/_api.py +++ b/open4d/io/_api.py @@ -491,10 +491,16 @@ def _write_frame( path: Path, frame: Frame, suffix: str, *, allow_lossy: bool = False ) -> Path: mesh = frame.geometry - present = {"positions", "triangles"} + # Not every representation has connectivity: a PointCloud carries positions + # and nothing to join them with, and `getattr` rather than attribute access + # is what lets one writer serve both without asking which it has. + triangles = getattr(mesh, "triangles", None) + present = {"positions"} + if triangles is not None: + present.add("triangles") present.update( name for name in ("colors", "normals", "texture_coordinates") - if getattr(mesh, name) is not None + if getattr(mesh, name, None) is not None ) present.update(mesh.attributes) unsupported = sorted(present - _OUTPUT_FIELDS[suffix]) @@ -509,13 +515,11 @@ def _write_frame( ) try: if suffix == ".obj": - return _mesh.write_obj(path, mesh.positions, mesh.triangles) + return _mesh.write_obj(path, mesh.positions, triangles) if suffix == ".ply": - return _mesh.write_ply( - path, mesh.positions, mesh.triangles, mesh.colors - ) + return _mesh.write_ply(path, mesh.positions, triangles, mesh.colors) return _mesh.write_with_trimesh( - path, mesh.positions, mesh.triangles, mesh.colors + path, mesh.positions, triangles, mesh.colors ) except ImportError as error: raise MissingDependencyError(str(error)) from error diff --git a/open4d/io/tests/test_public_stream_dispatch.py b/open4d/io/tests/test_public_stream_dispatch.py new file mode 100644 index 00000000..180f84a5 --- /dev/null +++ b/open4d/io/tests/test_public_stream_dispatch.py @@ -0,0 +1,38 @@ +"""Public TCP compatibility and optional browser streaming coexist.""" +from concurrent.futures import ThreadPoolExecutor +import sys + +import numpy as np +import pytest + +import open4d +from open4d.demo import mesh_sequence + + +@pytest.mark.parametrize("entry", ["send", "stream"]) +@pytest.mark.parametrize("address_style", ["positional", "keyword"]) +def test_public_tcp_round_trip_without_browser_dependency(monkeypatch, entry, address_style): + monkeypatch.setitem(sys.modules, "streamer", None) + with mesh_sequence(side=3, frames=2) as source: + with open4d.receive(port=0, timeout=3) as receiver, ThreadPoolExecutor(1) as pool: + if address_style == "positional": + result = pool.submit(getattr(open4d, entry), source, *receiver.address, + realtime=False, timeout=3) + else: + result = pool.submit(getattr(open4d, entry), source, + host=receiver.address[0], port=receiver.address[1], timeout=3) + restored = list(receiver) + assert result.result(timeout=3) == 2 + for expected, actual in zip(source, restored): + np.testing.assert_array_equal(actual.geometry.positions, expected.geometry.positions) + np.testing.assert_array_equal(actual.geometry.triangles, expected.geometry.triangles) + assert actual.timestamp == expected.timestamp + + +@pytest.mark.parametrize("use_path", [True, False]) +def test_browser_requests_report_missing_optional_streamer(monkeypatch, tmp_path, use_path): + monkeypatch.setitem(sys.modules, "streamer", None) + with mesh_sequence(side=3, frames=1) as frames: + source = tmp_path/'capture.ply' if use_path else frames + with pytest.raises(open4d.StreamerDependencyError, match="pip install -e open4d/streamer"): + open4d.stream(source, name="capture", out_dir=tmp_path/'bundle', open_browser=False, block=False) diff --git a/open4d/reconstruction/gs_tools/.gitignore b/open4d/reconstruction/gs_tools/.gitignore new file mode 100644 index 00000000..20f01264 --- /dev/null +++ b/open4d/reconstruction/gs_tools/.gitignore @@ -0,0 +1,9 @@ +# Upstream checkouts, cloned rather than vendored: 223 MB of embedded git +# repositories that `git add -A` in this directory otherwise sweeps into the +# index as gitlinks. That has happened twice, and both times the commit had to +# be amended, so it is stated here rather than remembered. +# +# These are *checkouts* with their own .git, unlike the other methods' upstream +# trees (`vega/`, `rerf/upstream/`) which are vendored copies and are tracked. +upstream/3dgstream/ +upstream/queen/ diff --git a/open4d/reconstruction/gs_tools/README.md b/open4d/reconstruction/gs_tools/README.md index 991f825c..7885a5e3 100644 --- a/open4d/reconstruction/gs_tools/README.md +++ b/open4d/reconstruction/gs_tools/README.md @@ -30,7 +30,170 @@ Then build the five CUDA extensions from `open4d/reconstruction/`, all with Build on ext4. On an ntfs3 mount ninja deadlocks in `ntfs_file_write_iter`. -## The viewer +## Comparing methods: Explore and Compare + +`gs-tools view` is a comparison tool, not a file browser. Clips are grouped by +**scene** and **method**, several methods share one viewport, and there are two +modes because there are two honest ways to put reconstructions side by side: + +**Explore** — a free camera, shared by every pane. Only methods that ship +Gaussians can appear, because there is nothing else to aim a free camera at. Any +difference you see between panes is the reconstruction, not the viewpoint. + +**Compare** — the scene's own **capture rig** as the camera path, with camera and +time scrubbed independently. Gaussian methods are rendered at the selected rig +pose; ReRF is rendered there too (`--rig-views`); and the captured photograph is +shown as a third method. This is the mode a volumetric representation and a +photograph can both join, and the one a PSNR/SSIM number could be attached to. +`A|B` wipes between two panes with a draggable handle. + +The camera path is not invented. `gs_tools/cameras.py` reads it out of the ORBIT +corpus's own `transforms.json` — the eight cameras that captured the scene — +which is what makes ground truth available at every station and means nothing +has to be registered: Vega's Gaussians are already in ORBIT world coordinates, +and ReRF is rendered at its own training views, which *are* those cameras. + + # everything, in one page: 9 Vega objects, ReRF at 8 rig cameras, the photographs + gs-tools view \ + -i results/vega-gaussian/prepared-final \ + /media/frozzzen/DataDrive/ORBIT_datasets_gaussian \ + ~/nevo_runs/g_basketball ~/nevo_runs/g_dancer ... \ + --bitstream rerf --rig-views 0 1 2 3 4 5 6 7 --render + + # just Vega, free camera + gs-tools view -i results/vega-gaussian/prepared-final --objects basketball + +The whole view state lives in the URL fragment, so a particular comparison is a +link: + + #scene=basketball&mode=compare&camera=0&methods=vega,rerf,captured&wipe=1 + +### What the comparison does and does not license + +- **No held-out view.** All eight cameras were training views for both Vega and + ReRF (`nevo_corpus.json` records no holdout, and Vega refines against all + eight). This measures reconstruction, not generalisation. A real held-out view + means re-preparing the corpus with a holdout and retraining. +- **Vega colour is baked** — see below. Its *geometry* at a rig pose is exact. +- **ReRF's framing differs from the captured pane.** The pose matches; the crop + does not, because ReRF renders at its training images' size and intrinsics + (1920x1080) while the corpus captured 4:3. Compare content, not pixel + positions. Aligning the intrinsics is the obvious next step and is not done. +- **The rig is eight coplanar cameras.** Neither silhouette carving nor a + photometric fit recovers what no camera saw, so a path far off that plane + makes every method look broken for reasons belonging to the capture. + `cameras.ring_path` stays on the ring. + +## Viewing output: Vega and ReRF + +`SIBR_gaussianViewer_app` below opens a 3DGS run directory on a Linux box with a +display. Two of the things in `open4d/reconstruction` cannot be opened that way +at all, for different reasons, and `gs-tools export` / `gs-tools view` are what +make them viewable: + +- **Vega** stores per-object `frame_XXXX.pt` chunks holding geometry only, with + colour in a hierarchical hash grid queried per Gaussian per view direction at + render time. Nothing but Vega can read it. The exporter drives Vega's own + `StreamingPlayer` and colour model and writes **one 3DGS PLY per frame**, so + the result opens in the bundled viewer, in SuperSplat, or in SIBR. +- **ReRF** (what the NeVo baseline vendors and streams) stores a DCT-coded, + arithmetic-coded feature voxel grid -- not Gaussians, so there is no PLY to + write. Its only decoder is its own, and that only runs under Python 3.8. The + adapter runs `rerf_render.py` and bundles the **images** it produces. + +- **QUEEN and 3DGStream** already store 3DGS PLYs, so the exporter copies them + and writes a manifest -- no decode, no conversion. Three unrelated on-disk + layouts are resolved by `gs_tools.outputs.gaussian_frames`: QUEEN's + `frames/NNNN/`, 3DGStream's per-frame `frameNNNNNN/point_cloud/iteration_N/`, + and a single frame's `point_cloud/iteration_N/`. The highest iteration is the + frame; 3DGStream's `added/` is skipped, since it holds only that frame's new + Gaussians rather than the scene. + +All land in the same shape -- a directory of frames plus a `view.json` -- and +`gs-tools view` serves it to the browser client in `streamer.client`, which is +the point: the GPU box usually has no display, and SIBR needs one (X11 +forwarding does not help, see below). The streaming half of this -- the manifest, +the server and that client -- lives in `../streamer`; this module produces +bundles and does not serve them. + + # a 3DGS run, its own PLYs, copied + gs-tools view -i ~/runs/coffee_martini --scene-name coffee_martini --method-name queen + + # the same run at 32 bytes per Gaussian instead of 248 + gs-tools export -i ~/runs/coffee_martini --frame-format splat -o /tmp/bundle + +`--frame-format splat` re-encodes to `gs_tools.io.splat`'s fixed 32 bytes: +measured 5.2x smaller on a degree-3 QUEEN run (73.8 MB -> 14.1 MB a frame). +Position and scale survive exactly; colour, opacity and rotation quantise to 8 +bits; **every SH band above degree 0 is dropped**, so appearance stops changing +with view direction. Worth it for delivery, wrong as an archive -- the PLY stays +the source of truth. + +`--scene-name` is what lets a 3DGS run share a `Compare` viewport with another +method's clip of the same subject. Without it the scene is named after the run +directory, because a 3DGS run records its subject nowhere reliable. Note that +`--scene-name` and `--method-name` apply to *every* source in one invocation, so +two runs of different subjects want two `export` calls. + + # what is this directory? + gs-tools inspect -i ~/nevo_runs/g_basketball + + # Vega: decode one object's 30 frames to PLY and serve them + gs-tools view -i results/vega-gaussian/prepared-final --objects basketball + + # every object in the catalog, as separate clips in one bundle + gs-tools view -i results/vega-gaussian/prepared-final + + # ReRF: bundle the renders already in the run (no GPU time) + gs-tools view -i ~/nevo_runs/g_basketball + + # ReRF: render a specific bitstream first (minutes of GPU time) + gs-tools view -i ~/nevo_runs/g_basketball --bitstream rerf --render + + # build a bundle without serving it, e.g. to copy elsewhere + gs-tools export -i ~/nevo_runs/g_basketball -o /tmp/bundle + +`view` binds loopback. Add `--host 0.0.0.0` to open it from another machine, and +note that the server has no authentication of any kind. Without `-o` the bundle +goes to a cache directory keyed by the source path, and a second `view` of the +same source reuses it; `--force` rebuilds. + +### What the exports do and do not preserve + +- **Vega colour is baked.** A PLY's `f_dc` is one colour per Gaussian; Vega's is + view-dependent. Colour is evaluated once from `--bake-azimuth` (default 0°) + and frozen, so orbiting in the viewer does not change appearance the way a + real Vega client would. Re-export at another azimuth to see it from + elsewhere. The export is `sh_degree 0` and says so in the viewer. +- **Vega geometry is exact.** Position, scale, rotation and opacity round-trip + bit-for-bit; verified against `diff_gaussian_rasterization` on the decoded + Gaussians. +- **ReRF gets no free camera.** The clip is whatever camera it was rendered at + -- upstream's orbit, or the rig cameras with `--rig-views`. Turning occupied + voxels into one Gaussian each would give a free camera, but its appearance + would not be what ReRF reconstructs, so it is not offered. +- **`--render_360 N` is not a full orbit.** Upstream computes + `angle = 2*pi*i/360`, so 30 frames sweep 29 degrees, and it advances time with + the camera -- the two cannot be separated. `--rig-views` renders at the capture + cameras instead, and advances the decode once per *timestep* rather than once + per image, so every view of one instant comes from the same decoded volume. +- **ReRF's codec settings are inferred, not remembered.** Upstream requires + `--pca`/`--pca_chs`/`--group_size` to match between compress and render and + nothing enforces it, so `gs_tools.methods.rerf.bitstream_info` reads them back + off the per-frame headers: entry count and channel split give the PCA + configuration, single-entry frames give the key frames. Override with + `--pca-chs`, `--group-size`, `--no-pca` if the inference is ever wrong. +- **`render_360_rerf_` does not name a bitstream.** Two bitstreams in one run + render to the same directory and the second overwrites the first. A bundle + keeps them apart; the run directory does not. This is why bundling existing + renders is the default and `--bitstream` is required to render one of several. + +ReRF runs in its own Python 3.8 environment (`conda activate nevo`; see +`../nevo/README.md`). `gs-tools` finds that interpreter as a sibling conda +environment of the current one -- override with `$OPEN4D_RERF_PYTHON` or +`--rerf-python`. + +## The SIBR viewer Linux only, and it needs a display with OpenGL 4.5. There is no macOS build, and X11 forwarding does not help: XQuartz offers indirect GLX at roughly OpenGL 2.1. diff --git a/open4d/reconstruction/gs_tools/gs_tools/cameras.py b/open4d/reconstruction/gs_tools/gs_tools/cameras.py new file mode 100644 index 00000000..72b8b89b --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/cameras.py @@ -0,0 +1,228 @@ +"""The capture rig, read back from the corpus, as the thing every method shares. + +Comparing two reconstructions means rendering them from the same place, and the +only place both were ever fit to is the rig that captured the scene. So the +canonical camera path here is not invented -- it is the eight ORBIT cameras, +read out of the corpus's own ``transforms.json``. Three things follow from that +choice, and all three are the reason for it: + +* **Ground truth exists at every station.** The captured image for a pose is + `frame_NNNNNN/images/view_NN.png`, so a comparison can put the real photograph + next to each method's render instead of only comparing methods to each other. +* **No convention translation.** ``camera_to_world_opencv`` is already the + (right, down, forward) frame the viewer's projection uses, so a pose can be + handed to the renderer as-is. The sibling ``transform_matrix`` is the OpenGL + form and is deliberately ignored. +* **Nothing has to be registered.** Vega's Gaussians are in ORBIT world + coordinates already, and ReRF's normalised volume maps back through + ``world_centre``/``world_scale`` in its corpus manifest, so both land in the + frame these poses are expressed in. + +The rig is eight coplanar cameras on a ring at subject height. That is a real +limitation on what a comparison can show -- neither silhouette carving nor a +photometric fit can recover what no camera saw -- so a path that climbs far off +that plane makes every method look broken for reasons that are the capture's, +not the method's. :func:`ring_path` stays on it. +""" + +from __future__ import annotations + +import json +import math +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +import numpy as np + + +@dataclass(frozen=True) +class Pose: + """One camera station, in ORBIT world coordinates. + + The three axes are the columns of the OpenCV camera-to-world rotation, named + for what they are on screen. `(right, down, forward)` is right-handed in that + order, which is what the viewer's projection assumes. + """ + + position: tuple[float, float, float] + right: tuple[float, float, float] + down: tuple[float, float, float] + forward: tuple[float, float, float] + #: Rig view index when this pose is a captured camera; None for a synthesised one. + view_id: int | None = None + + @classmethod + def from_c2w(cls, matrix, view_id: int | None = None) -> "Pose": + m = np.asarray(matrix, dtype=np.float64) + return cls( + position=tuple(float(v) for v in m[:3, 3]), + right=tuple(float(v) for v in m[:3, 0]), + down=tuple(float(v) for v in m[:3, 1]), + forward=tuple(float(v) for v in m[:3, 2]), + view_id=view_id, + ) + + def c2w(self) -> np.ndarray: + matrix = np.eye(4) + matrix[:3, 0] = self.right + matrix[:3, 1] = self.down + matrix[:3, 2] = self.forward + matrix[:3, 3] = self.position + return matrix + + def as_dict(self) -> dict[str, Any]: + return { + "position": list(self.position), + "right": list(self.right), + "down": list(self.down), + "forward": list(self.forward), + "view_id": self.view_id, + } + + +@dataclass +class Rig: + """A scene's capture rig and the volume it was pointed at.""" + + scene: str + width: int + height: int + fl_x: float + fl_y: float + cx: float + cy: float + poses: list[Pose] = field(default_factory=list) + bounds_min: list[float] | None = None + bounds_max: list[float] | None = None + + @property + def centre(self) -> np.ndarray: + if self.bounds_min is None or self.bounds_max is None: + return np.mean([pose.position for pose in self.poses], axis=0) + return (np.asarray(self.bounds_min) + np.asarray(self.bounds_max)) / 2.0 + + @property + def radius(self) -> float: + """Mean distance from the volume centre to a camera.""" + centre = self.centre + return float(np.mean([np.linalg.norm(np.asarray(p.position) - centre) for p in self.poses])) + + def fov_y(self) -> float: + """Vertical field of view in radians, which is what the viewer takes.""" + return 2.0 * math.atan(0.5 * self.height / self.fl_y) + + def as_dict(self) -> dict[str, Any]: + return { + "scene": self.scene, + "width": self.width, + "height": self.height, + "fov_y": self.fov_y(), + "bounds_min": self.bounds_min, + "bounds_max": self.bounds_max, + "poses": [pose.as_dict() for pose in self.poses], + } + + +def _first_transforms(scene_dir: Path) -> Path: + """A `transforms.json` for the scene -- the rig is fixed across frames. + + Per-frame first, because that is where the corpus actually writes them; the + object-level file `dataset.json` points at is the same rig. + """ + frames = sorted(p for p in scene_dir.glob("frame_*") if p.is_dir()) + for frame in frames: + candidate = frame / "transforms.json" + if candidate.is_file(): + return candidate + candidate = scene_dir / "transforms.json" + if candidate.is_file(): + return candidate + raise FileNotFoundError(f"no transforms.json under {scene_dir}") + + +def read_orbit_rig(scene_dir: Path | str) -> Rig: + """The rig for one ORBIT object, from its own `transforms.json`.""" + scene_dir = Path(scene_dir).expanduser().resolve() + data = json.loads(_first_transforms(scene_dir).read_text()) + views = sorted(data["frames"], key=lambda entry: entry.get("view_id", 0)) + poses = [ + # `camera_to_world_opencv`, not `transform_matrix`: the latter is the + # OpenGL form (y up, -z forward) and would silently flip the image. + Pose.from_c2w(entry["camera_to_world_opencv"], entry.get("view_id", index)) + for index, entry in enumerate(views) + ] + return Rig( + scene=scene_dir.name, + width=int(data["w"]), + height=int(data["h"]), + fl_x=float(data["fl_x"]), + fl_y=float(data["fl_y"]), + cx=float(data["cx"]), + cy=float(data["cy"]), + poses=poses, + bounds_min=data.get("bounds_min"), + bounds_max=data.get("bounds_max"), + ) + + +def ring_path(rig: Rig, samples: int) -> list[Pose]: + """`samples` poses evenly spaced on the rig's own ring. + + Used when a comparison wants a smoother sweep than eight stations. It stays + at the rig's radius and height on purpose: off that ring there is no captured + image to compare against, and below or above it there is no captured + *geometry* either -- see the module docstring. + + A sample that lands on a station keeps that station's `view_id`, so ground + truth is still addressable wherever it exists. + """ + if samples <= 0: + raise ValueError("samples must be positive") + centre = rig.centre + positions = np.asarray([pose.position for pose in rig.poses]) + height = float(np.mean(positions[:, 1])) + radius = float(np.mean(np.linalg.norm(positions[:, [0, 2]] - centre[[0, 2]], axis=1))) + + # Start where view 0 is, so sample 0 coincides with a real camera. + first = positions[0] - centre + start = math.atan2(float(first[0]), float(first[2])) + # Follow the rig's own direction of travel, so a sweep matches view order. + second = positions[1 % len(positions)] - centre + step = math.atan2(float(second[0]), float(second[2])) - start + step = (step + math.pi) % (2 * math.pi) - math.pi + direction = 1.0 if step >= 0 else -1.0 + + target = np.array([centre[0], height, centre[2]]) + stations = {index: pose for index, pose in enumerate(rig.poses)} + out: list[Pose] = [] + for index in range(samples): + # Land exactly on a station when the sampling divides the ring evenly. + on_station, remainder = divmod(index * len(rig.poses), samples) + if remainder == 0 and on_station in stations: + out.append(stations[on_station]) + continue + angle = start + direction * 2 * math.pi * index / samples + eye = np.array([ + centre[0] + radius * math.sin(angle), + height, + centre[2] + radius * math.cos(angle), + ]) + out.append(look_at(eye, target)) + return out + + +def look_at(eye, target, world_up=(0.0, 1.0, 0.0)) -> Pose: + """A pose looking from `eye` at `target`, in the (right, down, forward) frame.""" + eye = np.asarray(eye, dtype=np.float64) + forward = np.asarray(target, dtype=np.float64) - eye + forward /= np.linalg.norm(forward) + right = np.cross(forward, np.asarray(world_up, dtype=np.float64)) + right /= np.linalg.norm(right) + down = np.cross(forward, right) + return Pose( + position=tuple(float(v) for v in eye), + right=tuple(float(v) for v in right), + down=tuple(float(v) for v in down), + forward=tuple(float(v) for v in forward), + ) diff --git a/open4d/reconstruction/gs_tools/gs_tools/cli.py b/open4d/reconstruction/gs_tools/gs_tools/cli.py index b9115420..d95f051e 100644 --- a/open4d/reconstruction/gs_tools/gs_tools/cli.py +++ b/open4d/reconstruction/gs_tools/gs_tools/cli.py @@ -3,17 +3,40 @@ from __future__ import annotations import argparse +import hashlib import json +import os import sys from pathlib import Path -from . import env, paths, rast +from streamer import bundle +from streamer import server as view + +from . import env, outputs, paths, rast from .data import layouts from .io import manifest -from .methods import base, gstream, queen +from .methods import base, capture, gaussian, gstream, queen, rerf, vega METHODS = {"queen": queen, "3dgstream": gstream} +#: Exporter name -> the module implementing it. Which *kinds* each claims lives +#: in `gs_tools.outputs.EXPORTER_FOR`, next to the Kind enum, so that +#: `Detected.viewable` can answer "is there a path for this" without importing +#: this module -- and so the answer cannot disagree with what `--method auto` +#: picks. +EXPORTER_MODULES = { + "vega": vega, + "rerf": rerf, + "captured": capture, + "gaussian": gaussian, +} + +#: Kinds each exporter claims, derived so the two definitions cannot drift. +EXPORTERS = { + name: (module, tuple(k for k, v in outputs.EXPORTER_FOR.items() if v == name)) + for name, module in EXPORTER_MODULES.items() +} + def _spec(args: argparse.Namespace) -> base.RunSpec: return base.RunSpec( @@ -82,6 +105,188 @@ def _cmd_manifest(args: argparse.Namespace) -> int: return 0 +def _cmd_inspect(args: argparse.Namespace) -> int: + found = [outputs.detect(Path(entry).expanduser()) for entry in args.input] + for detected in found: + print(f"{detected.root}\n {outputs.describe(detected)}") + return 0 if all(detected.viewable for detected in found) else 1 + + +def _default_bundle_dir(sources: list[Path], kinds: list[outputs.Detected]) -> Path: + """Where a bundle goes when the user did not say. + + Not beside the sources: these outputs live on data mounts that are shared, + read-mostly, or (for a Vega catalog) somebody else's results directory, and + an export is derived data that can be regenerated. The cache directory is + keyed by the absolute source paths, so re-exporting the same set reuses the + same place instead of accumulating copies, and a different set gets its own. + """ + cache = Path(os.environ.get("XDG_CACHE_HOME", Path.home() / ".cache")) + digest = hashlib.sha1("\n".join(str(source) for source in sources).encode()).hexdigest()[:8] + if len(sources) == 1: + name = f"{kinds[0].kind.value}-{sources[0].name}-{digest}" + else: + name = f"mixed-{len(sources)}-sources-{digest}" + return cache / "open4d-gs-tools" / "view" / name + + +def _exporter(found: outputs.Detected, requested: str): + """The module that can export this output, or None if it needs none.""" + if found.kind is outputs.Kind.BUNDLE: + return None + if requested != "auto": + return EXPORTER_MODULES[requested] + claimed = outputs.EXPORTER_FOR.get(found.kind) + return EXPORTER_MODULES.get(claimed) if claimed else None + + +def _options_for(module, args: argparse.Namespace): + """Translate the shared flag set into the exporter's own options.""" + if module is capture: + return capture.CaptureOptions( + objects=tuple(args.objects or ()), + views=tuple(int(v) for v in args.views) if args.views else (), + frames=args.frames, + max_width=args.capture_width, + fps=args.fps, + ) + if module is gaussian: + return gaussian.GaussianExportOptions( + frames=args.frames, + frame_format=args.frame_format, + scene=args.scene_name, + method=args.method_name, + fps=args.fps, + ) + if module is vega: + return vega.VegaExportOptions( + objects=tuple(args.objects or ()), + frames=args.frames, + bake_azimuth_deg=args.bake_azimuth, + frame_format=args.frame_format, + device=args.device, + fps=args.fps, + ) + return rerf.RerfRenderOptions( + frames=args.frames, + depth=not args.no_depth, + bitstream=args.bitstream, + pca=None if args.pca is None else args.pca, + pca_chs=tuple(int(n) for n in args.pca_chs.split(",")) if args.pca_chs else None, + group_size=args.group_size, + config=Path(args.config).expanduser() if args.config else None, + fps=args.fps, + dry_run=getattr(args, "dry_run", False), + ) + + +def _export(args: argparse.Namespace) -> tuple[Path, list[outputs.Detected]]: + """Build (or reuse) one viewable bundle covering every `args.input`. + + Several sources land in one bundle rather than one each because comparing + them is the point: a Vega object and the ReRF render of the same subject are + two clips in one viewer, not two browser tabs. + """ + sources = [Path(entry).expanduser().resolve() for entry in args.input] + found = [outputs.detect(source) for source in sources] + for detected in found: + print(f"{detected.root}\n {outputs.describe(detected)}") + + unusable = [d.root for d in found if not d.viewable] + if unusable: + raise SystemExit("nothing viewable at " + ", ".join(str(path) for path in unusable)) + + bundles = [d for d in found if d.kind is outputs.Kind.BUNDLE] + if bundles: + if len(found) > 1: + raise SystemExit( + f"{bundles[0].root} is already a bundle; it cannot be combined with " + "other sources. Re-export from the original outputs instead." + ) + if getattr(args, "output", None): + print(" already a bundle; --output ignored") + return sources[0], found + + modules = [] + for detected in found: + module = _exporter(detected, args.method) + if module is None: + raise SystemExit( + f"no exporter for {detected.kind.value} at {detected.root}; --method takes " + + ", ".join(sorted(EXPORTERS)) + ) + modules.append(module) + + out_dir = ( + Path(args.output).expanduser().resolve() if args.output + else _default_bundle_dir(sources, found) + ) + existing = outputs.detect(out_dir) + if existing.kind is outputs.Kind.BUNDLE and not args.force: + # `view` calls this on every invocation, and decoding a Vega sequence or + # rendering a ReRF one is not something to repeat for a second look. + print(f" reusing bundle at {out_dir} ({outputs.describe(existing)}); --force to rebuild") + return out_dir, found + + titles: list[str] = [] + clips: list[bundle.Clip] = [] + detail: dict[str, object] = {} + for source, module in zip(sources, modules): + print(f" exporting {source.name} with {module.name} -> {out_dir}") + title, produced, produced_detail = module.build_clips( + source, out_dir, _options_for(module, args) + ) + titles.append(title) + clips += produced + detail.setdefault("sources", []).append( + {"path": str(source), "exporter": module.name, "detail": produced_detail} + ) + + if getattr(args, "dry_run", False) and not clips: + return out_dir, found + + # The rigs come from the corpus, not from any exporter: a shared camera is + # only shared if every method is handed the same one. + scenes = sorted({clip.scene for clip in clips if clip.scene}) + rigs = capture.rigs_for(args.corpus, scenes) if scenes else {} + missing = [scene for scene in scenes if scene not in rigs] + if missing: + print( + f" note: no rig in {args.corpus} for {', '.join(missing)} — those scenes " + "get no shared camera, so they are explore-only" + ) + bundle.write( + out_dir, + title=titles[0] if len(titles) == 1 else f"{len(clips)} clips from {len(sources)} sources", + source=", ".join(str(source) for source in sources), + clips=clips, + fps=args.fps, + scenes=rigs, + detail=detail, + ) + return out_dir, found + + +def _cmd_export(args: argparse.Namespace) -> int: + out_dir, _ = _export(args) + if getattr(args, "dry_run", False): + return 0 + print(f"\nbundle: {out_dir}\nview it with: gs-tools view -i {out_dir}") + return 0 + + +def _cmd_view(args: argparse.Namespace) -> int: + out_dir, _ = _export(args) + print() + view.serve( + out_dir, + host=args.host, + port=args.port, + open_browser=args.browser, + ) + return 0 + + def _cmd_depth_prior(args: argparse.Namespace) -> int: # Deliberately not implemented yet: it runs in the separate open4d-gs-midas # environment (timm==0.6.13), and wiring it before phase 1 has trained @@ -154,6 +359,84 @@ def build_parser() -> argparse.ArgumentParser: render.add_argument("passthrough", nargs="*", help=argparse.SUPPRESS) render.set_defaults(func=_cmd_render) + inspect = sub.add_parser("inspect", help="report what kind of output a directory holds") + inspect.add_argument("-i", "--input", required=True, nargs="+") + inspect.set_defaults(func=_cmd_inspect) + + def add_export_arguments(target: argparse.ArgumentParser) -> None: + target.add_argument("-i", "--input", required=True, nargs="+", + help="one or more of: a Vega bitstream/catalog, a ReRF run or " + "bitstream, a rendered image sequence, or (alone) an existing " + "bundle. Several sources become clips in one bundle.") + target.add_argument("-o", "--output", help="bundle directory (default: a cache directory)") + target.add_argument("--method", default="auto", choices=["auto", *sorted(EXPORTERS)]) + target.add_argument("--frames", type=int, help="export only the first N frames") + target.add_argument("--fps", type=int, default=30, help="playback rate recorded in the bundle") + # Vega + target.add_argument("--objects", nargs="+", + help="object names to export, for a Vega catalog or an ORBIT " + "corpus (default: all of them)") + target.add_argument("--corpus", default=str(capture.DEFAULT_CORPUS), + help="ORBIT corpus the capture rigs are read from, which is what " + "gives every method a shared camera") + target.add_argument("--views", nargs="+", + help="captured only: rig view ids to export (default: all 8)") + target.add_argument("--capture-width", type=int, default=1024, dest="capture_width", + help="captured only: longest edge of the exported images") + target.add_argument("--bake-azimuth", type=float, default=0.0, dest="bake_azimuth", + help="Vega only: camera azimuth in degrees that colour is baked from") + # 3DGS runs + target.add_argument("--frame-format", default="ply", choices=gaussian.FORMATS, + dest="frame_format", + help="Gaussian frames: 'ply' writes 3DGS PLY; 'splat' " + "re-encodes to 32 bytes per Gaussian, about half the " + "bytes but degree-0 colour only. For a Vega export that " + "is free — it bakes colour to degree 0 regardless") + target.add_argument("--scene-name", dest="scene_name", + help="3DGS runs only: subject name, so this clip lines up in " + "Compare with another method's clip of the same subject " + "(default: the run directory's name)") + target.add_argument("--method-name", dest="method_name", + help="3DGS runs only: method label in the viewer " + "(default: the run manifest's, else 'gaussian')") + target.add_argument("--device", help="Vega only: torch device for the colour decode") + # ReRF + target.add_argument("--force", action="store_true", + help="rebuild the bundle even if one is already there") + target.add_argument("--no-depth", action="store_true", dest="no_depth", + help="ReRF only: skip ReRF's depth maps") + target.add_argument("--bitstream", + help="ReRF only: which bitstream directory in the run to consider, " + "when a run holds several and none has been rendered. " + "Rendering one is `python -m rerf_stream.export`.") + target.add_argument("--config", help="ReRF only: ReRF config (default: /config.py)") + target.add_argument("--group-size", type=int, dest="group_size", + help="ReRF only: override the inferred ReRF key-frame interval") + target.add_argument("--pca-chs", dest="pca_chs", + help="ReRF only: override the inferred PCA channel split, e.g. 7,13") + target.add_argument("--pca", dest="pca", action="store_true", default=None, + help="ReRF only: force PCA decode on") + target.add_argument("--no-pca", dest="pca", action="store_false", + help="ReRF only: force PCA decode off") + target.add_argument("--dry-run", action="store_true", dest="dry_run", + help="ReRF only: print the render command without running it") + target.add_argument("passthrough", nargs="*", help=argparse.SUPPRESS) + + export = sub.add_parser( + "export", + help="convert a method's output into a viewable bundle (3DGS PLY or images)", + ) + add_export_arguments(export) + export.set_defaults(func=_cmd_export) + + viewer = sub.add_parser("view", help="export if needed, then serve the bundle to a browser") + add_export_arguments(viewer) + viewer.add_argument("--port", type=int, default=view.DEFAULT_PORT) + viewer.add_argument("--host", default="127.0.0.1", + help="0.0.0.0 to reach it from another machine (no authentication)") + viewer.add_argument("--browser", action="store_true", help="open a browser here") + viewer.set_defaults(func=_cmd_view) + show = sub.add_parser("manifest", help="print a run's manifest") show.add_argument("-m", "--run", required=True) show.set_defaults(func=_cmd_manifest) diff --git a/open4d/reconstruction/gs_tools/gs_tools/io/__init__.py b/open4d/reconstruction/gs_tools/gs_tools/io/__init__.py index 981ef85b..01007bc5 100644 --- a/open4d/reconstruction/gs_tools/gs_tools/io/__init__.py +++ b/open4d/reconstruction/gs_tools/gs_tools/io/__init__.py @@ -1,5 +1,15 @@ -"""Run manifests and representation IO.""" +"""Run manifests and representation IO. -from . import manifest +The viewable-bundle manifest moved to `streamer.bundle`: it is the contract +between a playback client and a server, so it belongs to the streaming module +rather than to a module that produces reconstructions. +""" -__all__ = ["manifest"] +from . import manifest, ply, splat + +#: The formats a Gaussian frame can be written in. Here rather than in one +#: exporter because two of them offer the choice, and two tuples of format +#: names is one that goes stale -- adding a format has to reach both. +GAUSSIAN_FORMATS = ("ply", "splat") + +__all__ = ["GAUSSIAN_FORMATS", "manifest", "ply", "splat"] diff --git a/open4d/reconstruction/gs_tools/gs_tools/io/ply.py b/open4d/reconstruction/gs_tools/gs_tools/io/ply.py new file mode 100644 index 00000000..f34bd68e --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/io/ply.py @@ -0,0 +1,234 @@ +"""The 3DGS PLY interchange format, read and written. + +Every Gaussian-splatting tool in the field reads the PLY that INRIA's original +3DGS writes, so it is the one format that lets a viewer -- SIBR, SuperSplat, +the bundled WebGL one in ``streamer.client`` -- open output from a method that +has never heard of it. That is what this module is for: the baselines in +``open4d/reconstruction`` each store Gaussians in their own container (Vega +ships ``frame_XXXX.pt`` chunks whose colour lives in a hash grid), and +normalising them to this PLY is what makes them viewable at all. + +The layout, from ``GaussianModel.construct_list_of_attributes``, is a +``binary_little_endian`` vertex element whose float32 properties are, in order:: + + x y z nx ny nz f_dc_0..2 f_rest_0..N opacity scale_0..2 rot_0..3 + +Three things about it are easy to get wrong, and all three are decisions this +module makes once: + +* **The values are raw, not activated.** ``opacity`` is a logit, ``scale`` is a + log, ``rot`` is an unnormalised quaternion. A viewer applies + sigmoid/exp/normalize itself. Writing activated values produces a file that + loads and renders as fog. +* **Colour is a spherical-harmonic coefficient, not RGB.** ``f_dc`` is the + degree-0 band, so ``rgb = 0.5 + C0 * f_dc``; :func:`rgb_to_sh_dc` is the + inverse. Storing RGB directly shifts and rescales every colour. +* **``f_rest`` is channel-major.** Upstream flattens + ``features_rest.transpose(1, 2)``, so the index is ``channel * n_bands + + band``, not the band-major order the name suggests. Only matters above + degree 0, which is why writing ``sh_rest=None`` is both allowed and the + common case here: a baked colour has no view-dependent bands to carry. +""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any + +import numpy as np + +#: The degree-0 spherical-harmonic constant, ``1 / (2 * sqrt(pi))``. +SH_C0 = 0.28209479177387814 + +_HEADER_MAGIC = b"ply" + + +def rgb_to_sh_dc(rgb: np.ndarray) -> np.ndarray: + """Linear RGB in [0, 1] -> the ``f_dc`` coefficients that decode back to it.""" + return (np.asarray(rgb, dtype=np.float32) - 0.5) / SH_C0 + + +def sh_dc_to_rgb(f_dc: np.ndarray) -> np.ndarray: + """``f_dc`` -> linear RGB, the inverse of :func:`rgb_to_sh_dc`.""" + return 0.5 + SH_C0 * np.asarray(f_dc, dtype=np.float32) + + +def attribute_names(n_rest: int = 0) -> list[str]: + """The property names, in file order, for ``n_rest`` per-channel SH bands.""" + names = ["x", "y", "z", "nx", "ny", "nz", "f_dc_0", "f_dc_1", "f_dc_2"] + names += [f"f_rest_{i}" for i in range(3 * n_rest)] + names += ["opacity", "scale_0", "scale_1", "scale_2"] + names += [f"rot_{i}" for i in range(4)] + return names + + +def _as_2d(array, name: str, width: int) -> np.ndarray: + values = np.asarray(array, dtype=np.float32) + if values.ndim == 1 and width == 1: + values = values[:, None] + if values.ndim != 2 or values.shape[1] != width: + raise ValueError(f"{name}: expected (N, {width}), got {tuple(values.shape)}") + return values + + +def write( + path: Path | str, + *, + xyz, + scale_raw, + rot_raw, + opacity_raw, + sh_dc, + sh_rest=None, + normals=None, +) -> Path: + """Write one 3DGS PLY. Values are raw (pre-activation); see the module docstring. + + ``sh_dc`` is ``(N, 3)`` degree-0 coefficients -- pass + ``rgb_to_sh_dc(rgb)`` if what you have is colour. ``sh_rest`` is + ``(N, n_bands, 3)`` and may be ``None``. + """ + path = Path(path) + xyz = _as_2d(xyz, "xyz", 3) + count = xyz.shape[0] + scale_raw = _as_2d(scale_raw, "scale_raw", 3) + rot_raw = _as_2d(rot_raw, "rot_raw", 4) + opacity_raw = _as_2d(opacity_raw, "opacity_raw", 1) + sh_dc = _as_2d(sh_dc, "sh_dc", 3) + + for name, values in ( + ("scale_raw", scale_raw), + ("rot_raw", rot_raw), + ("opacity_raw", opacity_raw), + ("sh_dc", sh_dc), + ): + if values.shape[0] != count: + raise ValueError(f"{name}: {values.shape[0]} rows, but xyz has {count}") + + if normals is None: + # Upstream writes zeros here and no renderer reads them; a Gaussian has + # no surface normal to record in the first place. + normals = np.zeros((count, 3), dtype=np.float32) + else: + normals = _as_2d(normals, "normals", 3) + + columns = [xyz, normals, sh_dc] + n_rest = 0 + if sh_rest is not None: + rest = np.asarray(sh_rest, dtype=np.float32) + if rest.ndim != 3 or rest.shape[0] != count or rest.shape[2] != 3: + raise ValueError(f"sh_rest: expected (N, bands, 3), got {tuple(rest.shape)}") + n_rest = rest.shape[1] + # Channel-major, matching upstream's `transpose(1, 2).flatten(start_dim=1)`. + columns.append(np.ascontiguousarray(rest.transpose(0, 2, 1)).reshape(count, -1)) + columns += [opacity_raw, scale_raw, rot_raw] + + table = np.concatenate(columns, axis=1).astype(" NumPy codes, for the mixed-type headers real +#: producers write. QUEEN adds an ``int vertex_id`` column to its Gaussians, so +#: assuming every property is float32 makes its output unreadable. +_PLY_TYPES = { + "float": "f4", "float32": "f4", "float64": "f8", "double": "f8", + "char": "i1", "int8": "i1", "uchar": "u1", "uint8": "u1", + "short": "i2", "int16": "i2", "ushort": "u2", "uint16": "u2", + "int": "i4", "int32": "i4", "uint": "u4", "uint32": "u4", +} + + +def _read_header(handle) -> tuple[int, list[tuple[str, str]], bool]: + """(vertex count, [(property, NumPy code)], little-endian) from an open file.""" + if handle.read(3) != _HEADER_MAGIC: + raise ValueError("not a PLY file") + handle.seek(0) + count: int | None = None + names: list[tuple[str, str]] = [] + little = True + in_vertex = False + while True: + line = handle.readline() + if not line: + raise ValueError("PLY header has no end_header") + text = line.decode("ascii", "replace").strip() + if text.startswith("format"): + if "ascii" in text: + raise ValueError("ascii PLY is not supported; 3DGS writes binary") + little = "little_endian" in text + elif text.startswith("element "): + parts = text.split() + in_vertex = parts[1] == "vertex" + if in_vertex: + count = int(parts[2]) + elif text.startswith("property ") and in_vertex: + parts = text.split() + if parts[1] == "list": + raise ValueError( + f"list properties are not supported in a vertex element: {text!r}" + ) + code = _PLY_TYPES.get(parts[1]) + if code is None: + raise ValueError(f"unsupported property type in {text!r}") + names.append((parts[2], code)) + elif text == "end_header": + break + if count is None: + raise ValueError("PLY header declares no vertex element") + return count, names, little + + +def count(path: Path | str) -> int: + """The Gaussian count, from the header alone -- no payload is read.""" + with Path(path).open("rb") as handle: + return _read_header(handle)[0] + + +def read(path: Path | str) -> dict[str, Any]: + """Read a 3DGS PLY back into raw arrays, the inverse of :func:`write`.""" + path = Path(path) + with path.open("rb") as handle: + n, properties, little = _read_header(handle) + order = "<" if little else ">" + # A structured dtype rather than one flat float32 block: the properties + # are not all the same width, and reading them as if they were shifts + # every column after the first odd one. + record = np.dtype([(name, order + code) for name, code in properties]) + rows = np.frombuffer(handle.read(n * record.itemsize), dtype=record, count=n) + names = [name for name, _ in properties] + present = set(names) + + def take(keys: list[str]) -> np.ndarray: + missing = [key for key in keys if key not in present] + if missing: + raise ValueError(f"{path.name} is missing {', '.join(missing)}") + return np.stack( + [rows[key].astype(np.float32) for key in keys], axis=-1 + ) + + n_rest = sum(1 for name in names if name.startswith("f_rest_")) // 3 + result: dict[str, Any] = { + "count": n, + "xyz": take(["x", "y", "z"]), + "sh_dc": take(["f_dc_0", "f_dc_1", "f_dc_2"]), + "opacity_raw": take(["opacity"]), + "scale_raw": take(["scale_0", "scale_1", "scale_2"]), + "rot_raw": take([f"rot_{i}" for i in range(4)]), + "sh_degree": int(round((n_rest + 1) ** 0.5)) - 1 if n_rest else 0, + } + if n_rest: + rest = take([f"f_rest_{i}" for i in range(3 * n_rest)]) + result["sh_rest"] = rest.reshape(n, 3, n_rest).transpose(0, 2, 1) + return result diff --git a/open4d/reconstruction/gs_tools/gs_tools/io/splat.py b/open4d/reconstruction/gs_tools/gs_tools/io/splat.py new file mode 100644 index 00000000..21095e3a --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/io/splat.py @@ -0,0 +1,160 @@ +"""The 32-byte-per-Gaussian ``.splat`` frame format. + +A 3DGS PLY is an interchange format, not a delivery one. It stores every +attribute as float32 and every spherical-harmonic band it was trained with, so +one frame of a degree-3 run is 248 bytes per Gaussian -- and a 30-frame clip of +a few hundred thousand Gaussians is hundreds of megabytes, which is the reason +a bundle cannot simply be downloaded. This is the same content quantised to a +fixed 32 bytes: + +=========== ====== ======================================================= +byte range type meaning +=========== ====== ======================================================= +``0..11`` 3f32 position, world space +``12..23`` 3f32 scale, **activated** -- world-space standard deviations +``24..27`` 4u8 colour RGB and opacity, ``round(v * 255)`` +``28..31`` 4u8 rotation ``(w, x, y, z)``, ``round(q * 128) + 128`` +=========== ====== ======================================================= + +The rotation encoding is not symmetric: ``q = 1`` would be 256, which clamps to +255 and comes back as 0.992. A reader must therefore **renormalise** the +quaternion, which both readers here do -- building a covariance from a non-unit +one scales every Gaussian by about 1.6%. + +Three consequences worth stating plainly, because each is a real loss: + +* **Every SH band above degree 0 is dropped.** View-dependent appearance goes + with them, so a frame written here looks the same from every direction. That + is the same caveat the Vega exporter already carries for its baked colour -- + and for Vega it costs nothing, because that colour was baked before it ever + reached a PLY. For a QUEEN or 3DGStream run it is a genuine reduction, which + is why writing this is opt-in rather than the default. +* **Values are stored activated**, the opposite of the PLY convention. A reader + must *not* apply exp/sigmoid/normalize; :func:`from_ply` is where that + happens, once. Getting this backwards renders as fog either way, so the two + formats are deliberately not interchangeable byte-for-byte. +* **Rotation keeps 8 bits per component.** Enough for a unit quaternion at this + scale, but it is quantisation, and re-exporting from a ``.splat`` would + compound it. Bundles keep the PLY as the source of truth. + +The byte layout matches the widely used ``.splat`` files by field order and +size. Interoperability with other tools that read the extension is *not* +verified here -- nothing in this repository reads one -- so the layout above, +not any external tool, is the specification this module implements. +""" + +from __future__ import annotations + +from pathlib import Path + +import numpy as np +from open4d.core import GaussianCloud + +from . import ply + +#: Bytes per Gaussian. Fixed, so a frame's Gaussian count is its file size / 32. +SPLAT_BYTES = 32 + +#: Rotation is stored as ``round(q * QUANT_SCALE) + QUANT_OFFSET``. +QUANT_SCALE = 128.0 +QUANT_OFFSET = 128 + + +def encode(cloud: GaussianCloud) -> bytes: + """One frame of Gaussians as ``.splat`` bytes. + + Takes `open4d.core.GaussianCloud` rather than loose arrays because that type + already guarantees what this encoding assumes and cannot check cheaply: + scales activated and nonnegative, quaternions unit, opacity in [0, 1]. + """ + if not isinstance(cloud, GaussianCloud): + raise TypeError("cloud must be an open4d.core.GaussianCloud") + count = len(cloud.positions) + if cloud.colors is None: + raise ValueError( + "cloud has no colors; .splat carries a single RGB per Gaussian, so " + "there is nothing to write without one" + ) + + payload = np.empty((count, SPLAT_BYTES), dtype=np.uint8) + floats = payload[:, :24].view(np.float32).reshape(count, 6) + floats[:, :3] = cloud.positions + floats[:, 3:] = cloud.scales + + payload[:, 24:27] = np.clip(np.rint(cloud.colors * 255.0), 0, 255).astype(np.uint8) + payload[:, 27] = np.clip(np.rint(cloud.opacities * 255.0), 0, 255).astype(np.uint8) + payload[:, 28:32] = np.clip( + np.rint(cloud.rotations * QUANT_SCALE) + QUANT_OFFSET, 0, 255 + ).astype(np.uint8) + return payload.tobytes() + + +def write(path: Path | str, cloud: GaussianCloud) -> Path: + """Write ``cloud`` as a ``.splat`` frame and return the path.""" + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_bytes(encode(cloud)) + return path + + +def count(path: Path | str) -> int: + """Gaussians in a ``.splat`` file, from its size alone.""" + size = Path(path).stat().st_size + if size % SPLAT_BYTES: + raise ValueError( + f"{path} is {size} bytes, not a multiple of {SPLAT_BYTES}; " + "this is not a .splat frame" + ) + return size // SPLAT_BYTES + + +def decode(data: bytes) -> GaussianCloud: + """``.splat`` bytes back to a `GaussianCloud`, for tests and round-trips.""" + if len(data) % SPLAT_BYTES: + raise ValueError(f"{len(data)} bytes is not a multiple of {SPLAT_BYTES}") + payload = np.frombuffer(data, dtype=np.uint8).reshape(-1, SPLAT_BYTES) + floats = payload[:, :24].copy().view(np.float32).reshape(-1, 6) + # Renormalised, because the quantisation is not symmetric: w = 1 encodes as + # round(128) + 128 = 256, which clamps to 255 and decodes to 0.992. Without + # this the covariance is built from a non-unit quaternion and every Gaussian + # is scaled by ~1.6% -- small, wrong, and invisible until measured. + rotations = (payload[:, 28:32].astype(np.float32) - QUANT_OFFSET) / QUANT_SCALE + norms = np.linalg.norm(rotations, axis=1, keepdims=True) + rotations = np.divide( + rotations, norms, out=np.zeros_like(rotations), where=norms > 0 + ) + return GaussianCloud( + positions=floats[:, :3].copy(), + scales=floats[:, 3:].copy(), + rotations=rotations, + opacities=payload[:, 27].astype(np.float32) / 255.0, + colors=payload[:, 24:27].astype(np.float32) / 255.0, + ) + + +def from_ply(path: Path | str) -> GaussianCloud: + """A 3DGS PLY as a `GaussianCloud`, with the activations applied once. + + This is the bridge every Gaussian producer in this repository crosses to + reach core's canonical form: PLY stores raw training parameters, core stores + activated ones, and doing the conversion here means no exporter has to + remember which is which. + + Only the degree-0 band is read. ``f_rest`` is left on the floor, which is + the loss this format's docstring describes. + """ + fields = ply.read(path) + scales = np.exp(fields["scale_raw"]) + opacities = 1.0 / (1.0 + np.exp(-fields["opacity_raw"].reshape(-1))) + rotations = fields["rot_raw"] + norms = np.linalg.norm(rotations, axis=1, keepdims=True) + # A zero quaternion is a corrupt file rather than something to normalise; the + # GaussianCloud constructor rejects it, and its message says so. + rotations = np.divide(rotations, norms, out=np.zeros_like(rotations), where=norms > 0) + return GaussianCloud( + positions=fields["xyz"], + scales=scales, + rotations=rotations, + opacities=opacities, + colors=np.clip(ply.sh_dc_to_rgb(fields["sh_dc"].reshape(-1, 3)), 0.0, 1.0), + ) diff --git a/open4d/reconstruction/gs_tools/gs_tools/methods/__init__.py b/open4d/reconstruction/gs_tools/gs_tools/methods/__init__.py index 16a080fc..8ccc9544 100644 --- a/open4d/reconstruction/gs_tools/gs_tools/methods/__init__.py +++ b/open4d/reconstruction/gs_tools/gs_tools/methods/__init__.py @@ -3,8 +3,24 @@ Each adapter translates Open4D's arguments into an upstream CLI and runs it as a subprocess. They cannot share a process: both upstreams use flat imports and both define `scene`, `utils`, and `arguments`. + +`queen` and `gstream` wrap trainers. `vega` and `rerf` wrap the *output* side +only -- both are produced by their own CLIs under +`open4d/reconstruction/{vega,nevo}`, and what was missing was any way to look at +what those produced. See their module docstrings for why one exports Gaussians +and the other cannot. """ -from . import base, gstream, queen +from importlib import import_module + +__all__ = ["base", "gstream", "queen", "rerf", "vega"] + -__all__ = ["base", "gstream", "queen"] +def __getattr__(name): + # Building a QUEEN/3DGStream command must not import the optional browser + # bundle exporters, which depend on the separately installed streamer. + if name not in __all__: + raise AttributeError(f"module {__name__!r} has no attribute {name!r}") + module = import_module(f".{name}", __name__) + globals()[name] = module + return module diff --git a/open4d/reconstruction/gs_tools/gs_tools/methods/capture.py b/open4d/reconstruction/gs_tools/gs_tools/methods/capture.py new file mode 100644 index 00000000..bdbcbb39 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/methods/capture.py @@ -0,0 +1,197 @@ +"""The captured images, as a method in their own right. + +Every other adapter here reconstructs; this one just reads the photographs back +out of the corpus. It is what turns a comparison from "which of these two +renders do you prefer" into "which is closer to what the camera saw", so it is +worth having even though it computes nothing. + +One clip per rig station, because a captured image only exists where a camera +was. That is also why `gs_tools.cameras` makes the rig the canonical path: at +those eight poses, and only there, every method can be lined up against the +truth. + +A caveat that matters for how far these numbers can be pushed: for the runs in +this repository *all eight cameras were training views* for both Vega and ReRF +(`nevo_corpus.json` records no held-out view, and Vega's photometric refinement +fits all eight). So a comparison here measures reconstruction, not +generalisation. A real held-out view means re-preparing the corpus with a +holdout and retraining. +""" + +from __future__ import annotations + +import json +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +from .. import cameras +from streamer import bundle + +name = "captured" + +#: Where the ORBIT Gaussian-training corpus lives on the machine that built it. +DEFAULT_CORPUS = Path("/media/frozzzen/DataDrive/ORBIT_datasets_gaussian") + + +@dataclass +class CaptureOptions: + """What to pull out of the corpus.""" + + #: Object names, when the input is a whole corpus; empty means all of them. + objects: tuple[str, ...] = () + #: Rig view ids to export; empty means every station. + views: tuple[int, ...] = () + #: How many frames; None means all of them. + frames: int | None = None + #: Longest edge of the exported image. The corpus is 4096x3072, which is + #: both far more than a comparison pane shows and ~1.5 MB a frame. + max_width: int = 1024 + quality: int = 90 + fps: int = 30 + extra: dict[str, Any] = field(default_factory=dict) + + +def scenes_in(corpus: Path | str) -> list[str]: + """Object names in an ORBIT corpus, from its own index.""" + corpus = Path(corpus).expanduser().resolve() + index = corpus / "dataset.json" + if index.is_file(): + return [entry["name"] for entry in json.loads(index.read_text()).get("objects", [])] + return sorted(p.name for p in corpus.iterdir() if p.is_dir()) + + +def frame_dirs(scene_dir: Path) -> list[Path]: + """The corpus's per-frame directories, in time order. + + Named `frame_000001` upward -- one-based, unlike every index this module + hands out, so the offset is resolved here and nowhere else. + """ + return sorted( + (p for p in scene_dir.glob("frame_*") if p.is_dir() and p.name[6:].isdigit()), + key=lambda p: int(p.name[6:]), + ) + + +def build_clips( + source: Path | str, out_dir: Path | str, options: CaptureOptions | None = None +) -> tuple[str, list[bundle.Clip], dict[str, Any]]: + """Copy captured views into ``out_dir``, one clip per scene per station. + + ``source`` is either one object's directory or a whole corpus, in which case + every object in it is exported (or the subset in ``options.objects``). + """ + options = options or CaptureOptions() + source = Path(source).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + + if (source / "dataset.json").is_file(): + wanted_objects = set(options.objects) if options.objects else None + scenes = [ + source / entry + for entry in scenes_in(source) + if wanted_objects is None or entry in wanted_objects + ] + if not scenes: + raise ValueError( + f"{source} has none of {', '.join(sorted(options.objects))}; it has " + + ", ".join(scenes_in(source)) + ) + clips: list[bundle.Clip] = [] + for scene in scenes: + _, produced, _ = _build_scene(scene, out_dir, options) + clips += produced + title = "Captured — " + ", ".join(scene.name for scene in scenes) + return title, clips, {"baseline": name, "corpus": str(source)} + + return _build_scene(source, out_dir, options) + + +def _build_scene( + scene_dir: Path, out_dir: Path, options: CaptureOptions +) -> tuple[str, list[bundle.Clip], dict[str, Any]]: + """One object's captured views, one clip per rig station.""" + from PIL import Image + + scene_dir = Path(scene_dir) + out_dir = Path(out_dir) + rig = cameras.read_orbit_rig(scene_dir) + frames = frame_dirs(scene_dir) + if not frames: + raise FileNotFoundError(f"{scene_dir} holds no frame_NNNNNN directories") + if options.frames is not None: + frames = frames[: options.frames] + + wanted = set(options.views) if options.views else {pose.view_id for pose in rig.poses} + clips: list[bundle.Clip] = [] + for pose in rig.poses: + if pose.view_id not in wanted: + continue + frames_at = bundle.frame_dir(out_dir, f"{rig.scene}-captured-view{pose.view_id:02d}") + written: list[str] = [] + for index, frame in enumerate(frames): + source = frame / "images" / f"view_{pose.view_id:02d}.png" + if not source.is_file(): + raise FileNotFoundError(f"missing captured view: {source}") + image = Image.open(source).convert("RGB") + if image.width > options.max_width: + height = round(image.height * options.max_width / image.width) + image = image.resize((options.max_width, height), Image.LANCZOS) + destination = frames_at / f"frame_{index:04d}.jpg" + image.save(destination, quality=options.quality) + written.append(str(destination.relative_to(out_dir))) + clips.append( + bundle.Clip( + name=frames_at.name, + representation="pixels", + scene=rig.scene, + method=name, + camera=pose.view_id, + frames=written, + notes=[ + f"captured by rig camera {pose.view_id}, downscaled to " + f"{options.max_width}px — this is the reference, not a reconstruction", + "a training view for both Vega and ReRF: this measures " + "reconstruction, not generalisation", + ], + detail={"source": str(scene_dir), "view_id": pose.view_id}, + ) + ) + print(f" {frames_at.name}: {len(written)} frames", flush=True) + + return f"Captured — {rig.scene}", clips, {"baseline": name, "corpus": str(scene_dir.parent)} + + +def rigs_for(corpus: Path | str, scenes: list[str]) -> dict[str, Any]: + """Each named scene's rig, for the bundle's `scenes` map. + + Looked up here rather than by each exporter: Vega and ReRF know which subject + they reconstruct but not which cameras captured it, and the whole point of a + shared camera is that the answer does not depend on who is asking. + """ + corpus = Path(corpus).expanduser().resolve() + found: dict[str, Any] = {} + for scene in scenes: + directory = corpus / scene + if not directory.is_dir(): + continue + try: + found[scene] = cameras.read_orbit_rig(directory).as_dict() + except (FileNotFoundError, KeyError, ValueError): + continue + return found + + +def export(source: Path | str, out_dir: Path | str, options: CaptureOptions | None = None) -> Path: + """Export captured views into a viewable bundle at ``out_dir``.""" + options = options or CaptureOptions() + out_dir = Path(out_dir).expanduser().resolve() + source = Path(source).expanduser().resolve() + title, clips, detail = build_clips(source, out_dir, options) + corpus = source if (source / "dataset.json").is_file() else source.parent + scenes = sorted({clip.scene for clip in clips if clip.scene}) + bundle.write( + out_dir, title=title, source=source, clips=clips, fps=options.fps, + scenes=rigs_for(corpus, scenes), detail=detail, + ) + return out_dir diff --git a/open4d/reconstruction/gs_tools/gs_tools/methods/gaussian.py b/open4d/reconstruction/gs_tools/gs_tools/methods/gaussian.py new file mode 100644 index 00000000..8f361663 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/methods/gaussian.py @@ -0,0 +1,209 @@ +"""3DGS runs, made viewable -- QUEEN, 3DGStream, or any plain 3DGS output. + +The easiest of the exporters here, and the last to be written, which is the +wrong way round: these runs already store Gaussians in the one format every +splat viewer reads, so making them viewable is a copy plus a manifest. Vega +needed its colour decoded through a hash grid first and ReRF cannot be decoded +in this process at all; this needs neither. Until it existed, the bundled +client's free camera showed Vega and nothing else -- not because the other +methods lacked geometry, but because nothing exported theirs. + +Deliberately not a trainer wrapper. `queen` and `gstream` in this package run +the upstream trainers; this reads whatever they left on disk, so it works on a +run produced by any 3DGS implementation, including ones this repository has +never heard of. `gs_tools.outputs.gaussian_frames` is what absorbs the three +unrelated on-disk layouts. + +Two things it does not attempt: + +* **The scene name is guessed from the directory.** A 3DGS run records no + subject name anywhere reliable, so a run called ``verify`` becomes a scene + called ``verify``. Pass ``scene`` to put it beside another method's clip of + the same subject -- which is the only way `Compare` can line them up. +* **No rig.** These runs carry their own ``cameras.json`` in formats that differ + per implementation, and the bundle's shared camera comes from the ORBIT corpus + via `gs_tools.cameras`. A 3DGS run of an ORBIT object gets a rig like any + other clip; one of some other capture is explore-only. +""" + +from __future__ import annotations + +import shutil +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +import numpy as np +from streamer import bundle + +from .. import io +from ..io import ply, splat +from ..outputs import Kind, detect, gaussian_frames + +name = "gaussian" +#: No upstream tree of its own: this reads output, and several trainers write it. +upstream = None + + +@dataclass +class GaussianExportOptions: + """What the export needs beyond the input and output paths.""" + + #: How many frames; None means every one the run holds. + frames: int | None = None + #: ``"ply"`` copies the run's own files. ``"splat"`` re-encodes to 32 bytes + #: per Gaussian, which is several times smaller and drops every + #: spherical-harmonic band above degree 0 -- see `gs_tools.io.splat`. + frame_format: str = "ply" + #: Subject name, shared with other methods' clips of the same thing. + scene: str | None = None + #: Method label in the viewer; defaults to the run's manifest or "gaussian". + method: str | None = None + fps: int = 30 + extra: dict[str, Any] = field(default_factory=dict) + + +#: Re-exported so `--frame-format` can list the choices without importing +#: `gs_tools.io` directly; defined in one place (see `gs_tools.io`). +FORMATS = io.GAUSSIAN_FORMATS + + +def _method_name(source: Path, found, options: GaussianExportOptions) -> str: + """What to label this in the viewer. + + A run's own manifest is the only honest source, since the directory name + says whatever its author felt like. Falling back to the representation name + rather than to the directory keeps two unrelated runs from both claiming to + be the method called ``output``. + """ + if options.method: + return options.method + recorded = (found.detail or {}).get("method") + return recorded or "gaussian" + + +def build_clips( + source: Path | str, + out_dir: Path | str, + options: GaussianExportOptions | None = None, +) -> tuple[str, list[bundle.Clip], dict[str, Any]]: + """Copy (or re-encode) a run's per-frame Gaussians into ``out_dir``.""" + options = options or GaussianExportOptions() + if options.frame_format not in FORMATS: + raise ValueError( + f"unknown frame format {options.frame_format!r}; expected one of " + + ", ".join(FORMATS) + ) + source = Path(source).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + + found = detect(source) + if found.kind is not Kind.GAUSSIAN_RUN: + raise ValueError( + f"{source} is {found.kind.value}, not a 3DGS run; expected one of the " + "layouts gs_tools.outputs.gaussian_frames resolves" + ) + + entries = gaussian_frames(source) + if options.frames is not None: + entries = entries[: options.frames] + if not entries: + raise ValueError(f"{source} has no frames to export") + + scene = options.scene or source.name + method = _method_name(source, found, options) + frames_at = bundle.frame_dir(out_dir, f"{source.name}-{method}") + clip_name = frames_at.name + + written: list[str] = [] + counts: list[int] = [] + lower = np.full(3, np.inf) + upper = np.full(3, -np.inf) + degrees: set[int] = set() + + for index, ply_path in entries: + if options.frame_format == "splat": + cloud = splat.from_ply(ply_path) + target = splat.write(frames_at / f"frame_{index:04d}.splat", cloud) + positions = cloud.positions + degrees.add(0) + else: + # Copied rather than rewritten: the run's PLY is already the + # interchange format, and re-encoding it would only add a chance to + # get an activation or an SH ordering wrong. + target = frames_at / f"frame_{index:04d}.ply" + shutil.copyfile(ply_path, target) + fields = ply.read(target) + positions = fields["xyz"] + degrees.add(int(fields["sh_degree"])) + + written.append(str(target.relative_to(out_dir))) + counts.append(len(positions)) + lower = np.minimum(lower, positions.min(axis=0)) + upper = np.maximum(upper, positions.max(axis=0)) + print( + f" {clip_name} frame {index:04d} {len(positions):7d} gaussians " + f"{target.stat().st_size / 1e6:5.2f} MB", + flush=True, + ) + + notes = [ + f"3DGS run, {len(written)} frames, read from {source.name} as-is — " + "not retrained or resampled", + ] + if options.frame_format == "splat": + notes.append( + "frames re-encoded to .splat (32 bytes per Gaussian): every " + "spherical-harmonic band above degree 0 is dropped, so appearance " + "no longer changes with view direction" + ) + else: + notes.append( + "sh_degree " + ", ".join(str(d) for d in sorted(degrees)) + + ": view-dependent bands preserved as the run wrote them" + ) + if options.scene is None: + notes.append( + f"scene name taken from the directory ({scene}); pass --scene to line " + "this up with another method's clip of the same subject" + ) + + clip = bundle.Clip( + name=clip_name, + representation="gaussians", + scene=scene, + method=method, + frames=written, + counts=counts, + bounds_min=lower.tolist(), + bounds_max=upper.tolist(), + notes=notes, + detail={ + "source": str(source), + "frame_format": options.frame_format, + "frame_indices": [index for index, _ in entries], + "sh_degrees": sorted(degrees), + "iterations": (found.detail or {}).get("iterations", []), + }, + ) + return f"3DGS — {source.name}", [clip], {"baseline": method} + + +def export( + source: Path | str, + out_dir: Path | str, + options: GaussianExportOptions | None = None, +) -> Path: + """Build a one-clip bundle from a 3DGS run and return its directory.""" + options = options or GaussianExportOptions() + out_dir = Path(out_dir).expanduser().resolve() + title, clips, detail = build_clips(source, out_dir, options) + bundle.write( + out_dir, + title=title, + source=str(source), + clips=clips, + fps=options.fps, + detail=detail, + ) + return out_dir diff --git a/open4d/reconstruction/gs_tools/gs_tools/methods/rerf.py b/open4d/reconstruction/gs_tools/gs_tools/methods/rerf.py new file mode 100644 index 00000000..df9b51fd --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/methods/rerf.py @@ -0,0 +1,359 @@ +"""ReRF, made viewable -- by bundling renders, not by producing them. + +ReRF stores a neural volume, so unlike every other method in this module its +output is not Gaussians and cannot be turned into a PLY. What a compressed +sequence holds is a DCT-coded, arithmetic-coded feature voxel grid +(`feature__.rerf*`), an occupancy mask, per-frame motion +vectors, a PCA basis per P-frame, and the shared colour MLP. The only decoder +for that is ReRF's, and it only runs under Python 3.8 -- its entropy coder +``ac_dc/`` ships as a CPython 3.8 binary with no sources. + +**Rendering moved out of this module.** It used to shell out to upstream's +``rerf_render.py`` through a wrapper in the vendored tree. That is now +`rerf_stream.export`, in ``open4d/reconstruction/rerf``, which does the job +better in three ways that matter: + +* it renders at the **corpus's own intrinsics**, where upstream's loader + letterboxes 4:3 footage into 16:9 -- the pose was right and the framing was + not, so a render could not be compared pixel-for-pixel against the + photograph beside it; +* it can write **quality rungs** as a clip's variants, which is what lets + anything downstream choose a rate; and +* it emits the **capture rig** and marks what each clip depicts, so depth maps + are not scored against colour photographs. + +So this module keeps what it is still the right home for -- reading a +bitstream's codec configuration, and turning an image sequence that already +exists into bundle clips -- and points at `rerf_stream` for the rest. +:func:`build_clips` raises with that instruction rather than rendering. + +What is inferred rather than asked for is the codec configuration, because +getting it wrong produces a silently wrong decode: upstream's README requires +``--pca``/``--pca_chs``/``--group_size`` to match between compress and render, +and nothing in the bitstream forces the issue. :func:`bitstream_info` reads it +back off the headers instead -- ``codec/compress.py`` writes one header per +frame whose entry count and channel split *are* the PCA configuration, and +whose single-entry frames are exactly the key frames. + +Runs anywhere: nothing here needs a GPU or Python 3.8 any more. +""" + +from __future__ import annotations + +import json +import os +import re +import shutil +import subprocess +import sys +import time +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +from .. import paths +from streamer import bundle +from ..outputs import Kind, detect + +name = "rerf" +#: The tree ReRF is vendored in, for `gs_tools.env`'s provenance record. It is +#: no longer executed from here -- see the module docstring -- but a manifest +#: should still say which commit of upstream the frames came from. +upstream = "rerf" + +#: Where rendering lives now, named in the error `build_clips` raises. +RENDERER = "rerf_stream.export" + + +@dataclass +class RerfRenderOptions: + """What bundling a ReRF output needs beyond the paths. + + Named for what it used to do. Kept as the name because `gs_tools.cli` + builds one for every method and the shape is part of that contract; the + render-only fields are gone, since nothing here renders. + """ + + #: Frames to bundle. None takes every frame present. + frames: int | None = None + #: Bundle ReRF's depth maps alongside the colour frames. + depth: bool = True + #: Which bitstream in a run to read the codec configuration from, by + #: directory name. None reads whichever one is present. + bitstream: str | None = None + #: Override the inferred codec configuration. None means infer. + pca: bool | None = None + pca_chs: tuple[int, ...] | None = None + group_size: int | None = None + #: ReRF config; defaults to the `config.py` in the run directory. + config: Path | None = None + fps: int = 30 + dry_run: bool = False + extra: dict[str, Any] = field(default_factory=dict) + + +def bitstream_info(rerf_dir: Path | str) -> dict[str, Any]: + """Read the codec configuration back off a ReRF bitstream's headers. + + `codec/compress.py` writes, per frame, either one header covering every + feature channel (a key frame, or PCA off) or two headers whose channel + counts are the PCA split -- the first at the requested quality and the + second one step below it. Both facts are recoverable, so neither has to be + remembered from whatever command produced the directory. + """ + rerf_dir = Path(rerf_dir).expanduser().resolve() + kwargs_path = rerf_dir / "model_kwargs.json" + if not kwargs_path.is_file(): + raise FileNotFoundError(f"{rerf_dir} has no model_kwargs.json; not a ReRF bitstream") + model_kwargs = json.loads(kwargs_path.read_text()) + + header_paths = sorted( + (p for p in rerf_dir.glob("header_*.json") if re.fullmatch(r"header_\d+\.json", p.name)), + key=lambda p: int(p.stem.split("_")[1]), + ) + if not header_paths: + raise FileNotFoundError(f"{rerf_dir} has no header_*.json; nothing was compressed") + + key_frames: list[int] = [] + pca_chs: tuple[int, ...] | None = None + feature_dim = 0 + qualities: set[int] = set() + for path in header_paths: + index = int(path.stem.split("_")[1]) + entries = json.loads(path.read_text())["headers"] + channels = [int(entry["origin_size"][0]) for entry in entries] + qualities.update(int(entry["quality"]) for entry in entries) + feature_dim = max(feature_dim, sum(channels)) + if len(entries) == 1: + key_frames.append(index) + elif pca_chs is None: + cumulative: list[int] = [] + total = 0 + for count in channels: + total += count + cumulative.append(total) + pca_chs = tuple(cumulative) + + frames = len(header_paths) + # Key frames are `frame_id % group_size == 0`, so consecutive key frames are + # one group apart; a single key frame means the group spans the sequence. + group_size = key_frames[1] - key_frames[0] if len(key_frames) > 1 else frames + grid = json.loads(header_paths[0].read_text())["headers"][0].get("origin_size", [])[1:] + + return { + "root": rerf_dir, + "frames": frames, + "key_frames": key_frames, + "group_size": group_size, + "pca": pca_chs is not None, + "pca_chs": pca_chs or (), + "feature_dim": feature_dim, + "quality": sorted(qualities, reverse=True), + "grid": [int(n) for n in grid], + "has_rgb_net": (rerf_dir / "rgb_net.tar").is_file(), + "xyz_min": model_kwargs.get("xyz_min"), + "xyz_max": model_kwargs.get("xyz_max"), + } + + +def _config_datadir(config: Path) -> str | None: + """``data.datadir`` out of a ReRF config, without importing mmcv.""" + match = re.search(r"^\s*datadir\s*=\s*['\"]([^'\"]*)['\"]", config.read_text(), re.M) + return match.group(1) if match else None + + +def scene_name(run_name: str) -> str: + """The subject a ReRF run reconstructs, from its run directory name. + + Runs are named `g_` after the corpus they were prepared from, and + the `g_` has to come off for the name to match what Vega and the captured + views call the same subject -- which is what lets the viewer line them up. + """ + return run_name[2:] if run_name.startswith("g_") else run_name + + +def collect( + image_dir: Path, + out_dir: Path, + clip_name: str, + options: RerfRenderOptions, + *, + scene: str | None = None, + method: str | None = None, + camera: int | None = None, + notes: list[str] | None = None, + detail: dict[str, Any] | None = None, +) -> list[bundle.Clip]: + """Copy a rendered image sequence into a bundle, colour and depth separately. + + Copied rather than symlinked so the bundle survives being moved or archived; + ReRF's 360 renders are tens of kilobytes a frame, so the duplication is not + worth avoiding. + """ + image_dir = Path(image_dir).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + colour: list[Path] = [] + depth: list[Path] = [] + for path in sorted(image_dir.iterdir()): + if path.suffix.lower() not in (".jpg", ".jpeg", ".png"): + continue + if re.fullmatch(r"\d+", path.stem): + colour.append(path) + elif re.fullmatch(r"\d+_depth", path.stem): + depth.append(path) + colour.sort(key=lambda p: int(p.stem)) + depth.sort(key=lambda p: int(p.stem.split("_")[0])) + if not colour: + raise FileNotFoundError(f"{image_dir} holds no numbered images") + + if options.frames is not None: + colour = colour[: options.frames] + depth = depth[: options.frames] + + clips: list[bundle.Clip] = [] + for suffix, sources in (("", colour), ("-depth", depth)): + if not sources or (suffix and not options.depth): + continue + frames_at = bundle.frame_dir(out_dir, f"{clip_name}{suffix}") + name = frames_at.name + frames: list[str] = [] + for index, source in enumerate(sources): + destination = frames_at / f"frame_{index:04d}{source.suffix.lower()}" + shutil.copyfile(source, destination) + frames.append(str(destination.relative_to(out_dir))) + clips.append( + bundle.Clip( + name=name, + representation="pixels", + scene=scene, + method=f"{method}-depth" if (suffix and method) else method, + camera=camera, + frames=frames, + notes=list(notes or []) + + (["ReRF's depth output, not colour"] if suffix else []), + detail={"source": str(image_dir), **(detail or {})}, + ) + ) + print(f" {name}: {len(frames)} frames from {image_dir.name}", flush=True) + return clips + + +def build_clips( + source: Path | str, out_dir: Path | str, options: RerfRenderOptions | None = None +) -> tuple[str, list[bundle.Clip], dict[str, Any]]: + """Render (or reuse) one ReRF output into ``out_dir`` and describe the clips. + + Separate from :func:`export` so several runs can be combined into one + bundle; see `gs_tools.methods.vega.build_clips`. + + ``source`` may be a run root (every bitstream in it), one bitstream + directory, or a directory of already-rendered images. + """ + options = options or RerfRenderOptions() + source = Path(source).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + found = detect(source) + clips: list[bundle.Clip] = [] + detail: dict[str, Any] = {"baseline": "rerf"} + + if found.kind is Kind.IMAGE_SEQUENCE: + clips = collect( + source, + out_dir, + source.name, + options, + scene=scene_name(source.parent.name), + method="rerf", + notes=["pre-rendered by ReRF; the camera is the one that render swept"], + ) + title = f"ReRF — {source.name}" + elif found.kind in (Kind.RERF_BITSTREAM, Kind.RERF_RUN): + if found.kind is Kind.RERF_BITSTREAM: + run_root, bitstreams, renders = source.parent, [source], [] + else: + run_root = source + available = list(found.detail["bitstreams"]) + renders = list(found.detail["renders"]) + if options.bitstream: + if options.bitstream not in available: + raise ValueError( + f"{source} has no bitstream {options.bitstream!r}; it has " + + (", ".join(available) or "none") + ) + bitstreams = [source / options.bitstream] + renders = [] + elif renders: + # Existing renders are named by whoever produced them, so they + # say which condition they are; a bitstream name plus a render + # command does not. Prefer the unambiguous evidence. + bitstreams = [] + elif len(available) == 1: + bitstreams = [source / available[0]] + else: + raise ValueError( + f"{source} holds {len(available)} bitstreams " + f"({', '.join(available)}) and no render; pass --bitstream to say " + "which one's codec configuration to record." + ) + # Order matters: a bitstream with no render is the case that moved out, + # and it deserves the instruction rather than "holds nothing". + if bitstreams: + # `bitstreams` is non-empty only when there is no render to prefer, + # so this is exactly the case that moved to `rerf_stream`. + # Rendering left this module; see the module docstring. Naming the + # command matters more than usual here, because the alternative is + # a person concluding ReRF cannot be bundled at all. + names = ", ".join(path.name for path in bitstreams) + raise RuntimeError( + f"{run_root} holds a ReRF bitstream ({names}) and no render. " + f"Rendering is {RENDERER}, which runs in the Python 3.8 " + "environment ReRF's entropy coder needs:\n" + f" python -m {RENDERER} --config {run_root / 'config.py'} " + f"--compression-path {bitstreams[0]} \\\n" + f" --out ~/rerf-clips --name {run_root.name} " + f"--scene {scene_name(run_root.name)} --depth --captured\n" + " python -m streamer.adopt ~/rerf-clips --bundle \n" + "It renders at the corpus's own intrinsics, can write quality " + "rungs, and emits the capture rig -- none of which the renderer " + "this used to call did." + ) + + if not renders: + raise ValueError(f"{source} holds no ReRF bitstream or render") + + for render_name in renders: + clips += collect( + run_root / render_name, + out_dir, + f"{run_root.name}-{render_name}", + options, + scene=scene_name(run_root.name), + method="rerf-whitebg" if render_name.endswith("_whitebg") else "rerf", + notes=[ + "existing ReRF render, reused as-is — not a free camera: ReRF stores " + "a feature voxel grid, not Gaussians", + "which bitstream produced it is not recorded in the directory", + ], + detail={"render": str(run_root / render_name)}, + ) + title = f"ReRF — {run_root.name}" + else: + raise ValueError( + f"{source} is {found.kind.value}, not a ReRF output; expected a run " + "root, a bitstream directory, or a rendered image sequence" + ) + + return title, clips, detail + + +def export(source: Path | str, out_dir: Path | str, options: RerfRenderOptions | None = None) -> Path: + """Export a ReRF output at ``source`` into a viewable bundle at ``out_dir``.""" + options = options or RerfRenderOptions() + out_dir = Path(out_dir).expanduser().resolve() + title, clips, detail = build_clips(source, out_dir, options) + if options.dry_run and not clips: + # Nothing was rendered, so there is nothing to index; writing an empty + # bundle would leave a directory that `view` would then fail on. + return out_dir + bundle.write(out_dir, title=title, source=source, clips=clips, fps=options.fps, detail=detail) + return out_dir diff --git a/open4d/reconstruction/gs_tools/gs_tools/methods/vega.py b/open4d/reconstruction/gs_tools/gs_tools/methods/vega.py new file mode 100644 index 00000000..f5e2e066 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/methods/vega.py @@ -0,0 +1,417 @@ +"""Vega, made viewable. + +Unlike QUEEN and 3DGStream this is not a trainer wrapper. Vega's own +`orbitvega.prepare` produces the run; what was missing was any way to *look* at +what it produced. Its bitstream is per-object ``frame_XXXX.pt`` chunks holding +geometry only -- position, scale, rotation, opacity -- with colour living in a +hierarchical hash grid (`vega.color_encoding`) that is queried per Gaussian per +view direction at render time. So there is no file in a Vega run that any +Gaussian-splatting viewer can open, and the fix is not a format shim: colour +has to be *decoded* first, which means running Vega's own model. + +This adapter does exactly that and nothing more. It drives +`vega.player.StreamingPlayer` -- the same client-side reassembly Vega's live +demo uses, so key/residual handling is upstream's, not a reimplementation -- +decodes colour through the object's own `HierarchicalColorModel`, and writes +one 3DGS PLY per frame. + +Two consequences of that route are worth being explicit about, because both are +visible in the output: + +* **Colour is baked, and the export is degree-0.** The hash grid is + view-dependent; a PLY's ``f_dc`` is not. Colour is therefore evaluated once, + from one camera azimuth, and frozen -- a free-camera viewer will show the + subject's appearance from that direction no matter where it is orbited to. + Re-export at a different ``--bake-azimuth`` to see the appearance from + somewhere else. Writing the full spherical-harmonic bands instead is not an + option: the hash grid is not an SH expansion, and fitting one per Gaussian + would be a different piece of work with its own error. +* **The grid is queried at each Gaussian's original position.** + `HierarchicalColorModel._normalize` maps position into [0, 1] against the + bbox the model was trained on and *clamps*, so feeding it anything outside + those bounds collapses colour onto the box faces. That is why the bake camera + is derived from the model's own bbox rather than from a scene layout. + +Reading the bitstream needs no server: `vega.player.BitstreamClient` fetches +over ``urllib``, and a ``file://`` base URL is a valid thing to hand it, so the +chunk server `orbitvega.scene_export` starts is skipped here. + +Runs in the ``open4d-gs`` environment: decoding colour needs torch and +tinycudann, the same as Vega itself. +""" + +from __future__ import annotations + +import json +import math +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +import numpy as np + +from .. import upstream_import +from streamer import bundle + +from .. import io +from ..io import ply +from ..outputs import Kind, detect + +name = "vega" +upstream = "vega" + + +@dataclass +class VegaExportOptions: + """What the export needs beyond the input and output paths.""" + + #: Object names to export; empty means every object in the catalog. + objects: tuple[str, ...] = () + #: How many frames per object; None means the whole encoded sequence. + frames: int | None = None + #: Camera azimuth, in degrees about the world up axis, that colour is + #: evaluated from. 0 looks along +Z towards the subject. + bake_azimuth_deg: float = 0.0 + #: Bake-camera height above the bbox centre, as a fraction of its radius. + #: Slightly above eye level, so the top of a head is not lit as if from below. + bake_elevation: float = 0.15 + #: ``"ply"`` writes 3DGS PLY. ``"splat"`` re-encodes to 32 bytes a + #: Gaussian, which for this exporter loses nothing structural: Vega's + #: colour is baked to a single band here anyway, so there are no + #: spherical-harmonic coefficients above degree 0 to drop. Roughly half the + #: bytes a frame, and half of what a client has to move. + frame_format: str = "ply" + #: "cuda", "cpu", or None to take CUDA when it is there. + device: str | None = None + fps: int = 30 + extra: dict[str, Any] = field(default_factory=dict) + + +FORMATS = io.GAUSSIAN_FORMATS + + +def _write_frame(path_without_suffix: Path, options: VegaExportOptions, **fields): + """One frame, in the format asked for. + + Shared by both export paths so they cannot disagree about it -- the + bitstream path and the pre-baked scene path write the same fields and + previously each called `ply.write` directly. + + `.splat` goes through the PLY rather than around it. Encoding straight from + the arrays would mean a second implementation of the same quantisation, and + the round trip is checked: `gs_tools.io.splat` keeps position and scale + exactly. The PLY is removed afterwards, since keeping both doubles the + export for a file no client asks for. + """ + if options.frame_format not in FORMATS: + raise ValueError( + f"frame_format {options.frame_format!r} is not one of " + + ", ".join(FORMATS) + ) + ply_path = ply.write(path_without_suffix.with_suffix(".ply"), **fields) + if options.frame_format != "splat": + return ply_path + cloud = io.splat.from_ply(ply_path) + written = io.splat.write(path_without_suffix.with_suffix(".splat"), cloud) + ply_path.unlink() + return written + + +def _format_notes(options: VegaExportOptions) -> list[str]: + """What the chosen format costs, said in the clip rather than assumed.""" + if options.frame_format != "splat": + return [] + return [ + "delivered as .splat (32 bytes a Gaussian) rather than 3DGS PLY: about " + "half the bytes a frame, and nothing structural is dropped because this " + "export bakes colour to degree 0 anyway — opacity, rotation and colour " + "are quantised to 8 bits, position and scale are exact", + ] + + +def _torch(): + import torch + + return torch + + +def _vega_modules(): + """Vega's player, camera helpers and view-direction helper. + + Imported from the vendored tree rather than a copy: the key/residual + reassembly in `StreamingPlayer.reconstruct` is the part that is easy to get + subtly wrong, and it is already written and tested next door. + """ + with upstream_import.on_path("vega"): + from vega.cameras import Camera, look_at_RT + from vega.player import BitstreamClient, StreamingPlayer + from vega.rasterize import view_directions + + return Camera, look_at_RT, BitstreamClient, StreamingPlayer, view_directions + + +def objects_in(source: Path) -> list[dict[str, Any]]: + """The exportable objects at ``source``, as ``{"name", "dir"}`` entries. + + A catalog lists them; a bare bitstream directory is itself one object. + """ + source = Path(source) + catalog_path = source / "catalog.json" + if catalog_path.is_file(): + catalog = json.loads(catalog_path.read_text()) + return [ + {"name": entry["name"], "dir": entry.get("dir", entry["name"]), "catalog": entry} + for entry in catalog.get("objects", []) + ] + return [{"name": source.name, "dir": ".", "catalog": {}}] + + +def _bake_camera(bbox_min, bbox_max, options: VegaExportOptions, device: str): + """A camera outside the object's bbox, used only for its centre position. + + Only the camera centre matters -- `view_directions` is the sole consumer -- + so the image size and field of view here are placeholders. Built through + Vega's own `look_at_RT` so the convention (+Z forward, y down) matches what + the colour model was trained against. + """ + Camera, look_at_RT, *_ = _vega_modules() + centre = (np.asarray(bbox_min) + np.asarray(bbox_max)) / 2.0 + radius = float(np.linalg.norm(np.asarray(bbox_max) - np.asarray(bbox_min))) or 1.0 + azimuth = math.radians(options.bake_azimuth_deg) + eye = centre + radius * np.array( + [math.sin(azimuth), options.bake_elevation, math.cos(azimuth)], dtype=np.float64 + ) + rotation, translation = look_at_RT( + eye.astype(np.float32), centre.astype(np.float32), np.array([0.0, 1.0, 0.0]) + ) + return Camera( + R=rotation, + T=translation, + fovx=1.0, + fovy=1.0, + width=1, + height=1, + device=device, + ) + + +def _export_bitstream( + object_dir: Path, + out_dir: Path, + clip_name: str, + options: VegaExportOptions, +) -> bundle.Clip: + """Decode one object's bitstream to a PLY per frame.""" + torch = _torch() + _, _, BitstreamClient, StreamingPlayer, view_directions = _vega_modules() + device = options.device or ("cuda" if torch.cuda.is_available() else "cpu") + + client = BitstreamClient(f"file://{object_dir}") + manifest = client.get_manifest() + color_model = client.get_color_model(device) + player = StreamingPlayer(color_model, device=device) + + bbox_min = color_model.bbox_min.detach().cpu().numpy() + bbox_max = color_model.bbox_max.detach().cpu().numpy() + camera = _bake_camera(bbox_min, bbox_max, options, device) + + entries = manifest["frames"] + if options.frames is not None: + entries = entries[: options.frames] + if not entries: + raise ValueError(f"{object_dir} has no frames to export") + + frames_at = bundle.frame_dir(out_dir, clip_name) + clip_name = frames_at.name + frames: list[str] = [] + counts: list[int] = [] + lower = np.full(3, np.inf) + upper = np.full(3, -np.inf) + + for entry in entries: + index = entry["frame_idx"] + chunk = client.get_frame_chunk(index) + gaussians = player.reconstruct(chunk) + with torch.no_grad(): + dirs = view_directions(camera, gaussians.get_xyz) + if chunk["frame_type"] == "key": + rgb = color_model.forward_key(gaussians.get_xyz, dirs) + else: + rgb = color_model.forward_residual(gaussians.get_xyz, dirs, index) + rgb = rgb.clamp(0.0, 1.0) + + xyz = gaussians.xyz.detach().cpu().numpy() + path = _write_frame( + frames_at / f"frame_{index:04d}", options, + xyz=xyz, + scale_raw=gaussians.scale_raw.detach().cpu().numpy(), + rot_raw=gaussians.rot_raw.detach().cpu().numpy(), + opacity_raw=gaussians.opacity_raw.detach().cpu().numpy(), + sh_dc=ply.rgb_to_sh_dc(rgb.cpu().numpy()), + ) + # Residual frames only carry the objects that changed, so the tiny hash + # for a frame already written is dead weight from here on. + color_model.drop_tiny_hash(index) + frames.append(str(path.relative_to(out_dir))) + counts.append(len(gaussians)) + lower = np.minimum(lower, xyz.min(axis=0)) + upper = np.maximum(upper, xyz.max(axis=0)) + print( + f" {clip_name} frame {index:04d} {chunk['frame_type']:8s} " + f"{len(gaussians):7d} gaussians {path.stat().st_size / 1e6:5.2f} MB", + flush=True, + ) + + return bundle.Clip( + name=clip_name, + representation="gaussians", + scene=clip_name, + method=name, + frames=frames, + counts=counts, + bounds_min=lower.tolist(), + bounds_max=upper.tolist(), + notes=[ + f"colour baked at azimuth {options.bake_azimuth_deg:g}° and frozen " + "(Vega's colour is a view-dependent hash grid; a PLY's f_dc is not)", + "sh_degree 0: no view-dependent bands", + *_format_notes(options), + ], + detail={ + "source": str(object_dir), + "bbox_min": bbox_min.tolist(), + "bbox_max": bbox_max.tolist(), + "frame_types": [entry["frame_type"] for entry in entries], + }, + ) + + +def _export_scene(scene_dir: Path, out_dir: Path, options: VegaExportOptions) -> bundle.Clip: + """Convert an `orbitvega.scene_export` directory, whose colour is already baked. + + Cheaper and lower-fidelity than :func:`_export_bitstream` in exactly one + way: the bake camera was chosen when that export ran, not here. It needs no + hash grid and no CUDA, which is the point -- the merged multi-object scene + is Vega's own tool's output, so this reads it rather than reproducing the + layout logic. + """ + torch = _torch() + scene = json.loads((scene_dir / "scene_manifest.json").read_text()) + entries = scene.get("frames") or [ + {"frame_idx": index, "file": path.name} + for index, path in enumerate(sorted(scene_dir.glob("frame_*.pt"))) + ] + if options.frames is not None: + entries = entries[: options.frames] + + frames_at = bundle.frame_dir(out_dir, scene_dir.name) + clip_name = frames_at.name + frames: list[str] = [] + counts: list[int] = [] + lower = np.full(3, np.inf) + upper = np.full(3, -np.inf) + + for entry in entries: + index = entry.get("frame_idx", len(frames)) + payload = torch.load(scene_dir / entry["file"], weights_only=False, map_location="cpu") + xyz = payload["xyz"].float().numpy() + path = _write_frame( + frames_at / f"frame_{index:04d}", options, + xyz=xyz, + scale_raw=payload["scale_raw"].float().numpy(), + rot_raw=payload["rot_raw"].float().numpy(), + opacity_raw=payload["opacity_raw"].float().numpy(), + sh_dc=ply.rgb_to_sh_dc(payload["rgb"].float().clamp(0, 1).numpy()), + ) + frames.append(str(path.relative_to(out_dir))) + counts.append(int(xyz.shape[0])) + lower = np.minimum(lower, xyz.min(axis=0)) + upper = np.maximum(upper, xyz.max(axis=0)) + print( + f" {clip_name} frame {index:04d} {xyz.shape[0]:7d} gaussians " + f"{path.stat().st_size / 1e6:5.2f} MB", + flush=True, + ) + + colour = scene.get("colour", {}) + return bundle.Clip( + name=clip_name, + representation="gaussians", + scene=clip_name, + method=name, + frames=frames, + counts=counts, + bounds_min=lower.tolist(), + bounds_max=upper.tolist(), + notes=[ + "colour baked by orbitvega.scene_export at azimuth " + f"{colour.get('bake_azimuth_deg', '?')}° and frozen", + f"layout={scene.get('layout')}: " + + ", ".join(entry["name"] for entry in scene.get("objects", [])), + *_format_notes(options), + ], + detail={"source": str(scene_dir), "scene_manifest": scene.get("layout")}, + ) + + +def build_clips( + source: Path | str, out_dir: Path | str, options: VegaExportOptions | None = None +) -> tuple[str, list[bundle.Clip], dict[str, Any]]: + """Write one Vega run's frames into ``out_dir`` and describe them. + + Separate from :func:`export` so several sources can be combined into one + bundle -- a Vega catalog next to a ReRF run is exactly the comparison this + is for, and it needs one index over both. + """ + options = options or VegaExportOptions() + source = Path(source).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + found = detect(source) + + if found.kind is Kind.VEGA_SCENE_EXPORT: + clips = [_export_scene(source, out_dir, options)] + title = f"Vega — {source.name} (merged scene)" + elif found.kind in (Kind.VEGA_CATALOG, Kind.VEGA_BITSTREAM): + entries = objects_in(source) + if options.objects: + wanted = set(options.objects) + missing = wanted - {entry["name"] for entry in entries} + if missing: + raise ValueError( + f"{source} has no object(s) {', '.join(sorted(missing))}; it has " + + ", ".join(entry["name"] for entry in entries) + ) + entries = [entry for entry in entries if entry["name"] in wanted] + clips = [ + _export_bitstream( + (source / entry["dir"]).resolve(), out_dir, entry["name"], options + ) + for entry in entries + ] + title = "Vega — " + ", ".join(clip.name for clip in clips) + else: + raise ValueError( + f"{source} is {found.kind.value}, not a Vega output; expected a " + "catalog directory, one object's bitstream, or a scene_export directory" + ) + + detail = { + "baseline": "vega", + "bake_azimuth_deg": options.bake_azimuth_deg, + "bake_elevation": options.bake_elevation, + } + return title, clips, detail + + +def export(source: Path | str, out_dir: Path | str, options: VegaExportOptions | None = None) -> Path: + """Export a Vega run at ``source`` into a viewable bundle at ``out_dir``. + + ``source`` may be a catalog directory (every object, or the subset named in + ``options.objects``), one object's bitstream directory, or an + `orbitvega.scene_export` directory. + """ + options = options or VegaExportOptions() + out_dir = Path(out_dir).expanduser().resolve() + title, clips, detail = build_clips(source, out_dir, options) + bundle.write(out_dir, title=title, source=source, clips=clips, fps=options.fps, detail=detail) + return out_dir diff --git a/open4d/reconstruction/gs_tools/gs_tools/outputs.py b/open4d/reconstruction/gs_tools/gs_tools/outputs.py new file mode 100644 index 00000000..71f2f9db --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools/outputs.py @@ -0,0 +1,356 @@ +"""Recognizing what kind of *output* a directory holds, without opening it fully. + +``data/layouts.py`` does this for scenes going in; this does it for results +coming out, and it exists because the four things this module has to view are +four unrelated containers. QUEEN and 3DGStream leave a 3DGS run directory. +Vega leaves per-object ``frame_XXXX.pt`` chunks whose colour is a hash grid, +so nothing can read it but Vega. ReRF leaves a compressed feature voxel grid -- +DCT-coded, not Gaussians at all -- which only its own decoder can open, and only +under Python 3.8. Telling them apart from the file names is what lets +``gs-tools view`` take a path and do the right thing with it. + +Detection is deliberately cheap: file existence and header-sized reads only, no +torch and no CUDA, so it stays usable on a laptop with nothing installed. +""" + +from __future__ import annotations + +import json +import re +from dataclasses import dataclass, field +from enum import Enum +from pathlib import Path +from typing import Any + +IMAGE_SUFFIXES = (".jpg", ".jpeg", ".png") + +#: The index file ``gs-tools export`` writes at the root of a viewable bundle. +BUNDLE_NAME = "view.json" + + +class Kind(str, Enum): + """The output containers ``gs-tools view`` and ``gs-tools export`` accept.""" + + #: A bundle this module's own exporter produced: `view.json` + frames. + BUNDLE = "bundle" + #: `orbitvega.prepare` output: catalog.json plus one directory per object. + VEGA_CATALOG = "vega-catalog" + #: One object's Vega bitstream: manifest.json + color_model.pt + frame_*.pt. + VEGA_BITSTREAM = "vega-bitstream" + #: `orbitvega.scene_export` output: the merged scene, colour already baked. + VEGA_SCENE_EXPORT = "vega-scene-export" + #: A ReRF training run root: config.py, with the bitstream in a subdirectory. + RERF_RUN = "rerf-run" + #: ReRF's compressed bitstream: model_kwargs.json + header_*.json + feature_*. + RERF_BITSTREAM = "rerf-bitstream" + #: A 3DGS run in any of the three layouts `gaussian_frames` resolves: + #: a single frame's `point_cloud/iteration_*/`, 3DGStream's per-frame + #: `frameNNNNNN/`, or QUEEN's `frames/NNNN/`. + GAUSSIAN_RUN = "gaussian-run" + #: The ORBIT Gaussian-training corpus: dataset.json listing every object. + ORBIT_CORPUS = "orbit-corpus" + #: One ORBIT object: frame_NNNNNN/ directories with transforms.json + images. + ORBIT_SCENE = "orbit-scene" + #: A directory of numbered images -- e.g. ReRF's own `render_360_rerf_N`. + IMAGE_SEQUENCE = "image-sequence" + UNKNOWN = "unknown" + + +@dataclass(frozen=True) +class Detected: + """What was found at a path, and the few facts a caller needs to act on it.""" + + root: Path + kind: Kind + detail: dict[str, Any] = field(default_factory=dict) + + @property + def viewable(self) -> bool: + """Whether `gs-tools export` has a path for this, not merely recognised it. + + These used to differ: anything but UNKNOWN called itself viewable, so a + 3DGS run passed `inspect` and then failed `export` with "no exporter for + gaussian-run". A kind with no entry in :data:`EXPORTER_FOR` is detected + but not yet exportable, and saying so here is the difference between a + clear refusal and a puzzling one. + """ + return self.kind is Kind.BUNDLE or self.kind in EXPORTER_FOR + + +#: Which exporter handles each kind. Names rather than modules because +#: `gs_tools.methods` imports this module, so the modules cannot be imported +#: from here; `gs_tools.cli` maps the names onto them and is checked against +#: this, so the two cannot drift. +EXPORTER_FOR: dict[Kind, str] = { + Kind.VEGA_CATALOG: "vega", + Kind.VEGA_BITSTREAM: "vega", + Kind.VEGA_SCENE_EXPORT: "vega", + Kind.RERF_RUN: "rerf", + Kind.RERF_BITSTREAM: "rerf", + Kind.IMAGE_SEQUENCE: "rerf", + Kind.ORBIT_CORPUS: "captured", + Kind.ORBIT_SCENE: "captured", + Kind.GAUSSIAN_RUN: "gaussian", +} + + +def _best_iteration_ply(directory: Path) -> Path | None: + """The trained result in a `point_cloud/iteration_*/` directory. + + The highest iteration, which is the finished model rather than a checkpoint + on the way to it. `added/point_cloud.ply`, which 3DGStream writes beside the + frame's own model, is skipped: it holds only that frame's *newly added* + Gaussians, so treating it as a frame would show a fraction of the scene. + """ + found = sorted( + directory.glob("point_cloud/iteration_*/point_cloud.ply"), + key=lambda path: int(path.parent.name.split("_")[-1]), + ) + return found[-1] if found else None + + +def gaussian_frames(root: Path) -> list[tuple[int, Path]]: + """Ordered ``(frame index, PLY)`` for a 3DGS run, in any of three layouts. + + The layouts are not variations on one convention, they are three unrelated + ones, so each is matched rather than globbed for generically: + + * ``/frames/NNNN/point_cloud.ply`` -- QUEEN, one directory per frame, + no iteration level. + * ``/frameNNNNNN/point_cloud/iteration_N/point_cloud.ply`` -- + 3DGStream, per-frame runs each with their own iterations. + * ``/point_cloud/iteration_N/point_cloud.ply`` -- a single frame, which + is what a static 3DGS run or 3DGStream's init step produces. + + Frame numbers come from the directory names, so a run whose frames start at + 2 keeps its own numbering instead of being silently renumbered from zero. + """ + queen = sorted( + (int(path.parent.name), path) + for path in root.glob("frames/*/point_cloud.ply") + if path.parent.name.isdigit() + ) + if queen: + return queen + + gstream: list[tuple[int, Path]] = [] + for directory in sorted(root.glob("frame*")): + if not directory.is_dir(): + continue + digits = directory.name[len("frame"):] + ply = _best_iteration_ply(directory) if digits.isdigit() else None + if ply is not None: + gstream.append((int(digits), ply)) + if gstream: + return sorted(gstream) + + single = _best_iteration_ply(root) + return [(0, single)] if single is not None else [] + + +def _frame_pts(root: Path) -> list[Path]: + return sorted( + p for p in root.glob("frame_*.pt") if re.fullmatch(r"frame_\d+\.pt", p.name) + ) + + +def _images(root: Path) -> list[Path]: + """Numbered images, excluding ReRF's `NNN_depth.jpg` companions.""" + found = [ + p + for p in root.iterdir() + if p.suffix.lower() in IMAGE_SUFFIXES and re.fullmatch(r"\d+", p.stem) + ] + return sorted(found, key=lambda p: int(p.stem)) + + +def _rerf_headers(root: Path) -> list[Path]: + headers = [ + p for p in root.glob("header_*.json") if re.fullmatch(r"header_\d+\.json", p.name) + ] + return sorted(headers, key=lambda p: int(p.stem.split("_")[1])) + + +def _config_value(config: Path, key: str) -> str | None: + """One top-level string assignment out of a ReRF ``config.py``. + + Read with a regex rather than by importing: the file is an mmcv config, and + importing it needs mmcv, which lives only in the Python 3.8 ``nevo`` + environment. Detection has to work without that. + """ + try: + text = config.read_text() + except OSError: + return None + # Indented on purpose: `datadir` is nested inside `data = dict(...)`. + match = re.search(rf"^\s*{re.escape(key)}\s*=\s*['\"]([^'\"]*)['\"]", text, re.M) + return match.group(1) if match else None + + +def detect(path: Path | str) -> Detected: + """Identify an output directory without modifying it.""" + root = Path(path).expanduser().resolve() + if not root.is_dir(): + return Detected(root, Kind.UNKNOWN, {"reason": "not a directory"}) + + if (root / BUNDLE_NAME).is_file(): + try: + index = json.loads((root / BUNDLE_NAME).read_text()) + except (OSError, json.JSONDecodeError) as error: + return Detected(root, Kind.UNKNOWN, {"reason": f"unreadable bundle: {error}"}) + clips = index.get("clips", []) + return Detected( + root, + Kind.BUNDLE, + { + "title": index.get("title"), + "clips": len(clips), + "frames": sum(len(clip.get("frames", [])) for clip in clips), + }, + ) + + index = root / "dataset.json" + if index.is_file(): + try: + dataset = json.loads(index.read_text()) + except (OSError, json.JSONDecodeError): + dataset = {} + if str(dataset.get("format", "")).startswith("orbit-"): + return Detected(root, Kind.ORBIT_CORPUS, { + "format": dataset.get("format"), + "objects": [entry["name"] for entry in dataset.get("objects", [])], + "frames": dataset.get("frame_limit"), + "views": dataset.get("views_per_frame"), + }) + + orbit_frames = sorted( + p for p in root.glob("frame_*") if p.is_dir() and p.name[6:].isdigit() + ) + if orbit_frames and (orbit_frames[0] / "transforms.json").is_file(): + return Detected(root, Kind.ORBIT_SCENE, { + "name": root.name, + "frames": len(orbit_frames), + "views": len(json.loads((orbit_frames[0] / "transforms.json").read_text()) + .get("frames", [])), + }) + + if (root / "catalog.json").is_file(): + try: + catalog = json.loads((root / "catalog.json").read_text()) + except (OSError, json.JSONDecodeError): + catalog = {} + objects = catalog.get("objects", []) + if objects: + return Detected( + root, + Kind.VEGA_CATALOG, + { + "baseline": catalog.get("baseline"), + "objects": [entry["name"] for entry in objects], + "frames": max((entry.get("frame_count", 0) for entry in objects), default=0), + }, + ) + + if (root / "scene_manifest.json").is_file() and _frame_pts(root): + try: + scene = json.loads((root / "scene_manifest.json").read_text()) + except (OSError, json.JSONDecodeError): + scene = {} + return Detected( + root, + Kind.VEGA_SCENE_EXPORT, + { + "frames": len(scene.get("frames") or _frame_pts(root)), + "layout": scene.get("layout"), + "objects": [entry["name"] for entry in scene.get("objects", [])], + }, + ) + + if (root / "color_model.pt").is_file() and (root / "manifest.json").is_file(): + try: + manifest = json.loads((root / "manifest.json").read_text()) + except (OSError, json.JSONDecodeError): + manifest = {} + frames = manifest.get("frames") or _frame_pts(root) + return Detected(root, Kind.VEGA_BITSTREAM, {"frames": len(frames), "name": root.name}) + + headers = _rerf_headers(root) + if headers and (root / "model_kwargs.json").is_file(): + return Detected( + root, + Kind.RERF_BITSTREAM, + {"frames": len(headers), "has_rgb_net": (root / "rgb_net.tar").is_file()}, + ) + + if (root / "config.py").is_file(): + bitstreams = sorted( + child.name + for child in root.iterdir() + if child.is_dir() and detect(child).kind is Kind.RERF_BITSTREAM + ) + renders = sorted( + child.name + for child in root.iterdir() + if child.is_dir() and child.name.startswith("render_") and _images(child) + ) + if bitstreams or renders: + return Detected( + root, + Kind.RERF_RUN, + { + "expname": _config_value(root / "config.py", "expname") or root.name, + "basedir": _config_value(root / "config.py", "basedir"), + "datadir": _config_value(root / "config.py", "datadir"), + "bitstreams": bitstreams, + "renders": renders, + }, + ) + + frames = gaussian_frames(root) + if frames: + return Detected( + root, + Kind.GAUSSIAN_RUN, + { + "frames": len(frames), + "first_frame": frames[0][0], + "last_frame": frames[-1][0], + "iterations": sorted( + { + int(ply.parent.name.split("_")[-1]) + for _, ply in frames + if ply.parent.name.startswith("iteration_") + } + ), + "method": (json.loads((root / "manifest.json").read_text()).get("method") + if (root / "manifest.json").is_file() else None), + }, + ) + + images = _images(root) + if images: + return Detected(root, Kind.IMAGE_SEQUENCE, {"frames": len(images)}) + + return Detected(root, Kind.UNKNOWN, {"reason": "no recognized output files"}) + + +def describe(found: Detected) -> str: + """One-line summary for the CLI.""" + parts = [f"kind={found.kind.value}"] + for key in ("baseline", "expname", "title", "name", "method", "layout"): + if found.detail.get(key): + parts.append(f"{key}={found.detail[key]}") + if found.detail.get("frames"): + parts.append(f"frames={found.detail['frames']}") + if found.detail.get("clips"): + parts.append(f"clips={found.detail['clips']}") + for key in ("objects", "bitstreams", "renders", "iterations"): + values = found.detail.get(key) + if values: + shown = ",".join(str(value) for value in values[:6]) + if len(values) > 6: + shown += f",+{len(values) - 6}" + parts.append(f"{key}=[{shown}]") + if found.detail.get("reason"): + parts.append(found.detail["reason"]) + return " ".join(parts) diff --git a/open4d/reconstruction/gs_tools/gs_tools_tests/test_compare.py b/open4d/reconstruction/gs_tools/gs_tools_tests/test_compare.py new file mode 100644 index 00000000..6ece0724 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools_tests/test_compare.py @@ -0,0 +1,185 @@ +"""The camera and comparison layer: rigs, captured views, and rig renders. + +Same discipline as `test_view_pipeline.py` -- no GPU, no torch, no corpus. What +is pinned here is the geometry and the plan, because those are what decide +whether two panes showing different things are showing them from the same place. +""" + +from __future__ import annotations + +import json +import math +from pathlib import Path + +import numpy as np +import pytest + +from gs_tools import cameras, outputs +from streamer import bundle +from gs_tools.methods import capture, rerf + + +def _ring_transforms(count: int = 8, radius: float = 3.0, height: float = 1.0) -> dict: + """A synthetic ORBIT `transforms.json`: a coplanar ring looking inward. + + Built the way the real corpus is -- OpenCV camera-to-world, x right, y down, + z forward -- so a convention slip in `read_orbit_rig` shows up as a pose that + is not orthonormal or not right-handed, which the tests below check for. + """ + centre = np.array([0.0, height, 0.0]) + frames = [] + for index in range(count): + angle = 2 * math.pi * index / count + eye = centre + np.array([radius * math.sin(angle), 0.0, radius * math.cos(angle)]) + forward = centre - eye + forward /= np.linalg.norm(forward) + right = np.cross(forward, np.array([0.0, 1.0, 0.0])) + right /= np.linalg.norm(right) + down = np.cross(forward, right) + c2w = np.eye(4) + c2w[:3, 0], c2w[:3, 1], c2w[:3, 2], c2w[:3, 3] = right, down, forward, eye + frames.append({ + "file_path": f"./images/view_{index:02d}.png", + # Deliberately wrong on purpose: the OpenGL matrix is what a reader + # must NOT pick up, so it is present and different. + "transform_matrix": np.eye(4).tolist(), + "camera_to_world_opencv": c2w.tolist(), + "view_id": index, + }) + return { + "camera_model": "OPENCV", "fl_x": 1000.0, "fl_y": 1000.0, + "cx": 512.0, "cy": 384.0, "w": 1024, "h": 768, + "bounds_min": [-0.5, 0.0, -0.5], "bounds_max": [0.5, 2.0, 0.5], + "frames": frames, + } + + +def _orbit_scene(root: Path, *, frames: int = 3, views: int = 8) -> Path: + root.mkdir(parents=True, exist_ok=True) + transforms = _ring_transforms(views) + for step in range(1, frames + 1): + frame = root / f"frame_{step:06d}" + (frame / "images").mkdir(parents=True, exist_ok=True) + (frame / "transforms.json").write_text(json.dumps(transforms)) + return root + + +# ----------------------------------------------------------------- cameras --- +def test_rig_poses_are_orthonormal_and_right_handed(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + assert len(rig.poses) == 8 + for pose in rig.poses: + basis = np.array([pose.right, pose.down, pose.forward]) + np.testing.assert_allclose(basis @ basis.T, np.eye(3), atol=1e-9) + # (right, down, forward) in that order, which is what the viewer's + # projection assumes; a flipped pair renders the scene mirrored. + np.testing.assert_allclose(np.cross(pose.right, pose.down), pose.forward, atol=1e-9) + + +def test_rig_reads_the_opencv_matrix_not_the_opengl_one(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + # The synthetic corpus sets transform_matrix to the identity; picking it up + # would put every camera at the origin. + assert not np.allclose(rig.poses[0].position, [0, 0, 0]) + + +def test_rig_geometry_matches_the_ring_it_was_built_from(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + np.testing.assert_allclose(rig.centre, [0, 1, 0], atol=1e-9) + assert rig.radius == pytest.approx(3.0, abs=1e-6) + assert math.degrees(rig.fov_y()) == pytest.approx( + 2 * math.degrees(math.atan(0.5 * 768 / 1000.0)), abs=1e-9) + + +def test_ring_path_lands_on_the_real_cameras(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + exact = cameras.ring_path(rig, 8) + assert [pose.view_id for pose in exact] == list(range(8)) + for a, b in zip(exact, rig.poses): + np.testing.assert_allclose(a.position, b.position, atol=1e-9) + + finer = cameras.ring_path(rig, 32) + assert len(finer) == 32 + # Every fourth sample is a station; the rest are synthesised and say so, so + # a caller can tell where ground truth exists. + assert [i for i, pose in enumerate(finer) if pose.view_id is not None] == list(range(0, 32, 4)) + for pose in finer: + assert np.linalg.norm(np.asarray(pose.position) - rig.centre) == pytest.approx(3.0, abs=1e-6) + + +def test_ring_path_rejects_a_nonsense_count(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + with pytest.raises(ValueError): + cameras.ring_path(rig, 0) + + +def test_pose_round_trips_through_its_matrix(tmp_path): + rig = cameras.read_orbit_rig(_orbit_scene(tmp_path / "subject")) + for pose in rig.poses: + again = cameras.Pose.from_c2w(pose.c2w(), pose.view_id) + np.testing.assert_allclose(again.position, pose.position, atol=1e-12) + np.testing.assert_allclose(again.forward, pose.forward, atol=1e-12) + + +# ----------------------------------------------------------------- corpus ---- +def test_detect_orbit_scene_and_corpus(tmp_path): + scene = _orbit_scene(tmp_path / "corpus" / "subject") + assert outputs.detect(scene).kind is outputs.Kind.ORBIT_SCENE + (tmp_path / "corpus" / "dataset.json").write_text(json.dumps( + {"format": "orbit-rgb-gaussian-training", "frame_limit": 3, + "views_per_frame": 8, "objects": [{"name": "subject"}]})) + found = outputs.detect(tmp_path / "corpus") + assert found.kind is outputs.Kind.ORBIT_CORPUS + assert found.detail["objects"] == ["subject"] + + +def test_rigs_for_skips_scenes_the_corpus_does_not_have(tmp_path): + _orbit_scene(tmp_path / "corpus" / "subject") + rigs = capture.rigs_for(tmp_path / "corpus", ["subject", "absent"]) + assert list(rigs) == ["subject"] + assert len(rigs["subject"]["poses"]) == 8 + + +def test_frame_dirs_are_time_ordered_past_nine(tmp_path): + scene = _orbit_scene(tmp_path / "subject", frames=12) + names = [p.name for p in capture.frame_dirs(scene)] + # Lexical order would put frame_000010 before frame_000002. + assert names[:3] == ["frame_000001", "frame_000002", "frame_000003"] + assert names[-1] == "frame_000012" + + +# ------------------------------------------------------------------ bundle --- +def test_frame_dir_never_lets_two_clips_share_a_directory(tmp_path): + first = bundle.frame_dir(tmp_path, "subject") + (first / "frame_0000.ply").write_bytes(b"x") + second = bundle.frame_dir(tmp_path, "subject") + assert first.name == "subject" and second.name == "subject-2" + + +def test_clip_carries_scene_method_and_camera(tmp_path): + clip = bundle.Clip(name="c", representation="pixels", scene="subject", method="rerf", camera=3, + frames=["c/frame_0000.jpg"]) + bundle.write(tmp_path, title="t", source="s", clips=[clip], + scenes={"subject": {"poses": []}}) + index = bundle.read(tmp_path) + stored = index["clips"][0] + assert (stored["scene"], stored["method"], stored["camera"]) == ("subject", "rerf", 3) + assert "subject" in index["scenes"] + + +# -------------------------------------------------------------- rig render --- +def _rerf_bitstream(root: Path, frames: int = 4) -> Path: + root.mkdir(parents=True, exist_ok=True) + (root / "model_kwargs.json").write_text(json.dumps({"xyz_min": [0] * 3, "xyz_max": [1] * 3})) + (root / "rgb_net.tar").write_bytes(b"") + for index in range(frames): + entries = ([{"origin_size": [13, 8, 16, 8], "quality": 99}] if index == 0 + else [{"origin_size": [7, 8, 16, 8], "quality": 99}, + {"origin_size": [6, 8, 16, 8], "quality": 98}]) + (root / f"header_{index}.json").write_text(json.dumps({"headers": entries})) + return root + + +def test_scene_name_strips_the_corpus_prefix(): + assert rerf.scene_name("g_basketball") == "basketball" + assert rerf.scene_name("basketball") == "basketball" diff --git a/open4d/reconstruction/gs_tools/gs_tools_tests/test_gaussian_export.py b/open4d/reconstruction/gs_tools/gs_tools_tests/test_gaussian_export.py new file mode 100644 index 00000000..a14a5eef --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools_tests/test_gaussian_export.py @@ -0,0 +1,235 @@ +"""The 3DGS exporter and the `.splat` delivery format.""" + +from __future__ import annotations + +import json + +import numpy as np +import pytest +from open4d.core import GaussianCloud, Representation + +from gs_tools.io import ply, splat +from gs_tools.methods import gaussian + +pytestmark = pytest.mark.cpu + + +def raw_gaussians(count: int = 5, *, n_rest: int = 0) -> dict: + """Arguments for `ply.write`, i.e. raw 3DGS training parameters.""" + rng = np.random.default_rng(count) + fields = { + "xyz": rng.normal(size=(count, 3)).astype(np.float32), + "scale_raw": rng.normal(size=(count, 3)).astype(np.float32) - 2.0, + "rot_raw": rng.normal(size=(count, 4)).astype(np.float32), + "opacity_raw": rng.normal(size=(count, 1)).astype(np.float32), + "sh_dc": rng.normal(size=(count, 3)).astype(np.float32) * 0.1, + } + if n_rest: + fields["sh_rest"] = rng.normal(size=(count, n_rest, 3)).astype(np.float32) * 0.01 + return fields + + +def queen_run(root, frames=(1, 2, 3), *, n_rest: int = 0): + for frame in frames: + target = root / "frames" / f"{frame:04d}" + target.mkdir(parents=True, exist_ok=True) + ply.write(target / "point_cloud.ply", **raw_gaussians(5, n_rest=n_rest)) + return root + + +def gstream_run(root, frames=(2, 3)): + for frame in frames: + target = root / f"frame{frame:06d}" / "point_cloud" / "iteration_150" + target.mkdir(parents=True, exist_ok=True) + ply.write(target / "point_cloud.ply", **raw_gaussians(4)) + return root + + +# ------------------------------------------------------------- the exporter --- + + +def test_exports_a_queen_run(tmp_path): + out = tmp_path / "bundle" + title, clips, _ = gaussian.build_clips(queen_run(tmp_path / "run"), out) + assert len(clips) == 1 + clip = clips[0] + assert clip.representation == "gaussians" + assert len(clip.frames) == 3 + assert clip.counts == [5, 5, 5] + for path in clip.frames: + assert (out / path).is_file() + + +def test_exports_a_gstream_run_keeping_its_frame_numbers(tmp_path): + out = tmp_path / "bundle" + _, clips, _ = gaussian.build_clips(gstream_run(tmp_path / "run"), out) + assert [name.split("_")[-1] for name in + (path.rsplit("/", 1)[-1].removesuffix(".ply") for path in clips[0].frames) + ] == ["0002", "0003"] + + +def test_the_clip_can_be_explored_because_gaussians_have_geometry(tmp_path): + _, clips, _ = gaussian.build_clips(queen_run(tmp_path / "run"), tmp_path / "b") + assert Representation(clips[0].representation).has_geometry is True + + +def test_frames_option_truncates(tmp_path): + _, clips, _ = gaussian.build_clips( + queen_run(tmp_path / "run"), + tmp_path / "b", + gaussian.GaussianExportOptions(frames=2), + ) + assert len(clips[0].frames) == 2 + + +def test_scene_and_method_can_be_set_to_line_up_with_another_method(tmp_path): + _, clips, _ = gaussian.build_clips( + queen_run(tmp_path / "run"), + tmp_path / "b", + gaussian.GaussianExportOptions(scene="basketball", method="queen"), + ) + assert (clips[0].scene, clips[0].method) == ("basketball", "queen") + + +def test_a_guessed_scene_name_says_so_in_the_notes(tmp_path): + _, clips, _ = gaussian.build_clips(queen_run(tmp_path / "myrun"), tmp_path / "b") + assert clips[0].scene == "myrun" + assert any("taken from the directory" in note for note in clips[0].notes) + + +def test_ply_frames_are_copied_byte_for_byte(tmp_path): + """Re-encoding would only add a chance to get an activation or SH order wrong.""" + run = queen_run(tmp_path / "run", frames=(1,), n_rest=15) + out = tmp_path / "b" + _, clips, _ = gaussian.build_clips(run, out) + original = (run / "frames" / "0001" / "point_cloud.ply").read_bytes() + assert (out / clips[0].frames[0]).read_bytes() == original + + +def test_sh_degree_is_reported_for_a_copied_run(tmp_path): + _, clips, _ = gaussian.build_clips( + queen_run(tmp_path / "run", frames=(1,), n_rest=15), tmp_path / "b" + ) + assert clips[0].detail["sh_degrees"] == [3] + assert any("sh_degree 3" in note for note in clips[0].notes) + + +def test_refuses_something_that_is_not_a_gaussian_run(tmp_path): + (tmp_path / "empty").mkdir() + with pytest.raises(ValueError, match="not a 3DGS run"): + gaussian.build_clips(tmp_path / "empty", tmp_path / "b") + + +def test_refuses_an_unknown_frame_format(tmp_path): + with pytest.raises(ValueError, match="unknown frame format"): + gaussian.build_clips( + queen_run(tmp_path / "run"), + tmp_path / "b", + gaussian.GaussianExportOptions(frame_format="obj"), + ) + + +def test_export_writes_a_readable_bundle(tmp_path): + out = gaussian.export(queen_run(tmp_path / "run"), tmp_path / "bundle") + index = json.loads((out / "view.json").read_text()) + assert index["clips"][0]["representation"] == "gaussians" + + +# -------------------------------------------------------------- .splat form --- + + +def test_splat_export_is_much_smaller_and_says_what_it_dropped(tmp_path): + run = queen_run(tmp_path / "run", frames=(1,), n_rest=15) + _, as_ply, _ = gaussian.build_clips(run, tmp_path / "p") + _, as_splat, _ = gaussian.build_clips( + run, tmp_path / "s", gaussian.GaussianExportOptions(frame_format="splat") + ) + ply_size = (tmp_path / "p" / as_ply[0].frames[0]).stat().st_size + splat_size = (tmp_path / "s" / as_splat[0].frames[0]).stat().st_size + assert splat_size < ply_size + assert as_splat[0].frames[0].endswith(".splat") + assert any("degree 0 is dropped" in note for note in as_splat[0].notes) + + +def test_a_splat_frame_is_exactly_32_bytes_per_gaussian(tmp_path): + _, clips, _ = gaussian.build_clips( + queen_run(tmp_path / "run", frames=(1,)), + tmp_path / "b", + gaussian.GaussianExportOptions(frame_format="splat"), + ) + path = tmp_path / "b" / clips[0].frames[0] + assert path.stat().st_size == 5 * splat.SPLAT_BYTES + assert splat.count(path) == 5 + + +def test_splat_round_trip_keeps_position_and_scale_exactly(tmp_path): + """Those are float32 in both forms; only colour, opacity and rotation quantise.""" + ply.write(tmp_path / "f.ply", **raw_gaussians(64)) + cloud = splat.from_ply(tmp_path / "f.ply") + back = splat.decode(splat.encode(cloud)) + assert np.array_equal(back.positions, cloud.positions) + assert np.array_equal(back.scales, cloud.scales) + assert np.abs(back.opacities - cloud.opacities).max() <= 1 / 255 + 1e-6 + assert np.abs(back.rotations - cloud.rotations).max() <= 1 / 128 + 1e-6 + assert np.abs(back.colors - cloud.colors).max() <= 1 / 255 + 1e-6 + + +def test_from_ply_activates_exactly_once(tmp_path): + fields = raw_gaussians(16) + ply.write(tmp_path / "f.ply", **fields) + cloud = splat.from_ply(tmp_path / "f.ply") + assert np.allclose(cloud.scales, np.exp(fields["scale_raw"]), atol=1e-6) + expected = 1 / (1 + np.exp(-fields["opacity_raw"].reshape(-1))) + assert np.allclose(cloud.opacities, expected, atol=1e-6) + assert np.allclose(np.linalg.norm(cloud.rotations, axis=1), 1.0, atol=1e-5) + + +def test_encode_requires_a_gaussian_cloud(): + with pytest.raises(TypeError, match="GaussianCloud"): + splat.encode({"positions": []}) + + +def test_encode_requires_colour(): + cloud = GaussianCloud( + positions=np.zeros((1, 3), dtype=np.float32), + scales=np.ones((1, 3), dtype=np.float32), + rotations=np.asarray([[1, 0, 0, 0]], dtype=np.float32), + opacities=np.asarray([0.5], dtype=np.float32), + ) + with pytest.raises(ValueError, match="no colors"): + splat.encode(cloud) + + +def test_count_rejects_a_file_that_is_not_a_splat(tmp_path): + path = tmp_path / "bad.splat" + path.write_bytes(bytes(33)) + with pytest.raises(ValueError, match="not a .splat"): + splat.count(path) + + +# ----------------------------------------------------------- the PLY reader --- + + +def test_ply_reader_handles_the_extra_int_column_queen_writes(tmp_path): + """QUEEN adds `property int vertex_id`; assuming all-float32 misreads every row.""" + fields = raw_gaussians(4) + source = tmp_path / "plain.ply" + ply.write(source, **fields) + original = ply.read(source) + + # Rebuild the same file with an int column appended to each vertex. + header, _, payload = source.read_bytes().partition(b"end_header\n") + header = header.replace(b"property float rot_3\n", + b"property float rot_3\nproperty int vertex_id\n") + stride = len(payload) // 4 + rows = [payload[i * stride:(i + 1) * stride] for i in range(4)] + mixed = header + b"end_header\n" + b"".join( + row + index.to_bytes(4, "little") for index, row in enumerate(rows) + ) + target = tmp_path / "mixed.ply" + target.write_bytes(mixed) + + parsed = ply.read(target) + assert parsed["count"] == 4 + assert np.array_equal(parsed["xyz"], original["xyz"]) + assert np.array_equal(parsed["rot_raw"], original["rot_raw"]) diff --git a/open4d/reconstruction/gs_tools/gs_tools_tests/test_representation_vocabulary.py b/open4d/reconstruction/gs_tools/gs_tools_tests/test_representation_vocabulary.py new file mode 100644 index 00000000..db5aa099 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools_tests/test_representation_vocabulary.py @@ -0,0 +1,171 @@ +"""`streamer`, `gs_tools` and `open4d.core` must name representations alike. + +These drifted apart once already: the viewer had `splats`/`images` while core +had only triangle meshes, and the cost was that adding a representation meant +editing every function that touched a clip. The point of the shared vocabulary +is that it stays shared, so it is asserted rather than trusted -- and now that +the producer and the streaming module are separate packages, nothing but a test +holds them to it. +""" + +from __future__ import annotations + +import re +from pathlib import Path + +import pytest + +from open4d.core import Representation + +from streamer import bundle +from streamer.client import viewer_path + +from gs_tools.methods import capture, rerf, vega + +CORE_VALUES = {member.value for member in Representation} + + +def _viewer_source() -> str: + return viewer_path().read_text() + + +def _worker_source() -> str: + """The decode worker, which is where the geometry codecs live.""" + return (viewer_path().parent / "worker.js").read_text() + + +def _codec_table(source: str) -> dict[str, set[str]]: + """The client's CODECS table: representation -> the suffixes it decodes.""" + body = re.search(r"const CODECS = \{(.*?)\n\};", source, re.DOTALL) + assert body, "CODECS table not found" + found: dict[str, set[str]] = {} + for name, entries in re.findall( + r"^ (\w+):\s*\{(.*?)\},$", body.group(1), re.MULTILINE + ): + found[name] = set(re.findall(r'"(\.[a-z0-9]+)"', entries)) + return found + + +def test_the_client_decodes_exactly_what_the_codec_registry_promises(): + """Python says a `.drc` mesh is client-decodable; the client must agree. + + Two hand-maintained tables in two languages -- a Python registry that tells a + producer what it may write, and a JavaScript one that decides what actually + parses. A promise on one side with no parser on the other is a blank pane, + which is the failure this registry exists to remove. + + Geometry is decoded in the worker and pixels on the main thread, so the two + halves are checked separately rather than against one table. That split is + deliberate: an `` needs the main thread, and a browser already decodes + one off it. + """ + from streamer import codecs + + worker = _codec_table(_worker_source()) + for representation in Representation: + promised = {spec.suffix for spec in codecs.client_decodable(representation)} + if representation.has_geometry: + assert worker.get(representation.value, set()) == promised, representation + else: + # Pixels never reach the worker. + assert representation.value not in worker + + +def test_the_page_decodes_pixels_and_delegates_the_rest(): + """The other half of the split, so neither side is left unasserted.""" + from streamer import codecs + + page = _viewer_source() + table = re.search(r"const REPRESENTATIONS = \{(.*?)\n\};", page, re.DOTALL) + assert table, "REPRESENTATIONS not found in the page" + for name, body in re.findall( + r"^ (\w+): \{(.*?)^ \},$", table.group(1), re.MULTILINE | re.DOTALL + ): + if Representation(name).has_geometry: + assert f'decodeFrame("{name}")' in body, name + else: + assert "decodeImage" in body, name + # And the suffixes Python promises for pixels are the ones decodeImage gets. + assert {spec.suffix for spec in codecs.client_decodable("pixels")} == { + ".jpg", ".jpeg", ".png" + } + + +def test_the_client_does_not_decode_what_python_says_is_server_side(): + """ReRF can never decode in a browser; a parser for it would be a lie.""" + from streamer import codecs + + worker = _codec_table(_worker_source()) + page = _viewer_source() + for spec in codecs.known(): + if spec.decodes == "server": + assert spec.suffix not in worker.get(spec.representation.value, set()) + assert spec.suffix not in page + + +def _registry_keys(source: str) -> set[str]: + """The keys of the viewer's REPRESENTATIONS object literal.""" + body = re.search( + r"const REPRESENTATIONS = \{(.*?)\n\};", source, re.DOTALL + ) + assert body, "REPRESENTATIONS literal not found in the viewer" + return set(re.findall(r"^ (\w+): \{", body.group(1), re.MULTILINE)) + + +def _legacy_map(source: str) -> dict[str, str]: + line = re.search(r"const LEGACY_KINDS = \{(.*?)\};", source) + assert line, "LEGACY_KINDS not found in the viewer" + return dict(re.findall(r"(\w+): \"(\w+)\"", line.group(1))) + + +def test_viewer_only_knows_representations_core_defines(): + """A key the viewer invents is a vocabulary fork; catch it here.""" + assert _registry_keys(_viewer_source()) <= CORE_VALUES + + +def test_the_viewer_renders_every_representation_core_defines(): + assert _registry_keys(_viewer_source()) == CORE_VALUES + + +def test_geometry_flags_agree_with_core(): + """`geometry:` in the viewer is `Representation.has_geometry`, not a second opinion.""" + source = _viewer_source() + body = re.search(r"const REPRESENTATIONS = \{(.*?)\n\};", source, re.DOTALL) + for name, flag in re.findall( + r"^ (\w+): \{\n\s*geometry: (true|false),", body.group(1), re.MULTILINE + ): + assert Representation(name).has_geometry is (flag == "true"), name + + +def test_legacy_kinds_migrate_onto_real_representations(): + mapping = _legacy_map(_viewer_source()) + assert mapping == {"splats": "gaussians", "images": "pixels"} + assert set(mapping.values()) <= CORE_VALUES + + +@pytest.mark.parametrize( + ("module", "expected"), + [(vega, "gaussians"), (rerf, "pixels"), (capture, "pixels")], +) +def test_every_exporter_declares_a_core_representation(module, expected): + source = Path(module.__file__).read_text() + found = set(re.findall(r'representation="(\w+)"', source)) + assert found == {expected} + assert found <= CORE_VALUES + + +def test_no_exporter_still_writes_the_v1_field(): + for module in (vega, rerf, capture): + assert 'kind="' not in Path(module.__file__).read_text() + + +def test_bundle_version_was_bumped_for_the_field_rename(): + """A v1 bundle carries `kind`; a reader has to be able to tell them apart.""" + assert bundle.VERSION >= 2 + + +def test_clip_has_no_kind_field(): + assert "kind" not in {field.name for field in bundle.dataclasses.fields(bundle.Clip)} + assert "representation" in { + field.name for field in bundle.dataclasses.fields(bundle.Clip) + } diff --git a/open4d/reconstruction/gs_tools/gs_tools_tests/test_vega_frame_format.py b/open4d/reconstruction/gs_tools/gs_tools_tests/test_vega_frame_format.py new file mode 100644 index 00000000..9c813fb1 --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools_tests/test_vega_frame_format.py @@ -0,0 +1,109 @@ +"""The format Vega's Gaussian frames are written in. + +`gs-tools export --method vega` wrote 3DGS PLY only, so the `.splat` clips in +the demo bundle came from an ad-hoc conversion that was not reproducible from +the tool. `--frame-format splat` now reaches this exporter too. + +For Vega specifically the choice is nearly free: this export bakes colour to a +single band before writing, so there are no spherical-harmonic coefficients +above degree 0 for `.splat` to drop. About half the bytes for nothing +structural -- which is why the note says so rather than leaving a reader to +assume a loss. + +Tested through `_write_frame` rather than through the exporter: the full path +needs Vega's own modules and CUDA, and the format decision does not. +""" +from __future__ import annotations + +import numpy as np +import pytest + +from gs_tools import io +from gs_tools.io import ply, splat +from gs_tools.methods import gaussian, vega + +pytestmark = pytest.mark.cpu + + +def _fields(count: int = 64) -> dict: + rng = np.random.default_rng(7) + return { + "xyz": rng.normal(size=(count, 3)).astype(np.float32), + "scale_raw": rng.normal(size=(count, 3)).astype(np.float32), + "rot_raw": rng.normal(size=(count, 4)).astype(np.float32), + "opacity_raw": rng.normal(size=(count, 1)).astype(np.float32), + "sh_dc": ply.rgb_to_sh_dc(rng.random((count, 3)).astype(np.float32)), + } + + +def test_the_format_list_is_defined_once(): + """Two exporters offer the choice, and two tuples of names is one that goes + stale -- adding a format would have to reach both.""" + assert vega.FORMATS is io.GAUSSIAN_FORMATS + assert gaussian.FORMATS is io.GAUSSIAN_FORMATS + + +def test_ply_is_still_the_default(tmp_path): + path = vega._write_frame(tmp_path / "frame_0000", + vega.VegaExportOptions(), **_fields()) + assert path.suffix == ".ply" + assert path.is_file() + + +def test_splat_is_written_and_the_ply_is_not_left_behind(tmp_path): + fields = _fields() + path = vega._write_frame(tmp_path / "frame_0000", + vega.VegaExportOptions(frame_format="splat"), + **fields) + assert path.suffix == ".splat" + # Keeping both would double the export for a file no client asks for. + assert not (tmp_path / "frame_0000.ply").exists() + # 32 bytes a Gaussian, exactly -- the client reads the count from the size. + assert path.stat().st_size == 32 * len(fields["xyz"]) + assert splat.count(path) == len(fields["xyz"]) + + +def test_splat_is_about_half_the_bytes(tmp_path): + fields = _fields() + as_ply = vega._write_frame(tmp_path / "a", vega.VegaExportOptions(), **fields) + as_splat = vega._write_frame( + tmp_path / "b", vega.VegaExportOptions(frame_format="splat"), **fields) + ratio = as_ply.stat().st_size / as_splat.stat().st_size + # Measured on the demo bundle: 3.85 MB a frame against 1.85. The bound is + # loose because a PLY header is a fixed cost and these frames are small. + assert 1.5 < ratio < 3.0, ratio + + +def test_position_and_scale_survive_the_conversion_exactly(tmp_path): + """What `.splat` quantises is colour, opacity and rotation. Position and + scale are float32 either way, so a geometry comparison against another + method is not being made against rounded coordinates.""" + fields = _fields() + path = vega._write_frame(tmp_path / "frame_0000", + vega.VegaExportOptions(frame_format="splat"), + **fields) + cloud = splat.decode(path.read_bytes()) + assert np.array_equal(cloud.positions, fields["xyz"]) + assert np.allclose(cloud.scales, np.exp(fields["scale_raw"]), rtol=1e-6) + + +def test_an_unknown_format_is_refused_before_anything_is_written(tmp_path): + with pytest.raises(ValueError, match="not one of ply, splat"): + vega._write_frame(tmp_path / "frame_0000", + vega.VegaExportOptions(frame_format="obj"), **_fields()) + assert not list(tmp_path.iterdir()) + + +def test_the_clip_says_what_the_format_costs(): + """A reader should not have to know what `.splat` drops.""" + assert vega._format_notes(vega.VegaExportOptions()) == [] + notes = vega._format_notes(vega.VegaExportOptions(frame_format="splat")) + assert len(notes) == 1 + note = notes[0] + assert "32 bytes" in note + # The claim that matters: nothing structural, because colour is already + # degree 0 here. Saying only "smaller" would understate it, and saying + # "lossless" would overstate it. + assert "nothing structural is dropped" in note + assert "quantised to 8 bits" in note + assert "position and scale are exact" in note diff --git a/open4d/reconstruction/gs_tools/gs_tools_tests/test_view_pipeline.py b/open4d/reconstruction/gs_tools/gs_tools_tests/test_view_pipeline.py new file mode 100644 index 00000000..c8db856e --- /dev/null +++ b/open4d/reconstruction/gs_tools/gs_tools_tests/test_view_pipeline.py @@ -0,0 +1,400 @@ +"""The parts of `gs-tools export` / `gs-tools view` that need no GPU. + +Deliberately torch-free and CUDA-free, so this runs on any machine: what it +covers is the format and detection logic, which is where a mistake is silent. +The decode paths themselves (Vega's colour model, ReRF's entropy coder) can only +be exercised on the training host, in two different environments, against data +that is not in the repository -- so they are checked there, by hand, and what is +pinned here is everything around them. +""" + +from __future__ import annotations + +import json +import urllib.request +from pathlib import Path + +import numpy as np +import pytest + +from gs_tools import outputs +from streamer import bundle + +from gs_tools.io import ply +from gs_tools.methods import rerf + + +# --------------------------------------------------------------------- PLY --- +def _gaussians(n: int, *, seed: int = 0) -> dict: + rng = np.random.default_rng(seed) + return { + "xyz": rng.normal(size=(n, 3)).astype(np.float32), + "scale_raw": rng.normal(-4, 1, size=(n, 3)).astype(np.float32), + "rot_raw": rng.normal(size=(n, 4)).astype(np.float32), + "opacity_raw": rng.normal(size=(n, 1)).astype(np.float32), + "sh_dc": rng.normal(size=(n, 3)).astype(np.float32), + } + + +def test_ply_round_trip_is_exact(tmp_path): + data = _gaussians(97) + path = ply.write(tmp_path / "frame.ply", **data) + read = ply.read(path) + assert read["count"] == 97 + assert read["sh_degree"] == 0 + for key, expected in data.items(): + np.testing.assert_array_equal(read[key], expected) + + +def test_ply_round_trip_carries_higher_sh_bands(tmp_path): + data = _gaussians(31, seed=1) + rng = np.random.default_rng(2) + # Degree 3 is 15 bands per channel, the shape both upstream trainers write. + rest = rng.normal(size=(31, 15, 3)).astype(np.float32) + path = ply.write(tmp_path / "frame.ply", **data, sh_rest=rest) + read = ply.read(path) + assert read["sh_degree"] == 3 + np.testing.assert_array_equal(read["sh_rest"], rest) + + +def test_ply_attribute_order_matches_upstream(): + # The order is what every other 3DGS reader indexes by name against; a + # reordering here would still round-trip through `read` and break everywhere + # else, so it is pinned rather than derived. + names = ply.attribute_names(0) + assert names[:9] == ["x", "y", "z", "nx", "ny", "nz", "f_dc_0", "f_dc_1", "f_dc_2"] + assert names[9:] == ["opacity", "scale_0", "scale_1", "scale_2", + "rot_0", "rot_1", "rot_2", "rot_3"] + assert ply.attribute_names(15)[9:12] == ["f_rest_0", "f_rest_1", "f_rest_2"] + + +def test_sh_dc_is_the_inverse_of_rgb(): + rgb = np.linspace(0, 1, 30, dtype=np.float32).reshape(10, 3) + np.testing.assert_allclose(ply.sh_dc_to_rgb(ply.rgb_to_sh_dc(rgb)), rgb, atol=1e-6) + + +def test_ply_count_reads_only_the_header(tmp_path): + path = ply.write(tmp_path / "frame.ply", **_gaussians(5)) + assert ply.count(path) == 5 + + +def test_ply_write_rejects_mismatched_row_counts(tmp_path): + data = _gaussians(10) + data["opacity_raw"] = data["opacity_raw"][:9] + with pytest.raises(ValueError, match="9 rows"): + ply.write(tmp_path / "frame.ply", **data) + + +def test_ply_read_rejects_ascii(tmp_path): + path = tmp_path / "ascii.ply" + path.write_text("ply\nformat ascii 1.0\nelement vertex 1\nproperty float x\nend_header\n0\n") + with pytest.raises(ValueError, match="ascii"): + ply.read(path) + + +# ----------------------------------------------------------------- outputs --- +def test_detect_vega_catalog(tmp_path): + (tmp_path / "catalog.json").write_text(json.dumps({ + "baseline": "Vega-ORBIT", + "objects": [{"name": "dancer", "dir": "dancer", "frame_count": 30}], + })) + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.VEGA_CATALOG + assert found.detail["objects"] == ["dancer"] + assert "vega-catalog" in outputs.describe(found) + + +def test_detect_vega_bitstream(tmp_path): + (tmp_path / "color_model.pt").write_bytes(b"") + (tmp_path / "manifest.json").write_text(json.dumps({"frames": [{"frame_idx": 0}]})) + assert outputs.detect(tmp_path).kind is outputs.Kind.VEGA_BITSTREAM + + +def test_detect_vega_scene_export(tmp_path): + (tmp_path / "scene_manifest.json").write_text(json.dumps( + {"layout": "row", "frames": [{"frame_idx": 0, "file": "frame_0000.pt"}], "objects": []})) + (tmp_path / "frame_0000.pt").write_bytes(b"") + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.VEGA_SCENE_EXPORT + assert found.detail["layout"] == "row" + + +def _rerf_bitstream(root: Path, *, frames: int = 4, pca: bool = True, group_size: int = 4) -> Path: + """A ReRF bitstream's headers, written the way `codec/compress.py` writes them.""" + root.mkdir(parents=True, exist_ok=True) + (root / "model_kwargs.json").write_text(json.dumps({"xyz_min": [0, 0, 0], "xyz_max": [1, 1, 1]})) + (root / "rgb_net.tar").write_bytes(b"") + for index in range(frames): + key = index % group_size == 0 + if key or not pca: + entries = [{"origin_size": [13, 8, 16, 8], "quality": 99}] + else: + entries = [{"origin_size": [7, 8, 16, 8], "quality": 99}, + {"origin_size": [6, 8, 16, 8], "quality": 98}] + (root / f"header_{index}.json").write_text(json.dumps({"headers": entries})) + return root + + +def test_detect_rerf_bitstream(tmp_path): + found = outputs.detect(_rerf_bitstream(tmp_path / "rerf")) + assert found.kind is outputs.Kind.RERF_BITSTREAM + assert found.detail["frames"] == 4 + + +def test_detect_rerf_run_lists_bitstreams_and_renders(tmp_path): + _rerf_bitstream(tmp_path / "rerf") + (tmp_path / "config.py").write_text( + "expname = 'g_x'\nbasedir = '/runs'\ndata = dict(\n datadir='/corpus/x',\n)\n") + render = tmp_path / "render_360_rerf_4" + render.mkdir() + for index in range(4): + (render / f"{index:03d}.jpg").write_bytes(b"") + (render / f"{index:03d}_depth.jpg").write_bytes(b"") + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.RERF_RUN + assert found.detail["expname"] == "g_x" + # Indented, because it is nested inside `data = dict(...)`. + assert found.detail["datadir"] == "/corpus/x" + assert found.detail["bitstreams"] == ["rerf"] + assert found.detail["renders"] == ["render_360_rerf_4"] + + +def test_detect_image_sequence_ignores_depth_companions(tmp_path): + for index in range(3): + (tmp_path / f"{index:03d}.jpg").write_bytes(b"") + (tmp_path / f"{index:03d}_depth.jpg").write_bytes(b"") + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.IMAGE_SEQUENCE + assert found.detail["frames"] == 3 + + +def test_detect_gaussian_run_single_frame(tmp_path): + """A static 3DGS run, or 3DGStream's init step: one frame, several checkpoints.""" + for iteration in (7000, 30000): + target = tmp_path / "point_cloud" / f"iteration_{iteration}" + target.mkdir(parents=True) + ply.write(target / "point_cloud.ply", **_gaussians(3)) + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.GAUSSIAN_RUN + assert found.detail["frames"] == 1 + # The finished model, not every checkpoint on the way to it. + assert found.detail["iterations"] == [30000] + assert outputs.gaussian_frames(tmp_path)[0][1].parent.name == "iteration_30000" + + +def test_detect_gaussian_run_queen_layout(tmp_path): + for frame in (1, 2, 3): + target = tmp_path / "frames" / f"{frame:04d}" + target.mkdir(parents=True) + ply.write(target / "point_cloud.ply", **_gaussians(3)) + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.GAUSSIAN_RUN + assert (found.detail["frames"], found.detail["first_frame"]) == (3, 1) + assert [index for index, _ in outputs.gaussian_frames(tmp_path)] == [1, 2, 3] + + +def test_detect_gaussian_run_gstream_layout(tmp_path): + for frame in (2, 3, 4): + target = tmp_path / f"frame{frame:06d}" / "point_cloud" / "iteration_150" + target.mkdir(parents=True) + ply.write(target / "point_cloud.ply", **_gaussians(3)) + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.GAUSSIAN_RUN + # Frame numbering is the run's own: a run starting at 2 is not renumbered. + assert [index for index, _ in outputs.gaussian_frames(tmp_path)] == [2, 3, 4] + assert found.detail["first_frame"] == 2 + + +def test_gstream_added_gaussians_are_not_a_frame(tmp_path): + """`added/` holds only that frame's new Gaussians, not the whole scene.""" + base = tmp_path / "frame000002" / "point_cloud" + (base / "iteration_150").mkdir(parents=True) + ply.write(base / "iteration_150" / "point_cloud.ply", **_gaussians(3)) + (base / "iteration_250" / "added").mkdir(parents=True) + ply.write(base / "iteration_250" / "added" / "point_cloud.ply", **_gaussians(1)) + frames = outputs.gaussian_frames(tmp_path) + assert len(frames) == 1 + assert frames[0][1].parent.name == "iteration_150" + + +def test_viewable_means_exportable_not_merely_recognised(tmp_path): + """These disagreed: `inspect` said viewable, then `export` refused.""" + assert not outputs.detect(tmp_path).viewable + for kind, name in outputs.EXPORTER_FOR.items(): + assert name in {"vega", "rerf", "captured", "gaussian"}, kind + + +def test_detect_unknown_and_missing(tmp_path): + assert outputs.detect(tmp_path).kind is outputs.Kind.UNKNOWN + assert not outputs.detect(tmp_path / "nope").viewable + + +# ------------------------------------------------------------------ bundle --- +def test_bundle_round_trip(tmp_path): + clip = bundle.Clip(name="dancer", representation="gaussians", frames=["dancer/frame_0000.ply"], + counts=[12], bounds_min=[0, 0, 0], bounds_max=[1, 1, 1], + notes=["colour baked"]) + bundle.write(tmp_path, title="Vega — dancer", source="/somewhere", clips=[clip], fps=24) + index = bundle.read(tmp_path) + assert index["version"] == bundle.VERSION + assert index["fps"] == 24 + assert index["clips"][0]["notes"] == ["colour baked"] + found = outputs.detect(tmp_path) + assert found.kind is outputs.Kind.BUNDLE + assert found.detail["frames"] == 1 + + +def test_bundle_read_of_a_plain_directory_is_empty(tmp_path): + assert bundle.read(tmp_path) == {} + + +# -------------------------------------------------------------------- ReRF --- +def test_bitstream_info_infers_the_pca_split_and_group_size(tmp_path): + info = rerf.bitstream_info(_rerf_bitstream(tmp_path / "rerf", frames=6, group_size=3)) + assert info["frames"] == 6 + assert info["key_frames"] == [0, 3] + assert info["group_size"] == 3 + assert info["pca"] is True + assert info["pca_chs"] == (7, 13) + assert info["feature_dim"] == 13 + assert info["quality"] == [99, 98] + assert info["grid"] == [8, 16, 8] + + +def test_bitstream_info_without_pca(tmp_path): + info = rerf.bitstream_info(_rerf_bitstream(tmp_path / "rerf", frames=3, pca=False)) + assert info["pca"] is False + assert info["pca_chs"] == () + # Every frame looks like a key frame with PCA off, so the group is one frame. + assert info["group_size"] == 1 + + +def test_bitstream_info_needs_a_bitstream(tmp_path): + with pytest.raises(FileNotFoundError, match="model_kwargs.json"): + rerf.bitstream_info(tmp_path) + + +def test_collect_separates_colour_from_depth(tmp_path): + images = tmp_path / "render_360_rerf_3" + images.mkdir() + for index in range(3): + (images / f"{index:03d}.jpg").write_bytes(b"colour") + (images / f"{index:03d}_depth.jpg").write_bytes(b"depth") + clips = rerf.collect(images, tmp_path / "out", "run-rerf", rerf.RerfRenderOptions()) + assert [clip.name for clip in clips] == ["run-rerf", "run-rerf-depth"] + assert all(clip.representation == "pixels" for clip in clips) + assert (tmp_path / "out" / "run-rerf" / "frame_0000.jpg").read_bytes() == b"colour" + assert (tmp_path / "out" / "run-rerf-depth" / "frame_0000.jpg").read_bytes() == b"depth" + + only_colour = rerf.collect(images, tmp_path / "out2", "run-rerf", + rerf.RerfRenderOptions(depth=False)) + assert [clip.name for clip in only_colour] == ["run-rerf"] + + +def test_export_of_an_image_sequence_needs_no_render(tmp_path): + images = tmp_path / "render_360_rerf_2" + images.mkdir() + for index in range(2): + (images / f"{index:03d}.jpg").write_bytes(b"colour") + out = rerf.export(images, tmp_path / "out", rerf.RerfRenderOptions()) + index = bundle.read(out) + assert index["title"].startswith("ReRF") + assert index["clips"][0]["frames"] == ["render_360_rerf_2/frame_0000.jpg", + "render_360_rerf_2/frame_0001.jpg"] + + +def test_export_needs_a_bitstream_choice_when_a_run_holds_several(tmp_path): + for name in ("rerf", "rerf_whitebg"): + _rerf_bitstream(tmp_path / name) + (tmp_path / "config.py").write_text("expname = 'g_x'\n") + with pytest.raises(ValueError, match="--bitstream"): + rerf.export(tmp_path, tmp_path / "out", rerf.RerfRenderOptions()) + + +# -------------------------------------------------------------------- view --- +def test_serve_hands_out_the_viewer_and_the_bundle(tmp_path): + from streamer import server as view + + frame = ply.write(tmp_path / "dancer" / "frame_0000.ply", **_gaussians(4)) + bundle.write(tmp_path, title="t", source="s", clips=[bundle.Clip( + name="dancer", representation="gaussians", + frames=[str(frame.relative_to(tmp_path))], counts=[4])]) + + server = view.serve(tmp_path, port=0, block=False) + base = f"http://127.0.0.1:{server.server_address[1]}" + try: + with urllib.request.urlopen(f"{base}/", timeout=10) as response: + assert response.headers["Content-Type"].startswith("text/html") + assert b"gs-tools view" in response.read() + with urllib.request.urlopen(f"{base}/view.json", timeout=10) as response: + assert json.loads(response.read())["clips"][0]["name"] == "dancer" + with urllib.request.urlopen(f"{base}/dancer/frame_0000.ply", timeout=10) as response: + # text/html here is the failure mode: the browser's PLY parse would + # succeed on the bytes but the fetch would be flagged as a mismatch. + assert response.headers["Content-Type"] == "application/octet-stream" + assert response.read()[:3] == b"ply" + finally: + server.shutdown() + server.server_close() + + +def test_serve_refuses_a_directory_that_is_not_a_bundle(tmp_path): + from streamer import server as view + + with pytest.raises(FileNotFoundError, match="view.json"): + view.serve(tmp_path, port=0, block=False) + + +# ------------------------------------------------- rendering moved out --- + + +def test_a_bitstream_with_no_render_names_the_renderer(tmp_path): + """Rendering left this module for `rerf_stream.export`, which renders at + the corpus's own intrinsics, can write quality rungs and emits the capture + rig. The error has to say so: the alternative is a person concluding that + ReRF cannot be bundled at all. + """ + _rerf_bitstream(tmp_path / "rerf") + (tmp_path / "config.py").write_text("expname = 'g_x'\n") + with pytest.raises(RuntimeError) as raised: + rerf.build_clips(tmp_path, tmp_path / "out", rerf.RerfRenderOptions()) + message = str(raised.value) + assert rerf.RENDERER in message + assert "streamer.adopt" in message + assert "--compression-path" in message + + +def test_the_module_no_longer_carries_a_renderer(): + """Structural: a leftover would shell out to a tree that is gone.""" + for name in ("render", "render_command", "render_dir", "render_at_rig", + "rig_render_dir", "collect_rig", "rerf_python"): + assert not hasattr(rerf, name), name + + +def test_it_points_at_the_tree_reef_is_actually_vendored_in(): + """`gs_tools.env` records the upstream commit a manifest's frames came + from, and it resolves this name. Pointing at the old tree made it None.""" + from gs_tools import paths + + assert rerf.upstream == "rerf" + assert (paths.upstream(rerf.upstream) / "upstream" / "run.py").is_file() + + +def test_an_existing_render_is_still_bundled(tmp_path): + """What this module is still the right home for: turning an image sequence + that already exists into clips.""" + run = tmp_path / "g_x" + images = run / "render_360_rerf_3" + images.mkdir(parents=True) + # config.py is what makes a directory a ReRF *run* rather than unknown. + (run / "config.py").write_text("expname = 'g_x'\n") + for index in range(3): + # collect() copies without decoding, so the bytes need not be an image. + (images / f"{index:03d}.jpg").write_bytes(b"colour") + _rerf_bitstream(run / "rerf") + title, clips, detail = rerf.build_clips( + run, tmp_path / "out", rerf.RerfRenderOptions(depth=False)) + assert clips and all(clip.representation == "pixels" for clip in clips) + # A run with a render prefers it and ignores the bitstream beside it: an + # existing render says which condition it is, a bitstream name does not. + assert "bitstreams" not in detail diff --git a/open4d/reconstruction/gs_tools/pyproject.toml b/open4d/reconstruction/gs_tools/pyproject.toml index 035cc648..eca37559 100644 --- a/open4d/reconstruction/gs_tools/pyproject.toml +++ b/open4d/reconstruction/gs_tools/pyproject.toml @@ -16,13 +16,17 @@ build-backend = "setuptools.build_meta" [project] name = "open4d-gs-tools" version = "0.1.0" -description = "Gaussian-splatting free-viewpoint-video reconstruction for Open4D (QUEEN, 3DGStream)" +description = "Gaussian-splatting free-viewpoint-video reconstruction for Open4D (QUEEN, 3DGStream), and viewing Vega and ReRF output" requires-python = ">=3.11,<3.13" # Not MIT, unlike the rest of Open4D: derived from QUEEN under the NVIDIA # License, which limits use to non-commercial research or evaluation. See # THIRD_PARTY.md. license = { text = "LicenseRef-NVIDIA-Non-Commercial AND MIT" } -dependencies = ["numpy"] +# `open4d` for `open4d.core.Representation`: the bundle manifest and the +# browser viewer name representations with core's vocabulary rather than +# their own, so this is a real dependency and not a convenience. Its own +# runtime requirement is numpy; everything heavy in Open4D is an extra. +dependencies = ["numpy", "open4d"] [project.scripts] gs-tools = "gs_tools.cli:main" @@ -30,3 +34,8 @@ gs-tools = "gs_tools.cli:main" [tool.setuptools.packages.find] where = ["."] include = ["gs_tools*"] + +# The viewer is data, not code, and `gs-tools view` reads it off disk at request +# time, so it has to be installed alongside the package. +[tool.setuptools.package-data] +"gs_tools.view" = ["viewer.html"] diff --git a/open4d/reconstruction/gs_tools/pytest.ini b/open4d/reconstruction/gs_tools/pytest.ini new file mode 100644 index 00000000..768c6d54 --- /dev/null +++ b/open4d/reconstruction/gs_tools/pytest.ini @@ -0,0 +1,8 @@ +# Pins pytest's rootdir inside this module, for the same reason the Vega and NeVo +# baselines do: without it pytest finds Open4D's root pyproject.toml, resolves +# this suite's package name upward through open4d/, and applies addopts and a +# strict marker list this tree never declared. +[pytest] +testpaths = gs_tools_tests +markers = + cpu: needs no GPU diff --git a/open4d/reconstruction/nevo/README.md b/open4d/reconstruction/nevo/README.md deleted file mode 100644 index 656d3dd1..00000000 --- a/open4d/reconstruction/nevo/README.md +++ /dev/null @@ -1,271 +0,0 @@ -# NeVo (ORBIT adaptation) - -An offline, trace-driven simulator for evaluating the streaming quality of - -> Nan Wu, Bo Chen, Ruizhi Cheng, Klara Nahrstedt, Bo Han. -> **"NeVo: Advancing Volumetric Video Streaming with Neural Content -> Representation."** ACM MobiCom 2025. - -NeVo itself has no released code. What it streams does: -[ReRF](https://github.com/aoliao12138/ReRF) (CVPR 2023), the streamable-NeRF -representation the paper builds on and benchmarks against, is vendored -unmodified under `rerf/` and is what every stage here loads, renders and -measures. `rerf/PATCHES.md` records exactly what was and was not touched. - -NeVo streams NeRF content rather than point clouds or meshes, so unlike every -other baseline in this repo it needs a *neural* volumetric video to stream. -ReRF's own dataset is licence-gated, so we train ReRF ourselves on -`ORBIT_datasets_gaussian` -- see `rerf/DATA.md`, and note the NeVo paper does -the same substitution for two of its six datasets. - -There are no sockets and no WebRTC anywhere in this baseline. It models byte -arrival: a bandwidth trace gives queueing delay, a loss trace gives drops, and -a deadline decides what counts as lost. That makes a run reproducible from a -seed and lets the ablations be exact rather than approximately re-measured. - -## Status - -Built and verified so far -- **steps 1 and 2 of the pipeline, plus the -importance CDF**, which is the evidence the rest of the design rests on: - -1. **Load** a ReRF feature voxel grid and its motion vectors (`nevo/sequence.py`). -2. **Score** every feature voxel's neural visibility by instrumenting ray - marching to emit `T_i * alpha_i` per sample and scatter-maxing it into a - per-voxel buffer (`nevo/importance.py`), then take the CDF (`nevo/cdf.py`). - -Plus a viewer (`orbitnevo/render_frames.py`, `orbitnevo/live_demo.py`, -`orbitnevo/report.py`) that plays the trained sequence back with the filtering -switchable, and two things needed to know whether step 2 is worth building -on: -`nevo/render.py` renders a reloaded frame against its training image (does -step 1 rebuild the checkpoint correctly?), and `nevo/filtering.py` + -`orbitnevo/filter_sweep.py` drop the sub-threshold voxels and score the result -(does the metric actually identify what is safe to discard?). - -Not built yet, deliberately, pending the verification below: - -3. Packetize surviving voxels (contiguous-block vs. interleaved mapping). -4. Simulate arrival: bandwidth trace -> queueing delay, loss trace -> drops, - plus RTT/2, with a 33 ms deadline. -5. Recover missing voxels (VRM: 3D CNN over 3x3x3 neighbours x 9 history - frames -> LSTM, with an availability-mask channel). -6. Render at the trace viewport; SSIM and LPIPS against the unfiltered grid. - -## What the verification says - -Full write-up in `RESULTS.md`; look at it before building on the paper's -numbers. The short version: - -- The instrumentation is faithful: our marched weights match ReRF's own - forward pass exactly, and reloaded frames render within a dB of what the - trainer logged. -- The long tail is real and the mechanism works, at roughly the scale claimed. - At the SSIM >= 0.98 bar the paper uses, filtering by neural visibility drops - **49-56%** of the non-empty feature voxels at ReRF's own 8^3 codec block - (two objects), **64%** at a 4^3 unit, and **69%** on a better-reconstructed - version of the same subject. -- **Caveat that matters for step 3:** the SSIM >= 0.98 bar those figures use - passes renders with plainly visible 8^3 block artefacts. SSIM forgives - spatially coherent error, which is exactly what dropping a block produces. - Fit the threshold against LPIPS, not SSIM alone. -- The specific figure "~60% of voxels below 0.025" is *not* an invariant. It - slides from 42% to 77% purely with the granularity at which a "feature - voxel" is defined, which the paper does not pin down -- and the paper itself - fits its threshold to an SSIM target rather than fixing it at 0.025 (on this - content the fitted value is 0.2, eight times the quoted one). Treat the - quality bar as the claim and the threshold as an output. - -## Layout - -``` -rerf/ vendored ReRF, unmodified (see rerf/PATCHES.md) - configs/nevo/ generated training configs -nevo/ the simulator, no ORBIT or harness dependencies - rerf_env.py make `import lib.dvgo` work outside upstream's wrapper - cameras.py rig geometry and the world -> normalised transform - sequence.py step 1: feature voxel grids + motion vectors - blocks.py ReRF's 8^3 codec block geometry and occupancy - importance.py step 2: instrumented ray marching -> per-voxel weights - viewports.py synthetic viewports and 6DoF trace decoding - cdf.py streaming CDF accumulation and the Figure 7 plot - render.py render a loaded frame; check it reloaded correctly - filtering.py drop the sub-threshold voxels and score what changed -orbitnevo/ ORBIT-corpus adapters and CLIs - prepare.py ORBIT -> ReRF-trainable NHR corpus - train.py drive ReRF training over a corpus - importance_cdf.py steps 1-2 end to end - filter_sweep.py what each threshold drops, and what it costs in SSIM - render_frames.py render the plain-ReRF and visibility-filtered conditions - live_demo.py stream those frames to a browser as MJPEG - rerf_cli.py run ReRF's own compress.py / rerf_render.py - rd_sweep.py bytes (real encoder) against quality, per threshold - report.py static page: the viewer, plus every other output -nevo_tests/ unit tests; the model-dependent ones skip without a run -``` - -## Environments - -This baseline needs **two**, and that is not incidental. ReRF's entropy coder -`ac_dc/` ships only as a CPython 3.8 binary with no sources, and importing -`lib.dvgo` pulls it in — so anything touching a ReRF model runs on 3.8, while -this repo itself requires 3.10+. - -| Stage | Environment | Why | -| --- | --- | --- | -| `orbitnevo/prepare.py` | repo env (`conda activate pytorch`, 3.10) | reuses DeltaStream's nvdiffrast rasteriser and this repo's 3.10-only type syntax | -| everything else | `conda activate nevo` (3.8) | ReRF's `ac_dc` binary | - -Creating the `nevo` environment: - -```bash -conda create -y -n nevo python=3.8 -conda activate nevo -pip install torch==2.4.1 torchvision==0.19.1 --index-url https://download.pytorch.org/whl/cu121 -pip install torch_scatter -f https://data.pyg.org/whl/torch-2.4.0+cu121.html -pip install "numpy<2" "mmcv==1.7.2" imageio imageio-ffmpeg opencv-python-headless \ - tqdm ipdb lpips pytorch_msssim bitarray scipy matplotlib einops pandas pytest -``` - -Torch 2.4.1 is the newest release that still builds for Python 3.8 *and* has -`sm_89` kernels for this box's RTX 4090s; upstream's pinned torch 1.12.1+cu116 -predates Ada and will not run here at all. `mmcv==1.7.2` is for `mmcv.Config`, -which `run.py` uses (mmcv 2.x moved it to mmengine). - -## Usage - -Build a ReRF-trainable corpus from `ORBIT_datasets_gaussian` (repo env). With -no `--objects` it prepares the scene configured in `vstream/config.py`, like -every other baseline here: - -```bash -conda activate pytorch -python -m baselines.NeVo.orbitnevo.prepare --output-dir ~/nevo_data_g -``` - -30 frames x 8 calibrated views per object, cropped to the subject's silhouette -and resampled to 1280x960. The crop matters: the corpus frames the whole -60-degree stage, so a standing subject is only ~25% of the frame height and -~6% of the pixels, and ReRF trains at 960x720. Cropping to the silhouette -union (3550x2662 for `basketball`) lifts that to ~9% of pixels and ~72% of -frame height without touching the calibration -- the crop goes into the -intrinsics. - -There is a second source, `--source mesh`, which rasterises ORBIT's textured -OBJ sequences on an arbitrary rig (48 views over four elevations by default) -using DeltaStream's nvdiffrast renderer. Better training data -- the prepared -corpus puts all 8 views on one horizontal ring, which leaves a NeRF free to -invent geometry above and below the subject -- but no longer the same pixels -the other baselines see. `RESULTS.md` reports both. - -Train the ReRF sequence (nevo env), one object at a time: - -```bash -conda activate nevo -python -m baselines.NeVo.orbitnevo.train \ - --corpus ~/nevo_data_g/basketball --expname basketball --frames 24 -``` - -Score the voxels and build the CDF: - -```bash -python -m baselines.NeVo.orbitnevo.importance_cdf \ - --config baselines/NeVo/rerf/configs/nevo/basketball.py \ - --out ~/nevo_results/basketball --viewports 300 --block-size 8 --verify -``` - -`--verify` checks that this module's transcription of ReRF's ray marching -returns exactly the weights the vendored model does; `--block-size` chooses -the granularity a "feature voxel" means (8 is ReRF's codec block, 1 is a -single grid entry). - -Sweep the filtering threshold against the quality bar: - -```bash -python -m baselines.NeVo.orbitnevo.filter_sweep \ - --config baselines/NeVo/rerf/configs/nevo/basketball.py \ - --out ~/nevo_results/basketball -``` - -### ReRF's own codec and renderer - -The baseline NeVo is measured against is plain ReRF, and upstream can produce it -end to end: `codec/compress.py` writes the compressed bitstream and -`rerf_render.py` decodes it and renders a 360-degree orbit. `orbitnevo/rerf_cli.py` -runs either one with the environment already set up, so no `LD_LIBRARY_PATH` -incantation is needed: - -```bash -python -m baselines.NeVo.orbitnevo.rerf_cli codec/compress.py \ - --model_path ~/nevo_runs/g_basketball --expr_name rerf \ - --quality 99 --pca --pca_chs 7,13 --frame_num 30 -python -m baselines.NeVo.orbitnevo.rerf_cli rerf_render.py \ - --config configs/nevo/g_basketball.py \ - --compression_path ~/nevo_runs/g_basketball/rerf \ - --render_360 30 --pca --pca_chs 7,13 -``` - -Three things worth knowing: - -- **`--pca` must match at both ends**, per upstream's README. It is ~13% smaller - on this content and is part of ReRF's published method, so it is worth passing. -- **`--render_360` must not exceed the number of compressed frames.** Frame ids - wrap on `cfg.frame_num`, but the decode stream is pulled sequentially, so - asking for more exhausts the iterator. -- **The bitstream is the whole deliverable.** A 30-frame object is ~25 MB of - `.rerf` files plus a 90 kB colour MLP, against ~19 GB of training checkpoints - -- roughly 800x. It decodes and renders without them, so the checkpoints can be - deleted once an object is compressed. Only NeVo's own measurements - (`importance_cdf`, `filter_sweep`, `rd_sweep`) still need them. - -Two upstream calls fail against modern dependencies and are patched at runtime by -`nevo/rerf_env.py:patch_dependencies` rather than by editing `rerf/`, which stays -byte-identical: `np.bool` (removed in numpy 1.24) in `codec.compress_utils.decode_pca`, -which every decode goes through, and `imageio.imwrite` on the single-channel depth -maps `rerf_render.py` writes. Without the first, `compress.py` truncates the -bitstream to one frame; without the second, rendering dies after the first frame. - -### Watch it - -Build and serve the page: - -```bash -python -m baselines.NeVo.orbitnevo.render_frames \ - --config baselines/NeVo/rerf/configs/nevo/g_basketball.py -python -m baselines.NeVo.orbitnevo.report --out ~/nevo_report --serve 8752 -``` - -`render_frames.py` writes the conditions the page compares -- plain ReRF and each -visibility-filtered threshold, plus the captured camera -- and `live_demo.py` -streams them as MJPEG (`--nevo-only` for just NeVo's output). - -Nothing renders NeRF in a browser and this simulator has no live path by -design, so the page indexes into precomputed renders along the two axes a viewer -moves along, time and viewpoint. - -Below the player the same page carries the diagnostics: the prepared views and -their mattes, each trained frame rendered from its reloaded checkpoint beside -its training image, a novel viewpoint the rig never saw, the filtering -difference amplified, and the CDF plots and threshold tables with the numbers -behind them. Pass `--skip-render` to build only the parts that need no GPU; -`--clean` rebuilds the thumbnails and leaves the playback grid alone. - -Tests: - -```bash -conda activate nevo -python -m pytest baselines/NeVo/nevo_tests -q -``` - -The model-dependent cases skip unless a trained sequence is present; point -`NEVO_TEST_CONFIG` at one to override which. - -## What this deliberately does not do - -- **No live streaming.** No V4DS wire protocol, no `system/Server`, no Quest - conditions in `scripts/user_study.py`. The point of this baseline is a - reproducible offline measurement of neural-content streaming quality; the - paper's own headset numbers come from a HoloLens 2 client that does not - exist here. -- **No 3DGS.** The paper explicitly scopes it out (section 2.2) on the grounds - that 3DGS needs more bandwidth than NeRF for comparable quality; this repo's - `baselines/Vega` covers 3DGS streaming. diff --git a/open4d/reconstruction/nevo/RESULTS.md b/open4d/reconstruction/nevo/RESULTS.md deleted file mode 100644 index f4d7d694..00000000 --- a/open4d/reconstruction/nevo/RESULTS.md +++ /dev/null @@ -1,304 +0,0 @@ -# Steps 1-2: does the importance long tail reproduce? - -**The mechanism reproduces, at roughly the scale claimed.** Judged by the -paper's own criterion -- the largest threshold whose *worst* viewport still -clears SSIM 0.98 -- filtering by neural visibility discards, at ReRF's own -83 codec block, **49.2%** of `g_dancer`'s non-empty feature voxels -and **55.8%** of `g_basketball`'s, with no visible change. At a 43 -unit it is **64.0%**, and on a better-reconstructed version of the same -subject **69.0%**. The paper claims ~60%; the honest reading is that ~60% is -the right order and the exact figure is a property of the setup. - -**Two caveats, both of which change how step 3 should be built.** First, the -SSIM 0.98 bar those figures are measured against passes renders with visible -block artefacts -- see below -- so they are an upper bound on what a viewer -would accept, not an estimate of it. Second, the route the ~60% is usually -quoted by does not survive contact: "~60% of -voxels below 0.025" slides from 42% to 77% purely with how big a "feature -voxel" is defined to be, which the paper never pins down; and the paper does -not hold 0.025 fixed either, since section 3.2 *fits* the threshold per video -against an SSIM target. On this content the fitted threshold is 0.2, eight -times the quoted one, and 0.025 leaves ~19 points of saving unclaimed. Treat -the quality bar as the claim and the threshold as an output. - -Primary corpus: `ORBIT_datasets_gaussian` -> `g_basketball` and `g_dancer`, 4 -trained ReRF frames each (1 I-frame + 3 P-frames). A mesh-rendered corpus of -the same subject is reported alongside as a robustness check. Raw output in `~/nevo_results/`; the same -material is browsable via `orbitnevo/report.py`. - -## First: is the instrumentation measuring ReRF, or measuring itself? - -Two checks, both run by `--verify`, both on an I-frame *and* a P-frame -- their -models are assembled differently (a P-frame's feature grid is a residual over -a motion-compensated predecessor), so one does not cover the other. - -| check | I-frame 0 | P-frame 1 | -| --- | --- | --- | -| our marched weights vs. `DirectVoxGO.forward`'s | identical, max abs diff **0.0** over 2,634,697 samples | identical, **0.0** over 3,444,294 samples | -| reloaded frame rendered at a training view | **38.22 dB** | **37.98 dB** | - -The first says `nevo/importance.py`'s transcription of ReRF's ray marching is -the same computation the vendored model does, so the weights being scattered -are the ones that actually rendered the frame. The second is the check that -matters more, because the first compares a model against *itself* and would -pass just as happily on a model reassembled wrongly from its checkpoint. - -## The distribution - -Per-viewport pooling -- every (voxel, viewport) pair is one sample, which is -what a single fetch can skip and therefore what converts to a bandwidth -saving. 300 sampled viewports per frame, importance = `max T_i * alpha_i` over -the samples in a voxel, over non-empty voxels only. - -| feature voxel | edge | non-empty | <0.01 | **<0.025** | <0.05 | never hit | median | -| --- | --- | --- | --- | --- | --- | --- | --- | -| 1 grid entry | 8.4 mm | 738k | 74.6% | **76.6%** | 79.4% | 52.1% | 0.0000 | -| 23 entries | 17 mm | 117k | 64.9% | **67.5%** | 70.6% | 43.8% | 0.0008 | -| 43 entries | 34 mm | 20.5k | 52.6% | **55.6%** | 58.9% | 33.3% | 0.0087 | -| 83 (ReRF's codec block) | 67 mm | 4.1k | 38.9% | **41.8%** | 44.6% | 22.9% | 0.0908 | - -![CDF](figures/importance_cdf_block4.png) - -Shape-wise this is Figure 7: a near-vertical rise off zero, a knee well below -0.1, then a flat tail to 1.0. The median voxel at entry granularity scores -0.0000 and at 43 scores 0.0087 -- two to three orders of magnitude -below the voxels carrying the image. - -### Why granularity decides the number - -Importance is a *max* over the samples in a voxel, so coarsening can only -raise it. An 83 block spans 67 mm of a 1.9 m subject and almost -always contains some front-facing surface; one visible sample makes the whole -block important. That is not an artefact of this setup, it is what the metric -is, and it means "N% of voxels are below a threshold" is a statement about the -filtering unit as much as about the content. - -Both ends are real engineering options, not just axis choices. Filtering at -83 is free because that is already the unit ReRF encodes and masks -(`codec/compress.py` splits the volume into 83 blocks and ships a -per-block bitfield). Filtering finer needs a sub-block mask the bitstream does -not currently carry. - -### What the tail is made of - -`never hit` is the share of voxels no surviving sample touched at all -- -outside the frustum, or fully behind an opaque surface. At 43 that -is 33.3 of the 55.6 points. The remaining ~22 points are voxels that *were* -sampled but contribute negligibly: the part a plain frustum-and-occlusion test -would miss, and the part that justifies computing neural visibility rather -than reusing ViVo's position-based test. - -## Second: what does filtering actually cost? - -The CDF only says voxels score low. This renders each viewport twice -- whole -grid vs. everything below a threshold dropped -- and scores the pair. Dropped -blocks are written back the way ReRF's decoder fills a block that never -arrived (raw density -4.1, zero features), not zeroed; raw density 0 activates -to a visible alpha and would paint fog. 83 blocks, -`orbitnevo/filter_sweep.py`. "Lossless" means the *worst* of the sampled -viewports still cleared SSIM 0.98, the bar the paper cites. - -`g_basketball`, gaussian corpus: - -| threshold | dropped | SSIM | worst SSIM | PSNR | | -| --- | --- | --- | --- | --- | --- | -| 0.01 | 33.5% | 0.9998 | 0.9997 | 56.4 dB | lossless | -| **0.025** (the quoted one) | **37.2%** | 0.9997 | 0.9993 | 54.1 dB | lossless | -| 0.05 | 41.0% | 0.9993 | 0.9983 | 49.1 dB | lossless | -| 0.1 | 46.0% | 0.9977 | 0.9957 | 41.2 dB | lossless | -| **0.2** | **55.8%** | **0.9921** | **0.9864** | 31.9 dB | **lossless** | -| 0.35 | 68.8% | 0.9828 | 0.9759 | 26.5 dB | degraded | -| 0.5 | 83.4% | 0.9715 | 0.9626 | 22.8 dB | degraded | -| 0.7 | 97.4% | 0.9620 | 0.9506 | 19.7 dB | degraded | - -`g_dancer`, same corpus and settings, for a second object: - -| threshold | dropped | SSIM | worst SSIM | | -| --- | --- | --- | --- | --- | -| 0.025 | 32.4% | 0.9997 | 0.9994 | lossless | -| 0.1 | 41.0% | 0.9978 | 0.9937 | lossless | -| **0.2** | **49.2%** | **0.9944** | **0.9873** | **lossless** | -| 0.35 | 62.5% | 0.9856 | 0.9760 | degraded | - -Same fitted threshold (0.2), six points less droppable. Content-dependent, as -the paper's own per-video threshold tuning implies. - -`g_basketball` again on the mesh corpus (48 views over four elevations instead -of 8 on one ring), which reconstructs the subject better: - -| threshold | dropped | SSIM | worst SSIM | | -| --- | --- | --- | --- | --- | -| 0.025 | 41.6% | 0.9999 | 0.9998 | lossless | -| 0.2 | 55.9% | 0.9976 | 0.9955 | lossless | -| **0.35** | **69.0%** | **0.9918** | **0.9881** | **lossless** | -| 0.5 | 86.7% | 0.9817 | 0.9775 | degraded | - -The gap between the two corpora -- 55.8% vs 69.0% on identical subject motion --- is the interesting part. It is -*not* that the sparse rig scores voxels differently; the CDFs are within three -points of each other everywhere. It is that an 8-view reconstruction carries -more marginal, low-confidence geometry that is nonetheless load-bearing for -the image, so pushing the threshold hurts sooner. - -You can see it directly. A viewpoint halfway between two of the corpus's eight -cameras -- one the model never saw: - -![novel view from the 8-view corpus](figures/novel_view_gaussian.jpg) - -The silhouette and the ball are clean, but the shorts and shins are mottled: -low-confidence volume the rig could not pin down, which still contributes to -the pixels. That is the material a higher threshold starts eating. - -**Filtering headroom is a property of the reconstruction, not just of the -metric** -- worth remembering before quoting any single bandwidth-saving -figure, and a reason `report.py` puts a novel view on the page next to the -training view. - -### Granularity buys headroom here too - -Running the same sweep at a 43 filtering unit on `g_basketball`: -the largest lossless threshold moves to 0.1 and drops **64.0%**, against -55.8% at 83. Consistent with the CDF -- a finer unit isolates the -invisible content instead of averaging it with a visible neighbour -- and it -puts a number on what a sub-block mask would be worth: about 8 points of extra -saving, in exchange for a mask ReRF's bitstream does not currently carry. - -| filtering unit | best lossless threshold | dropped | -| --- | --- | --- | -| 83 (what ReRF already masks) | 0.2 | 55.8% | -| 43 (needs a sub-block mask) | 0.1 | 64.0% | - -Note the fitted threshold *falls* as the unit gets finer (0.2 -> 0.1) while -the saving rises. Anyone carrying a single hard-coded threshold across -granularities will get both numbers wrong. - -### The SSIM 0.98 bar is too loose for this kind of error - -Worth stating plainly, because it undercuts the tidy answer above. Here is the -difference between the unfiltered and filtered render, amplified 8x, at the -quoted threshold and at the fitted one: - -| 0.025 -- SSIM 0.9999, 49.4% dropped | 0.2 -- SSIM 0.9967, 64.9% dropped | -| --- | --- | -| ![](figures/diff_threshold_0.025.jpg) | ![](figures/diff_threshold_0.2.jpg) | - -Both clear SSIM 0.98 comfortably, and by that measure they are 0.003 apart. -They are not remotely equivalent to look at. At 0.025 the residual is faint, -unstructured noise. At 0.2 it is *coherent 83 blocks* -- you can -count them on the torso, the forearm and the shins -- and they are visible in -the filtered render itself, not only in the amplified difference: - -![filtered at 0.2](figures/filtered_threshold_0.2.jpg) - -This is the classic failure mode of SSIM: it is computed over a local window -and forgives error that is spatially coherent, which is exactly the shape of -error that dropping a *block* produces. So the "largest lossless threshold" -figures above (55.8%, 64.0%, 69.0%) should be read as **what the paper's -stated criterion permits, not as what a viewer would accept**. Under a bar -that penalises structure -- LPIPS, which the paper also reports -- the fitted -threshold and the saving will both come down. - -This is not an argument against the mechanism. At 0.025 the filter is -genuinely invisible and still removes 37-49%. It is an argument against -fitting the threshold on SSIM alone, which is what step 6 was going to do. - -This is a diagnostic with perfect knowledge of the viewport being rendered. It -is an upper bound: the real system filters against a viewport *predicted* -several frames ahead, which is what step 4 introduces. - -## Sensitivity checks - -Where the viewer stands barely matters. Both 43, 300 viewports, -gaussian corpus: - -| viewer spread | radius (x rig) | elevation | <0.025 (gaussian) | <0.025 (mesh) | -| --- | --- | --- | --- | --- | -| tight -- orbiting at capture distance, near eye level | 0.95-1.15 | -10 to 25 deg | 54.5% | 57.0% | -| default | 0.75-1.45 | -25 to 55 deg | 55.6% | 57.8% | -| wide | 0.6-2.0 | -40 to 70 deg | 56.8% | 58.7% | - -Two points across a range of viewer positions far wider than anyone actually -watches from. Granularity is the whole story. - -Nor does the *reconstruction* move it much. The same sweep on a corpus built -by re-rendering the ORBIT meshes on a 48-camera, four-elevation rig (rather -than the prepared 8-view horizontal ring): - -| feature voxel | gaussian corpus, 8 views | mesh corpus, 48 views | -| --- | --- | --- | -| 1 grid entry | 76.6% | 79.5% | -| 23 | 67.5% | 70.5% | -| 43 | 55.6% | 57.8% | -| 83 | 41.8% | 42.4% | - -Two to three points, always in the same direction: the sparser rig produces a -slightly *tighter* reconstruction with fewer dim voxels to discard. Worth -knowing, not worth worrying about. - -Assignment rule, at 83: charging a sample to the voxel it sits in -gives 41.8% (gaussian) / 42.4% (mesh); charging it to all eight entries its -trilinear interpolation reads gives 38.1% / 38.8%. Trilinear is the stricter notion ("which -voxels does rendering this actually depend on") and finds fewer droppable. The -paper's wording -- "sampled points inside it" -- is the former, so that is the -default. - -## Per-frame pooling, for contrast - -Scoring each voxel by its best viewport over the whole sequence: 14.2% below -0.025 at entry granularity, 5.2% at 83. That is the fraction never -worth sending *from any angle* -- a storage question, not a streaming one. -Quoting it as the bandwidth saving would be a 30-to-60-point mistake, which is -why both are reported. - -## One thing that surprised us - -ReRF's two "is anything here" tests disagree by six orders of magnitude. The -codec keeps a block when some entry's raw density clears ~3.39 -(`softplus(d - 4.1) > 0.4`, `compress_utils.get_masks`); the renderer keeps a -*sample* when its alpha clears `fast_color_thres = 1e-4`, i.e. raw density -above about -4.6. Low-density haze renders but is never transmitted. At -83 that leaks ~1% of the total rendered weight, at entry -granularity ~14%. Upstream's behaviour, not something NeVo introduces, but it -caps what any block-level filtering can preserve; reported per frame in -`importance_cdf_*.json` as `weight_outside_codec_mask`. - -## Verdict for step 3 onwards - -Worth building on. Two things the rest of the simulator should carry rather -than assume: - -1. **Filtering granularity is a first-class config flag**, not a constant. A - bandwidth saving reported without it is meaningless. The honest default is - 83, because that is what ReRF can actually mask and drop -- but - 43 is worth ~8 more points if step 3 is willing to carry a - sub-block mask, and that is a real design decision, not a free parameter. -2. **The threshold is fitted, but not on SSIM alone.** The paper's own - `Loss = SSIM_T - SSIM_C` is the right shape, and hard-coding 0.025 leaves - ~19 points unclaimed on this content (37.2% vs 55.8%). But SSIM 0.98 - passes renders with plainly visible block artefacts (above), so step 6 - should fit against LPIPS or a structure-aware bound and treat SSIM as a - floor rather than the target. -3. **Report the reconstruction alongside the saving.** The same metric, the - same threshold policy and the same subject give 55.8% or 69.0% depending on - how well the NeRF was fit. A saving quoted without that is not comparable - to anyone else's. - -## Reproducing - -```bash -conda activate pytorch -python -m baselines.NeVo.orbitnevo.prepare --objects basketball --output-dir ~/nevo_data_g -conda activate nevo -python -m baselines.NeVo.orbitnevo.train --corpus ~/nevo_data_g/basketball \ - --expname g_basketball --frames 24 -for B in 1 2 4 8; do - python -m baselines.NeVo.orbitnevo.importance_cdf \ - --config baselines/NeVo/rerf/configs/nevo/g_basketball.py \ - --out ~/nevo_results/g_basketball --viewports 300 --block-size $B \ - --tag block$B --verify -done -python -m baselines.NeVo.orbitnevo.filter_sweep \ - --config baselines/NeVo/rerf/configs/nevo/g_basketball.py \ - --out ~/nevo_results/g_basketball -python -m baselines.NeVo.orbitnevo.report --out ~/nevo_report --serve 8752 -``` diff --git a/open4d/reconstruction/nevo/citation.txt b/open4d/reconstruction/nevo/citation.txt deleted file mode 100644 index 87790fc3..00000000 --- a/open4d/reconstruction/nevo/citation.txt +++ /dev/null @@ -1,17 +0,0 @@ -Nan Wu, Bo Chen, Ruizhi Cheng, Klara Nahrstedt, and Bo Han. -"NeVo: Advancing Volumetric Video Streaming with Neural Content Representation." -The 31st Annual International Conference on Mobile Computing and Networking -(ACM MobiCom '25), November 4-8, 2025, Hong Kong, China. pp. 267-282. -https://doi.org/10.1145/3680207.3723473 - -The content representation it streams, vendored under rerf/: - -Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, -Lan Xu, and Minye Wu. -"Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos." -IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), -pp. 76-87. https://github.com/aoliao12138/ReRF - -ReRF is released for non-commercial use only. Its code base derives from DVGO -(Sun et al., "Direct Voxel Grid Optimization", CVPR 2022) and borrows from NHR -(Wu et al., "Multi-view Neural Human Rendering", CVPR 2020) and torch-dct. diff --git a/open4d/reconstruction/nevo/figures/diff_threshold_0.025.jpg b/open4d/reconstruction/nevo/figures/diff_threshold_0.025.jpg deleted file mode 100644 index 94c4a179..00000000 Binary files a/open4d/reconstruction/nevo/figures/diff_threshold_0.025.jpg and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/diff_threshold_0.2.jpg b/open4d/reconstruction/nevo/figures/diff_threshold_0.2.jpg deleted file mode 100644 index 1fb157ea..00000000 Binary files a/open4d/reconstruction/nevo/figures/diff_threshold_0.2.jpg and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/filtered_threshold_0.2.jpg b/open4d/reconstruction/nevo/figures/filtered_threshold_0.2.jpg deleted file mode 100644 index 2a7af3da..00000000 Binary files a/open4d/reconstruction/nevo/figures/filtered_threshold_0.2.jpg and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/importance_cdf_block1.png b/open4d/reconstruction/nevo/figures/importance_cdf_block1.png deleted file mode 100644 index 99484329..00000000 Binary files a/open4d/reconstruction/nevo/figures/importance_cdf_block1.png and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/importance_cdf_block4.png b/open4d/reconstruction/nevo/figures/importance_cdf_block4.png deleted file mode 100644 index f4e0a447..00000000 Binary files a/open4d/reconstruction/nevo/figures/importance_cdf_block4.png and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/importance_cdf_block8.png b/open4d/reconstruction/nevo/figures/importance_cdf_block8.png deleted file mode 100644 index 3071e66a..00000000 Binary files a/open4d/reconstruction/nevo/figures/importance_cdf_block8.png and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/novel_view_gaussian.jpg b/open4d/reconstruction/nevo/figures/novel_view_gaussian.jpg deleted file mode 100644 index ab025fc1..00000000 Binary files a/open4d/reconstruction/nevo/figures/novel_view_gaussian.jpg and /dev/null differ diff --git a/open4d/reconstruction/nevo/figures/reload_iframe.png b/open4d/reconstruction/nevo/figures/reload_iframe.png deleted file mode 100644 index 0d0f99b7..00000000 Binary files a/open4d/reconstruction/nevo/figures/reload_iframe.png and /dev/null differ diff --git a/open4d/reconstruction/nevo/nevo/__init__.py b/open4d/reconstruction/nevo/nevo/__init__.py deleted file mode 100644 index 2262488e..00000000 --- a/open4d/reconstruction/nevo/nevo/__init__.py +++ /dev/null @@ -1 +0,0 @@ -"""Offline trace-driven simulator for NeRF volumetric video streaming.""" diff --git a/open4d/reconstruction/nevo/nevo/bitstream.py b/open4d/reconstruction/nevo/nevo/bitstream.py deleted file mode 100644 index 8787b745..00000000 --- a/open4d/reconstruction/nevo/nevo/bitstream.py +++ /dev/null @@ -1,279 +0,0 @@ -"""Real ReRF bytes on the wire, for a chosen subset of feature voxels. - -The question this answers is "how many bytes does frame *t* cost if we only -send the blocks whose importance clears a threshold", and it answers it by -running ReRF's own encoder rather than by estimating. - -Why not attribute the joint bitstream's bytes to individual blocks: measured on -this codec, encoding one 13x8x8x8 block on its own costs ~9 kB almost -regardless of content, because each of the 13 channels gets its own file and -header. A 400-block frame costs 824 kB jointly and 3.96 MB block-by-block -- -4.8x -- and rescaling those per-block numbers mispredicts a half-subset's real -cost by 30%. Per-block attribution is not recoverable from this encoder, so -instead the retained subset is encoded *as a set*, which is exactly what ReRF -would transmit. - -What a frame costs, mirroring ``codec/compress.py``: - -``feature__.rerf*`` - The DCT + entropy-coded payload of the retained blocks, one file per - channel. This is the part a threshold changes. -``mask_.rerf`` - A packed bitfield, one bit per block of the padded grid, telling the - decoder which blocks are present. Its size does not depend on the - threshold -- ``ceil(blocks / 8)`` bytes either way -- but it is on the wire - and is counted. -``deform_.npy`` + ``deform_mask_.rerf`` - P-frames only: the motion vectors, fp16 with the all-zero entries dropped, - plus their own bitfield. - -Not counted per frame, and reported separately: ``rgb_net.tar`` (the colour MLP -shared by the whole sequence) and ``model_kwargs.json``. Those are a one-time -startup payload, not a per-frame cost. - -The residual chain is advanced **unfiltered**. A P-frame codes a residual -against the previous frame's reconstruction, so filtering frame *t-1* would -change what frame *t* has to encode; letting that compound would mean every -threshold measured a different bitstream. Holding the chain fixed measures one -well-defined thing -- the delivery cost of a subset of ReRF's own stream -- -which is the quantity that compares like-for-like against another codec's -output. -""" -from __future__ import annotations - -import glob -import json -import os -import shutil -import tempfile -from dataclasses import dataclass -from pathlib import Path -from typing import Optional - -import numpy as np - -from . import rerf_env -from .blocks import RERF_BLOCK_SIZE - -DENSITY_ACT = -4.1 -"""``density_act`` in codec/compress.py: the raw density an absent voxel decodes -to, and the offset an I-frame's density is coded relative to.""" - -DEFAULT_QUALITY = 99 -"""compress.py's default ``--quality``.""" - - -@dataclass(frozen=True) -class FrameBytes: - """Per-frame wire cost at one threshold.""" - - frame: int - key_frame: bool - kept_blocks: int - occupied_blocks: int - total_blocks: int - feature_bytes: int - mask_bytes: int - motion_bytes: int - - @property - def total_bytes(self) -> int: - return self.feature_bytes + self.mask_bytes + self.motion_bytes - - @property - def kept_fraction(self) -> float: - return self.kept_blocks / self.occupied_blocks if self.occupied_blocks else 0.0 - - def as_dict(self) -> dict: - payload = { - "frame": self.frame, - "key_frame": self.key_frame, - "kept_blocks": self.kept_blocks, - "occupied_blocks": self.occupied_blocks, - "total_blocks": self.total_blocks, - "feature_bytes": self.feature_bytes, - "mask_bytes": self.mask_bytes, - "motion_bytes": self.motion_bytes, - "total_bytes": self.total_bytes, - "kept_fraction": self.kept_fraction, - } - return payload - - -class SequenceCoder: - """Walks a ReRF sequence, producing each frame's codable residual and cost. - - Use as an iterator over frames: each :meth:`advance` yields the state - needed to price any number of thresholds for that frame, then moves the - reconstruction chain on. - """ - - def __init__(self, sequence, quality: int = DEFAULT_QUALITY, - voxel_size: int = RERF_BLOCK_SIZE, workdir: Optional[str] = None): - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from codec import ( # noqa: F401 - decode_motion, - deform_warp, - encode_entropy_motion_npy, - encode_jpeg_huffman, - encode_motion, - get_masks, - grid_sampler, - quant_motion, - split_volume, - zero_pads, - ) - - self._torch = torch - self._codec = { - "encode_jpeg_huffman": encode_jpeg_huffman, - "get_masks": get_masks, - "split_volume": split_volume, - "zero_pads": zero_pads, - "encode_motion": encode_motion, - "decode_motion": decode_motion, - "quant_motion": quant_motion, - "encode_entropy_motion_npy": encode_entropy_motion_npy, - "deform_warp": deform_warp, - "grid_sampler": grid_sampler, - } - self.sequence = sequence - self.quality = quality - self.voxel_size = voxel_size - self._owns_workdir = workdir is None - self._workdir = Path(workdir or tempfile.mkdtemp(prefix="nevo-bitstream-")) - self._workdir.mkdir(parents=True, exist_ok=True) - self._former = None # previous frame's reconstruction, [1, 13, X, Y, Z] - - def close(self) -> None: - if self._owns_workdir and self._workdir.exists(): - shutil.rmtree(self._workdir, ignore_errors=True) - - def __enter__(self): - return self - - def __exit__(self, *_): - self.close() - - # ------------------------------------------------------------------ bytes - def _encoded_size(self, blocks, tag: str) -> int: - """Bytes ReRF's entropy coder writes for these blocks, across channels.""" - target = self._workdir / tag - for stale in glob.glob(str(target) + "*"): - os.remove(stale) - if blocks.shape[0] == 0: - return 0 - self._codec["encode_jpeg_huffman"](blocks, self.quality, str(target)) - return sum(os.path.getsize(path) for path in glob.glob(str(target) + "*")) - - @staticmethod - def _bitfield_bytes(count: int) -> int: - """``bitarray.pack`` + ``tofile`` rounds up to whole bytes.""" - return (count + 7) // 8 - - def _motion_bytes(self, frame) -> int: - if frame.motion is None: - return 0 - cube, _grid_size, _origin = self._codec["encode_motion"](frame.motion) - payload, mask = self._codec["encode_entropy_motion_npy"](cube) - path = self._workdir / "deform.npy" - np.save(str(path), payload) - size = os.path.getsize(path) - os.remove(path) - return size + self._bitfield_bytes(int(mask.shape[0])) - - # ------------------------------------------------------------------ chain - def _residual_volume(self, frame): - """``residual_full`` and the occupancy mask, as compress.py computes them.""" - torch = self._torch - density = frame.density - features = frame.features - if self._former is None: - residual = torch.cat([density - DENSITY_ACT, features], dim=1) - occupancy_source = [torch.zeros_like(density) + DENSITY_ACT, density] - return residual, occupancy_source - - # Motion-compensate the previous reconstruction, then code the delta. - motion = self._codec["quant_motion"](frame.motion) if frame.motion is not None else None - xyz_min, xyz_max = frame.xyz_min, frame.xyz_max - former = self._former - if motion is not None: - grid_xyz = torch.stack( - torch.meshgrid( - torch.linspace(xyz_min[0], xyz_max[0], motion.shape[2]), - torch.linspace(xyz_min[1], xyz_max[1], motion.shape[3]), - torch.linspace(xyz_min[2], xyz_max[2], motion.shape[4]), - ), - -1, - ) - warped = self._codec["deform_warp"](grid_xyz, motion, xyz_min, xyz_max) - former = self._codec["grid_sampler"](warped, xyz_min, xyz_max, former) - former = former.permute(3, 0, 1, 2).unsqueeze(0) - residual = torch.cat([density - former[:, :1], features - former[:, 1:]], dim=1) - return residual, [former[:, :1], density] - - def advance(self, frame): - """Price ``frame`` and move the chain on. Returns a per-frame pricer.""" - torch = self._torch - with torch.no_grad(): - residual, occupancy_source = self._residual_volume(frame) - masks = self._codec["get_masks"]( - torch.cat(occupancy_source, dim=1), self.voxel_size - ) - blocks, _grid = self._codec["split_volume"]( - self._codec["zero_pads"](residual, voxel_size=self.voxel_size), - voxel_size=self.voxel_size, - ) - blocks = blocks.cuda() - motion_bytes = self._motion_bytes(frame) - mask_bytes = self._bitfield_bytes(int(masks.shape[0])) - # The chain advances on the *unfiltered* frame; see the module - # docstring for why thresholds must not compound. - self._former = torch.cat([frame.density, frame.features], dim=1) - - occupied = int(masks.sum().item()) - total = int(masks.shape[0]) - - def price(keep) -> FrameBytes: - selected = keep & masks if keep is not None else masks - kept = int(selected.sum().item()) - feature_bytes = self._encoded_size(blocks[selected], "feature") - return FrameBytes( - frame=frame.index, - key_frame=frame.is_key_frame, - kept_blocks=kept, - occupied_blocks=occupied, - total_blocks=total, - feature_bytes=feature_bytes, - mask_bytes=mask_bytes, - motion_bytes=motion_bytes, - ) - - return price - - -def startup_bytes(sequence) -> dict: - """The one-time payload: the shared colour MLP and the model header. - - Reported separately from per-frame cost because it is sent once for the - whole sequence. Any comparison against another representation has to agree - on whether this is amortised in or not. - """ - run_dir = Path(sequence.run_dir) - mlp = run_dir / "rgb_net.tar" - header = { - "rgb_net_bytes": os.path.getsize(mlp) if mlp.is_file() else 0, - "model_kwargs_bytes": 0, - } - kwargs = run_dir / "model_kwargs.json" - if kwargs.is_file(): - header["model_kwargs_bytes"] = os.path.getsize(kwargs) - header["total_bytes"] = header["rgb_net_bytes"] + header["model_kwargs_bytes"] - return header - - -def write_json(path, payload) -> None: - with open(path, "w") as handle: - json.dump(payload, handle, indent=1) diff --git a/open4d/reconstruction/nevo/nevo/blocks.py b/open4d/reconstruction/nevo/nevo/blocks.py deleted file mode 100644 index abb20b2b..00000000 --- a/open4d/reconstruction/nevo/nevo/blocks.py +++ /dev/null @@ -1,169 +0,0 @@ -"""Block geometry of a ReRF feature volume. - -ReRF does not transmit individual grid entries. ``codec/compress.py`` stacks -density and the 12 colour features into a ``[1, 13, X, Y, Z]`` volume, -zero-pads each spatial axis up to a multiple of ``voxel_size`` (8), splits it -into ``8x8x8`` blocks, drops the blocks that hold no occupied entry, and -DCT+entropy-codes the survivors. The kept/dropped decision is shipped as a -packed bitfield (``mask_.rerf``). - -So the *feature voxel* that NeVo talks about -- the thing whose neural -visibility gets scored, that gets filtered out of a transmission, and that a -packet carries a piece of -- is one of these blocks, not one grid entry. That -is also what Figure 6 of the NeVo paper draws: a grid far coarser than the -250^3 feature grid, a few dozen cells across the subject. - -This module reproduces that geometry exactly, including ``split_volume``'s -x-major block ordering, so a block index here is the same integer as the bit -position in ReRF's own mask and the same slot in its bitstream. Nothing -downstream has to guess an ordering. - -``block_size=1`` degenerates to per-entry accounting, which is useful for -showing how much of the paper's long tail is an artefact of granularity. -""" -from __future__ import annotations - -from dataclasses import dataclass -from typing import Sequence, Tuple - -import torch - -RERF_BLOCK_SIZE = 8 -"""``voxel_size`` in codec/compress.py and codec/compress_utils.py.""" - -_DENSITY_SHIFT = 4.1 -_OCCUPANCY_THRESHOLD = 0.4 -"""The codec's own occupancy test, ``get_masks``: an entry counts as occupied -when ``softplus(density - 4.1) > 0.4``, i.e. raw density above ~3.39. Note -this is a hardcoded constant upstream, *not* the model's ``act_shift`` (which -is -4.595 for ``alpha_init=1e-2``); mirrored rather than corrected so block -occupancy here matches the bitfield ReRF would actually send.""" - - -@dataclass(frozen=True) -class BlockGrid: - """Maps grid entries and world points onto ReRF block indices.""" - - grid_shape: Tuple[int, int, int] - block_size: int = RERF_BLOCK_SIZE - - def __post_init__(self) -> None: - if self.block_size < 1: - raise ValueError("block_size must be >= 1") - if any(int(n) < 1 for n in self.grid_shape): - raise ValueError(f"degenerate grid shape {self.grid_shape}") - - @classmethod - def from_volume(cls, volume: torch.Tensor, block_size: int = RERF_BLOCK_SIZE) -> "BlockGrid": - """``volume`` is ``[1, C, X, Y, Z]`` as ReRF stores density and k0.""" - if volume.dim() != 5: - raise ValueError(f"expected a [1, C, X, Y, Z] volume, got {tuple(volume.shape)}") - return cls(tuple(int(n) for n in volume.shape[2:]), block_size) - - @property - def padded_shape(self) -> Tuple[int, int, int]: - block = self.block_size - return tuple(-(-int(n) // block) * block for n in self.grid_shape) - - @property - def blocks_shape(self) -> Tuple[int, int, int]: - block = self.block_size - return tuple(n // block for n in self.padded_shape) - - @property - def num_blocks(self) -> int: - x, y, z = self.blocks_shape - return x * y * z - - def block_index(self, ijk: torch.Tensor) -> torch.Tensor: - """Flat block index for integer entry coordinates ``ijk`` of shape [N, 3]. - - The stride order matches ``codec.utils.split_volume``, which appends - blocks in ``for x: for y: for z:`` order. - """ - if ijk.dim() != 2 or ijk.shape[1] != 3: - raise ValueError(f"expected [N, 3] entry coordinates, got {tuple(ijk.shape)}") - _, by, bz = self.blocks_shape - block = self.block_size - coords = ijk // block - return (coords[:, 0] * by + coords[:, 1]) * bz + coords[:, 2] - - def occupancy(self, density: torch.Tensor) -> torch.Tensor: - """Per-block occupancy, mirroring ``codec.compress_utils.get_masks``. - - ``density`` is the raw (pre-activation) density volume ``[1, 1, X, Y, Z]``. - Returns a bool tensor of length :attr:`num_blocks`. - """ - if density.dim() != 5 or density.shape[:2] != (1, 1): - raise ValueError(f"expected a [1, 1, X, Y, Z] density volume, got {tuple(density.shape)}") - occupied = torch.nn.functional.softplus( - density.detach() - _DENSITY_SHIFT - ) > _OCCUPANCY_THRESHOLD - return self.reduce_any(occupied[0, 0]) - - def reduce_any(self, entry_mask: torch.Tensor) -> torch.Tensor: - """Reduce an ``[X, Y, Z]`` bool grid to one bool per block.""" - if tuple(entry_mask.shape) != self.grid_shape: - raise ValueError( - f"mask shape {tuple(entry_mask.shape)} does not match grid {self.grid_shape}" - ) - block = self.block_size - padded = torch.zeros(self.padded_shape, dtype=torch.bool, device=entry_mask.device) - x, y, z = self.grid_shape - padded[:x, :y, :z] = entry_mask - bx, by, bz = self.blocks_shape - return ( - padded.reshape(bx, block, by, block, bz, block) - .permute(0, 2, 4, 1, 3, 5) - .reshape(bx, by, bz, -1) - .any(dim=-1) - .reshape(-1) - ) - - -def nearest_entry( - points: torch.Tensor, - xyz_min: torch.Tensor, - xyz_max: torch.Tensor, - grid_shape: Sequence[int], -) -> torch.Tensor: - """Nearest grid entry to each world point, as integer ``[N, 3]``. - - ReRF samples its grids with ``align_corners=True``, so entry ``(i, j, k)`` - sits exactly at ``xyz_min + (i, j, k) / (shape - 1) * (xyz_max - xyz_min)`` - -- the entries are the *corners* of the sampled domain, not cell centres. - """ - size = torch.tensor( - [float(n) for n in grid_shape], device=points.device, dtype=points.dtype - ) - normalised = (points - xyz_min) / (xyz_max - xyz_min) - index = torch.round(normalised * (size - 1.0)) - return index.clamp_(torch.zeros_like(size), size - 1.0).long() - - -def surrounding_entries( - points: torch.Tensor, - xyz_min: torch.Tensor, - xyz_max: torch.Tensor, - grid_shape: Sequence[int], -) -> torch.Tensor: - """The eight entries a trilinear sample at each point actually reads. - - Returns ``[8, N, 3]``. Use this when "which feature voxels does rendering - this point depend on" is the question being asked, rather than the paper's - "sampled points inside a voxel" -- at a block boundary the two differ. - """ - size = torch.tensor( - [float(n) for n in grid_shape], device=points.device, dtype=points.dtype - ) - normalised = (points - xyz_min) / (xyz_max - xyz_min) * (size - 1.0) - low = torch.floor(normalised) - corners = [] - for offset in range(8): - delta = torch.tensor( - [float((offset >> shift) & 1) for shift in range(3)], - device=points.device, - dtype=points.dtype, - ) - corners.append((low + delta).clamp(torch.zeros_like(size), size - 1.0).long()) - return torch.stack(corners) diff --git a/open4d/reconstruction/nevo/nevo/cameras.py b/open4d/reconstruction/nevo/nevo/cameras.py deleted file mode 100644 index 136a23fb..00000000 --- a/open4d/reconstruction/nevo/nevo/cameras.py +++ /dev/null @@ -1,187 +0,0 @@ -"""Camera rigs and world normalisation for the NeVo ReRF corpus. - -Two coordinate frames appear throughout this baseline and mixing them up is -the single easiest way to train a NeRF that renders noise: - -*world* ORBIT's own metric frame, metres, as the OBJ sequences store it. - The subjects stand tens of metres from the origin (a basketball - player sits at x~2.9, z~14.9). -*normalised* What ReRF trains in. The rig centre becomes the origin and the - rig radius becomes exactly ``NORMALISED_RADIUS``. - -The normalisation matters because ReRF inherits DVGO's scale-sensitive -defaults -- ``inward_nearfar_heuristic`` derives near/far purely from how far -apart the camera positions are, ``alpha_init``/``fast_color_thres`` are tuned -for a unit-ish scene, and upstream's own ``data_util.py`` does the same thing -(mean-centre the cameras, then scale so the furthest sits at 2.0). Rendering -is invariant under a similarity transform applied to cameras *and* geometry -together, so we rasterise in the world frame with real metric cameras and -write the *normalised* extrinsics next to the images. Nothing ever has to -transform a mesh. -""" -from __future__ import annotations - -import math -from dataclasses import dataclass -from typing import Iterable, List, Sequence, Tuple - -import numpy as np - -NORMALISED_RADIUS = 2.0 -"""Rig radius in the normalised frame. Matches upstream ReRF's data_util.py, -which scales camera positions so the furthest lands at 2.0.""" - - -@dataclass(frozen=True) -class Camera: - """A pinhole camera. ``c2w`` is OpenCV camera-to-world (x right, y down, - z forward), which is what ReRF's NHR loader expects when the config sets - ``inverse_y=True``.""" - - camera_id: int - width: int - height: int - fx: float - fy: float - cx: float - cy: float - c2w: np.ndarray - - @property - def intrinsic_matrix(self) -> np.ndarray: - return np.asarray( - ((self.fx, 0.0, self.cx), (0.0, self.fy, self.cy), (0.0, 0.0, 1.0)), - dtype=np.float64, - ) - - def scaled_translation(self, centre: np.ndarray, scale: float) -> np.ndarray: - """This camera's ``c2w`` re-expressed in the normalised frame.""" - out = np.array(self.c2w, dtype=np.float64, copy=True) - out[:3, 3] = (out[:3, 3] - centre) * scale - return out - - -def look_at_c2w(eye: np.ndarray, target: np.ndarray, world_up: np.ndarray) -> np.ndarray: - """OpenCV camera-to-world for a camera at ``eye`` aimed at ``target``.""" - forward = target - eye - norm = np.linalg.norm(forward) - if norm < 1e-9: - raise ValueError("camera and target coincide") - forward = forward / norm - right = np.cross(forward, world_up) - right_norm = np.linalg.norm(right) - if right_norm < 1e-6: - # Looking straight along the up axis: any right vector in the plane - # orthogonal to forward will do, so pick one deterministically rather - # than emitting a degenerate rotation. - fallback = np.asarray((1.0, 0.0, 0.0)) if abs(world_up[0]) < 0.9 else np.asarray((0.0, 0.0, 1.0)) - right = np.cross(forward, fallback) - right_norm = np.linalg.norm(right) - right = right / right_norm - down = np.cross(forward, right) - c2w = np.eye(4, dtype=np.float64) - c2w[:3, 0] = right - c2w[:3, 1] = down - c2w[:3, 2] = forward - c2w[:3, 3] = eye - return c2w - - -def fit_radius( - extent: np.ndarray, - width: int, - height: int, - hfov_degrees: float, - margin: float = 1.12, -) -> float: - """Smallest orbit radius that keeps the whole bbox inside every frame. - - ``margin`` leaves headroom so a subject that swings an arm mid-sequence - does not clip the border -- the bbox is the sequence-wide union, but the - projection of a corner is only bounded by this expression when the camera - looks at the centre. - """ - hfov = math.radians(hfov_degrees) - vfov = 2.0 * math.atan(math.tan(hfov / 2.0) * height / width) - # Worst case the bbox presents its diagonal to the camera, and the camera - # is only `radius - depth/2` away from the near face. - lateral = 0.5 * math.hypot(extent[0], extent[2]) - required = max( - lateral / math.tan(hfov / 2.0), - 0.5 * extent[1] / math.tan(vfov / 2.0), - ) - return float((required + lateral) * margin) - - -def orbit_rig( - bounds_min: Iterable[float], - bounds_max: Iterable[float], - width: int, - height: int, - *, - azimuths: int = 12, - elevations: Sequence[float] = (-15.0, 5.0, 25.0, 45.0), - hfov_degrees: float = 60.0, - up_axis: int = 1, - margin: float = 1.12, -) -> Tuple[List[Camera], np.ndarray, float]: - """Build a multi-elevation orbit rig around a bounding box. - - Returns ``(cameras, centre, scale)`` where ``centre``/``scale`` define the - world -> normalised map ``p' = (p - centre) * scale``. - - ORBIT's own 8-camera corpus (``ORBIT_datasets_gaussian``) puts every view - on a single horizontal ring, which leaves a NeRF free to invent geometry - above and below the subject -- concavities under a chin or between arm and - torso are unconstrained. Since we rasterise the source meshes ourselves - there is no reason to inherit that limitation, so the default rig spans - four elevations. The azimuths are offset per elevation so the views do not - stack into vertical columns. - """ - lower = np.asarray(bounds_min, dtype=np.float64) - upper = np.asarray(bounds_max, dtype=np.float64) - if lower.shape != (3,) or upper.shape != (3,): - raise ValueError("bounds must be 3-vectors") - if np.any(upper < lower): - raise ValueError("bounds_max must not be below bounds_min") - if azimuths < 3 or not elevations: - raise ValueError("need at least 3 azimuths and 1 elevation") - - centre = (lower + upper) * 0.5 - extent = upper - lower - radius = fit_radius(extent, width, height, hfov_degrees, margin) - fx = width / (2.0 * math.tan(math.radians(hfov_degrees) / 2.0)) - - up = np.zeros(3) - up[up_axis] = 1.0 - plane = [axis for axis in range(3) if axis != up_axis] - - cameras: List[Camera] = [] - for row, elevation_degrees in enumerate(elevations): - elevation = math.radians(elevation_degrees) - # Half-step offset on odd rows: with the rows aligned, a vertical - # column of cameras sees almost the same silhouette three times and - # the azimuths in between stay unconstrained. - offset = (row % 2) * math.pi / azimuths - for index in range(azimuths): - angle = 2.0 * math.pi * index / azimuths + offset - direction = np.zeros(3) - direction[plane[0]] = math.cos(angle) * math.cos(elevation) - direction[plane[1]] = math.sin(angle) * math.cos(elevation) - direction[up_axis] = math.sin(elevation) - eye = centre + direction * radius - cameras.append( - Camera( - camera_id=len(cameras), - width=width, - height=height, - fx=fx, - fy=fx, - cx=(width - 1) * 0.5, - cy=(height - 1) * 0.5, - c2w=look_at_c2w(eye, centre, up), - ) - ) - - scale = NORMALISED_RADIUS / radius - return cameras, centre, scale diff --git a/open4d/reconstruction/nevo/nevo/cdf.py b/open4d/reconstruction/nevo/nevo/cdf.py deleted file mode 100644 index 3ed97648..00000000 --- a/open4d/reconstruction/nevo/nevo/cdf.py +++ /dev/null @@ -1,202 +0,0 @@ -"""The per-voxel importance CDF, and the check the paper's claim rests on. - -NeVo justifies visibility-aware filtering with Figure 7: the CDF of non-empty -voxels' importance is long-tailed, and "with a threshold of 0.025, ~60% of -voxels could be removed from delivery" while SSIM stays above 0.98. - -Two poolings, because they are different claims and only one of them supports -filtering: - -per-viewport (default) - Every (voxel, viewport) pair is one sample. This is the fraction of the - grid a *single* fetch can skip, so it is the one that converts to a - bandwidth saving, and the one to compare against the paper's ~60%. - -per-frame - A voxel is scored by its best importance over all viewports tested. This - is the fraction that is unimportant from *every* angle -- content that - could be dropped from storage, not just from one delivery. Necessarily a - smaller fraction. -""" -from __future__ import annotations - -import json -from typing import Dict, List, Optional, Sequence - -import numpy as np - -PAPER_THRESHOLD = 0.025 -PAPER_REMOVABLE_FRACTION = 0.60 - -DEFAULT_THRESHOLDS = (0.005, 0.01, 0.015, 0.02, 0.025, 0.03, 0.04, 0.05, 0.1, 0.2, 0.5) - - -DEFAULT_BINS = 400_000 -"""Histogram resolution used by :class:`ImportanceAccumulator`. - -Weights live in [0, 1] by construction, so 400k uniform bins put every edge -2.5e-6 apart -- and, deliberately, put an exact edge on 0.025 (bin 10000), the -threshold the paper's claim is stated at. Nothing is estimated at that point. -""" - - -class ImportanceAccumulator: - """Streams per-frame scores into a CDF without holding them all. - - At ``block_size=1`` a single object's scores are ``frames x viewports x - occupied_entries`` floats -- tens of gigabytes for a 24-frame sequence at - 300 viewports. Only the distribution is wanted, so bin as we go. The - per-frame pooling keeps one max per voxel per frame instead, which is - small enough to hold exactly. - """ - - def __init__(self, bins: int = DEFAULT_BINS): - if bins < 1000: - raise ValueError("bins must be >= 1000 to resolve the tail") - self.bins = int(bins) - self.counts = np.zeros(self.bins, dtype=np.int64) - self.total = 0 - self.sum = 0.0 - self.maximum = 0.0 - - def add(self, values: np.ndarray) -> None: - flat = np.asarray(values, dtype=np.float64).reshape(-1) - if flat.size == 0: - return - # Weights cannot exceed 1 (they are a partition of a pixel's opacity), - # but clip rather than trust it: a NaN or a >1 from a pathological - # checkpoint would otherwise land out of range and vanish silently. - if not np.all(np.isfinite(flat)): - raise ValueError("importance scores contain non-finite values") - index = np.clip((flat * self.bins).astype(np.int64), 0, self.bins - 1) - self.counts += np.bincount(index, minlength=self.bins) - self.total += int(flat.size) - self.sum += float(flat.sum()) - self.maximum = max(self.maximum, float(flat.max())) - - def fraction_below(self, threshold: float) -> float: - """Exact when ``threshold * bins`` is an integer, else rounded down to - the nearest bin edge.""" - if self.total == 0: - return float("nan") - edge = int(np.floor(threshold * self.bins)) - edge = max(0, min(self.bins, edge)) - return float(self.counts[:edge].sum() / self.total) - - def quantile(self, q: float) -> float: - if self.total == 0: - return float("nan") - target = q * self.total - cumulative = np.cumsum(self.counts) - index = int(np.searchsorted(cumulative, target, side="left")) - return float(min(index, self.bins - 1) / self.bins) - - @property - def mean(self) -> float: - return self.sum / self.total if self.total else float("nan") - - @property - def never_hit_fraction(self) -> float: - """Share of scored voxels that no surviving sample ever touched. - - These land in bin 0 and are the *free* part of the removable fraction: - a voxel outside this viewport's frustum, or fully behind an opaque - surface, contributes nothing at all rather than a little. Splitting it - out matters because it is the part a cheap frustum test would already - find -- what neural visibility buys over that is the mass between bin 0 - and the threshold. - """ - if self.total == 0: - return float("nan") - return float(self.counts[0] / self.total) - - def curve(self, points: int = 512) -> Dict[str, List[float]]: - if self.total == 0: - return {"importance": [], "cdf": []} - cumulative = np.cumsum(self.counts) / self.total - # Sample the curve where it moves: uniform sampling in *probability* - # rather than in importance, so the near-vertical tail is not one point. - probabilities = np.linspace(0.0, 1.0, points) - positions = np.searchsorted(cumulative, probabilities, side="left") - positions = np.clip(positions, 0, self.bins - 1) - return { - "importance": (positions / self.bins).tolist(), - "cdf": cumulative[positions].tolist(), - } - - def summary( - self, - *, - pooling: str, - block_size: int, - assignment: str, - frames: int, - viewports: int, - occupied_blocks: int, - total_blocks: int, - thresholds: Sequence[float] = DEFAULT_THRESHOLDS, - extra: Optional[Dict[str, object]] = None, - ) -> Dict[str, object]: - observed = self.fraction_below(PAPER_THRESHOLD) - payload = { - "pooling": pooling, - "block_size": block_size, - "assignment": assignment, - "frames": frames, - "viewports_per_frame": viewports, - "total_blocks": total_blocks, - "occupied_blocks": occupied_blocks, - "scored_samples": self.total, - "histogram_bins": self.bins, - "fraction_below": {str(t): self.fraction_below(t) for t in thresholds}, - "quantiles": {str(q): self.quantile(q) for q in (0.1, 0.25, 0.5, 0.75, 0.9, 0.99)}, - "mean": self.mean, - "max": self.maximum, - "never_hit_fraction": self.never_hit_fraction, - "paper_check": { - "threshold": PAPER_THRESHOLD, - "paper_fraction_below": PAPER_REMOVABLE_FRACTION, - "observed_fraction_below": observed, - "difference": observed - PAPER_REMOVABLE_FRACTION, - "long_tailed": bool(observed >= 0.5), - }, - } - if extra: - payload.update(extra) - return payload - - -def write_report(path, summaries: Sequence[Dict[str, object]], curves: Dict[str, object]) -> None: - payload = {"summaries": list(summaries), "curves": curves} - with open(path, "w") as handle: - json.dump(payload, handle, indent=1) - - -def plot(curves: Dict[str, Dict[str, List[float]]], path, title: str) -> None: - """Write a Figure-7-style CDF plot. Optional: needs matplotlib.""" - import matplotlib - - matplotlib.use("Agg") - import matplotlib.pyplot as plt - - figure, axis = plt.subplots(figsize=(5.0, 3.2)) - for label, curve in curves.items(): - axis.plot(curve["importance"], curve["cdf"], label=label, linewidth=1.6) - axis.axvline(PAPER_THRESHOLD, color="0.5", linestyle="--", linewidth=1.0) - axis.annotate( - "0.025", - xy=(PAPER_THRESHOLD, 0.04), - xytext=(PAPER_THRESHOLD + 0.02, 0.04), - color="0.35", - fontsize=8, - ) - axis.set_xlabel("Voxel importance (max $T_i\\,\\alpha_i$)") - axis.set_ylabel("CDF") - axis.set_xlim(0.0, 1.0) - axis.set_ylim(0.0, 1.0) - axis.grid(alpha=0.25, linewidth=0.6) - axis.set_title(title, fontsize=9) - axis.legend(fontsize=7, loc="lower right") - figure.tight_layout() - figure.savefig(path, dpi=200) - plt.close(figure) diff --git a/open4d/reconstruction/nevo/nevo/filtering.py b/open4d/reconstruction/nevo/nevo/filtering.py deleted file mode 100644 index 7b02eb3b..00000000 --- a/open4d/reconstruction/nevo/nevo/filtering.py +++ /dev/null @@ -1,139 +0,0 @@ -"""Drop the unimportant voxels and look at what changed. - -Step 2 produces a number per voxel. The claim that hangs off it -- that most of -them can go without anyone noticing -- is a claim about pixels, so it is worth -checking as pixels rather than only as a CDF. - -:func:`preview` renders a viewport twice from the same frame, once with the -whole grid and once with every block below a threshold removed, and scores the -pair. "Removed" means what ReRF's decoder means by a block that never arrived: -``rerf_render.py`` initialises missing blocks to a raw density of -4.1 and zero -features, so that is what gets written in, rather than a zero density (which -is *not* empty -- it activates to a substantial alpha). - -This is a diagnostic for step 2, not the streaming simulation. It filters a -frame in place with perfect knowledge of the viewport being rendered; the real -system has to decide from a *predicted* viewport several frames ahead, which -is what step 4 is for. -""" -from __future__ import annotations - -import contextlib -from dataclasses import dataclass -import numpy as np - -from . import rerf_env -from .blocks import BlockGrid -from .cameras import Camera - -ABSENT_DENSITY = -4.1 -"""What ReRF's decoder fills an undelivered block's density with. - -``rerf_render.py``: ``former_rec[:, 0, :] = former_rec[:, 0, :] - 4.1`` over a -zero-initialised buffer. Not 0.0 -- raw density 0 activates through -``softplus(0 - 4.595)`` to a visible alpha, so zeroing a block would paint fog -where the paper's filter leaves nothing. -""" - - -def entry_mask_from_blocks(grid: BlockGrid, block_mask): - """Expand a per-block bool to a per-entry bool over the unpadded grid.""" - bx, by, bz = grid.blocks_shape - block = grid.block_size - expanded = ( - block_mask.reshape(bx, by, bz) - .repeat_interleave(block, dim=0) - .repeat_interleave(block, dim=1) - .repeat_interleave(block, dim=2) - ) - x, y, z = grid.grid_shape - return expanded[:x, :y, :z] - - -@contextlib.contextmanager -def voxels_dropped(frame, grid: BlockGrid, keep_blocks): - """Temporarily strip the blocks outside ``keep_blocks`` from ``frame.model``. - - A context manager because the model is shared with the caller's - :class:`~nevo.sequence.ReRFSequence`; leaving a filtered grid behind would - silently corrupt every later render of the same frame. - """ - model = frame.model - keep = entry_mask_from_blocks(grid, keep_blocks) - original_density = model.density.data.clone() - original_k0 = model.k0.k0.data.clone() - original_former = None - if hasattr(model.k0, "former_k0_cur"): - original_former = model.k0.former_k0_cur.clone() - try: - drop = ~keep - model.density.data[0, 0][drop] = ABSENT_DENSITY - model.k0.k0.data[0, :, drop] = 0.0 - if original_former is not None: - # A P-frame renders `former_k0_cur + k0`; zeroing only the residual - # would leave the predecessor's features behind, which is not what - # a dropped block looks like to the decoder. - model.k0.former_k0_cur[0, :, drop] = 0.0 - yield model - finally: - model.density.data.copy_(original_density) - model.k0.k0.data.copy_(original_k0) - if original_former is not None: - model.k0.former_k0_cur.copy_(original_former) - - -@dataclass -class FilterPreview: - full: np.ndarray - filtered: np.ndarray - difference: np.ndarray - psnr: float - ssim: float - kept_blocks: int - total_blocks: int - threshold: float - - @property - def dropped_fraction(self) -> float: - return 1.0 - self.kept_blocks / self.total_blocks - - -def preview( - sequence, - frame, - camera: Camera, - scores, - grid: BlockGrid, - threshold: float, - occupancy=None, -) -> FilterPreview: - """Render ``camera`` with and without the sub-threshold blocks.""" - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from lib import utils - - from .render import psnr, render_view - - keep = scores >= threshold - if occupancy is not None: - # Blocks the codec would not have sent are absent either way, so they - # are not part of what filtering chooses to drop. - keep = keep & occupancy - considered = occupancy if occupancy is not None else torch.ones_like(keep) - - full = render_view(sequence, frame, camera) - with voxels_dropped(frame, grid, keep): - filtered = render_view(sequence, frame, camera) - - difference = np.abs(full - filtered) - return FilterPreview( - full=full, - filtered=filtered, - difference=difference, - psnr=psnr(filtered, full), - ssim=float(utils.rgb_ssim(filtered, full, max_val=1.0)), - kept_blocks=int(keep.sum().item()), - total_blocks=int(considered.sum().item()), - threshold=threshold, - ) diff --git a/open4d/reconstruction/nevo/nevo/importance.py b/open4d/reconstruction/nevo/nevo/importance.py deleted file mode 100644 index 9cb963b2..00000000 --- a/open4d/reconstruction/nevo/nevo/importance.py +++ /dev/null @@ -1,243 +0,0 @@ -"""Step 2: per-voxel neural visibility from instrumented ray marching. - -NeVo's central observation (paper section 3.2) is that the ray-marching weight - - w_i = T_i * alpha_i, alpha_i = 1 - exp(-sigma_i * delta_i), - T_i = prod_{j None: - if self.assignment not in ASSIGNMENTS: - raise ValueError(f"assignment must be one of {ASSIGNMENTS}") - if self.ray_chunk < 1 or self.render_factor < 1: - raise ValueError("ray_chunk and render_factor must be positive") - - -def _rays_for(dvgo, camera: Camera, render_kwargs: dict, torch): - c2w = torch.tensor(camera.c2w, dtype=torch.float32, device="cuda") - intrinsics = torch.tensor(camera.intrinsic_matrix, dtype=torch.float32, device="cuda") - rays_o, rays_d, viewdirs = dvgo.get_rays_of_a_view( - H=camera.height, - W=camera.width, - K=intrinsics, - c2w=c2w, - ndc=False, - inverse_y=render_kwargs["inverse_y"], - flip_x=render_kwargs["flip_x"], - flip_y=render_kwargs["flip_y"], - ) - return rays_o.flatten(0, -2), rays_d.flatten(0, -2) - - -def march_weights(model, dvgo, rays_o, rays_d, render_kwargs): - """Sample positions and their ray-marching weights ``T_i * alpha_i``. - - A transcription of ``lib.dvgo.DirectVoxGO.forward`` up to the point where - it would query colour. The order of the two ``fast_color_thres`` prunes is - load-bearing: transmittance is accumulated over the alpha-pruned sample - list, so pruning by weight first would change every downstream ``T_i``. - """ - count = len(rays_o) - ray_pts, ray_id, step_id = model.sample_ray( - rays_o=rays_o, - rays_d=rays_d, - near=render_kwargs["near"], - far=render_kwargs["far"], - stepsize=render_kwargs["stepsize"], - is_train=False, - ) - interval = render_kwargs["stepsize"] * model.voxel_size_ratio - - if model.mask_cache is not None: - if model.use_deform: - ray_pts = model.deform_warp(ray_pts, model.deformation_field) - keep = model.mask_cache(ray_pts) - ray_pts, ray_id = ray_pts[keep], ray_id[keep] - - density = model.grid_sampler(ray_pts, model.density) - alpha = model.activate_density(density, interval) - if model.fast_color_thres > 0: - keep = alpha > model.fast_color_thres - ray_pts, ray_id, alpha = ray_pts[keep], ray_id[keep], alpha[keep] - - weights, _ = dvgo.Alphas2Weights.apply(alpha, ray_id, count) - if model.fast_color_thres > 0: - keep = weights > model.fast_color_thres - ray_pts, weights = ray_pts[keep], weights[keep] - return ray_pts, weights - - -def scatter_max( - grid: BlockGrid, - points, - weights, - xyz_min, - xyz_max, - grid_shape, - assignment: str, - into, -): - """Fold sample weights into a per-block maximum, in place.""" - if points.numel() == 0: - return into - if assignment == "nearest": - entries = nearest_entry(points, xyz_min, xyz_max, grid_shape).unsqueeze(0) - source = weights.unsqueeze(0) - else: - entries = surrounding_entries(points, xyz_min, xyz_max, grid_shape) - source = weights.unsqueeze(0).expand(entries.shape[0], -1) - index = grid.block_index(entries.reshape(-1, 3)) - into.scatter_reduce_(0, index, source.reshape(-1), reduce="amax") - return into - - -class ImportanceScorer: - """Scores one frame's voxels against arbitrary viewports.""" - - def __init__(self, sequence, frame, config: ImportanceConfig = ImportanceConfig()): - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from lib import dvgo - - self._torch = torch - self._dvgo = dvgo - self.config = config - self.frame = frame - self.model = frame.model - self.render_kwargs = sequence.render_kwargs() - self.grid_shape = frame.grid_shape - self.grid = BlockGrid(self.grid_shape, config.block_size) - self.occupancy = self.grid.occupancy(frame.density) - - @property - def num_blocks(self) -> int: - return self.grid.num_blocks - - @property - def occupied_blocks(self) -> int: - return int(self.occupancy.sum().item()) - - def score(self, camera: Camera): - """Per-block importance for one viewport, as a ``[num_blocks]`` tensor.""" - torch = self._torch - with torch.no_grad(): - rays_o, rays_d = _rays_for(self._dvgo, camera, self.render_kwargs, torch) - scores = torch.zeros(self.grid.num_blocks, dtype=torch.float32, device="cuda") - for begin in range(0, len(rays_o), self.config.ray_chunk): - chunk = slice(begin, begin + self.config.ray_chunk) - points, weights = march_weights( - self.model, - self._dvgo, - rays_o[chunk].contiguous(), - rays_d[chunk].contiguous(), - self.render_kwargs, - ) - scatter_max( - self.grid, - points, - weights, - self.frame.xyz_min, - self.frame.xyz_max, - self.grid_shape, - self.config.assignment, - scores, - ) - return scores - - def score_many(self, cameras: Iterable[Camera], progress=None): - """Stack :meth:`score` over viewports into ``[views, num_blocks]``.""" - torch = self._torch - rows = [] - for index, camera in enumerate(cameras): - rows.append(self.score(camera).cpu()) - if progress is not None: - progress(index) - return torch.stack(rows) - - -def check_against_rerf(sequence, frame, camera: Camera, tolerance: float = 1e-5) -> dict: - """Assert :func:`march_weights` matches ReRF's own forward pass. - - ``DirectVoxGO.forward`` returns the weights it used but not the sample - positions, so the guarantee we need -- that our transcription produces the - same weight sequence -- has to be checked rather than assumed. Compares the - sorted weight vectors and the resulting pixel opacity. - """ - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from lib import dvgo - - model = frame.model - render_kwargs = sequence.render_kwargs() - with torch.no_grad(): - rays_o, rays_d = _rays_for(dvgo, camera, render_kwargs, torch) - viewdirs = rays_d / rays_d.norm(dim=-1, keepdim=True) - reference = model(rays_o, rays_d, viewdirs, **render_kwargs) - ours_points, ours_weights = march_weights(model, dvgo, rays_o, rays_d, render_kwargs) - theirs = reference["weights"] - matched = ours_weights.numel() == theirs.numel() - difference = float("inf") - if matched: - difference = float((ours_weights - theirs).abs().max().item()) - return { - "samples_ours": int(ours_weights.numel()), - "samples_rerf": int(theirs.numel()), - "max_abs_difference": difference, - "agrees": bool(matched and difference <= tolerance), - } diff --git a/open4d/reconstruction/nevo/nevo/metrics.py b/open4d/reconstruction/nevo/nevo/metrics.py deleted file mode 100644 index 8bb48a68..00000000 --- a/open4d/reconstruction/nevo/nevo/metrics.py +++ /dev/null @@ -1,109 +0,0 @@ -"""PSNR / SSIM / LPIPS against a held-out capture, scored on the subject. - -Two choices here decide whether the numbers mean anything, and both have to -hold for anything this gets compared against: - -**The reference is a captured image, not our own render.** Scoring a filtered -render against the unfiltered one measures how well filtering preserves *our -reconstruction*, which flatters it -- errors the NeRF already had cancel out. -Scoring against the real camera measures how good the delivered content is, -which is the question. - -**Metrics are computed on the subject's bounding box, not the whole frame.** -The corpus renders a body over a background that is identically white in both -images. Including it inflates PSNR without bound (~90% of the frame is a -perfect match) and pins SSIM near 1, so the interesting variation disappears -into the average. The crop is taken from the reference matte, unioned over the -frames being scored so it does not move between them. -""" -from __future__ import annotations - -from dataclasses import dataclass -from typing import Iterable, Optional, Sequence, Tuple - -import numpy as np - -from . import rerf_env - - -@dataclass(frozen=True) -class Quality: - psnr: float - ssim: float - lpips: float - - def as_dict(self) -> dict: - return {"psnr": self.psnr, "ssim": self.ssim, "lpips": self.lpips} - - -def silhouette_box(mattes: Iterable[np.ndarray], pad: int = 8) -> Tuple[int, int, int, int]: - """``(top, left, bottom, right)`` covering every matte, with a small pad. - - ``pad`` keeps a margin of background inside the crop so a filtering - artefact that spills just outside the silhouette still lands in the score. - """ - top, left = np.inf, np.inf - bottom, right = -np.inf, -np.inf - height = width = 0 - for matte in mattes: - height, width = matte.shape[:2] - rows = np.flatnonzero(matte.any(axis=1)) - columns = np.flatnonzero(matte.any(axis=0)) - if rows.size == 0 or columns.size == 0: - continue - top, bottom = min(top, rows[0]), max(bottom, rows[-1]) - left, right = min(left, columns[0]), max(right, columns[-1]) - if not np.isfinite(top): - raise ValueError("every matte is empty; nothing to score") - return ( - int(max(top - pad, 0)), - int(max(left - pad, 0)), - int(min(bottom + pad + 1, height)), - int(min(right + pad + 1, width)), - ) - - -def crop(image: np.ndarray, box: Tuple[int, int, int, int]) -> np.ndarray: - top, left, bottom, right = box - return image[top:bottom, left:right] - - -class QualityScorer: - """Scores renders against references. Holds the LPIPS network.""" - - def __init__(self, net: str = "alex", device: str = "cuda"): - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from lib import utils - - self._torch = torch - self._utils = utils - self._device = device - self._net = net - - def score(self, prediction: np.ndarray, reference: np.ndarray, - box: Optional[Sequence[int]] = None) -> Quality: - if prediction.shape != reference.shape: - raise ValueError( - f"shape mismatch: rendered {prediction.shape} vs reference {reference.shape}" - ) - if box is not None: - prediction = crop(prediction, tuple(box)) - reference = crop(reference, tuple(box)) - prediction = np.clip(prediction.astype(np.float64), 0.0, 1.0) - reference = np.clip(reference.astype(np.float64), 0.0, 1.0) - - error = float(np.mean((prediction - reference) ** 2)) - psnr = float("inf") if error <= 0.0 else float(-10.0 * np.log10(error)) - # ReRF's own SSIM and LPIPS, so these are the numbers upstream reports. - ssim = float(self._utils.rgb_ssim(prediction, reference, max_val=1.0)) - lpips = float( - self._utils.rgb_lpips( - prediction.astype(np.float32), - reference.astype(np.float32), - net_name=self._net, - device=self._device, - ) - ) - return Quality(psnr=psnr, ssim=ssim, lpips=lpips) diff --git a/open4d/reconstruction/nevo/nevo/render.py b/open4d/reconstruction/nevo/nevo/render.py deleted file mode 100644 index a03829ae..00000000 --- a/open4d/reconstruction/nevo/nevo/render.py +++ /dev/null @@ -1,179 +0,0 @@ -"""Render a loaded frame, to check that step 1 loaded it correctly. - -:func:`check_against_rerf` in ``nevo.importance`` proves the marching -transcription matches the vendored model -- but it compares a model against -*itself*, so it would pass just as happily on a model reassembled wrongly from -its checkpoint. Reconstructing a ``DirectVoxGO`` from ReRF's files takes half a -dozen fixups (the shared colour MLP arrives in a separate file, the -residual/deform switches come from the config rather than the checkpoint, and -a P-frame's feature grid is a residual that only means anything once -``former_k0_cur`` is wired up), and getting any of them wrong yields a model -that still renders, just not the trained scene. - -So: render a training viewpoint and compare against the image ReRF was trained -on. A correctly reassembled frame lands within a dB or so of the PSNR the -trainer logged; a mis-wired one is tens of dB off. -""" -from __future__ import annotations - -import json -from pathlib import Path -from typing import Optional, Tuple - -import numpy as np - -from . import rerf_env -from .cameras import Camera - - -def render_view(sequence, frame, camera: Camera, chunk: int = 1 << 19) -> np.ndarray: - """Render one viewport as an ``[H, W, 3]`` float array in [0, 1].""" - rerf_env.activate() - with rerf_env.rerf_cwd(): - import torch - from lib import dvgo - - render_kwargs = sequence.render_kwargs() - with torch.no_grad(): - c2w = torch.tensor(camera.c2w, dtype=torch.float32, device="cuda") - intrinsics = torch.tensor( - camera.intrinsic_matrix, dtype=torch.float32, device="cuda" - ) - rays_o, rays_d, viewdirs = dvgo.get_rays_of_a_view( - H=camera.height, - W=camera.width, - K=intrinsics, - c2w=c2w, - ndc=False, - inverse_y=render_kwargs["inverse_y"], - flip_x=render_kwargs["flip_x"], - flip_y=render_kwargs["flip_y"], - ) - rays_o = rays_o.flatten(0, -2) - rays_d = rays_d.flatten(0, -2) - viewdirs = viewdirs.flatten(0, -2) - pieces = [ - frame.model( - rays_o[begin : begin + chunk].contiguous(), - rays_d[begin : begin + chunk].contiguous(), - viewdirs[begin : begin + chunk].contiguous(), - **render_kwargs, - )["rgb_marched"] - for begin in range(0, len(rays_o), chunk) - ] - image = torch.cat(pieces).reshape(camera.height, camera.width, 3) - return image.clamp(0.0, 1.0).cpu().numpy() - - -def training_view(sequence, frame_index: int, view: int = 0) -> Tuple[Camera, np.ndarray]: - """A camera and its ground-truth image, straight from the corpus. - - The image is composited onto white, matching ``white_bkgd=True`` in the - config: ``lib.load_data`` does ``rgb * alpha + (1 - alpha)`` before the - trainer ever sees it. - """ - with open(sequence.corpus_dir / ("cams_%d.json" % frame_index)) as handle: - frames = sorted(json.load(handle)["frames"], key=lambda d: d["file"]) - if view >= len(frames): - raise IndexError( - f"cams_{frame_index}.json lists {len(frames)} training cameras; {view} is not one " - "of them (a held-out camera is absent by design -- use held_out_view)" - ) - entry = frames[view] - truth = _composite(entry["file"], entry["mask"]) - - intrinsic = np.asarray(entry["intrinsic"], dtype=np.float64) - height, width = truth.shape[:2] - camera = Camera( - camera_id=view, - width=width, - height=height, - fx=float(intrinsic[0, 0]), - fy=float(intrinsic[1, 1]), - cx=float(intrinsic[0, 2]), - cy=float(intrinsic[1, 2]), - c2w=np.asarray(entry["extrinsic"], dtype=np.float64), - ) - return camera, truth - - -def _composite(image_path, mask_path) -> np.ndarray: - """Load a corpus view over white, the way ``lib.load_data`` does. - - ``rgb * alpha + (1 - alpha)`` with ``white_bkgd=True`` -- the trainer never - sees the black-backed PNG, so neither should anything scored against it. - """ - from PIL import Image - - rgb = np.asarray(Image.open(image_path).convert("RGB"), dtype=np.float32) / 255.0 - alpha = ( - np.asarray(Image.open(mask_path).convert("L"), dtype=np.float32)[..., None] / 255.0 - ) - return rgb * alpha + (1.0 - alpha) - - -def held_out_view(sequence, frame_index: int, view: Optional[int] = None): - """The camera kept out of training, and its captured image. - - Deliberately not read from ``cams_.json``: that file *is* the - training set (ReRF's NHR loader trains on every entry), so the held-out - camera is absent from it by construction. Its calibration lives in the - corpus manifest and its pixels are on disk beside the training views. - """ - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - if view is None: - view = manifest.get("holdout_view") - if view is None: - raise ValueError( - f"{sequence.corpus_dir} has no held-out camera; it was prepared without " - "--holdout-view, so every view was trained on" - ) - entry = next( - (item for item in manifest["cameras"] if int(item["camera_id"]) == int(view)), None - ) - if entry is None: - raise ValueError(f"camera {view} is not in the corpus manifest") - - directory = sequence.corpus_dir - truth = _composite( - directory / "image" / str(frame_index) / ("img_%04d.png" % view), - directory / "mask" / str(frame_index) / ("img_%04d.png" % view), - ) - height, width = truth.shape[:2] - camera = Camera( - camera_id=int(view), - width=width, - height=height, - fx=float(entry["fx"]), - fy=float(entry["fy"]), - cx=float(entry["cx"]), - cy=float(entry["cy"]), - c2w=np.asarray(entry["c2w_normalised"], dtype=np.float64), - ) - return camera, truth - - -def psnr(prediction: np.ndarray, truth: np.ndarray) -> float: - error = float(np.mean((prediction.astype(np.float64) - truth.astype(np.float64)) ** 2)) - if error <= 0.0: - return float("inf") - return float(-10.0 * np.log10(error)) - - -def check_reload(sequence, frame, view: int = 0, save_to: Optional[Path] = None) -> dict: - """Render a training view of ``frame`` and score it against the ground truth.""" - camera, truth = training_view(sequence, frame.index, view) - rendered = render_view(sequence, frame, camera) - score = psnr(rendered, truth) - if save_to is not None: - from PIL import Image - - side_by_side = np.concatenate([rendered, truth], axis=1) - Image.fromarray((side_by_side * 255.0).astype(np.uint8)).save(save_to) - return { - "frame": frame.index, - "view": view, - "psnr": score, - "size": [camera.width, camera.height], - } diff --git a/open4d/reconstruction/nevo/nevo/rerf_env.py b/open4d/reconstruction/nevo/nevo/rerf_env.py deleted file mode 100644 index a23ba95a..00000000 --- a/open4d/reconstruction/nevo/nevo/rerf_env.py +++ /dev/null @@ -1,159 +0,0 @@ -"""Import ReRF from anywhere, without upstream's wrapper script. - -Upstream runs everything as ``python run.py`` from the ReRF root behind:: - - LD_LIBRARY_PATH=./ac_dc:$LD_LIBRARY_PATH PYTHONPATH=./ac_dc/:$PYTHONPATH - -which is not just convention. Importing ``lib.dvgo`` transitively imports -``codec``, and ``codec.encoder_jpeg`` - -* imports the prebuilt ``ncvv_ac_dc`` extension, whose ``NEEDED`` - ``libcode_library.so`` is only findable via ``LD_LIBRARY_PATH`` -- its - recorded ``RUNPATH`` is ``/home/ubuntu/pybind11_numpy/build``, a path on the - machine that built it; and -* evaluates ``gen_3d_quant_tbl()`` at import time, which does - ``np.load("./codec/quant.npy")`` -- a *relative* path, so the process CWD - has to be the ReRF root at that moment. - -``LD_LIBRARY_PATH`` is read by the dynamic loader at exec and cannot be set -from inside a running process, so the obvious fix is to re-exec. Don't: under -pytest that silently restarts the whole session. Instead :func:`activate` -``dlopen``s the two libraries by absolute path with ``RTLD_GLOBAL``, which -puts their symbols in the global namespace so the extension's by-name lookup -resolves. Same effect, no re-exec, and it works when NeVo is imported as a -library rather than run as a script. - -``ncvv_ac_dc`` ships only as a CPython 3.8 binary with no sources, which is -why this baseline lives in its own ``nevo`` conda environment rather than the -repo's 3.10+ one. See ``baselines/NeVo/README.md``. -""" -from __future__ import annotations - -import contextlib -import ctypes -import os -import sys -from pathlib import Path - -RERF_ROOT = Path(__file__).resolve().parents[1] / "rerf" - -AC_DC_LIBRARIES = ("libcode_library.so", "libjfif_library.so") - -DEFAULT_CUDA_HOME = "/usr/local/cuda-12.4" -DEFAULT_HOST_COMPILER = "/usr/bin/gcc-11" -"""ReRF JIT-compiles DVGO's CUDA kernels through torch.utils.cpp_extension on -first import. Ubuntu 24.04's default GCC 13 is newer than CUDA 12.4's nvcc -accepts, so point it at the 11 toolchain this box also has -- the same -workaround DeltaStream's renderer applies.""" - -_activated = False -_preloaded = [] - - -def _prepend(name: str, value: str) -> None: - parts = [part for part in os.environ.get(name, "").split(os.pathsep) if part] - if value not in parts: - os.environ[name] = os.pathsep.join([value] + parts) - - -def patch_dependencies() -> None: - """Reconcile upstream's code with the dependency versions this env installs. - - Two upstream calls fail outright, and both are version drift rather than - logic, so they are patched at runtime here instead of by editing ``rerf/`` - -- which stays byte-identical to upstream (see ``rerf/PATCHES.md``). - - ``np.bool`` and friends - Removed in numpy 1.24. ``codec.compress_utils.decode_pca`` uses - ``np.bool``, and *every* decode goes through it. Upstream's - ``compress.py`` decodes each frame to build the next frame's reference, - so without this it raises after writing frame 0 and the bitstream is - silently truncated to one frame. - - ``imageio.imwrite`` on a ``(H, W, 1)`` array - Newer imageio/Pillow raise "Can't write images with one color channel". - ``rerf_render.py`` writes its depth maps that way, so it dies after the - first frame. Squeezing the trailing axis is what older imageio did - implicitly. - - Idempotent, and called by :func:`activate`. - """ - import numpy - - for name, builtin in (("bool", bool), ("object", object), ("int", int), - ("float", float), ("complex", complex), ("str", str)): - if not hasattr(numpy, name): - setattr(numpy, name, builtin) - - try: - import imageio - except ImportError: # only rerf_render.py needs it - return - if getattr(imageio.imwrite, "_nevo_squeezes_gray", False): - return - original = imageio.imwrite - - def imwrite(uri, image, **kwargs): - array = numpy.asarray(image) - if array.ndim == 3 and array.shape[-1] == 1: - array = array[..., 0] - return original(uri, array, **kwargs) - - imwrite._nevo_squeezes_gray = True - imageio.imwrite = imwrite - - -def activate(*, cuda_home: str = None, arch_list: str = "8.9") -> Path: - """Make ``import lib.dvgo`` work in this process. Idempotent.""" - global _activated - root = RERF_ROOT - if not (root / "run.py").is_file(): - raise RuntimeError(f"vendored ReRF is missing from {root}") - if _activated: - return root - patch_dependencies() - - for name in AC_DC_LIBRARIES: - path = root / "ac_dc" / name - if not path.is_file(): - raise RuntimeError(f"ReRF's entropy coder is incomplete: {path} is missing") - _preloaded.append(ctypes.CDLL(str(path), mode=ctypes.RTLD_GLOBAL)) - - home = cuda_home or os.environ.get("CUDA_HOME") or DEFAULT_CUDA_HOME - os.environ["CUDA_HOME"] = home - _prepend("PATH", str(Path(home) / "bin")) - os.environ.setdefault("TORCH_CUDA_ARCH_LIST", arch_list) - if Path(DEFAULT_HOST_COMPILER).exists(): - os.environ.setdefault("CC", DEFAULT_HOST_COMPILER) - os.environ.setdefault("CXX", DEFAULT_HOST_COMPILER.replace("gcc", "g++")) - - for path in (str(root), str(root / "ac_dc")): - if path not in sys.path: - sys.path.insert(0, path) - _activated = True - return root - - -@contextlib.contextmanager -def rerf_cwd(): - """Run a block with the CWD at the ReRF root. - - Needed around the *first* ``import lib.*``, because ``codec.quant`` reads - ``./codec/quant.npy`` at import time, and around any later call into - ``codec`` or ``run`` that resolves a relative path of its own. - """ - previous = os.getcwd() - os.chdir(activate()) - try: - yield Path(previous) - finally: - os.chdir(previous) - - -def import_rerf(): - """Import and return ReRF's ``(dvgo, dvgo_video, utils)`` modules.""" - activate() - with rerf_cwd(): - from lib import dvgo, dvgo_video, utils - - return dvgo, dvgo_video, utils diff --git a/open4d/reconstruction/nevo/nevo/sequence.py b/open4d/reconstruction/nevo/nevo/sequence.py deleted file mode 100644 index b433f4bb..00000000 --- a/open4d/reconstruction/nevo/nevo/sequence.py +++ /dev/null @@ -1,213 +0,0 @@ -"""Step 1: load a trained ReRF sequence's feature voxel grids and motion vectors. - -What ReRF leaves on disk per frame, under ``//``: - -``fine_last_0.tar`` - The I-frame. ``model_state_dict['density']`` is the raw density grid - ``[1, 1, X, Y, Z]`` and ``['k0.k0']`` the 12-channel feature grid. -``fine_last__deform.tar`` - A P-frame's **motion vectors**: ``deformation_field``, a ``[1, 3, X, Y, Z]`` - grid of per-entry displacements that warps frame ``n-1`` towards frame - ``n``. This is the cheap part of a P-frame -- upstream quantises it to fp16 - and drops the all-zero entries (``codec.encoder_motion``). -``fine_last_.tar`` - The P-frame proper. ``['k0.k0']`` is a *residual* over ``['k0.former_k0']`` - (the motion-compensated previous grid), so the absolute feature grid this - frame renders from is the sum of the two. ``TensorDVGORes.compute_features`` - does exactly that at sample time. - -:class:`ReRFFrame` hands back the absolute grids, because that is what the -importance pass (step 2) and the renderer need; the residual/motion split is -kept alongside because that is what the *network* carries and what steps 3-6 -have to packetize, drop and reconstruct. - -Reconstructing a usable model from these checkpoints means repeating the -fixups upstream applies in ``run.py``'s render path -- the saved -``model_kwargs`` does not round-trip on its own (the shared colour MLP lives -in a separate file, and ``use_res``/``use_deform`` come from the config rather -than the checkpoint). :meth:`ReRFSequence.frame` is that sequence, in one -place. -""" -from __future__ import annotations - -import json -from dataclasses import dataclass -from pathlib import Path -from typing import List, Optional, Tuple - -import numpy as np - -from . import rerf_env - - -@dataclass -class ReRFFrame: - """One decoded frame of a ReRF sequence.""" - - index: int - model: object - density: "object" - features: "object" - residual: Optional["object"] - motion: Optional["object"] - xyz_min: "object" - xyz_max: "object" - is_key_frame: bool - - @property - def grid_shape(self) -> Tuple[int, int, int]: - return tuple(int(n) for n in self.density.shape[2:]) - - @property - def feature_dim(self) -> int: - return int(self.features.shape[1]) - - def raw_bytes(self) -> int: - """Uncompressed size of the density + feature grids, fp32. - - The number the NeVo paper quotes as "~800 MB" for a frame of neural - content. - """ - entries = int(np.prod(self.grid_shape)) - return entries * (1 + self.feature_dim) * 4 - - -class ReRFSequence: - """A trained ReRF sequence on disk, indexed by frame.""" - - def __init__(self, config_path, device: str = "cuda"): - rerf_env.activate() - with rerf_env.rerf_cwd(): - import mmcv - import torch - from lib import dvgo, dvgo_video - - self._torch = torch - self._dvgo = dvgo - self._dvgo_video = dvgo_video - self.device = torch.device(device) - self.config_path = Path(config_path).expanduser().resolve() - self.cfg = mmcv.Config.fromfile(str(self.config_path)) - self.run_dir = Path(self.cfg.basedir).expanduser() / self.cfg.expname - self.corpus_dir = Path(self.cfg.data["datadir"]).expanduser() - if not self.run_dir.is_dir(): - raise FileNotFoundError(f"no trained run at {self.run_dir}") - - # One shared colour MLP for the whole sequence (`fix_rgbnet=True`), so - # build the container model once and reuse its rgbnet for every frame. - self._video_model = dvgo_video.DirectVoxGO_Video() - self._video_model.current_frame_id = 0 - self._video_model.load_rgb_net(self.cfg) - - # ------------------------------------------------------------------ paths - def _fine_path(self, index: int) -> Path: - return self.run_dir / ("fine_last_%d.tar" % index) - - def _deform_path(self, index: int) -> Path: - return self.run_dir / ("fine_last_%d_deform.tar" % index) - - def available_frames(self) -> List[int]: - """Frame indices whose fine-stage checkpoint exists, in order.""" - found = [] - for index in range(int(self.cfg.frame_num)): - if self._fine_path(index).is_file(): - found.append(index) - return found - - # ------------------------------------------------------------- near / far - def near_far(self) -> Tuple[float, float]: - """Reproduce ``lib.load_data.inward_nearfar_heuristic`` for this corpus. - - Read off ``cams_0.json`` rather than by loading the corpus: the - heuristic only looks at camera positions, and decoding 48 views of - every frame to learn two scalars costs a minute and several GB. - """ - with open(self.corpus_dir / "cams_0.json") as handle: - frames = json.load(handle)["frames"] - # load_NHR sorts views by file path before stacking, so the positions - # here are the same set the trainer saw, whatever the json order. - positions = np.asarray( - [np.asarray(f["extrinsic"], dtype=np.float64)[:3, 3] - for f in sorted(frames, key=lambda d: d["file"])] - ) - distance = np.linalg.norm(positions[:, None] - positions, axis=-1) - far = float(distance.max() * 1.4) - return far * 0.05, far - - def render_kwargs(self) -> dict: - near, far = self.near_far() - return { - "near": near, - "far": far, - "bg": 1 if self.cfg.data["white_bkgd"] else 0, - "stepsize": self.cfg.fine_model_and_render["stepsize"], - "inverse_y": self.cfg.data["inverse_y"], - "flip_x": self.cfg.data["flip_x"], - "flip_y": self.cfg.data["flip_y"], - } - - # ----------------------------------------------------------------- frames - def frame_density(self, index: int): - """Just the raw density grid of a frame, no model construction. - - Block occupancy only needs density, and taking a union of it over a - sequence before scoring saves rebuilding every DirectVoxGO twice. - """ - path = self._fine_path(index) - if not path.is_file(): - raise FileNotFoundError(f"frame {index} is not trained: {path}") - checkpoint = self._torch.load(str(path), map_location=self.device) - return checkpoint["model_state_dict"]["density"].detach() - - def frame(self, index: int) -> ReRFFrame: - torch = self._torch - path = self._fine_path(index) - if not path.is_file(): - raise FileNotFoundError(f"frame {index} is not trained: {path}") - checkpoint = torch.load(str(path), map_location=self.device) - kwargs = dict(checkpoint["model_kwargs"]) - state = dict(checkpoint["model_state_dict"]) - - deform_path = self._deform_path(index) - is_key_frame = not deform_path.is_file() - - # Mirrors run.py's render path: the checkpoint carries neither the - # shared MLP nor the residual/deform switches, and `deform_res_mode == - # "separate"` means every frame's grid is a residual model even though - # `cfg.use_res` is left False for the trainer's own bookkeeping. - kwargs["rgbnet"] = self._video_model.rgbnet - kwargs["cfg"] = self.cfg - kwargs["use_res"] = bool(self.cfg.use_res) or self.cfg.deform_res_mode == "separate" - kwargs["use_deform"] = "" - kwargs["rgbfeat_sigmoid"] = self.cfg.codec["rgbfeat_sigmoid"] - - model = self._dvgo.DirectVoxGO(**kwargs) - if kwargs["use_res"] and "k0.former_k0" not in state: - # The I-frame has no predecessor to residual against, so its k0 *is* - # the absolute grid and former_k0 stays at the zeros the constructor - # made. - state["k0.former_k0"] = model.k0.former_k0 - model.load_state_dict(state, strict=False) - if kwargs["use_res"]: - model.k0.former_k0_cur = model.k0.former_k0 - model = model.to(self.device).eval() - - residual = state["k0.k0"].detach() - former = state["k0.former_k0"].detach() - features = residual + former - motion = None - if not is_key_frame: - deform = torch.load(str(deform_path), map_location=self.device) - motion = deform["model_state_dict"]["deformation_field"].detach() - - return ReRFFrame( - index=index, - model=model, - density=state["density"].detach(), - features=features, - residual=None if is_key_frame else residual, - motion=motion, - xyz_min=model.xyz_min.detach(), - xyz_max=model.xyz_max.detach(), - is_key_frame=is_key_frame, - ) diff --git a/open4d/reconstruction/nevo/nevo/viewports.py b/open4d/reconstruction/nevo/nevo/viewports.py deleted file mode 100644 index 780b315c..00000000 --- a/open4d/reconstruction/nevo/nevo/viewports.py +++ /dev/null @@ -1,292 +0,0 @@ -"""Viewports to score neural visibility from. - -Two sources, both producing :class:`nevo.cameras.Camera` in the corpus's -normalised frame: - -:func:`sample_viewports` - Synthetic viewports drawn around the content. The NeVo paper's importance - CDF is measured "with each frame tested against 300 different viewports" - (section 3.2), and no recorded trajectory exists for the ORBIT objects, so - this is what the CDF is built on. - -:func:`read_quest_trace` / :func:`trace_viewports` - A real 6DoF trace in this repo's Quest logger format - (``system/QuestClient/.../ViewportTraceLogger.cs``), which is what the - end-to-end simulation replays. Positions are metres in the composed scene - frame and rotations are yaw/pitch/roll degrees; both the measured pose and - the pose the on-device predictor guessed a window earlier are recorded, so - a trace supplies the *predicted* viewport the edge would have fetched for - and the *actual* viewport the client reprojects to. -""" -from __future__ import annotations - -import csv -import math -from dataclasses import dataclass -from pathlib import Path -from typing import Iterable, List, Optional, Sequence - -import numpy as np - -from .cameras import Camera, look_at_c2w - -TRACE_FEATURES = ("x", "y", "z", "yaw", "pitch", "roll") - - -@dataclass(frozen=True) -class ViewportSpread: - """Where a viewer is assumed to be, relative to the content. - - Defaults describe someone walking around a body-scale subject on a headset: - a little inside to a little outside the capture rig, eye height ranging - from below the subject's chest to looking down on it, and always roughly - facing it. - """ - - radius_scale: tuple = (0.75, 1.45) - elevation_degrees: tuple = (-25.0, 55.0) - #: Fraction of the bbox half-extent the look-at point may wander by, so - #: the rays are not all funnelled through one point. - aim_jitter: float = 0.25 - - -def sample_viewports( - count: int, - xyz_min: Sequence[float], - xyz_max: Sequence[float], - *, - reference_radius: float, - width: int, - height: int, - focal: float, - spread: ViewportSpread = ViewportSpread(), - up_axis: int = 1, - seed: int = 0, -) -> List[Camera]: - """Draw ``count`` viewports around the content bounding box.""" - if count < 1: - raise ValueError("count must be positive") - lower = np.asarray(xyz_min, dtype=np.float64) - upper = np.asarray(xyz_max, dtype=np.float64) - centre = (lower + upper) * 0.5 - half_extent = (upper - lower) * 0.5 - rng = np.random.default_rng(seed) - - up = np.zeros(3) - up[up_axis] = 1.0 - plane = [axis for axis in range(3) if axis != up_axis] - - cameras: List[Camera] = [] - for index in range(count): - radius = reference_radius * rng.uniform(*spread.radius_scale) - azimuth = rng.uniform(0.0, 2.0 * math.pi) - elevation = math.radians(rng.uniform(*spread.elevation_degrees)) - direction = np.zeros(3) - direction[plane[0]] = math.cos(azimuth) * math.cos(elevation) - direction[plane[1]] = math.sin(azimuth) * math.cos(elevation) - direction[up_axis] = math.sin(elevation) - aim = centre + rng.uniform(-1.0, 1.0, size=3) * half_extent * spread.aim_jitter - eye = centre + direction * radius - cameras.append( - Camera( - camera_id=index, - width=width, - height=height, - fx=focal, - fy=focal, - cx=(width - 1) * 0.5, - cy=(height - 1) * 0.5, - c2w=look_at_c2w(eye, aim, up), - ) - ) - return cameras - - -def downscale(cameras: Iterable[Camera], factor: int) -> List[Camera]: - """Shrink every camera by an integer factor, intrinsics included. - - Importance is a max over the samples that land in a voxel, so it is far - less ray-density-sensitive than a rendered image -- halving the resolution - cuts the marching cost 4x. It is not free, though: thin structures can stop - being hit at all, which biases the CDF towards "unimportant". Verify before - relying on it. - """ - if factor < 1: - raise ValueError("factor must be >= 1") - if factor == 1: - return list(cameras) - scaled = [] - for camera in cameras: - width = max(1, camera.width // factor) - height = max(1, camera.height // factor) - scaled.append( - Camera( - camera_id=camera.camera_id, - width=width, - height=height, - fx=camera.fx / factor, - fy=camera.fy / factor, - cx=(camera.cx + 0.5) / factor - 0.5, - cy=(camera.cy + 0.5) / factor - 0.5, - c2w=camera.c2w, - ) - ) - return scaled - - -# --------------------------------------------------------------- Quest traces -@dataclass -class ViewportTrace: - """A recorded 6DoF trajectory plus the predictions made during it.""" - - times: np.ndarray # [N] seconds since the run started - actual: np.ndarray # [N, 6] x y z yaw pitch roll - predicted: np.ndarray # [N, 6], NaN where no prediction resolved - predicted_target: np.ndarray # [N] seconds the prediction aimed at, NaN if none - - def __len__(self) -> int: - return int(self.times.shape[0]) - - @property - def prediction_error(self) -> np.ndarray: - """Translational error in metres, NaN where unresolved.""" - return np.linalg.norm(self.actual[:, :3] - self.predicted[:, :3], axis=-1) - - -def read_quest_trace(path, streaming_only: bool = True) -> ViewportTrace: - """Read this repo's Quest viewport CSV. - - Columns come from ``ViewportTraceLogger.CsvHeader``. ``streaming_only`` - keeps the rows recorded while content was actually being streamed, which is - the same filter ``system/QuestClient/Tools/plot_viewport_traces.py`` applies. - """ - times: List[float] = [] - actual: List[List[float]] = [] - predicted: List[List[float]] = [] - targets: List[float] = [] - with open(Path(path), newline="", encoding="utf-8") as handle: - for row in csv.DictReader(handle): - if streaming_only and row.get("streaming") != "1": - continue - times.append(float(row["t_s"])) - actual.append([float(row[name]) for name in TRACE_FEATURES]) - if row.get("pred_valid") == "1": - predicted.append([float(row["pred_" + name]) for name in TRACE_FEATURES]) - targets.append(float(row["pred_target_s"])) - else: - predicted.append([math.nan] * len(TRACE_FEATURES)) - targets.append(math.nan) - if not times: - raise ValueError(f"{path} holds no usable rows") - return ViewportTrace( - np.asarray(times, dtype=np.float64), - np.asarray(actual, dtype=np.float64), - np.asarray(predicted, dtype=np.float64), - np.asarray(targets, dtype=np.float64), - ) - - -CONVENTIONS = ("unity", "right_handed") - - -def pose_to_c2w(pose: Sequence[float], convention: str = "unity") -> np.ndarray: - """OpenCV camera-to-world for one ``(x, y, z, yaw, pitch, roll)`` trace row. - - Rotations are degrees, applied yaw (Y) then pitch (X) then roll (Z), which - is Unity's ``Quaternion.Euler`` order and what ``ViewportPredictor`` - interpolates in. - - ``convention`` says what frame the *trace* is in: - - ``unity`` - Left-handed, Y up, +Z forward -- the frame a Unity ``Transform`` - reports, and therefore what ``ViewportTraceLogger`` writes. Mapping it - into the right-handed frame this codebase renders in flips Z in the - world (``W``) and, separately, flips Y in the camera's own axes to get - OpenCV's y-down (``L``): ``c2w = W . R . L``. Both flips are needed; - applying either alone leaves a mirrored (determinant -1) rotation that - renders a plausible but laterally-inverted view. - ``right_handed`` - Y up, -Z forward (glTF / the frame ``scene_layout.json`` describes for - the composed ORBIT scene). No world flip; the camera's own axes still - move, since OpenCV wants y down and z forward. - - Which one a given trace needs depends on where it was recorded, and no - recorded trace exists in this repo to settle it -- so it is an explicit - argument rather than a silent default baked into the pipeline. It is also - only half the job: putting a headset pose into a *corpus* frame needs the - object's placement in the scene too (``scene_layout.json``), which arrives - with the end-to-end simulation. - """ - if convention not in CONVENTIONS: - raise ValueError(f"convention must be one of {CONVENTIONS}") - x, y, z, yaw, pitch, roll = (float(v) for v in pose) - cy, sy = math.cos(math.radians(yaw)), math.sin(math.radians(yaw)) - cp, sp = math.cos(math.radians(pitch)), math.sin(math.radians(pitch)) - cr, sr = math.cos(math.radians(roll)), math.sin(math.radians(roll)) - rotation_y = np.asarray(((cy, 0.0, sy), (0.0, 1.0, 0.0), (-sy, 0.0, cy))) - rotation_x = np.asarray(((1.0, 0.0, 0.0), (0.0, cp, -sp), (0.0, sp, cp))) - rotation_z = np.asarray(((cr, -sr, 0.0), (sr, cr, 0.0), (0.0, 0.0, 1.0))) - rotation = rotation_y @ rotation_x @ rotation_z - - position = np.asarray((x, y, z)) - if convention == "unity": - # OpenCV local axes expressed in Unity local axes: right stays, down is - # -up, forward stays (+Z). Then flip the world's Z to make it - # right-handed. det = (-1) * (+1) * (-1) = +1. - rotation = np.diag((1.0, 1.0, -1.0)) @ rotation @ np.diag((1.0, -1.0, 1.0)) - position = position * np.asarray((1.0, 1.0, -1.0)) - else: - # Already right-handed with -Z forward, so only the camera's own axes - # move: down is -up and forward is -backward. det = +1, no world flip. - rotation = rotation @ np.diag((1.0, -1.0, -1.0)) - - c2w = np.eye(4) - c2w[:3, :3] = rotation - c2w[:3, 3] = position - return c2w - - -def trace_viewports( - trace: ViewportTrace, - *, - width: int, - height: int, - focal: float, - centre: Sequence[float], - scale: float, - use_predicted: bool = False, - convention: str = "unity", - indices: Optional[Sequence[int]] = None, -) -> List[Camera]: - """Turn trace rows into cameras in the corpus's normalised frame. - - ``centre``/``scale`` are ``nevo_corpus.json``'s ``world_centre`` and - ``world_scale``: the same world -> normalised map the training extrinsics - went through. - """ - poses = trace.predicted if use_predicted else trace.actual - rows = range(len(trace)) if indices is None else indices - origin = np.asarray(centre, dtype=np.float64) - cameras: List[Camera] = [] - for index in rows: - pose = poses[index] - if not np.all(np.isfinite(pose)): - continue - c2w = pose_to_c2w(pose, convention) - c2w[:3, 3] = (c2w[:3, 3] - origin) * scale - cameras.append( - Camera( - camera_id=int(index), - width=width, - height=height, - fx=focal, - fy=focal, - cx=(width - 1) * 0.5, - cy=(height - 1) * 0.5, - c2w=c2w, - ) - ) - if not cameras: - raise ValueError("no rows in the trace carried a finite pose") - return cameras diff --git a/open4d/reconstruction/nevo/nevo_tests/conftest.py b/open4d/reconstruction/nevo/nevo_tests/conftest.py deleted file mode 100644 index 8e513b1c..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/conftest.py +++ /dev/null @@ -1,80 +0,0 @@ -"""Shared gating for the tests that need a real ReRF model. - -Most of this suite is pure tensor/geometry logic and always runs. A handful of -cases load a trained sequence and render it, which needs both a checkpoint on -disk and a GPU with room to work in. - -"A GPU is present" is not the right condition. This box trains ReRF on both -cards for hours at a time, and a test that starts while 20 GB of the 24 is -already committed does not fail meaningfully -- it raises ``CUDA error: out of -memory`` somewhere inside DVGO's kernels and reports a red suite that says -nothing about the code. So the gate is *free* memory, and falling short of it -skips. -""" -from __future__ import annotations - -import os -from pathlib import Path - -import pytest - -REQUIRED_FREE_BYTES = 3 << 30 -"""Headroom a trained frame needs: its density and feature grids, the -occupancy pass, and a viewport's worth of ray samples.""" - -DEFAULT_CONFIG = Path(__file__).resolve().parents[1] / "rerf" / "configs" / "nevo" - - -def _first_trained_config() -> Path | None: - """The config named by ``NEVO_TEST_CONFIG``, else any run with a frame.""" - override = os.environ.get("NEVO_TEST_CONFIG") - if override: - candidate = Path(override) - return candidate if candidate.is_file() else None - if not DEFAULT_CONFIG.is_dir(): - return None - for candidate in sorted(DEFAULT_CONFIG.glob("*.py")): - return candidate - return None - - -def gpu_headroom() -> int: - try: - import torch - except ImportError: - return 0 - if not torch.cuda.is_available(): - return 0 - try: - free, _total = torch.cuda.mem_get_info() - except Exception: - return 0 - return int(free) - - -def _decide() -> "tuple[Path | None, str]": - free = gpu_headroom() - if free < REQUIRED_FREE_BYTES: - return None, ( - f"needs {REQUIRED_FREE_BYTES >> 30} GB of free GPU memory, " - f"{free / (1 << 30):.1f} GB available (training probably has the cards)" - ) - config = _first_trained_config() - if config is None: - return None, "needs a trained ReRF sequence; point NEVO_TEST_CONFIG at one" - return config, "" - - -# Decided once, at collection. Deciding per call would let the skip marker and -# the test body disagree -- the marker is evaluated at import, and free GPU -# memory moves while a suite runs, so a re-check inside the test can come back -# None on a test the marker already let through. -_CONFIG, _REASON = _decide() - - -def trained_config() -> "Path | None": - """The config the model tests should use, or None if they must skip.""" - return _CONFIG - - -needs_sequence = pytest.mark.skipif(_CONFIG is None, reason=_REASON) diff --git a/open4d/reconstruction/nevo/nevo_tests/test_blocks.py b/open4d/reconstruction/nevo/nevo_tests/test_blocks.py deleted file mode 100644 index 89fcbb0a..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_blocks.py +++ /dev/null @@ -1,161 +0,0 @@ -"""Block indexing has to agree with ReRF's, bit for bit. - -A block index in this codebase is meant to be the same integer as the bit -position in ReRF's ``mask_.rerf`` and the same slot in its bitstream. -Nothing enforces that at runtime, so it is pinned here against a transcription -of upstream's ``codec.utils.split_volume`` and ``compress_utils.get_masks``. -""" -from __future__ import annotations - -import numpy as np -import pytest - -torch = pytest.importorskip("torch") - -from nevo.blocks import ( # noqa: E402 - BlockGrid, - nearest_entry, - surrounding_entries, -) - - -def reference_split(data: torch.Tensor, voxel_size: int): - """``codec.utils.zero_pads`` + ``split_volume``, transcribed. - - Deliberately the slow triple loop upstream uses, so the ordering under - test is compared against the ordering that actually ships rather than - against another vectorised guess at it. - """ - if data.size(0) == 1: - data = data.squeeze(0) - size = list(data.size()) - padded_size = list(size) - for axis in range(1, 4): - if padded_size[axis] % voxel_size: - padded_size[axis] = (padded_size[axis] // voxel_size + 1) * voxel_size - padded = torch.zeros(padded_size, dtype=data.dtype) - padded[:, : size[1], : size[2], : size[3]] = data - blocks = [] - for x in range(padded_size[1] // voxel_size): - for y in range(padded_size[2] // voxel_size): - for z in range(padded_size[3] // voxel_size): - blocks.append( - padded[ - :, - x * voxel_size : (x + 1) * voxel_size, - y * voxel_size : (y + 1) * voxel_size, - z * voxel_size : (z + 1) * voxel_size, - ] - ) - return torch.stack(blocks) - - -@pytest.mark.parametrize("shape", [(9, 17, 5), (16, 16, 16), (126, 26, 13)]) -@pytest.mark.parametrize("block_size", [1, 2, 8]) -def test_block_index_matches_split_volume_ordering(shape, block_size): - """Give every entry a unique value, then find where each one landed.""" - volume = torch.arange(int(np.prod(shape)), dtype=torch.float64).reshape(1, 1, *shape) - blocks = reference_split(volume, block_size) - - grid = BlockGrid(shape, block_size) - assert grid.num_blocks == blocks.shape[0] - - entries = torch.stack( - torch.meshgrid( - *[torch.arange(n) for n in shape], - indexing="ij", - ), - dim=-1, - ).reshape(-1, 3) - predicted = grid.block_index(entries) - - for entry, block in zip(entries.tolist(), predicted.tolist()): - value = float(volume[0, 0, entry[0], entry[1], entry[2]]) - assert value in set(blocks[block].reshape(-1).tolist()), (entry, block) - - -@pytest.mark.parametrize("shape,block_size", [((9, 17, 5), 8), ((32, 32, 32), 8), ((7, 7, 7), 4)]) -def test_occupancy_matches_the_codec(shape, block_size): - torch.manual_seed(0) - density = torch.randn(1, 1, *shape) * 4.0 - grid = BlockGrid(shape, block_size) - ours = grid.occupancy(density).cpu().numpy() - - # codec.compress_utils.get_masks, on the single-channel case. - blocks = reference_split(density, block_size) - occupied = torch.nn.functional.softplus(blocks - 4.1) > 0.4 - theirs = occupied.reshape(blocks.shape[0], -1).any(dim=-1).numpy() - - assert np.array_equal(ours, theirs) - - -def test_occupancy_ignores_padding(): - """Padding is zeros; softplus(0 - 4.1) is far below the threshold, so a - partly-padded block must be occupied only if a real entry says so.""" - shape = (9, 9, 9) - density = torch.full((1, 1, *shape), -10.0) - density[0, 0, 8, 8, 8] = 100.0 - grid = BlockGrid(shape, 8) - occupancy = grid.occupancy(density) - assert int(occupancy.sum()) == 1 - corner = grid.block_index(torch.tensor([[8, 8, 8]])) - assert bool(occupancy[corner.item()]) - - -def test_padded_and_block_shapes(): - grid = BlockGrid((126, 256, 125), 8) - assert grid.padded_shape == (128, 256, 128) - assert grid.blocks_shape == (16, 32, 16) - assert grid.num_blocks == 16 * 32 * 16 - - -def test_nearest_entry_hits_the_grid_corners(): - """align_corners=True: the first and last entries sit exactly on the bbox.""" - shape = (5, 9, 3) - xyz_min = torch.tensor([-1.0, -2.0, 0.0]) - xyz_max = torch.tensor([1.0, 2.0, 4.0]) - points = torch.stack([xyz_min, xyz_max, (xyz_min + xyz_max) / 2]) - entries = nearest_entry(points, xyz_min, xyz_max, shape) - assert entries[0].tolist() == [0, 0, 0] - assert entries[1].tolist() == [4, 8, 2] - assert entries[2].tolist() == [2, 4, 1] - - -def test_nearest_entry_clamps_points_outside_the_box(): - shape = (4, 4, 4) - xyz_min = torch.zeros(3) - xyz_max = torch.ones(3) - points = torch.tensor([[-5.0, -5.0, -5.0], [5.0, 5.0, 5.0]]) - entries = nearest_entry(points, xyz_min, xyz_max, shape) - assert entries[0].tolist() == [0, 0, 0] - assert entries[1].tolist() == [3, 3, 3] - - -def test_surrounding_entries_bracket_the_sample(): - shape = (8, 8, 8) - xyz_min = torch.zeros(3) - xyz_max = torch.ones(3) - point = torch.tensor([[0.31, 0.52, 0.77]]) - corners = surrounding_entries(point, xyz_min, xyz_max, shape).squeeze(1) - assert corners.shape == (8, 3) - exact = point[0] * 7.0 - assert torch.all(corners.min(dim=0).values.float() <= exact) - assert torch.all(corners.max(dim=0).values.float() >= exact) - assert len({tuple(row) for row in corners.tolist()}) == 8 - - -def test_surrounding_entries_agree_with_nearest_at_a_grid_point(): - shape = (8, 8, 8) - xyz_min = torch.zeros(3) - xyz_max = torch.ones(3) - point = torch.tensor([[3.0 / 7.0, 3.0 / 7.0, 3.0 / 7.0]]) - nearest = nearest_entry(point, xyz_min, xyz_max, shape)[0] - corners = surrounding_entries(point, xyz_min, xyz_max, shape).squeeze(1) - assert nearest.tolist() in corners.tolist() - - -def test_rejects_a_non_volume(): - with pytest.raises(ValueError): - BlockGrid.from_volume(torch.zeros(4, 4, 4)) - with pytest.raises(ValueError): - BlockGrid((8, 8, 8), 0) diff --git a/open4d/reconstruction/nevo/nevo_tests/test_cameras.py b/open4d/reconstruction/nevo/nevo_tests/test_cameras.py deleted file mode 100644 index 2bffd821..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_cameras.py +++ /dev/null @@ -1,148 +0,0 @@ -"""Rig geometry: the part that silently ruins a NeRF if it is wrong. - -A camera convention error does not crash -- it trains a plausible-looking -model of the wrong scene -- so these check the invariants directly: rotations -stay orthonormal, the rig frames the whole bounding box, and normalisation -lands every camera on the radius ReRF's near/far heuristic assumes. -""" -from __future__ import annotations - -import math - -import numpy as np -import pytest - -from nevo.cameras import ( - NORMALISED_RADIUS, - Camera, - fit_radius, - look_at_c2w, - orbit_rig, -) - -BOUNDS_MIN = (-327.3, -480.0, -447.9) -BOUNDS_MAX = (331.5, 1395.0, 197.4) -"""A real ORBIT frame: basketball_player_fr0001, in native millimetres.""" - - -def test_look_at_is_a_rigid_transform(): - c2w = look_at_c2w( - np.asarray((3.0, 1.0, 0.0)), np.asarray((0.0, 1.0, 0.0)), np.asarray((0.0, 1.0, 0.0)) - ) - rotation = c2w[:3, :3] - assert np.allclose(rotation @ rotation.T, np.eye(3), atol=1e-12) - assert np.isclose(np.linalg.det(rotation), 1.0, atol=1e-12) - - -def test_look_at_uses_opencv_axes(): - """Column 2 points at the target, column 1 points down, column 0 right.""" - eye = np.asarray((0.0, 0.0, -5.0)) - target = np.zeros(3) - up = np.asarray((0.0, 1.0, 0.0)) - c2w = look_at_c2w(eye, target, up) - assert np.allclose(c2w[:3, 2], (0.0, 0.0, 1.0), atol=1e-12) - assert np.allclose(c2w[:3, 1], (0.0, -1.0, 0.0), atol=1e-12) - assert np.dot(np.cross(c2w[:3, 0], c2w[:3, 1]), c2w[:3, 2]) > 0 - - -def test_look_at_survives_a_straight_down_camera(): - """The naive right = forward x up is degenerate when they are parallel.""" - c2w = look_at_c2w( - np.asarray((0.0, 5.0, 0.0)), np.zeros(3), np.asarray((0.0, 1.0, 0.0)) - ) - assert np.all(np.isfinite(c2w)) - assert np.allclose(c2w[:3, :3] @ c2w[:3, :3].T, np.eye(3), atol=1e-9) - - -def _project(camera: Camera, point: np.ndarray) -> np.ndarray: - world_to_camera = np.linalg.inv(camera.c2w) - local = world_to_camera[:3, :3] @ point + world_to_camera[:3, 3] - assert local[2] > 0, "point fell behind the camera" - return camera.intrinsic_matrix @ (local / local[2]) - - -def test_rig_frames_the_whole_bounding_box(): - cameras, _, _ = orbit_rig(BOUNDS_MIN, BOUNDS_MAX, 1280, 960) - lower, upper = np.asarray(BOUNDS_MIN), np.asarray(BOUNDS_MAX) - corners = np.asarray( - [ - (lower[0] if x else upper[0], lower[1] if y else upper[1], lower[2] if z else upper[2]) - for x in (0, 1) - for y in (0, 1) - for z in (0, 1) - ] - ) - for camera in cameras: - for corner in corners: - pixel = _project(camera, corner) - assert -0.5 <= pixel[0] <= camera.width - 0.5, (camera.camera_id, pixel) - assert -0.5 <= pixel[1] <= camera.height - 0.5, (camera.camera_id, pixel) - - -def test_rig_actually_fills_the_frame(): - """Framing has to be tight enough to be worth 48 renders. - - A rig that fits the bbox by sitting a kilometre away also "frames" it. - """ - cameras, _, _ = orbit_rig(BOUNDS_MIN, BOUNDS_MAX, 1280, 960) - lower, upper = np.asarray(BOUNDS_MIN), np.asarray(BOUNDS_MAX) - centre = (lower + upper) * 0.5 - top = centre.copy() - top[1] = upper[1] - bottom = centre.copy() - bottom[1] = lower[1] - for camera in cameras: - height = abs(_project(camera, top)[1] - _project(camera, bottom)[1]) - assert height > 0.5 * camera.height, (camera.camera_id, height) - - -def test_normalisation_puts_every_camera_on_the_reference_radius(): - cameras, centre, scale = orbit_rig(BOUNDS_MIN, BOUNDS_MAX, 1280, 960) - radii = [ - np.linalg.norm(camera.scaled_translation(centre, scale)[:3, 3]) for camera in cameras - ] - assert np.allclose(radii, NORMALISED_RADIUS, atol=1e-9) - - -def test_normalisation_is_a_similarity_so_projection_is_unchanged(): - cameras, centre, scale = orbit_rig(BOUNDS_MIN, BOUNDS_MAX, 1280, 960) - probe = np.asarray((100.0, 300.0, -50.0)) - for camera in cameras[:6]: - world_pixel = _project(camera, probe) - normalised = Camera( - camera.camera_id, - camera.width, - camera.height, - camera.fx, - camera.fy, - camera.cx, - camera.cy, - camera.scaled_translation(centre, scale), - ) - normalised_pixel = _project(normalised, (probe - centre) * scale) - assert np.allclose(world_pixel, normalised_pixel, atol=1e-6) - - -def test_elevation_rows_are_offset_in_azimuth(): - """Stacked rows would see nearly the same silhouette three times over.""" - cameras, centre, _ = orbit_rig( - BOUNDS_MIN, BOUNDS_MAX, 1280, 960, azimuths=8, elevations=(0.0, 30.0) - ) - def azimuth(camera): - offset = camera.c2w[:3, 3] - centre - return math.atan2(offset[2], offset[0]) % (2 * math.pi) - - assert not np.isclose(azimuth(cameras[0]), azimuth(cameras[8]), atol=1e-6) - - -def test_fit_radius_grows_with_the_subject(): - small = fit_radius(np.asarray((0.5, 1.0, 0.5)), 1280, 960, 60.0) - large = fit_radius(np.asarray((1.0, 2.0, 1.0)), 1280, 960, 60.0) - assert large > small - assert np.isclose(large / small, 2.0, rtol=1e-9) - - -@pytest.mark.parametrize("azimuths,elevations", [(2, (0.0,)), (8, ())]) -def test_rig_rejects_degenerate_requests(azimuths, elevations): - with pytest.raises(ValueError): - orbit_rig(BOUNDS_MIN, BOUNDS_MAX, 1280, 960, azimuths=azimuths, elevations=elevations) diff --git a/open4d/reconstruction/nevo/nevo_tests/test_cdf.py b/open4d/reconstruction/nevo/nevo_tests/test_cdf.py deleted file mode 100644 index 80bd3184..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_cdf.py +++ /dev/null @@ -1,123 +0,0 @@ -"""The CDF is the artefact the whole first stage is judged on. - -Its one number that matters -- the share of voxels below 0.025 -- comes out of -a histogram rather than a sort, so the binning has to be exact at that -threshold and the pooling has to mean what it claims. -""" -from __future__ import annotations - -import numpy as np -import pytest - -from nevo import cdf as cdf_module - - -def test_fraction_below_is_exact_at_the_paper_threshold(): - """0.025 lands on a bin edge by construction, so no sample is misbinned.""" - rng = np.random.default_rng(0) - values = rng.random(200_000) - accumulator = cdf_module.ImportanceAccumulator() - accumulator.add(values) - expected = float((values < cdf_module.PAPER_THRESHOLD).mean()) - assert accumulator.fraction_below(cdf_module.PAPER_THRESHOLD) == pytest.approx( - expected, abs=1e-9 - ) - - -def test_default_bins_put_an_edge_on_the_threshold(): - edge = cdf_module.PAPER_THRESHOLD * cdf_module.DEFAULT_BINS - assert edge == int(edge) - - -def test_accumulation_is_order_independent(): - rng = np.random.default_rng(1) - chunks = [rng.random((7, 13)) for _ in range(5)] - first = cdf_module.ImportanceAccumulator() - second = cdf_module.ImportanceAccumulator() - for chunk in chunks: - first.add(chunk) - for chunk in reversed(chunks): - second.add(chunk) - assert np.array_equal(first.counts, second.counts) - assert first.total == second.total - - -def test_quantiles_and_mean_track_the_data(): - values = np.linspace(0.0, 1.0, 100_001) - accumulator = cdf_module.ImportanceAccumulator() - accumulator.add(values) - assert accumulator.mean == pytest.approx(0.5, abs=1e-6) - assert accumulator.quantile(0.5) == pytest.approx(0.5, abs=1e-4) - assert accumulator.quantile(0.9) == pytest.approx(0.9, abs=1e-4) - - -def test_non_finite_scores_are_refused(): - accumulator = cdf_module.ImportanceAccumulator() - with pytest.raises(ValueError): - accumulator.add(np.asarray([0.1, np.nan])) - - -def test_curve_is_monotone_and_spans_the_distribution(): - rng = np.random.default_rng(2) - accumulator = cdf_module.ImportanceAccumulator() - accumulator.add(rng.random(50_000) ** 3) - curve = accumulator.curve(points=128) - importance = np.asarray(curve["importance"]) - probability = np.asarray(curve["cdf"]) - assert np.all(np.diff(importance) >= -1e-12) - assert np.all(np.diff(probability) >= -1e-12) - assert probability[-1] == pytest.approx(1.0, abs=1e-6) - - -def test_empty_accumulator_reports_nan_rather_than_dividing_by_zero(): - accumulator = cdf_module.ImportanceAccumulator() - assert np.isnan(accumulator.fraction_below(0.025)) - assert np.isnan(accumulator.mean) - assert accumulator.curve() == {"importance": [], "cdf": []} - - -def test_excluding_empty_blocks_is_the_callers_job_and_changes_the_answer(): - """Empty blocks are never in the bitstream, so counting them would inflate - the removable fraction for free. `importance_cdf.py` subsets to the - occupancy union before feeding the accumulator; this pins what that is - worth.""" - scores = np.zeros((4, 10)) - scores[:, :5] = 0.9 - occupancy = np.zeros(10, dtype=bool) - occupancy[:5] = True - - occupied_only = cdf_module.ImportanceAccumulator() - occupied_only.add(scores[:, occupancy]) - everything = cdf_module.ImportanceAccumulator() - everything.add(scores) - - assert occupied_only.total == 4 * 5 - assert occupied_only.fraction_below(0.025) == 0.0 - assert everything.fraction_below(0.025) == 0.5 - - -def test_per_frame_pooling_takes_the_best_viewport(): - """Per-viewport asks "can this fetch skip it"; per-frame asks "is it ever - worth sending". The second is always the smaller fraction.""" - scores = np.asarray([[0.0, 0.5], [0.8, 0.0]]) - per_viewport = cdf_module.ImportanceAccumulator() - per_viewport.add(scores) - per_frame = cdf_module.ImportanceAccumulator() - per_frame.add(scores.max(axis=0)) - - assert per_viewport.fraction_below(0.025) == 0.5 - assert per_frame.fraction_below(0.025) == 0.0 - assert per_frame.total == 2 - - -def test_never_hit_fraction_counts_only_the_zero_bin(): - """Voxels no ray touched are the free part of the removable fraction, so - they have to be separable from voxels that were merely dim.""" - accumulator = cdf_module.ImportanceAccumulator() - accumulator.add(np.concatenate([np.zeros(300), np.full(100, 0.001), np.full(600, 0.5)])) - assert accumulator.never_hit_fraction == pytest.approx(0.3) - assert accumulator.fraction_below(0.025) == pytest.approx(0.4) - - -def test_never_hit_fraction_is_nan_when_empty(): - assert np.isnan(cdf_module.ImportanceAccumulator().never_hit_fraction) diff --git a/open4d/reconstruction/nevo/nevo_tests/test_filtering.py b/open4d/reconstruction/nevo/nevo_tests/test_filtering.py deleted file mode 100644 index cea3603c..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_filtering.py +++ /dev/null @@ -1,124 +0,0 @@ -"""Dropping voxels the way ReRF's decoder sees a block that never arrived.""" -from __future__ import annotations - -import numpy as np -import pytest - -torch = pytest.importorskip("torch") - -from nevo.blocks import BlockGrid # noqa: E402 -from nevo.filtering import ABSENT_DENSITY, entry_mask_from_blocks # noqa: E402 -from nevo_tests.conftest import needs_sequence, trained_config # noqa: E402 - - -def test_entry_mask_expands_each_block_over_its_entries(): - grid = BlockGrid((4, 4, 4), 2) # 2x2x2 blocks - keep = torch.zeros(grid.num_blocks, dtype=torch.bool) - keep[grid.block_index(torch.tensor([[0, 0, 0]]))] = True - mask = entry_mask_from_blocks(grid, keep) - assert mask.shape == (4, 4, 4) - assert int(mask.sum()) == 8 - assert bool(mask[0, 0, 0]) and bool(mask[1, 1, 1]) - assert not bool(mask[2, 0, 0]) - - -def test_entry_mask_trims_the_padding_off_a_ragged_grid(): - """A 9-entry axis pads to 16; the expansion must come back at 9, not 16, - or it will not line up with the density grid it indexes.""" - grid = BlockGrid((9, 5, 3), 8) - keep = torch.ones(grid.num_blocks, dtype=torch.bool) - mask = entry_mask_from_blocks(grid, keep) - assert mask.shape == (9, 5, 3) - assert bool(mask.all()) - - -def test_entry_mask_round_trips_every_block_index(): - grid = BlockGrid((6, 10, 4), 2) - for block in range(grid.num_blocks): - keep = torch.zeros(grid.num_blocks, dtype=torch.bool) - keep[block] = True - mask = entry_mask_from_blocks(grid, keep) - entries = torch.nonzero(mask) - if entries.numel() == 0: - continue # a block that is entirely padding - assert set(grid.block_index(entries).tolist()) == {block} - - -def test_absent_density_is_the_decoders_value_not_zero(): - """Zeroing a dropped block would paint fog: raw density 0 activates through - softplus(0 - 4.595) to a clearly visible alpha, whereas -4.1 (what - rerf_render.py fills an undelivered block with) does not.""" - act_shift = -4.595119850134584 - interval = 1.0 - def alpha(density): - return 1.0 - np.exp(-float(np.log1p(np.exp(density + act_shift))) * interval) - - assert alpha(0.0) > 1e-2 - assert alpha(ABSENT_DENSITY) < 1e-3 - - -# ------------------------------------------------------- against a real model -@needs_sequence -def test_dropping_voxels_restores_the_model_exactly(): - """The model is shared with the caller's sequence. A leaked mutation would - silently corrupt every later render of the same frame, and the symptom -- a - frame that renders slightly wrong the second time -- is nasty to trace.""" - from nevo import rerf_env - - rerf_env.activate() - from nevo.filtering import voxels_dropped - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(trained_config()) - frames = sequence.available_frames() - if not frames: - pytest.skip("sequence has no trained frames yet") - frame = sequence.frame(frames[0]) - grid = BlockGrid(frame.grid_shape, 8) - before_density = frame.model.density.data.clone() - before_k0 = frame.model.k0.k0.data.clone() - - keep = torch.zeros(grid.num_blocks, dtype=torch.bool, device=before_k0.device) - keep[: grid.num_blocks // 2] = True - with voxels_dropped(frame, grid, keep) as model: - assert not torch.equal(model.density.data, before_density), "nothing was dropped" - assert float(model.density.data.min()) <= ABSENT_DENSITY + 1e-6 - - assert torch.equal(frame.model.density.data, before_density) - assert torch.equal(frame.model.k0.k0.data, before_k0) - - -@needs_sequence -def test_dropping_nothing_leaves_the_render_unchanged(): - """Keeping every block must be a no-op on the pixels. - - Not *bit*-identical: DVGO accumulates a pixel's colour with - ``torch_scatter.segment_coo``, whose atomic adds run in whatever order the - GPU schedules, so two renders of the same model differ in the last couple - of float bits. The bound below is far tighter than any filtering effect and - is measured against a repeat render of the untouched model, so it fails if - the mask leaks rather than if the GPU is merely non-deterministic. - """ - from nevo import rerf_env - - rerf_env.activate() - from nevo.filtering import voxels_dropped - from nevo.render import render_view, training_view - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(trained_config()) - frames = sequence.available_frames() - if not frames: - pytest.skip("sequence has no trained frames yet") - frame = sequence.frame(frames[0]) - grid = BlockGrid(frame.grid_shape, 8) - camera, _ = training_view(sequence, frames[0], 0) - reference = render_view(sequence, frame, camera) - repeated = render_view(sequence, frame, camera) - noise = float(np.abs(reference - repeated).max()) - - keep = torch.ones(grid.num_blocks, dtype=torch.bool, device=frame.density.device) - with voxels_dropped(frame, grid, keep): - kept_everything = render_view(sequence, frame, camera) - difference = float(np.abs(reference - kept_everything).max()) - assert difference <= max(noise, 1e-5), (difference, noise) diff --git a/open4d/reconstruction/nevo/nevo_tests/test_importance.py b/open4d/reconstruction/nevo/nevo_tests/test_importance.py deleted file mode 100644 index 63694aea..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_importance.py +++ /dev/null @@ -1,208 +0,0 @@ -"""The importance pass. - -The scatter half is pure tensor arithmetic and always runs. The marching half -needs a trained ReRF sequence and a GPU with room to work in; see -``conftest.py`` for how that is gated. -""" -from __future__ import annotations - -import numpy as np -import pytest - -torch = pytest.importorskip("torch") - -from nevo.blocks import BlockGrid # noqa: E402 -from nevo_tests.conftest import needs_sequence, trained_config # noqa: E402 - - -# ------------------------------------------------------------------- scatter -def test_scatter_max_keeps_the_largest_weight_per_block(): - from nevo.importance import scatter_max - - grid = BlockGrid((8, 8, 8), 4) # 2x2x2 = 8 blocks - xyz_min = torch.zeros(3) - xyz_max = torch.ones(3) - # Two samples in the origin block, one in the far corner block. - points = torch.tensor([[0.0, 0.0, 0.0], [0.05, 0.05, 0.05], [1.0, 1.0, 1.0]]) - weights = torch.tensor([0.2, 0.7, 0.4]) - scores = torch.zeros(grid.num_blocks) - scatter_max(grid, points, weights, xyz_min, xyz_max, (8, 8, 8), "nearest", scores) - assert scores[0].item() == pytest.approx(0.7) - assert scores[-1].item() == pytest.approx(0.4) - assert scores.gt(0).sum().item() == 2 - - -def test_scatter_max_accumulates_across_calls(): - """One viewport is marched in ray chunks, so the running max has to survive - being fed a piece at a time.""" - from nevo.importance import scatter_max - - grid = BlockGrid((4, 4, 4), 4) - xyz_min, xyz_max = torch.zeros(3), torch.ones(3) - scores = torch.zeros(grid.num_blocks) - for weight in (0.1, 0.9, 0.3): - scatter_max( - grid, - torch.tensor([[0.5, 0.5, 0.5]]), - torch.tensor([weight]), - xyz_min, - xyz_max, - (4, 4, 4), - "nearest", - scores, - ) - assert scores.max().item() == pytest.approx(0.9) - - -def test_scatter_max_handles_an_empty_sample_list(): - from nevo.importance import scatter_max - - grid = BlockGrid((4, 4, 4), 4) - scores = torch.zeros(grid.num_blocks) - result = scatter_max( - grid, - torch.zeros(0, 3), - torch.zeros(0), - torch.zeros(3), - torch.ones(3), - (4, 4, 4), - "nearest", - scores, - ) - assert result.abs().sum().item() == 0.0 - - -def test_trilinear_assignment_covers_at_least_what_nearest_does(): - """A sample's nearest entry is always one of the eight it interpolates - from, so trilinear can only ever mark more blocks, never fewer.""" - from nevo.importance import scatter_max - - torch.manual_seed(0) - shape = (16, 16, 16) - grid = BlockGrid(shape, 4) - xyz_min, xyz_max = torch.zeros(3), torch.ones(3) - points = torch.rand(500, 3) - weights = torch.rand(500) - nearest = torch.zeros(grid.num_blocks) - trilinear = torch.zeros(grid.num_blocks) - scatter_max(grid, points, weights, xyz_min, xyz_max, shape, "nearest", nearest) - scatter_max(grid, points, weights, xyz_min, xyz_max, shape, "trilinear", trilinear) - assert torch.all(trilinear >= nearest - 1e-12) - assert trilinear.gt(0).sum() >= nearest.gt(0).sum() - - -def test_config_rejects_an_unknown_assignment(): - from nevo.importance import ImportanceConfig - - with pytest.raises(ValueError): - ImportanceConfig(assignment="bilinear") - with pytest.raises(ValueError): - ImportanceConfig(ray_chunk=0) - - -# ------------------------------------------------------- against a real model -@pytest.fixture(scope="module") -def scored(): - config = trained_config() - if config is None: - pytest.skip("no trained sequence") - from nevo import rerf_env - - rerf_env.activate() - import json - - from nevo.importance import ImportanceConfig - from nevo.sequence import ReRFSequence - from nevo.viewports import sample_viewports - - sequence = ReRFSequence(config) - frames = sequence.available_frames() - if not frames: - pytest.skip("sequence has no trained frames yet") - frame = sequence.frame(frames[0]) - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - reference = manifest["cameras"][0] - radius = float(np.linalg.norm(np.asarray(reference["c2w_normalised"])[:3, 3])) - cameras = sample_viewports( - 4, - manifest["xyz_min"], - manifest["xyz_max"], - reference_radius=radius, - width=320, - height=240, - focal=float(reference["fx"]) / 4.0, - seed=11, - ) - return sequence, frame, cameras, ImportanceConfig - - -@needs_sequence -def test_marching_reproduces_rerfs_own_forward_pass(scored): - """The whole of step 2 rests on this: our transcription of ReRF's ray - marching produces the identical weight sequence, so the weights we scatter - are the ones that actually rendered the frame.""" - from nevo.importance import check_against_rerf - - sequence, frame, cameras, _ = scored - result = check_against_rerf(sequence, frame, cameras[0]) - assert result["agrees"], result - assert result["samples_ours"] > 0 - - -@needs_sequence -def test_importance_is_a_weight_so_it_stays_in_the_unit_interval(scored): - from nevo.importance import ImportanceScorer - - sequence, frame, cameras, ImportanceConfig = scored - scores = ImportanceScorer(sequence, frame, ImportanceConfig(block_size=8)).score(cameras[0]) - assert float(scores.min()) >= 0.0 - assert float(scores.max()) <= 1.0 + 1e-6 - assert float(scores.max()) > 0.1, "no voxel was visible at all" - - -@needs_sequence -def test_codec_empty_blocks_can_still_render(scored): - """ReRF's two "is there anything here" tests disagree, on purpose. - - The codec keeps a block when some entry has raw density above ~3.39 - (``softplus(d - 4.1) > 0.4``). The renderer keeps a *sample* when its alpha - clears ``fast_color_thres = 1e-4``, which for ``act_shift = -4.595`` is raw - density above about -4.6. Six orders of magnitude apart: low-density haze - renders but is never transmitted. - - That is ReRF's own behaviour, not something NeVo introduces, and it bounds - what block-level filtering can preserve -- so it is pinned rather than - asserted away. :func:`unscored_weight_outside_codec_mask` reports how much - of it there is; if this ever starts failing, the codec mask and the - renderer have been reconciled and the CDF's denominator should be revisited. - """ - from nevo.importance import ImportanceScorer - - sequence, frame, cameras, ImportanceConfig = scored - scorer = ImportanceScorer(sequence, frame, ImportanceConfig(block_size=8)) - scores = scorer.score(cameras[0]).cpu().numpy() - occupied = scorer.occupancy.cpu().numpy() - assert float(scores[occupied].max(initial=0.0)) > 0.5, "the subject did not render" - outside = scores[~occupied] - assert float(outside.max(initial=0.0)) > 0.0, ( - "the codec mask and the renderer now agree; revisit the CDF denominator" - ) - # Whatever leaks past the codec mask is a minority of the visible weight. - assert outside.sum() < scores[occupied].sum() - - -@needs_sequence -def test_finer_blocks_leave_a_longer_tail(scored): - """Importance is a max over a block, so coarsening can only raise it. This - is why the fraction below a threshold is a statement about granularity as - much as about content.""" - from nevo.importance import ImportanceScorer - - sequence, frame, cameras, ImportanceConfig = scored - fractions = {} - for block_size in (1, 8): - scorer = ImportanceScorer(sequence, frame, ImportanceConfig(block_size=block_size)) - scores = scorer.score(cameras[0]).cpu().numpy()[scorer.occupancy.cpu().numpy()] - fractions[block_size] = float((scores < 0.025).mean()) - assert fractions[1] > fractions[8] diff --git a/open4d/reconstruction/nevo/nevo_tests/test_viewports.py b/open4d/reconstruction/nevo/nevo_tests/test_viewports.py deleted file mode 100644 index ab9ebb42..00000000 --- a/open4d/reconstruction/nevo/nevo_tests/test_viewports.py +++ /dev/null @@ -1,183 +0,0 @@ -"""Viewport sampling and 6DoF trace decoding.""" -from __future__ import annotations - -import math -import textwrap - -import numpy as np -import pytest - -from nevo import viewports -from nevo.cameras import Camera - -BOUNDS_MIN = (-0.3, -0.6, -0.3) -BOUNDS_MAX = (0.3, 0.6, 0.3) - -TRACE = textwrap.dedent( - """\ - t_s,scene,streaming,broadcast_id,x,y,z,yaw,pitch,roll,pred_valid,pred_made_at_s,pred_target_s,pred_x,pred_y,pred_z,pred_yaw,pred_pitch,pred_roll - 0.000,stage,0,b0,0.0,1.6,-2.0,0,0,0,0,,,,,,,, - 0.033,stage,1,b0,0.01,1.6,-2.0,5,1,0,1,0.000,0.033,0.02,1.61,-2.01,6,1,0 - 0.066,stage,1,b0,0.03,1.6,-1.98,9,2,0,0,,,,,,,, - """ -) - - -def test_sampled_viewports_stay_in_their_shell_and_face_the_content(): - cameras = viewports.sample_viewports( - 64, BOUNDS_MIN, BOUNDS_MAX, reference_radius=2.0, width=64, height=48, focal=55.0 - ) - centre = (np.asarray(BOUNDS_MIN) + np.asarray(BOUNDS_MAX)) * 0.5 - spread = viewports.ViewportSpread() - for camera in cameras: - eye = camera.c2w[:3, 3] - radius = np.linalg.norm(eye - centre) - assert 2.0 * spread.radius_scale[0] - 1e-9 <= radius <= 2.0 * spread.radius_scale[1] + 1e-9 - forward = camera.c2w[:3, 2] - assert np.dot(forward, centre - eye) > 0 - assert np.allclose(camera.c2w[:3, :3] @ camera.c2w[:3, :3].T, np.eye(3), atol=1e-9) - - -def test_sampling_is_reproducible_and_seed_dependent(): - common = dict( - xyz_min=BOUNDS_MIN, xyz_max=BOUNDS_MAX, reference_radius=2.0, - width=64, height=48, focal=55.0, - ) - first = viewports.sample_viewports(8, seed=3, **common) - again = viewports.sample_viewports(8, seed=3, **common) - other = viewports.sample_viewports(8, seed=4, **common) - assert np.allclose([c.c2w for c in first], [c.c2w for c in again]) - assert not np.allclose([c.c2w for c in first], [c.c2w for c in other]) - - -def test_elevation_spread_actually_leaves_the_horizontal_ring(): - cameras = viewports.sample_viewports( - 200, BOUNDS_MIN, BOUNDS_MAX, reference_radius=2.0, width=64, height=48, focal=55.0 - ) - centre = (np.asarray(BOUNDS_MIN) + np.asarray(BOUNDS_MAX)) * 0.5 - heights = np.asarray([camera.c2w[1, 3] - centre[1] for camera in cameras]) - assert heights.min() < -0.2 - assert heights.max() > 0.8 - - -def _ray_direction(camera: Camera, u: float, v: float) -> np.ndarray: - """Direction of the ray through a normalised image coordinate, OpenCV.""" - x = (u * camera.width - 0.5 - camera.cx) / camera.fx - y = (v * camera.height - 0.5 - camera.cy) / camera.fy - direction = camera.c2w[:3, :3] @ np.asarray((x, y, 1.0)) - return direction / np.linalg.norm(direction) - - -def test_downscale_preserves_the_view_frustum(): - """Halving resolution must not quietly change what the camera sees.""" - original = viewports.sample_viewports( - 4, BOUNDS_MIN, BOUNDS_MAX, reference_radius=2.0, width=64, height=48, focal=55.0 - ) - smaller = viewports.downscale(original, 2) - for full, half in zip(original, smaller): - assert (half.width, half.height) == (32, 24) - for u, v in ((0.0, 0.0), (0.5, 0.5), (1.0, 1.0), (0.25, 0.75)): - assert np.allclose(_ray_direction(full, u, v), _ray_direction(half, u, v), atol=1e-9) - - -def test_downscale_by_one_is_identity(): - original = viewports.sample_viewports( - 2, BOUNDS_MIN, BOUNDS_MAX, reference_radius=2.0, width=64, height=48, focal=55.0 - ) - assert [c.c2w.tolist() for c in viewports.downscale(original, 1)] == [ - c.c2w.tolist() for c in original - ] - - -def test_read_quest_trace_keeps_streaming_rows_and_marks_missing_predictions(tmp_path): - path = tmp_path / "viewport.csv" - path.write_text(TRACE) - trace = viewports.read_quest_trace(path) - assert len(trace) == 2 - assert trace.times.tolist() == [0.033, 0.066] - assert np.isfinite(trace.predicted[0]).all() - assert np.isnan(trace.predicted[1]).all() - assert trace.prediction_error[0] == pytest.approx( - math.sqrt(0.01 ** 2 + 0.01 ** 2 + 0.01 ** 2), rel=1e-6 - ) - - -def test_read_quest_trace_can_keep_the_non_streaming_rows(tmp_path): - path = tmp_path / "viewport.csv" - path.write_text(TRACE) - assert len(viewports.read_quest_trace(path, streaming_only=False)) == 3 - - -POSES = ((0, 1.6, -2, 0, 0, 0), (1, 1.5, 0.5, 37, -12, 5), (-1, 1, 2, 180, 0, 0)) - - -@pytest.mark.parametrize("convention", viewports.CONVENTIONS) -def test_pose_to_c2w_is_rigid_and_right_handed(convention): - """A determinant of -1 is the failure mode that matters here: it still - looks like a rotation, but renders a mirrored view.""" - for pose in POSES: - c2w = viewports.pose_to_c2w(pose, convention) - rotation = c2w[:3, :3] - assert np.allclose(rotation @ rotation.T, np.eye(3), atol=1e-9) - assert np.isclose(np.linalg.det(rotation), 1.0, atol=1e-9) - - -@pytest.mark.parametrize("convention", viewports.CONVENTIONS) -def test_identity_pose_faces_negative_z_with_y_down(convention): - """Both source frames put an unrotated camera looking down world -Z once - they are in OpenCV axes -- Unity because the world Z flips, glTF because - that is where its camera already points.""" - c2w = viewports.pose_to_c2w((0.0, 0.0, 0.0, 0.0, 0.0, 0.0), convention) - assert np.allclose(c2w[:3, 2], (0.0, 0.0, -1.0), atol=1e-12) - assert np.allclose(c2w[:3, 1], (0.0, -1.0, 0.0), atol=1e-12) - - -def test_a_left_handed_yaw_is_a_right_handed_yaw_of_the_other_sign(): - """The point of the world flip. If this ever stops holding, one of the two - branches has picked up a mirror.""" - unity = viewports.pose_to_c2w((0.0, 0.0, 0.0, 90.0, 0.0, 0.0), "unity") - right_handed = viewports.pose_to_c2w((0.0, 0.0, 0.0, -90.0, 0.0, 0.0), "right_handed") - assert np.allclose(unity[:3, :3], right_handed[:3, :3], atol=1e-12) - - -def test_unity_position_z_is_flipped(): - c2w = viewports.pose_to_c2w((1.0, 2.0, 3.0, 0.0, 0.0, 0.0), "unity") - assert np.allclose(c2w[:3, 3], (1.0, 2.0, -3.0)) - plain = viewports.pose_to_c2w((1.0, 2.0, 3.0, 0.0, 0.0, 0.0), "right_handed") - assert np.allclose(plain[:3, 3], (1.0, 2.0, 3.0)) - - -def test_unknown_convention_is_refused(): - with pytest.raises(ValueError): - viewports.pose_to_c2w(POSES[0], "opengl") - - -def test_trace_viewports_applies_the_corpus_normalisation(tmp_path): - path = tmp_path / "viewport.csv" - path.write_text(TRACE) - trace = viewports.read_quest_trace(path) - centre = (0.5, 1.5, 0.0) - scale = 4.0 - cameras = viewports.trace_viewports( - trace, width=32, height=24, focal=30.0, centre=centre, scale=scale - ) - assert len(cameras) == 2 - raw = viewports.pose_to_c2w(trace.actual[0], "unity")[:3, 3] - assert np.allclose(cameras[0].c2w[:3, 3], (raw - np.asarray(centre)) * scale) - - -def test_trace_viewports_skips_rows_without_a_prediction(tmp_path): - path = tmp_path / "viewport.csv" - path.write_text(TRACE) - trace = viewports.read_quest_trace(path) - cameras = viewports.trace_viewports( - trace, width=32, height=24, focal=30.0, centre=(0, 0, 0), scale=1.0, use_predicted=True - ) - assert [camera.camera_id for camera in cameras] == [0] - - -def test_empty_trace_is_an_error(tmp_path): - path = tmp_path / "viewport.csv" - path.write_text(TRACE.splitlines()[0] + "\n") - with pytest.raises(ValueError): - viewports.read_quest_trace(path) diff --git a/open4d/reconstruction/nevo/orbitnevo/__init__.py b/open4d/reconstruction/nevo/orbitnevo/__init__.py deleted file mode 100644 index 70835734..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/__init__.py +++ /dev/null @@ -1 +0,0 @@ -"""ORBIT-corpus adapters and CLI entry points for the NeVo baseline.""" diff --git a/open4d/reconstruction/nevo/orbitnevo/filter_sweep.py b/open4d/reconstruction/nevo/orbitnevo/filter_sweep.py deleted file mode 100644 index 30d4411f..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/filter_sweep.py +++ /dev/null @@ -1,178 +0,0 @@ -"""What does filtering by neural visibility actually cost in pixels? - -The CDF answers half of the paper's section 3.2 claim -- that most voxels score -below 0.025. The other half is the part that makes it a *saving* rather than a -loss: "The SSIM consistently exceeds 0.98 (i.e., visually lossless), indicating -that removing these voxels does not impact visual quality." - -This sweeps thresholds and measures both ends together: how many non-empty -blocks each threshold drops, and the SSIM and PSNR of the resulting render -against the same viewport rendered from the whole grid. Filtering is done -per-viewport with that viewport's own scores, which is the decision the edge -would make for a *correctly predicted* viewport -- an upper bound on the real -system, where the fetch is chosen from a prediction several frames early. - - python -m orbitnevo.filter_sweep \\ - --config baselines/NeVo/rerf/configs/nevo/g_basketball.py \\ - --out ~/nevo_results/g_basketball - -Runs in the ``nevo`` environment. -""" -from __future__ import annotations - -import argparse -import json -import sys -import time -from pathlib import Path - -import numpy as np - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import rerf_env, viewports # noqa: E402 - -VISUALLY_LOSSLESS_SSIM = 0.98 -"""The bar the paper cites (Cuervo et al., Kahawai, MobiSys'15).""" - - -def run(args) -> dict: - from nevo.blocks import BlockGrid - from nevo.filtering import preview - from nevo.importance import ImportanceConfig, ImportanceScorer - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(args.config) - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - trained = sequence.available_frames() - if not trained: - raise RuntimeError(f"no trained frames under {sequence.run_dir}") - frames = trained[: args.frames] if args.frames > 0 else trained - - reference = manifest["cameras"][0] - radius = float(np.linalg.norm(np.asarray(reference["c2w_normalised"])[:3, 3])) - cameras = viewports.sample_viewports( - args.viewports, - manifest["xyz_min"], - manifest["xyz_max"], - reference_radius=radius, - width=int(manifest["width"]), - height=int(manifest["height"]), - focal=float(reference["fx"]), - # A different seed from importance_cdf.py's default, so the thresholds - # are scored on viewports the CDF did not also report. - seed=args.seed, - ) - cameras = viewports.downscale(cameras, args.render_factor) - print(f"{sequence.cfg.expname}: {len(frames)} frames x {len(cameras)} viewports at " - f"{cameras[0].width}x{cameras[0].height}, block {args.block_size}^3", flush=True) - - config = ImportanceConfig(block_size=args.block_size, assignment=args.assignment) - samples = {threshold: [] for threshold in args.thresholds} - for frame_index in frames: - started = time.time() - frame = sequence.frame(frame_index) - scorer = ImportanceScorer(sequence, frame, config) - grid = BlockGrid(frame.grid_shape, args.block_size) - for camera in cameras: - scores = scorer.score(camera) - for threshold in args.thresholds: - result = preview( - sequence, frame, camera, scores, grid, threshold, scorer.occupancy - ) - samples[threshold].append( - (result.dropped_fraction, result.ssim, result.psnr) - ) - print(f"frame {frame_index:3d} {'I' if frame.is_key_frame else 'P'} " - f"{time.time() - started:.0f}s", flush=True) - del frame, scorer - - rows = [] - for threshold in args.thresholds: - values = np.asarray(samples[threshold]) - rows.append( - { - "threshold": threshold, - "dropped_mean": float(values[:, 0].mean()), - "ssim_mean": float(values[:, 1].mean()), - "ssim_min": float(values[:, 1].min()), - "psnr_mean": float(values[:, 2].mean()), - "psnr_min": float(values[:, 2].min()), - "visually_lossless": bool(values[:, 1].min() >= VISUALLY_LOSSLESS_SSIM), - } - ) - row = rows[-1] - print( - f"threshold {threshold:<7} dropped {row['dropped_mean'] * 100:5.1f}% " - f"SSIM {row['ssim_mean']:.4f} (worst {row['ssim_min']:.4f}) " - f"PSNR {row['psnr_mean']:5.1f} dB (worst {row['psnr_min']:5.1f}) " - f"{'lossless' if row['visually_lossless'] else 'DEGRADED'}", - flush=True, - ) - - best = max( - (row for row in rows if row["visually_lossless"]), - key=lambda row: row["dropped_mean"], - default=None, - ) - if best is not None: - print( - f"\nlargest visually-lossless threshold: {best['threshold']}, " - f"dropping {best['dropped_mean'] * 100:.1f}% of non-empty blocks", - flush=True, - ) - else: - print("\nno threshold stayed above SSIM 0.98", flush=True) - - report = { - "object": sequence.cfg.expname, - "config": str(sequence.config_path), - "frames": frames, - "viewports": len(cameras), - "viewport_size": [cameras[0].width, cameras[0].height], - "block_size": args.block_size, - "assignment": args.assignment, - "seed": args.seed, - "ssim_bar": VISUALLY_LOSSLESS_SSIM, - "rows": rows, - "best_lossless": best, - } - out_dir = Path(args.out).expanduser().resolve() - out_dir.mkdir(parents=True, exist_ok=True) - path = out_dir / f"filter_sweep_block{args.block_size}.json" - with open(path, "w") as handle: - json.dump(report, handle, indent=1) - print(f"wrote {path}", flush=True) - return report - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--config", required=True) - parser.add_argument("--out", required=True) - parser.add_argument("--frames", type=int, default=0, help="0 = every trained frame") - parser.add_argument("--viewports", type=int, default=8) - parser.add_argument("--render-factor", type=int, default=1) - parser.add_argument("--block-size", type=int, default=8) - parser.add_argument("--assignment", default="nearest", choices=("nearest", "trilinear")) - parser.add_argument("--thresholds", type=float, nargs="+", - default=[0.01, 0.025, 0.05, 0.1, 0.2, 0.35, 0.5, 0.7], - help="ranges well past the paper's 0.025 on purpose: the paper " - "*fits* its threshold to an SSIM target rather than fixing " - "it, so what matters is where quality actually breaks") - parser.add_argument("--seed", type=int, default=101) - return parser.parse_args(argv) - - -def main(argv=None) -> int: - args = parse_args(argv) - rerf_env.activate() - run(args) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/importance_cdf.py b/open4d/reconstruction/nevo/orbitnevo/importance_cdf.py deleted file mode 100644 index f29b4edf..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/importance_cdf.py +++ /dev/null @@ -1,303 +0,0 @@ -"""Steps 1-2 end to end: load a ReRF sequence, score its voxels, plot the CDF. - - python -m orbitnevo.importance_cdf \\ - --config baselines/NeVo/rerf/configs/nevo/basketball.py \\ - --viewports 300 --out ~/nevo_results/basketball - -Writes ``importance_cdf.json`` (summaries plus the plottable curves), -``importance_cdf.png``, and per-frame ``scores_.npy`` so later stages do -not have to re-march. ``--verify`` additionally checks that this module's -transcription of ReRF's ray marching returns the same weights the vendored -model does. - -Must run in the ``nevo`` conda environment; see baselines/NeVo/README.md. -""" -from __future__ import annotations - -import argparse -import json -import sys -import time -from pathlib import Path - -import numpy as np - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import cdf as cdf_module # noqa: E402 -from nevo import rerf_env, viewports # noqa: E402 - -# Importing lib.dvgo has import-time side effects (a JIT CUDA build, a chdir -# requirement, torch's default tensor type flipped to cuda), so it happens -# through rerf_env and only once the arguments have been parsed. - - -def _corpus_manifest(sequence) -> dict: - path = sequence.corpus_dir / "nevo_corpus.json" - if not path.is_file(): - raise FileNotFoundError(f"{path} not found; was this corpus made by orbitnevo.prepare?") - with open(path) as handle: - return json.load(handle) - - -def run(args) -> dict: - from nevo.importance import ( - ImportanceConfig, - ImportanceScorer, - check_against_rerf, - ) - from nevo.render import check_reload - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(args.config) - manifest = _corpus_manifest(sequence) - trained = sequence.available_frames() - if not trained: - raise RuntimeError(f"no trained frames under {sequence.run_dir}") - frames = trained[: args.frames] if args.frames > 0 else trained - print(f"{sequence.cfg.expname}: {len(trained)} trained frames, scoring {len(frames)}", - flush=True) - - out_dir = Path(args.out).expanduser().resolve() - out_dir.mkdir(parents=True, exist_ok=True) - - reference = manifest["cameras"][0] - focal = float(reference["fx"]) - width, height = int(manifest["width"]), int(manifest["height"]) - radius = float(np.linalg.norm(np.asarray(reference["c2w_normalised"])[:3, 3])) - - spread = viewports.ViewportSpread( - radius_scale=tuple(args.radius_scale), - elevation_degrees=tuple(args.elevation), - aim_jitter=args.aim_jitter, - ) - cameras = viewports.sample_viewports( - args.viewports, - manifest["xyz_min"], - manifest["xyz_max"], - reference_radius=radius, - width=width, - height=height, - focal=focal, - spread=spread, - seed=args.seed, - ) - cameras = viewports.downscale(cameras, args.render_factor) - print(f"{len(cameras)} viewports at {cameras[0].width}x{cameras[0].height}, " - f"rig radius {radius:.3f}", flush=True) - - config = ImportanceConfig( - block_size=args.block_size, - assignment=args.assignment, - render_factor=1, # already applied to the cameras - ray_chunk=args.ray_chunk, - ) - - # Pre-pass for the occupancy union. Which blocks are non-empty shifts as the - # subject moves, and pooling a CDF over frames needs one fixed column set; - # taking the union means a block occupied in any frame is scored in all of - # them, rather than silently changing the denominator per frame. - from nevo.blocks import BlockGrid - - occupancy = None - grid = None - for frame_index in frames: - density = sequence.frame_density(frame_index) - if grid is None: - grid = BlockGrid.from_volume(density, args.block_size) - frame_occupancy = grid.occupancy(density).cpu().numpy() - occupancy = frame_occupancy if occupancy is None else np.logical_or(occupancy, frame_occupancy) - del density - occupied_index = np.flatnonzero(occupancy) - print( - f"occupancy union: {occupied_index.size}/{grid.num_blocks} blocks " - f"({occupied_index.size / grid.num_blocks * 100:.1f}%), " - f"grid {grid.grid_shape} in {grid.blocks_shape} blocks of {args.block_size}^3", - flush=True, - ) - if occupied_index.size == 0: - raise RuntimeError("no occupied blocks; is this sequence trained?") - - per_viewport = cdf_module.ImportanceAccumulator(args.bins) - per_frame = cdf_module.ImportanceAccumulator(args.bins) - verification = {} - reload_checks = {} - per_frame_notes = [] - store_bytes = 0 - - for frame_index in frames: - started = time.time() - frame = sequence.frame(frame_index) - scorer = ImportanceScorer(sequence, frame, config) - if scorer.num_blocks != grid.num_blocks: - raise RuntimeError("grid resolution changes across frames; scores cannot be pooled") - # Verify once on an I-frame and once on a P-frame: their models are - # assembled differently (a P-frame's feature grid is a residual over a - # motion-compensated predecessor), so one check does not cover both. - kind = "I" if frame.is_key_frame else "P" - if args.verify and kind not in verification: - result = check_against_rerf(sequence, frame, cameras[0]) - print(f"marching check vs. ReRF ({kind}-frame {frame.index}): {result}", flush=True) - if not result["agrees"]: - raise RuntimeError("instrumented marching disagrees with ReRF's forward pass") - verification[kind] = result - if args.verify and kind not in reload_checks: - reload = check_reload( - sequence, frame, save_to=out_dir / f"reload_{kind.lower()}frame.png" - ) - print(f"reload check ({kind}-frame {frame.index}): " - f"{reload['psnr']:.2f} dB against the training view", flush=True) - if reload["psnr"] < args.min_reload_psnr: - raise RuntimeError( - f"frame {frame.index} reloaded to {reload['psnr']:.2f} dB, below " - f"{args.min_reload_psnr} -- the checkpoint was not reassembled correctly" - ) - reload_checks[kind] = reload - - dense = scorer.score_many(cameras).numpy() - # How much visible weight sits in blocks ReRF's codec never sends. Its - # occupancy test (raw density > ~3.39) is far stricter than the - # renderer's alpha prune, so low-density haze renders but is not in the - # bitstream. That is upstream's behaviour, but it caps what *any* - # block-level filtering can preserve, so it is worth a number. - outside = float(dense[:, ~occupancy].sum()) - inside = float(dense[:, occupancy].sum()) - scores = dense[:, occupied_index] - del dense - per_viewport.add(scores) - per_frame.add(scores.max(axis=0)) - if store_bytes + scores.nbytes <= args.max_store_bytes: - np.save(out_dir / f"scores_{frame_index}.npy", scores.astype(np.float32)) - store_bytes += scores.nbytes - - note = { - "frame": frame_index, - "key_frame": frame.is_key_frame, - "grid_shape": list(frame.grid_shape), - "raw_bytes": frame.raw_bytes(), - "occupied_blocks_this_frame": scorer.occupied_blocks, - "weight_outside_codec_mask": outside / (inside + outside) if inside + outside else 0.0, - "has_motion_vectors": frame.motion is not None, - "seconds": time.time() - started, - } - per_frame_notes.append(note) - print( - f"frame {frame_index:3d} " - f"{'I' if frame.is_key_frame else 'P'} " - f"grid {tuple(frame.grid_shape)} " - f"occupied {scorer.occupied_blocks} " - f"below {cdf_module.PAPER_THRESHOLD}: " - f"{float((scores < cdf_module.PAPER_THRESHOLD).mean()) * 100:.1f}% " - f"{note['seconds']:.1f}s", - flush=True, - ) - del scores, scorer, frame - - shared = dict( - block_size=args.block_size, - assignment=args.assignment, - frames=len(frames), - viewports=len(cameras), - occupied_blocks=int(occupied_index.size), - total_blocks=int(grid.num_blocks), - extra={"object": sequence.cfg.expname}, - ) - summaries = [ - per_viewport.summary(pooling="per-viewport", **shared), - per_frame.summary(pooling="per-frame", **shared), - ] - curves = { - f"{sequence.cfg.expname} (per-viewport)": per_viewport.curve(), - f"{sequence.cfg.expname} (per-frame)": per_frame.curve(), - } - - report = { - "object": sequence.cfg.expname, - "config": str(sequence.config_path), - "run_dir": str(sequence.run_dir), - "viewports": len(cameras), - "viewport_size": [cameras[0].width, cameras[0].height], - "render_factor": args.render_factor, - "block_size": args.block_size, - "assignment": args.assignment, - "seed": args.seed, - "viewport_spread": { - "radius_scale": list(args.radius_scale), - "elevation_degrees": list(args.elevation), - "aim_jitter": args.aim_jitter, - }, - "verification": verification, - "reload_checks": reload_checks, - "frames": per_frame_notes, - "summaries": summaries, - } - name = args.tag or f"block{args.block_size}_{args.assignment}" - with open(out_dir / f"importance_cdf_{name}.json", "w") as handle: - json.dump({"report": report, "curves": curves}, handle, indent=1) - if not args.no_plot: - cdf_module.plot( - curves, - out_dir / f"importance_cdf_{name}.png", - f"{sequence.cfg.expname}: voxel importance " - f"({len(frames)} frames x {len(cameras)} viewports, block {args.block_size})", - ) - - for summary in summaries: - check = summary["paper_check"] - print( - f"[{summary['pooling']}] below {check['threshold']}: " - f"{check['observed_fraction_below'] * 100:.1f}% " - f"(paper ~{check['paper_fraction_below'] * 100:.0f}%), " - f"median {summary['quantiles']['0.5']:.4f}, " - f"never hit {summary['never_hit_fraction'] * 100:.1f}%, " - f"{summary['scored_samples']} samples", - flush=True, - ) - print(f"wrote {out_dir}/importance_cdf_{name}.json", flush=True) - return report - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--config", required=True, help="the ReRF config used for training") - parser.add_argument("--out", required=True) - parser.add_argument("--frames", type=int, default=0, help="0 = every trained frame") - parser.add_argument("--viewports", type=int, default=300) - parser.add_argument("--render-factor", type=int, default=1, - help="integer downscale of each viewport before marching") - parser.add_argument("--block-size", type=int, default=8, - help="8 = ReRF's codec block; 1 = single grid entries") - parser.add_argument("--assignment", default="nearest", choices=("nearest", "trilinear")) - parser.add_argument("--ray-chunk", type=int, default=1 << 18) - parser.add_argument("--seed", type=int, default=0) - parser.add_argument("--radius-scale", type=float, nargs=2, default=(0.75, 1.45), - metavar=("MIN", "MAX"), - help="viewer distance, as a multiple of the capture rig radius") - parser.add_argument("--elevation", type=float, nargs=2, default=(-25.0, 55.0), - metavar=("MIN", "MAX"), help="viewer elevation band, degrees") - parser.add_argument("--aim-jitter", type=float, default=0.25, - help="how far the look-at point wanders, as a fraction of the bbox") - parser.add_argument("--verify", action="store_true", - help="check the marching against ReRF's forward pass, and check that " - "each frame reloads to a sane render of its training view") - parser.add_argument("--min-reload-psnr", type=float, default=25.0, - help="fail if a reloaded frame renders worse than this") - parser.add_argument("--no-plot", action="store_true") - parser.add_argument("--bins", type=int, default=cdf_module.DEFAULT_BINS) - parser.add_argument("--tag", default="", help="suffix for the output files") - parser.add_argument("--max-store-bytes", type=int, default=2 << 30, - help="stop writing per-frame scores_*.npy past this total") - return parser.parse_args(argv) - - -def main(argv=None) -> int: - args = parse_args(argv) - rerf_env.activate() - run(args) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/live_demo.py b/open4d/reconstruction/nevo/orbitnevo/live_demo.py deleted file mode 100644 index 84c967c0..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/live_demo.py +++ /dev/null @@ -1,242 +0,0 @@ -"""Play NeVo's output in a browser, the way `orbitvega/live_demo.py` does. - -Same last hop as Vega: frames are pushed as JPEG into -`vega.streaming.mjpeg_server`'s `FrameBuffer` and served over -`multipart/x-mixed-replace`, behind a dark single-image status page. Reusing -Vega's server rather than writing another one means the two baselines are -watched through identical machinery, so nothing about the viewer can account -for a difference between them. - -The one deliberate difference is what feeds it. Vega renders live on the GPU; -a NeRF frame takes ~0.5 s to ray-march at this resolution, which is not a -playback rate, and holding a GPU for the life of a browser tab is antisocial -on a box that is also training. So this loops the frames -`orbitnevo/render_frames.py` already wrote. The status line says so. - -Each streamed frame is a horizontal composite of the conditions -- plain ReRF, -NeVo's visibility-filtered reconstruction, and the captured camera -- at the -same instant and viewpoint, which is the comparison the whole baseline exists -to show. - - python -m orbitnevo.live_demo --object g_basketball --port 8752 -""" -from __future__ import annotations - -import argparse -import io -import json -import sys -import threading -import time -from pathlib import Path - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from vega.streaming.mjpeg_server import FrameBuffer, serve_forever # noqa: E402 - -LABEL_HEIGHT = 26 - - -def load_clip(directory: Path, nevo_only: bool = False) -> dict: - with open(directory / "manifest.json") as handle: - manifest = json.load(handle) - conditions = list(manifest.get("conditions") or []) - if nevo_only: - # Just the stream NeVo would send: drop plain ReRF and the captured - # camera, which are the comparison rather than the output. - conditions = [c for c in conditions if c.get("threshold") is not None] - else: - conditions.append( - { - "name": "capture", - "prefix": "reference", - "label": f"captured camera {manifest['view']}", - "threshold": None, - "kept_fraction": None, - } - ) - manifest["conditions"] = [ - condition - for condition in conditions - if (directory / f"{condition['prefix']}_{manifest['frames'][0]:03d}.png").is_file() - ] - manifest["directory"] = directory - manifest["crop"] = None - return manifest - - -def subject_box(manifest: dict, pad: float = 0.12): - """Crop rectangle around the subject, unioned over the clip's frames. - - Same motivation as Vega's ``--subject-fill``: the corpus frames a whole - stage, so a 1280x960 view downscaled to a 420 px panel leaves the body a - thumbnail and the comparison unreadable. Taken from the *captured* frames - (background is pure white there) and applied identically to every panel, so - the conditions stay pixel-aligned. - """ - import numpy as np - from PIL import Image - - directory = manifest["directory"] - top, left = np.inf, np.inf - bottom, right = -np.inf, -np.inf - height = width = 0 - for frame in manifest["frames"]: - path = directory / f"reference_{frame:03d}.png" - if not path.is_file(): - continue - pixels = np.asarray(Image.open(path).convert("L")) - height, width = pixels.shape - rows = np.flatnonzero((pixels < 245).any(axis=1)) - columns = np.flatnonzero((pixels < 245).any(axis=0)) - if rows.size == 0 or columns.size == 0: - continue - top, bottom = min(top, rows[0]), max(bottom, rows[-1]) - left, right = min(left, columns[0]), max(right, columns[-1]) - if not np.isfinite(top): - return None - margin_y = (bottom - top) * pad - margin_x = (right - left) * pad - return ( - int(max(left - margin_x, 0)), - int(max(top - margin_y, 0)), - int(min(right + margin_x + 1, width)), - int(min(bottom + margin_y + 1, height)), - ) - - -def compose(manifest: dict, frame: int, panel_width: int): - """One row: every condition at this instant, captioned.""" - from PIL import Image, ImageDraw - - panels = [] - for condition in manifest["conditions"]: - path = manifest["directory"] / f"{condition['prefix']}_{frame:03d}.png" - image = Image.open(path).convert("RGB") - box = manifest.get("crop") - if box: - image = image.crop(box) - height = max(1, round(image.height * panel_width / image.width)) - panels.append((condition, image.resize((panel_width, height), Image.LANCZOS))) - - panel_height = max(image.height for _, image in panels) - sheet = Image.new( - "RGB", (panel_width * len(panels), panel_height + LABEL_HEIGHT), (17, 17, 17) - ) - draw = ImageDraw.Draw(sheet) - for index, (condition, image) in enumerate(panels): - left = index * panel_width - sheet.paste(image, (left, LABEL_HEIGHT)) - caption = condition["label"] - kept = condition.get("kept_fraction") - if condition.get("threshold") is not None and kept is not None: - caption += f" ({kept * 100:.0f}% of blocks)" - draw.text((left + 8, 7), caption, fill=(230, 230, 230)) - return sheet - - -def to_jpeg(image, quality: int) -> bytes: - buffer = io.BytesIO() - image.save(buffer, format="JPEG", quality=quality) - return buffer.getvalue() - - -def run(args) -> int: - root = Path(args.output_root).expanduser().resolve() - directories = ( - [root / args.object] - if args.object - else sorted(path.parent for path in root.glob("*/manifest.json")) - ) - clips = [ - load_clip(directory, nevo_only=args.nevo_only) - for directory in directories - if (directory / "manifest.json").is_file() - ] - if not args.no_crop: - for clip in clips: - clip["crop"] = subject_box(clip, args.crop_pad) - if not clips: - raise SystemExit( - f"no rendered output under {root}. Run orbitnevo.render_frames first." - ) - - names = ", ".join(clip["name"] for clip in clips) - total_frames = sum(len(clip["frames"]) for clip in clips) - reference = clips[0] - conditions = " | ".join(condition["label"] for condition in reference["conditions"]) - width = args.panel_width * len(reference["conditions"]) - probe = compose(reference, reference["frames"][0], args.panel_width) - height = probe.height - - frame_buffer = FrameBuffer() - - def status_html(): - return ( - "NeVo output" - "" - f"

NeVo — {names} (precomputed frames, looping at {args.fps} fps)

" - f"

{conditions}

" - f"

camera {reference['view']} · " - f"{reference['width']}x{reference['height']} · " - f"{total_frames} frames · rendered from ReRF feature voxels

" - # Cache-busting token, same reason as Vega's demo: a browser holding - # an open /stream from an earlier run will keep showing that clip. - f'' - "" - ) - - serve_forever(frame_buffer, args.port, status_html_fn=status_html) - print(f"serving on http://:{args.port}/", flush=True) - print(f"clips: {names}", flush=True) - - def loop(): - interval = 1.0 / max(args.fps, 1) - while True: - for clip in clips: - for frame in clip["frames"]: - started = time.time() - frame_buffer.update( - to_jpeg(compose(clip, frame, args.panel_width), args.quality) - ) - remaining = interval - (time.time() - started) - if remaining > 0: - time.sleep(remaining) - - thread = threading.Thread(target=loop, daemon=True) - thread.start() - try: - while True: - time.sleep(3600) - except KeyboardInterrupt: - print("stopped", flush=True) - return 0 - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--output-root", default="~/nevo_output") - parser.add_argument("--object", default="", help="one clip; default cycles through all") - parser.add_argument("--port", type=int, default=8752) - parser.add_argument("--fps", type=int, default=8) - parser.add_argument("--panel-width", type=int, default=420) - parser.add_argument("--quality", type=int, default=88) - parser.add_argument("--nevo-only", action="store_true", - help="stream only NeVo's visibility-filtered output, without the " - "plain-ReRF and captured-camera panels beside it") - parser.add_argument("--no-crop", action="store_true", - help="show the full frame instead of cropping to the subject") - parser.add_argument("--crop-pad", type=float, default=0.12) - return parser.parse_args(argv) - - -def main(argv=None) -> int: - return run(parse_args(argv)) - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/objects.py b/open4d/reconstruction/nevo/orbitnevo/objects.py deleted file mode 100644 index c76985cb..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/objects.py +++ /dev/null @@ -1,39 +0,0 @@ -"""Which ORBIT objects this baseline serves. - -Vendored from `baselines/objects.py` in the 4DVideoStreaming repository, reduced -to the one function `prepare.py` uses and with the dependency on that repo's -`vstream.config` made lazy. - -There, `vstream.config.OBJECTS` was the single source of truth for the scene the -comparison streamed, so every baseline read the same list and none could silently -serve a different scene. That package is not part of Open4D, so the fallback can -no longer be resolved here: passing `--objects` works exactly as before, and -omitting it now raises instead of quietly reading a list that does not exist. -""" -from __future__ import annotations - -from typing import Iterable - - -def configured_object_names() -> tuple[str, ...]: - """Names of the objects enabled in the comparison's shared scene config. - - Imported lazily and by name so that this module -- and therefore - `orbitnevo.prepare` -- imports without the 4DVideoStreaming `vstream` - package present. Only the no-`--objects` path needs it. - """ - try: - from vstream import config # type: ignore[import-not-found] - except ImportError as exc: - raise RuntimeError( - "the configured ORBIT scene lives in vstream/config.py in the " - "4DVideoStreaming repository, which is not vendored here. Pass " - "--objects to name the objects explicitly." - ) from exc - return tuple(spec.name for spec in config.OBJECTS) - - -def resolve_object_names(requested: Iterable[str] | None) -> tuple[str, ...]: - """The explicitly requested names if there are any, else the configured set.""" - names = tuple(requested or ()) - return names or configured_object_names() diff --git a/open4d/reconstruction/nevo/orbitnevo/prepare.py b/open4d/reconstruction/nevo/orbitnevo/prepare.py deleted file mode 100644 index 19f9074f..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/prepare.py +++ /dev/null @@ -1,594 +0,0 @@ -"""Turn an ORBIT object into a ReRF-trainable NHR corpus. - -Why this step exists: NeVo is a *streaming* system layered on ReRF, so it -needs ReRF feature-voxel sequences as input, and ReRF's own dataset is only -released after a signed licence agreement returned to ShanghaiTech -- see -``rerf/DATA.md``. The ORBIT corpus this repo already carries has to stand in, -which is what the NeVo paper itself did for two of its six datasets: "We -render the 8i and V-SENSE datasets' high-quality point clouds to images from -different viewports and use them to train NeRF videos" (section 5.1). - -Two sources: - -``--source gaussian`` (default) - ``ORBIT_datasets_gaussian``: the repo's prepared multi-view RGB corpus, 30 - frames x 8 calibrated views per object, already rendered over black with - exact OPENCV poses in ``transforms.json``. No rendering needed, and it is - the same input `baselines/Vega` trains on, so a NeVo-vs-Vega comparison is - over identical pixels. - -``--source mesh`` - Rasterise ORBIT's textured OBJ sequences directly with DeltaStream's - nvdiffrast renderer, on an arbitrary rig. The prepared corpus puts all 8 - of its views on one horizontal ring, which leaves a NeRF free to invent - geometry above and below the subject; this path can put 48 views on four - elevations instead. Better training data, but no longer the same pixels - the other baselines see. - -Either way the output is the layout ``rerf/lib/load_NHR.py`` reads -- what -upstream's ``data_util.py`` produces, minus the mp4 round-trip, since we hold -exact calibration already and have nothing to recover from ``CamPose.inf``: - - /image//img_%04d.png RGB over black, one per camera - /mask//img_%04d.png coverage mask, 0 or 255 - /cams_.json per-view extrinsic (4x4) + intrinsic - /bbox.json xyz_min / xyz_max, read by run.py - /nevo_corpus.json rig + world->normalised transform - -Runs in the repo's main environment (Python 3.10+), *not* the ``nevo`` conda -environment: it uses this repo's 3.10-only type syntax and, for -``--source mesh``, DeltaStream's nvdiffrast rasteriser. Everything downstream -of it runs under ``nevo``. See ``baselines/NeVo/README.md``. -""" -from __future__ import annotations - -import argparse -import dataclasses -import json -import os -import sys -from dataclasses import dataclass -from pathlib import Path - -import numpy as np -from PIL import Image - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo.cameras import NORMALISED_RADIUS, Camera, orbit_rig # noqa: E402 -from orbitnevo.objects import resolve_object_names # noqa: E402 - -DEFAULT_GAUSSIAN_ROOT = "/media/frozzzen/DataDrive/ORBIT_datasets_gaussian" -DEFAULT_MESH_ROOT = "/media/frozzzen/DataDrive/ORBIT_datasets" - -SILHOUETTE_THRESHOLD = 12 -"""Per-channel 0-255 level above which a pixel counts as foreground. - -The corpus renders over pure black, so the matte is exact rather than learned; -the threshold only rejects texture-filtering ringing at the silhouette. Same -value `baselines/Vega/vega/datasets/orbit_gaussian.py` carves with, so the two -baselines agree on where the subject ends.""" - -# dataset -> (file prefix, frame-number format, first frame on disk). Mirrors -# vstream.config.ALL_OBJECTS / scripts/orbit_datasets.sh, kept literal so this -# script does not have to import the server config to learn a filename. -MESH_SEQUENCES = { - "basketball": ("basketball_player", "fr%04d", 1), - "dancer": ("dancer", "fr%04d", 1), - "mitch": ("mitch", "fr%04d", 1), - "thomas": ("thomas", "fr%04d", 618), - "UMA0": ("UMA0", "%06d", 900), - "UMA1": ("UMA1", "%06d", 1800), - "UMA2": ("UMA2", "%06d", 3900), - "UMA3": ("UMA3", "%06d", 4900), - "UMA4": ("UMA4", "%06d", 900), -} -# The Gaussian corpus renamed the UMA objects; everything else matches. -GAUSSIAN_ALIASES = {f"UMA{i}": f"UMA{i}" for i in range(5)} - - -@dataclass(frozen=True) -class _RendererCamera: - """Duck-typed stand-in for DeltaStream's ``CameraCalibration``. - - ``gpu_renderer.CudaRgbdRenderer`` only reads ``camera_to_world`` and the - four intrinsic scalars, so adapting is cheaper than round-tripping our rig - through that dataclass's validation.""" - - camera_to_world: tuple - fx: float - fy: float - cx: float - cy: float - - -# ------------------------------------------------------------------- shared -def _write_png(path: Path, array: np.ndarray) -> None: - path.parent.mkdir(parents=True, exist_ok=True) - Image.fromarray(array).save(path, optimize=False, compress_level=1) - - -def _cams_json(out_dir: Path, frame: int, cameras: list[Camera], - centre: np.ndarray, scale: float, holdout: int = -1) -> None: - """Write the per-frame camera file ReRF trains from. - - ``holdout`` names a camera to omit. ReRF's NHR loader sets train = val = - test = every entry in this file (``lib/load_NHR.py``'s ``i_split``), so a - view listed here is a view the model fits; the only way to keep one back - for evaluation is to leave it out. Its image is still written to disk -- - it is the reference the renders get scored against. - """ - frames = [ - { - "file": str(out_dir / "image" / str(frame) / ("img_%04d.png" % camera.camera_id)), - "mask": str(out_dir / "mask" / str(frame) / ("img_%04d.png" % camera.camera_id)), - "extrinsic": camera.scaled_translation(centre, scale).tolist(), - "intrinsic": camera.intrinsic_matrix.tolist(), - } - for camera in cameras - if camera.camera_id != holdout - ] - with open(out_dir / ("cams_%d.json" % frame), "w") as handle: - json.dump({"frames": frames}, handle, indent=1) - - -def _write_bbox(out_dir: Path, lower: np.ndarray, upper: np.ndarray) -> None: - with open(out_dir / "bbox.json", "w") as handle: - json.dump({"xyz_min": lower.tolist(), "xyz_max": upper.tolist()}, handle, indent=1) - - -def _write_manifest(out_dir: Path, payload: dict) -> None: - with open(out_dir / "nevo_corpus.json", "w") as handle: - json.dump(payload, handle, indent=1) - print(f"wrote {out_dir}/nevo_corpus.json", flush=True) - - -def bbox_corners(lower: np.ndarray, upper: np.ndarray) -> np.ndarray: - return np.asarray( - [ - (lower[0] if x else upper[0], lower[1] if y else upper[1], lower[2] if z else upper[2]) - for x in (0, 1) - for y in (0, 1) - for z in (0, 1) - ] - ) - - -# ------------------------------------------------------- gaussian corpus path -def _load_transforms(root: Path, obj: str) -> dict: - path = root / obj / "transforms.json" - if not path.is_file(): - raise FileNotFoundError( - f"{path} not found. The Gaussian corpus is one transforms.json per object; " - f"available: {sorted(p.name for p in root.iterdir() if p.is_dir())}" - ) - with open(path) as handle: - return json.load(handle) - - -def _gaussian_rig(transforms: dict) -> tuple[list[np.ndarray], int]: - """One c2w per view id, checked to be static across the sequence. - - The corpus rig is fixed (``dataset.json`` describes a single 8-camera set), - but nothing in the file format enforces it -- and a rig that quietly moved - would train a NeRF on inconsistent geometry, so verify rather than assume. - """ - by_view: dict[int, np.ndarray] = {} - for entry in transforms["frames"]: - view = int(entry["view_id"]) - c2w = np.asarray(entry["camera_to_world_opencv"], dtype=np.float64) - if view not in by_view: - by_view[view] = c2w - elif not np.allclose(by_view[view], c2w, atol=1e-9): - raise ValueError( - f"view {view} moves during the sequence; this corpus is not a static rig" - ) - views = sorted(by_view) - if views != list(range(len(views))): - raise ValueError(f"view ids are not contiguous from 0: {views}") - return [by_view[view] for view in views], len(views) - - -def _project(c2w: np.ndarray, intrinsic: np.ndarray, points: np.ndarray) -> np.ndarray: - world_to_camera = np.linalg.inv(c2w) - local = points @ world_to_camera[:3, :3].T + world_to_camera[:3, 3] - if np.any(local[:, 2] <= 1e-6): - raise ValueError("the bounding box crosses the camera plane") - return (local / local[:, 2:3]) @ intrinsic.T - - -def _silhouette_bounds(path: Path) -> tuple[int, int, int, int] | None: - """Pixel bbox of the non-black pixels in one render, or None if empty.""" - with Image.open(path) as handle: - pixels = np.asarray(handle.convert("RGB"), dtype=np.uint8).max(axis=2) - rows = np.flatnonzero(pixels.max(axis=1) > SILHOUETTE_THRESHOLD) - columns = np.flatnonzero(pixels.max(axis=0) > SILHOUETTE_THRESHOLD) - if rows.size == 0 or columns.size == 0: - return None - return int(columns[0]), int(rows[0]), int(columns[-1]), int(rows[-1]) - - -def _crop_windows( - root: Path, - obj: str, - by_frame: dict, - frames: int, - view_count: int, - source_size: tuple[int, int], - aspect: float, - margin: float, -) -> list[tuple[int, int, int, int]]: - """One static crop window per view, at ``aspect``, framing the subject. - - Measured from the silhouettes rather than by projecting the bounding box. - The box is a loose cover -- its nearest corner projects to nearly the full - frame height on a camera 3 m from a 1.9 m subject, which clamps the crop - back to the whole image and defeats the point. The renders are over pure - black, so the true silhouette is exact and free to read. - - The window is the union over all frames of the sequence, so it is static: - the rig stays a fixed rig and the intrinsics stay constant across frames, - which is what `_gaussian_rig` has already checked for the extrinsics. - """ - width, height = source_size - unions: list[list[int] | None] = [None] * view_count - for frame in range(frames): - for view in range(view_count): - found = _silhouette_bounds(root / obj / by_frame[frame][view].lstrip("./")) - if found is None: - continue - if unions[view] is None: - unions[view] = list(found) - else: - current = unions[view] - current[0] = min(current[0], found[0]) - current[1] = min(current[1], found[1]) - current[2] = max(current[2], found[2]) - current[3] = max(current[3], found[3]) - windows = [] - for view, bounds in enumerate(unions): - if bounds is None: - raise ValueError(f"{obj} view {view} is empty in every frame") - left, top, right, bottom = bounds - centre_x = (left + right + 1) * 0.5 - centre_y = (top + bottom + 1) * 0.5 - crop_height = max((bottom - top + 1), (right - left + 1) / aspect) * margin - crop_height = min(crop_height, height, width / aspect) - crop_width = crop_height * aspect - origin_x = int(round(min(max(centre_x - crop_width * 0.5, 0.0), width - crop_width))) - origin_y = int(round(min(max(centre_y - crop_height * 0.5, 0.0), height - crop_height))) - windows.append((origin_x, origin_y, int(round(crop_width)), int(round(crop_height)))) - return windows - - -def prepare_gaussian(args, obj: str) -> dict: - root = Path(args.dataset_root).expanduser().resolve() - out_dir = Path(args.output_dir).expanduser().resolve() / obj - transforms = _load_transforms(root, GAUSSIAN_ALIASES.get(obj, obj)) - - source_width, source_height = int(transforms["w"]), int(transforms["h"]) - intrinsic = np.asarray( - ( - (transforms["fl_x"], 0.0, transforms["cx"]), - (0.0, transforms["fl_y"], transforms["cy"]), - (0.0, 0.0, 1.0), - ) - ) - rig, view_count = _gaussian_rig(transforms) - available = int(transforms["source_frame_count"]) - frames = min(args.frames, available) if args.frames > 0 else available - - lower = np.asarray(transforms["bounds_min"], dtype=np.float64) - upper = np.asarray(transforms["bounds_max"], dtype=np.float64) - pad = (upper - lower) * args.bbox_pad - lower, upper = lower - pad, upper + pad - centre = (lower + upper) * 0.5 - radius = float(np.mean([np.linalg.norm(c2w[:3, 3] - centre) for c2w in rig])) - scale = NORMALISED_RADIUS / radius - - by_frame: dict[int, dict[int, str]] = {} - for entry in transforms["frames"]: - by_frame.setdefault(int(entry["frame_index"]), {})[int(entry["view_id"])] = entry["file_path"] - - aspect = args.width / args.height - windows = _crop_windows( - root, GAUSSIAN_ALIASES.get(obj, obj), by_frame, frames, view_count, - (source_width, source_height), aspect, args.crop_margin, - ) - cameras: list[Camera] = [] - for view, c2w in enumerate(rig): - left, top, crop_width, crop_height = windows[view] - factor = args.width / crop_width - cameras.append( - Camera( - camera_id=view, - width=args.width, - height=args.height, - fx=float(intrinsic[0, 0]) * factor, - fy=float(intrinsic[1, 1]) * factor, - cx=(float(intrinsic[0, 2]) - left + 0.5) * factor - 0.5, - cy=(float(intrinsic[1, 2]) - top + 0.5) * factor - 0.5, - c2w=c2w, - ) - ) - print( - f"{obj}: {view_count} views, {frames} frames, crop " - f"{windows[0][2]}x{windows[0][3]} of {source_width}x{source_height} " - f"-> {args.width}x{args.height}, rig radius {radius:.4f} m", - flush=True, - ) - - out_dir.mkdir(parents=True, exist_ok=True) - _write_bbox(out_dir, (lower - centre) * scale, (upper - centre) * scale) - - coverage = [] - resample = Image.LANCZOS if args.width < windows[0][2] else Image.BICUBIC - for frame in range(frames): - image_dir = out_dir / "image" / str(frame) - mask_dir = out_dir / "mask" / str(frame) - done = image_dir / ("img_%04d.png" % (view_count - 1)) - if done.is_file() and not args.overwrite: - print(f"frame {frame}: already prepared", flush=True) - continue - covered = 0 - for view, (left, top, crop_width, crop_height) in enumerate(windows): - source = root / GAUSSIAN_ALIASES.get(obj, obj) / by_frame[frame][view].lstrip("./") - with Image.open(source) as handle: - cropped = handle.convert("RGB").resize( - (args.width, args.height), - resample, - box=(left, top, left + crop_width, top + crop_height), - ) - rgb = np.asarray(cropped, dtype=np.uint8) - mask = rgb.max(axis=2) > SILHOUETTE_THRESHOLD - # Zero the sub-threshold pixels so the RGB we hand ReRF is exactly - # the matte's inside: it composites rgb*alpha + (1-alpha) onto - # white, and leftover dark ringing outside the mask would survive - # as a grey halo. - rgb = np.where(mask[..., None], rgb, 0) - _write_png(image_dir / ("img_%04d.png" % view), rgb) - _write_png( - mask_dir / ("img_%04d.png" % view), - np.repeat(mask[..., None].astype(np.uint8) * 255, 3, axis=2), - ) - covered += int(mask.sum()) - _cams_json(out_dir, frame, cameras, centre, scale, args.holdout_view) - fraction = covered / (view_count * args.width * args.height) - coverage.append(fraction) - print(f"frame {frame}: {view_count} views, foreground {fraction * 100:.1f}%", flush=True) - - manifest = _manifest( - args, obj, cameras, centre, scale, lower, upper, frames, coverage, - source="ORBIT_datasets_gaussian", - extra={"crop_windows": [list(w) for w in windows], - "source_size": [source_width, source_height]}, - ) - _write_manifest(out_dir, manifest) - return manifest - - -# ----------------------------------------------------------- mesh render path -def obj_paths(mesh_root: Path, obj: str, start_frame: int, frames: int) -> list[Path]: - if obj not in MESH_SEQUENCES: - raise ValueError(f"unknown ORBIT object {obj!r}; known: {sorted(MESH_SEQUENCES)}") - prefix, token, first = MESH_SEQUENCES[obj] - start = start_frame if start_frame > 0 else first - if start < first: - raise ValueError(f"{obj} starts at source frame {first}, not {start}") - paths = [] - for index in range(frames): - path = mesh_root / obj / f"{prefix}_{token % (start + index)}.obj" - if not path.is_file(): - raise FileNotFoundError(f"missing ORBIT frame: {path}") - paths.append(path) - return paths - - -def obj_bounds(path: Path) -> tuple[np.ndarray, np.ndarray]: - """Bounding box of an OBJ's vertices, without building a mesh. - - Reading the `v ` lines directly is far cheaper than a trimesh load, and the - sequence-wide bbox is needed before any rendering can start. - """ - lower = np.full(3, np.inf) - upper = np.full(3, -np.inf) - with open(path, "r") as handle: - for line in handle: - if not line.startswith("v "): - continue - parts = line.split() - point = np.asarray((float(parts[1]), float(parts[2]), float(parts[3]))) - np.minimum(lower, point, out=lower) - np.maximum(upper, point, out=upper) - if not np.all(np.isfinite(lower)): - raise ValueError(f"no vertices in {path}") - return lower, upper - - -def sequence_bounds(paths: list[Path]) -> tuple[np.ndarray, np.ndarray]: - lower = np.full(3, np.inf) - upper = np.full(3, -np.inf) - for path in paths: - frame_lower, frame_upper = obj_bounds(path) - np.minimum(lower, frame_lower, out=lower) - np.maximum(upper, frame_upper, out=upper) - return lower, upper - - -def prepare_mesh(args, obj: str) -> dict: - mesh_root = Path(args.dataset_root).expanduser().resolve() - out_dir = Path(args.output_dir).expanduser().resolve() / obj - frames = args.frames if args.frames > 0 else 30 - paths = obj_paths(mesh_root, obj, args.start_frame, frames) - - print(f"{obj}: scanning bounds over {len(paths)} frames", flush=True) - lower, upper = sequence_bounds(paths) - pad = (upper - lower) * args.bbox_pad - lower, upper = lower - pad, upper + pad - - cameras, centre, scale = orbit_rig( - lower, upper, args.width, args.height, - azimuths=args.azimuths, elevations=tuple(args.elevations), - hfov_degrees=args.hfov, margin=args.framing_margin, - ) - print(f"{obj}: rig of {len(cameras)} cameras, radius scale {scale:.6f}", flush=True) - - out_dir.mkdir(parents=True, exist_ok=True) - _write_bbox(out_dir, (lower - centre) * scale, (upper - centre) * scale) - - # Imported late: creating the CUDA context before the bounds scan would - # hold GPU memory doing nothing for a minute. - from baselines.DeltaStream.orbitstream.converter import load_textured_mesh - from baselines.DeltaStream.orbitstream.gpu_renderer import CudaRgbdRenderer - - os.environ.setdefault("TORCH_CUDA_ARCH_LIST", "8.9") - renderer = CudaRgbdRenderer( - args.gpu, args.width, args.height, texture_filter="linear", supersample=args.supersample - ) - # Rasterise in the *normalised* frame: DeltaStream's renderer exports depth - # as uint16 millimetres and refuses a frame that overflows it, which a - # millimetre-unit world does immediately. Normalised, the extrinsics we - # rasterise with are exactly the ones written to cams_.json. - renderer_cameras = tuple( - _RendererCamera( - tuple(tuple(float(v) for v in row) for row in camera.scaled_translation(centre, scale)), - camera.fx, camera.fy, camera.cx, camera.cy, - ) - for camera in cameras - ) - - coverage = [] - for frame, path in enumerate(paths): - image_dir = out_dir / "image" / str(frame) - mask_dir = out_dir / "mask" / str(frame) - done = image_dir / ("img_%04d.png" % cameras[-1].camera_id) - if done.is_file() and not args.overwrite: - print(f"frame {frame}: already rendered", flush=True) - continue - mesh = load_textured_mesh(path) - mesh = dataclasses.replace( - mesh, vertices=(np.asarray(mesh.vertices, np.float64) - centre) * scale - ) - # Chunked: all views at supersampled resolution at once would want many - # GB of a card this box shares with other work. - covered = 0 - total = 0 - for begin in range(0, len(cameras), args.view_chunk): - chunk = cameras[begin : begin + args.view_chunk] - result = renderer.render(mesh, renderer_cameras[begin : begin + args.view_chunk]) - mask = result.depth_mm > 0 - if result.rgb.shape[0] != len(chunk): - raise RuntimeError("renderer returned the wrong number of views") - for offset, camera in enumerate(chunk): - _write_png(image_dir / ("img_%04d.png" % camera.camera_id), result.rgb[offset]) - _write_png( - mask_dir / ("img_%04d.png" % camera.camera_id), - np.repeat(mask[offset][..., None].astype(np.uint8) * 255, 3, axis=2), - ) - covered += int(mask.sum()) - total += int(mask.size) - _cams_json(out_dir, frame, cameras, centre, scale, args.holdout_view) - coverage.append(covered / total) - print(f"frame {frame}: {path.name} -> {len(cameras)} views, " - f"foreground {coverage[-1] * 100:.1f}%", flush=True) - - manifest = _manifest( - args, obj, cameras, centre, scale, lower, upper, len(paths), coverage, - source="ORBIT_datasets", - extra={"azimuths": args.azimuths, "elevations": list(args.elevations), - "hfov_degrees": args.hfov, "supersample": args.supersample}, - ) - _write_manifest(out_dir, manifest) - return manifest - - -def _manifest(args, obj, cameras, centre, scale, lower, upper, frames, coverage, - *, source: str, extra: dict) -> dict: - payload = { - "format": "nevo-rerf-nhr-corpus", - "version": 1, - "source": source, - "object": obj, - "frames": frames, - "width": args.width, - "height": args.height, - "world_centre": centre.tolist(), - "world_scale": scale, - "world_bounds_min": lower.tolist(), - "world_bounds_max": upper.tolist(), - "xyz_min": ((lower - centre) * scale).tolist(), - "xyz_max": ((upper - centre) * scale).tolist(), - "cameras": [ - { - "camera_id": camera.camera_id, - "fx": camera.fx, "fy": camera.fy, "cx": camera.cx, "cy": camera.cy, - "c2w_world": camera.c2w.tolist(), - "c2w_normalised": camera.scaled_translation(centre, scale).tolist(), - } - for camera in cameras - ], - "mean_foreground_fraction": float(np.mean(coverage)) if coverage else None, - "holdout_view": args.holdout_view if args.holdout_view >= 0 else None, - "training_views": [ - camera.camera_id for camera in cameras if camera.camera_id != args.holdout_view - ], - } - payload.update(extra) - return payload - - -# ---------------------------------------------------------------------- CLI -def prepare(args) -> list[dict]: - objects = resolve_object_names(args.objects) - worker = prepare_gaussian if args.source == "gaussian" else prepare_mesh - return [worker(args, obj) for obj in objects] - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--source", default="gaussian", choices=("gaussian", "mesh")) - parser.add_argument("--dataset-root", default=None, - help="defaults to the root for --source") - parser.add_argument("--objects", nargs="*", default=(), - help="ORBIT objects; defaults to the scene in vstream/config.py") - parser.add_argument("--output-dir", required=True, - help="one subdirectory is written per object") - parser.add_argument("--frames", type=int, default=0, help="0 = every frame available") - parser.add_argument("--width", type=int, default=1280) - parser.add_argument("--height", type=int, default=960, - help="4:3 keeps ReRF's half_res crop a no-op (it crops h x 4h/3)") - parser.add_argument("--bbox-pad", type=float, default=0.04) - parser.add_argument("--overwrite", action="store_true") - parser.add_argument("--holdout-view", type=int, default=-1, - help="camera id kept out of training and reserved as the evaluation " - "reference; -1 trains on every view") - # gaussian only - parser.add_argument("--crop-margin", type=float, default=1.15, - help="how much room to leave around the subject's projected bbox") - # mesh only - parser.add_argument("--start-frame", type=int, default=0, help="0 = the object's first") - parser.add_argument("--azimuths", type=int, default=12) - parser.add_argument("--elevations", type=float, nargs="+", default=[-15.0, 5.0, 25.0, 45.0]) - parser.add_argument("--hfov", type=float, default=60.0) - parser.add_argument("--framing-margin", type=float, default=1.12) - parser.add_argument("--supersample", type=int, default=2, choices=(1, 2, 4)) - parser.add_argument("--gpu", type=int, default=0) - parser.add_argument("--view-chunk", type=int, default=6, - help="views rasterised per renderer call; bounds GPU memory") - args = parser.parse_args(argv) - if args.dataset_root is None: - args.dataset_root = ( - DEFAULT_GAUSSIAN_ROOT if args.source == "gaussian" else DEFAULT_MESH_ROOT - ) - return args - - -def main(argv=None) -> int: - prepare(parse_args(argv)) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/rd_sweep.py b/open4d/reconstruction/nevo/orbitnevo/rd_sweep.py deleted file mode 100644 index d3d81c0b..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/rd_sweep.py +++ /dev/null @@ -1,243 +0,0 @@ -"""Rate-distortion sweep: what a ReRF frame costs, and what it looks like. - -Produces the numbers a comparison against another volumetric representation -(Vega, say) needs, on both axes at once: - -* **bytes** -- ReRF's own encoder run over the retained blocks, plus the block - mask and, on P-frames, the motion vectors. See ``nevo.bitstream`` for exactly - what is and is not counted. -* **quality** -- the delivered content rendered at a camera the model never - trained on, scored against that camera's captured image, on the subject's - bounding box. See ``nevo.metrics``. - -Sweeping the importance threshold traces one system's rate-distortion curve. -The unfiltered end (threshold 0) is plain ReRF; everything above it is NeVo's -visibility filtering. - - python -m orbitnevo.rd_sweep \\ - --config baselines/NeVo/rerf/configs/nevo/h_basketball.py \\ - --out ~/nevo_results/rd_h_basketball - -Writes ``results.csv`` (one row per threshold x frame, plus per-threshold -means), ``summary.json``, and ``renders/`` holding every rendered frame next to -the held-out reference. - -Runs in the ``nevo`` environment. -""" -from __future__ import annotations - -import argparse -import csv -import json -import sys -import time -from pathlib import Path - -import numpy as np - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import rerf_env # noqa: E402 - -DEFAULT_THRESHOLDS = (0.0, 0.01, 0.05, 0.15, 0.35) -"""Five points spanning unfiltered to aggressive. 0.0 keeps every occupied -block and is plain ReRF -- the anchor the rest of the curve is read against.""" - - -def _save_png(path: Path, image: np.ndarray) -> None: - from PIL import Image - - path.parent.mkdir(parents=True, exist_ok=True) - Image.fromarray((np.clip(image, 0.0, 1.0) * 255.0).astype(np.uint8)).save(path) - - -def run(args) -> dict: - from nevo.bitstream import SequenceCoder, startup_bytes - from nevo.blocks import BlockGrid - from nevo.filtering import voxels_dropped - from nevo.importance import ImportanceConfig, ImportanceScorer - from nevo.metrics import QualityScorer, silhouette_box - from nevo.render import held_out_view, render_view - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(args.config) - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - holdout = manifest.get("holdout_view") - training_views = manifest.get("training_views") or [ - camera["camera_id"] for camera in manifest["cameras"] - ] - if args.view >= 0: - eval_view = args.view - elif holdout is not None: - eval_view = holdout - else: - # Corpora prepared before --holdout-view existed trained on every - # camera. Still usable for producing comparable renders and byte - # counts, but the quality figures flatter the model, so say so loudly - # here and carry the fact into the summary and the page. - eval_view = int(manifest["cameras"][-1]["camera_id"]) - held_out = eval_view not in training_views - if not held_out: - print( - f"WARNING: camera {eval_view} is in the training set " - f"{training_views}. Quality numbers are optimistic; bytes are unaffected.", - flush=True, - ) - trained = sequence.available_frames() - if not trained: - raise RuntimeError(f"no trained frames under {sequence.run_dir}") - frames = trained[: args.frames] if args.frames > 0 else trained - thresholds = list(args.thresholds) - - out_dir = Path(args.out).expanduser().resolve() - renders = out_dir / "renders" - renders.mkdir(parents=True, exist_ok=True) - print( - f"{sequence.cfg.expname}: {len(frames)} frames x {len(thresholds)} thresholds, " - f"camera {eval_view} ({'held out' if held_out else 'IN TRAINING SET'}), " - f"trained on {training_views}", - flush=True, - ) - - # One crop for the whole sweep, so every number is over the same pixels. - references = {} - mattes = [] - for frame_index in frames: - camera, truth = held_out_view(sequence, frame_index, view=eval_view) - references[frame_index] = (camera, truth) - mattes.append((truth < 0.999).any(axis=2)) - box = silhouette_box(mattes, pad=args.box_pad) - print(f"scoring box (top, left, bottom, right) = {box} of " - f"{mattes[0].shape[1]}x{mattes[0].shape[0]}", flush=True) - - scorer = QualityScorer(net=args.lpips_net) - rows = [] - with SequenceCoder(sequence, quality=args.quality) as coder: - for frame_index in frames: - started = time.time() - camera, truth = references[frame_index] - frame = sequence.frame(frame_index) - grid = BlockGrid(frame.grid_shape, args.block_size) - importance = ImportanceScorer( - sequence, frame, ImportanceConfig(block_size=args.block_size) - ) - # Scored at the camera being rendered: the filter gets a perfectly - # predicted viewport, which is the optimistic end of the design. - scores = importance.score(camera) - price = coder.advance(frame) - - if frame_index == frames[0]: - _save_png(renders / f"reference_f{frame_index:03d}.png", truth) - - for threshold in thresholds: - keep = (scores >= threshold) & importance.occupancy - cost = price(keep) - with voxels_dropped(frame, grid, keep): - rendered = render_view(sequence, frame, camera) - quality = scorer.score(rendered, truth, box) - _save_png( - renders / f"t{threshold}_f{frame_index:03d}.png", rendered - ) - if not args.keep_reference_once: - _save_png(renders / f"reference_f{frame_index:03d}.png", truth) - row = {"threshold": threshold, **cost.as_dict(), **quality.as_dict()} - rows.append(row) - print( - f" t={threshold:<5} frame {frame_index:3d} " - f"{cost.total_bytes / 1024:8.1f} kB " - f"kept {cost.kept_blocks:5d}/{cost.occupied_blocks:5d} " - f"PSNR {quality.psnr:5.2f} SSIM {quality.ssim:.4f} " - f"LPIPS {quality.lpips:.4f}", - flush=True, - ) - print(f"frame {frame_index} done in {time.time() - started:.0f}s", flush=True) - del frame, importance - - fields = list(rows[0].keys()) - with open(out_dir / "results.csv", "w", newline="") as handle: - writer = csv.DictWriter(handle, fieldnames=fields) - writer.writeheader() - writer.writerows(rows) - - per_threshold = [] - for threshold in thresholds: - subset = [row for row in rows if row["threshold"] == threshold] - per_threshold.append( - { - "threshold": threshold, - "frames": len(subset), - "bytes_per_frame": float(np.mean([r["total_bytes"] for r in subset])), - "feature_bytes_per_frame": float(np.mean([r["feature_bytes"] for r in subset])), - "mask_bytes_per_frame": float(np.mean([r["mask_bytes"] for r in subset])), - "motion_bytes_per_frame": float(np.mean([r["motion_bytes"] for r in subset])), - "kept_fraction": float(np.mean([r["kept_fraction"] for r in subset])), - "psnr": float(np.mean([r["psnr"] for r in subset])), - "ssim": float(np.mean([r["ssim"] for r in subset])), - "lpips": float(np.mean([r["lpips"] for r in subset])), - "mbps_at_30fps": float( - np.mean([r["total_bytes"] for r in subset]) * 8 * 30 / 1e6 - ), - } - ) - entry = per_threshold[-1] - print( - f"[t={entry['threshold']}] {entry['bytes_per_frame'] / 1024:8.1f} kB/frame " - f"({entry['mbps_at_30fps']:6.1f} Mbps @30fps) kept {entry['kept_fraction'] * 100:5.1f}% " - f"PSNR {entry['psnr']:5.2f} SSIM {entry['ssim']:.4f} LPIPS {entry['lpips']:.4f}", - flush=True, - ) - - summary = { - "object": sequence.cfg.expname, - "config": str(sequence.config_path), - "corpus": str(sequence.corpus_dir), - "eval_view": eval_view, - "held_out": held_out, - "holdout_view": holdout, - "training_views": training_views, - "frames": frames, - "block_size": args.block_size, - "quality": args.quality, - "render_size": [references[frames[0]][0].width, references[frames[0]][0].height], - "scoring_box_tlbr": list(box), - "lpips_net": args.lpips_net, - "startup_bytes": startup_bytes(sequence), - "per_threshold": per_threshold, - } - with open(out_dir / "summary.json", "w") as handle: - json.dump(summary, handle, indent=1) - print(f"wrote {out_dir}/results.csv and summary.json", flush=True) - return summary - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--config", required=True) - parser.add_argument("--out", required=True) - parser.add_argument("--frames", type=int, default=0, help="0 = every trained frame") - parser.add_argument("--thresholds", type=float, nargs="+", default=list(DEFAULT_THRESHOLDS)) - parser.add_argument("--block-size", type=int, default=8) - parser.add_argument("--quality", type=int, default=99, - help="ReRF codec quality; 99 is compress.py's default") - parser.add_argument("--box-pad", type=int, default=8) - parser.add_argument("--lpips-net", default="alex", choices=("alex", "vgg")) - parser.add_argument("--view", type=int, default=-1, - help="camera to render and score at; default is the corpus's " - "held-out camera, or the last one if none was held out") - parser.add_argument("--keep-reference-once", action="store_true", - help="write the reference image only for the first frame") - return parser.parse_args(argv) - - -def main(argv=None) -> int: - args = parse_args(argv) - rerf_env.activate() - run(args) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/render_frames.py b/open4d/reconstruction/nevo/orbitnevo/render_frames.py deleted file mode 100644 index aa4b3448..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/render_frames.py +++ /dev/null @@ -1,190 +0,0 @@ -"""Render a trained ReRF sequence to images, with and without NeVo's filtering. - -Two conditions come out of one pass: - -``rerf`` - The whole feature voxel grid rendered as trained. This is the system NeVo - is measured against, not NeVo. -``nevo`` - The same frame with every feature voxel whose neural visibility falls below - ``t`` removed -- NeVo's section 3.2, the visibility-aware optimisation. A - dropped block is written back the way ReRF's decoder fills a block that - never arrived (raw density -4.1, zero features), not zeroed. - -Those are the artefacts another representation's output gets compared against, -so the things that must match on both sides are what ``manifest.json`` records: -camera extrinsics and intrinsics, resolution, and the white background the -corpus composites onto. The captured image from the same camera is written -alongside. - - python -m orbitnevo.render_frames \\ - --config baselines/NeVo/rerf/configs/nevo/g_basketball.py \\ - --out ~/nevo_output --view 7 - -Two things this does *not* do, both from the paper's section 3.2 and both -noted in RESULTS.md: the threshold is passed in rather than fitted per video -against an SSIM target, and the selected set is not dilated by 20 cm to absorb -viewport-prediction error. Filtering here therefore sees a perfectly predicted -viewport, which is the optimistic end of the design. - -Runs in the ``nevo`` environment. -""" -from __future__ import annotations - -import argparse -import json -import sys -import time -from pathlib import Path - -import numpy as np - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import rerf_env # noqa: E402 - -PAPER_THRESHOLD = 0.025 -"""The importance threshold NeVo's section 3.2 quotes as its worked example.""" - - -def _save(path: Path, image: np.ndarray) -> None: - from PIL import Image - - path.parent.mkdir(parents=True, exist_ok=True) - Image.fromarray((np.clip(image, 0.0, 1.0) * 255.0).astype(np.uint8)).save(path) - - -def run(args) -> dict: - from nevo.blocks import BlockGrid - from nevo.filtering import voxels_dropped - from nevo.importance import ImportanceConfig, ImportanceScorer - from nevo.render import held_out_view, render_view - from nevo.sequence import ReRFSequence - - sequence = ReRFSequence(args.config) - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - frames = sequence.available_frames() - if not frames: - raise RuntimeError(f"no trained frames under {sequence.run_dir}") - if args.frames > 0: - frames = frames[: args.frames] - - view = args.view if args.view >= 0 else int(manifest["cameras"][-1]["camera_id"]) - training_views = manifest.get("training_views") or [ - camera["camera_id"] for camera in manifest["cameras"] - ] - out_dir = Path(args.out).expanduser().resolve() / sequence.cfg.expname - out_dir.mkdir(parents=True, exist_ok=True) - thresholds = list(args.thresholds) - - print( - f"{sequence.cfg.expname}: {len(frames)} frames at camera {view}, " - f"conditions: rerf + {['nevo@%s' % t for t in thresholds]}", - flush=True, - ) - started = time.time() - camera = None - kept = {threshold: [] for threshold in thresholds} - for frame_index in frames: - camera, truth = held_out_view(sequence, frame_index, view=view) - frame = sequence.frame(frame_index) - _save(out_dir / f"frame_{frame_index:03d}.png", render_view(sequence, frame, camera)) - _save(out_dir / f"reference_{frame_index:03d}.png", truth) - - if thresholds: - scorer = ImportanceScorer( - sequence, frame, ImportanceConfig(block_size=args.block_size) - ) - grid = BlockGrid(frame.grid_shape, args.block_size) - scores = scorer.score(camera) - occupied = int(scorer.occupancy.sum().item()) - for threshold in thresholds: - keep = (scores >= threshold) & scorer.occupancy - with voxels_dropped(frame, grid, keep): - rendered = render_view(sequence, frame, camera) - _save(out_dir / f"nevo{threshold}_{frame_index:03d}.png", rendered) - kept[threshold].append(int(keep.sum().item()) / max(occupied, 1)) - del scorer - print(f" frame {frame_index:3d}", flush=True) - del frame - elapsed = time.time() - started - - conditions = [ - { - "name": "rerf", - "prefix": "frame", - "label": "ReRF (all feature voxels)", - "threshold": None, - "kept_fraction": 1.0, - } - ] - for threshold in thresholds: - conditions.append( - { - "name": f"nevo{threshold}", - "prefix": f"nevo{threshold}", - "label": f"NeVo (visibility-filtered, t={threshold})", - "threshold": threshold, - "kept_fraction": float(np.mean(kept[threshold])), - } - ) - - payload = { - "name": sequence.cfg.expname, - "representation": "ReRF (NeRF feature voxels)", - "frames": frames, - "view": view, - "view_in_training_set": view in training_views, - "training_views": training_views, - "width": camera.width, - "height": camera.height, - "intrinsics": {"fx": camera.fx, "fy": camera.fy, "cx": camera.cx, "cy": camera.cy}, - "c2w": camera.c2w.tolist(), - "background": "white (corpus composites rgb*alpha + (1-alpha))", - "block_size": args.block_size, - "conditions": conditions, - "corpus": str(sequence.corpus_dir), - "run_dir": str(sequence.run_dir), - "seconds": elapsed, - } - with open(out_dir / "manifest.json", "w") as handle: - json.dump(payload, handle, indent=1) - for condition in conditions: - print( - f" {condition['label']}: {condition['kept_fraction'] * 100:.1f}% of " - f"non-empty blocks kept", - flush=True, - ) - print(f"wrote {out_dir} ({elapsed:.0f}s)", flush=True) - return payload - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--config", nargs="+", required=True) - parser.add_argument("--out", default="~/nevo_output") - parser.add_argument("--view", type=int, default=-1, - help="camera to render from; default is the corpus's last") - parser.add_argument("--frames", type=int, default=0, help="0 = every trained frame") - parser.add_argument("--thresholds", type=float, nargs="*", default=[PAPER_THRESHOLD], - help="one NeVo condition per importance threshold; empty = ReRF only") - parser.add_argument("--block-size", type=int, default=8, - help="filtering unit; 8 is ReRF's codec block") - return parser.parse_args(argv) - - -def main(argv=None) -> int: - args = parse_args(argv) - rerf_env.activate() - configs = list(args.config) - for config in configs: - args.config = config - run(args) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/report.py b/open4d/reconstruction/nevo/orbitnevo/report.py deleted file mode 100644 index 859cc95f..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/report.py +++ /dev/null @@ -1,1317 +0,0 @@ -"""Build a static page for eyeballing everything this baseline produces. - -Numbers in a JSON file are easy to be wrong about quietly. This renders the -things worth looking at directly: - -1. **Corpus** -- the prepared views and their mattes, per object and frame. - Catches a bad crop, an inverted mask, a mis-ordered rig. -2. **Reload** -- a trained frame rendered at a training viewpoint, beside the - image it was trained on. Catches a checkpoint reassembled wrongly. -3. **Filtering** -- the same viewpoint rendered with the whole feature grid and - with everything below an importance threshold dropped, plus the amplified - difference. This is what the CDF is actually claiming. -4. **Importance CDF** -- the Figure 7 plots and the numbers behind them. -5. **Rate-distortion** -- what a frame actually costs in ReRF's own encoder at - each threshold, against the quality delivered at the captured camera. The - axis a comparison against another representation is made on. - - python -m orbitnevo.report --out ~/nevo_report --serve 8752 - -Runs in the ``nevo`` environment; sections 2 and 3 need a trained sequence and -a GPU and are skipped (with a note on the page) when there is none. -""" -from __future__ import annotations - -import argparse -import csv -import dataclasses -import datetime -import glob -import html -import json -import re -import shutil -import sys -from pathlib import Path - -import numpy as np - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import rerf_env # noqa: E402 - -THUMBNAIL_WIDTH = 260 -VIEW_WIDTH = 420 - - -def _save(image, path: Path, width: int) -> str: - """Write a JPEG scaled to ``width`` and return its path relative to the page.""" - from PIL import Image - - if isinstance(image, np.ndarray): - array = np.clip(image, 0.0, 1.0) if image.dtype != np.uint8 else image - if array.dtype != np.uint8: - array = (array * 255.0).astype(np.uint8) - handle = Image.fromarray(array) - else: - handle = image - handle = handle.convert("RGB") - height = max(1, round(handle.height * width / handle.width)) - handle = handle.resize((width, height), Image.LANCZOS) - path.parent.mkdir(parents=True, exist_ok=True) - handle.save(path, quality=88) - # Relative to the page, not to the assets directory: index.html sits beside - # `assets/`, so a bare filename resolves to the wrong place and every - # thumbnail renders as a broken image. - for parent in path.parents: - if parent.name == "assets": - return str(path.relative_to(parent.parent)) - return path.name - - -def _pick(values: list, count: int) -> list: - """An evenly spaced subset of ``values``, always including the first. - - Used to thin the corpus contact sheet. A 48-camera rig over two frames over - a dozen objects is several hundred thumbnails, which is not a page anyone - can read -- and the point of the section is to spot a bad crop or an - inverted matte, which a spread of six views shows as well as all of them. - """ - if count <= 0 or len(values) <= count: - return list(values) - step = len(values) / count - return [values[int(index * step)] for index in range(count)] - - -def _output_clips(output_root, assets: Path, width: int) -> list: - """NeVo's own output, as player clips. - - Only the visibility-filtered conditions ``render_frames.py`` produced -- the - stream NeVo would actually send. Plain ReRF and the captured camera are - deliberately left out: they are the *comparison*, and they are already shown - frame-by-frame with their numbers in the rate-distortion section, which is - the honest place for a comparison. The player is for watching the output. - """ - from PIL import Image - - clips = [] - for root in output_root: - for manifest_path in sorted(Path(root).expanduser().glob("*/manifest.json")): - with open(manifest_path) as handle: - manifest = json.load(handle) - directory = manifest_path.parent - name = manifest["name"] - conditions = list(manifest.get("conditions") or [ - {"name": "render", "prefix": "frame", "label": "reconstruction", - "threshold": None, "kept_fraction": 1.0} - ]) - conditions = [c for c in conditions if c.get("threshold") is not None] - for condition in conditions: - target = assets / "play" / f"{name}_{condition['name']}" - frames = [] - for frame in manifest["frames"]: - source = directory / f"{condition['prefix']}_{frame:03d}.png" - if not source.is_file(): - continue - out = target / "rgb" / f"v00_f{frame:03d}.jpg" - if not out.is_file(): - _save(Image.open(source), out, width) - frames.append(frame) - if not frames: - continue - kept = condition.get("kept_fraction") - # Plain text, not markup: the player writes this with - # textContent, so an HTML entity would show up verbatim. - detail = ( - f" \u2014 {kept * 100:.0f}% of non-empty blocks delivered" - if kept is not None and condition.get("threshold") is not None - else "" - ) - clips.append( - { - "name": f"{name} \u2014 {condition['label']}", - "path": f"assets/play/{name}_{condition['name']}", - "frames": frames, - "views": [0], - "source": f"camera {manifest['view']}{detail}", - "size": [manifest["width"], manifest["height"]], - "no_matte": True, - } - ) - return clips - - -def _player_section(corpora, assets: Path, views_shown: int, width: int, - output_clips=()) -> str: - """A plain playback of the prepared corpus: pick an object and a camera, press play. - - This is an *input* check. A still tells you a crop is centred; only motion - tells you it is centred on every frame, that the sequence is in order, and - that the matte does not flicker. Nothing here is rendered or trained -- - these are the exact PNGs handed to ReRF, scaled down. - """ - from PIL import Image - - clips = list(output_clips) - if not corpora and not clips: - return "

Nothing rendered or prepared yet.

" - - for corpus in corpora: - with open(corpus / "nevo_corpus.json") as handle: - manifest = json.load(handle) - name = manifest["object"] - frames = sorted( - int(path.name) for path in (corpus / "image").iterdir() if path.name.isdigit() - ) - views = _pick(list(range(len(manifest["cameras"]))), views_shown) - target = assets / "play" / name - for view in views: - for frame in frames: - for kind, folder in (("image", "rgb"), ("mask", "matte")): - source = corpus / kind / str(frame) / ("img_%04d.png" % view) - if not source.is_file(): - continue - out = target / folder / f"v{view:02d}_f{frame:03d}.jpg" - if not out.is_file(): - _save(Image.open(source), out, width) - clips.append( - { - "name": name, - "path": f"assets/play/{name}", - "frames": frames, - "views": views, - "source": manifest["source"], - "size": [manifest["width"], manifest["height"]], - } - ) - - options = "".join( - f"" - for clip in clips - ) - return ( - "

Pick a clip and press play. <name> — render is " - "NeVo/ReRF's reconstruction, one rendered frame per input frame -- that is the output " - "to compare against another representation. — capture is the real camera " - "for the same frames. The bare object names are the prepared input ReRF trained on.

" - "
" - "
corpus playback" - "
loading…
" - "
" - "
" - "" - f"" - "" - "" - "" - "
" - "
" - "" - "
" - "

 

" - "
" - f"" - ) - - -def _corpus_section(corpora: list[Path], assets: Path, frames_shown: int, - views_shown: int) -> str: - from PIL import Image - - if not corpora: - return "

No prepared corpus found.

" - blocks = [] - for corpus in corpora: - with open(corpus / "nevo_corpus.json") as handle: - manifest = json.load(handle) - # Two corpus roots can hold the same object (the prepared multi-view - # set and a mesh re-render of it), so the asset names have to carry the - # root or one silently overwrites the other's thumbnails. - name = manifest["object"] - slug = f"{corpus.parent.name}_{name}" - available = sorted( - int(p.name) for p in (corpus / "image").iterdir() if p.name.isdigit() - ) - picked = _pick(available, frames_shown) - views = _pick(list(range(len(manifest["cameras"]))), views_shown) - rows = [] - for frame in picked: - cells = [] - for view in views: - image = corpus / "image" / str(frame) / ("img_%04d.png" % view) - mask = corpus / "mask" / str(frame) / ("img_%04d.png" % view) - if not image.is_file(): - continue - rgb = _save(Image.open(image), assets / f"{slug}_f{frame}_v{view}.jpg", - THUMBNAIL_WIDTH) - alpha = _save(Image.open(mask), assets / f"{slug}_f{frame}_v{view}_m.jpg", - THUMBNAIL_WIDTH) - cells.append( - f"
" - f"" - f"
view {view}
" - ) - rows.append(f"

frame {frame}

{''.join(cells)}
") - crop = manifest.get("crop_windows") - detail = ( - f"crop {crop[0][2]}x{crop[0][3]} of {manifest['source_size'][0]}x" - f"{manifest['source_size'][1]}" - if crop - else f"{manifest.get('azimuths', '?')} azimuths x " - f"{len(manifest.get('elevations', []))} elevations" - ) - total_views = len(manifest["cameras"]) - subset = "" if len(views) == total_views else ( - f" · showing {len(views)} of {total_views} views" - ) - blocks.append( - f"

{html.escape(name)} " - f"({html.escape(corpus.parent.name)})

" - f"

{manifest['source']} · {total_views} views " - f"· {manifest['frames']} frames · {detail} → " - f"{manifest['width']}x{manifest['height']} · foreground " - f"{(manifest['mean_foreground_fraction'] or 0) * 100:.1f}%{subset}

" - f"{''.join(rows)}
" - ) - return ( - "

Hover a thumbnail to swap between the render and its matte. " - "The matte should hug the subject with no halo and no holes.

" - + "".join(blocks) - ) - - -def _reload_and_filter_section(runs: list[Path], assets: Path, args) -> tuple[str, str]: - from nevo.blocks import BlockGrid - from nevo.filtering import preview - from nevo.importance import ImportanceConfig, ImportanceScorer - from nevo.cameras import look_at_c2w - from nevo.render import psnr, render_view, training_view - from nevo.sequence import ReRFSequence - - reload_blocks = [] - filter_blocks = [] - for config_path in runs: - try: - sequence = ReRFSequence(config_path) - trained = sequence.available_frames() - except Exception as error: # a half-written run should not sink the page - reload_blocks.append( - f"

{html.escape(config_path.stem)}

" - f"

could not open: {html.escape(str(error))}

" - ) - continue - if not trained: - continue - name = sequence.cfg.expname - for frame_index in trained[: args.frames_checked]: - frame = sequence.frame(frame_index) - kind = "I" if frame.is_key_frame else "P" - camera, truth = training_view(sequence, frame_index, args.view) - rendered = render_view(sequence, frame, camera) - score = psnr(rendered, truth) - left = _save(rendered, assets / f"{name}_reload_{frame_index}_r.jpg", VIEW_WIDTH) - right = _save(truth, assets / f"{name}_reload_{frame_index}_t.jpg", VIEW_WIDTH) - reload_blocks.append( - f"

{html.escape(name)} · frame " - f"{frame_index} ({kind})

" - f"

rendered from the reloaded checkpoint vs. the training " - f"image · {score:.2f} dB

" - f"
" - f"
reloaded
" - f"
ground truth
" - f"
" - ) - - if frame_index != trained[0]: - continue - - # A training view says nothing about reconstruction quality -- the - # model was fit to it. Render a viewpoint the rig never saw, halfway - # between two of its cameras, so the cost of a sparse rig (this - # corpus has 8 views on one horizontal ring) is visible rather than - # implied. - novel = _between_rig_cameras(sequence, camera, look_at_c2w) - if novel is not None: - path = _save( - render_view(sequence, frame, novel), - assets / f"{name}_novel.jpg", - VIEW_WIDTH, - ) - reload_blocks.append( - f"

{html.escape(name)} · novel view" - f"

a viewpoint midway between two rig cameras, " - f"which the model never saw. Floaters and smeared geometry show up " - f"here, not in the training view above.

" - f"
" - f"
novel view
" - ) - scorer = ImportanceScorer( - sequence, frame, ImportanceConfig(block_size=args.block_size) - ) - scores = scorer.score(camera) - grid = BlockGrid(frame.grid_shape, args.block_size) - cards = [] - for threshold in args.thresholds: - result = preview( - sequence, frame, camera, scores, grid, threshold, scorer.occupancy - ) - full = _save(result.full, assets / f"{name}_filter_full.jpg", VIEW_WIDTH) - filtered = _save( - result.filtered, - assets / f"{name}_filter_{threshold}.jpg", - VIEW_WIDTH, - ) - # The difference is invisible at 1x when the filter is working, - # which is the point -- amplify it so it can be judged. - amplified = _save( - np.clip(result.difference * args.difference_gain, 0.0, 1.0), - assets / f"{name}_filter_{threshold}_d.jpg", - VIEW_WIDTH, - ) - cards.append( - f"

threshold {threshold} " - f"· dropped {result.dropped_fraction * 100:.1f}% of " - f"{result.total_blocks} non-empty blocks · " - f"SSIM {result.ssim:.4f} · {result.psnr:.2f} dB

" - f"
" - f"
all voxels
" - f"
filtered
" - f"
" - f"
difference x{args.difference_gain}
" - f"
" - ) - filter_blocks.append( - f"

{html.escape(name)} · frame " - f"{frame_index}, {args.block_size}3 blocks

" - + "".join(cards) - + "
" - ) - reload_html = "".join(reload_blocks) or "

No trained sequence found.

" - filter_html = "".join(filter_blocks) or "

No trained sequence found.

" - return reload_html, filter_html - - -def _sweep_section(results: list[Path]) -> str: - """Table of the threshold sweep: what each threshold drops, and what it costs.""" - rows = [] - for path in sorted(results): - with open(path) as handle: - report = json.load(handle) - for row in report["rows"]: - verdict = ( - "lossless" if row["visually_lossless"] - else "degraded" - ) - rows.append( - "" - f"{html.escape(report['object'])}" - f"{report['block_size']}3" - f"{row['threshold']}" - f"{row['dropped_mean'] * 100:.1f}%" - f"{row['ssim_mean']:.4f}" - f"{row['ssim_min']:.4f}" - f"{row['psnr_mean']:.1f}" - f"{verdict}" - ) - if not rows: - return "

No threshold sweep found.

" - return ( - "

Each row filters with that viewport's own scores and renders, " - "against the same viewport rendered from the whole grid. lossless means the " - "worst viewport still cleared SSIM 0.98, the bar the paper cites. The paper " - "fits its threshold to that bar rather than fixing it at 0.025, so the row that " - "matters is the last lossless one.

" - "" - "" - + "".join(rows) - + "
objectblockthresholddroppedSSIMworst SSIMPSNR
" - ) - - -def _rd_per_frame(directory: Path) -> dict: - """``{(threshold, frame): (psnr, ssim)}`` from the sweep's per-row CSV.""" - path = directory / "results.csv" - if not path.exists(): - return {} - scores = {} - with open(path) as handle: - for row in csv.DictReader(handle): - key = (float(row["threshold"]), int(row["frame"])) - scores[key] = (float(row["psnr"]), float(row["ssim"])) - return scores - - -def _rd_strip(directory: Path, summary: dict, assets: Path, height: int = 340) -> str: - """Two strips: how well the frame reconstructs, then what filtering costs it. - - Both are cropped to ``scoring_box_tlbr`` -- the same subject bounding box the - PSNR/SSIM/LPIPS in the table were computed over. At full frame the subject is - 5-9% of the pixels, so an uncropped thumbnail shows neither the artefacts nor - the region the numbers describe. - - The first strip pairs frame 0 with frame 1 deliberately. Frame 0 is the - I-frame; frame 1 is the first P-frame, which codes a residual over a - motion-compensated predecessor and so fails differently -- on a held-out - camera, much harder. - """ - from PIL import Image - - renders = directory / "renders" - if not renders.is_dir(): - return "" - scores = _rd_per_frame(directory) - box = summary.get("scoring_box_tlbr") - - def cell(path: Path, key: str, label: str, frame: int, threshold=None) -> str: - image = Image.open(path) - if box: - top, left, bottom, right = box - pad = 12 - image = image.crop((max(left - pad, 0), max(top - pad, 0), - min(right + pad, image.width), - min(bottom + pad, image.height))) - width = max(1, round(image.width * height / image.height)) - # The filename comes from ``key``, never from ``label``: a label carries - # spaces, "%" and "=", and a "%" in a URL path is an escape prefix. - slug = re.sub(r"[^a-z0-9]+", "_", key.lower()).strip("_") - src = _save(image, assets / f"{summary['object']}_rd_f{frame}_{slug}.jpg", width) - if threshold is not None and (threshold, frame) in scores: - psnr, ssim = scores[(threshold, frame)] - label = f"{label} · {psnr:.1f} dB / {ssim:.3f}" - return (f"
{html.escape(label)}" - f"
{html.escape(label)}
") - - frames = (summary.get("frames") or [])[:2] - thresholds = [row["threshold"] for row in summary["per_threshold"]] - base = thresholds[0] if thresholds else 0.0 - out = [] - - cells = [] - for frame in frames: - reference = renders / f"reference_f{frame:03d}.png" - rendered = renders / f"t{base}_f{frame:03d}.png" - kind = "I-frame" if frame == 0 else "P-frame" - if reference.exists(): - cells.append(cell(reference, "captured", f"captured · f{frame} ({kind})", frame)) - if rendered.exists(): - cells.append( - cell(rendered, "rerf", f"ReRF · f{frame} ({kind})", frame, base) - ) - if cells: - out.append("

Reconstruction — unfiltered ReRF against the camera

" - "
" + "".join(cells) + "
") - - frame = frames[0] if frames else 0 - dropped = {row["threshold"]: 1.0 - row["kept_fraction"] for row in summary["per_threshold"]} - cells = [ - cell(renders / f"t{threshold}_f{frame:03d}.png", - f"t{threshold}", - f"t={threshold} · {dropped[threshold] * 100:.0f}% dropped", - frame, threshold) - for threshold in thresholds - if (renders / f"t{threshold}_f{frame:03d}.png").exists() - ] - if cells: - out.append(f"

Filtering — frame {frame} at each threshold

" - "
" + "".join(cells) + "
") - return "".join(out) - - -def _rd_section(results_roots, assets: Path) -> str: - """Rate-distortion: real ReRF bytes on one axis, delivered quality on the other.""" - summaries = [] - for root in results_roots: - for path in sorted(Path(root).expanduser().glob("*/summary.json")): - with open(path) as handle: - summary = json.load(handle) - if summary.get("per_threshold"): - summaries.append((path.parent, summary)) - if not summaries: - return "

No rate-distortion sweep found.

" - - blocks = [] - for directory, summary in summaries: - held_out = summary.get("held_out") - badge = ( - "held out" if held_out - else "in the training set" - ) - base = summary["per_threshold"][0]["bytes_per_frame"] - rows = [] - for row in summary["per_threshold"]: - saved = 1.0 - row["bytes_per_frame"] / base if base else 0.0 - rows.append( - "" - f"{row['threshold']}" - f"{row['kept_fraction'] * 100:.1f}%" - f"{row['bytes_per_frame'] / 1000:.1f}" - f"{saved * 100:.1f}%" - f"{row['mbps_at_30fps']:.1f}" - f"{row['psnr']:.2f}" - f"{row['ssim']:.4f}" - f"{row['lpips']:.4f}" - ) - note = "" - if not held_out: - note = ("

Camera " - f"{summary['eval_view']} was trained on, so these quality numbers " - "measure how well the NeRF memorised its own training image. The bytes " - "are unaffected.

") - blocks.append( - f"

{html.escape(summary['object'])} · camera " - f"{summary['eval_view']}, {badge}

" - f"

{len(summary['frames'])} frames · " - f"{summary['block_size']}3 blocks · encoder quality " - f"{summary['quality']} · startup " - f"{summary['startup_bytes']['total_bytes'] / 1000:.0f} kB (the colour MLP, " - "sent once, not per frame)

" - + note + - "" - "" - "" - + "".join(rows) + "
thresholdblocks keptkB/framebytes savedMbps @30fpsPSNRSSIMLPIPS
" - + _rd_strip(directory, summary, assets) - ) - return ( - "

Bytes are ReRF's own encoder run over the retained blocks, plus the " - "block mask and the P-frames' motion vectors — measured, not estimated. Note " - "that dropping blocks does not save bytes proportionally: this encoder is sub-linear " - "in block count, so the blocks kept and bytes saved columns disagree, " - "and the second is the one a network sees.

" - "

Quality here is scored against the captured camera on the " - "subject's bounding box, which is stricter than section 4's filtered-vs-unfiltered " - "comparison: errors the NeRF already had no longer cancel out, and the identical " - "white background no longer dilutes the average.

" - + "".join(blocks) - ) - - -def _subject_box(images, pad: int = 24): - """Union bounding box of the non-white pixels across ``images``.""" - box = None - for image in images: - array = np.asarray(image.convert("RGB")) - rows = np.where((array < 250).any(axis=(1, 2)))[0] - cols = np.where((array < 250).any(axis=(0, 2)))[0] - if not rows.size or not cols.size: - continue - here = (rows[0], cols[0], rows[-1], cols[-1]) - box = here if box is None else ( - min(box[0], here[0]), min(box[1], here[1]), - max(box[2], here[2]), max(box[3], here[3]), - ) - if box is None: - return None - height, width = np.asarray(images[0]).shape[:2] - return (max(int(box[0]) - pad, 0), max(int(box[1]) - pad, 0), - min(int(box[2]) + pad, height), min(int(box[3]) + pad, width)) - - -def _rerf_section(runs_root, assets: Path, height: int = 300, - frames_shown: int = 6) -> str: - """ReRF's own bitstream, decoded and rendered by its own ``rerf_render.py``. - - This is the baseline NeVo is measured *against*, produced entirely by - upstream: ``codec/compress.py`` writes the bitstream, ``rerf_render.py`` - decodes it and renders a 360 orbit. Nothing in ``nevo/`` participates, which - is the point -- these bytes and these pixels are ReRF's, not ours. - """ - from PIL import Image - - blocks, rows = [], [] - for root in runs_root: - for run in sorted(Path(root).expanduser().glob("*/rerf")): - directory = run.parent - headers = sorted(run.glob("header_*.json")) - if not headers: - continue - frames = len(headers) - startup = sum( - (run / name).stat().st_size - for name in ("rgb_net.tar", "model_kwargs.json") - if (run / name).exists() - ) - stream = sum( - path.stat().st_size for path in run.iterdir() - if path.is_file() and path.name not in ("rgb_net.tar", "model_kwargs.json") - ) - checkpoints = sum(p.stat().st_size for p in directory.glob("*.tar")) - per_frame = stream / frames - rows.append( - "" - f"{html.escape(directory.name)}" - f"{frames}" - f"{(stream + startup) / 1e6:.1f} MB" - f"{per_frame / 1000:.1f}" - f"{per_frame * 8 * 30 / 1e6:.1f}" - f"{checkpoints / 2 ** 30:.1f} GB" - f"{checkpoints / (stream + startup):.0f}×" - ) - - rendered = sorted( - path for directory_360 in directory.glob("render_360_rerf_*") - for path in directory_360.glob("*.jpg") - if not path.stem.endswith("_depth") - ) - if not rendered: - continue - picked = _pick(rendered, min(frames_shown, len(rendered))) - opened = [Image.open(path) for path in picked] - box = _subject_box(opened) - figures = [] - for path, image in zip(picked, opened): - if box: - top, left, bottom, right = box - image = image.crop((left, top, right, bottom)) - width = max(1, round(image.width * height / image.height)) - src = _save(image, assets / f"rerf_{directory.name}_{path.stem}.jpg", width) - figures.append( - f"
frame {path.stem}" - f"
frame {path.stem}
" - ) - depth = next(iter(sorted( - path for directory_360 in directory.glob("render_360_rerf_*") - for path in directory_360.glob("*_depth.jpg") - )), None) - if depth is not None: - image = Image.open(depth) - if box: - top, left, bottom, right = box - image = image.crop((left, top, right, bottom)) - width = max(1, round(image.width * height / image.height)) - src = _save(image, assets / f"rerf_{directory.name}_depth.jpg", width) - figures.append( - f"
depth" - f"
depth (frame {depth.stem.split('_')[0]})
" - "
" - ) - blocks.append( - f"

{html.escape(directory.name)}

" - f"

{frames} frames · " - f"{(stream + startup) / 1e6:.1f} MB of bitstream · " - f"{per_frame * 8 * 30 / 1e6:.1f} Mbps at 30 fps

" - "
" + "".join(figures) + "
" - ) - - if not rows: - return "

No ReRF bitstream found. Run codec/compress.py.

" - return ( - "

Produced entirely by upstream ReRF: " - "codec/compress.py writes the bitstream (PCA on, " - "--pca_chs 7,13, quality 99) and rerf_render.py " - "decodes it and renders a 360° orbit. Nothing in nevo/ is " - "involved, so these are ReRF's own bytes and pixels — the baseline " - "NeVo is measured against. 150–234 Mbps matches the paper's " - "\"150+ Mbps\" for ReRF.

" - "

The ratio column is the point of the exercise: the " - "bitstream is ~700–1000× smaller than the training checkpoints " - "it came from, because the checkpoints are dense fp32 grids that are ~90% " - "empty and carry no quantisation, DCT or entropy coding.

" - "" - "" - "" + "".join(rows) + "
runframesbitstreamkB/frameMbps @30fpscheckpointsratio
" - "

Framing is rerf_render.py's own synthesised " - "orbit at 1920×1080, cropped to the subject here for legibility. It is " - "not the corpus camera, so these pixels are not comparable frame-for-frame " - "with the sections above.

" - + "".join(blocks) - ) - - -def _between_rig_cameras(sequence, camera, look_at_c2w): - """A camera halfway between two of the corpus rig's, at the same radius.""" - try: - with open(sequence.corpus_dir / "nevo_corpus.json") as handle: - manifest = json.load(handle) - except OSError: - return None - positions = np.asarray( - [np.asarray(entry["c2w_normalised"])[:3, 3] for entry in manifest["cameras"]] - ) - if len(positions) < 2: - return None - centre = (np.asarray(manifest["xyz_min"]) + np.asarray(manifest["xyz_max"])) * 0.5 - radius = float(np.linalg.norm(positions[0] - centre)) - midpoint = (positions[0] + positions[1]) * 0.5 - offset = midpoint - centre - eye = centre + offset / np.linalg.norm(offset) * radius - return dataclasses.replace( - camera, c2w=look_at_c2w(eye, centre, np.asarray((0.0, 1.0, 0.0))) - ) - - -def _cdf_section(results: list[Path], assets: Path) -> str: - rows = [] - plots = [] - for path in sorted(results): - with open(path) as handle: - payload = json.load(handle) - report = payload["report"] - summary = [s for s in report["summaries"] if s["pooling"] == "per-viewport"][0] - below = summary["fraction_below"] - # `.get` on the derived stats: result files written before a diagnostic - # existed should still show up in the table rather than sink the page. - never_hit = summary.get("never_hit_fraction") - rows.append( - "" - f"{html.escape(report['object'])}" - f"{summary['block_size']}3" - f"{summary['assignment']}" - f"{summary['frames']}" - f"{summary['viewports_per_frame']}" - f"{summary['occupied_blocks']:,}" - f"{below['0.01'] * 100:.1f}%" - f"{below['0.025'] * 100:.1f}%" - f"{below['0.05'] * 100:.1f}%" - f"{never_hit * 100:.1f}%" if never_hit is not None else "—" - f"{summary['quantiles']['0.5']:.4f}" - "" - ) - image = path.with_suffix(".png") - if image.is_file(): - target = assets / image.name - shutil.copyfile(image, target) - plots.append( - f"
" - f"
{html.escape(image.stem)}
" - ) - if not rows: - return "

No CDF results found.

" - return ( - "

Per-viewport pooling: every (voxel, viewport) pair is one sample. " - "The highlighted column is the paper's 0.025 threshold.

" - "" - "" - "" - + "".join(rows) - + "
objectblockassignmentframesviewportsnon-empty<0.01<0.025<0.05never hitmedian
" - + "".join(plots) - + "
" - ) - - -STYLE = """ -:root { color-scheme: light dark; --line: #8883; --hi: #f0b90020; } -* { box-sizing: border-box; } -body { margin: 0 auto; padding: 2rem 1.5rem 6rem; max-width: 1500px; - font: 15px/1.55 ui-sans-serif, system-ui, -apple-system, sans-serif; } -h1 { font-size: 1.5rem; margin: 0 0 .2rem; } -h2 { font-size: 1.15rem; margin: 2.5rem 0 .6rem; padding-bottom: .3rem; - border-bottom: 1px solid var(--line); } -h3 { font-size: 1rem; margin: 1.4rem 0 .3rem; } -h4 { font-size: .85rem; font-weight: 500; opacity: .65; margin: .8rem 0 .3rem; } -p.meta, p.hint, p.none { font-size: .85rem; opacity: .75; margin: .2rem 0 .6rem; } -p.none { font-style: italic; } -.strip { display: flex; gap: .5rem; overflow-x: auto; padding-bottom: .4rem; } -.strip.wrap { flex-wrap: wrap; overflow: visible; } -figure { margin: 0; flex: 0 0 auto; } -figure img { display: block; border: 1px solid var(--line); border-radius: 4px; - background: #7772; } -figcaption { font-size: .75rem; opacity: .6; text-align: center; margin-top: .15rem; } -.card { border: 1px solid var(--line); border-radius: 6px; padding: .7rem .8rem; - margin-bottom: .7rem; } -table { border-collapse: collapse; font-size: .85rem; margin: .5rem 0 1.2rem; } -th, td { border: 1px solid var(--line); padding: .25rem .55rem; text-align: right; } -th:first-child, td:first-child, td:nth-child(3) { text-align: left; } -thead th { font-weight: 600; opacity: .8; } -td.hi { background: var(--hi); font-weight: 600; } -td.ok { color: #1a7f37; text-align: left; } -td.bad { color: #b3261e; text-align: left; } -p.meta.bad { color: #b3261e; opacity: 1; font-weight: 500; } -nav a { margin-right: 1rem; font-size: .85rem; } -.mono { font-variant-numeric: tabular-nums; font-family: ui-monospace, monospace; - font-size: .8rem; opacity: .75; } -.player { border: 1px solid var(--line); border-radius: 8px; overflow: hidden; - max-width: 560px; } -.stage { position: relative; background: #7771; display: flex; - justify-content: center; align-items: center; min-height: 200px; } -.stage img { display: block; width: 100%; height: auto; } -.loading { position: absolute; font-size: .8rem; opacity: .6; } -.loading.hidden { display: none; } -.controls { padding: .6rem .8rem .8rem; border-top: 1px solid var(--line); } -.controls .row { display: flex; gap: .9rem; align-items: center; flex-wrap: wrap; - margin-bottom: .45rem; font-size: .82rem; } -.controls .row:last-of-type { margin-bottom: 0; } -.controls label { display: flex; gap: .35rem; align-items: center; } -.controls label.grow { flex: 1 1 240px; } -.controls label.grow input[type=range] { flex: 1; } -button { font: inherit; font-size: .82rem; padding: .25rem .7rem; cursor: pointer; - border: 1px solid var(--line); border-radius: 5px; background: #7771; - color: inherit; } -button.primary { min-width: 5.6rem; } -button.on { background: #4a90d922; border-color: #4a90d9; font-weight: 600; } -""" - -PLAYER_SCRIPT = """ -(function () { - const clips = window.NEVO_CLIPS || []; - if (!clips.length) return; - - const image = document.getElementById('stage-image'); - const loading = document.getElementById('stage-loading'); - const clipPicker = document.getElementById('clip'); - const viewPicker = document.getElementById('view'); - const frameSlider = document.getElementById('frame'); - const frameLabel = document.getElementById('frame-value'); - const playButton = document.getElementById('play'); - const fpsSlider = document.getElementById('fps'); - const fpsLabel = document.getElementById('fps-value'); - const matte = document.getElementById('matte'); - const note = document.getElementById('clip-note'); - - let clip = clips[0]; - let view = clip.views[0]; - let frame = 0; - let playing = false; - let lastTick = 0; - - const source = (index) => - clip.path + '/' + (matte.checked && !clip.no_matte ? 'matte' : 'rgb') + '/v' + - String(view).padStart(2, '0') + '_f' + - String(clip.frames[index]).padStart(3, '0') + '.jpg'; - - // Decode the whole clip before playing. At 10 fps a browser fetching each - // frame on demand shows a blank stage for the first pass, which is exactly - // when you are looking for a hitch in the input. - function preload() { - loading.classList.remove('hidden'); - let done = 0; - const total = clip.frames.length; - return Promise.all(clip.frames.map((_, index) => new Promise((resolve) => { - const probe = new Image(); - const finish = () => { - done += 1; - loading.textContent = 'loading ' + Math.round((done / total) * 100) + '%'; - resolve(); - }; - probe.onload = finish; - probe.onerror = finish; - probe.src = source(index); - }))).then(function () { loading.classList.add('hidden'); }); - } - - function paint() { - image.src = source(frame); - frameLabel.textContent = clip.frames[frame]; - } - - function load(next) { - clip = next; - view = clip.views.indexOf(view) >= 0 ? view : clip.views[0]; - viewPicker.textContent = ''; - for (const candidate of clip.views) { - const option = document.createElement('option'); - option.value = String(candidate); - option.textContent = 'camera ' + candidate; - viewPicker.appendChild(option); - } - viewPicker.value = String(view); - // A rendered clip is one camera with no matte; hiding the controls that do - // not apply is clearer than leaving them to 404. - viewPicker.parentElement.style.display = clip.views.length > 1 ? '' : 'none'; - matte.parentElement.style.display = clip.no_matte ? 'none' : ''; - frame = 0; - frameSlider.max = String(clip.frames.length - 1); - frameSlider.value = '0'; - note.textContent = clip.source + ' \\u2014 ' + clip.frames.length + - ' frames, ' + clip.size[0] + 'x' + clip.size[1] + ' as prepared'; - paint(); - preload(); - } - - function tick(now) { - if (!playing) return; - if (now - lastTick >= 1000 / Number(fpsSlider.value)) { - lastTick = now; - frame = (frame + 1) % clip.frames.length; - frameSlider.value = String(frame); - paint(); - } - requestAnimationFrame(tick); - } - - playButton.addEventListener('click', function () { - playing = !playing; - playButton.innerHTML = playing ? '❙❙ pause' : '▶ play'; - playButton.className = playing ? 'primary on' : 'primary'; - if (playing) { lastTick = 0; requestAnimationFrame(tick); } - }); - frameSlider.addEventListener('input', function () { - frame = Number(frameSlider.value); paint(); - }); - fpsSlider.addEventListener('input', function () { fpsLabel.textContent = fpsSlider.value; }); - matte.addEventListener('change', function () { paint(); preload(); }); - viewPicker.addEventListener('change', function () { - view = Number(viewPicker.value); paint(); preload(); - }); - clipPicker.addEventListener('change', function () { - const found = clips.find(function (item) { return item.name === clipPicker.value; }); - if (found) load(found); - }); - - load(clip); -})(); -""" - -SCRIPT = """ -(function () { - const clips = window.NEVO_CLIPS || []; - if (!clips.length) return; - - const image = document.getElementById('stage-image'); - const loading = document.getElementById('stage-loading'); - const clipPicker = document.getElementById('clip'); - const frameSlider = document.getElementById('frame'); - const azimuthSlider = document.getElementById('azimuth'); - const frameLabel = document.getElementById('frame-value'); - const azimuthLabel = document.getElementById('azimuth-value'); - const conditionRow = document.getElementById('conditions'); - const conditionNote = document.getElementById('condition-note'); - const playButton = document.getElementById('play'); - const fpsSlider = document.getElementById('fps'); - const fpsLabel = document.getElementById('fps-value'); - const autoOrbit = document.getElementById('auto-orbit'); - - let clip = clips[0]; - let condition = clip.conditions[0]; - let frame = 0; - let azimuth = 0; - let playing = false; - let lastTick = 0; - - const source = (aClip, aCondition, frameIndex, azimuthIndex) => - `${aClip.path}/${aCondition}/f${String(aClip.frames[frameIndex]).padStart(3, '0')}` + - `_a${String(azimuthIndex).padStart(2, '0')}.jpg`; - - // Decode every cell up front. The grid is a few hundred small JPEGs, and a - // player that stutters on the first pass through a viewpoint is useless for - // judging whether filtering is visible. - function preload(aClip) { - loading.classList.remove('hidden'); - const urls = []; - for (const aCondition of aClip.conditions) - for (let f = 0; f < aClip.frames.length; f += 1) - for (let a = 0; a < aClip.azimuths; a += 1) - urls.push(source(aClip, aCondition, f, a)); - let done = 0; - return Promise.all(urls.map((url) => new Promise((resolve) => { - const probe = new Image(); - const finish = () => { - done += 1; - loading.textContent = `loading ${Math.round((done / urls.length) * 100)}%`; - resolve(); - }; - probe.onload = finish; - probe.onerror = finish; - probe.src = url; - }))).then(() => loading.classList.add('hidden')); - } - - function paint() { - image.src = source(clip, condition, frame, azimuth); - frameLabel.textContent = clip.frames[frame]; - azimuthLabel.textContent = `${Math.round((azimuth / clip.azimuths) * 360)}\u00b0`; - } - - function describe() { - if (condition === 'full') { - conditionNote.textContent = - `every non-empty feature voxel delivered \u2014 ${clip.block_size}\u00b3 blocks`; - return; - } - const stat = (clip.stats || {})[condition]; - conditionNote.textContent = stat - ? `importance \u2265 ${condition}: ${(stat.dropped_mean * 100).toFixed(1)}% of ` + - `non-empty ${clip.block_size}\u00b3 blocks dropped, ` + - `${stat.psnr_mean.toFixed(1)} dB against the full render` - : `importance \u2265 ${condition}`; - } - - function buildConditions() { - conditionRow.textContent = ''; - for (const name of clip.conditions) { - const button = document.createElement('button'); - button.textContent = name === 'full' ? 'all voxels' : `filtered \u2265 ${name}`; - button.className = name === condition ? 'on' : ''; - button.addEventListener('click', () => { - condition = name; - for (const other of conditionRow.children) other.className = ''; - button.className = 'on'; - describe(); - paint(); - }); - conditionRow.appendChild(button); - } - } - - function load(aClip) { - clip = aClip; - condition = clip.conditions.includes(condition) ? condition : clip.conditions[0]; - frame = 0; - azimuth = 0; - frameSlider.max = String(clip.frames.length - 1); - frameSlider.value = '0'; - azimuthSlider.max = String(clip.azimuths - 1); - azimuthSlider.value = '0'; - buildConditions(); - describe(); - paint(); - preload(clip); - } - - function tick(now) { - if (!playing) return; - const interval = 1000 / Number(fpsSlider.value); - if (now - lastTick >= interval) { - lastTick = now; - frame = (frame + 1) % clip.frames.length; - frameSlider.value = String(frame); - if (autoOrbit.checked) { - azimuth = (azimuth + 1) % clip.azimuths; - azimuthSlider.value = String(azimuth); - } - paint(); - } - requestAnimationFrame(tick); - } - - playButton.addEventListener('click', () => { - playing = !playing; - playButton.innerHTML = playing ? '❙❙ pause' : '▶ play'; - playButton.className = playing ? 'primary on' : 'primary'; - if (playing) { lastTick = 0; requestAnimationFrame(tick); } - }); - frameSlider.addEventListener('input', () => { frame = Number(frameSlider.value); paint(); }); - azimuthSlider.addEventListener('input', () => { - azimuth = Number(azimuthSlider.value); paint(); - }); - fpsSlider.addEventListener('input', () => { fpsLabel.textContent = fpsSlider.value; }); - clipPicker.addEventListener('change', () => { - const found = clips.find((item) => item.name === clipPicker.value); - if (found) load(found); - }); - - // Drag to orbit: one full sweep of the image width covers the whole ring. - let dragging = false; - let dragStartX = 0; - let dragStartAzimuth = 0; - const beginDrag = (x) => { dragging = true; dragStartX = x; dragStartAzimuth = azimuth; }; - const moveDrag = (x) => { - if (!dragging) return; - const span = image.clientWidth || 1; - const steps = Math.round(((x - dragStartX) / span) * clip.azimuths); - azimuth = ((dragStartAzimuth + steps) % clip.azimuths + clip.azimuths) % clip.azimuths; - azimuthSlider.value = String(azimuth); - paint(); - }; - image.addEventListener('pointerdown', (event) => { - beginDrag(event.clientX); - image.setPointerCapture(event.pointerId); - }); - image.addEventListener('pointermove', (event) => moveDrag(event.clientX)); - image.addEventListener('pointerup', () => { dragging = false; }); - image.addEventListener('pointercancel', () => { dragging = false; }); - - load(clip); -})(); -""" - -SCRIPT = """ -for (const figure of document.querySelectorAll('figure.swap')) { - const image = figure.querySelector('img'); - const caption = figure.querySelector('figcaption'); - const label = caption.textContent; - figure.addEventListener('mouseenter', () => { - image.src = figure.dataset.b; caption.textContent = label + ' (matte)'; - }); - figure.addEventListener('mouseleave', () => { - image.src = figure.dataset.a; caption.textContent = label; - }); -} -""" - - -def build(args) -> Path: - out_dir = Path(args.out).expanduser().resolve() - assets = out_dir / "assets" - if assets.exists() and args.clean: - # Only the thumbnails. `player/` is built by a separate tool and is the - # expensive artefact here (hundreds of renders); wiping it on every - # report rebuild would be a trap. - shutil.rmtree(assets) - assets.mkdir(parents=True, exist_ok=True) - - corpora = sorted( - path.parent - for root in args.corpus_root - for path in Path(root).expanduser().glob("*/nevo_corpus.json") - ) - runs = sorted(Path(args.config_dir).expanduser().glob("*.py")) - results = [ - Path(path) - for root in args.results_root - for path in glob.glob(str(Path(root).expanduser() / "*" / "importance_cdf_*.json")) - ] - sweeps = [ - Path(path) - for root in args.results_root - for path in glob.glob(str(Path(root).expanduser() / "*" / "filter_sweep_*.json")) - ] - - print(f"corpora: {[c.name for c in corpora]}", flush=True) - print(f"runs: {[r.stem for r in runs]}", flush=True) - print(f"cdf results: {len(results)}, threshold sweeps: {len(sweeps)}", flush=True) - - output_clips = _output_clips(args.output_root, assets, VIEW_WIDTH) - print(f"rendered clips: {[clip['name'] for clip in output_clips]}", flush=True) - # No corpus clips: the player shows NeVo's output, and the prepared corpus - # has its own section below. - player_html = _player_section([], assets, args.player_views, args.player_width, - output_clips) - corpus_html = _corpus_section(corpora, assets, args.frames_shown, args.views_shown) - if args.skip_render: - reload_html = filter_html = "

Skipped (--skip-render).

" - else: - rerf_env.activate() - try: - reload_html, filter_html = _reload_and_filter_section(runs, assets, args) - except Exception as error: - # Rendering wants a GPU, and on this box the report is often built - # while training is using both. Losing two sections is a much better - # outcome than losing the page, especially since the other three - # need no GPU at all. - note = ( - f"

Could not render: {html.escape(type(error).__name__)}: " - f"{html.escape(str(error).splitlines()[0])}
" - f"Re-run once a GPU is free, or pass --skip-render.

" - ) - print(f"render sections failed: {type(error).__name__}: {error}", flush=True) - reload_html = filter_html = note - cdf_html = _cdf_section(results, assets) - sweep_html = _sweep_section(sweeps) - rd_html = _rd_section(args.results_root, assets) - rerf_html = _rerf_section(args.runs_root, assets) - - stamp = datetime.datetime.now().strftime("%Y-%m-%d %H:%M") - page = f""" - - -NeVo baseline — output check -

NeVo baseline — output check

-

generated {stamp} · the paper's section 3.2 (neural visibility, -visibility-aware filtering) measured end to end, in bytes and in delivered quality. -Sections 3.3 (loss recovery) and 3.4 (reprojection) are not built.

- - -

1. Player — NeVo's output

-{player_html} - -

2. Corpus — what ReRF was trained on

-{corpus_html} - -

3. Reload — did step 1 rebuild the frame correctly?

-

A correctly reassembled checkpoint lands near the PSNR the trainer -logged. A mis-wired one still renders, just not this scene.

-{reload_html} - -

4. Filtering — what does dropping the low-importance voxels cost?

-

Dropped blocks are written back the way ReRF's decoder fills a block -that never arrived (raw density −4.1, zero features), not zeroed. The difference -image is amplified; at 1x it is invisible, which is the claim being checked.

-{filter_html} - -

5. Importance CDF

-{cdf_html} - -

6. Threshold sweep — how far can filtering go?

-{sweep_html} - -

7. Rate–distortion — bytes against delivered quality

-{rd_html} - -

8. ReRF itself — its own bitstream, decoded and rendered

-{rerf_html} - - -""" - index = out_dir / "index.html" - index.write_text(page) - print(f"wrote {index}", flush=True) - return index - - -def parse_args(argv=None) -> argparse.Namespace: - parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) - parser.add_argument("--out", default="~/nevo_report") - parser.add_argument("--corpus-root", nargs="+", default=["~/nevo_data_g"], - help="prepared corpora to show; pass several to compare") - parser.add_argument("--config-dir", - default=str(MODULE_ROOT / "rerf/configs/nevo")) - parser.add_argument("--results-root", nargs="+", default=["~/nevo_results"]) - parser.add_argument("--runs-root", nargs="+", default=["~/nevo_runs"], - help="trained runs; read for ReRF's own bitstream and renders") - parser.add_argument("--output-root", nargs="+", default=["~/nevo_output"], - help="rendered reconstruction frames from render_frames.py") - parser.add_argument("--frames-shown", type=int, default=2) - parser.add_argument("--views-shown", type=int, default=6, - help="views per object in the contact sheet; 0 = all") - parser.add_argument("--player-views", type=int, default=4, - help="cameras selectable in the player; 0 = all") - parser.add_argument("--player-width", type=int, default=360) - parser.add_argument("--frames-checked", type=int, default=2, - help="trained frames to reload-check per object") - parser.add_argument("--view", type=int, default=0, help="which training view to render") - parser.add_argument("--block-size", type=int, default=8) - parser.add_argument("--thresholds", type=float, nargs="+", default=[0.025, 0.2], - help="0.025 is the figure the paper is usually quoted by; 0.2 is " - "what its own SSIM-target fitting actually lands on for this " - "content. Showing both is the point.") - parser.add_argument("--difference-gain", type=float, default=8.0) - parser.add_argument("--skip-render", action="store_true") - parser.add_argument("--clean", action="store_true", - help="rebuild the thumbnails from scratch; leaves player/ alone") - parser.add_argument("--serve", type=int, default=0, metavar="PORT") - return parser.parse_args(argv) - - -def main(argv=None) -> int: - args = parse_args(argv) - index = build(args) - if args.serve: - import http.server - import socket - - directory = str(index.parent) - - class Handler(http.server.SimpleHTTPRequestHandler): - def __init__(self, *handler_args, **handler_kwargs): - super().__init__(*handler_args, directory=directory, **handler_kwargs) - - def end_headers(self): - # Regenerating the report renames assets. A browser holding a - # cached index.html then asks for files that no longer exist and - # shows a page of broken images, which looks exactly like the - # pipeline having produced nothing. Never cache. - self.send_header("Cache-Control", "no-store, must-revalidate") - super().end_headers() - - def log_message(self, *_): - pass - - http.server.ThreadingHTTPServer.allow_reuse_address = True - with http.server.ThreadingHTTPServer(("0.0.0.0", args.serve), Handler) as server: - host = socket.gethostbyname(socket.gethostname()) - print(f"serving {directory} at http://{host}:{args.serve}/ (ctrl-c to stop)", - flush=True) - server.serve_forever() - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/open4d/reconstruction/nevo/orbitnevo/rerf_cli.py b/open4d/reconstruction/nevo/orbitnevo/rerf_cli.py deleted file mode 100644 index 66ad7296..00000000 --- a/open4d/reconstruction/nevo/orbitnevo/rerf_cli.py +++ /dev/null @@ -1,99 +0,0 @@ -"""Run a vendored ReRF script with this repo's environment already set up. - -Upstream's README asks you to launch its scripts as - - LD_LIBRARY_PATH=./ac_dc:$LD_LIBRARY_PATH PYTHONPATH=./ac_dc/:$PYTHONPATH \\ - python codec/compress.py --model_path ... - -from inside ``rerf/``. That is easy to get wrong and it is not enough on its own: -the entropy coder has to be preloaded, the CWD has to be the ReRF root (``codec`` -resolves ``./codec/quant.npy`` at import), and two upstream calls break against -modern numpy/imageio. :mod:`nevo.rerf_env` handles all of it, so this just wires -it to a command line and execs the script in-process. - -Compress a trained sequence into ReRF's own bitstream, then render it back: - - python -m orbitnevo.rerf_cli codec/compress.py \\ - --model_path ~/nevo_runs/g_basketball --expr_name rerf \\ - --quality 99 --pca --pca_chs 7,13 --frame_num 30 - - python -m orbitnevo.rerf_cli rerf_render.py \\ - --config configs/nevo/g_basketball.py \\ - --compression_path ~/nevo_runs/g_basketball/rerf \\ - --render_360 30 --pca --pca_chs 7,13 - -Two things upstream will not tell you: - -* ``--pca``/``--group_size`` must match between compress and render, per - upstream's README. PCA is ~13% smaller here and part of ReRF's published - method, so it is worth passing at both ends. -* ``--render_360`` must not exceed the number of *compressed* frames. Frame ids - wrap on ``cfg.frame_num``, but the decode stream is pulled sequentially, so - asking for more frames than were compressed exhausts the iterator. - -Paths in arguments are expanded and made absolute before the CWD changes, so -``~`` and relative paths behave the way the shell led you to expect. - -Runs in the ``nevo`` environment (Python 3.8), like everything else that touches -a ReRF model. -""" -from __future__ import annotations - -import argparse -import os -import sys -from pathlib import Path - -MODULE_ROOT = Path(__file__).resolve().parents[1] -if str(MODULE_ROOT) not in sys.path: - sys.path.insert(0, str(MODULE_ROOT)) - -from nevo import rerf_env # noqa: E402 - -SCRIPTS = ("codec/compress.py", "rerf_render.py", "run.py", "tools/vis_volume.py") - - -def _absolute(argument: str) -> str: - """Expand a path-looking argument while the CWD is still the caller's.""" - if argument.startswith("-") or "/" not in argument: - return argument - expanded = Path(argument).expanduser() - # Only rewrite things that exist, or whose parent does: leaves values like - # `configs/nevo/x.py` (relative to the ReRF root) and `7,13` alone. - if expanded.exists() or (expanded.parent.exists() and expanded.is_absolute()): - return str(expanded.resolve()) - if expanded.is_absolute(): - return str(expanded) - return argument - - -def main(argv=None) -> int: - parser = argparse.ArgumentParser( - description=__doc__.split("\n", 1)[0], - usage="%(prog)s diff --git a/open4d/streamer/streamer/client/worker.js b/open4d/streamer/streamer/client/worker.js new file mode 100644 index 00000000..e5198f87 --- /dev/null +++ b/open4d/streamer/streamer/client/worker.js @@ -0,0 +1,498 @@ +/* The decode worker. + * + * Parsing a frame is the one part of playback that is unavoidably expensive and + * unavoidably synchronous: measured on this repository's content, 16.5 ms for a + * 3DGS PLY, 32.1 ms for a 439k-Gaussian .splat, 17.8 ms for a Draco mesh. On + * the main thread each of those is a frame's worth of `requestAnimationFrame` + * missed, so playback stuttered in proportion to how much geometry it was + * showing -- and the panes that were not decoding stuttered too, because there + * is only one main thread. + * + * So the codecs live here and nowhere else. Not shared with the page: sharing + * would mean either duplicating them, which drifts, or a build step, which this + * client does not have. The page holds renderers and UI; this holds every byte + * that has to be turned into geometry. + * + * `decodeImage` is the exception and stays on the main thread. It needs `Image`, + * and a browser already decodes an image off the main thread, so moving it here + * would buy nothing and cost the ImageBitmap-to- problem. + * + * Results come back with their buffers *transferred*, not copied -- a 4.3 MB + * frame parses to several megabytes of typed arrays, and copying those across + * the boundary would give back much of what the worker saved. + */ + +const SH_C0 = 0.28209479177387814; + +/* ------------------------------------------------------------------ PLY --- + * The 3DGS PLY: a binary_little_endian vertex element of float32 properties + * named x y z, nx ny nz, f_dc_*, f_rest_*, opacity, scale_*, rot_*. Values are + * raw, so opacity is a logit, scale a log, and rot an unnormalised quaternion. + * f_rest is ignored: a view-dependent band cannot be evaluated without knowing + * the direction the exporter baked colour from, and the exports this viewer is + * built for are degree 0 anyway. + */ +function parsePly(buffer) { + const bytes = new Uint8Array(buffer); + const limit = Math.min(bytes.length, 1 << 16); + const head = new TextDecoder("ascii").decode(bytes.subarray(0, limit)); + const marker = head.indexOf("end_header"); + if (marker < 0) throw new Error("not a PLY: no end_header in the first 64 KiB"); + const dataStart = head.indexOf("\n", marker) + 1; + const lines = head.slice(0, marker).split("\n").map((line) => line.trim()); + + if (!lines.some((line) => line.startsWith("format binary_little_endian"))) { + throw new Error("only binary_little_endian PLY is supported"); + } + let count = 0, inVertex = false; + const names = []; + for (const line of lines) { + if (line.startsWith("element ")) { + const parts = line.split(/\s+/); + inVertex = parts[1] === "vertex"; + if (inVertex) count = parseInt(parts[2], 10); + } else if (line.startsWith("property ") && inVertex) { + names.push(line.split(/\s+/)[2]); + } + } + if (!count) throw new Error("PLY declares no vertices"); + + const stride = names.length; + // Copied rather than viewed in place: the header length is whatever it is, so + // `dataStart` is rarely a multiple of 4 and a Float32Array cannot be created + // at an unaligned offset. + const table = new Float32Array(buffer.slice(dataStart, dataStart + count * stride * 4)); + const at = {}; + names.forEach((name, index) => { at[name] = index; }); + for (const need of ["x", "y", "z", "opacity", "scale_0", "rot_0", "f_dc_0"]) { + if (at[need] === undefined) throw new Error(`PLY is missing property ${need}`); + } + + // Three RGBA32F texels per Gaussian: (xyz, opacity), (c00 c01 c02 c11), + // (c12 c22 _ _). Colour rides in a separate RGBA8 texture so neither has to + // be bit-packed. + const data = new Float32Array(count * 12); + const colors = new Uint8Array(count * 4); + const positions = new Float32Array(count * 3); + + for (let i = 0; i < count; i++) { + const row = i * stride; + const px = table[row + at.x], py = table[row + at.y], pz = table[row + at.z]; + positions[i * 3] = px; positions[i * 3 + 1] = py; positions[i * 3 + 2] = pz; + + const sx = Math.exp(table[row + at.scale_0]); + const sy = Math.exp(table[row + at.scale_1]); + const sz = Math.exp(table[row + at.scale_2]); + + let qw = table[row + at.rot_0], qx = table[row + at.rot_1]; + let qy = table[row + at.rot_2], qz = table[row + at.rot_3]; + const norm = Math.hypot(qw, qx, qy, qz) || 1; + qw /= norm; qx /= norm; qy /= norm; qz /= norm; + + // 3DGS's build_rotation, with (w, x, y, z) as stored. + const r00 = 1 - 2 * (qy * qy + qz * qz), r01 = 2 * (qx * qy - qw * qz), r02 = 2 * (qx * qz + qw * qy); + const r10 = 2 * (qx * qy + qw * qz), r11 = 1 - 2 * (qx * qx + qz * qz), r12 = 2 * (qy * qz - qw * qx); + const r20 = 2 * (qx * qz - qw * qy), r21 = 2 * (qy * qz + qw * qx), r22 = 1 - 2 * (qx * qx + qy * qy); + + // Sigma = (R diag(s)) (R diag(s))^T, upper triangle only. + const m00 = r00 * sx, m01 = r01 * sy, m02 = r02 * sz; + const m10 = r10 * sx, m11 = r11 * sy, m12 = r12 * sz; + const m20 = r20 * sx, m21 = r21 * sy, m22 = r22 * sz; + + const base = i * 12; + data[base] = px; data[base + 1] = py; data[base + 2] = pz; + data[base + 3] = 1 / (1 + Math.exp(-table[row + at.opacity])); + data[base + 4] = m00 * m00 + m01 * m01 + m02 * m02; + data[base + 5] = m00 * m10 + m01 * m11 + m02 * m12; + data[base + 6] = m00 * m20 + m01 * m21 + m02 * m22; + data[base + 7] = m10 * m10 + m11 * m11 + m12 * m12; + data[base + 8] = m10 * m20 + m11 * m21 + m12 * m22; + data[base + 9] = m20 * m20 + m21 * m21 + m22 * m22; + + for (let c = 0; c < 3; c++) { + const value = 0.5 + SH_C0 * table[row + at["f_dc_" + c]]; + colors[i * 4 + c] = Math.max(0, Math.min(255, Math.round(value * 255))); + } + colors[i * 4 + 3] = 255; + } + return { count, data, colors, positions }; +} + +/* One frame of `.splat`: 32 bytes per Gaussian, the quantised delivery form + * (see gs_tools/io/splat.py for the byte layout and what it costs). + * + * The values are stored ALREADY ACTIVATED, unlike a PLY -- scales are + * world-space standard deviations, opacity is in [0, 1], the quaternion is + * unit. So this applies no exp, no sigmoid and no normalise, and the one bug to + * watch for is doing so anyway: it renders as fog, exactly as writing activated + * values into a PLY does. Returns the same shape parsePly does, so nothing + * downstream can tell which format a frame arrived in. */ +function parseSplat(buffer) { + const SPLAT_BYTES = 32; + if (buffer.byteLength % SPLAT_BYTES) { + throw new Error(`not a .splat: ${buffer.byteLength} bytes is not a multiple of 32`); + } + const count = buffer.byteLength / SPLAT_BYTES; + if (!count) throw new Error(".splat frame is empty"); + const bytes = new Uint8Array(buffer); + // Copied, not viewed: byteOffset 0 is aligned here, but a Float32Array view + // over the same buffer would still need a stride this layout does not have. + const floats = new Float32Array(buffer); + + const data = new Float32Array(count * 12); + const colors = new Uint8Array(count * 4); + const positions = new Float32Array(count * 3); + + for (let i = 0; i < count; i++) { + const f = i * 8; // 8 float32 slots per record + const b = i * SPLAT_BYTES; + const px = floats[f], py = floats[f + 1], pz = floats[f + 2]; + positions[i * 3] = px; positions[i * 3 + 1] = py; positions[i * 3 + 2] = pz; + + const sx = floats[f + 3], sy = floats[f + 4], sz = floats[f + 5]; + + // Renormalised: the quantisation is not symmetric, so q = 1 encodes as 256, + // clamps to 255 and arrives as 0.992. Skipping this scales every Gaussian by + // ~1.6% -- small enough to look fine and still be wrong. + let qw = (bytes[b + 28] - 128) / 128; + let qx = (bytes[b + 29] - 128) / 128; + let qy = (bytes[b + 30] - 128) / 128; + let qz = (bytes[b + 31] - 128) / 128; + const norm = Math.hypot(qw, qx, qy, qz) || 1; + qw /= norm; qx /= norm; qy /= norm; qz /= norm; + + const r00 = 1 - 2 * (qy * qy + qz * qz), r01 = 2 * (qx * qy - qw * qz), r02 = 2 * (qx * qz + qw * qy); + const r10 = 2 * (qx * qy + qw * qz), r11 = 1 - 2 * (qx * qx + qz * qz), r12 = 2 * (qy * qz - qw * qx); + const r20 = 2 * (qx * qz - qw * qy), r21 = 2 * (qy * qz + qw * qx), r22 = 1 - 2 * (qx * qx + qy * qy); + + const m00 = r00 * sx, m01 = r01 * sy, m02 = r02 * sz; + const m10 = r10 * sx, m11 = r11 * sy, m12 = r12 * sz; + const m20 = r20 * sx, m21 = r21 * sy, m22 = r22 * sz; + + const base = i * 12; + data[base] = px; data[base + 1] = py; data[base + 2] = pz; + data[base + 3] = bytes[b + 27] / 255; // opacity, already activated + data[base + 4] = m00 * m00 + m01 * m01 + m02 * m02; + data[base + 5] = m00 * m10 + m01 * m11 + m02 * m12; + data[base + 6] = m00 * m20 + m01 * m21 + m02 * m22; + data[base + 7] = m10 * m10 + m11 * m11 + m12 * m12; + data[base + 8] = m10 * m20 + m11 * m21 + m12 * m22; + data[base + 9] = m20 * m20 + m21 * m21 + m22 * m22; + + colors[i * 4] = bytes[b + 24]; // RGB, not an SH coefficient + colors[i * 4 + 1] = bytes[b + 25]; + colors[i * 4 + 2] = bytes[b + 26]; + colors[i * 4 + 3] = 255; + } + return { count, data, colors, positions }; +} + +/* A mesh or point-cloud frame, as `open4d.io.write_sequence` writes it: a + * binary_little_endian PLY with a vertex element and, for a mesh, a face + * element of `list uchar int vertex_indices`. + * + * A general-enough PLY reader rather than the fixed-stride one parsePly uses, + * because these headers genuinely vary: colour is float in Open4D's canon (see + * open4d.core.dtypes) but uchar in most files from elsewhere, and a face list + * is variable-length, so neither element has a stride known up front. + * + * Normals are not read. Open4D's PLY writer refuses to store them, and the + * renderer derives a face normal from screen-space derivatives instead, which + * works for any mesh regardless of what its producer chose to carry. */ +function parseMeshPly(buffer) { + const bytes = new Uint8Array(buffer); + const limit = Math.min(bytes.length, 1 << 16); + const head = new TextDecoder("ascii").decode(bytes.subarray(0, limit)); + const marker = head.indexOf("end_header"); + if (marker < 0) throw new Error("not a PLY: no end_header in the first 64 KiB"); + const dataStart = head.indexOf("\n", marker) + 1; + const lines = head.slice(0, marker).split("\n").map((line) => line.trim()); + if (!lines.some((line) => line.startsWith("format binary_little_endian"))) { + throw new Error("only binary_little_endian PLY is supported"); + } + + const SIZES = { char: 1, int8: 1, uchar: 1, uint8: 1, short: 2, int16: 2, + ushort: 2, uint16: 2, int: 4, int32: 4, uint: 4, uint32: 4, + float: 4, float32: 4, double: 8, float64: 8 }; + const elements = []; + for (const line of lines) { + const parts = line.split(/\s+/); + if (parts[0] === "element") { + elements.push({ name: parts[1], count: parseInt(parts[2], 10), properties: [] }); + } else if (parts[0] === "property" && elements.length) { + const element = elements[elements.length - 1]; + if (parts[1] === "list") { + element.properties.push({ list: true, countType: parts[2], type: parts[3], name: parts[4] }); + } else { + element.properties.push({ list: false, type: parts[1], name: parts[2] }); + } + } + } + + const view = new DataView(buffer); + let offset = dataStart; + const read = (type) => { + switch (type) { + case "char": case "int8": return view.getInt8(offset++); + case "uchar": case "uint8": return view.getUint8(offset++); + case "short": case "int16": { const v = view.getInt16(offset, true); offset += 2; return v; } + case "ushort": case "uint16": { const v = view.getUint16(offset, true); offset += 2; return v; } + case "int": case "int32": { const v = view.getInt32(offset, true); offset += 4; return v; } + case "uint": case "uint32": { const v = view.getUint32(offset, true); offset += 4; return v; } + case "float": case "float32": { const v = view.getFloat32(offset, true); offset += 4; return v; } + case "double": case "float64": { const v = view.getFloat64(offset, true); offset += 8; return v; } + default: throw new Error(`unsupported PLY property type ${type}`); + } + }; + + let count = 0; + let positions = null, colors = null; + const faces = []; + + for (const element of elements) { + if (element.name === "vertex") { + count = element.count; + const names = element.properties.map((property) => property.name); + if (!["x", "y", "z"].every((need) => names.includes(need))) { + throw new Error("PLY vertex element has no x/y/z"); + } + // Colour is float in [0, 1] when Open4D wrote it and uchar in [0, 255] + // when most other tools did; the property type says which. + const colourNames = names.includes("red") ? ["red", "green", "blue"] + : names.includes("r") ? ["r", "g", "b"] : null; + const colourType = colourNames + ? element.properties.find((property) => property.name === colourNames[0]).type + : null; + const byteScale = colourType && SIZES[colourType] === 1 ? 1 : 255; + positions = new Float32Array(count * 3); + if (colourNames) colors = new Uint8Array(count * 4); + for (let i = 0; i < count; i++) { + const row = {}; + for (const property of element.properties) { + if (property.list) { + const n = read(property.countType); + for (let k = 0; k < n; k++) read(property.type); + } else { + row[property.name] = read(property.type); + } + } + positions[i * 3] = row.x; positions[i * 3 + 1] = row.y; positions[i * 3 + 2] = row.z; + if (colourNames) { + for (let c = 0; c < 3; c++) { + const value = row[colourNames[c]] * byteScale; + colors[i * 4 + c] = Math.max(0, Math.min(255, Math.round(value))); + } + colors[i * 4 + 3] = 255; + } + } + } else { + for (let i = 0; i < element.count; i++) { + for (const property of element.properties) { + if (property.list) { + const n = read(property.countType); + const indices = []; + for (let k = 0; k < n; k++) indices.push(read(property.type)); + // Triangle fan, so a quad or an n-gon renders rather than being + // dropped. Open4D only writes triangles; other producers do not. + for (let k = 2; k < indices.length; k++) { + faces.push(indices[0], indices[k - 1], indices[k]); + } + } else { + read(property.type); + } + } + } + } + } + + if (!count) throw new Error("PLY declares no vertices"); + const indices = faces.length ? new Uint32Array(faces) : null; + + const lower = [Infinity, Infinity, Infinity]; + const upper = [-Infinity, -Infinity, -Infinity]; + for (let i = 0; i < count; i++) { + for (let c = 0; c < 3; c++) { + const value = positions[i * 3 + c]; + if (value < lower[c]) lower[c] = value; + if (value > upper[c]) upper[c] = value; + } + } + return { count, positions, colors, indices, bounds: [lower, upper] }; +} + +/* ------------------------------------------------------------------ draco --- + * A compressed mesh frame, decoded here rather than on the server. + * + * This is the first format in this client that is actually a *compression* of + * geometry rather than an interchange dump of it: measured on the basketball + * sequence, 761 kB of PLY becomes 59 kB of Draco at 14-bit quantisation, which + * is 1.8 MB/s at 30 fps instead of 23. That is the difference between a link + * and a LAN, and it is why this exists. + * + * Two costs, both real and both bounded. Position is quantised -- at 14 bits the + * worst vertex moved 0.0046% of the model's diagonal on that sequence -- and + * duplicate vertices are merged, so the decoded point count is lower than the + * encoder was given (20,672 -> 19,747 there, because the source OBJ splits + * vertices at seams). Neither changes what the surface looks like; both mean a + * `.drc` frame is a delivery form and the PLY stays the source of truth. + * + * The decoder is Google's, vendored under client/vendor/draco and served from + * this origin -- never a CDN, which is what keeps the page free of external + * dependencies. It is a WASM module, so loading is asynchronous and happens once + * on first use: `Scheduler.decode` is awaited, which is what makes an async + * decoder possible at all here. + */ +const DRACO_PATH = "vendor/draco"; + +let dracoModule = null; + + +/* One Draco frame, in the shape parseMeshPly returns so the renderer cannot + * tell which format a frame arrived in. */ +async function parseDraco(buffer) { + const draco = await loadDraco(); + const decoder = new draco.Decoder(); + const input = new draco.DecoderBuffer(); + input.Init(new Int8Array(buffer), buffer.byteLength); + + let mesh = null; + try { + if (decoder.GetEncodedGeometryType(input) !== draco.TRIANGULAR_MESH) { + throw new Error("not a Draco triangular mesh"); + } + mesh = new draco.Mesh(); + const status = decoder.DecodeBufferToMesh(input, mesh); + if (!status.ok()) throw new Error(status.error_msg()); + + const count = mesh.num_points(); + const faces = mesh.num_faces(); + + const positions = new Float32Array(count * 3); + const attribute = decoder.GetAttribute( + mesh, decoder.GetAttributeId(mesh, draco.POSITION)); + const values = new draco.DracoFloat32Array(); + decoder.GetAttributeFloatForAllPoints(mesh, attribute, values); + for (let i = 0; i < count * 3; i++) positions[i] = values.GetValue(i); + draco.destroy(values); + + // Through the heap rather than GetFaceFromMesh per face: 39k calls across + // the WASM boundary is most of the decode time on this content. + const bytes = faces * 3 * 4; + const pointer = draco._malloc(bytes); + decoder.GetTrianglesUInt32Array(mesh, bytes, pointer); + const indices = new Uint32Array(draco.HEAPF32.buffer, pointer, faces * 3).slice(); + draco._free(pointer); + + let colors = null; + const colourId = decoder.GetAttributeId(mesh, draco.COLOR); + if (colourId >= 0) { + const colourAttribute = decoder.GetAttribute(mesh, colourId); + const raw = new draco.DracoFloat32Array(); + decoder.GetAttributeFloatForAllPoints(mesh, colourAttribute, raw); + const channels = colourAttribute.num_components(); + colors = new Uint8Array(count * 4); + for (let i = 0; i < count; i++) { + for (let c = 0; c < 3; c++) { + const value = raw.GetValue(i * channels + Math.min(c, channels - 1)); + // Draco keeps whatever range it was given; Open4D's canon is [0, 1]. + colors[i * 4 + c] = Math.max(0, Math.min(255, Math.round( + value <= 1.0001 ? value * 255 : value))); + } + colors[i * 4 + 3] = 255; + } + draco.destroy(raw); + } + + const lower = [Infinity, Infinity, Infinity]; + const upper = [-Infinity, -Infinity, -Infinity]; + for (let i = 0; i < count; i++) { + for (let c = 0; c < 3; c++) { + const value = positions[i * 3 + c]; + if (value < lower[c]) lower[c] = value; + if (value > upper[c]) upper[c] = value; + } + } + return { count, positions, colors, indices, bounds: [lower, upper] }; + } finally { + if (mesh) draco.destroy(mesh); + draco.destroy(input); + draco.destroy(decoder); + } +} + +/* ----------------------------------------------------------------- codecs --- + * What is on the wire, keyed the way `streamer.codecs` keys it: by + * (representation, suffix), not by suffix. The same `.ply` is a 3DGS cloud or a + * mesh depending on which representation is asking, and they need different + * parsers -- which is why this was previously two hand-written dispatchers, one + * per representation, each sniffing an extension. + * + * Adding a codec is an entry here and a parser. Nothing else in this file + * changes, and nothing outside it learns the format's name. */ +function suffixOf(url) { + const path = url.split("?", 1)[0]; + const dot = path.lastIndexOf("."); + return dot < 0 ? "" : path.slice(dot).toLowerCase(); +} + +const CODECS = { + mesh: { ".ply": parseMeshPly, ".drc": parseDraco }, + points: { ".ply": parseMeshPly, ".drc": parseDraco }, + gaussians: { ".ply": parsePly, ".splat": parseSplat }, + // No pixels: an needs the main thread, and the browser already + // decodes one off it. +}; + +function codecFor(representation, url) { + const table = CODECS[representation] || {}; + const suffix = suffixOf(url); + const parse = table[suffix]; + if (!parse) { + const offered = Object.keys(table).join(", ") || "nothing"; + throw new Error( + `no decoder for a ${representation} frame ending ${suffix || "(none)"}; ` + + `this client decodes ${offered}`); + } + return parse; +} + + +/* `importScripts` rather than a script tag, and paths relative to this file + * rather than to the bundle root. Same module and same wasm as the page used to + * fetch, from the same origin. */ +async function loadDraco() { + if (dracoModule) return dracoModule; + dracoModule = (async () => { + if (!self.DracoDecoderModule) { + importScripts(`${DRACO_PATH}/draco_wasm_wrapper.js`); + } + const wasmBinary = await (await fetch(`${DRACO_PATH}/draco_decoder.wasm`)) + .arrayBuffer(); + return self.DracoDecoderModule({ wasmBinary }); + })(); + return dracoModule; +} + +/* Every typed array in a parsed frame, so they can be transferred rather than + * copied. Collected by inspection rather than by a fixed list: a parser that + * grows a field should not have to remember to declare it here. */ +function transferables(parsed) { + const buffers = []; + for (const value of Object.values(parsed)) { + if (value && value.buffer instanceof ArrayBuffer) buffers.push(value.buffer); + } + return buffers; +} + +self.onmessage = async (event) => { + const { id, representation, url, buffer } = event.data; + try { + const parse = codecFor(representation, url); + const parsed = await parse(buffer, url); + self.postMessage({ id, parsed }, transferables(parsed)); + } catch (error) { + // The message, not the Error: an Error does not survive structured cloning + // in every browser, and losing it would turn a clear failure into silence. + self.postMessage({ id, error: String((error && error.message) || error) }); + } +}; diff --git a/open4d/streamer/streamer/codecs.py b/open4d/streamer/streamer/codecs.py new file mode 100644 index 00000000..4c6fc5b6 --- /dev/null +++ b/open4d/streamer/streamer/codecs.py @@ -0,0 +1,267 @@ +"""What is on the wire, and who can decode it. + +`representations` answers what a *decoded* frame is. This answers the question +one level down, which is the one that decides whether a module can be streamed +at all: in what format do its frames travel, and can the client turn them back +into that representation? + +Until this existed the two questions were conflated, and the cost was concrete. +Open4D's codecs produce ``.d4d``, ``.v4d``, ``.q4d`` and six more; the client +could decode ``.ply``, ``.splat``, ``.jpg`` and ``.png``. The two sets did not +intersect at all, and nothing in the code said so — a producer found out by +watching a pane stay blank. A registry that names the wire format and where it +decodes makes that a lookup instead of a discovery. + +**The key is (representation, suffix), not suffix.** ``.ply`` is claimed by three +representations here and needs two different parsers: a 3DGS PLY and a mesh PLY +share an extension and nothing else. Keying on the extension alone is what +forced the client to carry a hand-written dispatcher per representation, which +is exactly the per-format branching this registry removes. + +``decodes`` is the axis that makes the platform honest about heterogeneous +modules: + +``client`` + The browser can turn the bytes back into geometry. This is what buys a free + camera, and it needs a decoder shipped in `streamer.client`. +``server`` + It cannot, and no amount of work here will change that for some formats — + ReRF's entropy coder exists only as a CPython 3.8 binary. Those are decoded + and rendered where the GPU is, and what reaches the client is ``pixels``. + That is the codec's output representation as the platform sees it, which is + why ``.rerf`` is registered as producing ``pixels`` rather than a volume it + never delivers. +""" + +from __future__ import annotations + +from dataclasses import dataclass + +from open4d.core import Representation + +_OCTET = "application/octet-stream" + +#: Where a codec's frames are turned back into a representation. +DECODERS = ("client", "server") + + +@dataclass(frozen=True) +class CodecSpec: + """One wire format, and what the platform can do with it.""" + + #: Stable name, used in a clip's ``detail`` and in error messages. + name: str + #: Frame-file suffix, lowercase, with the dot. + suffix: str + #: What a decoded frame of this codec is. + representation: Representation + #: ``Content-Type`` a server must send it as. + media_type: str = _OCTET + #: ``"client"`` or ``"server"``; see the module docstring. + decodes: str = "client" + #: Whether decoding loses anything relative to what was encoded. + lossy: bool = False + #: What it costs, in one line, for a human reading a table of these. Empty + #: for a lossless interchange format, where there is nothing to warn about. + cost: str = "" + + def __post_init__(self) -> None: + if not isinstance(self.representation, Representation): + raise TypeError("representation must be an open4d.core.Representation") + if not self.suffix.startswith(".") or self.suffix != self.suffix.lower(): + raise ValueError(f"suffix {self.suffix!r} must be lowercase and dotted") + if self.decodes not in DECODERS: + raise ValueError( + f"decodes must be one of {', '.join(DECODERS)}; got {self.decodes!r}" + ) + if self.lossy and not self.cost: + raise ValueError( + f"{self.name}: a lossy codec must say what it costs — that line is " + "what reaches the person looking at the render" + ) + + @property + def key(self) -> tuple[Representation, str]: + return (self.representation, self.suffix) + + +_REGISTRY: dict[tuple[Representation, str], CodecSpec] = {} + + +def register(spec: CodecSpec, *, replace: bool = False) -> CodecSpec: + """Add ``spec`` to the registry and return it.""" + if not isinstance(spec, CodecSpec): + raise TypeError("spec must be a CodecSpec") + if spec.key in _REGISTRY and not replace: + existing = _REGISTRY[spec.key] + raise ValueError( + f"{spec.representation.value} {spec.suffix} is already registered as " + f"{existing.name!r}; pass replace=True to override" + ) + _REGISTRY[spec.key] = spec + return spec + + +def known() -> tuple[CodecSpec, ...]: + """Every registered codec, ordered by representation then suffix.""" + return tuple( + _REGISTRY[key] + for key in sorted( + _REGISTRY, key=lambda k: (list(Representation).index(k[0]), k[1]) + ) + ) + + +def for_frame(representation: Representation | str, suffix: str) -> CodecSpec: + """The codec a frame of this representation and suffix is in.""" + key = ( + representation + if isinstance(representation, Representation) + else Representation(representation) + ) + try: + return _REGISTRY[(key, suffix.lower())] + except KeyError: + offered = ", ".join( + spec.suffix for spec in known() if spec.representation is key + ) + raise KeyError( + f"no codec for a {key.value} frame ending {suffix!r}" + + (f"; registered: {offered}" if offered else "") + ) from None + + +def by_name(name: str) -> CodecSpec: + """A codec by its stable name.""" + for spec in known(): + if spec.name == name: + return spec + raise KeyError(f"no codec named {name!r}") + + +def for_representation(representation: Representation | str) -> tuple[CodecSpec, ...]: + """Every codec that produces this representation.""" + key = ( + representation + if isinstance(representation, Representation) + else Representation(representation) + ) + return tuple(spec for spec in known() if spec.representation is key) + + +def client_decodable( + representation: Representation | str | None = None, +) -> tuple[CodecSpec, ...]: + """Codecs a browser can decode, optionally for one representation. + + The answer to "can I stream this module to a browser at all", which used to + require reading the client's source. + """ + found = ( + known() + if representation is None + else for_representation(representation) + ) + return tuple(spec for spec in found if spec.decodes == "client") + + +def media_types() -> dict[str, str]: + """Every registered suffix, merged, for a server's extension map. + + A suffix shared between representations is fine as long as they agree on the + type, which is checked rather than assumed: disagreement would make a + frame's ``Content-Type`` depend on registration order. + """ + merged: dict[str, str] = {} + for spec in known(): + existing = merged.get(spec.suffix) + if existing is not None and existing != spec.media_type: + raise ValueError( + f"{spec.suffix!r} is registered as both {existing!r} and " + f"{spec.media_type!r}; a suffix must have one Content-Type" + ) + merged[spec.suffix] = spec.media_type + return merged + + +# ------------------------------------------------------------- the defaults --- +# Registered here rather than by the modules that produce them, so a server +# started against a bundle it did not write still knows how to send its frames +# and a reader can see the whole picture in one place. + +register(CodecSpec( + name="mesh-ply", + suffix=".ply", + representation=Representation.MESH, + # Interchange, not delivery: readable by anything, 13x larger than Draco. +)) +register(CodecSpec( + name="mesh-draco", + suffix=".drc", + representation=Representation.MESH, + lossy=True, + cost="positions quantised (0.005% of the diagonal at 14 bits) and duplicate " + "vertices merged", +)) +register(CodecSpec( + name="points-ply", + suffix=".ply", + representation=Representation.POINTS, +)) +register(CodecSpec( + name="points-draco", + suffix=".drc", + representation=Representation.POINTS, + lossy=True, + cost="positions quantised and duplicate points merged", +)) +register(CodecSpec( + name="3dgs-ply", + suffix=".ply", + representation=Representation.GAUSSIANS, + # The format every splat tool reads, and the reason a Vega export opens in + # SuperSplat or SIBR. Stores raw training parameters at float32. +)) +register(CodecSpec( + name="splat", + suffix=".splat", + representation=Representation.GAUSSIANS, + lossy=True, + cost="every spherical-harmonic band above degree 0 dropped, so appearance " + "stops changing with view direction; colour, opacity and rotation " + "quantised to 8 bits", +)) +register(CodecSpec( + name="jpeg", + suffix=".jpg", + representation=Representation.PIXELS, + media_type="image/jpeg", + lossy=True, + cost="lossy image compression, and a fixed viewpoint: pixels carry no " + "geometry, so there is no free camera", +)) +register(CodecSpec( + name="jpeg", + suffix=".jpeg", + representation=Representation.PIXELS, + media_type="image/jpeg", + lossy=True, + cost="lossy image compression, and a fixed viewpoint", +)) +register(CodecSpec( + name="png", + suffix=".png", + representation=Representation.PIXELS, + media_type="image/png", + cost="a fixed viewpoint: pixels carry no geometry", +)) +register(CodecSpec( + name="rerf", + suffix=".rerf", + representation=Representation.PIXELS, + decodes="server", + lossy=True, + cost="DCT and arithmetic coded, and undecodable in a browser at all — its " + "entropy coder ships only as a CPython 3.8 binary, so it is rendered " + "where the GPU is and arrives as pixels at a fixed viewpoint", +)) diff --git a/open4d/streamer/streamer/export.py b/open4d/streamer/streamer/export.py new file mode 100644 index 00000000..76738d1c --- /dev/null +++ b/open4d/streamer/streamer/export.py @@ -0,0 +1,221 @@ +"""Any `open4d.Sequence` as a bundle, so the client can play it. + +This is where the two halves of the repository meet. Open4D's own sequence model +reads a dozen mesh formats and every codec artifact it ships -- `open4d.load` +resolves them all, and #47 showed the pattern by making a raw V-DMC bitstream a +playable source without adding a viewer. What was missing was a way to hand one +of those sequences to a browser: `gs_tools` exports Gaussians, and nothing +exported meshes at all, so Open4D's mesh sequences could only be seen through +the Qt window that needs a display the GPU machine does not have. + +The conversion is small because the shapes already match. A `Sequence` is frames +in order; a bundle clip is frames in order plus a declared representation. The +frames themselves are written by `open4d.io.write_sequence`, which already emits +one ``frame_NNNNNN.ply`` per frame -- so this chooses the representation, records +the bounds and hands the rest to code that already existed. + +Lives here rather than in `gs_tools` because a mesh has nothing to do with +Gaussian splatting, and because `streamer` already depends on `open4d` for the +representation vocabulary. The dependency direction is unchanged: this imports +Open4D's public API, and no reconstruction module. +""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any + +from open4d.core import Representation, Sequence + +from . import bundle + +#: Interchange form: Open4D's only first-party mesh format needing no optional +#: dependency, and readable by anything. +FRAME_FORMAT = "ply" + +#: Delivery form. Draco compresses this repository's mesh sequence 12.9x -- 761 +#: kB a frame to 59 kB, which is 1.8 MB/s at 30 fps rather than 23 -- and the +#: client decodes it with the WASM decoder vendored under `client/vendor/draco`. +#: Lossy in two bounded ways: positions are quantised (at 14 bits the worst +#: vertex moved 0.0046% of the model's diagonal on that sequence) and duplicate +#: vertices are merged. Delivery, not archive. +DRACO_FORMAT = "draco" +FORMATS = (FRAME_FORMAT, DRACO_FORMAT) + +#: Position quantisation. 14 is DracoPy's own default and the knee of the curve +#: measured here: 11 bits saves a further 20% for eight times the error. +DRACO_QUANTIZATION_BITS = 14 + + +def representation_of(sequence: Sequence) -> Representation: + """Whether a sequence is a mesh or a point cloud, from its first frame. + + Read from the geometry rather than the file extension, because the same + ``.ply`` carries either. A sequence whose frames disagree is not something + this guesses at -- the first frame decides, and the bundle records it. + """ + if not len(sequence): + raise ValueError("sequence has no frames") + return sequence[0].geometry.representation + + +def _write_draco(sequence: Sequence, frames_at: Path, bits: int) -> list[Path]: + """One ``.drc`` per frame, encoded with the same Draco this repository vendors. + + Per frame rather than one container for the sequence, because a streaming + client fetches frames: `open4d.save(..., codec="draco")` writes a `.d4d` + holding the whole sequence, which is the right shape for an archive and the + wrong one for a wire. + """ + try: + import DracoPy + except ImportError as error: # pragma: no cover - depends on the environment + raise RuntimeError( + "Draco frames need the DracoPy binding: pip install 'open4d[draco]'" + ) from error + + written: list[Path] = [] + for index in range(len(sequence)): + geometry = sequence[index].geometry + triangles = getattr(geometry, "triangles", None) + payload = DracoPy.encode( + geometry.positions.astype("float32"), + None if triangles is None else triangles.astype("uint32"), + quantization_bits=bits, + ) + target = frames_at / f"frame_{index:06d}.drc" + target.write_bytes(payload) + written.append(target) + return written + + +def from_sequence( + sequence: Sequence, + out_dir: Path | str, + *, + name: str, + frame_format: str = FRAME_FORMAT, + quantization_bits: int = DRACO_QUANTIZATION_BITS, + scene: str | None = None, + method: str | None = None, + notes: list[str] | None = None, + detail: dict[str, Any] | None = None, +) -> bundle.Clip: + """Write ``sequence`` into ``out_dir`` as one clip and describe it. + + ``frame_format`` is ``"ply"`` for the interchange form or ``"draco"`` for the + compressed one; see :data:`DRACO_FORMAT` for what the second costs. + """ + from open4d.io import write_sequence + + if frame_format not in FORMATS: + raise ValueError( + f"unknown frame format {frame_format!r}; expected one of " + + ", ".join(FORMATS) + ) + out_dir = Path(out_dir).expanduser().resolve() + representation = representation_of(sequence) + frames_at = bundle.frame_dir(out_dir, name) + clip_name = frames_at.name + + if frame_format == DRACO_FORMAT: + paths = _write_draco(sequence, frames_at, quantization_bits) + written = sorted(str(path.relative_to(out_dir)) for path in paths) + else: + write_sequence(sequence, frames_at, format=FRAME_FORMAT, overwrite=True) + written = sorted( + str(path.relative_to(out_dir)) + for path in frames_at.glob(f"frame_*.{FRAME_FORMAT}") + ) + if not written: + raise ValueError(f"{name} produced no frames") + + lower = [float("inf")] * 3 + upper = [float("-inf")] * 3 + counts: list[int] = [] + for index in range(len(sequence)): + positions = sequence[index].geometry.positions + counts.append(int(len(positions))) + for axis in range(3): + lower[axis] = min(lower[axis], float(positions[:, axis].min())) + upper[axis] = max(upper[axis], float(positions[:, axis].max())) + + return bundle.Clip( + name=clip_name, + representation=representation.value, + scene=scene or clip_name, + method=method or representation.value, + # Carried through from the provider rather than assumed: a Sequence whose + # frames need a key frame first says so, and the client is then able to + # plan a seek instead of requesting a frame and hoping. Writing whole + # frames per file makes this independent in practice today -- the field + # is the path by which that stops being the only option. + dependency=bundle.dependency_field(sequence.dependency), + # Vertices per frame. The field is named for the Gaussian case that + # needed it first; for a mesh it is the vertex count, which is the same + # thing the viewer reports. + counts=counts, + frames=written, + bounds_min=lower, + bounds_max=upper, + notes=(notes or []) + ([ + f"frames are Draco at {quantization_bits}-bit position quantisation: " + "a delivery form, decoded in the browser, lossy in position and in " + "merging duplicate vertices", + ] if frame_format == DRACO_FORMAT else []), + detail={ + **(detail or {}), + "frame_format": frame_format, + **({"quantization_bits": quantization_bits} + if frame_format == DRACO_FORMAT else {}), + }, + ) + + +def from_source( + source: Path | str, + out_dir: Path | str, + *, + fps: float | None = None, + name: str | None = None, + frame_format: str = FRAME_FORMAT, + scene: str | None = None, + method: str | None = None, +) -> Path: + """Load whatever `open4d.load` accepts at ``source`` and serve it as a bundle. + + ``fps`` applies only to a source that carries no timing of its own -- a bare + directory of per-frame meshes. A source that does carry it, such as one + `open4d.io.write_sequence` wrote, keeps its own timestamps and `fps` is + ignored rather than made an error: asking to play a clip at a given rate is + a reasonable thing to say, and refusing the whole export over it would not + be. + """ + import open4d + from open4d.io import inspect_sequence + + source = Path(source).expanduser().resolve() + out_dir = Path(out_dir).expanduser().resolve() + declared = inspect_sequence(source).timing_source + with open4d.load(source, fps=fps if declared == "default" else None) as sequence: + clip = from_sequence( + sequence, + out_dir, + name=name or source.stem or source.name, + frame_format=frame_format, + scene=scene, + method=method, + notes=[ + f"{len(sequence)} frames loaded from {source.name} through " + "open4d.load — whatever format it was in", + ], + detail={"source": str(source)}, + ) + bundle.write( + out_dir, + title=f"{clip.representation} — {clip.name}", + source=str(source), + clips=[clip], + fps=int(fps) if fps else 30, + ) + return out_dir diff --git a/open4d/streamer/streamer/link.py b/open4d/streamer/streamer/link.py new file mode 100644 index 00000000..f7700919 --- /dev/null +++ b/open4d/streamer/streamer/link.py @@ -0,0 +1,365 @@ +"""A link with a capacity, so a measurement means something. + +Every transport in this package has run on loopback, which is not a network: no +bottleneck, no queue, no round trip worth the name. That made three things +untestable at once. `monitor` could say how many bytes moved but never how long +they *should* have taken; `policy` could choose a rung for a budget but nothing +could say what the budget was; and a claim like "a decoded Gaussian frame needs +129 MB/s" described what the content demands rather than what a link delivers. +This is the missing constraint. + +**Shaped here rather than with ``tc``.** Kernel shaping on the loopback +interface is more faithful -- it is the real queue, with real congestion +control above it -- and it was rejected for three reasons. It needs root. It is +machine-wide, so it perturbs everything else on a shared box, including the +other demos running on this one. And it cannot be exercised by a test, which +for a measurement instrument is disqualifying: an unreproducible bandwidth +figure is not a result. NeVo makes the same trade for the same reason, modelling +byte arrival from a trace rather than shaping a real interface. + +So be clear about what this is. It is a **single-queue bottleneck model**: +bytes leave in the order they were requested, at the capacity in force when +they leave, after a propagation delay. Queueing delay emerges from contention, +which is the behaviour that matters when four panes share one pipe. What it does +*not* reproduce is TCP: no congestion window, no slow start, no retransmission +timers. Loss is charged as the delay a retransmission would cost, not as a +failed request, because that is what an application above TCP actually +experiences. Anything whose answer depends on congestion-control dynamics needs +a real link, and this will mislead. + + link = Link(capacity=20e6, latency=0.020) # 20 Mbit/s, 20 ms one-way + server = serve(bundle_dir, link=link, block=False) + ... + link.observed() # what it actually carried + +**How accurate it is, measured** -- pulling 30 ReRF frames (1.0 MB) off a real +bundle through a real socket: + +=============== ================== ============== +configured delivered error +=============== ================== ============== +5 Mbit/s 4.9 Mbit/s 2.6% +20 Mbit/s 18.5 Mbit/s 7.4% +50 Mbit/s 42.5 Mbit/s 15.0% +5 ms one-way 5.8 ms per request under 1 ms +20 ms one-way 20.9 ms per request under 1 ms +=============== ================== ============== + +The error is a fixed cost per write, so it grows with the rate: ``time.sleep`` +overshoots by a fraction of a millisecond, and at 50 Mbit/s a 64 kB chunk is +only 10 ms of link time. It always errs the same way -- **under**-delivering, +never letting more through than configured -- which is the safe direction for a +shaper, since a measurement is then pessimistic rather than flattering. + +The useful range is where the link is genuinely the constraint. One pane of +this repository's content is 0.7-8 Mbit/s and one scene's whole ladder spans +6-71 Mbit/s, so the interesting decisions all sit in the band where this is +accurate to single digits. Above about 100 Mbit/s it is measuring Python. + +A `Link` is shared by every connection a server has open, deliberately: the +contention is the point. Reservations are serialised, so concurrent panes queue +behind one another exactly as they would behind a real bottleneck. +""" +from __future__ import annotations + +import random +import threading +import time +from bisect import bisect_right +from dataclasses import dataclass +from pathlib import Path +from typing import Any + +#: Ceiling for a link with no capacity limit, in bits per second. Large enough +#: to be unreachable, finite so the arithmetic never has to special-case it. +UNLIMITED = 1e15 + + +@dataclass(frozen=True) +class Trace: + """Capacity over time, as a step function. + + ``at`` and ``capacity`` are parallel: capacity ``capacity[i]`` holds from + ``at[i]`` until ``at[i + 1]``, and the last value holds for ever. Written + this way rather than as one rate per second so a trace can carry a sharp + drop at a known instant, which is the case an adaptive client is judged on. + + Replayed from the moment the link starts, and looping by default: a + measurement usually outlasts the trace, and stopping at the end would + quietly hand the client an unlimited link for the rest of the run. + + ``period`` is where a looping trace wraps, and it must be *past* the last + point rather than at it. Wrapping at the last point would give that point's + rate zero width, so a two-point trace would replay only its first rate for + ever -- silently, and looking like a working trace. Left unset it is one + more sample-interval past the end, which for the uniformly sampled traces + real measurements come in is exactly right. + """ + + at: tuple[float, ...] + capacity: tuple[float, ...] + loop: bool = True + period: float | None = None + + def __post_init__(self) -> None: + if len(self.at) != len(self.capacity): + raise ValueError("at and capacity must be the same length") + if not self.at: + raise ValueError("a trace needs at least one point") + if self.at[0] != 0.0: + raise ValueError("a trace must start at t=0") + if list(self.at) != sorted(self.at): + raise ValueError("trace times must be non-decreasing") + if any(rate <= 0 for rate in self.capacity): + raise ValueError("trace capacities must be positive") + if self.period is not None and self.period <= self.at[-1]: + raise ValueError( + f"period {self.period} must be past the last point {self.at[-1]}; " + "wrapping at it would give that rate no time at all" + ) + + @property + def duration(self) -> float: + """How long one pass takes, including the last point's own stretch.""" + if self.period is not None: + return self.period + if len(self.at) < 2: + return 0.0 # a constant trace has no period + return self.at[-1] + (self.at[-1] - self.at[-2]) + + def at_time(self, elapsed: float) -> float: + """Capacity in force ``elapsed`` seconds after the link started.""" + if elapsed < 0: + elapsed = 0.0 + if self.loop and self.duration > 0: + elapsed %= self.duration + index = bisect_right(self.at, elapsed) - 1 + return self.capacity[max(0, index)] + + @classmethod + def steps(cls, *pairs: tuple[float, float], loop: bool = True, + period: float | None = None) -> "Trace": + """``Trace.steps((0, 20e6), (5, 3e6))`` -- 20 Mbit/s, then 3 after 5 s.""" + return cls(at=tuple(p[0] for p in pairs), + capacity=tuple(p[1] for p in pairs), loop=loop, period=period) + + @classmethod + def read(cls, path: Path | str, *, loop: bool = True, + period: float | None = None) -> "Trace": + """A two-column text trace: `` `` per line. + + Blank lines and ``#`` comments are skipped, so a trace can carry a note + about where it came from -- which for a bandwidth trace is the most + important thing about it. + """ + points = [] + for line in Path(path).read_text().splitlines(): + line = line.split("#", 1)[0].strip() + if not line: + continue + when, rate = line.split() + points.append((float(when), float(rate))) + if not points: + raise ValueError(f"{path} holds no trace points") + return cls.steps(*points, loop=loop, period=period) + + +@dataclass +class Reservation: + """When a write may start and when it will have finished.""" + + bytes: int + starts_at: float + finishes_at: float + #: Seconds of the total that were queueing behind other traffic rather than + #: transmission. The signal that a link is saturated. + queued: float + #: Seconds charged for retransmission, when loss was configured. + retransmit: float + + +class Link: + """A shared bottleneck: a capacity, a delay, and optionally a trace. + + ``capacity`` is bits per second, ignored when a ``trace`` is given. + ``latency`` is the one-way propagation delay in seconds -- a response pays + it once, not per chunk. ``loss`` is a per-chunk probability; a lost chunk is + charged one round trip, the cost of noticing and resending. + + ``seed`` makes loss reproducible. A bandwidth figure that changes between + runs is not a measurement, so the default is seeded rather than random. + """ + + def __init__( + self, + *, + capacity: float = UNLIMITED, + latency: float = 0.0, + loss: float = 0.0, + trace: Trace | None = None, + seed: int = 0, + clock=time.monotonic, + ) -> None: + if capacity <= 0: + raise ValueError("capacity must be positive") + if latency < 0: + raise ValueError("latency must not be negative") + if not 0.0 <= loss < 1.0: + raise ValueError("loss must be a probability in [0, 1)") + self.capacity = float(capacity) + self.latency = float(latency) + self.loss = float(loss) + self.trace = trace + self._clock = clock + self._random = random.Random(seed) + self._lock = threading.Lock() + self._started = clock() + # When the link next falls idle. Reservations queue behind this, which + # is what makes concurrent connections contend rather than each + # believing it has the whole pipe. + self._free_at = self._started + self._bytes = 0 + self._reservations = 0 + self._queued = 0.0 + self._retransmit = 0.0 + self._busy = 0.0 + + # ----------------------------------------------------------- the model --- + def capacity_at(self, when: float) -> float: + """Capacity in force at an absolute clock time.""" + if self.trace is None: + return self.capacity + return self.trace.at_time(when - self._started) + + def reserve(self, size: int) -> Reservation: + """Book ``size`` bytes of link time, and say when they land. + + Serialised: the reservation starts when the link is next free, so a + second caller queues behind the first. That queueing *is* the effect a + bottleneck has, and modelling each connection as independent would make + contention -- the whole reason a budget has to be shared -- invisible. + """ + if size < 0: + raise ValueError("size must not be negative") + now = self._clock() + with self._lock: + starts_at = max(now, self._free_at) + queued = starts_at - now + rate = self.capacity_at(starts_at) + duration = (size * 8) / rate + retransmit = 0.0 + if self.loss and size and self._random.random() < self.loss: + # One round trip: the time to notice nothing arrived and send + # it again. Not a failure -- TCP hides loss from the + # application as delay, and pretending a frame vanished would + # model something the client never sees. + retransmit = 2 * self.latency + self._free_at = starts_at + duration + self._bytes += size + self._reservations += 1 + self._queued += queued + self._retransmit += retransmit + self._busy += duration + return Reservation( + bytes=size, starts_at=starts_at, + finishes_at=self._free_at + retransmit, + queued=queued, retransmit=retransmit, + ) + + def send(self, size: int, *, propagate: bool = False) -> float: + """Reserve ``size`` bytes and block until they would have arrived. + + ``propagate`` adds the one-way delay, and is set for the first write of + a response rather than every chunk: propagation is paid once by a + stream of bytes already in flight. + """ + reservation = self.reserve(size) + wait = reservation.finishes_at - self._clock() + if propagate: + wait += self.latency + if wait > 0: + time.sleep(wait) + return max(0.0, wait) + + # ------------------------------------------------------ what it carried --- + def observed(self) -> dict[str, Any]: + """What the link actually delivered. + + ``bits_per_second`` is the rate while it was *delivering*, not averaged + over the link's lifetime. The distinction is the whole usefulness of the + number and getting it wrong is not subtle: a client that fetches eight + frames and then waits has an idle link most of the time, and dividing + by wall clock reported 0.3 Mbit/s for a 25 Mbit/s pipe. A chooser + handed that drops every pane. + + Both are **cumulative since the link started**, which makes this a + summary of a run and not a current reading. A link on a trace has no + single rate, and once the trace has cycled these average every phase + of it together. A client that has to track a changing link needs a + recent-window estimate instead -- an exponentially weighted mean over + its own last few transfers, which is what ``RateMeter`` in + `streamer.client` is for. Use :meth:`reset` to scope this to a phase. + + So two numbers, because they answer different questions: + + ``bits_per_second`` + Bytes over the seconds the link spent transmitting them. An + estimate of *capacity*, which is what a budget wants. + ``utilisation`` + The fraction of the elapsed time it was transmitting at all. How + much of the pipe was being used, which is what says whether the + link is the constraint. + """ + with self._lock: + elapsed = self._clock() - self._started + busy = self._busy + payload = { + "bytes": self._bytes, + "reservations": self._reservations, + "elapsed_seconds": round(elapsed, 4), + "busy_seconds": round(busy, 4), + "queued_seconds": round(self._queued, 4), + "retransmit_seconds": round(self._retransmit, 4), + "configured_bits_per_second": ( + None if self.trace else round(self.capacity, 1) + ), + "latency_seconds": self.latency, + "loss": self.loss, + } + payload["bits_per_second"] = ( + round(payload["bytes"] * 8 / busy, 1) if busy > 0 else None + ) + payload["utilisation"] = round(busy / elapsed, 4) if elapsed > 0 else None + # Saturation: the share of the time bytes were in the system that they + # spent waiting for the link rather than using it. High means the queue + # is the constraint, which is when a chooser has to act. + total = self._queued + self._retransmit + payload["queueing_fraction"] = ( + round(total / (total + busy), 4) if total + busy > 0 else None + ) + return payload + + def reset(self) -> None: + """Zero the counters and restart the trace, e.g. between two runs.""" + with self._lock: + self._started = self._clock() + self._free_at = self._started + self._bytes = 0 + self._reservations = 0 + self._queued = 0.0 + self._retransmit = 0.0 + self._busy = 0.0 + + +def described(link: Link | None) -> str: + """One line naming a link's constraint, for a server to print.""" + if link is None: + return "unshaped: loopback, so any rate measured through it is the disk's" + if link.trace is not None: + rates = link.trace.capacity + return (f"trace: {min(rates) / 1e6:.1f}-{max(rates) / 1e6:.1f} Mbit/s over " + f"{link.trace.duration:.0f}s" + + (", looping" if link.trace.loop else "") + + f", {link.latency * 1000:.0f} ms one-way") + return (f"{link.capacity / 1e6:.1f} Mbit/s, {link.latency * 1000:.0f} ms " + f"one-way" + (f", {link.loss:.1%} loss" if link.loss else "")) diff --git a/open4d/streamer/streamer/live.py b/open4d/streamer/streamer/live.py new file mode 100644 index 00000000..16c73eb1 --- /dev/null +++ b/open4d/streamer/streamer/live.py @@ -0,0 +1,145 @@ +"""Clips that are rendered as they are watched. + +Everything else in a bundle is a directory of frames someone decoded earlier. +That is the wrong shape for two cases this repository already has working, and +both matter: + +* **A representation that cannot be decoded in a browser at all.** ReRF's + entropy coder ships only as a CPython 3.8 binary, so nothing client-side will + ever read it. Rendering it where the GPU is and sending pixels is not a + fallback, it is the only transport it has. +* **Bandwidth.** Measured on this repository's own content, one rendered view + costs about 47 kB a frame against 4.3 MB for a decoded Gaussian frame -- 1.4 + MB/s against 129 MB/s at 30 fps. For a fixed viewpoint, sending pixels wins by + two orders of magnitude, and it is the only transport here that works over a + link rather than a LAN. + +So a live clip carries a URL instead of a frame list, and the client points an +```` at it. MJPEG because it needs no JavaScript, no codec and no +negotiation -- ``multipart/x-mixed-replace`` is a browser feature -- and because +`vega.streaming.mjpeg_server` already serves it, which is what makes this +wiring rather than new machinery. + +**Live clips get their own scene.** A live renderer chooses its own camera, so +putting one beside a rig pose in Compare would break exactly the guarantee that +mode exists to make: every pane at the same pose. Its own scene means it is +never presented as comparable to something it is not. +""" + +from __future__ import annotations + +from pathlib import Path +from urllib.parse import urlparse + +from . import bundle + +#: Path prefix the bundle server proxies live streams under. +ROUTE_PREFIX = "live" + + +#: Where a stream's pixels come from at the moment they are requested. +#: +#: The transport does not say. MJPEG carries a decode-and-render-on-demand loop +#: and a slideshow of files equally well, and this repository has both -- Vega's +#: wall demo runs the full client path per frame, while NeVo's loops PNGs a +#: renderer wrote hours earlier because a NeRF frame takes about half a second +#: to ray-march. Presenting the second as live would be a claim the software +#: does not support, so a producer has to say which it is. +ORIGINS = { + #: Decoded and rendered per frame, on demand. The encode may well be offline + #: -- that is true of any video -- but the decode is happening now. + "rendered": "live", + #: Frames prepared earlier and replayed over the same transport. Nothing is + #: being computed while you watch. + "replay": "replay", +} + + +def mjpeg( + url: str, + *, + name: str, + origin: str, + scene: str | None = None, + method: str = "live", + notes: list[str] | None = None, + detail: dict | None = None, +) -> bundle.Clip: + """A clip fed by an MJPEG endpoint. + + ``origin`` is required and has no default: see :data:`ORIGINS`. A default + would let a producer mislabel a slideshow as live by saying nothing, which + is exactly the mistake this argument exists to prevent. + + ``url`` is where the renderer actually is, and it is recorded as the clip's + *upstream*. What the manifest hands the client is a path on the bundle + server instead -- ``live/``, which `streamer.server` proxies. + + That indirection is the whole point, and skipping it is a mistake worth + naming: a manifest that gives the browser ``http://127.0.0.1:8768/stream`` + is telling it to fetch port 8768 *on the machine the browser is running on*. + Open the page through a tunnel and that is the laptop, which has nothing + there, so the pane stays blank with no error. Proxying makes the bundle + server the single origin, so one forwarded port is enough and the bundle + stays portable. + """ + parsed = urlparse(url) + if parsed.scheme not in ("http", "https"): + raise ValueError(f"{url!r} is not an http(s) URL") + if "/" in name: + raise ValueError(f"{name!r} must be a single path segment") + if origin not in ORIGINS: + raise ValueError( + f"origin must be one of {', '.join(sorted(ORIGINS))}; got {origin!r}" + ) + + return bundle.validate( + bundle.Clip( + name=name, + representation="pixels", + scene=scene or name, + method=method, + frames=[], + stream={ + "url": f"{ROUTE_PREFIX}/{name}", + "protocol": "mjpeg", + "upstream": url, + "origin": origin, + }, + notes=(notes or []) + [ + "rendered on demand: decoded and drawn per frame while you watch" + if origin == "rendered" + else "replay: frames were rendered earlier and are being looped — " + "nothing is being computed while you watch", + "no frame list and nothing to scrub: the stream sets its own pace", + f"proxied by the bundle server from {parsed.netloc}, so the page " + "needs no access to that port itself", + ], + detail={**(detail or {}), "upstream": url}, + ) + ) + + +def attach(bundle_dir, clip, *, replace: bool = False) -> Path: + """Add a live clip to an existing bundle. See `bundle.add`. + + Kept as a name here because starting a renderer and pointing a bundle at it + is one action from a producer's point of view, and this is where the other + half of it lives. + """ + return bundle.add(bundle_dir, clip, replace=replace) + + +def upstreams(index: dict) -> dict[str, str]: + """Clip name -> upstream URL, for every live clip in a manifest. + + Read from the manifest rather than passed in separately, so a server started + against a bundle it did not write proxies exactly what that bundle declares. + """ + found: dict[str, str] = {} + for clip in index.get("clips", []): + stream = clip.get("stream") or {} + upstream = stream.get("upstream") + if upstream: + found[clip["name"]] = upstream + return found diff --git a/open4d/streamer/streamer/metrics.py b/open4d/streamer/streamer/metrics.py new file mode 100644 index 00000000..5b6f481a --- /dev/null +++ b/open4d/streamer/streamer/metrics.py @@ -0,0 +1,557 @@ +"""What the viewer actually received, scored against the real thing. + +`monitor` answers how many bytes moved. This answers the other half, and the +one a comparison needs: how good was the picture. Without it a bundle can show +two methods side by side but cannot say which is better, and "which is better" +is the question the whole platform exists to make answerable. + +**It works on a bundle, not on a method.** Any method whose output reaches a +bundle can be scored the same way, by the same code, against the same +reference -- which is the only arrangement under which two numbers are +comparable. A metric shipped inside each method would be nine metrics. + +How a pair is found: a scene's clips each name a ``method`` and a ``camera`` +station. The clip whose method is ``captured`` at a station is the reference -- +it is the photograph, not a reconstruction -- and every other clip at that same +station is scored against it. Same subject, same instant, same pose, which is +what the scene and camera fields are for. + +What it cannot do yet, stated rather than silently skipped: + +* **Geometry clips.** A Gaussian or mesh clip has no pixels until something + renders it from the reference's pose. That is a renderer, not a metric, and + it belongs with whatever can rasterise the representation. Those clips are + reported as unmeasured, with the reason. +* **Live clips.** Nothing to score frame-by-frame against; a stream has no + frame list. + +SSIM is implemented here rather than imported. `streamer` depends on `open4d` +and nothing else, and a metric is a poor reason to add scikit-image to a +package whose job is transport. It is the Wang et al. (2004) formulation with a +Gaussian window, and ``streamer_tests/test_metrics.py`` holds it to +scikit-image's implementation to 1e-4 wherever that library happens to be +installed -- so the shortcut is checked rather than trusted. +""" +from __future__ import annotations + +import json +import math +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + +from . import bundle, representations + +#: The method whose clips are the reference rather than a reconstruction. +REFERENCE_METHOD = "captured" + +#: What a clip's pixels are of, read from ``detail["depicts"]``. Only a clip +#: depicting the scene's *appearance* is attempting to reproduce the reference +#: photograph, so only that one can be scored against it. +#: +#: This exists because the alternative is worse than useless. A depth map +#: scored against a colour photograph comes out at 0.4 dB -- which reads as a +#: method that catastrophically failed, rather than as a comparison that was +#: never meaningful. "Not applicable" has to be expressible. +APPEARANCE = "appearance" + + +def depicts(clip) -> str: + """What a clip's pixels are of. Absent means appearance, so older bundles + and any method that never thought about this are measured as before.""" + return (clip.get("detail") or {}).get("depicts", APPEARANCE) + +#: Score every Nth frame by default. A 30-frame clip at 5 gives six samples, +#: which is enough for a mean that does not move between runs and cheap enough +#: to score 200 clips. Deterministic, not random: a research number that +#: changes when you re-run it is not a number. +EVERY = 5 + +# SSIM constants, from the paper and matching scikit-image's defaults. +K1, K2 = 0.01, 0.03 +WINDOW, SIGMA, TRUNCATE = 11, 1.5, 3.5 + + +def psnr(prediction, truth, *, data_range: float = 1.0) -> float: + """Peak signal-to-noise ratio in dB. Infinite for identical images.""" + import numpy as np + + error = float(np.mean((np.asarray(prediction, np.float64) + - np.asarray(truth, np.float64)) ** 2)) + if error <= 0.0: + return float("inf") + return float(10.0 * math.log10(data_range ** 2 / error)) + + +def _gaussian_kernel(sigma: float, truncate: float): + import numpy as np + + radius = int(truncate * sigma + 0.5) + x = np.arange(-radius, radius + 1, dtype=np.float64) + kernel = np.exp(-(x ** 2) / (2.0 * sigma ** 2)) + return kernel / kernel.sum() + + +def _blur_scipy(image): + """The same blur via scipy, when it happens to be installed. + + Measured on a 1280x960 colour frame: 287 ms against 696 ms for the numpy + fallback, because scipy's filter is compiled and this one builds large + sliding-window views. Identical output -- ``sigma`` and ``truncate`` give + the same 11-tap kernel, and ``reflect`` is what the fallback pads with -- + and ``test_both_blur_backends_agree`` holds them together. + + Not a declared dependency: `streamer` requires `open4d` and nothing else, + so this is used when present and skipped when not. + """ + from scipy.ndimage import gaussian_filter + + sigma = (SIGMA, SIGMA, 0) if image.ndim == 3 else (SIGMA, SIGMA) + return gaussian_filter(image, sigma=sigma, truncate=TRUNCATE, mode="reflect") + + +def _blur_numpy(image, kernel): + """Separable Gaussian blur with reflect padding, as scipy's default does. + + Two decisions here are both about speed, on images where it matters -- a + 1280x960 frame is 1.2 million pixels and one SSIM needs five blurs of it. + + A 2-D Gaussian is the outer product of two 1-D ones, so this is two dot + products over sliding windows rather than one 2-D convolution: 2n + multiplies per pixel instead of n squared. + + And a colour image is blurred in one pass with the channel axis carried + along, rather than once per channel. Same arithmetic, a third of the Python + and a third of the passes over memory. + + ``image`` is ``[H, W]`` or ``[H, W, C]``; the result has the same shape. + """ + import numpy as np + from numpy.lib.stride_tricks import sliding_window_view + + radius = (len(kernel) - 1) // 2 + spatial = ((radius, radius), (radius, radius)) + padding = spatial + (((0, 0),) if image.ndim == 3 else ()) + # "symmetric", not numpy's "reflect": scipy's `reflect` -- what + # scikit-image's SSIM filters with -- repeats the edge sample + # (a b c d -> d c b a | a b c d), whereas numpy's `reflect` omits it + # (a b c d -> d c b | a b c d). The two agree everywhere except a + # five-pixel border, which is why this was invisible on a 1280x960 frame + # at 1e-8 and only showed up on a 40x56 one. + padded = np.pad(image, padding, mode="symmetric") + rows = sliding_window_view(padded, len(kernel), axis=1) @ kernel + return sliding_window_view(rows, len(kernel), axis=0) @ kernel + + +def _blur_for(image): + """The fastest blur available, as a one-argument callable.""" + try: + import scipy.ndimage # noqa: F401 + except ImportError: + kernel = _gaussian_kernel(SIGMA, TRUNCATE) + return lambda array: _blur_numpy(array, kernel) + return _blur_scipy + + +def ssim(prediction, truth, *, data_range: float = 1.0) -> float: + """Structural similarity, Gaussian-windowed, averaged over channels. + + Colour images are scored per channel and averaged, which is what + scikit-image does with ``channel_axis``. The border is cropped by the + window radius because a window that hangs off the edge is measuring the + padding. + """ + import numpy as np + + a = np.asarray(prediction, np.float64) + b = np.asarray(truth, np.float64) + if a.shape != b.shape: + raise ValueError(f"shapes differ: {a.shape} vs {b.shape}") + + blur = _blur_for(a) + c1, c2 = (K1 * data_range) ** 2, (K2 * data_range) ** 2 + + ux, uy = blur(a), blur(b) + uxx, uyy, uxy = blur(a * a), blur(b * b), blur(a * b) + # Population covariance, not sample: scikit-image's + # use_sample_covariance=False, which is what the paper specifies. + vx, vy = uxx - ux * ux, uyy - uy * uy + vxy = uxy - ux * uy + + numerator = (2.0 * ux * uy + c1) * (2.0 * vxy + c2) + denominator = (ux * ux + uy * uy + c1) * (vx + vy + c2) + similarity = numerator / denominator + + # Cropped by the window radius, because a window hanging off the edge is + # measuring the padding. Then averaged over everything left, which for a + # colour image means over the channels too -- the same figure + # scikit-image's `channel_axis` produces. + pad = (WINDOW - 1) // 2 + return float(similarity[pad:-pad, pad:-pad].mean()) + + +def _load(path: Path, size=None): + """A frame as float [0, 1], resized to ``size`` if it does not match.""" + import numpy as np + from PIL import Image + + image = Image.open(path).convert("RGB") + if size is not None and image.size != size: + image = image.resize(size, Image.LANCZOS) + return np.asarray(image, np.float32) / 255.0 + + +@dataclass +class ClipScore: + """One rendition of one reconstruction clip, scored against its reference.""" + + scene: str + method: str + clip: str + #: Which rung this is. ``None`` for the clip's default rendition, which is + #: what a clip with a single quality level has. + variant: str | None + reference: str + camera: int | None + frames: int + psnr: float + ssim: float + worst_psnr: float + #: Set when the clip and its reference are different sizes and one was + #: resampled to match. Worth surfacing: a resize is not free, and a + #: comparison across scenes where only some were resized is not level. + resized: str | None = None + + def as_dict(self) -> dict[str, Any]: + payload = { + "scene": self.scene, "method": self.method, "clip": self.clip, + "variant": self.variant, + "reference": self.reference, "camera": self.camera, + "frames": self.frames, "psnr": round(self.psnr, 3), + "ssim": round(self.ssim, 5), "worst_psnr": round(self.worst_psnr, 3), + } + if self.resized: + payload["resized"] = self.resized + return payload + + +@dataclass +class Report: + """Everything that was scored, and everything that could not be.""" + + scores: list = field(default_factory=list) + unmeasured: list = field(default_factory=list) + + def by_scene(self) -> dict: + grouped: dict = {} + for score in self.scores: + grouped.setdefault(score.scene, []).append(score) + return grouped + + def by_method(self) -> dict: + grouped: dict = {} + for score in self.scores: + grouped.setdefault(score.method, []).append(score) + return grouped + + def as_dict(self) -> dict[str, Any]: + return { + "scores": [score.as_dict() for score in self.scores], + "unmeasured": self.unmeasured, + } + + +def _references(clips) -> dict: + """``(scene, camera) -> clip name`` for every reference clip.""" + found = {} + for clip in clips: + if clip.get("method") == REFERENCE_METHOD and clip.get("frames"): + found[(clip.get("scene"), clip.get("camera"))] = clip["name"] + return found + + +def measure( + bundle_dir: Path | str, + *, + scene: str | None = None, + every: int = EVERY, + limit: int | None = None, +) -> Report: + """Score every measurable clip in a bundle against its reference.""" + if every < 1: + raise ValueError("every must be at least 1") + root = Path(bundle_dir).expanduser().resolve() + index = bundle.read(root) + if not index: + raise FileNotFoundError(f"{root} has no {bundle.INDEX_NAME}") + + from PIL import Image + + clips = index.get("clips", []) + references = _references(clips) + report = Report() + + for clip in clips: + name = clip["name"] + if scene is not None and clip.get("scene") != scene: + continue + if clip.get("method") == REFERENCE_METHOD: + continue + if clip.get("stream"): + report.unmeasured.append( + {"clip": name, "why": "a live stream has no frame list to score"} + ) + continue + subject = depicts(clip) + if subject != APPEARANCE: + report.unmeasured.append({ + "clip": name, + "why": f"depicts {subject}, not the scene's appearance, so the " + "captured photograph is not a reference for it", + }) + continue + spec = representations.spec(clip["representation"]) + if spec.has_geometry: + report.unmeasured.append({ + "clip": name, + "why": f"{clip['representation']} has no pixels until something " + "renders it from the reference's pose", + }) + continue + reference = references.get((clip.get("scene"), clip.get("camera"))) + if reference is None: + report.unmeasured.append({ + "clip": name, + "why": f"no {REFERENCE_METHOD} clip at scene " + f"{clip.get('scene')!r} camera {clip.get('camera')}", + }) + continue + + truth_frames = next( + entry["frames"] for entry in clips if entry["name"] == reference + ) + sampled = list(range(0, min(len(clip["frames"]), len(truth_frames)), every)) + if limit: + sampled = sampled[:limit] + if not sampled: + report.unmeasured.append({"clip": name, "why": "no overlapping frames"}) + continue + + size = Image.open(root / truth_frames[0]).size + rendered_size = Image.open(root / clip["frames"][0]).size + resized = ( + None if rendered_size == size + else f"{rendered_size[0]}x{rendered_size[1]} -> {size[0]}x{size[1]}" + ) + + # Every rendition, not just the default: a ladder whose rungs have no + # measured quality is a ladder nothing can choose sensibly between. + renditions = [(None, clip["frames"], resized)] + for rung in bundle.variants_of(clip): + rung_size = Image.open(root / rung.frames[0]).size + renditions.append(( + rung.name, rung.frames, + None if rung_size == size + else f"{rung_size[0]}x{rung_size[1]} -> {size[0]}x{size[1]}", + )) + + for rung_name, frames, note in renditions: + peaks, structures = [], [] + for position in sampled: + if position >= len(frames): + break + prediction = _load(root / frames[position], size) + truth = _load(root / truth_frames[position]) + peaks.append(psnr(prediction, truth)) + structures.append(ssim(prediction, truth)) + if not peaks: + continue + finite = [value for value in peaks if math.isfinite(value)] + report.scores.append(ClipScore( + scene=clip.get("scene"), method=clip.get("method"), clip=name, + variant=rung_name, reference=reference, + camera=clip.get("camera"), frames=len(peaks), + psnr=sum(finite) / len(finite) if finite else float("inf"), + ssim=sum(structures) / len(structures), + worst_psnr=min(peaks), + resized=note, + )) + + return report + + +def _mean(values) -> float: + values = list(values) + return sum(values) / len(values) if values else float("nan") + + +def render_table(report: Report, *, per_clip: bool = False) -> str: + """The report as text: per scene and method, or per clip.""" + lines = [] + def rung(score): + return score.variant or "default" + + if per_clip: + lines.append(f"{'clip':<34}{'method':<11}{'rung':<9}{'cam':>4}" + f"{'PSNR':>8}{'SSIM':>9}{'worst':>8}") + lines.append("-" * 83) + for score in report.scores: + lines.append( + f"{score.clip:<34}{score.method:<11}{rung(score):<9}" + f"{'' if score.camera is None else score.camera:>4}" + f"{score.psnr:>8.2f}{score.ssim:>9.4f}{score.worst_psnr:>8.2f}" + ) + else: + # Rolled up per rung as well as per method: the whole point of a ladder + # is what each step costs and gains, and a mean across rungs would + # average that away into one meaningless number. + lines.append(f"{'scene':<18}{'method':<11}{'rung':<9}{'clips':>6}" + f"{'PSNR':>8}{'SSIM':>9}{'worst':>8}") + lines.append("-" * 69) + grouped: dict = {} + for score in report.scores: + grouped.setdefault( + (str(score.scene), score.method, rung(score)), [] + ).append(score) + for (scene, method, name), group in sorted(grouped.items()): + lines.append( + f"{scene:<18}{method:<11}{name:<9}{len(group):>6}" + f"{_mean(s.psnr for s in group):>8.2f}" + f"{_mean(s.ssim for s in group):>9.4f}" + f"{min(s.worst_psnr for s in group):>8.2f}" + ) + lines.append("-" * 69) + overall: dict = {} + for score in report.scores: + overall.setdefault((score.method, rung(score)), []).append(score) + for (method, name), group in sorted(overall.items()): + lines.append( + f"{'all':<18}{method:<11}{name:<9}{len(group):>6}" + f"{_mean(s.psnr for s in group):>8.2f}" + f"{_mean(s.ssim for s in group):>9.4f}" + f"{min(s.worst_psnr for s in group):>8.2f}" + ) + + # A lower rung is smaller by design, so only flag a resize on a default + # rendition -- there it means two methods were compared at different sizes, + # which is the thing worth knowing. + resized = [score for score in report.scores + if score.resized and score.variant is None] + if resized: + lines.append("") + lines.append(f"{len(resized)} clip(s) were resampled to their reference's " + "size; those numbers are not level with the rest:") + for score in resized[:5]: + lines.append(f" {score.clip}: {score.resized}") + if report.unmeasured: + lines.append("") + lines.append(f"{len(report.unmeasured)} clip(s) not measured:") + seen = {} + for entry in report.unmeasured: + seen.setdefault(entry["why"], []).append(entry["clip"]) + for why, names in seen.items(): + lines.append(f" {len(names)}x {why}") + lines.append(f" e.g. {names[0]}") + return "\n".join(lines) + + +def write_back(bundle_dir: Path | str, report: Report) -> Path: + """Store measured quality into the bundle, beside the bytes. + + This is what closes the loop. An exporter knows a rung's *size* -- it wrote + the files -- but not its quality, because quality is a comparison against a + reference the exporter has no opinion about. So the rung arrives with + ``bytes`` filled in and ``quality`` empty, and something has to put the + other half in. Until it does, a chooser reading the ladder can see what + each rung costs and not what it buys, which is half a decision. + + A rung's own ``quality`` mapping holds it. The default rendition has no + rung entry to put it in, so it goes in the clip's ``detail["quality"]`` -- + the same shape, one level up. + + The default's **byte count** is recorded here too, as + ``detail["bytes_per_frame"]``. A variant states its own size because an + exporter wrote it; the default's size is only on disk, so a consumer that + cannot stat the files -- a browser, which is the consumer that matters -- + had no way to know what the rendition it plays by default actually costs. + Without it a client can compare rungs it might switch to and not the one it + is already on. + + Only clips this report actually scored are touched: a partial run should + fill in what it measured and leave the rest alone rather than blanking it. + """ + root = Path(bundle_dir).expanduser().resolve() + index = bundle.read(root) + if not index: + raise FileNotFoundError(f"{root} has no {bundle.INDEX_NAME}") + + scored: dict = {} + for score in report.scores: + scored[(score.clip, score.variant)] = { + "psnr": round(score.psnr, 3), "ssim": round(score.ssim, 5), + } + + clips = [] + for entry in index.get("clips", []): + clip = bundle.Clip(**entry) + measured = scored.get((clip.name, None)) + if measured: + clip.detail = dict(clip.detail, quality=measured) + if clip.frames and not clip.stream: + total = sum((root / path).stat().st_size for path in clip.frames) + clip.detail = dict( + clip.detail, + bytes_per_frame=round(total / len(clip.frames), 1), + ) + if clip.variants: + updated = [] + for raw in clip.variants: + found = scored.get((clip.name, raw.get("name"))) + updated.append(dict(raw, quality=found) if found else raw) + clip.variants = updated + clips.append(clip) + + return bundle.write( + root, + title=index.get("title", root.name), + source=index.get("source", str(root)), + clips=clips, + fps=index.get("fps", 30), + scenes=index.get("scenes") or {}, + detail=index.get("detail"), + ) + + +def main(argv=None) -> int: + import argparse + + parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) + parser.add_argument("bundle", help="a bundle directory holding view.json") + parser.add_argument("--scene", help="only this scene") + parser.add_argument("--every", type=int, default=EVERY, + help="score every Nth frame; deterministic, so the number " + "does not move between runs") + parser.add_argument("--limit", type=int, help="at most this many frames per clip") + parser.add_argument("--per-clip", action="store_true", + help="one row per clip instead of a per-method rollup") + parser.add_argument("--json", action="store_true", help="machine-readable output") + parser.add_argument("--write", action="store_true", + help="store the measured quality into the bundle, beside each " + "rung's byte count, so a chooser can read both") + args = parser.parse_args(argv) + + report = measure(args.bundle, scene=args.scene, every=args.every, + limit=args.limit) + if args.write: + write_back(args.bundle, report) + print(f"wrote quality for {len(report.scores)} renditions into " + f"{args.bundle}\n") + if args.json: + print(json.dumps(report.as_dict(), indent=2)) + else: + print(render_table(report, per_clip=args.per_clip)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/open4d/streamer/streamer/monitor.py b/open4d/streamer/streamer/monitor.py new file mode 100644 index 00000000..c4eb726f --- /dev/null +++ b/open4d/streamer/streamer/monitor.py @@ -0,0 +1,133 @@ +"""What actually went over the wire. + +A streaming project should not have to guess at its own transport. NeVo models +byte arrival offline -- a bandwidth trace gives queueing delay, a loss trace +gives drops -- and this is the live counterpart at the other end: what was +requested, how big it was, and how long it took. It is the same measurement, +taken rather than simulated, and it is what makes a claim about a +representation's cost checkable instead of asserted. + +Deliberately counters and nothing more. No bitrate adaptation, no policy, no +history beyond a small ring for the recent tail: a monitor that decided things +would be a second scheduler, and there is not yet a second transport to adapt +between. Totals, a per-clip rollup, and the last few transfers answer "is this +bundle heavy, and where", which is the question actually being asked. + +Thread-safe because the server is threaded: one handler thread per connection, +all recording into one instance. +""" + +from __future__ import annotations + +import threading +import time +from collections import deque +from dataclasses import dataclass +from typing import Any + + +@dataclass(frozen=True) +class Transfer: + """One completed response.""" + + path: str + status: int + bytes: int + seconds: float + + @property + def bytes_per_second(self) -> float | None: + """None rather than infinity when the transfer was too fast to time.""" + return self.bytes / self.seconds if self.seconds > 0 else None + + +def clip_of(path: str) -> str: + """The clip a frame path belongs to. + + Bundle paths are ``/``, so the first segment is the clip. Used + to roll bytes up per clip, which is the granularity a bundle's weight is + actually discussed at -- "the Vega clip is 128 MB" rather than a list of 30 + frame sizes. + """ + trimmed = path.lstrip("/").split("?", 1)[0] + head, _, tail = trimmed.partition("/") + return head if tail else "" + + +class Monitor: + """Counters for one server's traffic.""" + + def __init__(self, *, tail: int = 32) -> None: + if tail < 0: + raise ValueError("tail must be nonnegative") + self._lock = threading.Lock() + self._started = time.monotonic() + self._requests = 0 + self._bytes = 0 + self._errors = 0 + self._by_clip: dict[str, dict[str, int]] = {} + self._tail: deque[Transfer] = deque(maxlen=tail) if tail else deque(maxlen=1) + self._keep_tail = tail > 0 + + def record(self, path: str, status: int, size: int, seconds: float) -> Transfer: + """Note one completed response and return it.""" + transfer = Transfer(path=path, status=int(status), bytes=int(size), + seconds=float(seconds)) + clip = clip_of(path) + with self._lock: + self._requests += 1 + self._bytes += transfer.bytes + if transfer.status >= 400: + self._errors += 1 + rollup = self._by_clip.setdefault(clip, {"requests": 0, "bytes": 0}) + rollup["requests"] += 1 + rollup["bytes"] += transfer.bytes + if self._keep_tail: + self._tail.append(transfer) + return transfer + + def snapshot(self) -> dict[str, Any]: + """A JSON-ready view of the counters. + + Taken under the lock so the totals and the rollup describe the same + instant; a reader that saw bytes from one moment and per-clip figures + from another would not add up, and that is exactly the kind of + discrepancy that gets blamed on the measurement. + """ + with self._lock: + elapsed = time.monotonic() - self._started + by_clip = { + name: dict(counts) + for name, counts in sorted( + self._by_clip.items(), key=lambda item: -item[1]["bytes"] + ) + } + tail = list(self._tail) if self._keep_tail else [] + requests, total, errors = self._requests, self._bytes, self._errors + return { + "requests": requests, + "bytes": total, + "errors": errors, + "elapsed_seconds": round(elapsed, 3), + "bytes_per_second": round(total / elapsed, 1) if elapsed > 0 else None, + "by_clip": by_clip, + "recent": [ + { + "path": item.path, + "status": item.status, + "bytes": item.bytes, + "seconds": round(item.seconds, 4), + } + for item in tail + ], + } + + def reset(self) -> None: + """Zero the counters, e.g. between two measured playbacks.""" + with self._lock: + self._started = time.monotonic() + self._requests = 0 + self._bytes = 0 + self._errors = 0 + self._by_clip.clear() + self._tail.clear() diff --git a/open4d/streamer/streamer/playback.py b/open4d/streamer/streamer/playback.py new file mode 100644 index 00000000..abd6f215 --- /dev/null +++ b/open4d/streamer/streamer/playback.py @@ -0,0 +1,574 @@ +"""Playing a bundle over a link, and scoring what the viewer got. + +`policy` chooses a rung for a budget at an instant. That is one decision, and a +playback is thousands of them, so the questions a comparison actually asks +cannot be answered by it alone: did this method stall, how often did its +quality jump about, and which of two methods survives a bad link better. Those +need the thing between a decision and a picture -- a **buffer**. + +A buffer is why streaming is not just downloading. Frames arrive at whatever +rate the link gives and are consumed at a fixed one, so the occupancy is the +integral of the difference. When it reaches zero playback stops, and *that* -- +not a low bitrate -- is what a viewer reports as the stream being bad. Every +adaptation decision is really about keeping this number off the floor. + +Three things are told apart here, because collapsing them hides the behaviour +worth measuring: + +**Stall.** The buffer emptied and playback halted. Unintentional, and the most +expensive thing that can happen. + +**Freeze.** A clip was *deliberately* not fetched, so it holds its last frame +while the others keep playing. A policy decision under deficit, not a failure, +and cheaper than a stall -- but it grows more expensive the longer it lasts, +or a chooser would freeze one pane for ever to protect the rest. + +**Switch.** The rung changed. Visible, so it is charged; the point of charging +it is that a marginal quality gain should not buy a visible jump. + +Simulated on a virtual clock rather than by sleeping. `Link.reserve` is pure +bookkeeping given a clock, which is what makes this exact and fast: a +sixty-second playback of nine panes runs in milliseconds and gives the same +answer every time. A run that has to be repeated to be believed is not a +measurement. + +Two things this model found about its own subject, both of which contradicted +an assumption made while writing it: + +* **The rate estimator decides whether stalls happen at all.** Estimating from + `Link.observed`, which is cumulative, a 30 -> 2.5 Mbit/s trace produced 30.4 + seconds of stall and not one freeze -- the average of both phases never + looked like a deficit, so nothing ever told the chooser to act. Swapping in + the same windowed estimate the browser uses turned that into 0 seconds of + stall and 56 of freeze: the same shortfall, taken as a decision instead of a + failure. +* **Scaling the budget by occupancy is nearly free of benefit here**, worth + 0.7 dB and no stalls, because the deficit handling already prevents them -- + while costing up to 1224 rung changes in thirty seconds if decisions are not + spaced out. See `buffer_budget` and ``decision_interval``. +""" +from __future__ import annotations + +from dataclasses import dataclass, field +from typing import Any, Callable, Mapping, Sequence + +from .link import Link +from .policy import Rung + +#: What a second of stalled playback costs, in the same units as quality (dB). +#: Large: the ABR literature puts a stall at several dB-seconds and viewers +#: agree -- a stall is reported as the stream being broken, a low rung as it +#: being soft. +STALL_PENALTY = 40.0 + +#: What a second of deliberate freeze costs. Cheaper than a stall, because the +#: rest of the view keeps playing, but not free. +FREEZE_PENALTY = 12.0 + +#: What one rung change costs. Charged per switch rather than per second: the +#: jump is the visible event. +SWITCH_PENALTY = 1.0 + +#: Samples the rate estimate halves over. The same rule and horizon as +#: ``RateMeter`` in the client, deliberately: this model exists to predict what +#: that client will do, and giving it a better estimator than the thing it +#: models would flatter every result. +RATE_HALF_LIFE = 8 + + +class RateEstimate: + """An exponentially weighted mean of recent transfer rates. + + Not `Link.observed`, which is cumulative and therefore cannot see a change: + on a trace that starts fast and collapses, a cumulative estimate keeps + reporting the average of both phases, the chooser never registers a + deficit, and the buffers drain while it holds a rung it cannot afford. + Measured on a 30 -> 2.5 Mbit/s trace: 30.4 seconds of stall and not one + freeze, because nothing ever told the chooser the link had gone. + """ + + def __init__(self, half_life: int = RATE_HALF_LIFE) -> None: + self.half_life = half_life + self.bits_per_second = 0.0 + self.samples = 0 + + def record(self, size: int, seconds: float) -> None: + if size <= 0 or seconds <= 0: + return + rate = size * 8 / seconds + self.samples += 1 + if self.samples == 1: + self.bits_per_second = rate + return + alpha = 1 - 0.5 ** (1 / self.half_life) + self.bits_per_second += alpha * (rate - self.bits_per_second) + + +@dataclass +class Buffer: + """Frames decoded and waiting to be shown, for one clip. + + Held in **seconds** rather than frames so a clip at a different frame rate + is comparable, and because the quantity a viewer experiences is time. + """ + + clip: str + #: Seconds of playback currently buffered. + seconds: float = 0.0 + #: Seconds this clip has spent stalled, and how many separate stalls. + stalled: float = 0.0 + stalls: int = 0 + #: Seconds spent frozen by decision, and how many separate freezes. + frozen: float = 0.0 + freezes: int = 0 + #: Rung changes, and the rung in force. + switches: int = 0 + rung: str | None = None + #: Frames fetched, the playback seconds they bought, and the sum of + #: quality x seconds. Mean quality is the ratio, so a rung that was on + #: screen briefly counts briefly -- an unweighted mean over rungs would + #: score a one-frame excursion the same as a minute of it. + frames: int = 0 + filled_seconds: float = 0.0 + quality_seconds: float = 0.0 + #: True while deliberately not being fetched. + freezing: bool = False + _stalling: bool = False + _froze: bool = False + + def fill(self, seconds: float, quality: float) -> None: + """A frame arrived: it buys ``seconds`` of playback at ``quality``.""" + self.seconds += seconds + self.frames += 1 + self.filled_seconds += seconds + self.quality_seconds += seconds * quality + self._stalling = False + + def drain(self, seconds: float) -> float: + """Consume ``seconds`` of playback; return the seconds spent stalled. + + A buffer that empties mid-interval stalls for the remainder, which is + why this returns a duration rather than a flag: charging a whole + interval would make the penalty depend on how finely time was stepped. + """ + if self.freezing: + self.frozen += seconds + if not self._froze: + self.freezes += 1 + self._froze = True + return 0.0 + self._froze = False + played = min(self.seconds, seconds) + self.seconds -= played + short = seconds - played + if short > 0: + self.stalled += short + if not self._stalling: + self.stalls += 1 + self._stalling = True + return short + + def use(self, rung: str | None) -> None: + if self.rung is not None and rung != self.rung: + self.switches += 1 + self.rung = rung + + @property + def mean_quality(self) -> float: + """Quality delivered, weighted by the playback time it bought.""" + if self.filled_seconds <= 0: + return 0.0 + return self.quality_seconds / self.filled_seconds + + +@dataclass +class Report: + """What a playback delivered.""" + + seconds: float + buffers: tuple[Buffer, ...] + #: What the link carried, from `Link.observed`. + link: Mapping[str, Any] + #: Seconds of playback the client asked for but could not show. + stalled: float + frozen: float + switches: int + mean_quality: float + score: float + + def as_dict(self) -> dict[str, Any]: + return { + "seconds": round(self.seconds, 3), + "stalled_seconds": round(self.stalled, 3), + "frozen_seconds": round(self.frozen, 3), + "switches": self.switches, + "mean_quality": round(self.mean_quality, 3), + "score": round(self.score, 3), + "link": dict(self.link), + "clips": [ + { + "clip": buffer.clip, "frames": buffer.frames, + "stalled": round(buffer.stalled, 3), "stalls": buffer.stalls, + "frozen": round(buffer.frozen, 3), "freezes": buffer.freezes, + "switches": buffer.switches, "rung": buffer.rung, + } + for buffer in self.buffers + ], + } + + +def buffer_budget(rate: float, occupancy: float, target: float) -> float: + """The rate to spend, given what the link gives and how full the buffer is. + + Rate alone is a poor guide and known to be: an estimate is a lagging + average, so a client that trusts it keeps buying the rung the link *used* + to afford and drains its buffer doing so. Occupancy is the state that + actually says whether the last few decisions were affordable, which is why + every buffer-based algorithm since BBA leans on it. + + The rule here is the simplest one with the right shape: scale the budget by + how full the buffer is against its target, clamped to [0.5, 1.5]. Below + target it spends under the estimate and refills; above, it spends over and + banks quality. Clamped both ways -- unclamped, an empty buffer would pick + the floor and stay there, and a full one would overshoot the link and empty + itself. + + **Measured, it earns much less here than the literature would suggest.** + On a 30 -> 6 -> 2.5 -> 14 Mbit/s trace it bought 0.7 dB of mean quality and + removed no stalls at all, against a flat rate-only budget. The reason is + that stall prevention is already being done by the deficit handling in + `Playback._even_split`, which refuses to commit to more clips than the rate + can carry -- so this is a second-order correction on top of a first-order + fix, and it costs churn to get. + + It is also what makes the budget noisy enough to need + ``decision_interval``: occupancy jitters every frame, and where the scaled + budget can reach across a rung boundary that jitter becomes 1224 rung + changes in thirty seconds. + + Kept anyway, because it is the standard technique and a result measured + against it should be comparable to published ones. But a reader deciding + whether their own client needs it should know it was worth 0.7 dB on this + content, at that cost, and not assume otherwise. + """ + if target <= 0: + return rate + return rate * max(0.5, min(1.5, occupancy / target)) + + +class Playback: + """Play clips over a link on a virtual clock, and score the result. + + ``ladders`` is one sequence of `policy.Rung` per clip, as + `policy.measured_rungs` returns. ``chooser`` decides a rung per clip from a + budget; the default is the even split the viewer uses, since what is being + compared is usually methods rather than choosers. + + One fetch at a time, matching the single queue `Link` models. A real client + opens several connections, which changes the queueing and not the + arithmetic -- and a serial fetcher is the honest simple case rather than an + optimistic one. + """ + + def __init__( + self, + ladders: Sequence[Sequence[Rung]], + link: Link, + *, + fps: int = 30, + target_buffer: float = 4.0, + #: How often to reconsider, in seconds of playback. + #: + #: Hysteresis for a noisy input, not a segment boundary. The budget is + #: scaled by occupancy, which jitters as frames arrive and drain, so a + #: per-frame decision chases the jitter -- but only when the scaled + #: budget can reach across a rung boundary. Measured on an eight-pane + #: ladder over thirty seconds: + #: + #: =========== ========== ========== + #: capacity per frame every 4 s + #: =========== ========== ========== + #: 14 Mbit/s 1224 24 + #: 18 Mbit/s 548 8 + #: 19 Mbit/s 8 8 + #: =========== ========== ========== + #: + #: At 19 the share already sits above the middle rung and nothing can + #: flip; below it the [0.5, 1.5] scaling straddles the boundary and + #: every frame is a coin toss. So the interval costs nothing where it + #: is not needed and is worth fifty-fold where it is, which is why it + #: defaults to on rather than being offered as a tuning knob. + decision_interval: float = 2.0, + metric: str = "psnr", + clock: Callable[[], float] | None = None, + chooser: Callable[..., Mapping[str, Rung]] | None = None, + stall_penalty: float = STALL_PENALTY, + freeze_penalty: float = FREEZE_PENALTY, + switch_penalty: float = SWITCH_PENALTY, + ) -> None: + if fps <= 0: + raise ValueError("fps must be positive") + if not ladders: + raise ValueError("nothing to play") + self.ladders = {rungs[0].clip: tuple(rungs) for rungs in ladders if rungs} + self.link = link + self.fps = fps + self.frame_seconds = 1.0 / fps + self.target_buffer = target_buffer + self.decision_interval = decision_interval + self.metric = metric + self.chooser = chooser or self._even_split + self.stall_penalty = stall_penalty + self.freeze_penalty = freeze_penalty + self.switch_penalty = switch_penalty + self.buffers = {clip: Buffer(clip=clip) for clip in self.ladders} + self.rate = RateEstimate() + self._now = 0.0 + self._decided_at = float("-inf") + self._picked: dict = {} + self._external_clock = clock + + # ------------------------------------------------------------ the clock --- + def _tick(self) -> float: + return self._now if self._external_clock is None else self._external_clock() + + def _advance(self, to: float) -> None: + """Move the clock forward, draining every buffer by the interval.""" + step = max(0.0, to - self._now) + self._now = to + if step <= 0: + return + for buffer in self.buffers.values(): + buffer.drain(step) + + # ---------------------------------------------------------- the decision --- + def _even_split(self, rate: float, buffers: Mapping[str, Buffer]) -> dict: + """The viewer's rule: equal shares, best rung each share affords. + + Budget per clip is scaled by that clip's own occupancy, so a pane that + has fallen behind buys cheaper frames until it catches up rather than + being punished for the average. + """ + # Nothing measured yet: fetch everything at the floor to find out. + # + # Without this a cold start deadlocks, and silently. With no estimate + # the rate is zero, so the deficit loop below freezes every clip; a + # frozen clip is never fetched; and an estimate only ever comes from a + # fetch. Measured before the guard existed: 30 seconds of playback, 30 + # seconds frozen, mean quality 0.0, and no error anywhere. A player + # starts at its lowest rung for exactly this reason. + if not getattr(self.rate, "samples", 0): + chosen = {} + for clip, rungs in self.ladders.items(): + buffers[clip].freezing = False + chosen[clip] = rungs[0] + return chosen + + # Deficit first: when the link cannot carry every clip's cheapest rung, + # something has to give, and freezing some so the rest play is better + # than starving all of them equally into a stall. + # + # Two criteria, in order. **Already frozen** comes first, so a freeze is + # sticky: a chooser that reconsidered from scratch each interval would + # thaw one pane and freeze another, and a viewer would see panes + # flickering in and out rather than a stable subset playing. Then + # **fullest buffer**, because the pane with least ahead of it is the one + # about to stall and so the one worth protecting. + # + # A frozen pane stays frozen for the rest of the run unless the rate + # recovers. That is deliberate, and it is why `FREEZE_PENALTY` is + # charged per second rather than per event: showing seven panes well + # beats showing eight badly, but not for ever, and the score has to say + # so rather than treating a permanent freeze as free. + floors = {clip: rungs[0].bits_per_second + for clip, rungs in self.ladders.items()} + playing = sorted( + self.ladders, + key=lambda clip: (buffers[clip].freezing, buffers[clip].seconds), + ) + while playing and sum(floors[clip] for clip in playing) > rate: + playing.pop() # the stickiest, best-buffered one + frozen = set(self.ladders) - set(playing) + + chosen: dict = {} + share = rate / max(1, len(playing)) if playing else 0.0 + for clip, rungs in self.ladders.items(): + buffers[clip].freezing = clip in frozen + if clip in frozen: + chosen[clip] = rungs[0] # nothing is fetched for it + continue + budget = buffer_budget(share, buffers[clip].seconds, self.target_buffer) + pick = rungs[0] + for rung in rungs: + if rung.bits_per_second <= budget: + pick = rung + chosen[clip] = pick + return chosen + + # ------------------------------------------------------------- the run --- + def run(self, seconds: float, *, warm: bool = True) -> Report: + """Play for ``seconds`` of wall time and report what was delivered. + + ``warm`` fills each buffer to target before the clock starts, which is + what a player's startup does. Without it every run begins with a stall + that says nothing about the link. + """ + if warm: + self._prefill() + horizon = self._now + seconds + + while self._now < horizon: + # The rate a client would have estimated: what the link has + # delivered so far. Cumulative here rather than windowed, which is + # a deliberate simplification -- a windowed estimate is what the + # browser does, and modelling its lag belongs with modelling the + # browser. + if self._now - self._decided_at >= self.decision_interval: + self._picked = self.chooser(self.rate.bits_per_second, self.buffers) + for clip, rung in self._picked.items(): + if not self.buffers[clip].freezing: + self.buffers[clip].use(rung.name) + self._decided_at = self._now + picked = self._picked + + # Fetch for whichever clip is closest to running dry, ignoring the + # frozen ones. That is the scheduling decision a buffer model + # exists to make: with one connection, who gets it next decides who + # stalls. + live = [clip for clip in self.buffers if not self.buffers[clip].freezing] + if not live: + self._advance(min(self._now + self.decision_interval, horizon)) + continue + clip = min(live, key=lambda name: self.buffers[name].seconds) + buffer = self.buffers[clip] + rung = picked[clip] + size = int(rung.bits_per_second / (8 * self.fps)) + + booked = self.link.reserve(size) + arrives = booked.finishes_at + self.link.latency + # Timed as a client would: from asking to having it, so queueing + # and propagation are in the estimate. A model that timed only + # transmission would estimate the link's capacity rather than the + # rate this client can actually achieve, and then overshoot it. + self.rate.record(size, max(1e-9, arrives - self._now)) + self._advance(min(arrives, horizon)) + if arrives <= horizon: + buffer.fill(self.frame_seconds, rung.utility(self.metric)) + if arrives >= horizon: + break + + self._advance(horizon) + return self._report(seconds) + + def _prefill(self) -> None: + """Buy each clip its target buffer before the clock matters.""" + picked = {clip: rungs[0] for clip, rungs in self.ladders.items()} + for clip, rung in picked.items(): + buffer = self.buffers[clip] + buffer.use(rung.name) + while buffer.seconds < self.target_buffer: + size = int(rung.bits_per_second / (8 * self.fps)) + booked = self.link.reserve(size) + arrives = booked.finishes_at + self.link.latency + self.rate.record(size, max(1e-9, arrives - self._now)) + self._now = max(self._now, arrives) + buffer.fill(self.frame_seconds, rung.utility(self.metric)) + + def _report(self, seconds: float) -> Report: + buffers = tuple(self.buffers.values()) + stalled = sum(buffer.stalled for buffer in buffers) + frozen = sum(buffer.frozen for buffer in buffers) + switches = sum(buffer.switches for buffer in buffers) + filled = sum(buffer.filled_seconds for buffer in buffers) + quality = (sum(buffer.quality_seconds for buffer in buffers) / filled + if filled else 0.0) + score = ( + quality + - self.stall_penalty * (stalled / max(1, len(buffers))) + - self.freeze_penalty * (frozen / max(1, len(buffers))) + - self.switch_penalty * switches / max(1, len(buffers)) + ) + return Report( + seconds=seconds, buffers=buffers, link=self.link.observed(), + stalled=stalled, frozen=frozen, switches=switches, + mean_quality=quality, score=score, + ) + + +def render_table(rows: Sequence[tuple[str, Report]]) -> str: + """One row per condition: what the viewer got, and what it scored.""" + lines = [f"{'condition':<18}{'stalled':>9}{'frozen':>9}{'switches':>10}" + f"{'quality':>9}{'score':>9}"] + lines.append("-" * len(lines[0])) + for label, report in rows: + lines.append( + f"{label:<18}{report.stalled:>8.1f}s{report.frozen:>8.1f}s" + f"{report.switches:>10}{report.mean_quality:>9.2f}" + f"{report.score:>9.2f}" + ) + return "\n".join(lines) + + +def main(argv=None) -> int: + import argparse + import json + + from .link import Link, Trace + from .policy import measured_rungs + + parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) + parser.add_argument("bundle") + parser.add_argument("--scene", help="only this scene's clips") + parser.add_argument("--seconds", type=float, default=30.0) + parser.add_argument("--capacity", type=float, action="append", default=None, + help="Mbit/s to play at; repeat to sweep") + parser.add_argument("--trace", help="a two-column bandwidth trace to replay") + parser.add_argument("--latency", type=float, default=0.02, + help="one-way delay in seconds") + parser.add_argument("--target-buffer", type=float, default=3.0) + parser.add_argument("--decision-interval", type=float, default=2.0) + parser.add_argument("--json", action="store_true") + args = parser.parse_args(argv) + + ladders = measured_rungs(args.bundle, scene=args.scene) + if not ladders: + print("no clip here has more than one rendition; nothing to adapt") + return 0 + + def play(link): + return Playback( + ladders, link, target_buffer=args.target_buffer, + decision_interval=args.decision_interval, + # A frozen virtual clock: the model advances time itself, and a real + # one would let the machine's scheduling leak into the result. + clock=None, + ).run(args.seconds) + + rows = [] + if args.trace: + trace = Trace.read(args.trace) + rows.append((f"trace {args.trace.split('/')[-1]}", + play(Link(trace=trace, latency=args.latency, + clock=lambda: 0.0)))) + else: + floor = sum(rungs[0].bits_per_second for rungs in ladders) / 1e6 + ceiling = sum(rungs[-1].bits_per_second for rungs in ladders) / 1e6 + rates = args.capacity or [ceiling * 1.2, ceiling * 0.5, floor, floor * 0.5] + for mbit in rates: + rows.append((f"{mbit:.1f} Mbit/s", + play(Link(capacity=mbit * 1e6, latency=args.latency, + clock=lambda: 0.0)))) + + if args.json: + print(json.dumps([{"condition": label, **report.as_dict()} + for label, report in rows], indent=2)) + return 0 + print(f"{len(ladders)} panes, {args.seconds:.0f}s each, " + f"target buffer {args.target_buffer:.0f}s, " + f"deciding every {args.decision_interval:.0f}s\n") + print(render_table(rows)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/open4d/streamer/streamer/policy.py b/open4d/streamer/streamer/policy.py new file mode 100644 index 00000000..bf999913 --- /dev/null +++ b/open4d/streamer/streamer/policy.py @@ -0,0 +1,519 @@ +"""Choosing a rung: what to send when you cannot send everything. + +The last piece of the loop. `bundle` lets a clip carry quality rungs, `metrics` +measures what each one costs and buys, and this decides. Without it a ladder is +data nobody reads: playback takes the default rendition every time, however +little bandwidth there is. + +The problem is a **multiple-choice knapsack**. Several clips play at once -- a +comparison view is four panes, a wall is nine -- each offering one rung from +its own ladder, and the total has to fit a budget. Maximise the weighted +quality of what is chosen. + +Solved exactly, by dynamic programming over a quantised budget, rather than +greedily. Greedy "best quality per byte" is the standard approximation and it +is wrong in a way that matters here: with rungs this coarse (3.2x and 10.6x +apart) a greedy pass will spend everything on the first clip it considers and +leave the rest at their floor, which is exactly the lopsided allocation a +comparison view must not have. The problem size makes exactness free -- nine +clips times three rungs is nothing. + +What this deliberately does **not** model is buffering. A real client also +tracks per-clip buffer occupancy, distinguishes a deliberate freeze from a +stall, and penalises churn across segments. Those need a link that can starve +one, and this repository has no constrained transport yet: everything runs on +loopback. Building stall accounting now would mean writing policy against a +situation that cannot be produced or measured, so what is here is the +allocation, and `switch_penalty` is the one piece of dynamics that can be +tested without a network. + +Utility is quality in dB by default, which is a choice worth naming. Summing +decibels is not physically meaningful -- they are a log scale -- but it is what +the ABR literature optimises and it has the right shape: the gain from a bad +rung to a mediocre one exceeds the gain from a good one to a slightly better +one. Pass ``metric="ssim"`` for a bounded alternative, or a callable for +anything else. +""" +from __future__ import annotations + +import math +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Callable, Iterable, Mapping, Sequence + +from . import bundle + +#: Budget resolution for the dynamic program, in bits per second. Finer costs +#: proportionally more time for a decision no one could act on: 10 kbit/s is +#: well below the difference between any two rungs here. +QUANTUM = 10_000 + +#: Utility of a rendition with no measured quality. Chosen so an unmeasured +#: rung is never preferred to a measured one, rather than defaulting to zero +#: and being silently avoided -- which would look like a policy decision. +UNMEASURED = float("-inf") + + +@dataclass(frozen=True) +class Rung: + """One rendition of one clip, as a chooser sees it.""" + + clip: str + #: ``None`` for the clip's default rendition. + variant: str | None + bits_per_second: float + #: Measured, from `metrics`. Empty when nothing has scored this rendition. + quality: Mapping[str, float] + + @property + def name(self) -> str: + return self.variant or "default" + + def utility(self, metric: str | Callable[[Mapping[str, float]], float]) -> float: + if callable(metric): + return metric(self.quality) + value = self.quality.get(metric) + return UNMEASURED if value is None else float(value) + + +@dataclass(frozen=True) +class Choice: + """The rung picked for one clip.""" + + clip: str + variant: str | None + bits_per_second: float + utility: float + weight: float + + @property + def name(self) -> str: + return self.variant or "default" + + +@dataclass(frozen=True) +class Selection: + """What a chooser decided, and what it could not fit.""" + + choices: tuple[Choice, ...] + #: Clips left out because even their cheapest rung did not fit. Named + #: rather than silently omitted: a viewer showing three of four panes has + #: to be able to say why the fourth is missing. + dropped: tuple[str, ...] + budget: float + bits_per_second: float + utility: float + + @property + def headroom(self) -> float: + return self.budget - self.bits_per_second + + def of(self, clip: str) -> Choice: + for choice in self.choices: + if choice.clip == clip: + return choice + raise KeyError(f"{clip!r} was not chosen for; dropped: {self.dropped}") + + def as_dict(self) -> dict[str, Any]: + return { + "budget_mbit_s": round(self.budget / 1e6, 4), + "spent_mbit_s": round(self.bits_per_second / 1e6, 4), + "utility": round(self.utility, 4), + "dropped": list(self.dropped), + "choices": [ + { + "clip": choice.clip, "rung": choice.name, + "mbit_s": round(choice.bits_per_second / 1e6, 4), + "utility": round(choice.utility, 3), + "weight": choice.weight, + } + for choice in self.choices + ], + } + + +def rungs_of(clip: bundle.Clip | Mapping[str, Any], fps: int) -> tuple[Rung, ...]: + """Every rendition of a clip, cheapest first, as `Rung` objects. + + The default rendition is one of them. It has no variant entry, so its rate + comes from the clip's own frames and its quality from + ``detail["quality"]`` -- where `metrics.write_back` puts it. + """ + name = clip.name if isinstance(clip, bundle.Clip) else clip["name"] + frames = clip.frames if isinstance(clip, bundle.Clip) else clip.get("frames") or [] + detail = clip.detail if isinstance(clip, bundle.Clip) else clip.get("detail") or {} + found = [] + default_bytes = detail.get("bytes_per_frame") + if frames and default_bytes: + found.append(Rung( + clip=name, variant=None, + bits_per_second=float(default_bytes) * 8 * fps, + quality=dict(detail.get("quality") or {}), + )) + for variant in bundle.variants_of(clip): + found.append(Rung( + clip=name, variant=variant.name, + bits_per_second=variant.bitrate(fps), + quality=dict(variant.quality), + )) + return tuple(sorted(found, key=lambda rung: rung.bits_per_second)) + + +def measured_rungs( + bundle_dir: Path | str, *, scene: str | None = None, method: str | None = None +) -> tuple[tuple[Rung, ...], ...]: + """Every laddered clip's rungs, read off a bundle. + + A clip's default rendition needs a byte count that the manifest does not + carry, so it is measured off disk here -- once, and only for clips that + have a ladder at all. A clip with one rendition is not a choice and is left + out: including it would let the budget be spent on something a chooser + cannot change. + """ + root = Path(bundle_dir).expanduser().resolve() + index = bundle.read(root) + if not index: + raise FileNotFoundError(f"{root} has no {bundle.INDEX_NAME}") + fps = index.get("fps", 30) + + ladders = [] + for entry in index.get("clips", []): + if not entry.get("variants"): + continue + if scene is not None and entry.get("scene") != scene: + continue + if method is not None and entry.get("method") != method: + continue + frames = entry.get("frames") or [] + detail = dict(entry.get("detail") or {}) + if frames and "bytes_per_frame" not in detail: + total = sum((root / path).stat().st_size for path in frames) + detail["bytes_per_frame"] = total / len(frames) + ladders.append(rungs_of(dict(entry, detail=detail), fps)) + return tuple(ladder for ladder in ladders if len(ladder) > 1) + + +def choose( + ladders: Iterable[Sequence[Rung]], + *, + budget: float, + weights: Mapping[str, float] | None = None, + metric: str | Callable[[Mapping[str, float]], float] = "psnr", + previous: Mapping[str, str | None] | None = None, + switch_penalty: float = 0.0, + quantum: int = QUANTUM, +) -> Selection: + """Pick one rung per ladder, maximising weighted utility under ``budget``. + + ``budget`` is bits per second. ``weights`` scales a clip's utility by how + much it matters -- which pane is being looked at, in a viewer. ``previous`` + and ``switch_penalty`` discourage changing a clip's rung from one decision + to the next: churn is visible, so a marginal gain should not buy it. + + Returns the selection. When even the cheapest rung of everything exceeds + the budget, the lowest-weight clips are dropped until it fits, and they are + named in ``Selection.dropped`` -- there is no rung cheap enough, and + pretending otherwise would overrun the budget silently. + """ + ladders = [tuple(ladder) for ladder in ladders if ladder] + if budget < 0: + raise ValueError("budget must not be negative") + if quantum <= 0: + raise ValueError("quantum must be positive") + weights = dict(weights or {}) + previous = dict(previous or {}) + + def weight(clip: str) -> float: + return float(weights.get(clip, 1.0)) + + def score(rung: Rung) -> float: + value = rung.utility(metric) + if value == UNMEASURED: + return UNMEASURED + value *= weight(rung.clip) + if switch_penalty and rung.clip in previous: + if previous[rung.clip] != rung.variant: + value -= switch_penalty * weight(rung.clip) + return value + + def quanta(rate: float) -> int: + """A rate as whole budget quanta, rounded up. + + Up, not nearest: a rung must never be costed below what it spends, or a + selection can overrun the budget by a quantum per clip. + """ + return int(math.ceil(rate / quantum)) + + # Deficit: not even the floor fits. Drop by lowest weight, then by the most + # expensive floor, so what goes is what matters least and costs most. + # + # Costed in *quanta*, matching the dynamic program below. Comparing exact + # rates here and rounded costs there lets the two disagree at the margin: + # the floor would look affordable, the program would find it infeasible, + # and every clip would come back dropped with no reason given. + dropped: list[str] = [] + every_ladder = list(ladders) + slots = int(budget // quantum) + while ladders and sum(quanta(ladder[0].bits_per_second) for ladder in ladders) > slots: + victim = min( + ladders, + key=lambda ladder: (weight(ladder[0].clip), + -ladder[0].bits_per_second, + ladder[0].clip), + ) + dropped.append(victim[0].clip) + ladders = [ladder for ladder in ladders if ladder is not victim] + + # No early return when everything was dropped: the repair pass below is + # what puts back a clip that the rounded-up floor made look unaffordable, + # and with one clip in the ladder that is *every* clip. Skipping to a + # result here reported an empty selection for a budget that fits. + + # Exact DP over the quantised budget. cell[b] is the best utility using at + # most b quanta, with a back-pointer per clip so the choice can be recovered. + best = [0.0] * (slots + 1) + taken: list[list[Rung | None]] = [[None] * (slots + 1)] + for ladder in ladders: + nxt = [-math.inf] * (slots + 1) + picks: list[Rung | None] = [None] * (slots + 1) + for rung in ladder: + value = score(rung) + if value == UNMEASURED: + continue + cost = quanta(rung.bits_per_second) + for spend in range(cost, slots + 1): + candidate = best[spend - cost] + value + if candidate > nxt[spend]: + nxt[spend] = candidate + picks[spend] = rung + # Monotone fill: a larger budget can always do at least as well, and + # carrying the better cell forward is what makes the back-pointers + # recoverable without storing the whole table per clip. + for spend in range(1, slots + 1): + if nxt[spend - 1] > nxt[spend]: + nxt[spend] = nxt[spend - 1] + picks[spend] = picks[spend - 1] + best = nxt + taken.append(picks) + + # Walk the back-pointers from the fullest cell. + choices: list[Choice] = [] + spend = slots + for index in range(len(ladders), 0, -1): + rung = taken[index][spend] + if rung is None: + # No affordable measured rung for this clip; it is dropped rather + # than shown at an unknown quality. + dropped.append(ladders[index - 1][0].clip) + continue + choices.append(Choice( + clip=rung.clip, variant=rung.variant, + bits_per_second=rung.bits_per_second, + utility=rung.utility(metric), weight=weight(rung.clip), + )) + spend -= quanta(rung.bits_per_second) + + choices.reverse() + choices, dropped = _repair( + choices, every_ladder, budget, score, metric, dropped, weight) + return Selection( + choices=tuple(choices), + dropped=tuple(sorted(set(dropped))), + budget=budget, + bits_per_second=sum(choice.bits_per_second for choice in choices), + utility=sum(choice.utility * choice.weight for choice in choices), + ) + + +def _repair( + choices: list[Choice], + ladders: Sequence[Sequence[Rung]], + budget: float, + score: Callable[[Rung], float], + metric, + dropped: Sequence[str], + weight: Callable[[str], float], +) -> tuple[tuple[Choice, ...], list[str]]: + """Spend headroom the quantisation hid, then stop. + + Costs are rounded **up** to whole quanta so a selection can never overrun + the budget. The price is up to one quantum of phantom cost per clip, and + across eight clips it accumulates into a decision that is visibly wrong at + the boundaries. Both were measured on `g_thomas`: + + * at exactly the all-default budget, the program left one clip a rung down, + giving up 2.4 dB to save 80 kbit/s that did not exist; + * at exactly the all-lowest budget, the deficit check dropped a clip whose + cheapest rung did fit, showing seven panes instead of eight. + + So two passes against the *exact* rates: put back a clip that fits after + all, then upgrade while anything still fits. Greedy is safe here in a way + it is not for the allocation itself -- each step strictly increases utility + and strictly decreases headroom, so it terminates, and neither pass can do + worse than what the program returned. + """ + by_clip = {ladder[0].clip: ladder for ladder in ladders} + dropped = list(dropped) + + # Re-add, cheapest first, so the most recoverable clip goes back first. + still_dropped: list[str] = [] + for clip in sorted(dropped, key=lambda name: ( + by_clip[name][0].bits_per_second if name in by_clip else 0.0)): + ladder = by_clip.get(clip) + spent = sum(choice.bits_per_second for choice in choices) + if ladder is None or spent + ladder[0].bits_per_second > budget: + still_dropped.append(clip) + continue + floor_rung = ladder[0] + if floor_rung.utility(metric) == UNMEASURED: + still_dropped.append(clip) # unmeasured: no basis to show it + continue + choices.append(Choice( + clip=clip, variant=floor_rung.variant, + bits_per_second=floor_rung.bits_per_second, + utility=floor_rung.utility(metric), weight=weight(clip), + )) + dropped = still_dropped + if not choices: + return (), dropped + # The rung each clip is currently on, from the ladder itself rather than + # rebuilt from the Choice: `score` may apply a weight and a switch penalty, + # and a reconstructed rung would be scored against the wrong history. + current = { + choice.clip: next( + rung for rung in by_clip[choice.clip] if rung.variant == choice.variant + ) + for choice in choices + if choice.clip in by_clip + } + picked = {choice.clip: choice for choice in choices} + + while True: + spent = sum(choice.bits_per_second for choice in picked.values()) + best_gain, best_upgrade = 0.0, None + for clip, choice in picked.items(): + if clip not in current: + continue + here = score(current[clip]) + for rung in by_clip[clip]: + extra = rung.bits_per_second - choice.bits_per_second + if extra <= 0 or spent + extra > budget: + continue + gain = score(rung) - here + if gain > best_gain: + best_gain, best_upgrade = gain, (clip, rung) + if best_upgrade is None: + break + clip, rung = best_upgrade + current[clip] = rung + picked[clip] = Choice( + clip=clip, variant=rung.variant, + bits_per_second=rung.bits_per_second, + utility=rung.utility(metric), weight=picked[clip].weight, + ) + + order = [choice.clip for choice in choices] + return tuple(picked[clip] for clip in order), dropped + + +def _short(names: Sequence[str]) -> dict[str, str]: + """Clip names with their shared prefix removed. + + Eight stations of one object share 19 of 21 characters, so a table of full + names is a wall in which only the last two digits carry information. + """ + if len(names) < 2: + return {name: name for name in names} + shared = 0 + for position in range(min(len(name) for name in names)): + if len({name[position] for name in names}) > 1: + break + shared = position + 1 + # Back off to the last separator, so a name is cut at a boundary rather + # than mid-token. + prefix = names[0][:shared] + cut = max(prefix.rfind("-"), prefix.rfind("_"), prefix.rfind("/")) + 1 + return {name: (name[cut:] or name) for name in names} + + +def render_table(selections: Sequence[Selection]) -> str: + """One row per budget, showing which rung each clip landed on.""" + if not selections: + return "nothing to choose between" + clips: list[str] = [] + for selection in selections: + for choice in selection.choices: + if choice.clip not in clips: + clips.append(choice.clip) + short = _short(clips) + width = max(9, max((len(short[name]) for name in clips), default=9)) + 2 + + header = (f"{'budget':>9}{'spent':>9}{'mean':>8} " + + "".join(f"{short[name]:<{width}}" for name in clips)) + lines = [header, "-" * len(header)] + for selection in selections: + picked = {choice.clip: choice for choice in selection.choices} + mean = ( + sum(choice.utility for choice in selection.choices) / len(selection.choices) + if selection.choices else float("nan") + ) + row = (f"{selection.budget / 1e6:>8.1f}M{selection.bits_per_second / 1e6:>8.1f}M" + f"{mean:>8.2f} ") + for name in clips: + choice = picked.get(name) + row += f"{(choice.name if choice else '-'):<{width}}" + lines.append(row.rstrip()) + dropped = {name for selection in selections for name in selection.dropped} + if dropped: + lines.append("") + lines.append("a dash is a clip dropped at that budget: even its cheapest " + "rung did not fit") + return "\n".join(lines) + + +def main(argv=None) -> int: + import argparse + import json + + parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) + parser.add_argument("bundle", help="a bundle directory holding view.json") + parser.add_argument("--scene", help="only this scene's clips") + parser.add_argument("--method", help="only this method's clips") + parser.add_argument("--budget", type=float, action="append", default=None, + help="Mbit/s to spend; repeat to sweep several") + parser.add_argument("--metric", default="psnr", choices=("psnr", "ssim")) + parser.add_argument("--switch-penalty", type=float, default=0.0, + help="utility charged for moving a clip off its previous rung") + parser.add_argument("--json", action="store_true") + args = parser.parse_args(argv) + + ladders = measured_rungs(args.bundle, scene=args.scene, method=args.method) + if not ladders: + print("no clip in this bundle has more than one rendition; nothing to choose") + return 0 + + floor = sum(ladder[0].bits_per_second for ladder in ladders) + ceiling = sum(ladder[-1].bits_per_second for ladder in ladders) + budgets = ( + [value * 1e6 for value in args.budget] if args.budget + else [floor * 0.5, floor, (floor + ceiling) / 2, ceiling, ceiling * 1.5] + ) + + selections = [ + choose(ladders, budget=budget, metric=args.metric, + switch_penalty=args.switch_penalty) + for budget in budgets + ] + if args.json: + print(json.dumps([s.as_dict() for s in selections], indent=2)) + return 0 + + print(f"{len(ladders)} laddered clips, " + f"{floor / 1e6:.1f}-{ceiling / 1e6:.1f} Mbit/s between all-lowest and " + f"all-highest\n") + print(render_table(selections)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/open4d/streamer/streamer/representations.py b/open4d/streamer/streamer/representations.py new file mode 100644 index 00000000..4b19e7af --- /dev/null +++ b/open4d/streamer/streamer/representations.py @@ -0,0 +1,156 @@ +"""What the streamer needs to know about a representation, and nothing else. + +The point of this module is that adding a representation should not require +editing the server, the manifest writer, or any exporter. Before it existed the +server carried a hardcoded suffix-to-MIME table and the client a hardcoded pair +of clip kinds, so a new representation meant touching unrelated files and +discovering the omissions at runtime -- a ``.obj`` arriving as ``text/html``, +say, which fails as a parse error rather than as a missing registration. + +A spec is deliberately small, and got smaller when `streamer.codecs` arrived. +It answers one question -- whether the bundled client can *render* this +representation today -- and derives the rest: + +* ``has_geometry``, whether a free camera is meaningful, is + `open4d.core.Representation`'s to answer and is read from there rather than + copied, because a second copy is a second opinion. +* ``media_types`` used to be listed here per representation, which duplicated + what the codec registry knows and let the two disagree. It is now derived + from `streamer.codecs`: a suffix is a property of a codec, not of a + representation, and the same ``.ply`` is three different codecs depending on + which representation is asking. + +``playable`` is a property of this repository's client, not of the +representation: a bundle declaring something the client cannot draw is reported +as unplayable rather than shown as an empty pane. Every representation core +defines is playable today, so nothing sets it False -- it stays because the next +representation will arrive before its renderer does, and that gap should be +stated rather than discovered. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from types import MappingProxyType +from typing import Mapping + +from open4d.core import Representation + +_OCTET = "application/octet-stream" + + +@dataclass(frozen=True) +class RepresentationSpec: + """What the client can do with one representation.""" + + representation: Representation + #: Whether the client packaged in `streamer.client` can render it today. + playable: bool = True + + def __post_init__(self) -> None: + if not isinstance(self.representation, Representation): + raise TypeError("representation must be an open4d.core.Representation") + if not isinstance(self.playable, bool): + raise TypeError("playable must be bool") + + @property + def name(self) -> str: + """The wire value, as it appears in ``view.json`` and in URLs.""" + return self.representation.value + + @property + def has_geometry(self) -> bool: + """Read from core, never stored: one definition, one answer.""" + return self.representation.has_geometry + + @property + def media_types(self) -> Mapping[str, str]: + """Suffix -> ``Content-Type``, from the codecs that produce this. + + Derived rather than declared: a suffix belongs to a codec. Listing them + here as well is how the registry and the server came to disagree about + what a ``.drc`` was. + """ + from . import codecs + + return MappingProxyType({ + spec.suffix: spec.media_type + for spec in codecs.for_representation(self.representation) + }) + + @property + def codecs(self) -> tuple: + """Every codec producing this representation; see `streamer.codecs`.""" + from . import codecs as registry + + return registry.for_representation(self.representation) + + +_REGISTRY: dict[Representation, RepresentationSpec] = {} + + +def register(spec: RepresentationSpec, *, replace: bool = False) -> RepresentationSpec: + """Add ``spec`` to the registry and return it. + + Refuses to shadow an existing registration unless asked, because two specs + for one representation is the failure this registry exists to prevent -- and + a silent overwrite would make which one wins depend on import order. + """ + if not isinstance(spec, RepresentationSpec): + raise TypeError("spec must be a RepresentationSpec") + if spec.representation in _REGISTRY and not replace: + raise ValueError( + f"{spec.name!r} is already registered; pass replace=True to override" + ) + _REGISTRY[spec.representation] = spec + return spec + + +def spec(representation: Representation | str) -> RepresentationSpec: + """The spec for ``representation``, by enum member or by wire value.""" + key = ( + representation + if isinstance(representation, Representation) + else Representation(representation) + ) + try: + return _REGISTRY[key] + except KeyError: + raise KeyError( + f"{key.value!r} has no RepresentationSpec; register one with " + "streamer.representations.register" + ) from None + + +def known() -> tuple[RepresentationSpec, ...]: + """Every registered spec, in the order core declares the representations.""" + return tuple( + _REGISTRY[member] for member in Representation if member in _REGISTRY + ) + + +def playable() -> tuple[RepresentationSpec, ...]: + """The specs the packaged client can actually render.""" + return tuple(item for item in known() if item.playable) + + +def media_types() -> dict[str, str]: + """Every suffix any codec produces, for a server's extension map. + + Kept here as well as in `streamer.codecs` because this is where the server + has always asked; it is now a delegation rather than a second list. + """ + from . import codecs + + return codecs.media_types() + + +# ------------------------------------------------------------- the defaults --- +# One line each, now that suffixes live with the codecs that produce them. What +# is left is the single claim this module makes: the packaged client can render +# all four. + +register(RepresentationSpec(representation=Representation.GAUSSIANS)) +register(RepresentationSpec(representation=Representation.PIXELS)) +register(RepresentationSpec(representation=Representation.MESH)) +register(RepresentationSpec(representation=Representation.POINTS)) diff --git a/open4d/streamer/streamer/sequence.py b/open4d/streamer/streamer/sequence.py new file mode 100644 index 00000000..8bc6c829 --- /dev/null +++ b/open4d/streamer/streamer/sequence.py @@ -0,0 +1,283 @@ +"""A whole clip as one file. + +A clip is a directory of frames, and fetching it a frame at a time is how this +started: thirty requests for thirty frames. On loopback that is free, which is +why it survived. Over a link it is thirty round trips before anything plays -- +600 ms of pure latency on a 20 ms connection -- and it is thirty chances for +one frame to arrive late and stall a clip that was otherwise complete. + +For **on demand** none of that buys anything. A viewer with a free camera has +to hold the whole sequence in memory anyway, because a frame it has thrown away +cannot be redrawn from a new angle. So the unit of transfer should be the +sequence, and there is no reason to have asked for it in pieces. + +Hence a container: a header naming the frames, then their bytes end to end. One +request, one response, one progress bar. Deliberately not a zip or a tar -- +both would need a decoder in the client and neither buys anything here, since +the frames are already compressed and an archive's own compression would only +spend CPU to save nothing. + + header magic "O4DSEQ\\0\\0" | version | frame count | suffix + then per frame: offset, length -- both uint32 little-endian + body each frame's bytes, in playback order, unpadded + +The offsets are absolute within the file, so a client that has the header can +slice any frame out without walking the ones before it -- which is what lets a +partially arrived download start playing from the beginning while the rest +lands. +""" +from __future__ import annotations + +import dataclasses +import struct +from dataclasses import dataclass +from pathlib import Path +from typing import Sequence + +from . import representations + +MAGIC = b"O4DSEQ\x00\x00" +VERSION = 1 +#: ``magic + version + count + suffix length``, before the suffix itself. +_PREAMBLE = struct.calcsize("<8sIII") + + +@dataclass(frozen=True) +class Entry: + """Where one frame sits inside the container.""" + + offset: int + length: int + + +@dataclass(frozen=True) +class Sequence: + """A container's header: what is in it and where.""" + + suffix: str + entries: tuple[Entry, ...] + + @property + def frames(self) -> int: + return len(self.entries) + + @property + def bytes(self) -> int: + return sum(entry.length for entry in self.entries) + + +def pack(frames: Sequence[Path | str], destination: Path | str) -> Sequence: + """Write ``frames`` into one container at ``destination``. + + Every frame must share a suffix. A container holding two formats would need + the client to switch decoders mid-sequence, and a clip whose frames are not + all one codec is a clip that should have been two. + """ + paths = [Path(f) for f in frames] + if not paths: + raise ValueError("a sequence needs at least one frame") + suffixes = {path.suffix.lower().lstrip(".") for path in paths} + if len(suffixes) != 1: + raise ValueError( + f"frames must share one suffix; got {', '.join(sorted(suffixes))}" + ) + suffix = suffixes.pop() + missing = [str(path) for path in paths if not path.is_file()] + if missing: + raise FileNotFoundError(f"{len(missing)} frame(s) missing, e.g. {missing[0]}") + + encoded = suffix.encode() + header = _PREAMBLE + len(encoded) + 8 * len(paths) + entries, offset = [], header + for path in paths: + length = path.stat().st_size + entries.append(Entry(offset=offset, length=length)) + offset += length + + destination = Path(destination) + destination.parent.mkdir(parents=True, exist_ok=True) + with open(destination, "wb") as out: + out.write(struct.pack("<8sIII", MAGIC, VERSION, len(paths), len(encoded))) + out.write(encoded) + for entry in entries: + out.write(struct.pack(" Sequence: + """The header of a container, from its first bytes. + + Takes bytes rather than a path so a client that has only the start of a + download can already know what is coming -- which is what a progress bar + counting frames needs. + """ + if len(data) < _PREAMBLE: + raise ValueError("not enough bytes for a sequence header") + magic, version, count, suffix_length = struct.unpack_from("<8sIII", data) + if magic != MAGIC: + raise ValueError(f"not a sequence container: magic is {magic!r}") + if version != VERSION: + raise ValueError(f"sequence version {version} is not {VERSION}") + at = _PREAMBLE + suffix = data[at:at + suffix_length].decode() + at += suffix_length + needed = at + 8 * count + if len(data) < needed: + raise ValueError(f"header needs {needed} bytes, got {len(data)}") + entries = tuple( + Entry(*struct.unpack_from(" bytes: + """One frame's bytes out of a container.""" + header = header or read_header(data) + entry = header.entries[index] + return data[entry.offset:entry.offset + entry.length] + + +def unpack(path: Path | str, destination: Path | str) -> list[Path]: + """Write a container back out as numbered frame files. + + The inverse, for checking a container round-trips and for anyone who wants + the frames on disk again. + """ + data = Path(path).read_bytes() + header = read_header(data) + destination = Path(destination) + destination.mkdir(parents=True, exist_ok=True) + written = [] + for index in range(header.frames): + target = destination / f"frame_{index:04d}.{header.suffix}" + target.write_bytes(frame(data, index, header)) + written.append(target) + return written + + +def pack_clip(bundle_dir, clip, *, keep_frames: bool = False) -> dict: + """Pack one clip's frames into a container beside the bundle root. + + Returns the mapping a `bundle.Clip` records as its ``sequence``. The frame + files are removed unless ``keep_frames``: leaving both doubles the bundle + on disk for no benefit, since a client given a container never asks for the + pieces. + """ + root = Path(bundle_dir) + relative = f"{clip.name}.seq" + header = pack([root / frame for frame in clip.frames], root / relative) + if not keep_frames: + directories = {Path(frame).parts[0] for frame in clip.frames} + for frame in clip.frames: + (root / frame).unlink(missing_ok=True) + for name in directories: + directory = root / name + if directory.is_dir() and not any(directory.iterdir()): + directory.rmdir() + packed = { + "url": relative, + "frames": header.frames, + "bytes": (root / relative).stat().st_size, + "suffix": header.suffix, + } + # The container is served as octet-stream, so a client decoding a frame out + # of it has no response header to read the frame's type from. Recorded here + # from the registry rather than mapped again in the client, which would be + # a second table to drift. + media_type = representations.media_types().get(f".{header.suffix}") + if media_type: + packed["media_type"] = media_type + return packed + + +def pack_bundle( + bundle_dir: Path | str, + *, + names: Sequence[str] | None = None, + keep_frames: bool = False, +) -> list[tuple[str, dict]]: + """Pack a bundle's clips into containers and rewrite its manifest. + + Live clips are skipped rather than refused: a bundle is usually a mix, and + a whole-bundle command that failed on the first stream would be unusable on + exactly the bundles that have both. Clips already packed are skipped too, + so running this twice is not an error. + """ + from . import bundle as _bundle + + root = Path(bundle_dir) + index = _bundle.read(root) + if not index: + raise FileNotFoundError(f"{root} has no manifest") + + clips = [_bundle.Clip(**entry) for entry in index.get("clips", [])] + wanted = set(names) if names else None + if wanted: + missing = wanted - {clip.name for clip in clips} + if missing: + raise KeyError(f"no clip named {', '.join(sorted(missing))}") + + packed, changed = [], [] + for clip in clips: + skip = ( + (wanted is not None and clip.name not in wanted) + or clip.stream is not None + or clip.sequence is not None + or not clip.frames + ) + if skip: + changed.append(clip) + continue + entry = pack_clip(root, clip, keep_frames=keep_frames) + changed.append(dataclasses.replace(clip, sequence=entry)) + packed.append((clip.name, entry)) + + if packed: + _bundle.write( + root, + title=index.get("title", root.name), + source=index.get("source", str(root)), + clips=changed, + fps=index.get("fps", 30), + scenes=index.get("scenes") or {}, + detail=index.get("detail"), + ) + return packed + + +def main(argv=None) -> int: + import argparse + + parser = argparse.ArgumentParser(description=__doc__.split("\n", 1)[0]) + parser.add_argument("bundle", type=Path, help="bundle directory to pack") + parser.add_argument( + "--clip", action="append", dest="clips", metavar="NAME", + help="pack only this clip; repeatable (default: every packable clip)", + ) + parser.add_argument( + "--keep-frames", action="store_true", + help="leave the frame files in place as well as the container", + ) + args = parser.parse_args(argv) + + packed = pack_bundle( + args.bundle, names=args.clips, keep_frames=args.keep_frames + ) + if not packed: + print("nothing to pack") + return 0 + for name, entry in packed: + print( + f"{name:36s} {entry['frames']:4d} frames" + f" {entry['bytes'] / 1e6:8.1f} MB {entry['url']}" + ) + total = sum(entry["bytes"] for _, entry in packed) + print(f"{len(packed)} clip(s), {total / 1e6:.1f} MB, one request each") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/open4d/streamer/streamer/server/__init__.py b/open4d/streamer/streamer/server/__init__.py new file mode 100644 index 00000000..19949e78 --- /dev/null +++ b/open4d/streamer/streamer/server/__init__.py @@ -0,0 +1,611 @@ +"""Serving a bundle to the playback client. + +Deliberately dumb: static files out of the bundle directory, with ``/`` rewritten +to the packaged client (`streamer.client`) so the bundle itself never has to +carry a copy of the page. It binds loopback unless told otherwise, since a +bundle is research output on a shared machine and this server has no +authentication of any kind. + +This is the transport a *local* bundle needs, and it is the simplest of the +transports the streaming model admits: every frame is a file, reachable in one +request, in any order. See `open4d.core.Dependency` for the ones that are not -- +a codec whose frames depend on a key frame, or whose decode stream cannot be +rewound, needs a scheduler over this rather than a different server. +""" + +from __future__ import annotations + +import errno +import http.server +import gzip +import json +import socket +import socketserver +import threading +import time +import urllib.error +import urllib.request +import webbrowser +from functools import partial +from pathlib import Path + +from .. import bundle, live, representations +from ..link import Link, described +from ..client import viewer_path +from ..monitor import Monitor + +DEFAULT_PORT = 8770 + + +VIEWER_ROUTES = ("/", "/index.html", "/viewer.html") +STATS_ROUTE = "/stats.json" +#: Live streams are proxied under this prefix; see `streamer.live`. +LIVE_PREFIX = f"/{live.ROUTE_PREFIX}/" +#: Assets that belong to the client package rather than to any bundle -- the +#: vendored Draco decoder, for one. Served from here so the page fetches them +#: from its own origin and never a CDN, and so a bundle does not have to carry +#: a copy of a decoder it did not choose. +CLIENT_PREFIX = "/client/" +#: Copy size for the proxy. Small, because a frame boundary can fall anywhere +#: and a large buffer would hold the tail of one frame back until the next. +PROXY_CHUNK = 8192 +#: Copy size for a range response. Larger than the proxy's: there is no frame +#: boundary to respect, and a 100 MB clip should not be a million writes. +RANGE_CHUNK = 1 << 16 + +#: Media types worth compressing. Deliberately a list of text-ish types rather +#: than "everything except a few": a frame is already compressed -- JPEG, or a +#: `.splat`'s quantised bytes -- so gzipping one spends CPU per request to save +#: almost nothing, and gzipping a 63 MB container would also make a Range +#: request meaningless, which is what the header prefetch relies on. +COMPRESSIBLE = frozenset({ + "application/json", + # Both spellings: `CLIENT_TYPES` serves the worker as text/javascript and + # `mimetypes` may answer application/javascript for the same suffix, so + # listing one silently left the other uncompressed -- which is how the + # 22 kB worker script went out whole while the page beside it did not. + "application/javascript", + "text/javascript", + "text/html", + "text/css", + "text/plain", +}) + +#: Below this, the gzip header and the round trip through zlib cost more than +#: they save. +COMPRESS_FLOOR = 1 << 10 + +#: Content types for client-package assets. `.wasm` matters: a browser refuses +#: to compile a module served as anything else through the streaming API. +CLIENT_TYPES = { + ".js": "text/javascript", + ".wasm": "application/wasm", + ".html": "text/html; charset=utf-8", + ".md": "text/plain; charset=utf-8", +} + + +def _parse_range(header: str, size: int): + """``(start, end)`` inclusive for a single byte range, or None if unusable. + + Handles the three forms that occur: ``bytes=0-99`` explicit, + ``bytes=500-`` open ended, and ``bytes=-500`` meaning the last 500. None + means answer 416, which is what a start past the end of the file deserves + -- returning the whole file there would let a resuming client silently + append a second copy to what it already had. + """ + if not header.startswith("bytes=") or "," in header: + return None + spec = header[len("bytes="):].strip() + first, _, last = spec.partition("-") + try: + if not first: # bytes=-N, the final N bytes + if not last: + return None + length = int(last) + if length <= 0: + return None + return max(0, size - length), size - 1 + start = int(first) + end = int(last) if last else size - 1 + except ValueError: + return None + if start < 0 or start >= size or end < start: + return None + return start, min(end, size - 1) + + +class _ShapedWriter: + """Wraps a socket's write file so every byte is paced by a `Link`. + + Wrapping the file rather than pacing each route is what makes the shaping + total: headers, bodies, range responses and the live proxy all leave + through here, and a route added later is shaped without knowing a link + exists. Pacing at the routes instead would have left whichever one was + written next unshaped, and an unshaped route in a measurement is not a + smaller effect -- it is the whole result, since that is where the bytes go. + + The propagation delay is charged on the first write of each response, not + per chunk: a stream of bytes already in flight pays the flight time once. + """ + + def __init__(self, wfile, link: Link) -> None: + self._wfile = wfile + self._link = link + self._fresh = True + + def restart(self) -> None: + """Called per request, so each response pays propagation once.""" + self._fresh = True + + def write(self, data): + if data: + self._link.send(len(data), propagate=self._fresh) + self._fresh = False + return self._wfile.write(data) + + def __getattr__(self, name): + return getattr(self._wfile, name) + + +class _Handler(http.server.SimpleHTTPRequestHandler): + """Static files from the bundle, plus the client page and the counters. + + ``monitor`` is set per-server by :func:`serve`; a handler class with no + monitor still works, which is what keeps this usable as a plain static + server. + + **HTTP/1.1, so a connection is reused across frames.** The base class + defaults to 1.0, which closes after every response -- one TCP handshake and + one fresh slow-start per frame. On loopback that is invisible, which is why + it survived; over a link with 20 ms of round trip a 59 kB Draco frame then + costs a setup RTT plus roughly three more while the congestion window + opens, so about 80 ms a frame and a ~12 fps ceiling *regardless of + bandwidth*. Any rate this server appears to sustain would be measuring that + rather than the network, which makes it the first thing to fix before + measuring anything. + + 1.1 requires every response to be self-delimiting, or a client waits for a + body that never ends. Two consequences, both handled below: every response + here carries a ``Content-Length``, and the one that cannot -- the live + proxy, whose body is endless -- says ``Connection: close`` and means it. + + It also means an idle client holds a thread until it goes away, so + ``timeout`` reaps connections that stop talking. + """ + + protocol_version = "HTTP/1.1" + + #: Turn off Nagle's algorithm. Not an optimisation -- without it keep-alive + #: is *slower* than the connection-per-request it replaces, and measurably: + #: 30 frames took 1.20 s against 0.01 s, 40 ms each, on loopback. + #: + #: The cause is Nagle meeting delayed ACK. A response leaves here as two + #: writes, headers then body, because the base class buffers the headers + #: and flushes them at ``end_headers``. Nagle holds the second small write + #: until the first is acknowledged; the client's delayed-ACK timer sits on + #: that acknowledgement for ~40 ms. Closing the connection used to mask it, + #: since the FIN pushes everything out at once -- so this only appeared + #: once keep-alive worked. + #: + #: 40 ms a frame is a 25 fps ceiling on loopback, which would have made the + #: transport worse than before while looking like progress. + disable_nagle_algorithm = True + + #: Seconds an idle keep-alive connection is held before it is dropped. + #: Without this a browser tab left open pins a thread indefinitely, and a + #: threaded server with unbounded idle connections eventually stops + #: accepting. + timeout = 30 + + monitor: Monitor | None = None + #: A shared bottleneck, or None for an unshaped loopback server. Shared on + #: purpose: concurrent panes have to contend for one pipe, or the budget a + #: chooser is given means nothing. + link: Link | None = None + #: Clip name -> upstream URL, from the manifest. Empty for a bundle with no + #: live clips, which makes the proxy route 404 rather than exist unused. + upstreams: dict = {} + + # SimpleHTTPRequestHandler guesses by extension and falls back to + # text/html, which makes a .ply arrive as markup and fail to parse. The + # suffixes come from the representation registry rather than a list here, so + # registering a representation is all it takes to serve its frames. + extensions_map = { + **http.server.SimpleHTTPRequestHandler.extensions_map, + **representations.media_types(), + ".json": "application/json", + # A packed clip is opaque: the frames inside keep their own type, which + # the manifest names, so guessing one for the container would only be + # wrong. See streamer/sequence.py. + ".seq": "application/octet-stream", + } + + def do_GET(self): # noqa: N802 - the base class names it + if self.path in VIEWER_ROUTES: + return self._send_viewer() + if self.path.split("?", 1)[0] == STATS_ROUTE: + return self._send_stats() + if self.path.startswith(LIVE_PREFIX): + return self._proxy_live(self.path[len(LIVE_PREFIX):]) + if self.path.startswith(CLIENT_PREFIX): + return self._send_client_asset(self.path[len(CLIENT_PREFIX):]) + if self.headers.get("Range"): + self._offer_ranges = True + return self._send_range() + compressible = self._compressible_file() + if compressible is not None: + path, content_type = compressible + # Not offering Accept-Ranges here: a range names bytes of the + # entity as sent, and the entity as sent is compressed. Offering + # ranges over one form and serving another is how a resumed + # download quietly reassembles garbage. + return self._send_payload(path.read_bytes(), content_type) + self._offer_ranges = True + return super().do_GET() + + def do_HEAD(self): # noqa: N802 + if self.path in VIEWER_ROUTES: + return self._send_viewer(body=False) + self._offer_ranges = True + return super().do_HEAD() + + def _compressible_file(self): + """The static file this request names, if it is worth compressing. + + Returns ``(path, content_type)`` or None. Read whole into memory by the + caller, which is why only the text types qualify: those are manifests + and scripts, megabytes at the outside, where a frame container is tens + of megabytes and must keep streaming off disk. + """ + if not self._accepts_gzip(): + return None + path = Path(self.translate_path(self.path)) + if not path.is_file(): + return None + content_type = self.guess_type(str(path)) + if content_type.split(";", 1)[0].strip() not in COMPRESSIBLE: + return None + if path.stat().st_size < COMPRESS_FLOOR: + return None + return path, content_type + + def _accepts_gzip(self) -> bool: + """Whether this client said it would take gzip. + + Only ever offered, never assumed: `streamer.transfer` and the tests + speak plain HTTP through `urllib`, which does not advertise gzip and + would be handed bytes it will not decode. + """ + offered = self.headers.get("Accept-Encoding", "") + return any( + token.split(";", 1)[0].strip() == "gzip" + for token in offered.split(",") + ) + + def _send_payload(self, payload: bytes, content_type: str, *, + cache: str | None = None, body: bool = True, + status: int = 200): + """One response, compressed when that is worth doing. + + The bundle manifest is what motivated this. It names every frame of + every clip, and the orbit exports put 216 clips in a scene, so nine + subjects come to 6.6 MB -- fetched before the page can draw anything, + which on a 20 Mbit/s link is two and a half seconds of blank. It is + also enormously repetitive (the same four notes on 216 clips, and + frame paths differing by four digits), and that is exactly what gzip + eats: measured at 20.3x, so 6.6 MB becomes 0.33 MB. + + Compressing rather than restructuring the manifest is the cheaper + answer and the more general one -- it helps the viewer and the worker + script too, and it does not change a format the client has to agree + about. + """ + base = content_type.split(";", 1)[0].strip() + compressed = ( + base in COMPRESSIBLE + and len(payload) >= COMPRESS_FLOOR + and self._accepts_gzip() + ) + if compressed: + payload = gzip.compress(payload, 6) + self.send_response(status) + self.send_header("Content-Type", content_type) + if compressed: + self.send_header("Content-Encoding", "gzip") + # Named whether or not this response is compressed: a cache holding the + # identity form must not serve it to a client that asked for gzip, and + # the header is what says the two differ. + self.send_header("Vary", "Accept-Encoding") + self.send_header("Content-Length", str(len(payload))) + if cache: + self.send_header("Cache-Control", cache) + self.end_headers() + if body: + self.wfile.write(payload) + + def end_headers(self): + # Advertised here because the base class ends its own headers, leaving + # no later point to add one. Only on the static-file path, since that + # is the only route that honours a Range. + if getattr(self, "_offer_ranges", False): + self.send_header("Accept-Ranges", "bytes") + self._offer_ranges = False + super().end_headers() + + def _send_range(self): + """Serve a byte range of a bundle file. + + What this buys is resumption: `streamer.transfer` fetching a 100 MB + Gaussian clip over a tunnel that drops can continue from where it + stopped instead of starting the file again. Without it the only + recovery is re-downloading, which on a big clip is the difference + between seconds and minutes. + + Only a single range is honoured. Multipart ranges exist in the + standard, are used by essentially nothing, and would need a different + body format; asking for several gets the whole file, which is a + response the standard permits and every client handles. + """ + path = Path(self.translate_path(self.path)) + if not path.is_file(): + return super().do_GET() # let the base class 404 or index + size = path.stat().st_size + span = _parse_range(self.headers.get("Range", ""), size) + if span is None: + self.send_response(416, "Requested Range Not Satisfiable") + self.send_header("Content-Range", f"bytes */{size}") + self.send_header("Content-Length", "0") + self.end_headers() + return + start, end = span + length = end - start + 1 + self.send_response(206, "Partial Content") + self.send_header("Content-Type", self.guess_type(str(path))) + self.send_header("Content-Range", f"bytes {start}-{end}/{size}") + self.send_header("Content-Length", str(length)) + self.end_headers() + with open(path, "rb") as handle: + handle.seek(start) + remaining = length + while remaining > 0: + block = handle.read(min(RANGE_CHUNK, remaining)) + if not block: + break + self.wfile.write(block) + remaining -= len(block) + + def _send_stats(self): + """What has gone over the wire, as JSON. Absent without a monitor.""" + if self.monitor is None: + self.send_error(404, "this server keeps no counters") + return + snapshot = self.monitor.snapshot() + if self.link is not None: + snapshot["link"] = self.link.observed() + payload = json.dumps(snapshot, indent=2).encode() + self._send_payload(payload, "application/json", cache="no-store") + + def _send_client_asset(self, relative: str): + """A file from the client package, e.g. the vendored Draco decoder.""" + root = viewer_path().parent + # Resolved and then checked to be inside the package: the path comes off + # a URL, so `../../etc/passwd` is a request this will receive eventually. + target = (root / relative.split("?", 1)[0]).resolve() + if not target.is_file() or root.resolve() not in target.parents: + self.send_error(404, f"no client asset {relative!r}") + return + payload = target.read_bytes() + self._send_payload( + payload, + CLIENT_TYPES.get(target.suffix.lower(), "application/octet-stream"), + # Immutable: these ship with the package, so a version of the page + # and a version of its decoder always arrive together. + cache="public, max-age=86400", + ) + + def _proxy_live(self, name: str): + """Relay a live stream from its renderer, so the page has one origin. + + The alternative is putting the renderer's own URL in the manifest, which + asks the *browser* to reach that port -- and a browser on a laptop + looking at a tunnelled page cannot, so the pane stays blank and nothing + reports why. Proxying costs a thread and a copy loop and removes the + whole class of problem. + + Copied through rather than buffered: this is a `multipart/x-mixed-replace` + body that never ends, so anything that waits for completion waits + forever. + """ + upstream = self.upstreams.get(name.split("?", 1)[0]) + if upstream is None: + self.send_error(404, f"no live stream named {name!r} in this bundle") + return + try: + source = urllib.request.urlopen(upstream, timeout=10) + except (urllib.error.URLError, OSError) as error: + # 502 with the reason, because "the renderer is not running" is the + # single most likely thing to be wrong and is not the bundle's fault. + self.send_error(502, f"cannot reach {upstream}: {error}") + return + + self.send_response(200) + for header in ("Content-Type", "Age", "Cache-Control", "Pragma"): + value = source.headers.get(header) + if value: + self.send_header(header, value) + self.send_header("Cache-Control", "no-store, no-cache, private") + # The one response here with no Content-Length: a + # multipart/x-mixed-replace body never ends. Under HTTP/1.1 that has to + # be delimited by the close, so say so and hold the connection for this + # stream alone rather than trying to reuse it afterwards. + self.send_header("Connection", "close") + self.close_connection = True + self.end_headers() + try: + while True: + block = source.read(PROXY_CHUNK) + if not block: + break + self.wfile.write(block) + except (BrokenPipeError, ConnectionResetError): + pass # the tab was closed or navigated away; expected + finally: + source.close() + + def send_response(self, code, message=None): + # Captured here because this is the one place every response passes + # through with its status known, including the base class's static-file + # path and its errors -- wrapping only do_GET would miss both. + self._status = code + super().send_response(code, message) + + def send_header(self, keyword, value): + if keyword.lower() == "content-length": + self._size = int(value) + super().send_header(keyword, value) + + def setup(self): + super().setup() + if self.link is not None: + self.wfile = _ShapedWriter(self.wfile, self.link) + + def handle_one_request(self): + if isinstance(self.wfile, _ShapedWriter): + self.wfile.restart() + self._status, self._size, started = 200, 0, time.monotonic() + super().handle_one_request() + if self.monitor is not None and self.path: + self.monitor.record( + self.path, self._status, self._size, time.monotonic() - started + ) + + def _send_viewer(self, *, body: bool = True): + try: + payload = viewer_path().read_bytes() + except OSError as error: + self.send_error(500, f"viewer is missing: {error}") + return + self._send_payload(payload, "text/html; charset=utf-8", + cache="no-store", body=body) + + def log_message(self, fmt, *args): + # One line per request would bury the URL the user needs; errors still + # surface through log_error, which the base class routes here with a + # different format string. + if not str(args[0] if args else "").startswith(("GET /frame", "GET /", "HEAD")): + super().log_message(fmt, *args) + + +class _Server(socketserver.ThreadingTCPServer): + daemon_threads = True + allow_reuse_address = True + + +def _reachable_address(host: str, port: int) -> str: + """The URL to print. For a wildcard bind, the LAN address is the useful one.""" + if host not in ("0.0.0.0", "::", ""): + return f"http://{host}:{port}/" + probe = socket.socket(socket.AF_INET, socket.SOCK_DGRAM) + try: + probe.connect(("8.8.8.8", 80)) + return f"http://{probe.getsockname()[0]}:{port}/" + except OSError: + return f"http://{socket.gethostname()}:{port}/" + finally: + probe.close() + + +def serve( + bundle_dir: Path | str, + *, + host: str = "127.0.0.1", + port: int = DEFAULT_PORT, + open_browser: bool = False, + block: bool = True, + monitor: Monitor | None = None, + link: Link | None = None, +) -> _Server: + """Serve ``bundle_dir`` with the client at ``/`` and counters at ``/stats.json``. + + Returns the server, with the `Monitor` it is recording into attached as + ``server.monitor`` -- which is how a caller measures a playback without + scraping the JSON back out of its own process. Pass ``monitor`` to share one + across several servers, or to keep it after the server is gone. + + A response is recorded once it has completed, because its size and duration + are not known before that. So a client that has already read a body may + observe the counters a moment before they include it; over a playback the + difference is one request, and the alternative -- counting responses as they + start -- would report bytes that had not been sent. + + With ``block=True`` it runs until interrupted; with ``block=False`` it serves + on a daemon thread, which is what a test wants. + """ + bundle_dir = Path(bundle_dir).expanduser().resolve() + index = bundle.read(bundle_dir) + if not index: + raise FileNotFoundError( + f"{bundle_dir} has no {bundle.INDEX_NAME}; export one first" + ) + + # One monitor per server, attached to a subclass rather than to `_Handler` + # itself: the class attribute is shared, so setting it on the base would + # make two servers in one process count into each other. + counters = Monitor() if monitor is None else monitor + handler_class = type( + "_BundleHandler", + (_Handler,), + {"monitor": counters, "link": link, "upstreams": live.upstreams(index)}, + ) + handler = partial(handler_class, directory=str(bundle_dir)) + try: + server = _Server((host, port), handler) + except OSError as error: + if error.errno != errno.EADDRINUSE: + raise + # This box runs several long-lived demo servers in the 87xx range, so a + # fixed default port collides often enough that failing outright would + # just be an obstacle. Falling back is visible, since the port is printed. + server = _Server((host, 0), handler) + print(f"port {port} is already in use; using {server.server_address[1]} instead") + server.monitor = counters + server.link = link + url = _reachable_address(host, server.server_address[1]) + + clips = index.get("clips", []) + print(f"{index.get('title', bundle_dir.name)}") + for clip in clips: + print( + f" {clip['name']:<28} {len(clip.get('frames', []))} frames" + f" ({clip.get('representation')})" + ) + print(f"\nserving {bundle_dir}\n {url}") + print(f" link: {described(link)}") + print(f" counters at {url.rstrip('/')}{STATS_ROUTE}") + for name, upstream in live.upstreams(index).items(): + print(f" live {name} <- {upstream}") + if host in ("0.0.0.0", "::", ""): + print(" bound to every interface, with no authentication: anyone who can") + print(" reach this port can read the bundle.") + if open_browser: + webbrowser.open(url) + + if not block: + # `shutdown()` waits for serve_forever to notice, so the default 0.5 s + # poll makes every programmatic stop cost half a second -- which a test + # suite or a measured run pays once per server for nothing. + threading.Thread( + target=server.serve_forever, kwargs={"poll_interval": 0.02}, daemon=True + ).start() + return server + + print(" Ctrl-C to stop") + try: + server.serve_forever() + except KeyboardInterrupt: + print() + finally: + server.shutdown() + server.server_close() + return server diff --git a/open4d/streamer/streamer/session.py b/open4d/streamer/streamer/session.py new file mode 100644 index 00000000..0f739769 --- /dev/null +++ b/open4d/streamer/streamer/session.py @@ -0,0 +1,258 @@ +"""Collect clips into a bundle, and write its index once. + +`export.from_sequence` writes one clip's frames and hands back a `Clip`; the +``view.json`` that makes them playable is a separate `bundle.write` with the +whole list. That split is right for a producer assembling clips across +processes -- which is what `bundle.add` exists for -- and wrong for the common +case of one caller and a few sequences, where forgetting the second call leaves +frames on disk that nothing will play. Nothing errors; the directory simply is +not a bundle. + +This is the common case: a builder that holds the list and writes the index +when the block ends. + + with Bundle("out/", title="Capture") as clips: + clips.add(sequence, name="capture", rungs=["draco", "draco@11"]) + server.serve("out/") + +Rungs are the other half. A single encode cannot be adapted between -- a client +with one rendition has nothing to switch to -- so `add` takes a list, writes the +first as the clip's default and the rest as `bundle.Variant` entries with their +sizes measured off disk. `quality` is left empty on purpose: this knows what a +rung cost, not what it was worth, and `metrics` is what fills that in. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Iterable, Sequence as TypingSequence + +from open4d.core import Sequence + +from . import bundle, export + +#: Separates a frame format from its quantisation in a rung spec: ``draco@11``. +RUNG_SEPARATOR = "@" + + +@dataclass(frozen=True) +class Rung: + """One quality level to encode, parsed from a spec string.""" + + #: What the variant is called in the manifest. The spec verbatim, so the + #: same rung of two different clips carries the same name and a consumer + #: can ask for "draco@11" of everything -- which is what `Variant.name` + #: asks for and what `policy` needs to compare panes. + id: str + frame_format: str + quantization_bits: int = export.DRACO_QUANTIZATION_BITS + + +def parse_rung(spec: str | Rung) -> Rung: + """``"ply"``, ``"draco"`` or ``"draco@11"`` as a `Rung`. + + Quantisation on a format that does not quantise is an error rather than an + ignored argument: ``ply@11`` is a caller believing they asked for something + smaller, and silently writing the same bytes at the same size would hide + that until someone compared the rungs and found them identical. + """ + if isinstance(spec, Rung): + return spec + text = str(spec).strip() + frame_format, _, bits = text.partition(RUNG_SEPARATOR) + frame_format = frame_format.strip() + if frame_format not in export.FORMATS: + raise ValueError( + f"unknown frame format {frame_format!r} in rung {spec!r}; " + f"expected one of {', '.join(export.FORMATS)}" + ) + if not bits: + return Rung(id=text, frame_format=frame_format) + if frame_format != export.DRACO_FORMAT: + raise ValueError( + f"rung {spec!r} sets quantisation on {frame_format!r}, which does " + f"not quantise; only {export.DRACO_FORMAT!r} does" + ) + try: + quantization_bits = int(bits) + except ValueError: + raise ValueError( + f"rung {spec!r} has a non-numeric quantisation {bits!r}" + ) from None + return Rung( + id=text, frame_format=frame_format, quantization_bits=quantization_bits + ) + + +def _measure(out_dir: Path, frames: Iterable[str]) -> int: + """Total bytes of ``frames``, which are relative to the bundle root. + + Measured rather than predicted, because the files already exist -- see the + note in `bundle.Variant`. + """ + return sum((out_dir / frame).stat().st_size for frame in frames) + + +class Bundle: + """Clips accumulating into one bundle directory. + + Usable as a context manager, in which case the index is written on a clean + exit and *not* written if the block raises. A half-built bundle with no + index is a directory of frames, which is recoverable; a half-built bundle + with an index is a manifest promising clips that are not all there, which + a client reports as missing frames rather than as a failed export. + """ + + def __init__( + self, + out_dir: Path | str, + *, + title: str | None = None, + source: Path | str | None = None, + fps: int = 30, + scenes: dict[str, Any] | None = None, + detail: dict[str, Any] | None = None, + ) -> None: + self.out_dir = Path(out_dir).expanduser().resolve() + self.title = title or self.out_dir.name + self.source = source + self.fps = fps + self.scenes = scenes + self.detail = detail + self.clips: list[bundle.Clip] = [] + + def __enter__(self) -> "Bundle": + return self + + def __exit__(self, exc_type, exc, traceback) -> None: + if exc_type is None: + self.write() + + @property + def index_path(self) -> Path: + return self.out_dir / bundle.INDEX_NAME + + def add_clip(self, clip: bundle.Clip) -> bundle.Clip: + """A clip some other exporter produced, taken into this bundle.""" + self.clips.append(clip) + return clip + + def add( + self, + sequence: Sequence, + *, + name: str, + rungs: TypingSequence[str | Rung] = (export.FRAME_FORMAT,), + scene: str | None = None, + method: str | None = None, + notes: list[str] | None = None, + detail: dict[str, Any] | None = None, + ) -> bundle.Clip: + """Write ``sequence`` at every rung, as one clip with variants. + + The first rung is the clip's default rendition -- the one a reader that + knows nothing about variants plays -- and the rest become `Variant` + entries beside it. Order is the caller's: this does not sort by size, + because which rendition should be the default is a delivery decision + (interchange? cheapest? middle?) and not one a byte count settles. + """ + parsed = [parse_rung(rung) for rung in rungs] + if not parsed: + raise ValueError(f"{name}: needs at least one rung") + seen: set[str] = set() + for rung in parsed: + if rung.id in seen: + raise ValueError(f"{name}: rung {rung.id!r} is listed twice") + seen.add(rung.id) + + default, *alternates = parsed + clip = export.from_sequence( + sequence, + self.out_dir, + name=name, + frame_format=default.frame_format, + quantization_bits=default.quantization_bits, + scene=scene, + method=method, + notes=notes, + detail={**(detail or {}), "rung": default.id}, + ) + for rung in alternates: + # A separate clip export per rung, whose frame list is then folded + # in as a variant and whose Clip is discarded. `from_sequence` is + # the only thing that knows how to write each format, and the + # alternative -- teaching it to write several at once -- would put + # the rung loop inside the exporter, where a caller exporting one + # rendition would pay for it. + rendition = export.from_sequence( + sequence, + self.out_dir, + name=f"{name}-{rung.id.replace(RUNG_SEPARATOR, '')}", + frame_format=rung.frame_format, + quantization_bits=rung.quantization_bits, + scene=scene or name, + method=method, + ) + clip.variants.append( + bundle.Variant( + name=rung.id, + frames=rendition.frames, + bytes=_measure(self.out_dir, rendition.frames), + detail={ + "frame_format": rung.frame_format, + **( + {"quantization_bits": rung.quantization_bits} + if rung.frame_format == export.DRACO_FORMAT + else {} + ), + }, + ).as_dict() + ) + return self.add_clip(clip) + + def add_source( + self, + source: Path | str, + *, + name: str | None = None, + fps: float | None = None, + **kwargs: Any, + ) -> bundle.Clip: + """Whatever `open4d.load` reads at ``source``, added as one clip. + + The `fps` rule is `export.from_source`'s, and deliberately the same: it + applies only to a source carrying no timing of its own, and is ignored + rather than rejected for one that does. Repeated here rather than + shared because that function is the single-call path -- load, export + and write in one -- and reaching into it for the middle third would + make the simple case depend on the general one. + """ + import open4d + from open4d.io import inspect_sequence + + source = Path(source).expanduser().resolve() + declared = inspect_sequence(source).timing_source + with open4d.load( + source, fps=fps if declared == "default" else None + ) as sequence: + return self.add( + sequence, name=name or source.stem or source.name, **kwargs + ) + + def write(self) -> Path: + """Write ``view.json`` for everything added so far.""" + if not self.clips: + raise ValueError( + f"{self.out_dir} has no clips; a bundle with an empty clip list " + "is a page that loads and shows nothing" + ) + return bundle.write( + self.out_dir, + title=self.title, + source=str(self.source) if self.source is not None else self.out_dir.name, + clips=self.clips, + fps=self.fps, + scenes=self.scenes, + detail=self.detail, + ) diff --git a/open4d/streamer/streamer/transfer.py b/open4d/streamer/streamer/transfer.py new file mode 100644 index 00000000..b43bacbd --- /dev/null +++ b/open4d/streamer/streamer/transfer.py @@ -0,0 +1,192 @@ +"""Receiving a bundle: the other half of the server's job. + +The server sends files; this pulls them. It exists because the common case is +awkward without it: a bundle is produced on the machine with the GPU, and the +person who wants to look at it is on a laptop somewhere else. Streaming it over +a tunnel works for a look, but the geometry clips are heavy -- a 30-frame +Gaussian clip is over 100 MB as PLY -- so anyone returning to the same bundle +twice wants a local copy, and a local copy is just the manifest plus every path +it names. + +It reads only `view.json`, so it needs to know nothing about representations: +the manifest already lists every frame, relative to the bundle root. That is the +same property that lets the client play a bundle it did not write. + +Not a sync tool. Existing files of the right size are skipped, and a partial +file is continued from where it stopped rather than started again -- that is +the whole policy. + +Resumption is the reason this cares about the transport at all. A 30-frame +Gaussian clip is over 100 MB; a tunnel that drops halfway through one frame +used to mean fetching that frame again from zero. With a byte range it costs +only what was actually missed. +""" + +from __future__ import annotations + +import json +import urllib.error +import urllib.parse +import urllib.request +from dataclasses import dataclass +from pathlib import Path +from typing import Callable, Iterable + +from . import bundle +from .monitor import Monitor + +#: Read size for streaming a body to disk. Large enough that a 100 MB frame is +#: not a million syscalls, small enough not to matter in memory. +CHUNK = 1 << 20 + + +@dataclass(frozen=True) +class FetchResult: + """What a fetch did.""" + + root: Path + fetched: tuple[str, ...] + skipped: tuple[str, ...] + bytes: int + + @property + def paths(self) -> tuple[str, ...]: + return self.fetched + self.skipped + + +def _get(url: str, timeout: float) -> bytes: + with urllib.request.urlopen(url, timeout=timeout) as response: + return response.read() + + +def _download(url: str, destination: Path, timeout: float) -> int: + """Stream ``url`` to ``destination``, returning bytes written this call. + + Written to a temporary neighbour and moved into place, so an interrupted + transfer leaves no short file that the size check would later mistake for a + complete one. + + A ``.partial`` left by an earlier attempt is *continued*, by asking for the + bytes after it with a ``Range`` header. Three things have to be true for + that to be safe, and all three are checked rather than assumed: + + * the server has to answer ``206`` -- a ``200`` means it ignored the range + and is sending the whole file, so the partial is discarded and this + becomes a plain download rather than appending a second copy; + * the range has to start where the partial ends, which is what was asked + for and is verified against ``Content-Range``; + * on any failure the partial survives, so the next attempt can try again. + + The file is only moved into place once the transfer completes, so a + ``.partial`` is always exactly the prefix that has arrived. + """ + destination.parent.mkdir(parents=True, exist_ok=True) + partial = destination.with_name(destination.name + ".partial") + have = partial.stat().st_size if partial.is_file() else 0 + + request = urllib.request.Request(url) + if have: + request.add_header("Range", f"bytes={have}-") + + written = 0 + with urllib.request.urlopen(request, timeout=timeout) as response: + resuming = have > 0 and response.status == 206 + if have and not resuming: + # The server sent the whole file despite the range. Start over + # rather than append: the alternative is a corrupt file that is + # exactly the right size for the check to accept. + have = 0 + with open(partial, "ab" if resuming else "wb") as handle: + while True: + block = response.read(CHUNK) + if not block: + break + handle.write(block) + written += len(block) + partial.replace(destination) + return written + + +def _remote_size(url: str, timeout: float) -> int | None: + """Content-Length from a HEAD, or None when the server will not say.""" + request = urllib.request.Request(url, method="HEAD") + try: + with urllib.request.urlopen(request, timeout=timeout) as response: + length = response.headers.get("Content-Length") + return int(length) if length is not None else None + except (urllib.error.URLError, ValueError, OSError): + return None + + +def frame_paths(index: dict) -> tuple[str, ...]: + """Every frame path a manifest names, in order, without duplicates.""" + seen: dict[str, None] = {} + for clip in index.get("clips", []): + for path in clip.get("frames", []): + seen.setdefault(path, None) + return tuple(seen) + + +def fetch( + url: str, + destination: Path | str, + *, + timeout: float = 30.0, + monitor: Monitor | None = None, + progress: Callable[[str, int, int], None] | None = None, + only: Iterable[str] | None = None, +) -> FetchResult: + """Copy the bundle served at ``url`` into ``destination``. + + ``only`` restricts the transfer to those clip names, which is how you take + one method off a 225-clip bundle instead of all of it. ``monitor`` records + the same counters the server keeps, so a transfer can be measured from + either end. ``progress`` is called as ``(path, done, total)``. + """ + base = url if url.endswith("/") else url + "/" + root = Path(destination).expanduser().resolve() + root.mkdir(parents=True, exist_ok=True) + + index = json.loads(_get(base + bundle.INDEX_NAME, timeout)) + if only is not None: + wanted = set(only) + index["clips"] = [ + clip for clip in index.get("clips", []) if clip.get("name") in wanted + ] + missing = wanted - {clip.get("name") for clip in index["clips"]} + if missing: + raise KeyError(f"no clip named {', '.join(sorted(missing))} in {base}") + + # Written first and last: first so the directory is recognisable as a bundle + # while frames arrive, last so a `--only` subset's manifest is the filtered + # one rather than the whole thing. + (root / bundle.INDEX_NAME).write_text(json.dumps(index, indent=2) + "\n") + + paths = frame_paths(index) + fetched: list[str] = [] + skipped: list[str] = [] + total_bytes = 0 + for position, path in enumerate(paths, start=1): + target = root / path + source = base + urllib.parse.quote(path) + expected = _remote_size(source, timeout) if target.is_file() else None + if target.is_file() and expected is not None and target.stat().st_size == expected: + skipped.append(path) + else: + # `written` counts what crossed the wire, not the file's size: a + # resumed frame reports only the tail, which is what a measurement + # of this transfer should say. + written = _download(source, target, timeout) + total_bytes += written + fetched.append(path) + if monitor is not None: + monitor.record(path, 200, written, 0.0) + if progress is not None: + progress(path, position, len(paths)) + + return FetchResult( + root=root, + fetched=tuple(fetched), + skipped=tuple(skipped), + bytes=total_bytes, + ) diff --git a/open4d/streamer/streamer_tests/__init__.py b/open4d/streamer/streamer_tests/__init__.py new file mode 100644 index 00000000..49a131ef --- /dev/null +++ b/open4d/streamer/streamer_tests/__init__.py @@ -0,0 +1,20 @@ +"""Shared fixtures locating repository data, wherever this package sits. + +`open4d_tree` exists because three tests counted ``parents[3]`` from their own +file to reach the `open4d` package directory, and moving `streamer` up one +level made all three resolve one directory too high. They did not fail: the +dataset path simply did not exist, so they skipped, reporting "not present" +about data that was there all along. A hop count is a claim about where this +package lives, and this package has now moved twice. +""" + +from __future__ import annotations + +from pathlib import Path + + +def open4d_tree() -> Path: + """The `open4d` package directory, from the installed package itself.""" + import open4d + + return Path(open4d.__file__).resolve().parent diff --git a/open4d/streamer/streamer_tests/test_adaptation.py b/open4d/streamer/streamer_tests/test_adaptation.py new file mode 100644 index 00000000..16ce41c5 --- /dev/null +++ b/open4d/streamer/streamer_tests/test_adaptation.py @@ -0,0 +1,354 @@ +"""The client's half of the loop: measure the link, pick a rung, fetch it. + +The pieces are cut out of `viewer.html` as shipped and run under Node, the same +way `test_scheduler.py` exercises the scheduler -- so what is tested is the +code the browser gets, not a copy of it. + +The rule the client applies is deliberately *not* `streamer.policy`'s. That one +maximises total weighted quality, which for a comparison view is the wrong +objective: the way to maximise a sum is to make the panes unequal. These tests +pin the even-share rule so that difference stays a decision rather than a +divergence. +""" +from __future__ import annotations + +import json +import shutil +import subprocess +import textwrap +from pathlib import Path + +import pytest + +from streamer import bundle +from streamer.client import viewer_path + +pytestmark = pytest.mark.cpu + +NODE = shutil.which("node") +requires_node = pytest.mark.skipif(NODE is None, reason="node is not installed") + + +def _extract(*names: str) -> str: + page = viewer_path().read_text() + chunks = [] + for name in names: + for prefix in (f"function {name}(", f"class {name} ", f"const {name} ="): + start = page.find(prefix) + if start >= 0: + break + else: + raise AssertionError(f"{name} is not defined in the viewer") + end = (page.index(";\n", start) + 2 if prefix.startswith("const") + else page.index("\n}\n", start) + 3) + chunks.append(page[start:end]) + return "\n".join(chunks) + + +def run_js(body: str, tmp_path: Path, name: str = "a.mjs") -> object: + script = tmp_path / name + script.write_text( + _extract("RateMeter", "rungsOf", "affordable") + "\n" + + textwrap.dedent(body) + ) + finished = subprocess.run( + [NODE, str(script)], capture_output=True, text=True, timeout=120 + ) + if finished.returncode: + raise AssertionError(finished.stderr) + return json.loads(finished.stdout) + + +# ------------------------------------------------------------- the estimate --- + + +@requires_node +def test_the_first_sample_is_taken_as_the_rate(tmp_path): + """With no history there is nothing to smooth towards, and starting from + zero would spend the first several frames climbing out of it.""" + result = run_js(""" + const meter = new RateMeter(); + meter.record(1_000_000, 1.0); + process.stdout.write(JSON.stringify(meter.bitsPerSecond)); + """, tmp_path) + assert result == pytest.approx(8_000_000) + + +@requires_node +def test_the_estimate_moves_towards_a_changed_rate(tmp_path): + result = run_js(""" + const meter = new RateMeter({ halfLife: 2 }); + meter.record(1_000_000, 1.0); // 8 Mbit/s + const before = meter.bitsPerSecond; + for (let i = 0; i < 8; i++) meter.record(1_000_000, 8.0); // 1 Mbit/s + process.stdout.write(JSON.stringify({ before, after: meter.bitsPerSecond })); + """, tmp_path) + assert result["before"] == pytest.approx(8_000_000) + assert 1_000_000 <= result["after"] < 2_500_000 + + +@requires_node +def test_a_tiny_response_is_ignored(tmp_path): + """A cache hit completes in microseconds and would read as a gigabit link, + which is exactly the over-estimate that then overshoots the real one.""" + result = run_js(""" + const meter = new RateMeter(); + meter.record(1_000_000, 1.0); + meter.record(300, 0.000001); // a 304, effectively + process.stdout.write(JSON.stringify(meter.bitsPerSecond)); + """, tmp_path) + assert result == pytest.approx(8_000_000) + + +@requires_node +def test_a_zero_duration_sample_is_ignored(tmp_path): + """Rather than dividing by zero and poisoning the estimate with Infinity.""" + result = run_js(""" + const meter = new RateMeter(); + meter.record(1_000_000, 0); + process.stdout.write(JSON.stringify( + { rate: meter.bitsPerSecond, samples: meter.samples })); + """, tmp_path) + assert result == {"rate": 0, "samples": 0} + + +@requires_node +def test_reset_forgets_everything(tmp_path): + result = run_js(""" + const meter = new RateMeter(); + meter.record(1_000_000, 1.0); + meter.reset(); + process.stdout.write(JSON.stringify( + { rate: meter.bitsPerSecond, samples: meter.samples })); + """, tmp_path) + assert result == {"rate": 0, "samples": 0} + + +# ---------------------------------------------------------------- the ladder --- + + +def a_clip(*, default_bytes=33_500, variants=(("low", 3_035), ("medium", 9_902))): + clip = { + "name": "c", "representation": "pixels", + "frames": [f"c/frame_{i:04d}.jpg" for i in range(30)], + "detail": {"bytes_per_frame": default_bytes, + "quality": {"psnr": 47.2}} if default_bytes else {}, + "variants": [ + {"name": name, "bytes": per * 30, + "frames": [f"c@{name}/frame_{i:04d}.jpg" for i in range(30)], + "quality": {"psnr": 40.0}} + for name, per in variants + ], + } + return clip + + +@requires_node +def test_the_ladder_is_read_cheapest_first(tmp_path): + result = run_js(f""" + const rungs = rungsOf({json.dumps(a_clip())}, 30); + process.stdout.write(JSON.stringify(rungs.map( + (r) => [r.name, Math.round(r.bitsPerSecond / 1000)]))); + """, tmp_path) + assert [name for name, _ in result] == ["low", "medium", "default"] + # 3035 B/frame x 8 x 30 fps = 728 kbit/s + assert result[0][1] == pytest.approx(728, abs=2) + assert result[-1][1] == pytest.approx(8040, abs=5) + + +@requires_node +def test_the_default_needs_its_recorded_size_to_be_a_rung(tmp_path): + """`streamer.metrics --write` records it. Without it the client can compare + the rungs it might move to and not the one it is already playing.""" + result = run_js(f""" + const rungs = rungsOf({json.dumps(a_clip(default_bytes=0))}, 30); + process.stdout.write(JSON.stringify(rungs.map((r) => r.name))); + """, tmp_path) + assert result == ["low", "medium"] + + +@requires_node +def test_a_clip_with_no_frames_has_no_ladder(tmp_path): + result = run_js(""" + process.stdout.write(JSON.stringify([ + rungsOf(null, 30).length, + rungsOf({ frames: [] }, 30).length, + ])); + """, tmp_path) + assert result == [0, 0] + + +@requires_node +def test_an_empty_variant_is_skipped(tmp_path): + """A rung with no frames is not playable, and picking it would blank the + pane rather than degrade it.""" + clip = a_clip() + clip["variants"].append({"name": "broken", "bytes": 10, "frames": []}) + result = run_js(f""" + const rungs = rungsOf({json.dumps(clip)}, 30); + process.stdout.write(JSON.stringify(rungs.map((r) => r.name))); + """, tmp_path) + assert "broken" not in result + + +# ------------------------------------------------------------- the decision --- + + +@requires_node +def test_the_best_affordable_rung_is_taken(tmp_path): + result = run_js(f""" + const rungs = rungsOf({json.dumps(a_clip())}, 30); + const at = (share) => affordable(rungs, share).name; + process.stdout.write(JSON.stringify({{ + starved: at(100e3), low: at(1e6), medium: at(3e6), + almost: at(8e6), plenty: at(20e6), + }})); + """, tmp_path) + assert result == { + "starved": "low", # nothing fits; the cheapest rather than nothing + "low": "low", + "medium": "medium", + "almost": "medium", # 8.04 Mbit/s default does not fit in 8 + "plenty": "default", + } + + +@requires_node +def test_a_starved_pane_shows_the_cheapest_rung_rather_than_nothing(tmp_path): + """A blank pane tells a viewer less than a coarse one, and the frames are + on disk either way -- this is a bundle, not a live encoder, so overrunning + costs lateness and not absence.""" + result = run_js(f""" + const rungs = rungsOf({json.dumps(a_clip())}, 30); + process.stdout.write(JSON.stringify(affordable(rungs, 0).name)); + """, tmp_path) + assert result == "low" + + +@requires_node +def test_no_rungs_means_no_choice(tmp_path): + result = run_js(""" + process.stdout.write(JSON.stringify(affordable([], 1e9))); + """, tmp_path) + assert result is None + + +# ------------------------------------------------- wired into the page --- + + +def test_the_scheduler_reports_to_a_meter(): + """Injected, not reached for as a global: a scheduler that read one could + not be run outside a page, and this one is.""" + page = viewer_path().read_text() + start = page.index("class Scheduler") + body = page[start:page.index("\n}\n", start)] + assert "meter = null" in body # injected with a default + assert "if (this.meter)" in body # and optional + assert "this.meter.record(" in body + + +def test_the_scheduler_fetches_the_chosen_rung(): + page = viewer_path().read_text() + start = page.index("class Scheduler") + body = page[start:page.index("\n}\n", start)] + assert "get frames()" in body + assert "this.root + this.frames[index]" in body + # Not the clip's own list, which would ignore the choice entirely. + assert "this.clip.frames[index]" not in body + + +def test_switching_rung_keeps_the_cache(): + """Renditions share a timeline, so a decoded frame is still the right + picture for its index. Dropping the cache would re-fetch what is in hand, + stalling exactly when the link is under pressure.""" + page = viewer_path().read_text() + start = page.index(" useRung(rung) {") + body = page[start:page.index("\n }\n", start)] + assert "cache" not in body + assert "release" not in body + + +def test_the_decision_runs_before_the_fetch(): + page = viewer_path().read_text() + start = page.index("function tick(") + body = page[start:page.index("\n}\n", start)] + assert body.index("chooseRungs()") < body.index("showFrame(app.frame + 1)") + + +def test_adaptation_can_be_pinned_off(): + """Comparing two methods usually means wanting both at their best whatever + the link can sustain; letting the link decide would quietly make the + comparison about bandwidth.""" + page = viewer_path().read_text() + start = page.index("function chooseRungs(") + body = page[start:page.index("\n}\n", start)] + assert "if (!app.adapt)" in body + assert "useRung(null)" in body + assert 'id="adapt"' in page + + +def test_the_client_splits_the_budget_evenly(): + """The decision that separates this from `streamer.policy`. If it ever + becomes a utility maximisation, this test should be the thing that objects. + + Split across the panes still *playing*, not all of them: a frozen pane is + fetching nothing, so counting it would leave its share unspent. + """ + page = viewer_path().read_text() + start = page.index("function chooseRungs(") + body = page[start:page.index("\n}\n", start)] + assert "rate / playing.length" in body + assert "utility" not in body and "weight" not in body + + +def test_the_client_freezes_under_deficit_rather_than_starving_everything(): + """Measured in `streamer.playback`: on a collapsing trace this turned 30 + seconds of stall into 56 of freeze -- the same shortfall taken as a + decision instead of a failure.""" + page = viewer_path().read_text() + start = page.index("function chooseRungs(") + body = page[start:page.index("\n}\n", start)] + assert "playing.pop()" in body + assert "pane.source.frozen = frozen.has(pane)" in body + + +def test_the_freeze_is_sticky(): + """Reconsidering from scratch would thaw one pane and freeze another, and a + viewer would see panes flickering rather than a stable subset playing.""" + page = viewer_path().read_text() + start = page.index("function chooseRungs(") + body = page[start:page.index("\n}\n", start)] + assert "froze(a) - froze(b)" in body + + +def test_a_frozen_pane_is_not_shown(): + """The difference between a freeze and a stall is that the rest of the view + stays in motion.""" + page = viewer_path().read_text() + start = page.index("async function showFrame(") + body = page[start:page.index("\n}\n", start)] + assert "pane.source.frozen" in body + assert ".filter(" in body + + +def test_the_client_budget_matches_the_measured_rule(): + """`bufferBudget` mirrors `streamer.playback.buffer_budget`, where the same + rule is measured. Two implementations of one rule drift, so the clamp is + pinned on both sides.""" + from streamer.playback import buffer_budget + + page = viewer_path().read_text() + start = page.index("function bufferBudget(") + body = page[start:page.index("\n}\n", start)] + assert "0.5" in body and "1.5" in body + # And the Python side agrees at both ends of the clamp. + assert buffer_budget(10e6, 0.0, 4.0) == pytest.approx(5e6) + assert buffer_budget(10e6, 99.0, 4.0) == pytest.approx(15e6) + + +def test_live_panes_are_left_alone(): + """A live stream has no frame list, so there is no rung to choose.""" + page = viewer_path().read_text() + start = page.index("function chooseRungs(") + body = page[start:page.index("\n}\n", start)] + assert "!isLive(pane.clip)" in body diff --git a/open4d/streamer/streamer_tests/test_adopt.py b/open4d/streamer/streamer_tests/test_adopt.py new file mode 100644 index 00000000..181b6c19 --- /dev/null +++ b/open4d/streamer/streamer_tests/test_adopt.py @@ -0,0 +1,339 @@ +"""Importing frames a method exported in another interpreter. + +`streamer.adopt` exists because some methods cannot be driven from this package +at all -- ReRF's entropy coder is a Python 3.8 binary, and this needs 3.10 -- +so the handoff is a directory plus a sidecar. These check the handoff, since a +malformed one is otherwise discovered as a pane that plays for two seconds and +then 404s. +""" +from __future__ import annotations + +import json +from pathlib import Path + +import pytest + +from streamer import adopt, bundle + +pytestmark = pytest.mark.cpu + + +def an_export(root, *, clips=2, frames=3, name="obj", missing=False): + root.mkdir(parents=True, exist_ok=True) + entries = [] + for index in range(clips): + clip = f"{name}-cam{index:02d}" + paths = [] + for frame in range(frames): + relative = f"{clip}/frame_{frame:04d}.jpg" + paths.append(relative) + if missing and frame == frames - 1 and index == 0: + continue # the export died partway through + target = root / relative + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(b"\xff\xd8\xff\xd9") + entries.append({ + "name": clip, "method": "rerf", "camera": index, "frames": paths, + "notes": ["a note the renderer wanted shown"], + "detail": {"view": index}, + }) + (root / "clips.json").write_text(json.dumps({ + "format": "rerf-clips", "version": 1, "scene": "basketball", + "representation": "pixels", "clips": entries, + })) + return root + + +def a_bundle(root): + root.mkdir(parents=True, exist_ok=True) + bundle.write(root, title="t", source="s", fps=24, + scenes={"basketball": {"poses": []}}, + clips=[bundle.Clip(name="already", representation="mesh", + frames=["already/f.ply"])]) + return root + + +def test_frames_are_copied_into_the_bundle(tmp_path): + """A bundle has to be servable and fetchable whole, so a clip may not point + outside it.""" + export = an_export(tmp_path / "export") + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + + index = bundle.read(root) + imported = [c for c in index["clips"] if c["name"].startswith("obj-")] + assert len(imported) == 2 + for clip in imported: + for relative in clip["frames"]: + assert (root / relative).is_file(), relative + + +def test_the_scene_and_representation_come_from_the_sidecar(tmp_path): + export = an_export(tmp_path / "export") + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + clip = next(c for c in bundle.read(root)["clips"] if c["name"] == "obj-cam00") + assert clip["scene"] == "basketball" + assert clip["representation"] == "pixels" + assert clip["camera"] == 0 + assert clip["notes"] == ["a note the renderer wanted shown"] + + +def test_existing_clips_survive(tmp_path): + export = an_export(tmp_path / "export") + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + names = [c["name"] for c in bundle.read(root)["clips"]] + assert names[0] == "already" + assert bundle.read(root)["fps"] == 24 + + +def test_a_half_written_export_is_refused(tmp_path): + """Better than importing it: a clip whose frames are partly there plays and + then 404s, and the manifest says nothing is wrong.""" + export = an_export(tmp_path / "export", missing=True) + root = a_bundle(tmp_path / "view") + with pytest.raises(FileNotFoundError, match="did not finish"): + adopt.adopt(export, root) + # And nothing was added. + assert [c["name"] for c in bundle.read(root)["clips"]] == ["already"] + + +def test_re_exporting_needs_replace(tmp_path): + export = an_export(tmp_path / "export") + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + with pytest.raises(ValueError, match="already has a clip named"): + adopt.adopt(export, root) + adopt.adopt(export, root, replace=True) + names = [c["name"] for c in bundle.read(root)["clips"]] + assert names.count("obj-cam00") == 1 + + +def test_a_missing_sidecar_names_the_exporter(tmp_path): + root = a_bundle(tmp_path / "view") + (tmp_path / "empty").mkdir() + with pytest.raises(FileNotFoundError, match="rerf_stream.export"): + adopt.adopt(tmp_path / "empty", root) + + +def test_an_unknown_format_is_refused(tmp_path): + export = an_export(tmp_path / "export") + payload = json.loads((export / "clips.json").read_text()) + payload["version"] = 99 + (export / "clips.json").write_text(json.dumps(payload)) + root = a_bundle(tmp_path / "view") + with pytest.raises(ValueError, match="not a format this understands"): + adopt.adopt(export, root) + + +def test_an_export_with_no_clips_is_refused(tmp_path): + export = an_export(tmp_path / "export") + (export / "clips.json").write_text(json.dumps({ + "format": "rerf-clips", "version": 1, "clips": []})) + with pytest.raises(ValueError, match="lists no clips"): + adopt.adopt(export, a_bundle(tmp_path / "view")) + + +# ------------------------------------------------------------ the scene rig --- + + +def a_rig(stations=3): + return { + "width": 1280, "height": 960, "fov_y": 0.7, + "bounds_min": [-1.0, 0.0, -1.0], "bounds_max": [1.0, 2.0, 1.0], + "poses": [ + {"position": [float(n), 0.0, 0.0], "right": [1.0, 0.0, 0.0], + "down": [0.0, 1.0, 0.0], "forward": [0.0, 0.0, 1.0]} + for n in range(stations) + ], + } + + +def with_rig(export, rig): + payload = json.loads((export / "clips.json").read_text()) + payload["rig"] = rig + (export / "clips.json").write_text(json.dumps(payload)) + return export + + +def test_the_rig_is_installed_for_a_scene_the_bundle_did_not_know(tmp_path): + """Without one, a viewer cannot offer station selection, so the scene's + panes are shown but not comparable to each other by pose.""" + export = with_rig(an_export(tmp_path / "export"), a_rig()) + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + + scenes = bundle.read(root)["scenes"] + assert "basketball" in scenes + assert len(scenes["basketball"]["poses"]) == 3 + assert scenes["basketball"]["scene"] == "basketball" + + +def test_an_existing_rig_is_not_overwritten(tmp_path): + """A rig that came from a geometry method is the authority: that is the + output which has to line up in 3D, and silently replacing it would move + every other method's camera.""" + root = a_bundle(tmp_path / "view") + index = bundle.read(root) + bundle.write( + root, title=index["title"], source=index["source"], + clips=[bundle.Clip(**c) for c in index["clips"]], fps=index["fps"], + scenes={"basketball": dict(a_rig(stations=8), scene="basketball", + origin="from-geometry")}, + ) + adopt.adopt(with_rig(an_export(tmp_path / "export"), a_rig(stations=3)), root) + + scene = bundle.read(root)["scenes"]["basketball"] + assert len(scene["poses"]) == 8 + assert scene["origin"] == "from-geometry" + + +def test_an_export_without_a_rig_still_imports(tmp_path): + """Not every method knows the rig, and clips are worth having regardless.""" + export = an_export(tmp_path / "export") + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + assert bundle.read(root)["scenes"] == {"basketball": {"poses": []}} + assert any(c["name"] == "obj-cam00" for c in bundle.read(root)["clips"]) + + +def test_installing_the_rig_keeps_the_clips_that_were_there(tmp_path): + """It rewrites the manifest, so the pre-existing clips have to survive.""" + export = with_rig(an_export(tmp_path / "export"), a_rig()) + root = a_bundle(tmp_path / "view") + adopt.adopt(export, root) + names = [c["name"] for c in bundle.read(root)["clips"]] + assert "already" in names + assert len(names) == 3 + + +# ------------------------------------------------ adopting a whole bundle --- +# `gs-tools export` writes a bundle, not a `clips.json` sidecar, so building one +# bundle from several exporters used to need a hand-rolled merge. A bundle is a +# superset of the sidecar, so it is read here instead. + + +def _bundle_at(directory: Path, clips, scenes=None) -> Path: + for clip in clips: + for relative in clip.frames: + target = directory / relative + target.parent.mkdir(parents=True, exist_ok=True) + target.write_bytes(b"x" * 16) + return bundle.write(directory, title="staged", source="somewhere", + clips=list(clips), scenes=scenes or {}) + + +def test_a_staged_bundle_can_be_adopted(tmp_path): + target = tmp_path / "target" + _bundle_at(target, [bundle.Clip(name="a", representation="pixels", scene="s", + method="rerf", camera=0, + frames=["a/frame_0000.jpg"])]) + staged = tmp_path / "staged" + _bundle_at(staged, [bundle.Clip(name="v", representation="gaussians", scene="s", + method="vega", counts=[57907], + frames=["v/frame_0000.splat"])]) + + adopt.adopt(staged, target) + index = bundle.read(target) + names = {clip["name"]: clip for clip in index["clips"]} + assert set(names) == {"a", "v"} + # Copied, not referenced: a bundle has to be servable and movable whole. + assert (target / "v/frame_0000.splat").is_file() + # And the fields a sidecar does not have survive. + assert names["v"]["counts"] == [57907] + assert names["v"]["representation"] == "gaussians" + + +def test_a_bundle_s_clips_keep_their_own_representation(tmp_path): + """A sidecar states one representation for the whole export; a bundle + carries a mix, and taking the top-level value would relabel a point cloud + as pixels -- which sends it to the image decoder and blanks the pane.""" + target = tmp_path / "target" + _bundle_at(target, [bundle.Clip(name="keep", representation="pixels", scene="s", + frames=["keep/frame_0000.jpg"])]) + staged = tmp_path / "staged" + _bundle_at(staged, [ + bundle.Clip(name="px", representation="pixels", scene="s", method="rerf", + camera=0, frames=["px/frame_0000.jpg"]), + bundle.Clip(name="pts", representation="points", scene="s", method="rerf", + frames=["pts/frame_0000.ply"]), + ]) + + adopt.adopt(staged, target) + got = {clip["name"]: clip["representation"] for clip in bundle.read(target)["clips"]} + assert got["px"] == "pixels" + assert got["pts"] == "points" + + +def test_a_directory_with_neither_names_both(tmp_path): + with pytest.raises(FileNotFoundError, match="neither clips.json nor a bundle"): + adopt.read_export(tmp_path) + + +def test_a_bundle_brings_a_rig_for_each_of_its_scenes(tmp_path): + """A sidecar names one rig for one scene; a bundle carries several.""" + target = tmp_path / "target" + _bundle_at(target, [bundle.Clip(name="keep", representation="pixels", scene="a", + frames=["keep/frame_0000.jpg"])]) + staged = tmp_path / "staged" + rig = {"width": 8, "height": 8, "fov_y": 1.0, + "poses": [{"position": [0, 0, 1], "right": [1, 0, 0], + "down": [0, 1, 0], "forward": [0, 0, -1]}]} + _bundle_at( + staged, + [bundle.Clip(name="a1", representation="gaussians", scene="a", + frames=["a1/frame_0000.splat"]), + bundle.Clip(name="b1", representation="gaussians", scene="b", + frames=["b1/frame_0000.splat"])], + scenes={"a": dict(rig), "b": dict(rig)}, + ) + + adopt.adopt(staged, target) + scenes = bundle.read(target)["scenes"] + assert set(scenes) == {"a", "b"} + assert all(len(scene["poses"]) == 1 for scene in scenes.values()) + + +# ------------------------------------- a camera has to index a real rig --- + + +def test_a_clip_numbered_past_its_rig_is_refused(tmp_path): + """Reachable by import order alone. + + An export installs a rig only into a scene that has none, so whichever + lands first wins. A Vega export brings the corpus's 8-camera capture rig; a + ReRF orbit export brings 216 stations. Import them the other way round and + the orbit clips are numbered against a rig that stops at 7 -- and most + Compare stations show nothing, silently, because an empty pane is a + legitimate state. + """ + target = tmp_path / "target" + small = {"width": 8, "height": 8, "fov_y": 1.0, + "poses": [{"position": [0, 0, 1], "right": [1, 0, 0], + "down": [0, 1, 0], "forward": [0, 0, -1]}]} + _bundle_at(target, [bundle.Clip(name="keep", representation="gaussians", + scene="s", frames=["keep/frame_0000.splat"])], + scenes={"s": small}) + staged = tmp_path / "staged" + _bundle_at(staged, [bundle.Clip(name="cam05", representation="pixels", scene="s", + method="rerf", camera=5, + frames=["cam05/frame_0000.jpg"])]) + + with pytest.raises(ValueError, match="camera 5, but scene 's' has 1 rig pose"): + adopt.adopt(staged, target) + + +def test_a_scene_with_no_rig_is_explore_only_not_an_error(tmp_path): + """No rig means no station selection, which the viewer reports. Only a rig + that exists and is too short is a mismatch.""" + target = tmp_path / "target" + _bundle_at(target, [bundle.Clip(name="keep", representation="gaussians", + scene="s", frames=["keep/frame_0000.splat"])]) + staged = tmp_path / "staged" + _bundle_at(staged, [bundle.Clip(name="cam09", representation="pixels", scene="s", + method="rerf", camera=9, + frames=["cam09/frame_0000.jpg"])]) + + adopt.adopt(staged, target) # no raise + assert len(bundle.read(target)["clips"]) == 2 diff --git a/open4d/streamer/streamer_tests/test_client_page.py b/open4d/streamer/streamer_tests/test_client_page.py new file mode 100644 index 00000000..4e347635 --- /dev/null +++ b/open4d/streamer/streamer_tests/test_client_page.py @@ -0,0 +1,261 @@ +"""Whole-page checks on the client, of the kind a browser would catch too late. + +The page is edited by hand and served as one file, so a syntax error or a +dangling reference ships silently: the server returns 200, the browser reports +it in a console nobody is watching, and every pane is blank. These are the two +cheapest checks that catch that class of mistake -- the script parses, and the +representation registry actually evaluates. +""" + +from __future__ import annotations + +import json +import shutil +import subprocess +from pathlib import Path + +import pytest +from open4d.core import Representation + +from streamer import representations +from streamer.client import viewer_path + +pytestmark = pytest.mark.cpu + +NODE = shutil.which("node") +requires_node = pytest.mark.skipif(NODE is None, reason="node is not installed") + + +def page() -> str: + return viewer_path().read_text() + + +def script() -> str: + """The page's script, which is where everything that can break lives.""" + source = page() + start = source.index("")] + + +def cut(name: str) -> str: + """One top-level definition, by the name the page gives it. + + `async function X(` is tried first because `function X(` is a substring of + it: matching the shorter one drops the `async` keyword, and the extracted + text then fails to parse on its own `await`. + """ + source = page() + for prefix in ( + f"async function {name}(", + f"function {name}(", + f"const {name} =", + f"let {name} =", + f"class {name} ", + ): + start = source.find(prefix) + if start < 0: + continue + if prefix.startswith(("const", "let")): + return source[start : source.index(";\n", start) + 2] + line = source[start : source.index("\n", start)] + if line.count("{") and line.count("{") == line.count("}"): + return line + return source[start : source.index("\n}\n", start) + 3] + raise AssertionError(f"{name} is not defined in the viewer") + + +@requires_node +def test_the_script_parses(tmp_path): + path = tmp_path / "viewer.js" + path.write_text(script()) + finished = subprocess.run( + [NODE, "--check", str(path)], capture_output=True, text=True, timeout=120 + ) + assert finished.returncode == 0, finished.stderr + + +@requires_node +def test_the_representation_registry_evaluates(tmp_path): + """It is a const literal that calls functions declared further down the file. + + That works only because those are function declarations and therefore + hoisted. Reordering them into `const` arrow functions would leave the + literal reading them in their temporal dead zone -- which fails at load, + before anything renders, and only in a browser. + """ + # Everything the literal reaches while it is being evaluated. Names the + # literal only calls later -- parseDraco, loadDraco -- are not needed here, + # which is why this is a list rather than the whole script: what is being + # tested is that the literal can be built at load, not that the page runs. + # The literal reaches `decodeFrame` and `decodeImage` while it is being + # evaluated. Both are function declarations, so both are hoisted; the parsers + # they eventually reach live in the worker and are not needed here. + probe = "\n".join( + cut(name) for name in ("REPRESENTATIONS", "decodeFrame", "decodeImage") + ) + """ + const out = {}; + for (const [name, spec] of Object.entries(REPRESENTATIONS)) { + out[name] = { + decode: typeof spec.decode, + geometry: spec.geometry, + cacheSize: spec.cacheSize, + }; + } + process.stdout.write(JSON.stringify(out)); + """ + path = tmp_path / "registry.mjs" + path.write_text(probe) + finished = subprocess.run( + [NODE, str(path)], capture_output=True, text=True, timeout=120 + ) + assert finished.returncode == 0, finished.stderr + entries = json.loads(finished.stdout) + + assert set(entries) == {member.value for member in Representation} + for name, entry in entries.items(): + assert entry["decode"] == "function", name + assert entry["geometry"] is Representation(name).has_geometry, name + assert isinstance(entry["cacheSize"], int) and entry["cacheSize"] > 0, name + + +@requires_node +def test_frames_display_in_order_under_varying_decode_latency(tmp_path): + """Why playback has to be buffer-driven, demonstrated on both patterns. + + Decode latency varies per frame -- that is what a fetch does -- and an + advance loop that fires on a wall clock without waiting overlaps its own + calls, which then resolve in whatever order they finish. The result is not + slow playback but *wrong* playback: the counter races ahead and the panes + show whichever decode landed last. Measured here at 3,1,5,0,6,4,2 for a + request of 0..6. + + This simulates the two patterns rather than driving the real `tick`, which + needs a DOM. `test_the_playback_loop_waits_for_the_frame_it_asked_for` + checks that the shipped loop still uses the pattern this one vindicates. + """ + script = tmp_path / "loop.mjs" + script.write_text( + """ + const LATENCY = [40, 15, 60, 10, 50, 12, 45, 18, 55, 11]; + const FPS = 30; + function makeShow(shown) { + return async (index) => { + await new Promise((r) => setTimeout(r, LATENCY[index % LATENCY.length])); + shown.push(index); + }; + } + async function wallClock(steps) { + const shown = []; const show = makeShow(shown); + let frame = 0, last = 0, now = 0; + for (let i = 0; i < steps; i++) { + now += 1000 / FPS; + if (now - last >= 1000 / FPS) { last = now; show(frame++); } + await new Promise((r) => setTimeout(r, 1)); + } + await new Promise((r) => setTimeout(r, 300)); + return shown; + } + async function bufferDriven(steps) { + const shown = []; const show = makeShow(shown); + let frame = 0, advancing = false; + for (let i = 0; i < steps * 12; i++) { + if (!advancing) { + advancing = true; + show(frame++).finally(() => { advancing = false; }); + } + await new Promise((r) => setTimeout(r, 1)); + if (shown.length >= steps) break; + } + return shown; + } + process.stdout.write(JSON.stringify({ + wallClock: (await wallClock(10)).slice(0, 7), + bufferDriven: (await bufferDriven(7)).slice(0, 7), + })); + """ + ) + finished = subprocess.run( + [NODE, str(script)], capture_output=True, text=True, timeout=120 + ) + assert finished.returncode == 0, finished.stderr + result = json.loads(finished.stdout) + + ordered = lambda seq: all(b > a for a, b in zip(seq, seq[1:])) + assert not ordered(result["wallClock"]), result["wallClock"] + assert ordered(result["bufferDriven"]), result["bufferDriven"] + assert result["bufferDriven"] == sorted(result["bufferDriven"]) + + +def test_the_playback_loop_waits_for_the_frame_it_asked_for(): + """Structural, because `tick` needs a DOM to run. + + Weak on its own, which is why it names what it is guarding: the invariant is + at most one advance in flight, and the next interval timed from when a frame + was actually shown. + """ + source = page() + start = source.index("function tick(") + body = source[start : source.index("\n}\n", start)] + assert "!app.advancing" in body, "the advance guard is gone" + assert "app.advancing = true" in body + assert ".finally(" in body, "the guard is never cleared" + assert "app.lastAdvance = performance.now()" in body, ( + "the interval is timed from the clock again, not from the frame" + ) + + +def test_the_page_carries_no_external_dependency(): + """No build step and no CDN is what makes the page servable as one file.""" + source = page() + for pattern in ("http://", "https://", "cdn.", " + + diff --git a/open4d/webclients/system/WebClient/public/compare.html b/open4d/webclients/system/WebClient/public/compare.html new file mode 100644 index 00000000..6c362bb1 --- /dev/null +++ b/open4d/webclients/system/WebClient/public/compare.html @@ -0,0 +1,62 @@ + + + + + +4DVideoStreaming + + + +
+

Pick a system

+

loading …

+
+
+
+ Shaping: sudo scripts/shape_web_demo.sh cascade-20. +
+ + + diff --git a/open4d/webclients/system/WebClient/public/index.html b/open4d/webclients/system/WebClient/public/index.html new file mode 100644 index 00000000..83f7809a --- /dev/null +++ b/open4d/webclients/system/WebClient/public/index.html @@ -0,0 +1,85 @@ + + + + + +4DVideoStreaming — browser client + + + + + + + + diff --git a/open4d/webclients/system/WebClient/public/nevo.html b/open4d/webclients/system/WebClient/public/nevo.html new file mode 100644 index 00000000..e4eb06f3 --- /dev/null +++ b/open4d/webclients/system/WebClient/public/nevo.html @@ -0,0 +1,82 @@ + + + + + +4DVideoStreaming — NeVo viewer + + + + + + + + diff --git a/open4d/webclients/system/WebClient/public/vega.html b/open4d/webclients/system/WebClient/public/vega.html new file mode 100644 index 00000000..8d007f53 --- /dev/null +++ b/open4d/webclients/system/WebClient/public/vega.html @@ -0,0 +1,84 @@ + + + + + +4DVideoStreaming — Vega splat viewer + + + + + + + + diff --git a/open4d/webclients/system/WebClient/src/baseline-client.js b/open4d/webclients/system/WebClient/src/baseline-client.js new file mode 100644 index 00000000..e0e9e3c0 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/baseline-client.js @@ -0,0 +1,397 @@ +'use strict'; + +/** + * Browser client for the V4DS point-cloud baselines. + * + * MetaStream, DeltaStream, ViVo, NAVA and LiVo all speak the same protocol, so + * one client serves all five; the mode arrives in the CONNECTION header rather + * than being configured here. + * + * This does NOT use `ClientCore`, and that is the right call rather than an + * omission. ClientCore implements *our* system: an HTTP segment loop, a + * published ladder, an MCKP selector choosing one representation per object per + * segment. A baseline pushes frames at its own cadence over a socket and makes + * its own adaptation decisions server-side. Wrapping that in a segment loop + * would model the baselines as something they are not, and the comparison would + * measure the wrapper. + * + * What it does: + * 1. connect to the bridge, which proxies the baseline's TCP socket + * 2. decode CONNECTION -> calibrations, stream mode, tile catalogue + * 3. per FRAME: decode every record's Draco point cloud in the worker, fold + * it through `ReconstructionState`, hand the world clouds to the renderer + * 4. send FEEDBACK at the content rate: displayed frame, measured fps, + * measured goodput, and the live camera + */ + +const { + decodeConnection, decodeFrame, encodeFeedback, messageType, MessageType +} = require('./v4ds-protocol'); +const { ReconstructionState, PointCloud } = require('./point-reconstruction'); +const { PointRenderer } = require('./point-renderer'); + +/** Rolling goodput estimate over the received WebSocket bytes. */ +class GoodputMeter { + constructor({ windowMs = 2000, now = () => performance.now() } = {}) { + this.windowMs = windowMs; + this._now = now; + this._samples = []; // { at, bytes } + } + + record(bytes) { + const at = this._now(); + this._samples.push({ at, bytes }); + const cutoff = at - this.windowMs; + while (this._samples.length && this._samples[0].at < cutoff) { + this._samples.shift(); + } + } + + /** Mbps over the window, or null before there is enough to divide by. */ + get mbps() { + if (this._samples.length < 2) return null; + const span = this._samples[this._samples.length - 1].at - this._samples[0].at; + if (span <= 0) return null; + const bytes = this._samples.reduce((sum, s) => sum + s.bytes, 0); + return (bytes * 8) / span / 1000; + } +} + +class BaselineClient { + /** + * @param {object} args + * @param {string} args.bridgeUrl ws:// address of v4ds-bridge + * @param {HTMLCanvasElement} args.canvas + * @param {string} [args.workerUrl] + * @param {string} [args.vendorBase] + * @param {number} [args.feedbackIntervalMs] + * @param {boolean} [args.strictOrder] throw on a frame gap (default false + * in the browser: a lossy link is normal, and resynchronising on the next + * keyframe beats aborting the run) + * @param {(event: object) => void} [args.onEvent] + */ + constructor({ + bridgeUrl, canvas, workerUrl = '/web/draco-worker.js', + vendorBase = '/web/vendor/draco', feedbackIntervalMs = 200, + strictOrder = false, onEvent = null, pointSize = 0.012 + }) { + this.bridgeUrl = bridgeUrl; + this.workerUrl = workerUrl; + this.vendorBase = vendorBase; + this.feedbackIntervalMs = feedbackIntervalMs; + this.strictOrder = strictOrder; + this._onEvent = onEvent; + + this.renderer = new PointRenderer({ canvas, pointSize }); + this.header = null; + this.state = null; + this.mode = null; + + this.stats = { + framesReceived: 0, framesReconstructed: 0, framesDropped: 0, + bytesReceived: 0, resyncs: 0, decodeFailures: 0, + lastFrameId: -1, displayedFrameId: -1, measuredFps: 0 + }; + + this._socket = null; + this._worker = null; + this._workerSeq = 0; + this._workerWaiters = new Map(); + this._feedbackTimer = null; + this._goodput = new GoodputMeter(); + this._frameTimestamps = []; + this._busy = false; + this._pendingFrame = null; + this._awaitingKeyframe = false; + } + + _emit(type, detail = {}) { + this._onEvent?.({ type, ...detail }); + } + + async start() { + this.renderer.start(); + this._worker = new Worker(this.workerUrl); + this._worker.onmessage = event => { + const waiter = this._workerWaiters.get(event.data.id); + this._workerWaiters.delete(event.data.id); + if (!waiter) return; + if (event.data.error) waiter.reject(new Error(event.data.error)); + else waiter.resolve(event.data.frames); + }; + this._worker.onerror = err => + this._emit('error', { message: `draco worker: ${err.message}` }); + + await this._connect(); + this._feedbackTimer = setInterval( + () => this._sendFeedback(), this.feedbackIntervalMs); + } + + _connect() { + return new Promise((resolve, reject) => { + const socket = new WebSocket(this.bridgeUrl); + socket.binaryType = 'arraybuffer'; + this._socket = socket; + + socket.onopen = () => { + this._emit('open', { url: this.bridgeUrl }); + resolve(); + }; + socket.onerror = () => { + const message = `could not reach the bridge at ${this.bridgeUrl}`; + this._emit('error', { message }); + reject(new Error(message)); + }; + socket.onclose = event => this._emit('closed', { + code: event.code, + reason: event.reason || '(no reason given)' + }); + socket.onmessage = event => this._onMessage(event.data); + }); + } + + async _onMessage(buffer) { + this.stats.bytesReceived += buffer.byteLength; + this._goodput.record(buffer.byteLength); + + let type; + try { + type = messageType(buffer); + } catch (err) { + this._emit('error', { message: `bad message: ${err.message}` }); + return; + } + + if (type === MessageType.CONNECTION) { + this._onConnectionHeader(buffer); + return; + } + if (type === MessageType.FRAME) { + // Only one frame is reconstructed at a time; a newer frame replaces + // any frame still waiting, because showing the freshest content + // matters more than showing every frame of a backlog. + if (this._busy) { + if (this._pendingFrame) this.stats.framesDropped++; + this._pendingFrame = buffer; + return; + } + await this._drainFrames(buffer); + return; + } + if (type === MessageType.LIVO_SEGMENT) { + // LiVo ships HEVC colour+depth rather than point clouds; it needs + // the texture decoder and an unprojection step, not this path. + this._emit('unsupported', { + message: 'LiVo segments need RGB-D unprojection, not implemented' + }); + return; + } + this._emit('error', { message: `unexpected message type ${type}` }); + } + + _onConnectionHeader(buffer) { + try { + this.header = decodeConnection(buffer); + } catch (err) { + this._emit('error', { message: `bad CONNECTION: ${err.message}` }); + return; + } + this.mode = this.header.mode; + this.state = new ReconstructionState(this.header); + this._emit('header', { + mode: this.header.mode, + objects: this.header.objects.map(o => ({ + objectId: o.objectId, name: o.name, + cameras: o.cameras.length, loopFrames: o.loopFrames + })), + fps: this.header.fps, + source: `${this.header.width}x${this.header.height}`, + blockSize: this.header.blockSize, + tiled: Boolean(this.header.tileAbr) + }); + } + + async _drainFrames(first) { + this._busy = true; + let buffer = first; + try { + while (buffer) { + await this._handleFrame(buffer); + buffer = this._pendingFrame; + this._pendingFrame = null; + } + } finally { + this._busy = false; + } + } + + async _handleFrame(buffer) { + if (!this.state) { + this._emit('error', { message: 'FRAME arrived before CONNECTION' }); + return; + } + let frame; + try { + frame = decodeFrame(buffer); + } catch (err) { + this._emit('error', { message: `bad FRAME: ${err.message}` }); + return; + } + this.stats.framesReceived++; + this.stats.lastFrameId = frame.frameId; + + // After a gap the delta chain is broken; wait for a keyframe rather + // than compounding the error into visible corruption. + if (this._awaitingKeyframe) { + if (frame.frameType !== 'keyframe') return; + this._awaitingKeyframe = false; + this.state.reset(); + } + + const clouds = await this._decodeRecords(frame); + if (!clouds) return; + + try { + const worlds = this.state.apply(frame, blob => { + const cloud = clouds.get(blobKey(blob)); + return cloud || PointCloud.empty(); + }, { strictOrder: this.strictOrder }); + this.renderer.update(worlds); + this.stats.framesReconstructed++; + this.stats.displayedFrameId = frame.frameId; + this._recordPresentation(); + } catch (err) { + this.stats.resyncs++; + this._awaitingKeyframe = true; + this._emit('resync', { message: err.message, frameId: frame.frameId }); + } + } + + /** + * Decode every record's Draco payload in one worker round trip. + * + * `ReconstructionState.apply` wants a synchronous decoder, so the payloads + * are decoded up front and looked up by identity. One round trip per frame + * also beats one per record: a nine-object scene with four cameras is 36 + * payloads, and 36 postMessage round trips would not fit in a frame budget. + */ + async _decodeRecords(frame) { + const withPayload = frame.records.filter(r => r.draco && r.draco.length); + if (withPayload.length === 0) return new Map(); + + // COPY, do not transfer. Transferring `record.draco.buffer` detaches it, + // after which `record.draco.length` reads 0 — and + // ReconstructionState.apply uses exactly that length to decide whether a + // record has a payload. It would silently treat every record as empty + // and then fail the point-count check. A memcpy of a few hundred KB is + // negligible beside the Draco decode it feeds. + const buffers = withPayload.map(r => r.draco.slice().buffer); + const id = ++this._workerSeq; + let decoded; + try { + decoded = await new Promise((resolve, reject) => { + this._workerWaiters.set(id, { resolve, reject }); + this._worker.postMessage( + { id, buffers, vendorBase: this.vendorBase }, buffers); + }); + } catch (err) { + this.stats.decodeFailures++; + this._emit('error', { message: `draco decode: ${err.message}` }); + return null; + } + + const clouds = new Map(); + withPayload.forEach((record, index) => { + const frameData = decoded[index]; + if (!frameData || frameData.kind !== 'cloud') { + this.stats.decodeFailures++; + clouds.set(blobKey(record.draco), PointCloud.empty()); + return; + } + clouds.set(blobKey(record.draco), new PointCloud( + frameData.positions, + frameData.colors || new Uint8Array(frameData.positions.length))); + }); + return clouds; + } + + _recordPresentation() { + const now = performance.now(); + this._frameTimestamps.push(now); + while (this._frameTimestamps.length + && this._frameTimestamps[0] < now - 2000) { + this._frameTimestamps.shift(); + } + if (this._frameTimestamps.length >= 2) { + const span = now - this._frameTimestamps[0]; + this.stats.measuredFps = + ((this._frameTimestamps.length - 1) * 1000) / span; + } + } + + _sendFeedback() { + if (!this._socket || this._socket.readyState !== WebSocket.OPEN) return; + if (!this.header) return; + const viewer = this.renderer.viewerState(); + const bandwidth = this._goodput.mbps; + + // The frustum tier requires the extended tier, and both are all-or- + // nothing; sending a partial one is a protocol error the server rejects. + const payload = { + displayedFrameId: Math.max(0, this.stats.displayedFrameId), + measuredFps: this.stats.measuredFps || this.header.fps, + repeatedFrames: this.stats.framesDropped + }; + if (bandwidth !== null) { + Object.assign(payload, { + viewPosition: viewer.position, + viewForward: viewer.forward, + bandwidthMbps: bandwidth, + viewUp: viewer.up, + verticalFovDegrees: viewer.verticalFovDegrees, + viewAspect: viewer.aspect, + viewNear: viewer.near, + viewFar: viewer.far + }); + } + try { + this._socket.send(encodeFeedback(payload)); + } catch (err) { + this._emit('error', { message: `feedback: ${err.message}` }); + } + } + + stop() { + if (this._feedbackTimer) clearInterval(this._feedbackTimer); + this._feedbackTimer = null; + try { this._socket?.close(1000, 'client stopped'); } catch (_) { /* closed */ } + this._worker?.terminate(); + this._worker = null; + this.renderer.stop(); + } + + inspect() { + return { + mode: this.mode, + header: this.header && { + objects: this.header.objects.length, + fps: this.header.fps, + tiled: Boolean(this.header.tileAbr) + }, + stats: { ...this.stats, goodputMbps: this._goodput.mbps }, + renderer: this.renderer.stats, + viewer: this.header ? this.renderer.viewerState() : null, + awaitingKeyframe: this._awaitingKeyframe + }; + } +} + +/** Identity key for a payload, so the sync decoder callback can look it up. */ +let blobCounter = 0; +const blobKeys = new WeakMap(); +function blobKey(blob) { + if (!blobKeys.has(blob)) blobKeys.set(blob, ++blobCounter); + return blobKeys.get(blob); +} + +module.exports = { BaselineClient, GoodputMeter }; diff --git a/open4d/webclients/system/WebClient/src/baseline-main.js b/open4d/webclients/system/WebClient/src/baseline-main.js new file mode 100644 index 00000000..4dc96a1c --- /dev/null +++ b/open4d/webclients/system/WebClient/src/baseline-main.js @@ -0,0 +1,147 @@ +'use strict'; + +/** + * Entry point for the V4DS baseline viewer. + * + * Separate from `main.js` because the two are genuinely different clients: that + * one drives our HTTP segment ladder through ClientCore, this one consumes a + * pushed socket stream. Sharing an entry would mean a mode flag that changes + * almost everything. + * + * /web/baseline.html?bridge=ws://host:8790 + * + * | parameter | default | meaning | + * |-----------|----------------------------|--------------------------------| + * | bridge | ws://:8790 | v4ds-bridge WebSocket address | + * | pointSize | 0.012 | point size in metres | + * | strict | 0 | 1 = abort on a frame gap | + */ + +const { BaselineClient } = require('./baseline-client'); +const { mountLinkRate } = require('./link-rate'); + +function readConfig() { + const params = new URLSearchParams(window.location.search); + const host = window.location.hostname || '127.0.0.1'; + const number = (name, fallback) => { + const value = Number(params.get(name)); + return params.get(name) !== null && Number.isFinite(value) ? value : fallback; + }; + return { + bridgeUrl: params.get('bridge') || `ws://${host}:8790`, + pointSize: number('pointSize', 0.012), + strictOrder: params.get('strict') === '1' + }; +} + +function createUi() { + const status = document.getElementById('status'); + const logPane = document.getElementById('log'); + const statsPane = document.getElementById('stats'); + return { + setStatus: text => { status.textContent = text; }, + log(level, message) { + const row = document.createElement('div'); + row.className = `log-line log-${level}`; + row.textContent = `[${new Date().toISOString().slice(11, 23)}] ${message}`; + logPane.appendChild(row); + while (logPane.childElementCount > 300) { + logPane.removeChild(logPane.firstChild); + } + logPane.scrollTop = logPane.scrollHeight; + }, + setStats(lines) { statsPane.textContent = lines.join('\n'); } + }; +} + +async function main() { + const config = readConfig(); + const ui = createUi(); + const stopLinkRate = mountLinkRate(document.getElementById('link')); + ui.setStatus(`connecting to ${config.bridgeUrl} …`); + + const client = new BaselineClient({ + bridgeUrl: config.bridgeUrl, + canvas: document.getElementById('view'), + pointSize: config.pointSize, + strictOrder: config.strictOrder, + onEvent: event => { + switch (event.type) { + case 'open': + ui.log('info', `bridge connected: ${event.url}`); + ui.setStatus('waiting for the stream header …'); + break; + case 'header': + ui.log('info', `${event.mode.toUpperCase()} · ` + + `${event.objects.length} objects · ${event.fps} fps · ` + + `source ${event.source}` + + (event.tiled ? ' · tiled/ABR' : '')); + for (const o of event.objects) { + ui.log('info', ` object ${o.objectId} ${o.name}: ` + + `${o.cameras} cameras, ${o.loopFrames} frames`); + } + ui.setStatus(`streaming ${event.mode}`); + break; + case 'resync': + ui.log('warn', `resynchronising: ${event.message}`); + break; + case 'unsupported': + ui.log('warn', event.message); + break; + case 'error': + ui.log('error', event.message); + break; + case 'closed': + ui.log('warn', `bridge closed (${event.code}): ${event.reason}`); + ui.setStatus('disconnected'); + break; + default: + break; + } + } + }); + + window.__vs4dBaseline = client; + + document.getElementById('stop').addEventListener('click', () => { + stopLinkRate(); + client.stop(); + ui.setStatus('stopped'); + document.getElementById('stop').disabled = true; + }); + + setInterval(() => { + const info = client.inspect(); + const s = info.stats; + // Four lines. Frame/resync/drop counters and byte totals were useful + // while bringing the protocol up and are noise now; they remain on + // window.__vs4dBaseline. Faults appear only when non-zero, so a clean + // run stays quiet and a broken one still says so. + const faults = [ + s.decodeFailures ? `${s.decodeFailures} decode` : null, + s.resyncs ? `${s.resyncs} resync` : null, + s.framesDropped ? `${s.framesDropped} dropped` : null + ].filter(Boolean); + ui.setStats([ + `system ${info.mode ?? '-'}`, + `rate ${s.measuredFps.toFixed(1)} fps`, + `goodput ${s.goodputMbps ? s.goodputMbps.toFixed(1) + ' Mbps' : '-'}`, + `points ${info.renderer.objects + .reduce((sum, o) => sum + o.points, 0).toLocaleString()}`, + ...(faults.length ? [`FAULTS ${faults.join(', ')}`] : []) + ]); + }, 500); + + try { + await client.start(); + } catch (err) { + ui.setStatus(`could not connect: ${err.message}`); + ui.log('error', 'Is the bridge running? ' + + 'node system/WebClient/bridge/v4ds-bridge.js --baseline-port '); + } +} + +main().catch(err => { + document.getElementById('status').textContent = `fatal: ${err.message}`; + console.error(err); +}); diff --git a/open4d/webclients/system/WebClient/src/browser-platform.js b/open4d/webclients/system/WebClient/src/browser-platform.js new file mode 100644 index 00000000..b88a4986 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/browser-platform.js @@ -0,0 +1,494 @@ +'use strict'; + +/** + * Browser implementation of the ClientPlatform contract + * (../../ClientCore/platform.js). + * + * Six of the seven capabilities live here; the renderer is ./webgl-renderer.js + * because it is much larger and needs a GPU. All streaming logic is in + * ClientCore and is shared verbatim with the Node desktop client. + * + * Browser globals are reached through an injected `env` rather than captured at + * module scope, so this file can be exercised in Node against stubs — which is + * how tests/test_browser_platform.js verifies the cache policy, the OPFS paths + * and the non-blocking telemetry buffer without a browser. + */ + +/** Default environment: the real browser globals. */ +function defaultEnv() { + return { + fetch: typeof fetch === 'function' ? fetch.bind(globalThis) : undefined, + navigator: typeof navigator !== 'undefined' ? navigator : undefined, + document: typeof document !== 'undefined' ? document : undefined, + console: typeof console !== 'undefined' ? console : undefined, + setInterval: globalThis.setInterval?.bind(globalThis), + clearInterval: globalThis.clearInterval?.bind(globalThis), + setTimeout: globalThis.setTimeout?.bind(globalThis), + now: () => Date.now(), + URL: globalThis.URL, + Blob: globalThis.Blob + }; +} + +// -------------------------------------------------------------------------- +// Asset stores +// -------------------------------------------------------------------------- + +/** + * Downloaded media held as compressed bytes in memory. + * + * This is the default, and it is the right default: what blows up a browser's + * memory is DECODED frames (a 1920-wide texture frame is 5.5 MB), not the + * compressed segment, which is tens of MB for a whole scene. The decode cache + * is the renderer's problem; this only has to hold what arrived off the wire + * until the renderer has consumed it. + */ +class MemoryAssetStore { + constructor() { + this.buffers = new Map(); // handle -> ArrayBuffer + } + + async put(handle, buffer) { this.buffers.set(handle, buffer); } + + /** The renderer reads bytes back out by handle. */ + get(handle) { return this.buffers.get(handle) || null; } + + has(handle) { return this.buffers.has(handle); } + + /** Drop everything under a handle prefix (an object-segment directory). */ + async releasePrefix(prefix) { + for (const key of [...this.buffers.keys()]) { + if (key === prefix || key.startsWith(`${prefix}/`)) { + this.buffers.delete(key); + } + } + } + + get byteLength() { + let total = 0; + for (const buffer of this.buffers.values()) total += buffer.byteLength || 0; + return total; + } +} + +/** + * Downloaded media written to the Origin Private File System. + * + * Slower than memory but survives a reload and does not compete with the + * decoder for heap. Note that `createSyncAccessHandle` is worker-only; this uses + * the async writable-stream API so it works on the main thread. + */ +class OpfsAssetStore { + constructor(rootDirectory) { + this.root = rootDirectory; + } + + static async create(env, name) { + const opfsRoot = await env.navigator.storage.getDirectory(); + const directory = await opfsRoot.getDirectoryHandle(name, { create: true }); + return new OpfsAssetStore(directory); + } + + async _dirFor(parts, { create }) { + let directory = this.root; + for (const part of parts) { + directory = await directory.getDirectoryHandle(part, { create }); + } + return directory; + } + + async put(handle, buffer) { + const parts = handle.split('/').filter(Boolean); + const filename = parts.pop(); + const directory = await this._dirFor(parts, { create: true }); + const fileHandle = await directory.getFileHandle(filename, { create: true }); + const writable = await fileHandle.createWritable(); + await writable.write(buffer); + await writable.close(); + } + + async get(handle) { + const parts = handle.split('/').filter(Boolean); + const filename = parts.pop(); + try { + const directory = await this._dirFor(parts, { create: false }); + const fileHandle = await directory.getFileHandle(filename); + return await (await fileHandle.getFile()).arrayBuffer(); + } catch (_) { + return null; + } + } + + async releasePrefix(prefix) { + const parts = prefix.split('/').filter(Boolean); + const name = parts.pop(); + try { + const directory = await this._dirFor(parts, { create: false }); + await directory.removeEntry(name, { recursive: true }); + } catch (_) { + // Contract: release must tolerate a handle that was never created. + } + } +} + +// -------------------------------------------------------------------------- +// Transport +// -------------------------------------------------------------------------- + +function createTransport({ serverUrl, assetStore, env }) { + async function call(method, apiPath, body) { + const res = await env.fetch(`${serverUrl}${apiPath}`, { + method, + // API responses must never come from the HTTP cache: the manifest + // changes every segment and a cached menu would silently pin the + // client to a stale ladder. + cache: 'no-store', + ...(body === undefined || body === null ? {} : { + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(body) + }) + }); + let parsed = null; + try { + parsed = await res.json(); + } catch (_) { + // Not every endpoint answers with JSON; telemetry posts do not. + } + return { ok: res.ok, status: res.status, body: parsed }; + } + + return { + getJson: apiPath => call('GET', apiPath, null), + postJson: (apiPath, body) => call('POST', apiPath, body), + + assetUrl(assetPath) { + if (!assetPath) return null; + if (/^https?:\/\//.test(assetPath)) return assetPath; + if (assetPath.startsWith('/files/')) return `${serverUrl}${assetPath}`; + const match = assetPath.match(/files\/(.+)/); + return match ? `${serverUrl}/files/${match[1]}` : null; + }, + + /** + * Fetch one media file. Never rejects: a failed asset is normal and is + * judged by download-plan's 90% rules. + * + * `cache: 'no-store'` is REQUIRED, not a nicety. The segment must keep + * generating real network load or the shaped-bandwidth experiment stops + * meaning anything — the same reason system/Client/decode_cache.py + * caches decodes but deliberately never caches downloads. A browser + * quietly serving a segment from its HTTP cache turns a bandwidth + * measurement into fiction. + */ + async fetchAsset(url, handle) { + const start = env.now(); + try { + const res = await env.fetch(url, { cache: 'no-store' }); + if (!res.ok) { + return { + success: false, size: 0, timeMs: 0, error: `HTTP ${res.status}` + }; + } + const buffer = await res.arrayBuffer(); + // A null handle means "count the bytes, keep nothing". + if (handle) await assetStore.put(handle, buffer); + return { + success: true, + size: buffer.byteLength, + timeMs: env.now() - start + }; + } catch (err) { + return { + success: false, size: 0, + timeMs: env.now() - start, error: err.message + }; + } + } + }; +} + +// -------------------------------------------------------------------------- +// Storage +// -------------------------------------------------------------------------- + +/** + * Buffered line sink. + * + * `write` MUST NOT block: render telemetry arrives at 30 Hz, and in Node a + * synchronous write per frame stalled the event loop badly enough to starve the + * renderer. Lines accumulate in an array and are flushed on an interval, so the + * hot path is one array push. + */ +class BufferedLineSink { + constructor({ name, flush, env, flushIntervalMs = 2000 }) { + this.name = name; + this.lines = []; + this._pending = []; + this._flush = flush; + this._env = env; + this._closed = false; + this._timer = flush + ? env.setInterval(() => { this._drain(); }, flushIntervalMs) + : null; + } + + write(line) { + if (this._closed) throw new Error(`write after close on ${this.name}`); + this.lines.push(line); + this._pending.push(line); + } + + _drain() { + if (this._pending.length === 0) return Promise.resolve(); + const batch = this._pending.splice(0, this._pending.length); + // The try/catch is load-bearing: a SYNCHRONOUS throw from the flush + // (an OPFS quota error, say) escapes before Promise.resolve can wrap + // it, so a trailing .catch alone would let it reject close() — which + // finish() awaits, taking down the metrics upload with it. Telemetry + // must never take the run down. + try { + return Promise.resolve(this._flush(batch)).catch(() => {}); + } catch (_) { + return Promise.resolve(); + } + } + + async close() { + this._closed = true; + if (this._timer) this._env.clearInterval(this._timer); + await this._drain(); + } + + text() { return this.lines.join('\n') + (this.lines.length ? '\n' : ''); } +} + +/** + * @param {object} args + * @param {MemoryAssetStore|OpfsAssetStore} args.assetStore + * @param {object} args.env + * @param {(filename: string, text: string) => void} [args.onArtifact] + * Called with the run's result and telemetry so the page can offer them as + * downloads. The browser has nowhere to "write a file" unprompted. + */ +function createStorage({ assetStore, env, onArtifact = null }) { + const artifacts = new Map(); // filename -> text + const sinks = new Map(); // name -> BufferedLineSink + let scratchCount = 0; + + return { + artifacts, + sinks, + + handle: (...parts) => parts.filter(p => p != null).join('/'), + + async createScratch() { + return `run-${++scratchCount}`; + }, + + async release(handle) { + await assetStore.releasePrefix(handle); + }, + + async writeText(name, text) { + artifacts.set(name, text); + onArtifact?.(name, text); + }, + + async writeResult(text) { + artifacts.set('metrics.json', text); + onArtifact?.('metrics.json', text); + }, + + async openAppendStream(name) { + const sink = new BufferedLineSink({ + name, + env, + // Keep the accumulated text available as a downloadable + // artifact; flushing to OPFS as well would double-store it for + // no benefit at these sizes. + flush: () => { artifacts.set(name, sink.text()); } + }); + sinks.set(name, sink); + return sink; + } + }; +} + +// -------------------------------------------------------------------------- +// Viewpoints +// -------------------------------------------------------------------------- + +/** + * Where the initial camera pose comes from. + * + * `initialPose` short-circuits the fetch, which is what an interactive page + * does: the user's camera is the pose, and the canned list only matters in + * simulated mode. + * + * Rejects when nothing is available, per the contract — a run with no initial + * pose would solve the first ladder against a default camera and silently + * invalidate the viewpoint-aware comparison. + */ +function createViewpoints({ serverUrl, indexPath, initialPose, env }) { + return { + async list() { + if (initialPose) { + return [{ filename: 'browser-initial-pose', data: initialPose }]; + } + if (!indexPath) { + throw new Error( + 'no viewpoints available: pass initialPose or viewpointIndexPath'); + } + const res = await env.fetch(`${serverUrl}${indexPath}`, + { cache: 'no-store' }); + if (!res.ok) { + throw new Error(`viewpoint index fetch failed: HTTP ${res.status}`); + } + const body = await res.json(); + const list = Array.isArray(body) ? body : body?.viewpoints; + if (!Array.isArray(list) || list.length === 0) { + throw new Error('viewpoint index contained no poses'); + } + return list.map((entry, index) => ({ + filename: entry.filename || `view_${String(index).padStart(2, '0')}.json`, + data: entry.data || entry + })); + } + }; +} + +// -------------------------------------------------------------------------- +// Clock, logger, lifecycle +// -------------------------------------------------------------------------- + +function createClock(env) { + return { + now: () => env.now(), + every: (ms, fn) => env.setInterval(fn, ms), + cancel: handle => env.clearInterval(handle), + delay: ms => new Promise(resolve => env.setTimeout(resolve, ms)) + }; +} + +/** + * @param {object} args + * @param {(entry: {level: string, line: string, data: object|null}) => void} [args.onLine] + * Page sink, e.g. an on-screen log pane. + * @param {number} [args.keep] ring-buffer size held for inspection + */ +function createLogger({ env, onLine = null, keep = 500 }) { + const lines = []; + return { + lines, + emit(level, line, data) { + lines.push({ level, line, data }); + if (lines.length > keep) lines.shift(); + if (data) env.console?.log(line, data); + else env.console?.log(line); + onLine?.({ level, line, data }); + } + }; +} + +/** + * Page lifecycle. + * + * `exit` MUST NOT navigate away. The final POST /api/results happens during + * shutdown, and unloading the page cancels it — the run's metrics would be lost + * exactly when they matter. So this resolves a promise the page can await and + * leaves the document alone. + */ +function createLifecycle({ env } = { env: defaultEnv() }) { + let resolveDone; + const done = new Promise(resolve => { resolveDone = resolve; }); + const handlers = []; + + return { + /** Resolves with the exit code once the run has finalized. */ + done, + exitCode: null, + + exit(code) { + this.exitCode = code; + resolveDone(code); + }, + + onShutdownRequest(fn) { + handlers.push(fn); + }, + + /** Wire a Stop control / pagehide to the registered finalizers. */ + requestShutdown() { + for (const fn of handlers) fn(); + }, + + get handlerCount() { return handlers.length; } + }; +} + +// -------------------------------------------------------------------------- +// Assembly +// -------------------------------------------------------------------------- + +/** + * Build the browser platform. + * + * @param {object} args + * @param {string} args.serverUrl + * @param {object|null} [args.renderer] a RendererAdapter, or null for simulated mode + * @param {object} [args.initialPose] + * @param {string} [args.viewpointIndexPath] + * @param {'memory'|'opfs'} [args.storageMode='memory'] + * @param {Function} [args.onArtifact] + * @param {Function} [args.onLogLine] + * @param {object} [args.env] injected globals, for tests + * @returns {Promise} the platform, plus `assetStore` and `lifecycle` + */ +async function createBrowserPlatform({ + serverUrl, + renderer = null, + initialPose = null, + viewpointIndexPath = null, + storageMode = 'memory', + onArtifact = null, + onLogLine = null, + env = defaultEnv() +}) { + if (!serverUrl) throw new Error('serverUrl is required'); + if (typeof env.fetch !== 'function') { + throw new Error('this environment has no fetch()'); + } + + const assetStore = storageMode === 'opfs' + ? await OpfsAssetStore.create(env, 'vs4d-client') + : new MemoryAssetStore(); + + return { + transport: createTransport({ serverUrl, assetStore, env }), + storage: createStorage({ assetStore, env, onArtifact }), + viewpoints: createViewpoints({ + serverUrl, indexPath: viewpointIndexPath, initialPose, env + }), + clock: createClock(env), + logger: createLogger({ env, onLine: onLogLine }), + renderer, + lifecycle: createLifecycle({ env }), + // Not part of the contract; the renderer needs to read downloaded bytes + // back out by handle. + assetStore + }; +} + +module.exports = { + createBrowserPlatform, + createTransport, + createStorage, + createViewpoints, + createClock, + createLogger, + createLifecycle, + MemoryAssetStore, + OpfsAssetStore, + BufferedLineSink, + defaultEnv +}; diff --git a/open4d/webclients/system/WebClient/src/camera-pose.js b/open4d/webclients/system/WebClient/src/camera-pose.js new file mode 100644 index 00000000..fb1b0689 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/camera-pose.js @@ -0,0 +1,169 @@ +'use strict'; + +/** + * Camera pose conversion: WebGL/Three.js camera -> Open3D + * `PinholeCameraParameters`. + * + * This is not a convenience. What the client POSTs to `/api/viewpoint` is + * written to disk verbatim and read back by + * `o3d.io.read_pinhole_camera_parameters` (see + * `vstream/ladder/ladder_service.py` and `create_ladder.py`), so the JSON must + * be exactly that schema or the ladder cannot parse the pose at all. When it + * cannot, `normalize_weights` necessarily returns equal weights and the ladder + * silently stops being viewpoint-aware — which is the entire contribution being + * measured. A wrong pose here does not crash anything; it quietly invalidates + * the experiment. + * + * Two conventions have to be bridged: + * + * 1. AXES. Three.js cameras look down -Z with +Y up. Open3D's camera frame is + * OpenCV-style: +X right, +Y DOWN, +Z FORWARD. So the world->camera matrix + * gets its Y and Z rows negated. The captured viewpoint files show this + * directly — their extrinsic diagonal is roughly (+1, -1, -1). + * + * 2. UNITS. `vstream/config.py` is explicit: "Baked OBJ/Draco coordinates are + * metres, while Open3D camera extrinsics and the quality model's + * mean-distance feature are millimetres", and + * VIEW_RAYCAST_UNITS_PER_METER defaults to 1000. The captured files agree — + * their translations are in the thousands. A browser scene in metres must + * therefore scale its translation by 1000, or the server raycasts from + * ~4 mm away, every ray misses, and the weights come back uniform. + * + * Only the translation scales. The extrinsic maps world points to camera + * space as `p_cam = R*p_world + t` with `t = -R*C`, so re-expressing the + * same pose in a millimetre world gives `t_mm = 1000 * t_m` and leaves R + * untouched. + */ + +/** Open3D stores the principal point at the pixel-grid centre. */ +function principalPoint(size) { + return (size - 1) / 2; +} + +/** + * Vertical focal length in pixels for a vertical field of view. + * + * Cross-check against the captured corpus: a 60 degree vertical FOV at 1920 px + * gives 960 / tan(30 deg) = 1662.7687752661222, which is exactly the value in + * `system/Client/viewpoints/view_00.json`. + */ +function focalLengthPx(fovDegrees, height) { + const fovRadians = (fovDegrees * Math.PI) / 180; + return (height / 2) / Math.tan(fovRadians / 2); +} + +/** + * Build an Open3D `PinholeCameraParameters` object. + * + * @param {object} args + * @param {number[]} args.viewMatrix world->camera matrix, COLUMN-major, 16 + * elements. In Three.js this is `camera.matrixWorldInverse.elements`, which is + * already column-major. + * @param {number} args.fovDegrees vertical field of view + * @param {number} args.width viewport width in pixels + * @param {number} args.height viewport height in pixels + * @param {number} [args.unitsPerMeter=1000] world units per metre on the + * SERVER side; must match config.VIEW_RAYCAST_UNITS_PER_METER + * @returns {object} JSON-ready PinholeCameraParameters + */ +function toOpen3DCameraParameters({ + viewMatrix, fovDegrees, width, height, unitsPerMeter = 1000 +}) { + if (!Array.isArray(viewMatrix) && !(viewMatrix instanceof Float32Array) + && !(viewMatrix instanceof Float64Array)) { + throw new TypeError('viewMatrix must be an array of 16 numbers'); + } + if (viewMatrix.length !== 16) { + throw new RangeError(`viewMatrix must have 16 elements, got ${viewMatrix.length}`); + } + if (!(width > 0) || !(height > 0)) { + throw new RangeError('width and height must be positive'); + } + if (!(fovDegrees > 0) || fovDegrees >= 180) { + throw new RangeError(`fovDegrees out of range: ${fovDegrees}`); + } + + // Column-major indexing: element(row, col) = viewMatrix[col * 4 + row]. + // Negating rows 1 and 2 applies diag(1, -1, -1, 1) on the left, which is + // the Y-up/-Z-forward -> Y-down/+Z-forward change of basis. + const extrinsic = new Array(16); + for (let col = 0; col < 4; col++) { + for (let row = 0; row < 4; row++) { + const index = col * 4 + row; + const flip = (row === 1 || row === 2) ? -1 : 1; + // `+ 0` normalizes -0 away. Negating a zero row entry yields -0, + // which is numerically harmless but does not survive a JSON round + // trip, so it would make the emitted pose fail an equality check + // against itself. + extrinsic[index] = flip * viewMatrix[index] + 0; + } + } + // Translation is the last column; scale it into the server's world units. + extrinsic[12] *= unitsPerMeter; + extrinsic[13] *= unitsPerMeter; + extrinsic[14] *= unitsPerMeter; + + const focal = focalLengthPx(fovDegrees, height); + + return { + class_name: 'PinholeCameraParameters', + extrinsic, + intrinsic: { + height, + // Column-major 3x3: [fx, 0, 0, 0, fy, 0, cx, cy, 1]. + // Square pixels: Three.js takes a vertical FOV and derives the + // horizontal extent from the aspect ratio, so fx == fy. + intrinsic_matrix: [ + focal, 0, 0, + 0, focal, 0, + principalPoint(width), principalPoint(height), 1 + ], + width + }, + version_major: 1, + version_minor: 0 + }; +} + +/** + * Convenience wrapper for a Three.js PerspectiveCamera. + * + * `updateMatrixWorld` then `matrixWorldInverse` rather than trusting whatever + * the render loop last computed: the pose is read on the segment tick, which is + * not synchronised with a frame. + */ +function fromThreeCamera(camera, { width, height, unitsPerMeter = 1000 }) { + camera.updateMatrixWorld(); + camera.updateProjectionMatrix(); + return toOpen3DCameraParameters({ + viewMatrix: Array.from(camera.matrixWorldInverse.elements), + fovDegrees: camera.fov, + width, + height, + unitsPerMeter + }); +} + +/** + * Map a world point into the camera frame described by these parameters. + * Used by the tests to assert the axis convention: a point the camera is + * looking at must land at POSITIVE z. + */ +function projectToCameraSpace(parameters, worldPointMetres) { + const e = parameters.extrinsic; + const scale = 1000; // parameters are in the server's units + const [x, y, z] = worldPointMetres.map(v => v * scale); + return [ + e[0] * x + e[4] * y + e[8] * z + e[12], + e[1] * x + e[5] * y + e[9] * z + e[13], + e[2] * x + e[6] * y + e[10] * z + e[14] + ]; +} + +module.exports = { + toOpen3DCameraParameters, + fromThreeCamera, + focalLengthPx, + principalPoint, + projectToCameraSpace +}; diff --git a/open4d/webclients/system/WebClient/src/chooser.js b/open4d/webclients/system/WebClient/src/chooser.js new file mode 100644 index 00000000..505ae0ac --- /dev/null +++ b/open4d/webclients/system/WebClient/src/chooser.js @@ -0,0 +1,227 @@ +'use strict'; + +/** + * The launcher: one row per system, so the list stays readable as methods are + * added. + * + * Each row is a link. Selecting objects is behind a per-row toggle rather than + * inline, because the object list is the one thing here that does not scale — + * nine rows of checkboxes under every method would bury the list it belongs to. + * + * Why an object picker exists at all: the ladder must publish at least one + * representation per object in the scene, so the scene sets an irreducible + * bitrate floor. All nine ORBIT objects floor at ~116 Mbps, and under that + * floor the MCKP can only buy the few highest-weighted objects at their + * cheapest rung and freezes the rest — the page then looks permanently starved + * however good the link is. Three objects floor near 21 Mbps. So each picker + * defaults to the cheapest few rather than leaving a viewer to discover the + * deficit, and Vega's defaults to the smallest clips because that page + * preloads whole clips before it plays them. + */ + +// One line per system, carrying the one thing a screenshot cannot show: +// whether it adapts, and where. A fixed-quality player and an adaptive one +// look identical on a fast link. +const DESCRIPTION = { + mesh: 'Textured meshes. Adapts in this browser, one representation per ' + + 'object per segment.', + vivo: 'Point clouds. Adapts server-side, per spatial tile.', + nava: 'Point clouds. Adapts server-side, one quality per object per segment.', + vega: '3D Gaussian splats. Fixed quality, no adaptation.', + nevo: 'Neural volumetric. Pre-rendered comparison panels.' +}; + +function el(tag, props = {}, children = []) { + const node = document.createElement(tag); + for (const [key, value] of Object.entries(props)) { + if (key === 'class') node.className = value; + else if (key === 'text') node.textContent = value; + else if (key.startsWith('on')) node.addEventListener(key.slice(2), value); + else node.setAttribute(key, value); + } + for (const child of [].concat(children)) if (child) node.appendChild(child); + return node; +} + +function mbps(value) { + return Number.isFinite(value) ? `${value.toFixed(1)} Mbps` : '—'; +} + +/** Where a row points, including whatever the page needs to find its data. */ +function launchUrl(system, objects) { + const url = new URL(system.page, window.location.origin); + if (objects && objects.length) { + url.searchParams.set('objects', objects.join(',')); + } + // Each point-cloud baseline has its own bridge, so the address is what + // selects ViVo vs NAVA — the page itself is the same. + if (system.bridge) url.searchParams.set('bridge', system.bridge); + return url.toString(); +} + +/** + * Which objects a system offers, and a sensible default selection. + * `cost` labels the single numeric column: bitrate for a ladder, size for a + * preloading viewer. + */ +function objectChoice(system) { + const objects = [...(system.objects || [])]; + if (!objects.length) return null; + + if (system.id === 'vega') { // sized by download, not by bitrate + return { + objects, + head: 'clip', + cell: o => `${(o.bytes / 1e6).toFixed(0)} MB`, + defaults: [...objects].sort((a, b) => a.bytes - b.bytes) + .slice(0, 2).map(o => o.name) + }; + } + const priced = objects.filter(o => o.floorMbps !== null); + return { + objects: objects.sort((a, b) => (b.weight || 0) - (a.weight || 0) + || a.name.localeCompare(b.name)), + head: 'floor', + cell: o => mbps(o.floorMbps), + defaults: (priced.length ? priced : objects).slice() + .sort((a, b) => (a.floorMbps ?? Infinity) - (b.floorMbps ?? Infinity)) + .slice(0, 3).map(o => o.name) + }; +} + +/** + * Ask the server to (re)start a point-cloud baseline with this object set. + * + * Needed because these baselines fix their scene with `--objects` at startup, + * so unlike the others the selection cannot be a query parameter — it is a new + * process. The server validates every name against the tile catalogue and + * waits until both the bridge and the baseline are listening before replying, + * so a resolved promise means the page can actually connect. + */ +async function restartBaseline(system, objects) { + const response = await fetch('/api/pointcloud/start', { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + cache: 'no-store', + body: JSON.stringify({ id: system.id, objects }) + }); + const body = await response.json().catch(() => ({})); + if (!response.ok) throw new Error(body.error || `HTTP ${response.status}`); + return body; +} + +/** One row: name, description, and an optional collapsed object picker. */ +function renderRow(system) { + const choice = objectChoice(system); + const chosen = new Set(choice ? choice.defaults : []); + const status = document.getElementById('status'); + + const link = el('a', { + class: 'name', + text: system.name, + // `detail` explains a system that cannot run. It is not on the row — + // that is what keeps the list scannable — but it is one hover away, + // and /api/systems still reports `ready`. + title: system.detail || '' + }); + const sync = () => { link.href = launchUrl(system, [...chosen]); }; + sync(); + + if (system.restartable) { + // Starting the process takes ~15 s, so navigating first would land on + // a page that cannot connect yet. Hold the click, start it, then go. + link.addEventListener('click', async (event) => { + if (event.metaKey || event.ctrlKey || event.button !== 0) return; + event.preventDefault(); + const objects = [...chosen]; + if (!objects.length) { + status.textContent = `${system.name}: choose at least one object`; + status.className = 'bad'; + return; + } + link.classList.add('busy'); + status.className = ''; + status.textContent = `starting ${system.name} with ` + + `${objects.join(', ')} …`; + try { + const started = await restartBaseline(system, objects); + window.location.href = launchUrl( + { ...system, bridge: started.bridge }, []); + } catch (error) { + status.textContent = `${system.name}: ${error.message}`; + status.className = 'bad'; + link.classList.remove('busy'); + } + }); + } + + const row = el('li', { class: 'row' }, [ + link, + el('span', { class: 'desc', text: DESCRIPTION[system.id] || '' }) + ]); + if (!choice) return row; + + const picker = el('div', { class: 'picker', hidden: 'hidden' }); + const toggle = el('button', { + class: 'toggle', + text: `${chosen.size} of ${choice.objects.length}`, + onclick: () => { + const open = picker.hasAttribute('hidden'); + if (open) picker.removeAttribute('hidden'); + else picker.setAttribute('hidden', 'hidden'); + toggle.classList.toggle('open', open); + } + }); + row.appendChild(toggle); + row.appendChild(picker); + + const table = el('table', {}, [el('tbody')]); + const body = table.querySelector('tbody'); + for (const object of choice.objects) { + const box = el('input', { type: 'checkbox' }); + box.checked = chosen.has(object.name); + const tr = el('tr', {}, [ + el('td', {}, box), + el('td', { text: object.name }), + el('td', { class: 'num', text: choice.cell(object) }) + ]); + box.addEventListener('change', () => { + if (box.checked) chosen.add(object.name); + else chosen.delete(object.name); + tr.classList.toggle('off', !box.checked); + toggle.textContent = `${chosen.size} of ${choice.objects.length}`; + sync(); + }); + tr.classList.toggle('off', !box.checked); + body.appendChild(tr); + } + picker.appendChild(table); + return row; +} + +async function main() { + const root = document.getElementById('systems'); + const status = document.getElementById('status'); + let payload; + try { + const response = await fetch('/api/systems', { cache: 'no-store' }); + if (!response.ok) throw new Error(`HTTP ${response.status}`); + payload = await response.json(); + } catch (error) { + status.textContent = `could not reach /api/systems: ${error.message}`; + status.className = 'bad'; + return; + } + status.textContent = ''; + + const list = el('ul', { class: 'systems' }); + for (const system of payload.systems || []) list.appendChild(renderRow(system)); + root.appendChild(list); +} + +main().catch(error => { + const status = document.getElementById('status'); + if (status) status.textContent = `fatal: ${error.message}`; + // eslint-disable-next-line no-console + console.error(error); +}); diff --git a/open4d/webclients/system/WebClient/src/decode-cache.js b/open4d/webclients/system/WebClient/src/decode-cache.js new file mode 100644 index 00000000..b67dac8f --- /dev/null +++ b/open4d/webclients/system/WebClient/src/decode-cache.js @@ -0,0 +1,166 @@ +'use strict'; + +/** + * Decoded-clip cache for the browser renderer. + * + * The browser analogue of system/Client/decode_cache.py, and it exists for the + * same measured reason. The desktop client originally re-ran the Draco decoder + * over 60 frames plus a texture decode for EVERY object of EVERY segment — an + * order of magnitude more work than a 2 s segment budget allows. Objects then + * never became playable and were reported missing. The browser has strictly + * less decode headroom, so this is built in from the start rather than added + * after the first starved run. + * + * `(objectName, repId)` is a COMPLETE cache key. Media is a fixed + * FRAMES_PER_SEGMENT-frame loop published once under + * `files/media///` and shared by every logical segment + * (ladderlib.build_mpd writes `source_segment: 0, loop: true`), so a clip + * decoded during segment 3 is still correct in segment 88. + * + * What is deliberately NOT cached: the downloads. The segment must keep + * generating real network load or the shaped-bandwidth experiment stops meaning + * anything. Only the decode is amortized. + * + * Eviction is by byte budget, least-recently-used first, because decoded frames + * are what actually exhaust a browser tab: a 1920-wide texture frame is + * w*h*1.5 = 5.5 MB however it is stored, so one 60-frame clip for five objects + * is 1.66 GB. `vstream/config.py` caps published texture width for the same + * reason. + */ + +const DEFAULT_BUDGET_BYTES = 512 * 1024 * 1024; + +/** `(objectName, repId)` -> cache key. */ +function clipKey(objectName, repId) { + return `${objectName}::${repId}`; +} + +class DecodeCache { + /** + * @param {object} [options] + * @param {number} [options.budgetBytes] evict once decoded bytes exceed this + * @param {(clip: object) => void} [options.onEvict] release GPU resources + * @param {() => number} [options.now] clock, for LRU ordering and tests + */ + constructor({ + budgetBytes = DEFAULT_BUDGET_BYTES, + onEvict = null, + now = () => Date.now() + } = {}) { + this.budgetBytes = budgetBytes; + this._onEvict = onEvict; + this._now = now; + /** key -> { clip, bytes, lastUsed, pinned } */ + this._entries = new Map(); + /** key -> Promise, so concurrent requests share one decode. */ + this._inFlight = new Map(); + this.stats = { hits: 0, misses: 0, shared: 0, evictions: 0 }; + } + + get bytes() { + let total = 0; + for (const entry of this._entries.values()) total += entry.bytes; + return total; + } + + get size() { return this._entries.size; } + + has(objectName, repId) { return this._entries.has(clipKey(objectName, repId)); } + + /** + * Fetch a decoded clip, decoding it at most once. + * + * Concurrent callers for the same key await the SAME decode rather than + * starting a second one. Without this, two segments selecting the same + * representation would each decode 60 frames — the exact duplication that + * `decodeShared` reports in the desktop client's telemetry. + * + * @param {string} objectName + * @param {string} repId + * @param {() => Promise<{clip: object, bytes: number}>} decode + * @returns {Promise<{clip: object, cacheHit: boolean, decodeShared: boolean}>} + */ + async get(objectName, repId, decode) { + const key = clipKey(objectName, repId); + + const entry = this._entries.get(key); + if (entry) { + entry.lastUsed = this._now(); + this.stats.hits++; + return { clip: entry.clip, cacheHit: true, decodeShared: false }; + } + + const pending = this._inFlight.get(key); + if (pending) { + this.stats.shared++; + const clip = await pending; + return { clip, cacheHit: false, decodeShared: true }; + } + + this.stats.misses++; + const work = (async () => { + const { clip, bytes } = await decode(); + this._entries.set(key, { + clip, bytes, lastUsed: this._now(), pinned: false + }); + this._evictToBudget(); + return clip; + })(); + this._inFlight.set(key, work); + try { + const clip = await work; + return { clip, cacheHit: false, decodeShared: false }; + } finally { + this._inFlight.delete(key); + } + } + + /** + * Protect a clip from eviction while it is on screen. + * + * Without pinning, a large scene can evict the clip the renderer is in the + * middle of presenting, which shows up as a frame reverting mid-playback + * rather than as an error. + */ + pin(objectName, repId) { + const entry = this._entries.get(clipKey(objectName, repId)); + if (entry) entry.pinned = true; + } + + unpin(objectName, repId) { + const entry = this._entries.get(clipKey(objectName, repId)); + if (entry) entry.pinned = false; + } + + /** Unpin every clip of an object except the one now in use. */ + pinOnly(objectName, repId) { + for (const [key, entry] of this._entries) { + if (key.startsWith(`${objectName}::`)) { + entry.pinned = (key === clipKey(objectName, repId)); + } + } + } + + _evictToBudget() { + if (this.bytes <= this.budgetBytes) return; + // Least-recently-used first, pinned clips last-resort only. + const candidates = [...this._entries.entries()] + .filter(([, entry]) => !entry.pinned) + .sort((a, b) => a[1].lastUsed - b[1].lastUsed); + + for (const [key, entry] of candidates) { + if (this.bytes <= this.budgetBytes) break; + this._entries.delete(key); + this.stats.evictions++; + this._onEvict?.(entry.clip); + } + } + + /** Drop everything, releasing GPU resources. */ + clear() { + for (const entry of this._entries.values()) this._onEvict?.(entry.clip); + this._entries.clear(); + } +} + +module.exports = { DecodeCache, clipKey, DEFAULT_BUDGET_BYTES }; diff --git a/open4d/webclients/system/WebClient/src/draco-worker.js b/open4d/webclients/system/WebClient/src/draco-worker.js new file mode 100644 index 00000000..04e0ade2 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/draco-worker.js @@ -0,0 +1,191 @@ +'use strict'; + +/** + * Draco decode worker. + * + * A segment is up to 60 Draco meshes per object. Decoding them on the main + * thread would compete with the render loop for exactly the window in which + * frames must be presented, so every decode happens here and only typed arrays + * cross back — transferred, not copied. + * + * Loads the official Draco JS/WASM decoder vendored under ../vendor/draco. + * That build is upstream's, unmodified. + * + * Handles BOTH Draco geometry kinds, because the two pipelines differ: + * + * TRIANGULAR_MESH the mesh ladder (`compress_geometry.py` -> textured OBJs) + * POINT_CLOUD every V4DS baseline (`draco_encoder -point_cloud -qp 11 + * -qg 8`), carrying POSITION plus 8-bit COLOR + * + * The kind is read from the payload rather than configured, so one worker + * serves both the mesh client and the baseline client. + * + * Protocol: + * in { id, buffers: ArrayBuffer[] } frames in order + * out { id, frames: [frame|null] } | { id, error } + * mesh frame = { kind: 'mesh', positions, normals, uvs, indices } + * cloud frame = { kind: 'cloud', positions, colors } + * + * A frame that fails to decode comes back as null in the array rather than + * failing the whole clip: download-plan already tolerates gaps, and the + * renderer holds the previous frame for one. + */ + +let decoderModulePromise = null; + +function loadDecoder(vendorBase) { + if (decoderModulePromise) return decoderModulePromise; + decoderModulePromise = new Promise((resolve, reject) => { + try { + self.importScripts(`${vendorBase}/draco_decoder.js`); + } catch (err) { + reject(new Error(`could not load draco_decoder.js: ${err.message}`)); + return; + } + // The emscripten build resolves its .wasm through locateFile. + self.DracoDecoderModule({ + locateFile: file => `${vendorBase}/${file}` + }).then(resolve, reject); + }); + return decoderModulePromise; +} + +/** + * Decode one Draco buffer into plain typed arrays. + * + * Every Draco object must be explicitly destroyed: the WASM heap is not + * garbage-collected from JS, and leaking one mesh per frame at 30 fps exhausts + * it within a minute. + */ +function decodeGeometry(draco, decoder, buffer) { + const dracoBuffer = new draco.DecoderBuffer(); + dracoBuffer.Init(new Int8Array(buffer), buffer.byteLength); + + let geometry = null; + try { + const geometryType = decoder.GetEncodedGeometryType(dracoBuffer); + + if (geometryType === draco.TRIANGULAR_MESH) { + geometry = new draco.Mesh(); + const status = decoder.DecodeBufferToMesh(dracoBuffer, geometry); + if (!status.ok() || geometry.ptr === 0) { + throw new Error(status.error_msg() || 'mesh decode failed'); + } + return { + kind: 'mesh', + positions: readFloatAttribute( + draco, decoder, geometry, draco.POSITION, 3), + normals: readFloatAttribute( + draco, decoder, geometry, draco.NORMAL, 3), + uvs: readFloatAttribute( + draco, decoder, geometry, draco.TEX_COORD, 2), + indices: readIndices(draco, decoder, geometry) + }; + } + + if (geometryType === draco.POINT_CLOUD) { + geometry = new draco.PointCloud(); + const status = decoder.DecodeBufferToPointCloud(dracoBuffer, geometry); + if (!status.ok() || geometry.ptr === 0) { + throw new Error(status.error_msg() || 'point cloud decode failed'); + } + return { + kind: 'cloud', + positions: readFloatAttribute( + draco, decoder, geometry, draco.POSITION, 3), + // Colours are quantised to 8 bits by the encoder's -qg 8, so + // read them as bytes. Reading them as floats yields 0..255 + // values that then have to be rediscovered downstream. + colors: readUint8Attribute( + draco, decoder, geometry, draco.COLOR, 3) + }; + } + + throw new Error(`unsupported draco geometry type ${geometryType}`); + } finally { + if (geometry) draco.destroy(geometry); + draco.destroy(dracoBuffer); + } +} + +function readFloatAttribute(draco, decoder, geometry, attributeType, components) { + const id = decoder.GetAttributeId(geometry, attributeType); + if (id < 0) return null; + const attribute = decoder.GetAttribute(geometry, id); + const count = geometry.num_points(); + const array = new draco.DracoFloat32Array(); + try { + decoder.GetAttributeFloatForAllPoints(geometry, attribute, array); + const out = new Float32Array(count * components); + for (let i = 0; i < out.length; i++) out[i] = array.GetValue(i); + return out; + } finally { + draco.destroy(array); + } +} + +function readUint8Attribute(draco, decoder, geometry, attributeType, components) { + const id = decoder.GetAttributeId(geometry, attributeType); + if (id < 0) return null; + const attribute = decoder.GetAttribute(geometry, id); + const count = geometry.num_points(); + const array = new draco.DracoUInt8Array(); + try { + decoder.GetAttributeUInt8ForAllPoints(geometry, attribute, array); + const out = new Uint8Array(count * components); + for (let i = 0; i < out.length; i++) out[i] = array.GetValue(i); + return out; + } finally { + draco.destroy(array); + } +} + +function readIndices(draco, decoder, mesh) { + const faceCount = mesh.num_faces(); + const out = new Uint32Array(faceCount * 3); + const face = new draco.DracoInt32Array(); + try { + for (let i = 0; i < faceCount; i++) { + decoder.GetFaceFromMesh(mesh, i, face); + out[i * 3] = face.GetValue(0); + out[i * 3 + 1] = face.GetValue(1); + out[i * 3 + 2] = face.GetValue(2); + } + return out; + } finally { + draco.destroy(face); + } +} + +self.onmessage = async event => { + const { id, buffers, vendorBase } = event.data; + try { + const draco = await loadDecoder(vendorBase || '/web/vendor/draco'); + const decoder = new draco.Decoder(); + const frames = []; + const transfer = []; + try { + for (const buffer of buffers) { + if (!buffer) { frames.push(null); continue; } + try { + const frame = decodeGeometry(draco, decoder, buffer); + frames.push(frame); + // Transfer rather than copy: a 60-frame clip is tens of MB. + for (const key of ['positions', 'normals', 'uvs', 'indices', + 'colors']) { + if (frame[key]) transfer.push(frame[key].buffer); + } + } catch (_) { + // One bad frame is a gap the renderer can hold through, not + // a reason to discard the whole object-segment. + frames.push(null); + } + } + } finally { + draco.destroy(decoder); + } + self.postMessage({ id, frames }, transfer); + } catch (err) { + self.postMessage({ id, error: err.message }); + } +}; diff --git a/open4d/webclients/system/WebClient/src/link-rate.js b/open4d/webclients/system/WebClient/src/link-rate.js new file mode 100644 index 00000000..59f3d7b7 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/link-rate.js @@ -0,0 +1,55 @@ +'use strict'; + +/** + * Live readout of the shaped link rate, polled from /api/shaping. + * + * The point of showing it is causality. A representation switch on its own is + * unreadable — it could be the ABR working or the ABR thrashing. Beside the + * rate the kernel is enforcing, the same switch becomes evidence. And when the + * link is unshaped this says so, which matters because an unshaped run looks + * like a broken adaptive system: every system just holds one operating point. + */ +function mountLinkRate(element, { fetchImpl = fetch, intervalMs = 2000 } = {}) { + if (!element) return () => {}; + let stopped = false; + + // Short text on screen, full detail in the tooltip. The distinctions still + // matter -- "cannot tell" must never read as "flat link" -- but they belong + // on hover rather than in the viewer's way. + const render = (state) => { + if (state.error) { + element.textContent = 'link: unknown'; + element.title = `could not read the shaped rate: ${state.error}`; + element.className = 'link unknown'; + return; + } + if (!state.shaped) { + element.textContent = 'link: unshaped'; + element.title = 'no root TBF installed, so nothing to adapt to — ' + + 'replay a trace with scripts/shape_web_demo.sh'; + element.className = 'link unshaped'; + return; + } + element.textContent = state.rateMbps === null + ? 'link: shaped' + : `link: ${state.rateMbps.toFixed(1)} Mbps`; + element.title = `enforced by tc on ${state.interface}`; + element.className = 'link shaped'; + }; + + const tick = async () => { + try { + const response = await fetchImpl('/api/shaping', { cache: 'no-store' }); + if (!response.ok) throw new Error(`HTTP ${response.status}`); + render(await response.json()); + } catch (error) { + render({ error: error.message }); + } + }; + + tick(); + const handle = setInterval(() => { if (!stopped) tick(); }, intervalMs); + return () => { stopped = true; clearInterval(handle); }; +} + +module.exports = { mountLinkRate }; diff --git a/open4d/webclients/system/WebClient/src/main.js b/open4d/webclients/system/WebClient/src/main.js new file mode 100644 index 00000000..5e7b0fb0 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/main.js @@ -0,0 +1,224 @@ +'use strict'; + +/** + * Browser client entry point. + * + * The counterpart of system/Client/client.js: read configuration, assemble the + * platform, hand it to the shared `StreamingClient`. No streaming logic here. + * + * Configuration comes from the query string so a run can be launched from a URL + * without a rebuild, mirroring how the Node client takes environment variables: + * + * /web/?server=http://host:3000&mode=interactive&segments=20 + * + * | parameter | default | meaning | + * |--------------|----------------------|--------------------------------------| + * | server | the page's own origin| server base URL | + * | mode | interactive | interactive \| simulated | + * | storage | memory | memory \| opfs asset store | + * | viewpoints | (none) | path to a viewpoint index JSON | + * | decodeBudget | 512 | decoded-clip cache budget, MB | + * | concurrency | 10 | parallel asset requests | + * | inflight | 2 | max concurrent segment downloads | + * | objects | (server's full scene)| comma-separated object subset | + * + * `?objects=` is the one parameter that changes what the experiment can show. + * The ladder publishes at least one representation per object in the scene, so + * nine ORBIT objects floor at ~116 Mbps; under that the MCKP can only buy the + * few highest-weighted objects at the cheapest rung and freezes the rest, and + * the page looks permanently starved however good the link is. Three objects + * floor near 21 Mbps, which leaves an ordinary link enough headroom to climb + * the ladder — which is the behaviour worth watching. + */ + +const { StreamingClient } = require('../../ClientCore/streaming-client'); +const { mountLinkRate } = require('./link-rate'); +const { createBrowserPlatform } = require('./browser-platform'); +const { WebGLRenderer } = require('./webgl-renderer'); + +function readConfig() { + const params = new URLSearchParams(window.location.search); + const number = (name, fallback) => { + const raw = params.get(name); + const value = Number(raw); + return raw !== null && Number.isFinite(value) ? value : fallback; + }; + return { + serverUrl: (params.get('server') || window.location.origin) + .replace(/\/+$/, ''), + mode: (params.get('mode') || 'interactive').toLowerCase(), + storageMode: (params.get('storage') || 'memory').toLowerCase(), + viewpointIndexPath: params.get('viewpoints'), + decodeBudgetBytes: number('decodeBudget', 512) * 1024 * 1024, + downloadConcurrency: number('concurrency', 10), + maxInflightSegments: number('inflight', 2), + runLabel: params.get('label') || 'web-client', + sceneObjects: (params.get('objects') || '') + .split(',').map(name => name.trim()).filter(Boolean) + }; +} + +/** Minimal page chrome: status line, log pane, artifact downloads. */ +function createUi() { + const status = document.getElementById('status'); + const logPane = document.getElementById('log'); + const artifactList = document.getElementById('artifacts'); + const stopButton = document.getElementById('stop'); + + return { + stopButton, + setStatus(text) { status.textContent = text; }, + appendLog({ level, line }) { + // DEBUG is for the console, not the pane. Lines like + // "mitch superseded" fire several times a segment and bury the + // INFO/WARN lines that actually tell you what the run is doing. + if (level.toUpperCase() === 'DEBUG') return; + const row = document.createElement('div'); + row.className = `log-line log-${level.toLowerCase()}`; + row.textContent = line; + logPane.appendChild(row); + while (logPane.childElementCount > 400) { + logPane.removeChild(logPane.firstChild); + } + logPane.scrollTop = logPane.scrollHeight; + }, + /** + * A browser cannot write a file unprompted, so every artifact the core + * "writes" is surfaced as a download link instead. + */ + offerArtifact(name, text) { + let link = artifactList.querySelector(`[data-name="${name}"]`); + if (!link) { + link = document.createElement('a'); + link.dataset.name = name; + link.textContent = name; + link.download = name; + artifactList.appendChild(link); + } + if (link.href) URL.revokeObjectURL(link.href); + link.href = URL.createObjectURL( + new Blob([text], { type: 'application/json' })); + } + }; +} + +async function main() { + const config = readConfig(); + const ui = createUi(); + // Kept so it can be stopped: a poller that outlives the run keeps hitting + // /api/shaping after Stop, and in a test harness the live interval stops + // the process exiting at all. + const stopLinkRate = mountLinkRate(document.getElementById('link')); + ui.setStatus(`connecting to ${config.serverUrl} …`); + + const interactive = config.mode === 'interactive'; + if (!['interactive', 'simulated'].includes(config.mode)) { + throw new Error( + `?mode must be "interactive" or "simulated", got "${config.mode}"`); + } + // Simulated mode replays canned poses, so it has no camera of its own. There + // is deliberately no default: solving the first ladder against an arbitrary + // pose would silently invalidate the viewpoint-aware comparison, which is + // the whole point of the experiment. + if (!interactive && !config.viewpointIndexPath) { + throw new Error( + 'simulated mode needs ?viewpoints=; ' + + 'interactive mode uses the live camera instead'); + } + let renderer = null; + + // The platform is built first so the renderer can read downloaded bytes + // back out of the asset store by handle. + const platform = await createBrowserPlatform({ + serverUrl: config.serverUrl, + storageMode: config.storageMode, + viewpointIndexPath: config.viewpointIndexPath, + // Interactive mode's pose IS the live camera, so no canned list is + // needed; the renderer's own pose is adopted immediately after start. + initialPose: interactive && !config.viewpointIndexPath + ? { objects: {} } : null, + onArtifact: (name, text) => ui.offerArtifact(name, text), + onLogLine: entry => ui.appendLog(entry) + }); + + if (interactive) { + renderer = new WebGLRenderer({ + canvas: document.getElementById('view'), + readAsset: handle => platform.assetStore.get(handle), + decodeBudgetBytes: config.decodeBudgetBytes + }); + platform.renderer = renderer; + } + + const client = new StreamingClient({ + platform, + config: { + clientMode: interactive ? 'interactive' : 'simulated', + downloadConcurrency: config.downloadConcurrency, + maxInflightSegments: config.maxInflightSegments, + runLabel: config.runLabel, + sceneObjects: config.sceneObjects + } + }); + + // Stop must finalize the run, not unload the page: the final + // POST /api/results happens during shutdown and navigating away cancels it. + ui.stopButton.addEventListener('click', () => { + ui.stopButton.disabled = true; + ui.setStatus('finishing …'); + platform.lifecycle.requestShutdown(); + }); + window.addEventListener('pagehide', () => platform.lifecycle.requestShutdown()); + + // Debug handle. Deliberate and documented: without it the only way to + // inspect a live run is to add logging and rebuild, and the renderer's + // state (what is on screen, where the camera is) is exactly what you need + // when the page looks wrong but every log line looks right. + window.__vs4d = { + client, platform, renderer, + inspect() { + const scene = renderer?._three?.scene; + const camera = renderer?._three?.camera; + return { + objects: [...(renderer?._objects || new Map())].map(([name, e]) => ({ + name, + visible: e.mesh.visible, + vertices: e.mesh.geometry.attributes.position?.count ?? 0, + frames: e.clip?.frameCount ?? 0, + textured: Boolean(e.material.map), + centre: e.mesh.geometry.boundingSphere + ? e.mesh.geometry.boundingSphere.center.toArray() + .map(v => Number(v.toFixed(2))) + : null + })), + playback: renderer?._playback, + camera: camera ? { + position: camera.position.toArray().map(v => Number(v.toFixed(2))), + target: renderer._three.controls.target.toArray() + .map(v => Number(v.toFixed(2))), + fov: camera.fov, near: camera.near, far: camera.far + } : null, + sceneChildren: scene?.children?.length ?? 0, + stats: renderer?.stats + }; + } + }; + + await client.run(); + ui.setStatus(`streaming · broadcast ${client.broadcastId ?? '(none)'}`); + + const code = await platform.lifecycle.done; + ui.setStatus(code === 0 + ? 'run complete — metrics.json is ready below' + : `run failed (exit ${code}) — see the log`); + ui.stopButton.disabled = true; + stopLinkRate(); + if (renderer) await renderer.stop(); +} + +main().catch(err => { + const status = document.getElementById('status'); + if (status) status.textContent = `fatal: ${err.message}`; + // eslint-disable-next-line no-console + console.error(err); +}); diff --git a/open4d/webclients/system/WebClient/src/nevo-client.js b/open4d/webclients/system/WebClient/src/nevo-client.js new file mode 100644 index 00000000..2623e318 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/nevo-client.js @@ -0,0 +1,274 @@ +'use strict'; + +/** + * NeVo viewer: plays pre-rendered ReRF / NeVo / captured-camera panels. + * + * A 2D canvas, not WebGL, because there is no geometry to draw — the frames are + * images. See nevo-manifest.js for why NeVo cannot be rendered client-side at + * all, and therefore why there is no camera control here. + */ + +const { + resolveConditions, frameFiles, frameFile, layoutPanels, conditionCaption, + clipSummary +} = require('./nevo-manifest'); + +class NevoClient { + /** + * @param {object} args + * @param {string} args.assetBase URL prefix serving the render output root + * @param {HTMLCanvasElement} args.canvas + * @param {string} args.object clip directory name, e.g. g_dancer + * @param {boolean} [args.nevoOnly] + * @param {number} [args.fps] + * @param {(event: object) => void} [args.onEvent] + */ + constructor({ + assetBase, canvas, object, nevoOnly = false, fps = 8, onEvent = null + }) { + this.assetBase = assetBase.replace(/\/+$/, ''); + this.canvas = canvas; + this.object = object; + this.nevoOnly = nevoOnly; + this.fps = fps; + this._onEvent = onEvent; + + this.manifest = null; + this.conditions = []; + this.images = new Map(); // file -> HTMLImageElement + this.crop = null; + this.frameIndex = 0; + this.playing = false; + + this.stats = { + imagesLoaded: 0, imagesFailed: 0, bytesUnknown: true, + framesPresented: 0, lastDrawMs: 0 + }; + this._context = canvas.getContext('2d'); + this._lastAdvance = 0; + } + + _emit(type, detail = {}) { this._onEvent?.({ type, ...detail }); } + + get clipBase() { return `${this.assetBase}/${this.object}`; } + + async start() { + const response = await fetch(`${this.clipBase}/manifest.json`, + { cache: 'no-store' }); + if (!response.ok) { + throw new Error( + `manifest fetch failed for ${this.object}: HTTP ${response.status}`); + } + this.manifest = await response.json(); + this.conditions = resolveConditions(this.manifest, + { nevoOnly: this.nevoOnly }); + this._emit('manifest', clipSummary(this.manifest, this.conditions)); + + this._installResize(); + await this._preload(); + + this.playing = true; + this._lastAdvance = performance.now(); + this._loop(); + } + + /** + * Load every image before playing. + * + * A clip is 30 images at most, so loading them all up front costs a second + * and removes any chance of the comparison showing one condition a frame + * behind another — which would be indistinguishable from a real difference + * between the conditions. + */ + async _preload() { + const entries = frameFiles(this.manifest, this.conditions); + await Promise.all(entries.map(entry => new Promise(resolve => { + const image = new Image(); + image.onload = () => { + this.images.set(entry.file, image); + this.stats.imagesLoaded++; + resolve(); + }; + image.onerror = () => { + this.stats.imagesFailed++; + this._emit('error', { message: `missing render ${entry.file}` }); + resolve(); + }; + image.src = `${this.clipBase}/${entry.file}`; + }))); + this.crop = this._contentCrop(); + this._emit('ready', { + loaded: this.stats.imagesLoaded, failed: this.stats.imagesFailed, + crop: this.crop + }); + } + + /** + * Crop to the subject, so three 4:3 panels side by side are not mostly + * empty background. + * + * Measured once from the captured-camera frame and then applied IDENTICALLY + * to every panel — identically is the point, because the conditions have to + * stay pixel-aligned or the comparison stops meaning anything. This mirrors + * what live_demo.py does on the server side. + * + * The renders are a subject over pure white, so "content" is anything below + * the white threshold. + */ + _contentCrop(pad = 0.06, threshold = 246) { + const reference = this.conditions.find(c => c.isReference) + || this.conditions[0]; + const image = this.images.get( + frameFile(reference.prefix, this.manifest.frames[0])); + if (!image) return null; + + const probe = document.createElement('canvas'); + probe.width = image.naturalWidth; + probe.height = image.naturalHeight; + const context = probe.getContext('2d', { willReadFrequently: true }); + context.drawImage(image, 0, 0); + let data; + try { + data = context.getImageData(0, 0, probe.width, probe.height).data; + } catch (_) { + return null; // tainted canvas; fall back to the full frame + } + + let minX = probe.width, minY = probe.height, maxX = -1, maxY = -1; + for (let y = 0; y < probe.height; y++) { + for (let x = 0; x < probe.width; x++) { + const i = (y * probe.width + x) * 4; + if (data[i] < threshold || data[i + 1] < threshold + || data[i + 2] < threshold) { + if (x < minX) minX = x; + if (x > maxX) maxX = x; + if (y < minY) minY = y; + if (y > maxY) maxY = y; + } + } + } + if (maxX < 0) return null; + + const padX = (maxX - minX) * pad; + const padY = (maxY - minY) * pad; + const left = Math.max(0, Math.round(minX - padX)); + const top = Math.max(0, Math.round(minY - padY)); + const right = Math.min(probe.width, Math.round(maxX + padX + 1)); + const bottom = Math.min(probe.height, Math.round(maxY + padY + 1)); + return { x: left, y: top, width: right - left, height: bottom - top }; + } + + _installResize() { + const apply = () => { + const ratio = Math.min(window.devicePixelRatio || 1, 2); + const width = this.canvas.clientWidth || 1280; + const height = this.canvas.clientHeight || 720; + this.canvas.width = Math.max(1, Math.round(width * ratio)); + this.canvas.height = Math.max(1, Math.round(height * ratio)); + this._draw(); + }; + apply(); + if (typeof ResizeObserver === 'function') { + this._resizeObserver = new ResizeObserver(apply); + this._resizeObserver.observe(this.canvas); + } else { + window.addEventListener('resize', apply); + } + } + + _draw() { + if (!this.manifest || !this._context) return; + const started = performance.now(); + const context = this._context; + const { width, height } = this.canvas; + + context.fillStyle = '#101014'; + context.fillRect(0, 0, width, height); + + const frame = this.manifest.frames[this.frameIndex]; + const first = this.images.get( + frameFile(this.conditions[0].prefix, frame)); + if (!first) return; + + const crop = this.crop + || { x: 0, y: 0, width: first.naturalWidth, height: first.naturalHeight }; + const layout = layoutPanels({ + panelCount: this.conditions.length, + sourceWidth: crop.width, + sourceHeight: crop.height, + canvasWidth: width, + canvasHeight: height, + labelHeight: Math.max(18, Math.round(height * 0.03)) + }); + + context.textBaseline = 'top'; + context.font = `${Math.max(10, Math.round(layout.labelHeight * 0.6))}px ` + + 'ui-monospace, Menlo, monospace'; + + this.conditions.forEach((condition, index) => { + const panel = layout.panels[index]; + const image = this.images.get(frameFile(condition.prefix, frame)); + if (image) { + context.drawImage(image, + crop.x, crop.y, crop.width, crop.height, + panel.x, panel.y, panel.width, panel.height); + } else { + context.fillStyle = '#1a1a22'; + context.fillRect(panel.x, panel.y, panel.width, panel.height); + } + // The captured camera is ground truth, so mark it differently from + // the two reconstructions it is there to judge. + context.fillStyle = condition.isReference ? '#8b8b98' : '#e6e6ea'; + context.fillText(conditionCaption(condition), + panel.x + 2, panel.labelY + 2, panel.width - 4); + }); + + this.stats.lastDrawMs = performance.now() - started; + this.stats.framesPresented++; + } + + _loop() { + this._rafHandle = requestAnimationFrame(() => this._loop()); + if (!this.playing || !this.manifest) return; + const interval = 1000 / this.fps; + const now = performance.now(); + if (now - this._lastAdvance >= interval) { + const steps = Math.floor((now - this._lastAdvance) / interval); + this._lastAdvance += steps * interval; + this.frameIndex = + (this.frameIndex + steps) % this.manifest.frames.length; + this._draw(); + } + } + + setPlaying(playing) { + this.playing = playing; + this._lastAdvance = performance.now(); + } + + stop() { + this.playing = false; + if (this._rafHandle) cancelAnimationFrame(this._rafHandle); + this._resizeObserver?.disconnect(); + } + + inspect() { + return { + object: this.object, + frameIndex: this.frameIndex, + frameCount: this.manifest ? this.manifest.frames.length : 0, + playing: this.playing, + conditions: this.conditions.map(c => ({ + prefix: c.prefix, label: c.label, + keptFraction: c.kept_fraction ?? null + })), + summary: this.manifest + ? clipSummary(this.manifest, this.conditions) : null, + crop: this.crop, + stats: { ...this.stats }, + canvas: { width: this.canvas.width, height: this.canvas.height } + }; + } +} + +module.exports = { NevoClient }; diff --git a/open4d/webclients/system/WebClient/src/nevo-main.js b/open4d/webclients/system/WebClient/src/nevo-main.js new file mode 100644 index 00000000..b1270d08 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/nevo-main.js @@ -0,0 +1,95 @@ +'use strict'; + +/** + * Entry point for the NeVo viewer. + * + * /web/nevo.html?object=g_dancer + * + * | parameter | default | meaning | + * |-----------|---------------|--------------------------------------------| + * | assets | /nevo-assets | URL prefix serving the render output root | + * | object | g_dancer | clip directory name | + * | fps | 8 | playback rate | + * | nevoOnly | 0 | 1 shows only NeVo's own filtered output | + */ + +const { NevoClient } = require('./nevo-client'); + +async function main() { + const params = new URLSearchParams(window.location.search); + const status = document.getElementById('status'); + const logPane = document.getElementById('log'); + const statsPane = document.getElementById('stats'); + + const log = (level, message) => { + const row = document.createElement('div'); + row.className = `log-line log-${level}`; + row.textContent = message; + logPane.appendChild(row); + while (logPane.childElementCount > 200) { + logPane.removeChild(logPane.firstChild); + } + logPane.scrollTop = logPane.scrollHeight; + }; + + const fps = Number(params.get('fps')); + const client = new NevoClient({ + assetBase: params.get('assets') || '/nevo-assets', + canvas: document.getElementById('view'), + object: params.get('object') || 'g_dancer', + nevoOnly: params.get('nevoOnly') === '1', + fps: Number.isFinite(fps) && fps > 0 ? fps : 8, + onEvent: event => { + if (event.type === 'manifest') { + log('info', `${event.name} · ${event.representation}`); + log('info', `${event.frames} frames · ${event.source} · ` + + `view ${event.view}` + + (event.viewInTrainingSet + ? ' (in the training set)' : ' (held out)')); + for (const kept of event.keptFractions) { + log('info', ` ${kept.label}: kept ` + + `${(kept.keptFraction * 100).toFixed(1)}% of voxels`); + } + status.textContent = 'loading renders …'; + } else if (event.type === 'ready') { + log('info', `${event.loaded} renders loaded` + + (event.failed ? `, ${event.failed} missing` : '')); + status.textContent = 'playing pre-rendered frames'; + } else if (event.type === 'error') { + log('error', event.message); + } + } + }); + window.__vs4dNevo = client; + + const playButton = document.getElementById('play'); + playButton.addEventListener('click', () => { + client.setPlaying(!client.playing); + playButton.textContent = client.playing ? 'Pause' : 'Play'; + }); + + setInterval(() => { + const info = client.inspect(); + statsPane.textContent = [ + `clip ${info.object}`, + `frame ${info.frameIndex + 1} / ${info.frameCount}`, + ...(info.stats.imagesFailed + ? [`MISSING ${info.stats.imagesFailed} renders`] : []) + ].join('\n'); + }, 500); + + try { + await client.start(); + } catch (error) { + status.textContent = `could not start: ${error.message}`; + log('error', error.message); + log('info', 'Renders come from orbitnevo/render_frames.py; point ' + + '?assets= at its output root (default ~/nevo_output, served at ' + + '/nevo-assets).'); + } +} + +main().catch(error => { + document.getElementById('status').textContent = `fatal: ${error.message}`; + console.error(error); +}); diff --git a/open4d/webclients/system/WebClient/src/nevo-manifest.js b/open4d/webclients/system/WebClient/src/nevo-manifest.js new file mode 100644 index 00000000..24c44392 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/nevo-manifest.js @@ -0,0 +1,158 @@ +'use strict'; + +/** + * NeVo clip manifest handling: conditions, frame filenames, panel layout. + * + * Pure, so the parts that decide WHAT is shown can be tested without a browser + * (see tests/test_nevo_manifest.js). The drawing itself is in nevo-client.js. + * + * NeVo is the one baseline that cannot run client-side at all. It is a NeRF + * (ReRF feature voxels) whose frames take roughly half a second each to + * ray-march on a workstation GPU, and its entropy decoder is CUDA and Python + * 3.8. So the browser shows PRE-RENDERED frames — exactly what + * `orbitnevo/live_demo.py` does, and its own docstring says the same. There is + * therefore no camera control here, and pretending otherwise would be + * misleading rather than convenient. + * + * What the panel does show is the comparison the baseline exists for: plain + * ReRF against NeVo's visibility-filtered reconstruction at the same instant + * and viewpoint, with the captured camera alongside as ground truth. + */ + +/** The captured-camera condition, appended the way live_demo.py appends it. */ +function referenceCondition(manifest) { + return { + name: 'capture', + prefix: 'reference', + label: `captured camera ${manifest.view}`, + threshold: null, + kept_fraction: null, + isReference: true + }; +} + +/** + * Resolve which conditions to show, in display order. + * + * @param {object} manifest parsed manifest.json + * @param {object} [options] + * @param {boolean} [options.nevoOnly] show only NeVo's own output — the + * conditions with a threshold — dropping plain ReRF and the captured camera, + * which are the comparison rather than the output. + * @param {boolean} [options.withReference=true] + */ +function resolveConditions(manifest, { nevoOnly = false, withReference = true } = {}) { + const declared = Array.isArray(manifest.conditions) ? manifest.conditions : []; + if (nevoOnly) { + return declared.filter(condition => condition.threshold !== null + && condition.threshold !== undefined); + } + const resolved = [...declared]; + if (withReference) resolved.push(referenceCondition(manifest)); + return resolved; +} + +/** `{prefix}_{frame:03d}.png`, matching what render_frames.py wrote. */ +function frameFile(prefix, frame) { + return `${prefix}_${String(frame).padStart(3, '0')}.png`; +} + +/** Every image a clip needs, so they can be preloaded before playback. */ +function frameFiles(manifest, conditions) { + const files = []; + for (const frame of manifest.frames) { + for (const condition of conditions) { + files.push({ frame, condition, file: frameFile(condition.prefix, frame) }); + } + } + return files; +} + +/** + * Lay out the condition panels in a row. + * + * The source renders are 1280x960 with the subject a small part of the frame, + * so a crop is applied identically to every panel — identically, because the + * conditions must stay pixel-aligned for the comparison to mean anything. + * + * @param {object} args + * @param {number} args.panelCount + * @param {number} args.sourceWidth width after cropping + * @param {number} args.sourceHeight height after cropping + * @param {number} args.canvasWidth space available + * @param {number} args.canvasHeight + * @param {number} [args.labelHeight] + * @param {number} [args.gap] + */ +function layoutPanels({ + panelCount, sourceWidth, sourceHeight, canvasWidth, canvasHeight, + labelHeight = 26, gap = 6 +}) { + if (panelCount <= 0) throw new RangeError('panelCount must be positive'); + if (!(sourceWidth > 0) || !(sourceHeight > 0)) { + throw new RangeError('source dimensions must be positive'); + } + const totalGap = gap * (panelCount - 1); + const aspect = sourceWidth / sourceHeight; + + // Fit by width, then shrink if the resulting height does not fit. Both + // constraints matter: a wide window is width-bound, a tall narrow one is + // height-bound, and picking only one leaves panels clipped. + let panelWidth = Math.max(1, Math.floor((canvasWidth - totalGap) / panelCount)); + let panelHeight = Math.round(panelWidth / aspect); + const available = canvasHeight - labelHeight; + if (panelHeight > available && available > 0) { + panelHeight = available; + panelWidth = Math.max(1, Math.round(panelHeight * aspect)); + } + + const rowWidth = panelWidth * panelCount + totalGap; + const originX = Math.max(0, Math.round((canvasWidth - rowWidth) / 2)); + const originY = Math.max(0, Math.round( + (canvasHeight - (panelHeight + labelHeight)) / 2)); + + return { + panelWidth, panelHeight, labelHeight, gap, rowWidth, originX, originY, + panels: Array.from({ length: panelCount }, (_, index) => ({ + index, + x: originX + index * (panelWidth + gap), + y: originY, + width: panelWidth, + height: panelHeight, + labelY: originY + panelHeight + })) + }; +} + +/** A caption line for a condition: what it is, and how much it kept. */ +function conditionCaption(condition) { + if (condition.kept_fraction === null || condition.kept_fraction === undefined) { + return condition.label; + } + const percent = (condition.kept_fraction * 100).toFixed(1); + return `${condition.label} · ${percent}% of voxels`; +} + +/** Headline facts about a clip, for the page's status area. */ +function clipSummary(manifest, conditions) { + const filtered = conditions.filter( + c => c.threshold !== null && c.threshold !== undefined); + return { + name: manifest.name, + representation: manifest.representation, + frames: manifest.frames.length, + view: manifest.view, + viewInTrainingSet: manifest.view_in_training_set === true, + source: `${manifest.width}x${manifest.height}`, + seconds: manifest.seconds, + conditions: conditions.length, + keptFractions: filtered.map(c => ({ + label: c.label, threshold: c.threshold, keptFraction: c.kept_fraction + })) + }; +} + +module.exports = { + resolveConditions, referenceCondition, frameFile, frameFiles, + layoutPanels, conditionCaption, clipSummary +}; diff --git a/open4d/webclients/system/WebClient/src/point-reconstruction.js b/open4d/webclients/system/WebClient/src/point-reconstruction.js new file mode 100644 index 00000000..661f1ca0 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/point-reconstruction.js @@ -0,0 +1,259 @@ +'use strict'; + +/** + * MetaStream / DeltaStream point-cloud reconstruction. + * + * A faithful port of `ReconstructionState` in + * `baselines/DeltaStream/orbitstream/reconstruction.py`, which is the + * authority. `tests/test_point_reconstruction.js` compares this against that + * implementation's output on the same input, because the delta model is fiddly + * enough that "looks about right on screen" is not evidence. + * + * The streamed points are in CAMERA space. The server encodes + * `cloud.positions` unchanged and only uses `camera_to_world` for tile + * assignment, so placing the points is the client's job — see `worldClouds()`. + * + * The delta model, per (object, camera) stream: + * + * keyframe the payload IS the whole cloud; replace. + * delta project the previous cloud back into the source image to recover + * each point's block, then + * 1. copy the points inside each motion's source block and + * translate them by that motion's delta, + * 2. drop every point whose block appears in removal_blocks, + * 3. append the residual payload. + * A motion source may overlap a removal block; the translated copy + * survives, matching the reference desktop client. + */ + +/** Camera-space cloud: interleaved xyz floats and rgb bytes. */ +class PointCloud { + constructor(positions, colors) { + this.positions = positions; // Float32Array, 3N + this.colors = colors; // Uint8Array, 3N + } + + static empty() { + return new PointCloud(new Float32Array(0), new Uint8Array(0)); + } + + get pointCount() { return this.positions.length / 3; } +} + +/** Concatenate clouds in order. */ +function concatClouds(clouds) { + const total = clouds.reduce((sum, c) => sum + c.pointCount, 0); + if (total === 0) return PointCloud.empty(); + const positions = new Float32Array(total * 3); + const colors = new Uint8Array(total * 3); + let offset = 0; + for (const cloud of clouds) { + positions.set(cloud.positions, offset); + colors.set(cloud.colors, offset); + offset += cloud.positions.length; + } + return new PointCloud(positions, colors); +} + +class ReconstructionState { + /** @param {object} header decoded CONNECTION message */ + constructor(header) { + this.header = header; + this.calibrations = new Map(); + for (const object of header.objects) { + for (const camera of object.cameras) { + this.calibrations.set(`${object.objectId}:${camera.cameraId}`, + { ...camera, objectId: object.objectId }); + } + } + this._cameraClouds = new Map(); // "obj:cam" -> PointCloud + this.lastFrameId = -1; + } + + reset() { + this._cameraClouds.clear(); + this.lastFrameId = -1; + } + + /** + * Fold one FRAME into the state and return the per-object world clouds. + * + * @param {object} frame decoded FRAME message + * @param {(draco: Uint8Array) => PointCloud} decodeDraco + * @param {{strictOrder?: boolean}} [options] + * strictOrder mirrors the Python implementation, which refuses a gap in + * the frame sequence. A browser over a real link may legitimately miss a + * frame, in which case the stream cannot be reconstructed and the caller + * must resynchronise on the next keyframe rather than render nonsense. + */ + apply(frame, decodeDraco, { strictOrder = true } = {}) { + if (strictOrder && frame.frameId !== this.lastFrameId + 1) { + throw new Error( + `dependency frame out of order: expected ${this.lastFrameId + 1}, ` + + `received ${frame.frameId}`); + } + if (frame.frameId === 0 && frame.frameType !== 'keyframe') { + throw new Error('a run must begin with a keyframe'); + } + + const seen = new Set(); + for (const record of frame.records) { + const key = `${record.objectId}:${record.cameraId}`; + if (seen.has(key)) throw new Error(`duplicate record for stream ${key}`); + if (!this.calibrations.has(key)) { + throw new Error(`unknown stream ${key}`); + } + seen.add(key); + + const residual = record.draco && record.draco.length + ? decodeDraco(record.draco) : PointCloud.empty(); + if (residual.pointCount !== record.pointCount) { + throw new Error( + `decoded point count differs for ${key}: ` + + `${residual.pointCount} != ${record.pointCount}`); + } + + if (frame.frameType === 'keyframe') { + if (record.removalBlocks.length || record.motions.length) { + throw new Error('keyframes cannot contain delta operations'); + } + this._cameraClouds.set(key, residual); + } else { + if (!this._cameraClouds.has(key)) { + throw new Error(`delta precedes keyframe for stream ${key}`); + } + this._cameraClouds.set(key, this._applyDelta( + this._cameraClouds.get(key), residual, record)); + } + } + + this.lastFrameId = frame.frameId; + return this.worldClouds(); + } + + /** Streams the header declares but this frame omitted. */ + missingStreams(seenKeys) { + return [...this.calibrations.keys()].filter(key => !seenKeys.has(key)); + } + + _applyDelta(previous, residual, record) { + const calibration = this.calibrations.get( + `${record.objectId}:${record.cameraId}`); + const { fx, fy, cx, cy } = calibration; + const blockSize = this.header.blockSize; + const blocksPerRow = Math.floor(this.header.width / blockSize); + const count = previous.pointCount; + const points = previous.positions; + + // Project each previous point back into its source image to recover the + // block it belongs to. A point behind the camera or outside the frame + // has no block, and must be excluded rather than clamped. + const u = new Float32Array(count); + const v = new Float32Array(count); + const inImage = new Uint8Array(count); + const blockIndices = new Int32Array(count).fill(-1); + + for (let i = 0; i < count; i++) { + const z = points[i * 3 + 2]; + if (!(z > 0)) { u[i] = -1; v[i] = -1; continue; } + const pu = (points[i * 3] / z) * fx + cx; + const pv = (points[i * 3 + 1] / z) * fy + cy; + u[i] = pu; + v[i] = pv; + if (pu >= 0 && pv >= 0 && pu < this.header.width + && pv < this.header.height) { + inImage[i] = 1; + blockIndices[i] = Math.floor(pv / blockSize) * blocksPerRow + + Math.floor(pu / blockSize); + } + } + + // 1. Motion: translated copies of the points in each source block. + const moved = []; + for (const motion of record.motions) { + const x0 = motion.sourceX - blockSize / 2; + const y0 = motion.sourceY - blockSize / 2; + const selected = []; + for (let i = 0; i < count; i++) { + if (!inImage[i]) continue; + if (u[i] >= x0 && u[i] < x0 + blockSize + && v[i] >= y0 && v[i] < y0 + blockSize) { + selected.push(i); + } + } + if (selected.length === 0) continue; + const positions = new Float32Array(selected.length * 3); + const colors = new Uint8Array(selected.length * 3); + selected.forEach((source, target) => { + positions[target * 3] = points[source * 3] + motion.deltaX; + positions[target * 3 + 1] = points[source * 3 + 1] + motion.deltaY; + positions[target * 3 + 2] = points[source * 3 + 2] + motion.deltaZ; + colors[target * 3] = previous.colors[source * 3]; + colors[target * 3 + 1] = previous.colors[source * 3 + 1]; + colors[target * 3 + 2] = previous.colors[source * 3 + 2]; + }); + moved.push(new PointCloud(positions, colors)); + } + + // 2. Removal: drop points whose block was republished. + let kept; + if (record.removalBlocks.length) { + const removed = new Set(Array.from(record.removalBlocks)); + const keepIndices = []; + for (let i = 0; i < count; i++) { + if (!removed.has(blockIndices[i])) keepIndices.push(i); + } + const positions = new Float32Array(keepIndices.length * 3); + const colors = new Uint8Array(keepIndices.length * 3); + keepIndices.forEach((source, target) => { + positions.set(points.subarray(source * 3, source * 3 + 3), target * 3); + colors.set(previous.colors.subarray(source * 3, source * 3 + 3), + target * 3); + }); + kept = new PointCloud(positions, colors); + } else { + kept = previous; + } + + // 3. Order matches the Python implementation: moved, residual, kept. + const parts = [...moved]; + if (residual.pointCount) parts.push(residual); + if (kept.pointCount) parts.push(kept); + return concatClouds(parts); + } + + /** + * Per-object clouds in world space. + * + * `cameraToWorldRowMajor` is row-major (the Python encoder writes + * `for row in matrix for v in row`), so the rotation is rows 0..2 columns + * 0..2 and the translation is column 3 — indices 3, 7, 11. + */ + worldClouds() { + const byObject = new Map(); + for (const [key, cloud] of this._cameraClouds) { + const calibration = this.calibrations.get(key); + const m = calibration.cameraToWorldRowMajor; + const count = cloud.pointCount; + const positions = new Float32Array(count * 3); + for (let i = 0; i < count; i++) { + const x = cloud.positions[i * 3]; + const y = cloud.positions[i * 3 + 1]; + const z = cloud.positions[i * 3 + 2]; + positions[i * 3] = m[0] * x + m[1] * y + m[2] * z + m[3]; + positions[i * 3 + 1] = m[4] * x + m[5] * y + m[6] * z + m[7]; + positions[i * 3 + 2] = m[8] * x + m[9] * y + m[10] * z + m[11]; + } + const transformed = new PointCloud(positions, cloud.colors); + const objectId = calibration.objectId; + byObject.set(objectId, byObject.has(objectId) + ? concatClouds([byObject.get(objectId), transformed]) + : transformed); + } + return byObject; + } + + get streamCount() { return this._cameraClouds.size; } +} + +module.exports = { ReconstructionState, PointCloud, concatClouds }; diff --git a/open4d/webclients/system/WebClient/src/point-renderer.js b/open4d/webclients/system/WebClient/src/point-renderer.js new file mode 100644 index 00000000..ee6d8373 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/point-renderer.js @@ -0,0 +1,246 @@ +'use strict'; + +/** + * Point-cloud renderer for the V4DS baselines. + * + * One `THREE.Points` per object, fed world-space clouds from + * `point-reconstruction.js`. Shared by MetaStream, DeltaStream, ViVo, NAVA and + * LiVo, which is the point: they differ in what they choose to send, not in how + * a received cloud is drawn, so putting the drawing in one place keeps the + * comparison about the algorithms. + * + * Deliberately separate from `webgl-renderer.js`. That one implements the + * ClientPlatform renderer contract for the mesh ladder — decode caching, + * per-frame clip playback, `object_ready` crediting. None of that applies here: + * a baseline pushes a fresh cloud whenever it likes, there is no segment + * buffer, and there is no ABR asking whether an object became playable. Forcing + * both through one class would mean a contract that fits neither. + * + * Geometry is reallocated only when a cloud outgrows its buffer. Point counts + * swing frame to frame (a DeltaStream residual adds a few thousand, a keyframe + * replaces everything), and allocating a new BufferGeometry per frame at 30 Hz + * makes the garbage collector the bottleneck. + */ + +const THREE = require('three'); +const { OrbitControls } = require('three/examples/jsm/controls/OrbitControls.js'); + +/** Extra headroom when growing a buffer, so small growth is not a realloc. */ +const GROWTH_FACTOR = 1.5; + +class PointRenderer { + /** + * @param {object} args + * @param {HTMLCanvasElement} args.canvas + * @param {number} [args.pointSize] world-space point size, metres + * @param {object} [args.scene] { background } + */ + constructor({ canvas, pointSize = 0.012, scene: sceneOptions = {} }) { + this.canvas = canvas; + this.pointSize = pointSize; + this._sceneOptions = sceneOptions; + this._objects = new Map(); // objectId -> { points, geometry, capacity } + this._userMovedCamera = false; + this._framedObjectCount = 0; + this._running = false; + this._presented = 0; + this._three = null; + } + + start() { + this._three = this._buildScene(); + this._running = true; + this._loop(); + return this; + } + + stop() { + this._running = false; + if (this._rafHandle) cancelAnimationFrame(this._rafHandle); + this._resizeObserver?.disconnect(); + for (const entry of this._objects.values()) { + entry.geometry.dispose(); + entry.points.material.dispose(); + } + this._objects.clear(); + this._three?.renderer.dispose(); + } + + _buildScene() { + const renderer = new THREE.WebGLRenderer({ canvas: this.canvas, antialias: false }); + const scene = new THREE.Scene(); + scene.background = new THREE.Color(this._sceneOptions.background ?? 0x101014); + const camera = new THREE.PerspectiveCamera(60, 1, 0.05, 500); + camera.position.set(0, 1.6, 4); + + const controls = new OrbitControls(camera, this.canvas); + controls.target.set(0, 1.0, 0); + controls.enableDamping = true; + controls.update(); + controls.addEventListener('start', () => { this._userMovedCamera = true; }); + + const apply = () => { + const width = this.canvas.clientWidth || this.canvas.width || 1280; + const height = this.canvas.clientHeight || this.canvas.height || 720; + if (width <= 0 || height <= 0) return; + renderer.setPixelRatio(Math.min(window.devicePixelRatio || 1, 2)); + renderer.setSize(width, height, false); + camera.aspect = width / height; + camera.updateProjectionMatrix(); + }; + apply(); + if (typeof ResizeObserver === 'function') { + this._resizeObserver = new ResizeObserver(apply); + this._resizeObserver.observe(this.canvas); + } else { + window.addEventListener('resize', apply); + } + + return { renderer, scene, camera, controls }; + } + + _entry(objectId, requiredPoints) { + let entry = this._objects.get(objectId); + if (!entry) { + const geometry = new THREE.BufferGeometry(); + const capacity = Math.max(1, Math.ceil(requiredPoints * GROWTH_FACTOR)); + geometry.setAttribute('position', + new THREE.BufferAttribute(new Float32Array(capacity * 3), 3)); + geometry.setAttribute('color', + new THREE.BufferAttribute(new Uint8Array(capacity * 3), 3, true)); + const material = new THREE.PointsMaterial({ + size: this.pointSize, sizeAttenuation: true, vertexColors: true + }); + const points = new THREE.Points(geometry, material); + points.frustumCulled = false; // bounds change every frame + this._three.scene.add(points); + entry = { points, geometry, capacity }; + this._objects.set(objectId, entry); + return entry; + } + if (requiredPoints > entry.capacity) { + const capacity = Math.ceil(requiredPoints * GROWTH_FACTOR); + entry.geometry.setAttribute('position', + new THREE.BufferAttribute(new Float32Array(capacity * 3), 3)); + entry.geometry.setAttribute('color', + new THREE.BufferAttribute(new Uint8Array(capacity * 3), 3, true)); + entry.capacity = capacity; + } + return entry; + } + + /** + * Show the per-object world clouds produced by a reconstruction step. + * + * @param {Map} clouds + */ + update(clouds) { + for (const [objectId, cloud] of clouds) { + const count = cloud.pointCount; + const entry = this._entry(objectId, count); + const position = entry.geometry.getAttribute('position'); + const color = entry.geometry.getAttribute('color'); + position.array.set(cloud.positions); + color.array.set(cloud.colors); + // drawRange, not a resize: the buffers are oversized on purpose and + // only the first `count` points are valid this frame. + entry.geometry.setDrawRange(0, count); + position.needsUpdate = true; + color.needsUpdate = true; + entry.points.visible = count > 0; + entry.lastCount = count; + } + // An object the stream stopped sending should disappear rather than + // freeze at its last cloud, which would read as a live object. + for (const [objectId, entry] of this._objects) { + if (!clouds.has(objectId)) entry.points.visible = false; + } + if (!this._userMovedCamera && this._objects.size !== this._framedObjectCount) { + this._framedObjectCount = this._objects.size; + this.frameScene(clouds); + } + } + + /** + * Point the camera at the content. + * + * Computed from the cloud data rather than Three.js bounding spheres: the + * buffers are oversized, so the unused tail would drag the bounds towards + * the origin. The ORBIT corpus is also baked into venue world coordinates + * with raised floor tiers, so a fixed pose frames empty air. + */ + frameScene(clouds) { + if (this._userMovedCamera || !this._three) return; + let minX = Infinity, minY = Infinity, minZ = Infinity; + let maxX = -Infinity, maxY = -Infinity, maxZ = -Infinity; + let total = 0; + for (const cloud of clouds.values()) { + for (let i = 0; i < cloud.pointCount; i++) { + const x = cloud.positions[i * 3]; + const y = cloud.positions[i * 3 + 1]; + const z = cloud.positions[i * 3 + 2]; + if (x < minX) minX = x; + if (y < minY) minY = y; + if (z < minZ) minZ = z; + if (x > maxX) maxX = x; + if (y > maxY) maxY = y; + if (z > maxZ) maxZ = z; + total++; + } + } + if (total === 0 || !Number.isFinite(minX)) return; + + const centre = new THREE.Vector3( + (minX + maxX) / 2, (minY + maxY) / 2, (minZ + maxZ) / 2); + const radius = Math.max( + 0.5, Math.hypot(maxX - minX, maxY - minY, maxZ - minZ) / 2); + const { camera, controls } = this._three; + const distance = (radius / Math.tan((camera.fov * Math.PI) / 360)) * 1.6; + controls.target.copy(centre); + camera.position.set(centre.x, centre.y + radius * 0.15, centre.z + distance); + camera.near = Math.max(radius / 20, 0.05); + camera.far = distance + radius * 30; + camera.updateProjectionMatrix(); + controls.update(); + } + + _loop() { + if (!this._running) return; + this._rafHandle = requestAnimationFrame(() => this._loop()); + this._three.controls.update(); + this._three.renderer.render(this._three.scene, this._three.camera); + this._presented++; + } + + /** Camera pose in the shape V4DS FEEDBACK wants. */ + viewerState() { + const { camera, controls } = this._three; + camera.updateMatrixWorld(); + const forward = new THREE.Vector3(); + camera.getWorldDirection(forward); + return { + position: camera.position.toArray(), + forward: forward.toArray(), + up: camera.up.toArray(), + target: controls.target.toArray(), + verticalFovDegrees: camera.fov, + aspect: camera.aspect, + near: camera.near, + far: camera.far + }; + } + + get stats() { + return { + framesPresented: this._presented, + objects: [...this._objects].map(([objectId, entry]) => ({ + objectId, + points: entry.lastCount ?? 0, + capacity: entry.capacity, + visible: entry.points.visible + })) + }; + } +} + +module.exports = { PointRenderer }; diff --git a/open4d/webclients/system/WebClient/src/splat-renderer.js b/open4d/webclients/system/WebClient/src/splat-renderer.js new file mode 100644 index 00000000..7a4a33a2 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/splat-renderer.js @@ -0,0 +1,422 @@ +'use strict'; + +/** + * 3D Gaussian splat renderer for the Vega baseline. + * + * Draws anisotropic Gaussians the way 3DGS does: project each splat's 3D + * covariance into a screen-space 2D covariance, size a camera-facing quad to + * its eigenvectors, and evaluate the Gaussian falloff in the fragment shader + * with premultiplied-alpha blending, back to front. + * + * Two design choices worth knowing: + * + * 1. SPLAT DATA LIVES IN A TEXTURE, and the only per-instance attribute is an + * index into it. Correct 3DGS needs back-to-front ordering, which changes + * every time the camera moves. Reordering the attribute buffers themselves + * would mean copying 14 floats per splat per sort — about 3 MB for a 58k + * frame. Reordering a single index attribute is 230 KB. + * + * 2. SORTING IS A COUNTING SORT over quantised depth, not a comparison sort. + * At 58k splats a comparison sort costs milliseconds of the frame budget; + * a 16-bit bucket pass is a few hundred microseconds and the ordering error + * within a bucket is far below what the blending can show. + * + * Depth testing is off and depth writing is off, as 3DGS requires: the ordering + * IS the depth resolution. Turning them on produces hard edges where splats + * should blend. + */ + +const THREE = require('three'); + +/** Texels per splat in the data texture: 4 x RGBA32F = 16 floats. */ +const TEXELS_PER_SPLAT = 4; +const TEXTURE_WIDTH = 1024; +/** Depth buckets for the counting sort. */ +const SORT_BUCKETS = 1 << 16; + +const VERTEX_SHADER = /* glsl */` +precision highp float; +precision highp int; + +// RawShaderMaterial does NOT inject Three.js's built-in uniforms, so they are +// declared here. GLSL 3 is required for transpose(), which GLSL ES 1.00 lacks. +uniform mat4 modelMatrix; +uniform mat4 viewMatrix; +uniform mat4 projectionMatrix; + +uniform sampler2D splatData; +uniform vec2 splatTextureSize; +uniform vec2 viewport; +uniform float splatScale; + +in vec2 quadPosition; // corner in [-1, 1] +in float splatIndex; + +out vec4 vColor; +out vec2 vGaussian; + +vec4 fetch(float texel) { + float index = splatIndex * ${TEXELS_PER_SPLAT}.0 + texel; + float x = mod(index, splatTextureSize.x); + float y = floor(index / splatTextureSize.x); + return texture(splatData, (vec2(x, y) + 0.5) / splatTextureSize); +} + +mat3 quaternionToMatrix(vec4 q) { + float x = q.x, y = q.y, z = q.z, w = q.w; + return mat3( + 1.0 - 2.0 * (y * y + z * z), 2.0 * (x * y + w * z), 2.0 * (x * z - w * y), + 2.0 * (x * y - w * z), 1.0 - 2.0 * (x * x + z * z), 2.0 * (y * z + w * x), + 2.0 * (x * z + w * y), 2.0 * (y * z - w * x), 1.0 - 2.0 * (x * x + y * y) + ); +} + +void main() { + vec4 centerAndOpacity = fetch(0.0); + vec4 logScale = fetch(1.0); + vec4 rotation = fetch(2.0); + vec4 color = fetch(3.0); + + vec3 center = centerAndOpacity.xyz; + float opacity = centerAndOpacity.w; + + mat4 modelView = viewMatrix * modelMatrix; + vec4 cam = modelView * vec4(center, 1.0); + vec4 clip = projectionMatrix * cam; + // Behind the camera: park the vertex outside the clip volume rather than + // letting a divide by a near-zero w throw the quad across the screen. + if (clip.w <= 0.0) { + gl_Position = vec4(0.0, 0.0, 2.0, 1.0); + vColor = vec4(0.0); + vGaussian = vec2(0.0); + return; + } + + vec3 scale = exp(logScale.xyz) * splatScale; + +#ifdef ISOTROPIC_SPLATS + // Isotropic path: size the quad from the mean scale projected through the + // focal length, with no covariance projection at all. + // + // Two reasons this is the default rather than a fallback. First, Vega's + // Gaussians are near-isotropic — the exported log scales sit within about + // 0.3 of each other — so the anisotropic projection buys very little here. + // Second, the full covariance path (mat3 transposes and products in the + // vertex stage) does not render under software GL: measured in headless + // Chrome with SwiftShader, a shader that merely CONTAINS that math draws + // nothing even when gl_Position does not use its result, while the same + // shader without it draws correctly. That is a driver-level failure, not a + // logic error, but it makes the anisotropic path unverifiable here. + float fxIso = projectionMatrix[0][0] * viewport.x * 0.5; + float fyIso = projectionMatrix[1][1] * viewport.y * 0.5; + float meanScale = (scale.x + scale.y + scale.z) / 3.0; + float depth = max(-cam.z, 1e-4); + // Two sigma, with a half-pixel floor so a distant splat still marks a pixel. + vec2 radiusPx = vec2(max(2.0 * fxIso * meanScale / depth, 0.5), + max(2.0 * fyIso * meanScale / depth, 0.5)); + vec2 offset = quadPosition * radiusPx; + vColor = vec4(color.rgb, opacity); + vGaussian = quadPosition * 2.0; + gl_Position = vec4( + clip.xy / clip.w + offset / viewport * 2.0, + clip.z / clip.w, 1.0); +#else + mat3 rotationMatrix = quaternionToMatrix(rotation); + mat3 scaled = mat3( + rotationMatrix[0] * scale.x, + rotationMatrix[1] * scale.y, + rotationMatrix[2] * scale.z); + mat3 covariance3d = scaled * transpose(scaled); + + // Focal lengths in pixels, recovered from the projection matrix so this + // follows whatever FOV and viewport the camera currently has. + float fx = projectionMatrix[0][0] * viewport.x * 0.5; + float fy = projectionMatrix[1][1] * viewport.y * 0.5; + + // Jacobian of the perspective projection at this splat's camera position. + mat3 jacobian = mat3( + fx / cam.z, 0.0, -(fx * cam.x) / (cam.z * cam.z), + 0.0, fy / cam.z, -(fy * cam.y) / (cam.z * cam.z), + 0.0, 0.0, 0.0); + mat3 world = transpose(mat3(modelView)); + mat3 transform = world * jacobian; + mat3 covariance2d = transpose(transform) * covariance3d * transform; + + // A low-pass term keeps sub-pixel splats from vanishing entirely, which is + // what 3DGS calls the dilation filter. + float a = covariance2d[0][0] + 0.3; + float b = covariance2d[0][1]; + float c = covariance2d[1][1] + 0.3; + float mid = 0.5 * (a + c); + float radius = length(vec2(0.5 * (a - c), b)); + float lambda1 = mid + radius; + float lambda2 = max(mid - radius, 0.1); + // A splat smaller than a pixel contributes nothing; skip its quad. + if (lambda1 < 0.02) { + gl_Position = vec4(0.0, 0.0, 2.0, 1.0); + vColor = vec4(0.0); + vGaussian = vec2(0.0); + return; + } + + // The eigenvector of the 2D covariance, guarded against the isotropic case. + // + // This guard is essential, not defensive. For a near-isotropic splat the + // off-diagonal b tends to 0 AND lambda1 tends to a, so the unguarded + // normalize(vec2(b, lambda1 - a)) is normalize(vec2(0, 0)) = NaN. A NaN + // gl_Position produces no primitive at all, silently: every intermediate + // value reads correct, 100k triangles are submitted, and not one pixel is + // shaded. Vega's Gaussians are close to isotropic (log scales within 0.3 of + // each other), so this is the common case here, not a rare one. + // + // When the covariance really is isotropic, any orthonormal basis is a + // correct pair of axes, so falling back to the x-axis loses nothing. + vec2 eigenDirection = vec2(b, lambda1 - a); + float eigenLength = length(eigenDirection); + vec2 majorAxis = eigenLength > 1e-6 + ? eigenDirection / eigenLength + : vec2(1.0, 0.0); + // Clamped so one degenerate splat cannot ask for a screen-filling quad. + vec2 axis1 = min(sqrt(2.0 * lambda1), 1024.0) * majorAxis; + vec2 axis2 = min(sqrt(2.0 * lambda2), 1024.0) * vec2(majorAxis.y, -majorAxis.x); + + vec2 offset = quadPosition.x * axis1 + quadPosition.y * axis2; + vColor = vec4(color.rgb, opacity); + // Two sigma across the quad, matching the 4.0 cutoff in the fragment stage. + vGaussian = quadPosition * 2.0; + + gl_Position = vec4( + clip.xy / clip.w + offset / viewport * 2.0, + clip.z / clip.w, 1.0); +#endif +} +`; + +const FRAGMENT_SHADER = /* glsl */` +precision highp float; + +in vec4 vColor; +in vec2 vGaussian; + +out vec4 fragColor; + +void main() { + float power = -dot(vGaussian, vGaussian); + // Beyond two sigma the contribution is under 2%; discarding there saves + // most of the fill cost with no visible change. + if (power < -4.0) discard; + float alpha = exp(0.5 * power) * vColor.a; + if (alpha < 1.0 / 255.0) discard; + // Premultiplied alpha, to match the ONE / ONE_MINUS_SRC_ALPHA blend. + fragColor = vec4(vColor.rgb * alpha, alpha); +} +`; + +/** + * One object's splat cloud: a data texture plus a sortable index attribute. + */ +class SplatObject { + /** + * @param {object} args + * @param {number} args.capacity + * @param {'isotropic'|'anisotropic'} [args.mode] quad sizing. Isotropic is + * the default; see the ISOTROPIC_SPLATS note in the vertex shader. + */ + constructor({ capacity, mode = 'isotropic' }) { + this.mode = mode; + this._bound = false; + this.capacity = 0; + this.count = 0; + this.geometry = new THREE.InstancedBufferGeometry(); + + // A unit quad, two triangles, shared by every instance. + this.geometry.setAttribute('quadPosition', + new THREE.BufferAttribute(new Float32Array([ + -1, -1, 1, -1, 1, 1, -1, 1 + ]), 2)); + this.geometry.setIndex([0, 1, 2, 0, 2, 3]); + + this.material = new THREE.RawShaderMaterial({ + vertexShader: VERTEX_SHADER, + fragmentShader: FRAGMENT_SHADER, + glslVersion: THREE.GLSL3, + defines: mode === 'isotropic' ? { ISOTROPIC_SPLATS: '' } : {}, + uniforms: { + splatData: { value: null }, + splatTextureSize: { value: new THREE.Vector2(1, 1) }, + viewport: { value: new THREE.Vector2(1, 1) }, + splatScale: { value: 1.0 } + }, + transparent: true, + // 3DGS blending: the ordering carries the depth information, so + // depth test and write must both be off or splats hard-clip each + // other instead of blending. + depthTest: false, + depthWrite: false, + blending: THREE.CustomBlending, + blendSrc: THREE.OneFactor, + blendDst: THREE.OneMinusSrcAlphaFactor, + blendSrcAlpha: THREE.OneFactor, + blendDstAlpha: THREE.OneMinusSrcAlphaFactor + }); + + this.mesh = new THREE.Mesh(this.geometry, this.material); + this.mesh.frustumCulled = false; // bounds change every frame + // Hidden until setFrame supplies data. Not cosmetic: a visible mesh is + // bound by the render loop, and being bound at a placeholder capacity + // is what caps the instance count forever (see _allocate). + this.mesh.visible = false; + this._allocate(Math.max(capacity, 1)); + + // Sort scratch, reused across frames. + this._depths = new Float32Array(0); + this._counts = new Uint32Array(SORT_BUCKETS); + this._sorted = new Float32Array(0); + } + + /** + * Grow the data texture and the sort index to hold `capacity` splats. + * + * Growing AFTER the geometry has been drawn once is a trap. Three caches + * `_maxInstanceCount` from the instanced attributes the first time it sets + * up the vertex bindings, and the draw count is + * `min(geometry.instanceCount, _maxInstanceCount)`. Swapping in a larger + * `splatIndex` attribute later does not always reset that cache, so a + * geometry first bound at capacity 1 keeps drawing ONE instance no matter + * what `instanceCount` says -- two triangles instead of two hundred + * thousand, which looks exactly like an empty canvas. + * + * Measured: identical draw calls, 213,224 triangles when the growth + * happened before the first bind and 4 when it happened after, varying + * per page load. Callers therefore pass the real capacity up front (see + * `VegaClient`), and the mesh stays hidden until it has a frame so it + * cannot be bound at a placeholder size. + */ + _allocate(capacity) { + if (capacity <= this.capacity) return; + if (this._bound) { + // Not reachable when the caller sized this correctly, and a loud + // failure beats silently rendering one splat out of sixty thousand. + throw new Error( + `SplatObject grew from ${this.capacity} to ${capacity} after it ` + + 'was first drawn; construct it with the maximum splat count ' + + 'for the sequence instead'); + } + this.capacity = capacity; + const texels = capacity * TEXELS_PER_SPLAT; + const height = Math.ceil(texels / TEXTURE_WIDTH); + this._textureData = new Float32Array(TEXTURE_WIDTH * height * 4); + this._texture?.dispose(); + this._texture = new THREE.DataTexture( + this._textureData, TEXTURE_WIDTH, height, + THREE.RGBAFormat, THREE.FloatType); + this._texture.needsUpdate = true; + this.material.uniforms.splatData.value = this._texture; + this.material.uniforms.splatTextureSize.value.set(TEXTURE_WIDTH, height); + + this._indexAttribute = new THREE.InstancedBufferAttribute( + new Float32Array(capacity), 1); + this._indexAttribute.setUsage(THREE.DynamicDrawUsage); + this.geometry.setAttribute('splatIndex', this._indexAttribute); + this._depths = new Float32Array(capacity); + this._sorted = new Float32Array(capacity); + } + + /** Upload a decoded VGS frame. */ + setFrame(frame) { + this._allocate(frame.count); + this.count = frame.count; + const data = this._textureData; + for (let i = 0; i < frame.count; i++) { + const base = i * TEXELS_PER_SPLAT * 4; + data[base] = frame.positions[i * 3]; + data[base + 1] = frame.positions[i * 3 + 1]; + data[base + 2] = frame.positions[i * 3 + 2]; + data[base + 3] = frame.opacities[i]; + + data[base + 4] = frame.scales[i * 3]; + data[base + 5] = frame.scales[i * 3 + 1]; + data[base + 6] = frame.scales[i * 3 + 2]; + data[base + 7] = 0; + + data[base + 8] = frame.rotations[i * 4]; + data[base + 9] = frame.rotations[i * 4 + 1]; + data[base + 10] = frame.rotations[i * 4 + 2]; + data[base + 11] = frame.rotations[i * 4 + 3]; + + data[base + 12] = frame.colors[i * 3]; + data[base + 13] = frame.colors[i * 3 + 1]; + data[base + 14] = frame.colors[i * 3 + 2]; + data[base + 15] = 0; + } + this._texture.needsUpdate = true; + this._positions = frame.positions; + this.geometry.instanceCount = frame.count; + this._sortedForKey = null; + this.mesh.visible = frame.count > 0; + // From here the geometry may be bound at this capacity, so growth is + // no longer safe. + this._bound = this._bound || frame.count > 0; + } + + /** + * Order splats back to front for the given camera. + * + * Counting sort over depth quantised to 16 bits. Re-sorting only when the + * camera or the frame actually changed keeps a static view free. + */ + sort(camera) { + if (this.count === 0 || !this._positions) return false; + const matrix = camera.matrixWorldInverse.elements; + // Third row of the view matrix gives camera-space z directly. + const m2 = matrix[2], m6 = matrix[6], m10 = matrix[10], m14 = matrix[14]; + const key = `${m2.toFixed(5)},${m6.toFixed(5)},${m10.toFixed(5)},` + + `${m14.toFixed(3)},${this.count}`; + if (this._sortedForKey === key) return false; + this._sortedForKey = key; + + const depths = this._depths; + let min = Infinity; + let max = -Infinity; + for (let i = 0; i < this.count; i++) { + const z = m2 * this._positions[i * 3] + + m6 * this._positions[i * 3 + 1] + + m10 * this._positions[i * 3 + 2] + m14; + depths[i] = z; + if (z < min) min = z; + if (z > max) max = z; + } + const span = max - min || 1; + const counts = this._counts.fill(0); + const scale = (SORT_BUCKETS - 1) / span; + // Camera-space z is negative in front of the camera, so ASCENDING z is + // far to near — exactly the back-to-front order the blend needs. + for (let i = 0; i < this.count; i++) { + counts[((depths[i] - min) * scale) | 0]++; + } + let running = 0; + for (let bucket = 0; bucket < SORT_BUCKETS; bucket++) { + const value = counts[bucket]; + counts[bucket] = running; + running += value; + } + const order = this._indexAttribute.array; + for (let i = 0; i < this.count; i++) { + order[counts[((depths[i] - min) * scale) | 0]++] = i; + } + this._indexAttribute.needsUpdate = true; + this._indexAttribute.updateRanges = [{ start: 0, count: this.count }]; + return true; + } + + dispose() { + this.geometry.dispose(); + this.material.dispose(); + this._texture?.dispose(); + } +} + +module.exports = { + SplatObject, VERTEX_SHADER, FRAGMENT_SHADER, + TEXELS_PER_SPLAT, TEXTURE_WIDTH, SORT_BUCKETS +}; diff --git a/open4d/webclients/system/WebClient/src/texture-decoder.js b/open4d/webclients/system/WebClient/src/texture-decoder.js new file mode 100644 index 00000000..c72da281 --- /dev/null +++ b/open4d/webclients/system/WebClient/src/texture-decoder.js @@ -0,0 +1,208 @@ +'use strict'; + +/** + * Texture clip decode: MP4 demux (mp4box) + WebCodecs VideoDecoder. + * + * WebCodecs rather than an HTMLVideoElement on purpose. Geometry arrives as one + * Draco mesh per frame, and a presented frame must pair frame N of the texture + * with frame N of the geometry. A `