Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,7 @@ Reference pages define stable inputs, outputs, configuration, and stored-data
contracts.

- [Embedded integration boundaries](./EMBEDDED_BOUNDARIES.md)
- [Local video preparation and measurements](./VIDEO_PRIMITIVES.md)
- [Canonical episode format](./FORMAT.md)
- [Catalog tables and curation API](./CATALOG.md)
- [Environment variables](./ENVIRONMENT.md)
Expand Down
72 changes: 72 additions & 0 deletions docs/VIDEO_PRIMITIVES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# Local video preparation and measurements

These file-level APIs do not require a catalog, scheduler, model service, or
persistent workspace. Callers retain source identity and own output lifetimes.
FFmpeg-based APIs use HFlow's binary policy and may download managed binaries
unless the caller configures an installed toolchain.

## MCAP camera export

`hflow.mcap_video.export_mcap_camera(source, output, camera_topic=None, limits=...)`
requires the optional `hflow[video]` dependency. It streams one camera from an MCAP
into MP4, returning `PreparedMcapVideo` with the output path, selected topic,
original first log timestamp, and duration in milliseconds. Multiple cameras
require an exact topic. Existing destinations are never overwritten.

Frame times are MCAP log times relative to the first selected frame, rounded to
microseconds. Each frame lasts until the next; the final frame repeats the prior
positive interval. At least two frames are required. Duplicate rounded timestamps,
unsupported formats, changing dimensions, B-frames, missing H.264 codec headers,
and gaps beyond the MP4 duration representation are rejected. H.264 access units
are remuxed with lossless AUD repair; supported JPEG/PNG and 8-bit raw images are
encoded as lossless RGB H.264. Raw recordings remain unchanged.

This is synchronous native decoding. A caller needing a hard execution deadline
must isolate it in a process and reap that process before deleting temporary media.
The existing canonical `camera_video` enrichment keeps its constant-rate contract.

## Direct model-video preparation

`hflow.importers.video.prepare_model_video(source, output, config, limits=...,
transform_config=...)` uses `VideoImportConfig` and `TransformConfig` to produce
canonical model-input pixels without first writing an MCAP. It shares the video
importer's fixed-rate JPEG rendering and canonical H.264 encoding, including the
single-frame cadence. Output is an atomically published caller-owned MP4.

Expected media failures return `UnreadableVideo` or `UnsupportedVideo`; operational
failures raise. Existing destinations raise `FileExistsError`. The JPEG intermediate
is intentional: bypassing it would change model-input pixels. This helper does not
change canonical transformation defaults or identities.

## Frame statistics

`hflow.video_statistics.measure_video_frame_statistics(video, settings=...,
toolchain=None, instrument_output_cache_path=None)` exposes file-level statistics,
settings, provenance, and errors as a public API. No persistent cache is created
unless explicitly requested.

`FrameStatisticsSettings.luma_range` accepts `LumaRangePolicy.PRESERVE` (the existing
instrument behavior) or `FULL` (convert the declared input range to full-range luma
inside the measurement filter graph). The graph and settings are recorded in
provenance. Measure unpadded source framing; black model-input borders would affect
pixel statistics. Full-range measurement needs no intermediate video encode.

## Raw blur and shake summaries

`hflow.blur.measure_video_blur(video, timeout_seconds=120, executable=None)` uses
FFmpeg's default `blurdetect` settings at the supplied frame cadence and resolution.
`BlurSummary` contains observed/scored frame counts and the mean finite raw score.
Nonfinite scores are unassessed, zero assessed frames produce a null mean, and
negative scores fail. `summarize_blur_scores` applies the same accounting to a
caller-supplied score stream.

`hflow.camera_motion.summarize_camera_shake(observations)` consumes filtered motion
observations in constant summary memory. Mean and RMS weight adjacent-pair duration;
maximum, assessment duration, and missing-context counts remain explicit. The final
frame has no following pair and contributes no extra duration. Missing observations
never become zero shake. Field of view is caller configuration, not calibration.

These measurements do not classify footage as blurry or unstable and are not
accuracy estimates. See [streaming camera motion](how-to/stream-camera-motion.md)
for extraction and filtering contracts, and [source sampling](how-to/sample-source-video.md)
for original-frame evidence selection.
19 changes: 19 additions & 0 deletions docs/how-to/sample-source-video.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,3 +96,22 @@ and selected FFmpeg build in their version or input contract. Changing selection
can change measured results. This is a new original-source API; existing
`Episode.frames()` and Build AI `FrameSampling` keep their declared frame-rate
behavior. No canonical transform version changes are needed to use these helpers.

## Nearest keyframes and unpadded evidence

`SourceSamplingMode.NEAREST_KEYFRAMES` selects unique codec keyframes nearest
caller-supplied `keyframe_positions`, expressed as fractions of the requested
window. Ties choose the earlier frame; results are chronological. Empty windows
remain empty and this mode has no uniform fallback. Position count must fit
`maximum_frames`.

Import `SourceFrameResize` from `hflow.source_sampling`. Set `resize=SourceFrameResize.FIT`
to fit within `width` and `height` without adding padding; the default `PAD` retains
the existing padded canvas. `scaling_algorithm` accepts `lanczos` (default) or
`bicubic`, and `jpeg_quality` accepts FFmpeg quality values 1 through 31 (default 5).
For example, a caller can choose positions `(0.15, 0.5, 0.85)`, a 960×960 fit box,
and JPEG quality 2. These are evidence settings, not a classification policy.

Nearest-keyframe probing shares the extraction deadline and is bounded by
`maximum_probe_bytes`. Returned timestamps retain the source's playback clock and
rational time base; subtract the window start explicitly for relative timestamps.
2 changes: 2 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,7 @@ mediapipe = ["mediapipe>=1.0.1"]
# dependency is real, and confined to the one check that needs it. Headless
# because a data pipeline runs where there is no display.
motion = ["opencv-python-headless>=5.0.0.93"]
video = ["av>=18.1.0"]
native-build = ["Cython>=3.3.0", "setuptools>=84.0.0"]
openai = ["openai>=3.3.1"]

Expand All @@ -67,6 +68,7 @@ hflow = "hflow.cli:main"

[dependency-groups]
dev = [
"av>=18.1.0", # the video extra, exercised by timestamp-preserving camera export tests
"Cython>=3.3.0", # the native-build extra, present so its real build path is tested
"hflow-server", # the workspace API package, present in dev so the suite tests it
"obstore>=0.11.1", # the bucket extra, present in dev so the suite tests it
Expand Down
15 changes: 14 additions & 1 deletion src/hflow/_video_measurements/_frame_statistics.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,13 @@ class UnsupportedVideoMeasurementToolchainError(RuntimeError):
"""The supplied FFmpeg build lacks a filter required by the measurement."""


class LumaRangePolicy(StrEnum):
"""Whether measurements retain decoded luma or normalize its declared range."""

PRESERVE = "preserve"
FULL = "full"


class LumaRangeEvidence(StrEnum):
"""What decoded luma samples show about the nominal limited range."""

Expand Down Expand Up @@ -76,8 +83,11 @@ class FrameStatisticsSettings:
freeze_noise_tolerance_decibels: float = -60.0
freeze_minimum_duration_seconds: float = 2.0
overexposed_average_luma_threshold: float = 235.0
luma_range: LumaRangePolicy = LumaRangePolicy.PRESERVE

def __post_init__(self) -> None:
if not isinstance(self.luma_range, LumaRangePolicy):
raise ValueError("luma_range must be a LumaRangePolicy")
require_int(
self.black_frame_minimum_pixel_share_percent,
"black_frame_minimum_pixel_share_percent",
Expand Down Expand Up @@ -595,7 +605,10 @@ def _temporary_instrument_cache_output(

def frame_statistics_filter_graph(settings: FrameStatisticsSettings) -> str:
"""Return the effective single-pass FFmpeg measurement graph."""
return (
normalization = (
"scale=in_range=auto:out_range=full," if settings.luma_range is LumaRangePolicy.FULL else ""
)
return normalization + (
"format=pix_fmts=yuv420p,"
f"blackframe=amount=0:threshold={settings.black_pixel_luma_threshold},"
"freezedetect="
Expand Down
105 changes: 105 additions & 0 deletions src/hflow/blur.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
"""Raw FFmpeg blur scores with explicit finite-score coverage, not quality labels."""

from __future__ import annotations

import math
import subprocess
import tempfile
from collections.abc import Iterable
from dataclasses import dataclass
from pathlib import Path

from hflow.ffmpeg import ffmpeg_path


@dataclass(frozen=True, slots=True)
class BlurSummary:
frame_count: int
scored_frame_count: int
mean_blur_score: float | None


def summarize_blur_scores(scores: Iterable[float]) -> BlurSummary:
"""Average finite raw frame scores; unavailable scores never become zero."""
frame_count = 0
scored_frame_count = 0
mean_blur_score = 0.0
for score in scores:
frame_count += 1
# blurdetect emits NaN when it cannot measure edges, e.g. a flat image.
if not math.isfinite(score):
continue
if score < 0:
raise RuntimeError("Blur analysis returned an invalid score")
scored_frame_count += 1
mean_blur_score += (score - mean_blur_score) / scored_frame_count
return BlurSummary(
frame_count=frame_count,
scored_frame_count=scored_frame_count,
mean_blur_score=mean_blur_score if scored_frame_count else None,
)


def measure_video_blur(
video_path: Path, *, timeout_seconds: float = 120.0, executable: Path | None = None
) -> BlurSummary:
"""Score a local video window at its supplied resolution and frame cadence.

Uses FFmpeg's default blurdetect settings on the first video stream. The
result is the frame-weighted mean of finite raw scores, not a percentage or
probability. No resizing, frame sampling, classification, or model call is
performed. Results contain raw scores and coverage, without affected-time percentages.
"""
if not math.isfinite(timeout_seconds) or timeout_seconds <= 0:
raise ValueError("timeout_seconds must be positive and finite")
try:
with tempfile.TemporaryFile() as metadata_output:
subprocess.run(
[
str(executable or ffmpeg_path()),
"-hide_banner",
"-loglevel",
"error",
"-nostdin",
"-xerror",
"-protocol_whitelist",
"file",
"-noautorotate",
"-threads",
"1",
"-i",
str(video_path.resolve()),
"-map",
"0:v:0",
"-an",
"-sn",
"-dn",
"-filter_threads",
"1",
"-vf",
"blurdetect,metadata=mode=print:key=lavfi.blur:file=-",
"-fps_mode",
"passthrough",
"-f",
"null",
"-",
],
stdin=subprocess.DEVNULL,
stdout=metadata_output,
stderr=subprocess.PIPE,
timeout=timeout_seconds,
check=True,
)
metadata_output.seek(0)
summary = summarize_blur_scores(
float(line.removeprefix(b"lavfi.blur="))
for line in metadata_output
if line.startswith(b"lavfi.blur=")
)
metadata_output.seek(0)
emitted_frame_count = sum(line.startswith(b"frame:") for line in metadata_output)
if summary.frame_count != emitted_frame_count:
raise RuntimeError("Blur analysis returned incomplete frame scores")
return summary
except Exception as error:
raise RuntimeError("Blur analysis failed") from error
109 changes: 108 additions & 1 deletion src/hflow/camera_motion.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,9 +5,12 @@
episodes, catalogs, or orchestration. See docs/how-to/stream-camera-motion.md.
"""

from collections.abc import Iterator
import math
from collections.abc import Iterable, Iterator
from contextlib import contextmanager
from dataclasses import dataclass
from pathlib import Path
from typing import assert_never

from hflow._video_measurement_toolchain import resolved_video_measurement_toolchain
from hflow._video_measurements._motion_fit import (
Expand Down Expand Up @@ -41,6 +44,7 @@
"CameraMotionTransform",
"CameraShakeObservation",
"CameraShakeSettings",
"CameraShakeSummary",
"MeasuredCameraMotion",
"MeasuredCameraShake",
"MotionFitEvidence",
Expand All @@ -50,6 +54,7 @@
"filter_camera_shake",
"iter_frame_motion",
"stream_camera_motion",
"summarize_camera_shake",
]


Expand All @@ -76,3 +81,105 @@ def stream_camera_motion(
yield observations
finally:
observations.close()


@dataclass(frozen=True, slots=True)
class CameraShakeSummary:
"""Continuous residual rates and their observed frame-pair coverage.

Angular rates use the caller's approximate field of view. A fitted motion
estimate can still be unreliable; coverage does not establish accuracy.
The final frame's display duration is outside the observed pair intervals.
"""

pair_count: int
measured_motion_pair_count: int
measured_shake_pair_count: int
insufficient_context_pair_count: int
unmeasured_context_pair_count: int
observed_seconds: float
assessed_seconds: float
mean_shake_degrees_per_second: float | None
rms_shake_degrees_per_second: float | None
maximum_shake_degrees_per_second: float | None

@property
def unassessed_seconds(self) -> float:
return self.observed_seconds - self.assessed_seconds

@property
def assessed_fraction(self) -> float | None:
if self.observed_seconds == 0:
return None
return self.assessed_seconds / self.observed_seconds


def summarize_camera_shake(observations: Iterable[CameraShakeObservation]) -> CameraShakeSummary:
"""Reduce a complete filtered motion stream with constant summary memory.

Missing motion and filter context stay unassessed, never zero shake. Rates
are weighted by the adjacent-pair durations; no video frames or rate history
are retained. Ordering and cadence belong to HFlow's filter contract.
"""
pair_count = 0
measured_motion_pair_count = 0
measured_shake_pair_count = 0
insufficient_context_pair_count = 0
unmeasured_context_pair_count = 0
observed_seconds = 0.0
assessed_seconds = 0.0
mean_shake_degrees_per_second = 0.0
rms_shake_degrees_per_second = 0.0
maximum_shake_degrees_per_second = 0.0

for observation in observations:
pair_count += 1
pair_seconds = observation.motion.end_seconds - observation.motion.start_seconds
observed_seconds += pair_seconds
if isinstance(observation.motion.measurement, MeasuredCameraMotion):
measured_motion_pair_count += 1

shake = observation.shake
if isinstance(shake, MeasuredCameraShake):
measured_shake_pair_count += 1
assessed_seconds += pair_seconds
shake_degrees_per_second = shake.residual.magnitude_degrees_per_second
duration_share = pair_seconds / assessed_seconds
mean_shake_degrees_per_second += duration_share * (
shake_degrees_per_second - mean_shake_degrees_per_second
)
# The weighted Euclidean norm avoids squaring large finite rates.
rms_shake_degrees_per_second = math.hypot(
rms_shake_degrees_per_second * math.sqrt(1 - duration_share),
shake_degrees_per_second * math.sqrt(duration_share),
)
maximum_shake_degrees_per_second = max(
maximum_shake_degrees_per_second, shake_degrees_per_second
)
else:
match shake.reason:
case "insufficient_context":
insufficient_context_pair_count += 1
case "unmeasured_context":
unmeasured_context_pair_count += 1
case unknown_reason:
assert_never(unknown_reason)

return CameraShakeSummary(
pair_count=pair_count,
measured_motion_pair_count=measured_motion_pair_count,
measured_shake_pair_count=measured_shake_pair_count,
insufficient_context_pair_count=insufficient_context_pair_count,
unmeasured_context_pair_count=unmeasured_context_pair_count,
observed_seconds=observed_seconds,
assessed_seconds=assessed_seconds,
mean_shake_degrees_per_second=(
mean_shake_degrees_per_second if measured_shake_pair_count else None
),
rms_shake_degrees_per_second=(
rms_shake_degrees_per_second if measured_shake_pair_count else None
),
maximum_shake_degrees_per_second=(
maximum_shake_degrees_per_second if measured_shake_pair_count else None
),
)
Loading
Loading