[REA-6855] Save each reported step as a folder of files - #238
Conversation
tempusfrangit
left a comment
There was a problem hiding this comment.
Requesting changes for result integrity: saved video should preserve the emitted media's playout rate, and an existing completion marker must not expose files while they are being rewritten. I also reproduced the existing stale-report message loss and admission/shutdown race. Please address these cases with regression tests. The error-step payload behavior is a separate contract clarification noted inline.
1c8fe61 to
f5bf9c5
Compare
f03b104 to
2333bb0
Compare
f5bf9c5 to
e6ae667
Compare
|
@tempusfrangit all addressed in 2333bb0, each with a regression test: the saved video takes the emission's rate (measured throughput for unpinned models); folders are written under a hidden name and swapped in whole, so a reused session id never exposes a half-rewritten step; error steps keep only |
2333bb0 to
d976a28
Compare
a510431 to
355b7ab
Compare
d976a28 to
2c56ea6
Compare
355b7ab to
d7e067d
Compare
2c56ea6 to
c0645cd
Compare
## Why
Whether each finished step is kept as a folder of files is the model's choice, the same way recording is: it decides what is worth keeping and how to encode it. So it belongs in the model's `reactor.yaml`, next to `runtime.recording`, and not in a session option a client could switch.
## What Changed
The manifest gains a `runtime.step_results` block:
```yaml
runtime:
step_results:
enabled: true
video: {codec: h264, preset: veryfast, crf: 23} # or codec: h265
audio: {codec: aac, bitrate_kbps: 128}
queue: 8
```
It names no tracks: every track of a step's output is kept as its own stream. `queue` bounds how many steps may wait to be saved, so a model that steps faster than its steps encode never waits on saving; a `queue` below 1 fails the load, because an unbounded queue is what that value would otherwise mean. A missing or malformed block leaves step results off at their defaults, and unknown keys are ignored, as for the recording block.
The block parses into a `StepResultsConfig` on `RuntimeConfig`, and the README documents it where the manifest is introduced, so the authoring surface lands with its documentation. Nothing is saved yet: the writer, the store, and the routes follow in the next PRs, and the session descriptor reports `"step_results"` only from #238, where steps are saved, so it never promises folders the runtime does not write.
When a model's manifest turns step results on, every step it reports is kept as a folder under the session's own id: the step's output as one output.mp4, the extra files the model kept with it, and result.json, which lists the files, the messages the model broadcast since the step before, the step's error, and how long it took to generate and to save. result.json is written last, through a rename, so a folder that has one is complete. Saving never makes the model wait. A step is offered to the store when it is reported, and one that finds the queue full (runtime.step_results .queue, eight by default) is not saved; the step_completed fact says which, so saved: true always means a folder follows. A step whose save fails still gets its result.json, with the reason in save_error and no files. One worker saves steps in order and journals step_result_ready with the session, the step, and the file names, after the session has ended too, so the last step of a session that stopped at its steps still arrives. Each folder is deleted five minutes after its result.json is written, so disk use follows the step rate rather than the session's length. On shutdown, pending saves get ten seconds to finish. Each step report now carries the model's playout rate, which the saved video plays at. Messages are taken on the model's thread when the step is reported, so each lands with the step it was sent before; the buffer keeps the latest 256 for a model that broadcasts without reporting steps. Only a session id that is a UUID names a folder. Signed-off-by: Dere-Wah <derexcontact@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
A saved step's video now plays at the rate its media was emitted at. For a model that does not pin its rate the default loop paces each emit by its measured throughput, so the report takes the same rate the emit chose rather than the declared fps; a step reported with complete_step() takes the rate of the model's last emission. A step's folder is written under a hidden name and renamed into place once complete, replacing any folder a reused session id left there, so a reader never sees a result.json beside files still being rewritten. A step that failed keeps only its result.json, as documented, even when its report carried output or files. Admission and close() now share a lock, so a step admitted while the store closes is queued before the worker is told to drain and is saved, and nothing is admitted after. A late report from an earlier session no longer takes the messages the current session has sent, so they land in its next step. The session descriptor reports step_results, now that steps are saved. Signed-off-by: Dere-Wah <derexcontact@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
c0645cd to
f45ade0
Compare

Why
A caller that runs a session for its output, rather than to watch it, needs each step's result kept somewhere it can collect it, including the last step of a session that closed itself at its
steps. Step reports and the manifest block are in place; this is the part that keeps what they describe.What Changed
When
runtime.step_resultsis on, every step the model reports is saved as a folder under the session's own id, the id/clips/chunksuses:{ "step": 1, "session_id": "9f1e2d3c-0000-4000-8000-00000000c0de", "files": [{"name": "output.mp4", "content_type": "video/mp4", "size": 5170}], "messages": [{"type": "progress", "data": {"percent": 50}}], "error": null, "save_error": null, "timings": {"generate_s": 0.012, "encode_s": 0.015} }Each folder is written under a hidden name,
result.jsonlast, and renamed to its step number once complete, replacing any folder a reused session id left there; a reader therefore never sees aresult.jsonbeside files still being written.messagesare the messages the model broadcast since the step before; they are taken on the model's thread when the step is reported, so each lands with the step it was sent before, a late report from an earlier session leaves them to the current one, and the buffer keeps the latest 256 for a model that broadcasts without reporting steps. Each step report now carries the rate its media was emitted at, which the saved video plays at: the measured throughput for a model that does not pin its rate, or the rate of the last emission for a step reported withcomplete_step().Saving never makes the model wait. A step is offered to the store when it is reported, and one that finds the queue full is not saved.
step_completedsays which, sosaved: truealways means a folder follows. A step that failed is saved asresult.jsonalone, whatever output or files its report carried, and so is a step that produced nothing. A step whose save fails still gets itsresult.json, with the reason insave_errorand no files, so an admitted step never leaves a gap. A step reported after the session stopped at itsstepsis saved too.One worker saves steps in order and journals
step_result_readywith{session_id, step, files}. The session id is in the detail because a folder can finish after its session has ended, and the next session may have started. Each folder is deleted five minutes after itsresult.jsonis written, while the session runs and after it ends, so disk use follows the step rate rather than the session's length; a folder still being written is never touched. Admission and shutdown share a lock, so a step admitted while the store closes is saved and nothing is admitted after; on shutdown, pending saves get ten seconds to finish. The session descriptor reports"step_results": {"enabled": ...}from here, now that steps are saved. Only a session id that is a UUID names a folder, so a caller-chosen id cannot point outside the root.