Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
62 commits
Select commit Hold shift + click to select a range
4ba9649
chore(openhuman-core): update tinyhumans-sdk vendor reference
senamakel Sep 23, 2026
3c17e75
chore(openhuman-core): register harness tools for tinyagents
senamakel Sep 23, 2026
7219a5d
fix(agent): handle empty agent list in tinyagents module
senamakel Sep 23, 2026
02a043d
chore(agent): remove unused tinyagents module
senamakel Sep 23, 2026
fd981b5
chore(openhuman-core): remove unused toolpack tools
senamakel Sep 23, 2026
f7d77ee
fix(use_skill_dispatch): restore skill dispatch for tiny agents
senamakel Sep 23, 2026
7029f83
chore(deps): bump tinyhumans-sdk submodule
senamakel Sep 23, 2026
bd7987e
fix(harness): register tools from harness config
senamakel Sep 23, 2026
738b9a1
chore(deps): update vendor submodules
senamakel Sep 23, 2026
7f80026
chore(deps): update tinyhumans-sdk submodule
senamakel Sep 23, 2026
9de3163
chore(deps): update tinyhumans-sdk submodule
senamakel Sep 23, 2026
fd9f0bd
fix(agent): update tinyagents vendor and fix skill dispatch tests
senamakel Sep 23, 2026
0c36291
chore(deps): update vendored submodules
senamakel Sep 23, 2026
a950c09
chore(openhuman-core): update model id schema for tinyhumans sdk
senamakel Sep 23, 2026
10d179c
test(agent): add tests for skill dispatch in tinyagents
senamakel Sep 23, 2026
29bceca
chore: add use_skill_dispatch tests
senamakel Sep 23, 2026
b90c665
chore(deps): update vendored submodules
senamakel Sep 23, 2026
54d0e37
chore: update tinyagents vendor and tier fallback logic
senamakel Sep 23, 2026
4ce8879
fix(config): remove unused import in schema types
senamakel Sep 23, 2026
406cda7
fix(config): remove unused import
senamakel Sep 23, 2026
19d30ff
feat(provider): add tier-based provider selection for inference
senamakel Sep 23, 2026
787e943
chore(vision-agent): add missing agent metadata
senamakel Sep 23, 2026
23722de
chore(agent): add agent.toml for image agent
senamakel Sep 23, 2026
024f318
feat(agent): add video agent configuration
senamakel Sep 23, 2026
4b28710
docs(image_agent): update prompt to reflect new default model
senamakel Sep 23, 2026
04ffa2d
docs(video-agent): update prompt to reference OpenRouter model
senamakel Sep 23, 2026
e221fa6
fix(model_context): restore removed test coverage
senamakel Sep 23, 2026
be515f2
fix(test): update model context tests for new inference behavior
senamakel Sep 23, 2026
7f3c8c4
chore(deps): add tinyagents vendor dependency
senamakel Sep 23, 2026
c960b4d
chore: add loader tests for specialist agents
senamakel Sep 23, 2026
a70d5b6
fix(media): handle missing provider gracefully in generation
senamakel Sep 23, 2026
cb00c83
fix(media): handle empty tool output in media generation
senamakel Sep 23, 2026
f0fe327
chore: files changed crates/openhuman-core/src/media/generation/mod.rs
senamakel Sep 23, 2026
e406968
refactor(media): remove local download logic and inline types in favo…
senamakel Sep 23, 2026
db8c0cf
chore(tests): remove unused test module
senamakel Sep 23, 2026
fe61b3a
chore(deps): add tinyinference image and video crates to lockfile
senamakel Sep 23, 2026
3b4ed18
feat(agent): add orchestrator agent configuration
senamakel Sep 23, 2026
4c59974
chore: fix test module name for builtin registration tests
senamakel Sep 24, 2026
aab1bc2
feat(media): make video wait policy configurable in media tools
senamakel Sep 24, 2026
a4dc335
fix(tests): update e2e test to match new media generation output
senamakel Sep 24, 2026
70bd753
test(cli): register end-to-end media generation test
senamakel Sep 24, 2026
44a5a79
feat(scripts): add media route to mock API
senamakel Sep 24, 2026
c4d51b9
fix(media): correct mock API route for media retrieval
senamakel Sep 24, 2026
046d222
feat(scripts): add media route handler to mock API server
senamakel Sep 24, 2026
332f8f5
style(mock-api): reformat long lines for readability
senamakel Sep 24, 2026
4a365d9
chore(deps): update tinyagents submodule
senamakel Sep 24, 2026
623bc7a
fix: reformat long method chains for readability
senamakel Sep 24, 2026
55c40e7
chore(deps): update tinyagents subproject commit
senamakel Sep 24, 2026
be2d7e9
Merge remote-tracking branch 'upstream/main' into media-openrouter
senamakel Sep 24, 2026
0debc67
feat(openhuman-app): add tinyinference-image and tinyinference-video …
senamakel Sep 24, 2026
032dc47
docs(coverage): add media generation and use_skill live-parent rows
senamakel Sep 24, 2026
a4535ac
chore(deps): update vendor submodules
senamakel Sep 24, 2026
64e5931
chore(deps): update tinyagents submodule
senamakel Sep 24, 2026
146c486
chore(vendor): bump tinyagents and sdk to reviewed heads
senamakel Sep 24, 2026
58ea338
feat(media): update image and video agent prompts for new API parameters
senamakel Sep 24, 2026
907e437
chore(vendor): bump tinyagents to 9a62be32
senamakel Sep 24, 2026
435f890
Merge remote-tracking branch 'upstream/main' into media-openrouter
senamakel Sep 24, 2026
fb00e09
chore(deps): update vendor submodules
senamakel Sep 24, 2026
e45674f
chore(vendor): pin tinyhumans-sdk to the sdk#35 merge commit f1e46de5
senamakel Sep 24, 2026
6a4cafa
chore: resolve tinyagents submodule merge
senamakel Sep 24, 2026
ad41667
chore: reformat long lines and improve error handling in media genera…
senamakel Sep 24, 2026
cf579b8
chore: merge latest main into media-openrouter
senamakel Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

31 changes: 31 additions & 0 deletions crates/openhuman-app/Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

5 changes: 5 additions & 0 deletions crates/openhuman-cli/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -251,6 +251,11 @@ path = "../../tests/mcp_registry_multi_server.rs"
name = "mcp_stdio_integration"
path = "../../tests/mcp_stdio_integration.rs"

[[test]]
name = "media_generation_e2e"
path = "../../tests/media_generation_e2e.rs"
Comment thread
coderabbitai[bot] marked this conversation as resolved.
required-features = ["media"]

[[test]]
name = "memory_roundtrip_e2e"
path = "../../tests/memory_roundtrip_e2e.rs"
Expand Down
15 changes: 8 additions & 7 deletions crates/openhuman-core/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -999,16 +999,17 @@ runtime-node = []
# and removing it from one without the other fails that lane.
contacts = []
# Media-generation + image domains: the `media_generate_*` agent tools
# (image/video via GMI through the backend) and the `openhuman::image` tool
# (image/video via OpenRouter through the backend's
# `/agent-integrations/openrouter` proxy) and the `openhuman::image` tool
# contracts scaffold. Default-ON. Slim builds opt out via
# `--no-default-features --features "<explicit list without media>"`.
# Composes with the runtime `DomainSet::media` flag (#4796).
# NOTE: this gate sheds no exclusive dependencies — media generation is
# backend-proxied (reqwest, shared). It is a surface-only gate (drops the tool
# code + module from the compile), not a dependency-shedding one. There are no
# controllers / stores / subscribers tagged `Media` (agent tools only), and
# `openhuman::image` is currently unwired scaffold (added #2997).
media = []
# Enables `tinyagents-harness/media`, which pulls in `tinyinference-image` and
# `tinyinference-video` (the wire contract, job loop and generic tools); both
# are light (serde + the already-shared reqwest), so this remains mostly a
# surface gate. There are no controllers / stores / subscribers tagged `Media`
# (agent tools only), and `openhuman::image` is currently unwired scaffold.
media = ["tinyagents-harness/media"]
# Flows domains: the `flows::` automation surface (saved tinyflows graphs —
# create/run/schedule + the workflow_builder / flow_discovery agents), the
# `tinyflows::` adapter seam, and the `rhai_workflows::` language-workflow tool.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -11,10 +11,16 @@ omit_safety_preamble = false
omit_profile = true
omit_memory_md = true

# Multimodal tier so the agent can review the images it generates (and any
# reference images) via the image_info / inline-image path, then iterate.
# Pinned to a dedicated OpenRouter passthrough model rather than the
# deprecated `hint:vision` (regression R4). This exact id is in
# `MANAGED_MULTIMODAL_MODELS`, so it keeps the multimodal tier needed to
# review the images it generates (and any reference images) via the
# image_info / inline-image path, then iterate. Actual image generation goes
# through `media_generate_image`, which defaults to
# `bytedance-seed/seedream-5-0-lite` on OpenRouter — a separate model from
# the one this agent's own turns run on.
[model]
hint = "vision"
exact = "openrouter/qwen/qwen3.7-flash"

[tools]
# media_generate_image submits the generation and returns a saved local path;
Expand Down
Original file line number Diff line number Diff line change
@@ -1,28 +1,32 @@
# Image-generation specialist

You are a focused **image-creation** sub-agent. You turn a delegating agent's
request into one or more finished image files using the hosted GMI image models
(Seedream for text-to-image, SeedEdit for edits). You run on a multimodal model,
so you can look at reference images and at the images you generate.
request into one or more finished image files using a hosted image-generation
model — the default is `bytedance-seed/seedream-5-0-lite` via OpenRouter for
text-to-image and edits, but the catalog offers other supported models too.
You run on a multimodal model, so you can look at reference images and at the
images you generate.

## Your job

- **Create** images from a text prompt (`media_generate_image`).
- **Edit / restyle** a supplied image by passing its URL(s) as `input_images`.
- **Pick the right model** when it matters — call `media_list_models` to see the
catalog (defaults are fine for most requests; `include_upstream` exposes the
full GMI list).
- **Edit / restyle** a supplied image by passing it in `references` (https
URLs, `data:` URLs, or workspace file paths).
- **Pick the right model** when it matters — call `media_list_models`
(`kind: "image"`, optional `search`) to see the catalog. The default suits
most requests.

## How to work

- Write a vivid, specific prompt. Translate a terse request into concrete visual
detail — subject, composition, lighting, style, mood, colour — but stay true
to what was asked. Don't invent requirements the user didn't state.
- Default the model and size unless the task calls for something specific. Use a
`size` like `1024x1024` (square), `1536x1024` (landscape), or `1024x1536`
(portrait) when the aspect ratio matters.
- For edits, pass the source image URL(s) in `input_images` and describe the
change precisely.
- Default the model and shape unless the task calls for something specific. Set
`aspect_ratio` (`1:1`, `16:9`, `9:16`, `4:3`, …) and optionally `resolution`
(`1K`, `2K`, `4K`); use `size` (e.g. `1536x1024`) only when exact pixels
matter. `n` asks for several variants; `seed` makes a result reproducible.
- For edits, pass the source image(s) in `references` and describe the change
precisely.
- Each generation **saves the image to the workspace and returns a local file
path**. Always report that path back so the deck/answer can reference the
concrete artifact. Do not paste raw base64 or invent URLs.
Expand All @@ -34,5 +38,6 @@ so you can look at reference images and at the images you generate.
- Report results to the delegating agent — you are not talking to the end user.
- If a request is unsafe or disallowed, decline rather than attempting a
work-around.
- If generation fails or times out, say so plainly and surface the request id;
don't fabricate a path or claim success.
- If generation fails, say so plainly and surface the request id; don't
fabricate a path or claim success. When the error says the call was billed,
do **not** call again — report it.
Original file line number Diff line number Diff line change
Expand Up @@ -224,12 +224,17 @@ fn every_builtin_is_stamped_builtin_source() {
}

#[test]
fn vision_agent_loads_on_vision_hint() {
// The vision sub-agent rides the multimodal `vision-v1` tier (via the
// `vision` hint) so its model is image-capable, and it must be reachable
// from the orchestrator's subagent allowlist.
fn vision_agent_loads_on_its_pinned_multimodal_model() {
// The vision sub-agent used to ride the multimodal `vision-v1` tier (via
// the `vision` hint), which is now deprecated — `vision-v1` silently
// falls back to the chat default on managed routes (regression R4). It
// is pinned to a dedicated OpenRouter passthrough model instead, and
// must remain reachable from the orchestrator's subagent allowlist.
let def = find("vision_agent");
assert!(matches!(def.model, ModelSpec::Hint(ref h) if h == "vision"));
assert!(matches!(
def.model,
ModelSpec::Exact(ref m) if m == crate::config::MODEL_MEDIA_UNDERSTANDING
));

let orchestrator = find("orchestrator");
assert!(
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -528,6 +528,44 @@ fn code_executor_has_curl_for_artifact_downloads() {
}
}

/// R4 regression: `hint:vision` is deprecated (`vision-v1` silently falls
/// back to the chat default on managed routes, with no error), so no
/// built-in agent may still declare `ModelSpec::Hint("vision")`.
#[test]
fn no_builtin_agent_declares_the_deprecated_vision_hint() {
for def in load_builtins().expect("built-ins load") {
assert!(
!matches!(&def.model, ModelSpec::Hint(h) if h == "vision"),
"`{}` still declares the deprecated `hint:vision` — pin an exact model instead",
def.id
);
}
}

/// The three media agents are pinned to their dedicated OpenRouter
/// passthrough models (regression R4), not left on `Inherit` or a `Hint`.
#[test]
fn media_agents_are_pinned_to_their_exact_models() {
use crate::config::{
MODEL_IMAGE_GENERATION_AGENT, MODEL_MEDIA_UNDERSTANDING, MODEL_VIDEO_GENERATION_AGENT,
};

for (agent_id, expected_model) in [
("vision_agent", MODEL_MEDIA_UNDERSTANDING),
("image_agent", MODEL_IMAGE_GENERATION_AGENT),
("video_agent", MODEL_VIDEO_GENERATION_AGENT),
] {
let def = find(agent_id);
match &def.model {
ModelSpec::Exact(model) => assert_eq!(
model, expected_model,
"{agent_id} must be pinned to `{expected_model}`, got `{model}`"
),
other => panic!("{agent_id} must use ModelSpec::Exact, got {other:?}"),
}
}
}

#[test]
fn orchestrator_does_not_get_curl() {
// Per design: curl is a `Write` permission tool that writes
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -121,8 +121,9 @@ allowlist = [
# the memory tree only when a message needs it, not before every turn.
"agent_memory",
# Image-understanding specialist. Route anything that hinges on the content
# of an attached or on-disk user-provided image file here — it rides the
# multimodal `hint:vision` tier, so it can actually see the image.
# of an attached or on-disk user-provided image file here — it is pinned to
# a dedicated multimodal model (`MODEL_MEDIA_UNDERSTANDING`), so it can
# actually see the image.
"vision_agent",
# Image-generation specialist. Synthesised into a `delegate_create_image`
# tool. Route make/generate/edit an image requests here — it owns prompt
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -11,10 +11,15 @@ omit_safety_preamble = false
omit_profile = true
omit_memory_md = true

# Multimodal tier so the agent can inspect a reference/first-frame image or the
# returned thumbnail when shaping an image-to-video request.
# Pinned to a dedicated OpenRouter passthrough model rather than the
# deprecated `hint:vision` (regression R4). This exact id is in
# `MANAGED_MULTIMODAL_MODELS`, so it keeps the multimodal tier needed to
# inspect a reference/first-frame image or the returned thumbnail when
# shaping an image-to-video request. Actual video generation goes through
# `media_generate_video`, which defaults to `bytedance/seedance-2.0-mini` on
# OpenRouter — a separate model from the one this agent's own turns run on.
[model]
hint = "vision"
exact = "openrouter/qwen/qwen3.7-flash"

[tools]
# media_generate_video submits the generation and returns a saved local path;
Expand Down
Original file line number Diff line number Diff line change
@@ -1,27 +1,33 @@
# Video-generation specialist

You are a focused **video-creation** sub-agent. You turn a delegating agent's
request into a finished video clip using the hosted GMI video models (Seedance
for fast clips, Veo for premium-tier output). You can do text-to-video or
animate a supplied first-frame/reference image (image-to-video).
request into a finished video clip using a hosted video-generation model — the
default is `bytedance/seedance-2.0-mini` via OpenRouter for fast clips, with
premium-tier models available in the catalog for higher-quality output. You
can do text-to-video or animate a supplied first-frame/reference image
(image-to-video).

## Your job

- **Create** a clip from a text prompt (`media_generate_video`).
- **Animate** a supplied image by passing its URL as `input_image`.
- **Pick the right model** when it matters — call `media_list_models` to see the
catalog (the fast Seedance default suits most requests; `include_upstream`
exposes the full GMI list, including premium tiers).
- **Animate** a supplied image by passing it as `first_frame` (and optionally
`last_frame`) — an https URL, `data:` URL, or workspace file path.
- **Pick the right model** when it matters — call `media_list_models`
(`kind: "video"`, optional `search`) to see the catalog. The fast default
suits most requests.

## How to work

- Write a concrete prompt describing the motion, subject, and scene — what
happens over the clip, not just a static description. Mention camera movement,
pacing, and style when relevant.
- Use `duration_seconds` and `aspect_ratio` (e.g. `16:9`, `9:16`, `1:1`) when the
task specifies them; otherwise let the model default.
- For image-to-video, pass the source image URL in `input_image` and describe
the motion you want applied to it.
- Use `duration` (seconds; the default model accepts 4–15), `aspect_ratio`
(e.g. `16:9`, `9:16`, `1:1`) and `resolution` (`480p`, `720p`) when the task
specifies them; otherwise let the model default. `generate_audio` adds a
soundtrack where supported.
- For image-to-video, pass the source image in `first_frame` and describe the
motion you want applied to it. `references` guide subject or style without
fixing a frame.
- Generation is **asynchronous and can take minutes** — the tool blocks until the
clip is ready, saves it to the workspace, and returns a local file path. Report
that path back. Set expectations: tell the delegating agent it may take a
Expand All @@ -34,5 +40,7 @@ animate a supplied first-frame/reference image (image-to-video).
- Report results to the delegating agent — you are not talking to the end user.
- If a request is unsafe or disallowed, decline rather than attempting a
work-around.
- If generation fails or times out, say so plainly and surface the request id;
don't fabricate a path or claim success.
- If generation fails, say so plainly and surface the job id; don't fabricate a
path or claim success. If it **times out**, call the tool again with
`resume_job_id` set to that job id to collect the clip — never submit a new
job for the same request, since each submit is billed.
Original file line number Diff line number Diff line change
Expand Up @@ -10,12 +10,14 @@ omit_identity = true
omit_memory_context = true
omit_safety_preamble = true

# Multimodal tier. `ModelSpec::Hint("vision")` resolves to `hint:vision`, which
# `oh_tier_supports_vision` reports as vision-capable — so this sub-agent's
# model is always treated as image-enabled (managed or BYOK), and the turn
# engine never strips the attached image at the vision gate.
# Pinned to a dedicated OpenRouter passthrough model rather than the
# deprecated `hint:vision` (`vision-v1` silently falls back to the chat
# default on managed routes, with no error — regression R4). This exact id
# is in `MANAGED_MULTIMODAL_MODELS`, so `oh_tier_supports_vision` still
# reports it as image-enabled and the turn engine never strips the attached
# image at the vision gate.
[model]
hint = "vision"
exact = "openrouter/qwen/qwen3.7-flash"

[tools]
# Attached images arrive inline in the sub-agent's context via the multimodal
Expand Down
Loading
Loading