Skip to content

Discussion: what is the intended shape of the media-type taxonomy? #95

Description

@mlcooper

Summary

ARD's four defined media types (§3.3) describe network-reachable services and one document kind:

Type Artifact
application/mcp-server-card+json a server you connect to
application/a2a-agent-card+json an agent service you call
application/ai-registry+json a registry you federate with
application/ai-skill+md a markdown document you load

Cataloging a real internal estate, I ran into three categories that don't fit, in different ways:

  1. Locally-executed command-line tools — no type applies at all.
  2. Subagent definitions — a type fits structurally but is too coarse to say how the artifact is consumed.
  3. Curated collections — no type applies, and separately, no field on an entry can express what the collection contains.

Since type is a MUST on every entry (§4.2), the first and third cases leave a publisher with no conformant option but omission. I'd like to discuss all three, and the broader question underneath them.

Case 1: command-line tools

Agents with shell access — Claude Code, Cursor, Codex, Aider, and others — invoke CLIs directly. For many teams this is the preferred integration shape rather than a legacy one:

  • No server lifecycle. Nothing to start, supervise, health-check, or reconnect to. The tool runs, returns, exits.
  • No protocol overhead. A CLI emitting structured JSON is directly consumable; there is no transport to negotiate.
  • One interface for humans, agents, scripts, and CI. The same binary serves an engineer at a terminal, an agent in a loop, and a pipeline step. A server serves only the agent.
  • Auth reuses existing local credential stores rather than a separate server-side path.
  • Distribution is already solved by the artifact registries and package managers teams run today.

CLIs are the largest category of tooling we want agents to discover, and the direct answer to the question users actually ask: "what do we have that could help with this?" Today there is no way to express one.

What a publisher can do today

Catalog the tool's documentation as application/ai-skill+md. Conformant, and useful where such a document exists — an agent reads it and learns the tool exists. But it catalogs the documentation about a capability rather than the capability, requires every tool to have an accompanying document, and gives the consumer no structured signal that an executable exists, what it is called, or how to obtain it. In our own catalog that relationship isn't expressible as data at all — it's a naming convention with real exceptions, so the mapping can't be generated reliably. This is what we chose for v1, and it leaves roughly a third of our tools uncatalogable.

Use a vendor-prefixed type (application/vnd.example.cli-tool+json). Legitimate IANA practice, works inside one organization, and invisible to every consumer that doesn't know the prefix — forfeiting exactly the federation ARD exists to provide.

Omit them. Leaves the largest category of capability undiscoverable.

Mislabel as a2a-agent-card+json. Worth naming to rule out. type is a protocol contract telling the consumer how to consume; labeling a local executable as an A2A agent promises a callable endpoint that doesn't exist, so the consumer fails rather than degrades.

Case 2: subagent definitions

We also publish subagent definitions: markdown documents with frontmatter that configure a separate agent instance — its own context window, its own tool permissions, sometimes its own model — which the host delegates a task to and collects a result from.

To be precise, and unlike case 1: a subagent definition genuinely is a loadable markdown document, so application/ai-skill+md structurally fits. The problem is that it's too coarse to convey consumption semantics.

Skill Subagent definition
Effect of loading augments the current agent's context instantiates a separate agent
Context shared isolated
Lifecycle lasts the conversation bounded task, returns a result
Tool access the current agent's often deliberately narrowed

A consumer receiving application/ai-skill+md cannot tell whether loading the artifact reshapes its own context or spawns a delegate. Those need different handling, and on a host without delegation support the second isn't usable at all — which the consumer should be able to determine from type rather than by parsing the document.

Case 3: curated collections

We also publish named collections: one action installs the eight or ten artifacts somebody in a given role actually needs. This is the most useful discovery answer we have, because the honest response to "based on my problem, what should I install?" is usually one collection rather than eight separate entries.

No type applies. A collection is not a service you call and not a document you load — nothing loads it; an installer expands it into other artifacts. The nearest-looking fit is application/ai-registry+json, and it's wrong in the specific way case 1 rules out a2a-agent-card+json: that type promises a registry you federate with, and §5.2 has an Agent Registry exposing a Search API. A collection has no endpoint and nothing to query, so the label makes a consumer fail rather than degrade.

But the type is the smaller half of this one. An entry cannot express "this entry is a set of those entries." The entry shape offers identifier, displayName, type, url/data, description, tags, capabilities, representativeQueries and metadata — all of which describe the entry itself. None points at another entry.

So the membership can only live in prose in description. That is the same "catalogs the documentation about a capability rather than the capability" problem as case 1, reached from a different direction — and here it removes the single property that makes the artifact worth discovering, since what one press installs is the whole point.

Worth noting how this compounds with representativeQueries: "what do I need for on-call work?" is a query a collection answers precisely and an individual entry answers badly. The artifact best suited to that field is the one that currently cannot be published.

Unlike cases 1 and 2, I don't think this one can be read as "your estate is unusual." Any registry that groups its offerings — a starter pack, a role-based set, a suite — hits it, and grouping becomes close to universal once a registry passes a few dozen entries. A consumer that could resolve a collection into its members would get better answers from every registry publishing one.

The question underneath all three

Rather than proposing three types, I'd rather ask the shape question, since the answer determines whether these are candidates or out of scope. Some possibilities, and I don't hold a strong view:

  1. The type set is intentionally minimal, and artifacts outside it are out of scope. Reasonable — but worth stating in the spec, since type being a MUST currently leaves publishers with no conformant option but omission.
  2. Types grow case by case as real needs appear. ADR-0008's rename of mcp-server+json suggests the set is actively curated, so these would be two candidates.
  3. type stays coarse and a secondary mechanism carries the detail — a capabilities convention, or a subtype parameter — letting consumers branch without the type list growing without bound.

Option 3 might address cases 1 and 2 at once and keep the taxonomy small, which is why I'd rather settle the shape than keep proposing individual types.

Case 3 is why I'd now separate two questions that the first two cases conflate. Cases 1 and 2 are both about how to consume one artifact, which is exactly what a subtype parameter or capabilities convention can carry. A collection's problem is a claim about other entries, and no per-entry mechanism reaches it — a subtype could say "this is a collection" and a consumer would still have no way to learn what is in it. So option 3 would settle cases 1 and 2 and leave case 3 where CLIs are today: no conformant option but omission. The taxonomy question and the expressiveness question look separable, and only the third case distinguishes them.

If a CLI type is the answer

A strawman, in the same family as the existing *-card+json descriptors — e.g. application/cli-tool-card+json:

  • Invocation name — the command as it appears on PATH
  • Acquisition — package registry, download URL, or installer command, with platform/architecture availability
  • Machine-readable interface discovery — a convention for retrieving the tool's own command/flag schema, analogous to how an MCP server advertises its tools. Many agent-oriented CLIs already expose something in this shape (a --json / --agent mode, or a schema subcommand)
  • Authentication requirements — whether credentials are needed and how they're supplied, consistent with §3.6's delegation of auth to the artifact protocol
  • Structured-output affordance — whether the tool can emit machine-readable output, which is what makes a CLI agent-usable rather than merely human-usable

Field names and shape are illustrative only.

If a collection type is the answer

The relation matters more than the type here, and the two are independent — a membership relation would be useful even if collections ended up sharing an existing type:

  • A membership array — members (or contains), holding identifiers of other entries in the same manifest. This is the part with no workaround.
  • Indexing semantics — whether a consumer is expected to fan a collection out into its members, or keep it as one indexable unit. Both are defensible and they produce very different search results, so leaving it unstated means two conformant consumers disagree on the same manifest.
  • Resolution across manifests — whether a member identifier may name an entry in a federated manifest rather than the local one, and what a consumer does when it cannot resolve one. Partial resolution seems like the common case in a federated setting.

Field names and shape are illustrative only.

Questions

  1. Are these in scope for ARD? The existing types suggest a scope of "callable services and loadable documents." Is excluding locally-executed artifacts deliberate, or a gap because the spec grew from MCP and A2A?
  2. If in scope, is a new media type the right mechanism, or is option 3 above closer to the intended design?
  3. Is a relation between entries in scope at all? Every field in §4.2 describes its own entry, so I may be asking for something deliberately outside the data model. If entry-to-entry references are out of scope, that's a useful answer on its own — it would mean collections are expected to be flattened by the publisher before publishing, with the grouping simply not surviving into ARD.
  4. Is there prior art in the working group I should read? I found ADR-0008 but nothing on executables, delegation artifacts, or entry relations.

Happy to draft a concrete schema and conformance-test cases if there's appetite. Following the contributing guidance, opening as a discussion issue before proposing a normative change.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions