Skip to content

[Platform] Complete structured output instance mode - #2448

Open
marco-jouwweb wants to merge 3 commits into
symfony:mainfrom
marco-jouwweb:platform-response-format-factory-instance
Open

marco-jouwweb wants to merge 3 commits into
symfony:mainfrom
marco-jouwweb:platform-response-format-factory-instance

Conversation

@marco-jouwweb

@marco-jouwweb marco-jouwweb commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor
Q A
Bug fix? no
New feature? yes
Docs? yes
Issues -
License MIT

Problem

response_format accepts an existing object instance, and the Platform uses it as the
serializer's object_to_populate, so the model's answer is written onto that very
instance (Populating Existing Object Instances in the Platform docs).

The schema half of that feature never sees the object.
ResponseFormatFactoryInterface::create() takes only a class-string, so the JSON
schema is always the open-ended shape the class allows, never the shape the instance
actually needs. The model is asked to invent values the instance already holds: wasted
tokens, and an invitation to contradict data the application is sure about.

The feature

A missing_properties_only option. With an instance as response_format, the schema is
built from what that instance still lacks rather than from everything its class allows:

final class City
{
    public function __construct(
        public ?string $name = null,
        public ?int $population = null,
        public ?string $country = null,
        public ?string $mayor = null,
    ) {
    }
}

$city = new City(name: 'Berlin');

$result = $platform->invoke($model, $messages, [
    'response_format' => $city,
    'missing_properties_only' => true,
]);

assert($city === $result->asObject());

The request carries a schema of population, country and mayor. name is absent,
because the application already knows the city is Berlin. Without the option the same call
sends all four properties and asks the model to restate a value it was given.

A property is left to the model when it is uninitialized, null or an empty array.
Every other value, including '', 0 and false, is taken as given and left out of the
schema. A nested object is decided by its own properties rather than by its presence: one
with nothing missing is left out entirely, a null one is described in full, and a
partially filled one is described with only its own gaps and populated in place. A
non-empty collection is left out, an empty one is described with its full item schema.

Without the option nothing changes: an instance still describes its whole class, exactly
as before.

How it works

The decision is made while the schema is built, not by post-processing a finished one,
so it composes with everything the describers already do:

  1. PlatformSubscriber consumes the option and hands the instance to the factory as
    $context['populate_instance']. The option itself is never forwarded to the provider.
  2. ResponseFormatFactoryInterface::create() gains an array $context = [] parameter and
    forwards it to Contract\JsonSchema\Factory::buildProperties(), which already took a
    describer context for serializer_groups. This is the BC break.
  3. PropertyInfoDescriber reads each property off the instance and skips the ones already
    filled, recursing into nested objects with the nested value as the new instance.

Values are read through the backing property by reflection regardless of visibility, so a
getter that would throw on an uninitialized property is never invoked; only a virtual
property without backing property falls back to its getter.

Using the option with anything but an instance as response_format — a class name, a raw
schema, or nothing at all — throws an InvalidArgumentException, as does an instance with
no missing properties left, before any request is sent.

Two fixes that came with it

  • Nested objects lost the describer context. TypeInfoDescriber and MethodDescriber
    built nested object schemas with an empty context, so serializer_groups silently
    stopped applying one level down. PropertySubject now carries the context and both
    describers propagate it. This is a pre-existing bug, independent of this feature.
  • Nested objects were replaced rather than populated. Deserializing onto an existing
    instance now uses DEEP_OBJECT_TO_POPULATE, in the buffered ResultConverter and in
    PartialObjectStreamListener alike. Without it a partial answer for a nested object
    would hand back a fresh instance and drop the values it already held — which is exactly
    what this feature produces.

BC impact

BC Break label + UPGRADE.md entry. Callers are unaffected. Implementors of
ResponseFormatFactoryInterface must add the parameter, otherwise PHP raises a
declaration-compatibility fatal error:

 final class MyResponseFormatFactory implements ResponseFormatFactoryInterface
 {
-    public function create(string $responseClass): array
+    public function create(string $responseClass, array $context = []): array
     {
         // ...
     }
 }

An implementation that accepts the parameter but ignores it keeps working and silently
disables missing_properties_only. In this repository the only implementations are
ResponseFormatFactory and a test double, both updated here. ai-bundle registers the
factory as a service and aliases the interface; nothing to change there.

Alternatives considered

(a) Widen the parameter to create(string|object $response). An earlier revision of
this PR. Equally a break for implementors, but it stops at the factory: the instance would
still have to be smuggled into buildProperties(), which is where the decision is
actually made. The context array reaches the describers, composes with serializer_groups
instead of sitting beside it, and keeps the public option a plain boolean.

(b) A separate createForInstance(object $response). Also breaks every implementor,
and leaves two methods to keep in sync for one concept.

(c) Post-process the finished schema. Cannot see what the describers saw, so nested
objects and recursion have to be re-derived from the schema rather than from the values.

(d) Do nothing; document decorating the factory. The status quo, and it only works by
keeping mutable per-invocation state on a service, which is not safe to recommend.

Follow-up (deliberately not here)

Contract\JsonSchema\Provider\SchemaProviderInterface::getSchemaFragment() still cannot
depend on the instance: SchemaAttributeDescriber calls it with the static array from
#[Schema(context: ...)]. Letting a provider fragment see the runtime subject raises its
own design questions and is better argued separately.

@carsonbot carsonbot added Feature New feature Platform Issues & PRs about the AI Platform component Status: Needs Review labels Aug 28, 2026
@chr-hertel

chr-hertel commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

Thanks @marco-jouwweb - makes sense as a feature, I'd say.

What do you think of converting this to an option that gets interpreted in the listener and does some post-processing of the schema there? the interface could stay the same.

brings in the issue of "how to detect if we should drop something" 🤔

@marco-jouwweb

Copy link
Copy Markdown
Contributor Author

Thanks @marco-jouwweb - makes sense as a feature, I'd say.

What do you think of converting this to an option that gets interpreted in the listener and does some post-processing of the schema there? the interface could stay the same.

brings in the issue of "how to detect if we should drop something" 🤔

Good to hear you are positive about the idea 🙂

About making it an option that gets interpreted by a listener; I am not fully sure what you mean. Do you mean that we pass it as new variable through $options, instead of distinguishing between instance & class scenarios through \is_object()? I can see that being a better option. Or possibly using \is_object() as fallback, in case the explicit option (e.g. use_instance) is not provided? Not sure what'd be cleanest here.

I don't know if I'm convinced by the post-processing listener suggestion, though. It seems like symptom treatment to first describe the class itself, just to alter it after the fact to represent the instance. I can see that becoming hard to manage real quickly. My gut feeling would say to properly describe instances from the start instead, so we don't have to solve that through post-processing of the JSON schema. Though I am not completely sure this is what you are talking about 🤔

@chr-hertel chr-hertel left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it makes sense to have that feature directly available with an option like this:

$platform->invoke($model, $messages, [
    'response_format' => new City(name: 'Berlin'),
    'missing_data_only' => true,
]);

but maybe missing_data_only is not that great of a name 😬


this would require tho to ship the code, and currently the documented MissingPropertiesResponseFormatFactory also only does post-processing of a decorated factory - but we need to change the interface for that.

If you'd promote that post-processing, that we already have, to be directly or triggered within the PlatformSubscriber, we don't need to change the interface :)

@marco-jouwweb

Copy link
Copy Markdown
Contributor Author

If you'd promote that post-processing, that we already have, to be directly or triggered within the PlatformSubscriber, we don't need to change the interface :)

I kinda trusted Claude a bit too much here, my bad. The example does encourage the thing I am not fond of, you are right. My misunderstanding was thinking that plumbing was the only missing part. It is not: even with the instance reaching the factory, the shipped factory still describes the full class. So the PR as it stands is an extension point, not the feature you describe.

That said, I still think post-processing is not the right place for it. Deciding which properties get described is a responsibility of the describer chain, and the chain already does exactly this kind of narrowing for serializer_groups: at describe-time, through the subject context, never by pruning afterwards. I'd like to propose doing the same here.

API

An explicit carrier option instead of sniffing the type of response_format:

$platform->invoke($model, $messages, [
    'response_format' => City::class,
    'populate_instance' => $city,
]);

Passing the instance as response_format keeps working as today (populate, full schema), so nothing changes for existing callers. Do we want to deprecate that path later in favor of the explicit option, or keep both?

Plumbing

PlatformSubscriber reads populate_instance, keeps it as the object to populate, and hands it to the factory through a context array:

public function create(string $responseClass, array $context = []): array;

ResponseFormatFactory forwards that context to Factory::buildProperties($responseClass, $context), which already accepts a context. So it is the same dictionary from options down to the describers, exactly like serializer_groups. Yes, it is still an interface change for implementors, but as you said, one is needed either way, and this shape is additive for callers.

Describer

PropertyInfoDescriber::describeObject() is the single place that decides which properties exist, and it already reads the per-call context there. If the context carries an instance and the property is already filled, the describer simply does not yield it. Everything downstream comes out right for free: required and additionalProperties are computed from the yielded properties, so strict: true stays consistent without rewriting anything. The describer also already resolved how to read each property ($readInfo: public property or getter), so reading the current value is a couple of lines.

Nested objects

For the fix to be complete, nested objects need the same treatment. Today TypeInfoDescriber recurses into a nested class with a fresh ObjectSubject and no context, and PropertySubject has no context at all. That is also why serializer_groups currently only reaches discriminated (anyOf) sub-schemas and not plain nested objects, despite what the docs say. I'd propose:

  • PropertySubject gets a context array, populated by PropertyInfoDescriber from the parent subject.
  • When recursing into an object-typed property, TypeInfoDescriber passes the parent context on, with the instance swapped for the property's current value. A null child describes the full child class; a partially filled child is narrowed the same way as the root.
  • This fixes the serializer_groups propagation gap as a side effect.
  • Collections: a schema has a single items definition, so per-item state cannot be expressed. First cut: a non-empty collection counts as filled and is skipped, an empty or null one is described in full. Per-item description (prefixItems) can be a follow-up if a provider supports it in strict mode (but I know OpenAI does not as of today).

"How to detect if we should drop something"

The describer has the reflector, the type and the read accessor at hand, so the rule can be simple and explicit: a property is to be filled when it is uninitialized or null. Anything else, including non-nullable defaults like '' or 0, counts as filled. I'd rather keep this strict than guess at empty-ish values, since 0 and false are legitimate data. Happy to discuss if you'd prefer treating '' as missing.

Edge case in either approach: when every property is filled, the schema collapses to nothing. I'd throw there, since sending a request with an empty schema is almost certainly not intended?

Why not post-process in the subscriber

  • The subscriber only sees a JSON array. It has no reflection info, no read accessors, no type info. The describer has all three.
  • It only works for the default schema shape. Any custom factory, anyOf, or nested structure needs a schema walker.
  • It has to keep required in sync by hand, and a mistake there fails silently: the model just gets asked the wrong questions.
  • It forecloses instance-aware decisions in the describers later, e.g. describing a recursive tree by the nodes that actually exist.

This increases the scope of the PR significantly. Could you please LMK what you think about the proposed solution, and if you'd like me to put all of this in the current PR or if you'd like it spread out over multiple smaller PR's?

Thanks for your time!

@chr-hertel

chr-hertel commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

hmm, I was writing with this code in mind and not sure how you pull that of in a describer

image

but sounds like you gave that more thought than I - however, let's try to avoid changing the existing response_format option and make the new behavior an explicit opt-in option missing_properties_only => true

Comment thread src/platform/src/Contract/JsonSchema/Describer/PropertyInfoDescriber.php Outdated
Comment thread src/platform/src/StructuredOutput/PlatformSubscriber.php Outdated
@marco-jouwweb
marco-jouwweb force-pushed the platform-response-format-factory-instance branch 2 times, most recently from 1cbcd9f to e96f524 Compare September 22, 2026 11:51
@marco-jouwweb marco-jouwweb changed the title [Platform] Pass the response_format instance to ResponseFormatFactoryInterface::create() [Platform] Complete structured output instance mode Sep 22, 2026
@marco-jouwweb

Copy link
Copy Markdown
Contributor Author

@chr-hertel I did the full implementation & updated the title and description. This extended the scope significantly, though most LOC are just tests.

Let me know if you need anything from me to make reviewing more bearable :-) but I don't expect that being necessary.

@chr-hertel chr-hertel left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still not convinced by that approach, sorry - let me explain:
We're adding a new feature, rather a specific one - maybe potential to grow, for now rather limited tho, but still meaningful. we could maybe make those cases up, but we don't have them yet.

we have the benefit that this feature only needs to go into an internal layer - we can be lazy. the only opinion people will have is about "how do i use it?" and maybe "how can i change it's behavior?" - the second question we're dropping for now - use-case based feature might change that.

with the current implementation this has impact one extension point and to unrelated implementations - so the change spreads.
compare the footprint of #2570, where the feature gets hooked in where it is used, but is isolated. IMO that makes it simpler to iterate and rework. (haven't tested the schema filter class tho yet.)

do you think i miss something here?

$context = [];
if (null !== $this->objectToPopulate) {
$context[AbstractNormalizer::OBJECT_TO_POPULATE] = $this->objectToPopulate;
// Populate nested objects in place too, so a partial answer keeps the values they already hold

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
// Populate nested objects in place too, so a partial answer keeps the values they already hold


if (null !== $this->objectToPopulate) {
$context[AbstractNormalizer::OBJECT_TO_POPULATE] = $this->objectToPopulate;
// Populate nested objects in place too, so a partial answer keeps the values they already hold

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
// Populate nested objects in place too, so a partial answer keeps the values they already hold

* }
*/
public function create(string $responseClass): array;
public function create(string $responseClass, array $context = []): array;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

to this is mostly the same contract change like before - not overloading the first parameter, but hiding the instance in the second. and array $context is more fuzzy and only used internally, right? i mean we'd have exactly the same functionality with create(string $responseClass, ?object $instanceToPopulate = null)

@marco-jouwweb

Copy link
Copy Markdown
Contributor Author

@chr-hertel Thanks for putting #2570 together, it made the comparison a lot more concrete than arguing in the abstract 🙂

I've thought about it some more and I want to split my answer in two, because I think we're actually disagreeing about one thing only.

The signature: you're right

array $context was me designing for a use-case we don't (yet) have. As the PR stands the subscriber only ever puts one key in it, so functionally it is ?object $instanceToPopulate. I'll change the interface to exactly that. It keeps the interface change minimal and explicit, and the describer side doesn't change at all since Factory::buildProperties() already takes the context internally.

One thing I want to flag so it's a conscious decision and not an accident: the day we want a per-call serializer_groups option for structured output (the natural sibling of validation_groups from #2527), the instance and the groups both have to reach the schema factory, and ?object can't carry the groups. That means a second change to create(), i.e. a second BC break for implementors, where array $context would have absorbed it. I'm fine with that trade since nobody has asked for per-call groups yet afaik and you'd rather not design for cases we don't have, I just want to ensure you are aware of that potential side-effect 🙂

I'd also take over the two rules #2570 gets right and mine doesn't: dropping properties that aren't writable on the instance (readonly / constructor-only), and stripping null from the type of a partially filled nested object so the model can't wipe it. Those changes are clear improvements.

The nested serializer_groups propagation fix (PropertySubject context, TypeInfoDescriber/MethodDescriber passing it on) stays in this PR on purpose: it is the same plumbing the feature needs to reach nested objects, so it is not really separable. It also shows the chain already has to carry per-call context one level down, the docs just claimed it did while it didn't. If you'd rather review it on its own I'm happy to split it out and rebase this on top, just say so.

The placement: I still think this belongs in the describer chain

Not because post-processing is ugly, but because of what the filter has to know. I put two failing tests on top of your branch in #2581 (illustration only, not meant to be merged):

  1. Polymorphic nested property. A SearchRequest holding an OrderFilter that still lacks userResponsible and departureDate. The property's schema is a discriminated anyOf without properties of its own, so the filter drops it and throws has no missing properties left to describe while two properties are missing. If the root had another gap it would drop filter silently instead. The describe-time version on this branch narrows the matching branch on the same input (root: filter; OrderFilter branch: userResponsible, departureDate + discriminator). Fixable in the filter, sure, but the fix needs the discriminator property and the class-to-value mapping, i.e. the serializer metadata SerializerDescriber already had in hand when it built the anyOf.

  2. Nested object behind $ref. Same outcome for a TreeNode whose child still lacks its own child, once the nested schema is a $ref into $defs. This one isn't fixable by patching: a $def is shared by every occurrence, while every instance has different gaps, so the only correct output is an inlined narrowed copy per occurrence. At that point the filter is rebuilding the schema. Describe-time narrowing gets this for free because it describes the occurrence with the instance in hand.

I don't think $defs/$ref is hypothetical. Factory::buildProperties() on a self-referential class currently overflows the stack, $defs is the only way to support recursion, OpenAI strict mode and Gemini both accept it, and I have a use-case for it myself. When it lands, every "which properties do we describe" decision that lives outside the chain has to learn about references. serializer_groups already lives inside it; I'd rather not have two mechanisms answering the same question in two places.

On "how can I change its behaviour": describers are already the customization point people have, so a future ''-counts-as-missing rule would land where they already look, instead of on a second surface.

Proposal

  • create(string $responseClass, ?object $instanceToPopulate = null), as you suggested, with the per-call groups caveat above on record.
  • Decision stays in PropertyInfoDescriber, with your readonly and null-stripping rules adopted.
  • serializer_groups propagation fix kept in here as part of the context plumbing, split out on request.
  • DEEP_OBJECT_TO_POPULATE as in both PRs.

That gets the footprint close to #2570 while keeping the one thing I think is worth the interface change. Would that work for you? If yes I'll rework this PR accordingly.

I think we've gave this quite some thought already and I don't want to take too much of your time, so whatever you decide I'll accept. Its a complex feature, thanks for the effort!

@chr-hertel

Copy link
Copy Markdown
Member

Love your endurance for the topic and your approach - really fun to wrangle around a topic and challenge ideas - thanks for that!

The new contract idea is not that "ugly" anymore - sure, give it a try. I don't have the chance to go after #2581 - in the morning at least.

@marco-jouwweb

marco-jouwweb commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Woah, I messed up a rebase and this caused a billion changes and this somehow automatically requested reviews from related code owners which was never my intention. Sorry for that! Fixing the Git spaghetti as we speak.

@marco-jouwweb
marco-jouwweb marked this pull request as draft September 23, 2026 09:03
@marco-jouwweb
marco-jouwweb force-pushed the platform-response-format-factory-instance branch from 74ba130 to 3f885d7 Compare September 23, 2026 09:05
`response_format` accepts an existing object instance, used as the serializer's
`object_to_populate`, but the schema half of the feature never saw it: the JSON schema
was always derived from the class, so the model was asked for every property again,
including the ones the instance already holds.

Add a `missing_properties_only` option. `PlatformSubscriber` hands the instance to
`ResponseFormatFactoryInterface::create()` through a new `?object $instanceToPopulate`
argument, and the shipped factory passes it into `Contract\JsonSchema\Factory::buildProperties()`
as the `populate_instance` describer context, so the decision which properties to describe
is made by the describers themselves rather than by post-processing a finished schema. A
property counts as missing when it is uninitialized, `null` or an empty array, and it can be
written onto the instance; nested objects are narrowed the same way, lose `null` from their
type and are populated in place through `DEEP_OBJECT_TO_POPULATE`.

`PropertySubject` now carries the describer context, which `TypeInfoDescriber` and
`MethodDescriber` propagate into nested object schemas. That also closes a gap for
`serializer_groups`, which previously only reached discriminated sub-schemas.

Co-authored-by: Cursor <cursoragent@cursor.com>
@marco-jouwweb
marco-jouwweb force-pushed the platform-response-format-factory-instance branch from 3f885d7 to ef6761f Compare September 23, 2026 09:10
@marco-jouwweb
marco-jouwweb marked this pull request as ready for review September 23, 2026 09:14
@marco-jouwweb

marco-jouwweb commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Love your endurance for the topic and your approach - really fun to wrangle around a topic and challenge ideas - thanks for that!

The new contract idea is not that "ugly" anymore - sure, give it a try. I don't have the chance to go after #2581 - in the morning at least.

Happy to hear, likewise! 🙂 Nice to see how you handle PRs in great depth, much appreciated.
I just pushed the changes as discussed, so this should be about it. Was able to make it a bit more compact and removed the excessive comments as well.


Edit: I thought about what the current contract means for a use case I have, and it makes me doubt ?object on its own.

We use domain objects that predate LLMs directly as structured output; a parallel family of DTOs did not survive the volume. We scope the fields the model sees with #[Groups], so the schema only offers what a call is meant to fill. The describer applies groups application-wide, through the PropertyInfoDescriber constructor, so a class without those groups gets an empty schema. What I need is per call:

$agent->call($messages, [
    'response_format' => $page,
    'serializer_groups' => ['ai-fill'],
]);

The describer half exists: Factory::buildProperties($class, $context) takes the context and PropertyInfoDescriber reads serializer_groups from it. Missing:

  1. PlatformSubscriber consumes the option.
  2. It hands the groups to create() so they reach buildProperties().
  3. ResultConverter puts them in its deserialization context.
  4. StructuredOutput\Serializer passes its ClassMetadataFactory to the ObjectNormalizer too. Today only the discriminator resolver gets it, so a groups key is ignored when deserializing. That one changes behaviour (serializer attributes start applying), so it needs its own look.

That is the sibling of validation_groups from #2527. This PR would only carry the groups to the schema side; the rest is a follow-up.

Which brings me back to the signature. With ?object $instanceToPopulate, adding this later means changing create() a second time, a second break for implementors. That is the trade-off I flagged earlier; it felt hypothetical then, now I have a concrete case. Two ways to avoid it, assuming you see value in my use-case example:

  • create(string $responseClass, ?object $instanceToPopulate = null, array $context = []), where $context is the contract buildProperties() already has.
  • Leave this PR as is, add serializer_groups in a follow-up, and accept the second break then.

What do you prefer? If the first, I prepare the signature here and send serializer_groups as a separate small PR, so this one doesn't grow.

@chr-hertel

Copy link
Copy Markdown
Member

okay, interesting, and you're right - we should zoom out: i want to control which properties are open for generation! maybe only missing ones, maybe a strategy like me defining groups, etc ... makes sense, but not easier right away 🤔

I think my main concern about not using the current factory-describer is about mixing data, behavior and schema. so the schema doesn't change by the data, it's only which parts we want to use from that - that's why my brain is pulling me into that filter idea over and over. encapsulating those concerns help with managing the complexity - unless there is a meaningful synergy in the first place - what part of your hypothesis is.

need to let that sink in for now and will come back to this ...

@marco-jouwweb

marco-jouwweb commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

okay, interesting, and you're right - we should zoom out: i want to control which properties are open for generation! maybe only missing ones, maybe a strategy like me defining groups, etc ... makes sense, but not easier right away 🤔

I think my main concern about not using the current factory-describer is about mixing data, behavior and schema. so the schema doesn't change by the data, it's only which parts we want to use from that - that's why my brain is pulling me into that filter idea over and over. encapsulating those concerns help with managing the complexity - unless there is a meaningful synergy in the first place - what part of your hypothesis is.

need to let that sink in for now and will come back to this ...

I agree, take all the time you need to let this sink in. I have trouble wrapping my mind around the complete picture as well 😅 I wanted to share the idea below, but please don't feel pressed to answer this today. For when you are ready:

I share your concern about mixing data, behaviour and schema. It blurs the lines and gets hard to follow quickly, so we should separate those properly. I understand why that makes the post-processing idea tempting again, but I still think it bites us in the long run (#2581), so here is a way to get the separation without it.

My first thought was to abstract the what-is-open logic (readValue() and shouldPopulate() in this PR) behind an interface:

interface PropertySelectorInterface
{
    /** Whether the model may generate this property of the object being described. */
    public function isOpen(ObjectSubject $object, PropertySubject $property): bool;

    /** The selector for a nested object property, null to describe the nested class in full. */
    public function forProperty(PropertySubject $property): ?self;
}

On its own that only moves the problem around: the runtime data would still be consulted inside a describer, a well-wrapped version of the same smell.

So the selection logic has to leave the describers. The natural call-site is the Describer orchestrator: its describeObject() loop takes every PropertySubject a describer yields and calls describeProperty() on it. At that moment we still know which PHP property a node belongs to, and no JSON exists yet. Roughly:

public function describeObject(ObjectSubject $subject, ?array &$schema): iterable
{
    $selector = $subject->getContext()[Factory::CONTEXT_SELECTOR] ?? null;
    $schema = $required = [];

    foreach ($this->objectDescribers as $describer) {
        foreach ($describer->describeObject($subject, $schema) as $property) {
            if ($selector instanceof PropertySelectorInterface) {
                if (!$selector->isOpen($subject, $property)) {
                    continue;
                }

                // The nested selector travels in the property's context; null describes the nested class in full
                $property = $property->withContext([Factory::CONTEXT_SELECTOR => $selector->forProperty($property)] + $property->getContext());
            }

            $this->describeProperty($property, $schema['properties'][$property->getName()]);
            // ...
        }
    }
}

(withContext() would be a small addition to PropertySubject.) TypeInfoDescriber already recurses with the property's context, so the nested selector reaches the next describeObject() without the describer knowing what it carries. Dropping null from the type of a nested object that is populated in place moves to the same spot in the orchestrator.

What this gives us:

  • Describers are a pure function of the class again. PropertyInfoDescriber goes back to almost what is on main, keeping only the context propagation.
  • The selector is behaviour and holds the data, in one class.
  • Class-level description stays cacheable. Per-instance selection is per call by definition, so nothing is lost there.
  • A getter that throws is the selector's problem, with the reflection-first read this PR already has.

To be upfront: the final schema of a call still depends on the instance, that is the feature. But the dependency is confined to one step in the orchestrator and one class, and the describers never see it.

Introduce Contract\JsonSchema\Selector\PropertySelectorInterface, applied by the Describer
orchestrator through the Factory::CONTEXT_SELECTOR context before a property is described,
so the describers stay a function of the class and the instance never enters them.
MissingPropertiesSelector holds the instance and the rule that was in PropertyInfoDescriber;
the nested selector travels in the property context via PropertySubject::withContext().

The orchestrator also drops a nested object populated in place when nothing on it is open,
and strips null from its type, where TypeInfoDescriber previously checked for the instance.
ResponseFormatFactory wraps the instance in the selector. Behaviour and tests are unchanged.
$schema = $required = [];
$selector = $subject->getContext()[Factory::CONTEXT_SELECTOR] ?? null;
if (!$selector instanceof PropertySelectorInterface) {
$selector = null;

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possibly throw here? Dunno, seems like a reason to fail fast because not narrowing can silently affect the output & costs. Better to fail loud & early?

* an existing instance never calls its constructor. A nested object is decided by its own properties, through the
* selector returned for it.
*
* @author Marco van Angeren <marco@jouwweb.nl>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude decided to leave me credits, feel free to remove it everywhere x]

@marco-jouwweb

marco-jouwweb commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

@chr-hertel Updated the PR to reflect my last comment: the selection now lives in the Describer orchestrator behind PropertySelectorInterface, and MissingPropertiesSelector holds the instance. PropertyInfoDescriber and TypeInfoDescriber are back to main plus the context propagation, so no describer sees runtime data anymore.

That also changed the contract: with the instance wrapped in a selector, ?object $instanceToPopulate had nothing left to do, so create() now takes the describer context itself, create(string $responseClass, array $context = []), the same array buildProperties() already accepts. PlatformSubscriber builds the selector and passes it under Factory::CONTEXT_SELECTOR.

Either shape works for me, ?object $objectToPopulate, array $context or the instance as part of $context, but I think we need $context on create() either way for the per-call serializer_groups from my earlier comment, and this PR is the place to add it so it stays one BC break instead of two. The serializer_groups option itself is not in here; that stays a separate small PR. LMK what you think when you are back on this topic :-)

With the instance wrapped in a MissingPropertiesSelector, the factory did nothing with it
but wrap it, so the `?object $instanceToPopulate` parameter is replaced by the describer
context itself: `create(string $responseClass, array $context = [])`, the contract
`Contract\JsonSchema\Factory::buildProperties()` already has. PlatformSubscriber builds the
selector and passes it under Factory::CONTEXT_SELECTOR; a per-call `serializer_groups`
option can later travel the same way without another change to the interface.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Feature New feature Platform Issues & PRs about the AI Platform component Status: Needs Review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants