Skip to content

[Feature]: Multioutput models #485

Description

@joein

What feature would you like to request?

In some cases our default model interface returning an array of numpy arrays is not enough
There are scenarios when we want to return more than that either in order to save resources or just because there is no other way.

An example of the first case is when models can return several types of embeddings using the same backbone but different head layers, e.g. baai/bge-m3.
Or when it is desired to obtain not only pooled embedding, but also token embeddings of the penultimate layer.

Another example is colpali-like models, it is not always possible to build index on colpali embeddings since of the huge resources consumption and thus it makes it difficult to use colpali as the first stage retriever.
Instead, first stage retrieval can be implemented via colpali embeddings pooled by either rows or columns.

In the particular case of colpali, it is possible to do the pooling after inference is completed and embeddings were returned, since the number of patches into which colpali preprocessor breaks an image is known beforehand.
However, with colpali-inspired models like colqwen, it is not possible to do the pooling afterwards, because the number of patches into which these models break an image is dynamic, thus, we need to know an image size to get the number of patches.

In order to handle these cases, we want to introduce new family of Multioutput classes with an interface allowing to return either dict or special model output objects.

Is there any additional information you would like to provide?

No response

Activity

  1. AyushAnand413 commented on Oct 6, 2026

    @AyushAnand413

    Hi @joein, I'd like to work on this.

    I see #756 recently handled selecting a single named output, but the broader multi-output return type is still open here. Before starting, I wanted to confirm the preferred interface: should embed() yield a dict mapping output names to arrays per document (e.g. {"dense": np.ndarray, "token_embeddings": np.ndarray}), or would you prefer a dedicated dataclass?

    On the implementation side, I'm planning to update the ONNX runner to pass all requested output names to session.run() and map each postprocessed result to its key, making sure batch/parallel workers handle the dict outputs correctly. Happy to start with a draft PR if this direction makes sense.

  2. self-assigned this
    on Oct 6, 2026
  3. joein commented on Oct 6, 2026

    @joein
    MemberAuthor

    Hey @AyushAnand413

    Thanks for your interest, though we don't yet know how we would like to implement it properly, we don't want a dict.
    We'd need to come up with a design doc first.

    This issue is not ready to be worked on now.

  4. AyushAnand413 commented on Oct 6, 2026

    @AyushAnand413

    Understood thanks for the quick clarification @joein!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions