Skip to content

Docs: Add design for CPU inference class support - #509

Open
RedbackThomson wants to merge 3 commits into
modelplaneai:mainfrom
RedbackThomson:push-txsyyqltylou
Open

RedbackThomson wants to merge 3 commits into
modelplaneai:mainfrom
RedbackThomson:push-txsyyqltylou

Conversation

@RedbackThomson

Copy link
Copy Markdown

Description of your changes

Modelplane, today, focuses strongly on GPU as the only hosting hardware for models. With GPUs being hard to come by, and slow to provision, these are impractical for quickly testing the deployment of a new Modelplane system. Instead, I offer that Modelplane supports CPU models through the dra-driver-cpu device classes, which we've already tested internally at Upbound to work with Modelplane.

This design outlines a plan for achieving parity with existing GPU inference classes using CPU-only nodes.

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes.
  • Signed off every commit with git commit -s.

Signed-off-by: Nicholas Thomson <RedbackThomson@users.noreply.github.com>
@RedbackThomson RedbackThomson changed the title Docs: Add design for CPU models Docs: Add design for CPU inference class support Oct 7, 2026
Comment thread design/cpu-inference.md
Comment on lines +21 to +22
3. The serving stack installs dra-driver-cpu on a cluster when one of its pools
uses a class with a `dra.cpu` device, and only on those pools.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reason not to install it everywhere? It seems like it could be useful for reasons other than CPU inference.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're talking about also installing it on the GPU pools?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right - what I'm picturing is being able to steer workloads based on CPU and GPU requirements there, e.g. for CPU offload.

That said iiuc you have to reserve CPU for to be managed by DRA? If so I can see why that wouldn't make sense.

Comment thread design/cpu-inference.md
Comment on lines +204 to +207
`reservedCPUs` decides what a class should declare as `capacity.cpu`: a node's
vCPUs less two. I propose a fixed value, documented with the InferenceClass,
over a per-pool one. A per-pool value would have to be kept in step with the
class by hand, which is the same problem in a less visible place.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reserved for inference? Any reason to pick vCPUs minus two?

(Little unclear in the current doc.)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two vCPUs for overhead like the DRA driver and any other critical DaemonSet like CSI or CNI drivers. But yeah I do feel this really will depend on how much is running on the node apart from the model

Comment thread design/cpu-inference.md
cpu: "4"
```

That means adding `capacity.requests` to the ModelDeployment and ModelReplica

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you know if anything apart from the CPU DRA driver supports capacity.requests? Just curious.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The NVIDIA GPU driver supports it through ConsumableShares - https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/dra-intro-install.html#capability-comparison

Other than that, nothing I could find

Comment thread design/cpu-inference.md

That means adding `capacity.requests` to the ModelDeployment and ModelReplica
device requests, carrying it through the scheduler into the
ResourceClaimTemplate's `exactly.capacity.requests`, and checking it against

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does exactly mean here? Are there other (inexact?) types of requests on the RCT?

I ask because the asymmetry makes me wonder why you're proposing this live under capacity.requests in our API, not exactly.capacity.requests.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, we already default to exactly. I wonder if we should change that and expose it in our APIs so we can support firstAvailable if we ever need to.

Comment thread design/cpu-inference.md
charges a full node to every member that claims a device. It would need to
account for capacity on shared devices. The cluster also needs KEP-5075's
`DRAConsumableCapacity` feature, and I haven't checked which Kubernetes
versions enable it by default. That is enough work to deserve its own design.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds like it's worth a future work section.

Comment thread design/cpu-inference.md Outdated
Comment on lines +249 to +251
- `status.gpuPools` on InferenceCluster will list CPU pools. Its content is
already generic. I propose loosening its description rather than renaming a
field the scheduler reads.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's rename it too so it doesn't say gpuPools when they're not GPU pools. Breaking changes are fine.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Proposing status.schedulablePools

Comment thread design/cpu-inference.md
Comment on lines +283 to +288
## Open questions

- What is the minimum Kubernetes version dra-driver-cpu 0.3.0 supports, given
it builds on 1.37, and does every provider Modelplane provisions offer it?
- Which Kubernetes versions on each provider enable `DRAConsumableCapacity`?
That decides whether capacity requests can be relied on.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I prefer to avoid open questions in design if we can. Are these easy to answer?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The second question is kind of hard to answer because I couldn't find any public documentation from Nebius with the word DRAConsumableCapacity. We can always spin up clusters at different versions and see when they enabled it, but I can't just google it.

@negz

negz commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Few questions but this seems like a good idea overall.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants