Repository navigation
Docs: Add design for CPU inference class support - #509
RedbackThomson wants to merge 3 commits into
Conversation
Signed-off-by: Nicholas Thomson <RedbackThomson@users.noreply.github.com>
8a33459 to
5fe6a53
Compare
| 3. The serving stack installs dra-driver-cpu on a cluster when one of its pools | ||
| uses a class with a `dra.cpu` device, and only on those pools. |
There was a problem hiding this comment.
Is there a reason not to install it everywhere? It seems like it could be useful for reasons other than CPU inference.
There was a problem hiding this comment.
You're talking about also installing it on the GPU pools?
There was a problem hiding this comment.
Right - what I'm picturing is being able to steer workloads based on CPU and GPU requirements there, e.g. for CPU offload.
That said iiuc you have to reserve CPU for to be managed by DRA? If so I can see why that wouldn't make sense.
| `reservedCPUs` decides what a class should declare as `capacity.cpu`: a node's | ||
| vCPUs less two. I propose a fixed value, documented with the InferenceClass, | ||
| over a per-pool one. A per-pool value would have to be kept in step with the | ||
| class by hand, which is the same problem in a less visible place. |
There was a problem hiding this comment.
Reserved for inference? Any reason to pick vCPUs minus two?
(Little unclear in the current doc.)
There was a problem hiding this comment.
Two vCPUs for overhead like the DRA driver and any other critical DaemonSet like CSI or CNI drivers. But yeah I do feel this really will depend on how much is running on the node apart from the model
| cpu: "4" | ||
| ``` | ||
|
|
||
| That means adding `capacity.requests` to the ModelDeployment and ModelReplica |
There was a problem hiding this comment.
Do you know if anything apart from the CPU DRA driver supports capacity.requests? Just curious.
There was a problem hiding this comment.
The NVIDIA GPU driver supports it through ConsumableShares - https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/dra-intro-install.html#capability-comparison
Other than that, nothing I could find
|
|
||
| That means adding `capacity.requests` to the ModelDeployment and ModelReplica | ||
| device requests, carrying it through the scheduler into the | ||
| ResourceClaimTemplate's `exactly.capacity.requests`, and checking it against |
There was a problem hiding this comment.
What does exactly mean here? Are there other (inexact?) types of requests on the RCT?
I ask because the asymmetry makes me wonder why you're proposing this live under capacity.requests in our API, not exactly.capacity.requests.
There was a problem hiding this comment.
Ah, we already default to exactly. I wonder if we should change that and expose it in our APIs so we can support firstAvailable if we ever need to.
| charges a full node to every member that claims a device. It would need to | ||
| account for capacity on shared devices. The cluster also needs KEP-5075's | ||
| `DRAConsumableCapacity` feature, and I haven't checked which Kubernetes | ||
| versions enable it by default. That is enough work to deserve its own design. |
There was a problem hiding this comment.
Sounds like it's worth a future work section.
| - `status.gpuPools` on InferenceCluster will list CPU pools. Its content is | ||
| already generic. I propose loosening its description rather than renaming a | ||
| field the scheduler reads. |
There was a problem hiding this comment.
Let's rename it too so it doesn't say gpuPools when they're not GPU pools. Breaking changes are fine.
There was a problem hiding this comment.
Proposing status.schedulablePools
| ## Open questions | ||
|
|
||
| - What is the minimum Kubernetes version dra-driver-cpu 0.3.0 supports, given | ||
| it builds on 1.37, and does every provider Modelplane provisions offer it? | ||
| - Which Kubernetes versions on each provider enable `DRAConsumableCapacity`? | ||
| That decides whether capacity requests can be relied on. |
There was a problem hiding this comment.
I prefer to avoid open questions in design if we can. Are these easy to answer?
There was a problem hiding this comment.
The second question is kind of hard to answer because I couldn't find any public documentation from Nebius with the word DRAConsumableCapacity. We can always spin up clusters at different versions and see when they enabled it, but I can't just google it.
|
Few questions but this seems like a good idea overall. |
Description of your changes
Modelplane, today, focuses strongly on GPU as the only hosting hardware for models. With GPUs being hard to come by, and slow to provision, these are impractical for quickly testing the deployment of a new Modelplane system. Instead, I offer that Modelplane supports CPU models through the
dra-driver-cpudevice classes, which we've already tested internally at Upbound to work with Modelplane.This design outlines a plan for achieving parity with existing GPU inference classes using CPU-only nodes.
I have:
nix flake check(or./nix.sh flake check) and made sure it passes.git commit -s.