Skip to content

[TP][sglang] When/how will TP be supported? #6

Description

@lvhoaa

Hi team,

This is a great work to address the tough CUDA graph init coldstart, and I believe it's going to be much more impactful if Foundry could support TP since a great number of LLM engine deployment uses TP.

I understand that supporting TP is not straightforward as other parallelism schemes (DP, EP) since it involves NCCL all-reduce buffer/communicator issues and VMM conflicts...

So I want to ask if the team has any concrete solution / strategy / plan to support TP in Foundry.

Having this capability would make this work truly impressive in covering almost all parallel deployment configs today!

Thanks!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions