Skip to content

perf(cuda): cap SM86 Stream-K grid to useful chunks - #583

Open
Amidwestnoob wants to merge 2 commits into
Luce-Org:mainfrom
Amidwestnoob:exp/sm86-streamk-pr
Open

perf(cuda): cap SM86 Stream-K grid to useful chunks#583
Amidwestnoob wants to merge 2 commits into
Luce-Org:mainfrom
Amidwestnoob:exp/sm86-streamk-pr

Conversation

@Amidwestnoob

@Amidwestnoob Amidwestnoob commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Caps the SM86 Stream-K grid at useful MMQ iteration chunks, avoiding CTAs that cannot receive work. Other architectures retain the existing schedule.

Validated on dual RTX 3090 SM86:

  • Host schedule tests: 17/17 Release and UBSan
  • IQ4_XS GPU correctness: 5/5
  • Current-upstream A/B/A: +1.46% wall throughput
  • All 15 outputs byte-identical

Review in cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 4 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread server/deps/llama.cpp/ggml/src/ggml-cuda/mmq-streamk-schedule.h Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant