RoPE has an fp64 torch.autograd.gradcheck and it passes; GeGLU does not have one. The GeGLU backward is only verified against PyTorch autograd at bf16/fp16 tolerances, which are loose enough to hide a genuinely wrong gradient term — a sign error on a small contribution can sit inside a 5e-2 bar.
Add an fp64 gradcheck on small shapes for both the activation-only path and the gate+up fused path. Small shapes are fine and fast; the point is exactness, not scale.
RoPE has an fp64
torch.autograd.gradcheckand it passes; GeGLU does not have one. The GeGLU backward is only verified against PyTorch autograd at bf16/fp16 tolerances, which are loose enough to hide a genuinely wrong gradient term — a sign error on a small contribution can sit inside a 5e-2 bar.Add an fp64 gradcheck on small shapes for both the activation-only path and the gate+up fused path. Small shapes are fine and fast; the point is exactness, not scale.