Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Triton Learning Notes & Experiments

I've been digging into OpenAI's Triton to understand how it compares to writing raw CUDA or just using PyTorch. Here are my notes and performance findings from implementing some basic kernels.

1. Vector Addition

The "Hello World" of GPU programming. I wanted to see if a simple Python-like kernel could actually match the performance of PyTorch's optimized C++ backend.

Implementation

The kernel is pretty straightforward (vec_add/vec_add.py). It calculates offsets based on the program ID and handles boundary conditions (masking) so we don't read/write out of bounds.

@triton.jit
def add_kernel(X_ptr, Y_ptr, output_ptr, num_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(0)
    # ... offset calculations ...
    x = tl.load(X_ptr + offsets, mask=mask, other=0.0)
    y = tl.load(Y_ptr + offsets, mask=mask, other=0.0)
    tl.store(output_ptr + offsets, x + y, mask=mask)

Findings

I benchmarked this across vector sizes ranging from 1KB to ~16MB.

Vector Addition Performance

Comparison of throughput (GB/s) vs Vector Size

The results were surprisingly tight. My Triton kernel (Blue) sits right on top of PyTorch (Orange) for almost the entire graph. Both peak around 158 GB/s on my machine.

This confirms that for bandwidth-bound operations like this, Triton generates code that saturates memory bandwidth just as well as native CUDA kernels, without the boilerplate.


2. Softmax Implementation

Softmax is more interesting because it requires a reduction (calculating max and sum across a dimension) before normalizing.

I implemented two versions in softmax/softmax_v1.py:

  1. Standard Softmax: Loads the entire row into SRAM/Registers, computes, and writes back.
  2. Online Softmax: Uses the online softmax algorithm (similar to Welford's for variance) to compute streaming updates for max and sum. This allows tiling the computation.

The "Aha!" Moment with Large Tensors

The performance comparison revealed a critical limitation of the naive approach.

Softmax Performance

Performance metric: GB/s vs Column Size (N)

Observations:

  • Small Columns (N < 8192): The naive Triton implementation (Blue) is extremely fast, matching or slightly beating PyTorch (Green). It hits about 156 GB/s.
  • The Cliff (N > 16384): Notice the massive drop in the Blue line at the end. For $N=65536$, the standard kernel drops to ~30 GB/s.
    • Why? The naive kernel tries to keep the entire row of data in registers/SRAM. When the row size exceeds the GPU's shared memory resources, it spills to global memory, killing performance.
  • Online Softmax (Red): The streaming/tiled implementation shines here. It maintains a steady ~105 GB/s even at the largest sizes where the naive version fails.

Takeaway

For "small" reductions that fit in L1/SRAM, the naive approach is great. But for production-grade kernels that need to handle arbitrary shapes, the Online Softmax approach is essential to maintain throughput and avoid register spilling.

About

Benchmark-driven Triton kernel experiments covering vector operations, reductions, online softmax, and GPU optimization.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages