Skip to content

Add elu activation - #6

Closed
haixuanTao wants to merge 2 commits into
dimforge:mainfrom
haixuanTao:feat/elu-activation
Closed

Add elu activation#6
haixuanTao wants to merge 2 commits into
dimforge:mainfrom
haixuanTao:feat/elu-activation

Conversation

@haixuanTao

Copy link
Copy Markdown
Contributor

Add multiple activation kernels!

haixuanTao and others added 2 commits May 26, 2026 16:20
Adds vortx::linalg::{Activation (tanh + tanh_backward), Adam} and their shaders,
the GPU building blocks for MLP training (used by nexus RL demos / zealot-rl).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add ELU to the activation kernels alongside the existing tanh:
- gpu_elu / gpu_elu_backward (backward reuses the cached output:
  grad *= y>0 ? 1 : y+1, same formulation as gpu_tanh_backward)
- gpu_elu_vec4 (4-wide ELU; host reinterprets the f32 buffer as
  glamx::Vec4) -- pulls in glamx as a dependency
- host Activation::elu / elu_backward / elu_vec4 wrappers
- export GpuElu, GpuEluBackward, GpuEluVec4

Needed for zealot's all-GPU PPO policy (ActorCritic uses ELU hidden
activations). Verified bit-exact vs the CPU ELU/elu_grad in zealot-rl
(~1e-7).

Note: the one-line glamx Cargo.toml dependency is shared with the
PPO-kernels and GEMM-vec4 branches; whichever lands first, the others
need a trivial rebase of that line.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@sebcrozet

Copy link
Copy Markdown
Member

Thank you for this PR!
I think this would be more suitable as part of inferi instead of Vortx, since inferi is about ML kernels. Some activations are already implemented there so additional ones can be added.

@sebcrozet

Copy link
Copy Markdown
Member

This is being merged as part of #11

sebcrozet added a commit that referenced this pull request Aug 27, 2026
* feat: move inferi shaders into an optional ml module for vortx

* chore: cargo fmt

* feat(ml): add tanh activation + Adam optimizer GPU ops

Replaces #5

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* refactor(ml): move the tanh-backward + Adam kernels into vortx::ml and drop the duplicated tanh forward

Completes #5

* feat(ml): add ELU activation GPU ops (fwd/backward/vec4)

Replaces #6

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* refactor(ml): keep only the ELU backward pass, the forward is already UnaryOp::Elu

Completes #6

* feat(ml): add PPO loss-gradient GPU kernels

Replaces #5

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* refactor(ml): wire the PPO kernels to vortx::ml and use StepRng for uniform control flow

Completes #5

* perf(linalg): vec4 GEMM (compute-FMA inner loop + vec4 global-load variant)

Replaces #7

Co-Authored-By: Haixuan Xavier Tao <tao.xavier@outlook.com>

* perf(linalg): keep the vec4 FMA inner loop, drop the unmeasured vec4 global-load GEMM variant

Completes #7

* feat(ml): gpu_ppo_stage_batch, building the PPO minibatch on device

* chore(ml): comment cleanups

* chore: CI fixes

---------

Co-authored-by: Haixuan Xavier Tao <tao.xavier@outlook.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants