Repository navigation
Replace transformer blocks config with building blocks - #470
Merged
Merged
Conversation
Models now write the block body as plain code, with Layers.Transformer.blocks/3 handling the cache and outputs, and shared helpers for attention and FFN. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Given the similarity between LLMs, we have a single transformer function that each model called with a bunch of options (and the options were passed a few levels down where they were actually relevant, so adding a new option involved some boilerplate). However, as each model introduces some changes or new approach, it grows the option list. In a few cases, we had to add callbacks to let the caller further override the behaviour. This makes it progressively hard to tell the overall shape of a given model.
On the other hand, we don't want to duplicate all code per model. Being self-contained has some advantages, but there are fundamental layer/math parts that we don't want to diverge and potentially re-review across models.
This PR changes it to a better middle ground, which is, we replace the one massive function with a few building blocks. This means each model has a bit more code, but it also has flexibility to arrange layers and their order as necessary, and that flow stays visible within the model. Models that change something more fundamentally, can break further break apart into using smaller building blocks.
We still have
Layers.Transformer.blocks/3function that covers the repeated block structure of most LLMs, but it now accepts a function to build each block in entirety, and that gets composed out of smaller pieces.