[NeurIPS 2025 Spotlight] TPA: Tensor ProducT ATTenTion Transformer (https://arxiv.org/abs/2501.06425)
-
Updated
Sep 4, 2026 - Python
[NeurIPS 2025 Spotlight] TPA: Tensor ProducT ATTenTion Transformer (https://arxiv.org/abs/2501.06425)
[ICLR 2026] GRAPE: Group Representational Position Encoding (https://arxiv.org/abs/2512.07805)
Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)
Implementing and training/testing popular model architectures on the CIFAR10 dataset.
A beginner's investigation into the world of neural networks, using the MNIST image dataset
Load a Hugging Face model onto a real GPU and inspect it live. See the real architecture, weights, and activations, and run causal interventions — dense and Mixture-of-Experts models alike.
To associate your repository with the model-architectures topic, visit your repo's landing page and select "manage topics."