Hi
Firstly thank you for this work - this is super interesting. I'm curious if youve looked at the newer Metal Performance Primitives for dispatching metal work for running ML workload within a command buffer / command stream?
https://machinelearning.apple.com/research/exploring-llms-mlx-m5
https://developer.apple.com/download/files/Metal-Performance-Primitives-Programming-Guide.pdf
Im wondering if MLX would provide any optimizations to make this run on M5 much faster?
Hi
Firstly thank you for this work - this is super interesting. I'm curious if youve looked at the newer Metal Performance Primitives for dispatching metal work for running ML workload within a command buffer / command stream?
https://machinelearning.apple.com/research/exploring-llms-mlx-m5
https://developer.apple.com/download/files/Metal-Performance-Primitives-Programming-Guide.pdf
Im wondering if MLX would provide any optimizations to make this run on M5 much faster?