Back to News
Models
GitHub - mizorewww/laya-mlx: Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Github.com·September 20, 2026·1 min read
AI Summary
Open-weight typed decisions, running natively on Apple Silicon. 13.
From the source
Open-weight typed decisions, running natively on Apple Silicon. 13.4 ms median end-to-end for a short English typed decision. 7.4 ms with the multilingual checkpoint. 0 output tokens. Local MLX inference, with no PyTorch, Transformers runtime, or cloud API. 中…
The full text couldn't be loaded here (the source may require a subscription).
View original at Github.comKeep reading
Was this useful?