laya-mlx
View on GitHubNative MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Native Apple Silicon MLX inference runtime for Laya typed-decision models. Returns choice probabilities, rubric scores and P(true) locally in ~7-14 ms per short question, with no PyTorch, tokenizer decoding or cloud API at runtime.
Use Cases
Local short-question classification on Apple SiliconSupport ticket department routingUrgency rubric scoring with expected scoreBoolean proposition probability estimationMultilingual language routing across checkpointsEmail triage helpersShortlisting hundreds of choice labels via embeddingsOffline inference without cloud API or output tokensBenchmarking inference latency and throughput vs upstream PyTorch portsTerminal Snake demo driving decisions per game move
Built With
- Language
- Python
- Frameworks
- MLX · Hugging Face Hub · Hugging Face Tokenizers · NumPy · PyTorch (reference) · Transformers (reference) · safetensors · pytest · Ruff · Hatchling · matplotlib
Tags
mlx · apple-silicon · inference-runtime · local-inference · on-device-ai · typed-decisions · modernbert · python · fp16 · structured-output · batch-inference · model-conversion · huggingface-hub · macos · low-latency · no-cloud