Vibe Coding Discover

AI Frameworks

Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.

★ 5.8K428 forksPythonApache-2.0mizorewww

Native Apple Silicon MLX inference runtime for Laya typed-decision models. Returns choice probabilities, rubric scores and P(true) locally in ~7-14 ms per short question, with no PyTorch, tokenizer decoding or cloud API at runtime.

Use Cases

Local short-question classification on Apple SiliconSupport ticket department routingUrgency rubric scoring with expected scoreBoolean proposition probability estimationMultilingual language routing across checkpointsEmail triage helpersShortlisting hundreds of choice labels via embeddingsOffline inference without cloud API or output tokensBenchmarking inference latency and throughput vs upstream PyTorch portsTerminal Snake demo driving decisions per game move

Built With

Language
Python
Frameworks
MLX · Hugging Face Hub · Hugging Face Tokenizers · NumPy · PyTorch (reference) · Transformers (reference) · safetensors · pytest · Ruff · Hatchling · matplotlib

Tags

mlx · apple-silicon · inference-runtime · local-inference · on-device-ai · typed-decisions · modernbert · python · fp16 · structured-output · batch-inference · model-conversion · huggingface-hub · macos · low-latency · no-cloud