Token-Print
View on GitHubInteractive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.
TokenPrint renders a real LLM forward pass in interactive 3D: attention heads, RoPE, GQA and the residual stream traced from actual tensors or a local GGUF file. Four modes (architecture, generation, walkthrough, debugger) with provenance tagging. Runs a FastAPI backend plus a Next.js/Three.js frontend.
Use Cases
Visualize attention heads and GQA grouping in real 3DInspect any local .gguf model in the browser without uploadingTrace real forward-pass tensors and residual stream valuesDebug per-layer activations with breakpoints and ablationTeach or learn transformer internals with cited explanationsCompare architecture across Qwen, Llama, Gemma, Mistral, GPT-2Watch token generation step by stepCheck provenance of each displayed number (real vs derived)
Built With
- Language
- TypeScript
- Frameworks
- Next.js · React · Three.js · WebGL · FastAPI · Uvicorn · PyTorch · Hugging Face Transformers · llama.cpp · Playwright · Vitest · Python
Tags
transformer-visualization · llm-interpretability · mechanistic-interpretability · 3d-visualization · threejs · webgl · gguf · attention-heads · forward-pass · activation-debugging · rope · gqa · model-inspector · websocket · huggingface