nanoGPT
View on GitHubThe simplest, fastest repository for training/finetuning medium-sized GPTs.
Minimal, readable ~300-line PyTorch repo for training and finetuning GPT-2-scale models; reproduces GPT-2 124M on OpenWebText. Author marks it deprecated in favor of the newer nanochat, but it remains a compact hacking/education base.
Use Cases
Train a GPT from scratchFinetune GPT-2 on custom textReproduce GPT-2 124M on OpenWebTextTrain character-level language modelsSample/generate text from trained checkpointsLearn transformer internals via readable codeMulti-GPU/multi-node DDP trainingStudy scaling laws and transformer sizingHack a minimal training loop for research
Built With
- Language
- Python
- Frameworks
- PyTorch · HuggingFace Transformers · HuggingFace Datasets · tiktoken · wandb · numpy · tqdm
Tags
gpt · llm-training · finetuning · transformer · pytorch · gpt-2 · text-generation · minimal · educational · ddp · character-level · openwebtext · shakespeare · reproduction · deprecated · nanochat