nanogpt
Provides nanoGPT, Karpathy’s minimal PyTorch GPT implementation (model.py and train.py), with workflows for training character-level Shakespeare on CPU, reproducing GPT-2 124M on OpenWebText with multi-GPU torchrun, fine-tuning pretrained GPT-2 checkpoints, and training on custom text. Use when learning how GPT and transformer blocks work from scratch. Use when running a small training experiment on CPU or a single GPU. Use when modifying a transformer variant in plain PyTorch. Use when preparing character-level or BPE binary datasets. Use when troubleshooting out-of-memory or slow training in nanoGPT. Not for production or large-scale distributed training; use HuggingFace Transformers or Megatron-LM instead.Category: ml-training · License: MIT · Version: 1.0.0
