Custom GPT-2 with memory-efficient FlashAttention implemented in Triton for faster training in PyTorch.