distributed-llm-pretraining-torchtitan
Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/distributed-llm-pretraining-torchtitan/SKILL.md
- skills/distributed-llm-pretraining-torchtitan/references/checkpoint.md
- skills/distributed-llm-pretraining-torchtitan/references/custom-models.md
- skills/distributed-llm-pretraining-torchtitan/references/float8.md
- skills/distributed-llm-pretraining-torchtitan/references/fsdp.md
Every link opens the file at its source, pinned to the revision this page describes.