deepspeed
Covers DeepSpeed for distributed deep learning training and I/O: ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8 training, 1-bit Adam, sparse attention, and DeepNVMe (aio_handle, gds_handle, async_io and gds operators, ds_nvme_tune, ds_report) for fast transfers between NVMe storage and host or GPU tensors. Use when configuring ZeRO stages for large-model training, enabling mixed precision or 1-bit Adam, offloading parameters or optimizer state to NVMe with ZeRO-Infinity, writing or reading tensors to files with blocking or non-blocking DeepNVMe calls, or tuning NVMe I/O settings such as block size, queue depth and parallelism. Not for plain PyTorch DistributedDataParallel without DeepSpeed.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/deepspeed/SKILL.md
- skills/deepspeed/references/08.md
- skills/deepspeed/references/09.md
- skills/deepspeed/references/2020.md
- skills/deepspeed/references/2023.md
- skills/deepspeed/references/assets.md
- skills/deepspeed/references/index.md
- skills/deepspeed/references/mii.md
- skills/deepspeed/references/other.md
- skills/deepspeed/references/tutorials.md
Every link opens the file at its source, pinned to the revision this page describes.