huggingface-tokenizers
Provides the HuggingFace Tokenizers library (Rust core with Python and Node.js bindings) for training and using BPE, WordPiece, and Unigram tokenizers. Covers the pipeline of normalizers, pre-tokenizers, models, post-processors and decoders, padding and truncation, batch encoding, alignment tracking, and conversion to transformers PreTrainedTokenizerFast. Use when training a custom tokenizer or vocabulary on a new corpus, tokenizing large text corpora quickly, mapping tokens back to character offsets for NER or question answering, or configuring normalization and special tokens. Use when wrapping a custom tokenizer for transformers. For SentencePiece models or tiktoken, use those tools instead; for loading a pretrained tokenizer only, AutoTokenizer is enough.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/huggingface-tokenizers/SKILL.md
- skills/huggingface-tokenizers/references/algorithms.md
- skills/huggingface-tokenizers/references/integration.md
- skills/huggingface-tokenizers/references/pipeline.md
- skills/huggingface-tokenizers/references/tokenization-algorithms.md
- skills/huggingface-tokenizers/references/training.md
Every link opens the file at its source, pinned to the revision this page describes.