text-and-embedding-features
Turn text into numbers for search, matching, deduplication, classification or clustering, and choose between TF-IDF, embeddings and simpler methods. Use whenever someone asks how to find similar tickets, products, documents, CVs or reviews by their text; detect near-duplicate records from names, addresses or descriptions; build TF-IDF vectors, n-grams or stop-word lists; choose word, sentence or document embeddings (word2vec, GloVe, fastText, sentence transformers or a hosted embedding API); decide whether embeddings are worth the cost over TF-IDF; handle tokenisation, subwords, spelling variants or several languages; measure cosine similarity; or store and search vectors for semantic search. Not for choosing how to cluster the resulting vectors into segments, and not for encoding ordinary categorical columns.
Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/text-and-embedding-features/SKILL.md
- skills/text-and-embedding-features/scripts/tfidf_similarity.py
Every link opens the file at its source, pinned to the revision this page describes.