molfeat
Converts SMILES strings or RDKit/datamol molecules into numerical features using molfeat (0.11.0), which provides calculators, scikit-learn compatible transformers, and pretrained embedding models. Covers fingerprints (ECFP, MACCS, MAP4), RDKit and Mordred descriptors, pharmacophore descriptors, and pretrained models such as ChemBERTa and GIN, with parallel processing and caching. Use when building QSAR or QSPR models from SMILES. Use when choosing among molecular featurizers for a property prediction task. Use when generating embeddings for virtual screening, similarity search, or chemical space clustering. Use when adding a featurizer to a scikit-learn pipeline. Not for general cheminformatics tasks such as reading or editing structures, which belong to RDKit itself.
- Version
- 1.2
- License
- Apache-2.0 license
- Compatibility
- Requires Python 3.9–3.10 (molfeat 0.11.0 does not support 3.11+). Requires datamol, PyTorch, and optional extras for GNN/transformer models.
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Write
- Edit
- Bash
Files
- skills/molfeat/SKILL.md
- skills/molfeat/references/api_reference.md
- skills/molfeat/references/available_featurizers.md
- skills/molfeat/references/choosing_a_featurizer.md
- skills/molfeat/references/examples.md
Every link opens the file at its source, pinned to the revision this page describes.