k-dense-ai/drug-discovery-agent-skills
Agent Skills for small-molecule and protein therapeutics: target validation and human genetics, bioactivity and chemical space, generative design and retrosynthesis, docking, free energy and dynamics, ADMET and PK translation, protein, antibody, degrader and oligonucleotide design, and the clinical and regulatory record.
How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports adaptyv, adaptyv_sdk, or FoundryClient, or references foundry-api-public.adaptyvbio.com.
Turn a set of structures into absorption, distribution, metabolism, excretion, and toxicity estimates with ADMET-AI, and read them as a developability verdict rather than a table of numbers. Use this skill to run batch prediction over a library, interpret each endpoint against its DrugBank-approved percentile, and flag the liabilities that stop a series — hERG blockade, CYP inhibition, poor Caco-2 permeability, high clearance, and plasma protein binding. Also trigger on ADMET-AI, admet_ai, Chemprop-RDKit, hERG liability, CYP3A4 inhibition, Caco-2, bioavailability prediction, or developability triage.
Number antibody variable domains, annotate CDRs, and assess developability from sequence. Use this skill to apply IMGT, Kabat, Chothia, Martin, or AHo numbering with ANARCI, delimit CDRs and framework regions, scan for chemical liabilities (N-glycosylation sequons, deamidation NG, isomerisation DG, oxidation, unpaired cysteine, fragmentation), compute pI, net charge, extinction coefficient and hydrophobicity, and plan humanisation by CDR grafting. Also trigger on antibody, nanobody, VHH, scFv, Fab, CDR, framework, ANARCI, abnumber, IgBLAST, OAS, SAbDab, humanization, Vernier residues, or developability.
Structure-based docking with AutoDock Vina, Vinardo, and AutoDock4 through the Meeko toolchain. Use this skill to define a docking box, prepare receptors and ligands as PDBQT, run single or batch docking, rescore, and interpret affinities, poses, and ligand efficiency. Covers box definition from a reference ligand or pocket residues, protonation and tautomer decisions, flexible side chains, exhaustiveness and seeds, redocking validation, and virtual screening over compound libraries. Also trigger on vina, smina, gnina, mk_prepare_ligand, mk_prepare_receptor, mk_export, scrub.py, PDBQT, autogrid4, docking box, or binding-pose prediction.
Decide whether a protein has a pocket worth targeting, and where it is, before committing to a docking or design campaign. Use this skill to run fpocket cavity detection, rank cavities by druggability and volume, compare apo and holo conformations to spot induced fit, identify allosteric and cryptic cavities that only open in simulation, and convert a chosen cavity into the search box coordinates a docking run needs. Also trigger on fpocket, cavity detection, druggability score, alpha sphere, cryptic pocket, allosteric site, pocket volume, hotspot mapping, or undruggable target assessment.
Cofold protein-ligand, protein-protein, and nucleic-acid complexes with Boltz-2, and predict binding affinity with its trained affinity head. Use this skill to build Boltz input YAML, run structure prediction with MSAs, pocket constraints, templates, and modified residues, screen compound libraries by cofolding, and interpret confidence scores (pLDDT, pTM, ipTM, PDE) and affinity output (binder probability and log10 IC50). Also trigger on Boltz, Boltz-1, Boltz-2, cofolding, boltz predict, affinity_pred_value, affinity_probability_binary, ipTM, or open-weights AlphaFold3 alternatives.
Query the ChEMBL database web services for measured bioactivity data, compound records and calculated properties, targets, assays, mechanisms of action, drug indications and warnings. Use this skill to build curated SAR or QSAR datasets for a target, look compounds up by SMILES, InChIKey, name, or ChEMBL id, run similarity and substructure searches, and check what chemistry is already known against a protein. Also trigger when a query mentions ChEMBL ids (CHEMBL...), pChEMBL values, IC50/Ki/Kd/EC50 retrieval, assay confidence scores, or ebi.ac.uk/chembl.
Navigate make-on-demand catalogues — ZINC-22 through CartBlanche and Enamine REAL Space — to find compounds that can actually be ordered. Use this skill to look substances up by ZINC identifier or structure, understand tranche partitioning by heavy-atom count and logP, and choose between screening an enumerated subset and searching a combinatorial synthon space with a fragment-growing method such as V-SYNTHES. Also trigger on ZINC22, CartBlanche, Enamine REAL, make-on-demand, tangible library, synthon, tranche, giga-scale enumeration, or ultra-large virtual screening.
Search the ClinicalTrials.gov registry through its version 2 REST API for interventional and observational studies, their phases, enrolment, endpoints, sponsors, and posted results. Use this skill to survey who is developing what against an indication, date a competitor's programme, read primary and secondary outcome measures, find eligibility criteria, and distinguish a study that completed from one that was terminated or withdrawn. Also trigger on ClinicalTrials.gov, NCT number, trial registry, study phase, enrolment, primary outcome measure, recruiting status, trial sponsor, or competitive landscape.
Pythonic wrapper around RDKit with a simplified interface and sensible defaults. Preferred for standard drug discovery work — SMILES/SELFIES/InChI conversion, molecule standardization and sanitization, descriptors, ECFP and other fingerprints, Tanimoto distance matrices, Butina clustering and diverse subset picking, Bemis-Murcko scaffolds and scaffold splits, BRICS/RECAP fragmentation, 3D conformer generation, SDF/CSV/Excel and cloud I/O, and parallel processing via n_jobs. Returns native rdkit.Chem.Mol objects, so it composes with RDKit throughout. Also trigger on datamol, import datamol as dm, dm.to_mol, dm.standardize_mol, dm.cluster_mols, or dm.pick_diverse. For advanced control or custom parameters, use the rdkit skill directly.
Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featurization options and MoleculeNet benchmarks. Best for quick experiments with pre-trained models, diverse molecular representations. For ready-made ADMET numbers without training a model use admet-prediction; for benchmark datasets and task-aware splits use pytdc.
Work on bifunctional degraders and molecular glues, where potency comes from a ternary complex rather than occupancy. Use this skill to apply the property rules that govern this beyond-rule-of-five space, reason about linker length, attachment vector and E3 ligase choice, prepare inputs for ternary complex structure prediction, and interpret degradation readouts — DC50, Dmax, cooperativity, and the hook effect that makes a dose-response curve turn over. Also trigger on PROTAC, molecular glue, targeted protein degradation, E3 ligase, cereblon, VHL, ternary complex, DC50, Dmax, hook effect, cooperativity, or PROTAC-DB.
Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), RNAi DEMETER2 scores, PRISM compound sensitivity, and gene effect profiles across the cell-line panel. Use for identifying cancer-selective vulnerabilities, separating pan-essential genes from selective ones, finding synthetic lethal interactions, correlating dependency with mutation, expression and copy number, and validating oncology drug targets. Also trigger on DepMap, Chronos gene effect, CRISPRGeneEffect.csv, DEMETER2, PRISM repurposing, co-essentiality, pan-essential, or ACH- cell line identifiers.
DiffDock and DiffDock-L diffusion-based molecular docking. Use for blind protein-small-molecule pose prediction from a PDB file or sequence plus SMILES/SDF/MOL2, batch docking over a CSV of complexes, virtual screening triage, sampling multiple poses per complex, and reading the confidence score correctly. Also trigger on DiffDock, DiffDock-L, inference.py, confidence_model, samples_per_complex, ESM embedding preparation for docking, or blind docking without a defined box. Not for binding affinity prediction — the confidence score ranks pose plausibility, not potency.
Protein language models through the EvolutionaryScale esm Python SDK. Generate and embed sequences with ESM3 (multimodal sequence, structure and function prompting), extract per-residue and mean-pooled embeddings with ESM C, fold sequences with ESMFold2, and run inference locally or against the Forge and Biohub hosted clients. Use this skill for protein representation learning, variant effect and mutational scanning from likelihoods, sequence generation and inpainting, structure prediction from sequence alone, and embedding features for downstream models. Also trigger on esm, ESM3, ESMC, ESM Cambrian, ESMFold2, from esm.models, ESMProtein, GenerationConfig, forge.evolutionaryscale.ai, biohub.ai, or ESM_API_KEY.
Compute relative and absolute binding free energies with the Open Free Energy toolkit — the rigorous alchemical alternative to docking scores when a congeneric series needs reliable potency ranking. Use this skill to plan a perturbation network over a ligand set, choose atom mappings, run hybrid-topology or separated-topology protocols, and analyse the result — per-edge ΔΔG with uncertainty, cycle-closure error, and mean unsigned error against measured affinities. Also trigger on OpenFE, alchemical transformation, thermodynamic cycle, RBFE, ABFE, SepTop, lambda window, MBAR, cycle closure, or perturbation map.
Generate and optimise novel small molecules with REINVENT 4 — de novo sampling from a chemical language model, scaffold decoration with LibInvent, fragment linking with LinkInvent, and similarity-constrained analogue generation with Mol2Mol. Use this skill to set up reinforcement-learning or curriculum runs, compose a multi-parameter scoring function from docking scores, predictive models, and physicochemical desirability, and read the resulting sampled sets. Also trigger on REINVENT, LibInvent, LinkInvent, Mol2Mol, scaffold hopping, R-group replacement, linker design, chemical language model, or reinforcement-learning molecule optimisation.
Analyze and engineer protein glycosylation. Scan sequences for canonical N-glycosylation sequons (N-X-S/T with X not proline, including overlapping sites), predict O-GalNAc hotspots, read glycan notation, and reach the curated external tooling (NetNGlyc, NetOGlyc, GlycoShield, GlycoWorkbench, GlyTouCan, GlyConnect). Use this skill for therapeutic antibody glycoengineering and afucosylation for ADCC, Fc glycan control, glycan shielding in vaccine immunogen design, sequon removal or insertion, and half-life engineering through sialylation. Also trigger on N-glycosylation, sequon, NXS/NXT, O-glycosylation, glycoform heterogeneity, afucosylation, high-mannose, GlyTouCan, or WURCS.
Estimate how likely a protein therapeutic is to provoke an anti-drug antibody response, and locate the sequence regions responsible. Use this skill to tile a sequence into peptides, predict class II MHC presentation across a population-representative allele panel, aggregate predicted binders into a per-region and whole-molecule risk score, compare a candidate against its closest human germline, and decide which liabilities are worth deimmunising. Also trigger on immunogenicity, anti-drug antibody, ADA, T-cell epitope, MHC class II, HLA-DRB1, NetMHCIIpan, NetMHCpan, deimmunisation, tregitope, or population coverage.
Medicinal chemistry filters for compound triage. Apply drug-likeness rules (Lipinski rule of five, Veber, Oprea, CNS, lead-like, rule of three), structural alert catalogs (PAINS a/b/c, NIBR screening-deck severity, Brenk, BMS, Glaxo, Dundee, ChEMBL common alerts), ZINC-15 percentile complexity metrics (Bertz, SAscore, QED, Whitlock, Barone), chemical-group detection, Lilly demerits, and the medchem query language (MATCHRULE, HASALERT, HASPROP, HASGROUP) for filtering a library at scale. Also trigger on medchem, import medchem as mc, RuleFilters, NIBRFilters, CommonAlertsFilters, NamedCatalogs, QueryFilter, PAINS filtering, or structural alerts.
Run and analyze molecular dynamics simulations with OpenMM and MDAnalysis. Set up protein and protein-ligand systems with PDBFixer, choose force fields and water models (AMBER14, CHARMM36m, ff19SB, GAFF2, TIP3P), solvate and add ions, run energy minimization, NVT/NPT equilibration and production MD on GPU, then analyze trajectories for RMSD, RMSF, radius of gyration, hydrogen bonds, native contacts, PCA and free energy surfaces. Use this skill for protein stability under mutation, ligand binding-mode and residence-time questions, conformational sampling, membrane proteins, and disordered ensembles. Also trigger on OpenMM, MDAnalysis, mdtraj, Simulation.step, LangevinMiddleIntegrator, PDBFixer, DCD or XTC trajectory, RMSD analysis, or production MD.
Molecular featurization hub with one consistent interface over 100+ featurizers. Fingerprints (ECFP/Morgan, MACCS, atom pair, topological torsion, Avalon, RDKit, ERG), RDKit and Mordred descriptor sets, pharmacophore and 3D shape descriptors, scaffold keys, and pretrained embeddings (ChemBERTa, ChemGPT, MolT5, GIN, Graphormer) through a common transformer API with caching and parallelism. Use this skill to convert SMILES into model-ready feature matrices for QSAR, virtual screening, and molecular ML, and to choose between featurizer families. Also trigger on molfeat, MoleculeTransformer, FPVecTransformer, PretrainedHFTransformer, molfeat model store, or featurizer selection.
Queries the NCATS Translator ARAX production API for bounded, typed, provenance-rich one-hop and endpoint-pinned two-hop biomedical knowledge-graph relationships. Use for Biolink-constrained RTX-KG2 lookup, explicit selected-provider ARAX federation, separate entity normalization, qualifier-aware graph traversal, and inspection of TRAPI edge bindings, publications, and knowledge-source provenance. Do not use for inference, ranking, open-ended pathfinding, clinical guidance, or sensitive queries.
Design small interfering RNA and antisense oligonucleotide sequences against a transcript, and screen them for the failure modes specific to nucleic-acid drugs. Use this skill to tile a target transcript, apply positional and thermodynamic selection rules including duplex asymmetry and nearest-neighbour melting temperature, scan candidates for seed-region complementarity to off-target transcripts, and lay out a chemical modification pattern — gapmer architecture, 2'-O-methyl and 2'-MOE wings, locked nucleic acid, and phosphorothioate placement. Also trigger on siRNA, antisense oligonucleotide, ASO, gapmer, RNase H, seed region, duplex asymmetry, 2'-MOE, locked nucleic acid, phosphorothioate, or GalNAc conjugate.
Query the Open Targets Platform GraphQL API for target-disease associations, genetic and clinical evidence, tractability and safety liabilities, target prioritisation metrics, known drugs and mechanisms of action, and disease ontology. Use this skill for target identification and validation, target-disease evidence review, druggability assessment, drug repurposing, and resolving gene, disease, and drug names to Ensembl, MONDO, and ChEMBL identifiers. Also trigger when a query mentions Open Targets, platform.opentargets.org, association scores, tractability buckets, or api.platform.opentargets.org.
Query the FDA's public openFDA APIs for post-market drug data — FAERS adverse-event reports, Drugs@FDA approval and submission history, Structured Product Labels including boxed warnings, the National Drug Code directory, recall enforcement reports, and drug shortages. Use this skill to check what a regulator has already concluded about a molecule or its class, to date an approval and count its efficacy supplements, to read an approved indication or boxed warning, and to score a drug-event pair for disproportionate reporting with PRR, ROR, and chi-squared. Also trigger on openFDA, api.fda.gov, FAERS, Drugs@FDA, SPL, NDC, pharmacovigilance, boxed warning, adverse event report, drug recall, or safety signal.
Find out whether a chemical series is already claimed, using SureChEMBL's patent-extracted compound corpus and, where a key is available, PatentsView for legal status and assignee history. Use this skill to trace a structure to the patent documents that disclose it, survey an assignee's filings around a target, and understand what the freedom-to-operate question requires that a structure search cannot answer. Also trigger on SureChEMBL, patent chemistry, Markush structure, freedom to operate, composition of matter, assignee, priority date, patent family, or PatentsView.
Turn in vitro potency and animal pharmacokinetics into a defensible human dose projection — the arithmetic that decides whether a compound can reach its target concentration safely. Use this skill for non-compartmental analysis of a concentration-time profile (AUC, Cmax, terminal half-life, clearance, volume of distribution), one- and two-compartment simulation of a dosing regimen, interspecies allometric scaling, human-equivalent dose conversion by body surface area, and the exposure margin between a projected therapeutic concentration and a toxicology no-effect level. Also trigger on non-compartmental analysis, AUC, clearance, volume of distribution, allometric scaling, human equivalent dose, first-in-human, NOAEL, therapeutic index, or exposure margin.
Query the Precision Medicine Knowledge Graph (PrimeKG) for multiscale biological relationships across genes and proteins, drugs, diseases, phenotypes, pathways, biological processes, exposures and anatomy. Use this skill to search entities by name, pull direct neighbours and their evidence types, summarise the local network around a disease, and find direct or two-hop drug-disease connections for repurposing hypotheses. Also trigger on PrimeKG, kg.csv, Harvard Dataverse knowledge graph, disease_protein, drug_protein, indication and contraindication edges, or network pharmacology over a biomedical knowledge graph.
Design new proteins that bind a chosen surface, using BindCraft's AlphaFold2-guided hallucination or the RFdiffusion backbone plus ProteinMPNN sequence pipeline. Use this skill to specify a target epitope by hotspot residue, trim a receptor to the region worth designing against, set up a design campaign, and filter the output on the in-silico metrics that predict experimental success — interface predicted TM-score, predicted aligned error at the interface, buried surface area, and shape complementarity. Also trigger on BindCraft, RFdiffusion, ProteinMPNN, minibinder, hallucination, inverse folding, hotspot residue, epitope targeting, ipTM, or de novo binder.
Use Therapeutics Data Commons through the PyTDC Python package for registry discovery, approved dataset access, task-aware splits (scaffold, cold-start, temporal, combination), evaluator metrics, benchmark groups, and bounded molecular-oracle workflows. Use this skill to find which TDC datasets exist for a therapeutic task, load them with a split that does not leak, score predictions with the task's own official metric rather than a generic one, and run benchmark groups reproducibly. Also trigger on PyTDC, Therapeutics Data Commons, tdc.single_pred, tdc.multi_pred, ADMET Benchmark Group, scaffold split, get_split, or molecular oracles such as GSK3B, JNK3 and DRD2.
Cheminformatics toolkit for fine-grained molecular control. Parse and write SMILES, SDF, MOL and InChI; compute descriptors (MW, LogP, TPSA, QED, Bertz); build fingerprints (Morgan/ECFP, RDKit, MACCS, atom pair, torsion) and score Tanimoto, Dice or cosine similarity; run SMARTS substructure search and reaction SMARTS; generate 2D depictions and ETKDG 3D conformers; extract Murcko scaffolds and canonical hashes; control sanitization and stereochemistry directly. Also trigger on rdkit, Chem.MolFromSmiles, rdFingerprintGenerator, SDMolSupplier, SMARTS query, ETKDG, or FilterCatalog. For standard workflows with a simpler interface use the datamol skill, which wraps RDKit; use rdkit for advanced control, custom sanitization, and specialized algorithms.
Plan synthetic routes and judge whether a proposed molecule can actually be made, using AiZynthFinder's Monte-Carlo tree search over template-derived reactions and a purchasable building-block stock. Use this skill to configure expansion and filter policies, choose a stock file, run route search over a candidate set, and read the returned trees — solved fraction, route depth, and which building blocks a route bottoms out in. Also trigger on AiZynthFinder, retrosynthetic tree search, synthetic accessibility, SAscore, RAscore, building-block stock, reaction template, or route scoring.
Rowan is a cloud-native molecular modeling and medicinal-chemistry workflow platform with a Python API. Use for pKa and macropKa prediction, conformer and tautomer ensembles, docking and analogue docking, protein-ligand cofolding, MSA generation, molecular dynamics, permeability, descriptor workflows, and related small-molecule or protein modeling tasks. Ideal for programmatic batch screening, multi-step chemistry pipelines, and workflows that would otherwise require maintaining local HPC/GPU infrastructure.
Access a collection of open-source molecular design and structural biology tools on the Tamarind Bio platform, via its REST API or MCP server — no local GPUs required. Tamarind bundles popular open-source models for structure prediction (AlphaFold, Boltz, Chai, ESMFold), protein, binder, and de novo design (RFdiffusion, ProteinMPNN, BoltzGen), antibody and nanobody design and developability, protein-ligand docking (DiffDock, Autodock Vina), binding-affinity prediction, MSA generation, and molecular dynamics. Use when the user mentions Tamarind or tamarind.bio, wants to run any of these open-source tools in the cloud, references app.tamarind.bio/api or the x-api-key header, or needs to submit batches of sequences for structural or biophysical characterization.
Assemble the human genetic evidence for and against a target before a programme commits to it — the evidence class that most improves the odds of surviving clinical development. Use this skill to pull gnomAD constraint metrics (LOEUF, pLI, observed/expected) that show whether loss of function is tolerated in people, retrieve GWAS Catalog associations and fine-mapped credible sets for a gene, and read a natural human knockout as a safety readout. Also trigger on gnomAD, LOEUF, pLI, loss-of-function intolerance, mutational constraint, GWAS Catalog, credible set, human knockout, genetic support, or target safety dossier.
Retrieve protein sequences, annotation, and structures from UniProtKB, the RCSB PDB, and AlphaFold DB. Use this skill to resolve a gene or protein name to a UniProt accession, pull sequences and FASTA files, find binding sites and domains, search the PDB by UniProt accession, sequence, ligand, or text, download mmCIF/PDB coordinates and biological assemblies, fetch AlphaFold models with their pLDDT confidence, and check whether a structure is actually usable before docking or simulating it. Also trigger on UniProt accessions, PDB ids, rest.uniprot.org, search.rcsb.org, files.rcsb.org, alphafold.ebi.ac.uk, id mapping, SEQRES, or missing residues.