Skip to content

kalarislabs/research-agent-skills

v1.1.1MIT

281 Agent Skills for researchers: scientific and research paper writing, journal formats, literature review, citations, data science, ML research and domain science.

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional isotopomers, parallel tracer experiments, and determining whether labeling data constrain a pathway flux. Distinguishes measured-label inference from COBRA flux balance analysis and flags experiments requiring nonstationary MFA.

abstract-and-title

Write and sharpen research paper titles, abstracts (structured and unstructured), keywords, highlights, significance statements, graphical-abstract text and lay summaries. Use when drafting or revising an abstract to a word limit, turning results into a one-sentence contribution, choosing between candidate titles, writing Cell/Elsevier highlights, a PNAS significance statement or a plain-language summary, or optimizing discoverability (keywords, search terms).

academic-plotting

Generates publication-quality figures for ML papers from research context. Given a paper section or description, extracts system components and relationships to generate architecture diagrams via Gemini. Given experiment results or data, auto-selects chart type and generates data-driven figures via matplotlib/seaborn. Use when creating any figure for a conference paper.

acm-sigconf

Format ACM conference papers and journal articles with the acmart LaTeX class (sigconf, sigplan, acmsmall, acmlarge, acmtog, manuscript/review/anonymous modes) or the ACM Word template, covering CCS concepts, keywords, ACM Reference Format, rights management and copyright blocks, anonymization for double-blind review, accessibility (alt text), the TAPS production workflow and common acmart errors. Use when writing for CHI, SIGGRAPH, KDD, SIGMOD, CCS, FAccT, WWW, SIGIR, CSCW, PLDI or any ACM venue or journal.

adaptyv

How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports adaptyv, adaptyv_sdk, or FoundryClient, or references foundry-api-public.adaptyvbio.com.

aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

alphagenome

Look up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.

analytical-method-validation

Plans and evaluates analytical method validation, verification, and transfer using Python scripts (plan_validation, check_response, check_accuracy_precision, check_detection_limits, check_bioanalytical_run, compare_methods) under ICH Q2(R2)/Q14, ICH M10, USP 1220/1225/1226, CLSI EP, or ISO/IEC 17025. Covers protocols, linearity, accuracy and precision, DL/QL, M10 run checks, TOST equivalence, Deming and Passing-Bablok. Use when designing a validation protocol with pre-stated acceptance criteria. Use when evaluating linearity, precision, or detection limit data. Use when checking an ICH M10 bioanalytical run or ISR. Use when comparing or transferring a method between labs. Does not decide a method is validated; USP, CLSI, ISO text is not supplied.

anndata

Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

apa7

Format papers, theses and references in APA Style 7th edition for psychology, education, social sciences, nursing and business, covering student vs professional papers, title page, abstract and keywords, heading levels, in-text citations, reference list entries (journal articles, books, chapters, datasets, software, webpages, AI tools), tables and figures, bias-free language, and APA via LaTeX (apa7 class + biblatex-apa), Word or Pandoc/CSL. Use when a user needs APA format, APA citations or references, or an APA-compliant manuscript or thesis.

ara-compiler

Compiles any research input — PDF papers, GitHub repositories, experiment logs, code directories, or raw notes — into a complete Agent-Native Research Artifact (ARA) with cognitive layer (claims, concepts, heuristics), physical layer (configs, code stubs), exploration graph, and grounded evidence. Use when ingesting a paper or codebase into a structured, machine-executable knowledge package, building an ARA from scratch, or converting research outputs into a falsifiable, agent-traversable form.

ara-research-manager

Records research provenance at the end of a coding or research session by scanning the conversation and writing decisions, experiments, dead ends, pivots, claims, and heuristics into the ara/ directory (exploration_tree.yaml, claims.md, heuristics.md, observations.yaml, session records), each tagged user, ai-suggested, ai-executed, or user-revised. Use when a task is finished and you want a session epilogue. Use when you need an auditable trace of how a project evolved. Use when decisions and dead ends need logging with provenance. Use when initializing an ara/ directory. Do not use during task execution, since it must run only after the request is complete.

ara-rigor-reviewer

Performs ARA Seal Level 2 semantic epistemic review of an Agent-Native Research Artifact directory, reading PAPER.md, logic/claims.md, logic/experiments.md, and trace/exploration_tree.yaml. Scores six dimensions (evidence relevance, falsifiability, scope calibration, argument coherence, exploration integrity, methodological rigor) from 1 to 5 and writes a severity-ranked level2_report.json with a Strong Accept to Reject grade. Use when Level 1 structural validation has passed and an ARA needs a critique before release. Use when checking whether claims are supported by their cited experiments. Use when auditing falsification criteria or scope over-claiming. Use when judging whether an exploration tree documents real dead ends. Not for structural or reference validation, which is Level 1.

arbor

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone arbor CLI tool see references/arbor-upstream.md.

arboreto

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.

arxiv-submission

Prepare and post preprints to arXiv without processing failures or leaks, covering TeX source packaging (.bbl, figures, case-sensitive paths), stripping private comments, choosing categories and license, endorsement, metadata (title, abstract, comments field), versioning and replacements, and linking the journal DOI later. Use when posting a paper to arXiv, when arXiv's TeX compilation fails, before uploading source, or when updating a preprint after acceptance. Includes a zero-dependency preflight checker.

astropy

Core Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.

audiocraft-audio-generation

PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation.

autogpt-agents

Autonomous AI agent platform for building and deploying continuous agents. Use when creating visual workflow agents, deploying persistent autonomous agents, or building complex multi-step AI automation systems.

autoresearch

Orchestrates end-to-end autonomous AI research projects using a two-loop architecture. The inner loop runs rapid experiment iterations with clear optimization targets. The outer loop synthesizes results, identifies patterns, and steers research direction. Routes to domain-specific skills for execution, supports continuous agent operation via Claude Code /loop and OpenClaw heartbeat, and produces research presentations and papers. Use when starting a research project, running autonomous experiments, or managing a multi-hypothesis research effort.

autoskill

Observe the user's screen via screenpipe, detect repeated research workflows, match them against existing research-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.

awq-quantization

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.

axolotl

Provides guidance for fine-tuning large language models with Axolotl, covering YAML training configs, LoRA and QLoRA, preference training with DPO, KTO, ORPO and GRPO, multimodal models, FSDP and multi-GPU setups, sequence (context) parallelism, NCCL bandwidth tests, dataset formats, compressed model saving for vLLM and llmcompressor, and custom integrations. Use when writing or debugging an Axolotl YAML config, choosing a dataset format for a fine-tuning run, setting up FSDP or context_parallel_size across GPUs, or running DPO, KTO, ORPO or GRPO training. Use when saving a compressed model for vLLM inference, or when writing a custom Axolotl plugin or integration. Not for inference serving or general Hugging Face Trainer scripts outside Axolotl.

benchling-integration

Benchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.

bgpt-paper-search

Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.

bibtex-hygiene

Clean, deduplicate and validate BibTeX/BibLaTeX bibliographies before submission. Use when a .bib file has duplicate keys or works, missing fields, broken DOIs, lost acronym capitalization (e.g. "bert" instead of "BERT"), inconsistent page dashes, arXiv preprints that now have published versions, or when LaTeX/biber emits bibliography warnings. Includes a zero-dependency linter and normalizer.

bids

Organizes, queries, validates, and converts neuroscience and biomedical datasets using the Brain Imaging Data Structure (BIDS) standard, covering MRI, PET, EEG, MEG, iEEG, EMG, NIRS, microscopy, motion capture, and behavioral data. Covers PyBIDS (BIDSLayout), the bids-validator, HeuDiConv, dcm2bids, BIDScoin, JSON sidecars, dataset_description.json, events.tsv, participants.tsv, .bidsignore, and derivatives. Use when arranging raw data into BIDS directories and filenames. Use when converting DICOM scans to BIDS. Use when validating a dataset before OpenNeuro or DANDI submission. Use when querying files or sidecar metadata with PyBIDS. Use when writing BIDS derivatives. For extracellular Neuropixels analysis, use neuropixels-analysis instead.

biopython

Provides Biopython (Bio.Seq, Bio.SeqIO, Bio.Align, Bio.Entrez, Bio.Blast, Bio.PDB, Bio.Phylo, Bio.motifs, Bio.SeqUtils, Bio.Restriction) for sequence handling, file parsing, NCBI access, BLAST, structures, and phylogenetic trees in Python. Reads, writes, and converts FASTA, GenBank, FASTQ, PDB, mmCIF, Newick, and NEXUS files. Use when manipulating or translating DNA, RNA, or protein sequences, converting sequence file formats, fetching records from NCBI via Entrez, running or parsing BLAST searches, doing pairwise or multiple sequence alignment, analyzing PDB structures, or building and editing phylogenetic trees. For quick lookups use gget; for multi-service integration use bioservices.

bioservices

Unified Python interface to 40+ bioinformatics services. Use when querying multiple databases (UniProt, KEGG, ChEMBL, Reactome) in a single workflow with consistent API. Best for cross-database analysis, ID mapping across services. For quick single-database lookups use gget; for sequence/file manipulation use biopython.

blip-2-vision-language

Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.

brainstorming-research-ideas

Guides researchers through structured ideation frameworks to discover high-impact research directions. Use when exploring new problem spaces, pivoting between projects, or seeking novel angles on existing work.

bulk-rnaseq

End-to-end bulk RNA-seq orchestrator — takes raw FASTQ reads through QC and trimming (FastQC, fastp/Trim Galore), alignment and quantification (STAR, Salmon, featureCounts), assembles a gene-level counts matrix, then hands off to differential expression (pydeseq2), pathway/GSEA enrichment (pathway-enrichment), and publication figures (scientific-visualization). Use whenever the user has bulk RNA-seq reads or quant output and wants a complete, reproducible differential-expression workflow — e.g. "analyze my RNA-seq", "FASTQ to DESeq2", "run nf-core/rnaseq", "STAR/Salmon quantification", "build a counts matrix for DESeq2", or "go from reads to differentially expressed genes and enriched pathways". Routes between an nf-core/rnaseq (Nextflow) path and a standalone STAR/Salmon path, and covers experimental design, strandedness, and QC gates. For single-cell RNA-seq use the scanpy skill instead.

cell-press

Prepare manuscripts for Cell Press journals (Cell, Molecular Cell, Neuron, Immunity, Cell Reports, Cell Systems, iScience, Cell Metabolism, Current Biology and others), covering Summary, Highlights, eTOC blurb and graphical abstract, STAR Methods with the Key Resources Table, resource availability (lead contact, materials, data and code), figure and statistics requirements, and Cell Press formatting at submission vs revision. Use when targeting a Cell Press journal, writing highlights or a graphical abstract, or converting methods into STAR Methods.

cellxgene-census

Query the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data. Use when you need population-scale cell metadata, gene expression slices, Census summary counts, source H5AD URIs/downloads, embeddings, spatial Census data, or reference atlas comparisons across organisms, tissues, diseases, assays, and cell types. For analyzing your own local single-cell data use scanpy, anndata, or scvi-tools.

chroma

Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects.

cirq

Google quantum computing framework. Use when targeting Google Quantum AI hardware, designing noise-aware circuits, or running quantum characterization experiments. Best for Google hardware, noise modeling, and low-level circuit design. For IBM hardware use qiskit; for quantum ML with autodiff use pennylane; for physics simulations use qutip.

citation-management

Searches OpenAlex, PubMed, and Google Scholar, extracts metadata from DOIs, PMIDs, PMCIDs, arXiv IDs, and URLs via CrossRef, PubMed, and arXiv, then formats, deduplicates, and validates BibTeX files using bundled Python scripts (search_openalex.py, extract_metadata.py, format_bibtex.py, validate_citations.py, doi_to_bibtex.py). Use when converting DOIs or PMIDs to BibTeX, finding papers on a topic, filling missing volume, pages, or DOI fields, cleaning or merging .bib files with duplicate entries, or checking a bibliography for errors before submission. Not for systematic review search methodology or synthesis; use literature-review for that.

citation-verification

Verify that every reference in a manuscript really exists and matches its metadata, catching hallucinated, corrupted or mismatched citations before submission. Use when checking a .bib file or reference list, after an AI assistant drafted citations, when a DOI may point to the wrong paper, when adding missing DOIs, or before camera-ready. Queries Crossref, OpenAlex and arXiv; zero dependencies.

clinical-decision-support

Prepare and validate research-only clinical decision-support evaluation, evidence-profile, cohort, survival, biomarker/model, privacy, and governance artifacts. Use for aggregate or synthetic research documentation and traceability—not patient care or live clinical operation.

clinical-reports

Generates fail-closed draft JSON templates and runs local deterministic structure and consistency checks for clinical reports: CARE case reports, radiology, pathology and lab scaffolds, CONSORT 2025, SPIRIT 2025, ICH E3 CSRs, aggregate adverse-event tables, provenance manifests, terminology schemas, and de-identification process checklists. Use when drafting a case report for publication from a verified source-fact manifest. Use when checking a trial results manuscript, protocol, or CSR draft against its reporting structure. Use when building an aggregate safety table from synthetic or de-identified data. Use when routing an artifact to the right guideline. Not for diagnosis, treatment advice, individual case safety report creation, or filing and submission.

clip

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

cobrapy

Runs constraint-based metabolic modeling with COBRApy (Python, import cobra) on genome-scale models in SBML, JSON, YAML, or MATLAB format. Covers FBA, pFBA, geometric FBA, FVA, flux sampling, gene and reaction knockouts, production envelopes, growth media, gapfilling, and building models. Use when loading or exporting a genome-scale metabolic model. Use when predicting growth or flux distributions with FBA or FVA. Use when screening gene or reaction knockouts. Use when tuning growth media or exchange constraints. Use when gap-filling an infeasible model or checking model consistency. Not for kinetic or ODE-based simulation of metabolism.

consciousness-council

Run a multi-perspective Mind Council deliberation on any question, decision, or creative challenge. Use this skill whenever the user wants diverse viewpoints, needs help making a tough decision, asks for a council/panel/board discussion, wants to explore a problem from multiple angles, requests devil's advocate analysis, or says things like "what would different experts think about this", "help me think through this from all sides", "council mode", "mind council", or "deliberate on this". Also trigger when the user faces a dilemma, trade-off, or complex choice with no obvious answer.

constitutional-ai

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.

cover-letter-to-editor

Write journal submission cover letters, presubmission inquiries, transfer requests, and reviewer suggestion or exclusion lists that help a manuscript get past editorial triage. Use when submitting a paper to a journal, pitching a manuscript to an editor (presubmission enquiry), transferring after rejection, explaining a resubmission or appeal, or preparing the editor-facing parts of a submission such as the significance pitch, suggested reviewers and conflicts. Not for job or fellowship application letters.

creative-thinking-for-research

Applies cognitive science frameworks for creative thinking to CS and AI research ideation. Use when seeking genuinely novel research directions by leveraging combinatorial creativity, analogical reasoning, constraint manipulation, and other empirically grounded creative strategies.

crewai-multi-agent

Multi-agent orchestration framework for autonomous AI collaboration. Use when building teams of specialized agents working together on complex tasks, when you need role-based agent collaboration with memory, or for production workflows requiring sequential/hierarchical execution. Built without LangChain dependencies for lean, fast execution.

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

database-lookup

Query documented public database APIs with explicit endpoints, filters, pagination, and provenance. Use when a scientific, regulatory, financial, or other database-backed fact must be retrieved reproducibly from a named source rather than inferred from general knowledge.

datalad

Retrieve, version, and publish scientific datasets with DataLad and git-annex, and capture computational provenance with datalad run, rerun, and containers-run. Use when cloning or fetching data from OpenNeuro, DANDI, datasets.datalad.org, or any DataLad dataset; when a file in a dataset reads as a broken symlink or a small pointer instead of real data; when an analysis needs a machine-readable record of how each output was produced so it can be re-executed; or when publishing a dataset to siblings such as a GitHub repository plus a storage remote. Also use to decide between DataLad and plain Git for a data-carrying repository.

datamol

Wraps RDKit through the datamol Python library (import datamol as dm) for molecular cheminformatics, returning native rdkit.Chem.Mol objects. Covers SMILES/SELFIES/InChI conversion, sanitization and standardization, descriptors, fingerprints and similarity, Butina clustering and diverse subset picking, Bemis-Murcko scaffolds, BRICS/RECAP fragmentation, 3D conformers, reactions, SDF/CSV/Excel I/O including cloud paths, and parallel batch processing. Use when parsing or standardizing SMILES from external sources, computing descriptors or fingerprints for a compound library, clustering or selecting diverse molecules, making scaffold-based train/test splits, or generating conformers. Use when running a molecule-processing pipeline with n_jobs. For fine-grained control or custom parameters, use RDKit directly instead.

deepchem

Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featurization options and MoleculeNet benchmarks. Best for quick experiments with pre-trained models, diverse molecular representations. For graph-first PyTorch workflows use torchdrug; for benchmark datasets use pytdc.

deepspeed

Covers DeepSpeed for distributed deep learning training and I/O: ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8 training, 1-bit Adam, sparse attention, and DeepNVMe (aio_handle, gds_handle, async_io and gds operators, ds_nvme_tune, ds_report) for fast transfers between NVMe storage and host or GPU tensors. Use when configuring ZeRO stages for large-model training, enabling mixed precision or 1-bit Adam, offloading parameters or optimizer state to NVMe with ZeRO-Infinity, writing or reading tensors to files with blocking or non-blocking DeepNVMe calls, or tuning NVMe I/O settings such as block size, queue depth and parallelism. Not for plain PyTorch DistributedDataParallel without DeepSpeed.

deeptools

Runs deepTools command-line programs on NGS alignment data: bamCoverage and bamCompare for BAM to bigWig/bedGraph with RPGC, CPM, RPKM or BPM normalization, multiBamSummary with plotCorrelation and plotPCA, plotFingerprint, computeMatrix with plotHeatmap and plotProfile, and alignmentSieve with --ATACshift. Use when converting BAM files to normalized coverage tracks. Use when checking ChIP-seq quality or comparing replicates. Use when plotting signal around TSS or peak regions. Use when comparing treatment versus control samples. Use when building ChIP-seq, RNA-seq or ATAC-seq coverage workflows. Not for peak calling or differential expression analysis.

depmap

Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), drug sensitivity data, and gene effect profiles. Use for identifying cancer-specific vulnerabilities, synthetic lethal interactions, and validating oncology drug targets.

dhdna-profiler

Extract cognitive patterns and thinking fingerprints from any text. Use this skill when the user wants to analyze how someone thinks, understand cognitive style, profile writing or speech patterns, compare thinking styles between people, asks "what's my thinking style", "analyze how this person reasons", "cognitive profile", "thinking pattern", "DHDNA", "digital DNA", or wants to understand the mind behind any text. Also trigger when the user provides text and wants deeper insight into the author's reasoning patterns, decision-making style, or cognitive signature.

diffdock

DiffDock and DiffDock-L molecular docking. Use for protein-small-molecule pose prediction from PDB or sequence plus SMILES/SDF/MOL2, batch docking, virtual screening, and pose-confidence interpretation. Not for binding affinity prediction.

distributed-llm-pretraining-torchtitan

Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.

dnanexus-integration

Build and operate reproducible genomics workloads on DNAnexus with the dx CLI, dxpy, apps/applets, native workflows, dxCompiler, and Nextflow. Use for DNAnexus data transfers, dxapp.json development, execution monitoring, workflow import, and project automation.

dspy

Builds and optimizes language model programs with DSPy (Stanford NLP), using Signatures, modules (Predict, ChainOfThought, ReAct, ProgramOfThought) and optimizers (BootstrapFewShot, MIPRO, BootstrapFinetune). Use when replacing hand-written prompts with declarative signatures, when automatically tuning prompts against training data and a metric, when building multi-stage RAG pipelines or ReAct agents, when building classifiers or structured-output programs, or when saving and evaluating optimized LM programs across providers such as Anthropic, OpenAI, or Ollama. Not for quick prototypes or simple chains; use manual prompting or LangChain instead.

elsevier-cas

Prepare submissions to Elsevier journals (including The Lancet family style notes, Cell-independent Elsevier titles, and thousands of society journals) using the elsarticle LaTeX class, the CAS single/double-column templates or Word, covering highlights, graphical abstracts, keywords, CRediT author statements, declarations of interest, generative-AI disclosure, data availability, reference styles per journal, Editorial Manager submission and "Your Paper Your Way" format-free first submission. Use when targeting an Elsevier journal or fixing elsarticle/CAS template issues.

esm

Covers the EvolutionaryScale/Biohub esm Python SDK: ESM3 generative protein design (sequence, structure and function tracks, chain-of-thought), ESMC embeddings, inverse folding, ESMFold2 structure prediction, and hosted Forge/Biohub inference with ESM_API_KEY. Use when generating or completing proteins with ESM3, extracting ESMC embeddings for classification or similarity, folding a sequence or designing sequences from a structure, or choosing between esm3-open, esmc_300m and Forge-hosted models. Use when batching async Forge or Biohub API calls. Not for other protein models such as AlphaFold or generic Hugging Face transformers usage.

etetoolkit

Analyze, manipulate, compare, annotate, and visualize phylogenetic or other hierarchical trees with ETE 4. Use for Newick/Nexus tree I/O, topology edits and pattern matching, Robinson-Foulds comparisons, gene-tree evolutionary events and reconciliation, NCBI/GTDB taxonomy, SmartView exploration, and publication rendering. Do not use it to infer trees from raw sequences; align sequences and infer a tree first.

evaluating-code-models

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

evaluating-cosmos-policy

Evaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting up cosmos-policy for robot manipulation evaluation, running headless GPU evaluations with EGL rendering, or profiling inference latency on cluster or local GPU machines.

evaluating-llms-harness

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

evolving-ai-agents

Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, optimizing agent prompts and skills against benchmarks, or implementing automated agent evaluation loops.

exa-search

Web toolkit powered by Exa, tuned for scientific and technical content. Use this skill when the user needs to search the web or fetch/extract URL content. Covers: web search (semantic lookups, research, current info — with optional research-paper category and academic domain filtering) and URL extraction (fetching pages, articles, academic PDFs in batch). Use this skill for web-related tasks when the user wants high-quality search or scholarly filtering via category=research paper. Triggers on requests to search, look up, fetch a page, or extract an article.

experiment-tracking-swanlab

Tracks ML experiments with SwanLab, an open-source tool covering swanlab.init, config and metric logging, scalar charts, and media logging (images, audio, text, GIFs, point clouds, molecules), in cloud, local (mode="local" with swanlab[dashboard]), or self-hosted setups. Includes integrations for PyTorch, Transformers, PyTorch Lightning, and Fastai. Use when logging metrics and hyperparameters for training runs, comparing runs across seeds, checkpoints, or hyperparameters, running an offline or self-hosted dashboard instead of managed SaaS, logging media alongside scalars, or adding tracking callbacks to a Transformers, Lightning, or Fastai training loop. Not for other trackers such as Weights & Biases or MLflow.

experimental-design

Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable. Use whenever someone is planning a study, asks how to assign subjects/samples to groups, mentions randomization, blocking, stratification, controls, factorial or fractional-factorial designs, design of experiments (DOE), screening many factors, response-surface optimization, crossover or repeated-measures or split-plot designs, cluster/group randomization, Latin squares, plate layouts, batch/run-order effects, replication vs. pseudoreplication, or sequential/adaptive/group-sequential designs. Trigger even for informal phrasings like "how should I set up this experiment", "how do I avoid confounding", "what's the best way to test these 6 factors", or "assign these mice to conditions". For computing the sample size or power once the design is chosen, use statistical-power; for analyzing data already collected, use statistical-analysis.

exploratory-data-analysis

Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain formats are reference-only and unknown formats fail closed.

faiss

Facebook's library for efficient similarity search and clustering of dense vectors. Supports billions of vectors, GPU acceleration, and various index types (Flat, IVF, HNSW). Use for fast k-NN search, large-scale vector retrieval, or when you need pure similarity search without metadata. Best for high-performance applications.

fictiv

Drives the Fictiv on-demand manufacturing web app (app.fictiv.com) in the user's browser, since there is no public API. Covers uploading CAD files (STEP, SLDPRT, STL), configuring process, material, finish, threads and tolerances, reading DFM feedback, instant or manual quotes, lead-time tiers, checkout by card or PO, order tracking and reorders. Use when the user mentions Fictiv. Use when they want a part CNC machined, 3D printed, sheet-metal fabricated, cast, or molded online. Use when they ask to get a quote or order parts from a STEP or STL file. Use when they need to check a Fictiv quote or order status. Use when a Fictiv upload, DFM warning, price or checkout fails. Not for ITAR-controlled parts, which Fictiv does not accept.

fine-tuning-openvla-oft

Fine-tunes and evaluates OpenVLA-OFT and OpenVLA-OFT+ policies for robot action generation with continuous action heads, LoRA adaptation, and FiLM conditioning on LIBERO simulation and ALOHA real-world setups. Use when reproducing OpenVLA-OFT paper results, training custom VLA action heads (L1 or diffusion), deploying server-client inference for ALOHA, or debugging normalization, LoRA merge, and cross-GPU issues.

fine-tuning-serving-openpi

Fine-tune and serve Physical Intelligence OpenPI models (pi0, pi0-fast, pi0.5) using JAX or PyTorch backends for robot policy inference across ALOHA, DROID, and LIBERO environments. Use when adapting pi0 models to custom datasets, converting JAX checkpoints to PyTorch, running policy inference servers, or debugging norm stats and GPU memory issues.

fine-tuning-with-trl

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

firecrawl-research-index

Query Firecrawl Research Index paper endpoints for topic discovery, source metadata, question-matched passages, and citation-neighbor expansion. Use when a researcher explicitly asks for Firecrawl Research or its hosted arXiv and biomedical paper records.

flowio

Read, inspect, and write Flow Cytometry Standard (FCS) 2.0, 3.0, and 3.1 files with FlowIO. Use for low-level FCS metadata and channel inspection, NumPy event extraction, multi-dataset files, table export, and FCS 3.1 creation; use FlowKit for compensation, cytometry transforms, gating, or FlowJo workspaces.

fluidsim

Plan, configure, inspect, restart, and analyze bounded FluidSim computational-fluid-dynamics simulations with explicit numerical-validity and HPC safety checks. Use for FluidSim solver selection, parameter review, FFT/MPI setup, output diagnostics, or restart compatibility.

folklore-variant-evidence

Retrieve ClinGen gene-disease validity assertions for a public gene or disease, and review source-linked public evidence and literature for one supported GRCh38 germline nuclear SNV or simple indel through Folklore Clinical Variant Interpretation MCP. Use when a scientific agent must branch deterministically on resolved, ambiguous, not-found, invalid, unsupported, or unavailable variant outcomes; chain a resolved public variant into related literature or publication details; or preserve evidence provenance without accepting patient, phenotype, family, segregation, or private case data.

generate-image

Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow). Use for photos, illustrations, artwork, concept art, visual assets, logos, and image editing or compositing from reference images. For flowcharts, circuits, pathways, and other technical diagrams, use the scientific-schematics skill instead.

geniml

Plans and audits local genomic-interval machine learning workflows with Geniml (0.8.4) and Gtars: validates BED files against chromosome sizes and assembly contracts, plans Region2Vec, scEmbed, and BEDspace runs, checks model, tokenizer, and universe.bed compatibility, and assesses consensus universes (CC, CCF, ML, HMM, assess-universe). Bundled scripts only validate or plan; they do not train. Use when checking BED coordinates, contigs, or assembly before analysis. Use when planning Region2Vec or scEmbed training on tokenized Parquet data. Use when verifying that a checkpoint, config.yaml, and universe match. Use when building or assessing a consensus peak universe. Use when reviewing BEDbase or Hugging Face download risks. Not for general BED manipulation or peak calling; use bedtools or similar tools.

genomic-coordinates

Convert genomic intervals between coordinate conventions, normalise and compare variant representations, and detect assembly or contig-naming mismatches before they corrupt an analysis. Use whenever coordinates cross a format, tool, or assembly boundary - converting between BED, GFF/GTF, VCF, SAM/BAM, WIG, PSL, genePred, Picard interval_list, or region strings; reconciling 0-based half-open with 1-based inclusive; left-aligning or trimming indels; checking whether two variant records describe the same change; mapping genomic to transcript, CDS, or protein positions; auditing a BED/GTF/VCF for convention violations; or diagnosing GRCh37 vs hg19 vs GRCh38 vs T2T, chr-prefix, and liftover problems. Triggers include "off by one", "0-based", "1-based", "half-open", "coordinate system", "left-align", "normalize variant", "bcftools norm", "chr prefix", "wrong genome build", "liftover", "REF mismatch", and "HGVS".

genomic-intelligence

Predict regulatory features, gene structure, and expression directly from DNA sequence using Genomic Intelligence's hosted transformer DNA language models — no local GPU or model weights. Six tasks over a REST API and a hosted MCP server (keyless public demo): promoter regions, splice donor/acceptor sites, enhancer activity, chromatin state, sequence-to-expression (log TPM), and de-novo gene annotation, plus a composite find-genes-then-predict-expression workflow. Use when the user has a gene symbol, a genomic region, or a DNA/FASTA sequence and wants any of these predictions, mentions Genomic Intelligence, genomicintelligence.ai, api.genomicintelligence.ai, or mcp.genomicintelligence.ai.

geomaster

Provides geospatial and Earth observation workflows using GeoPandas, Rasterio, GDAL, Xarray, Shapely, Laspy, PDAL, Google Earth Engine, and STAC/Planetary Computer, covering Sentinel, Landsat, MODIS, SAR and hyperspectral imagery, spectral indices, terrain and network analysis, point clouds, COGs, CRS handling, and spatial ML, with code in Python, R, Julia, JavaScript, C++, Java, Go, and Rust. Use when computing NDVI or other indices from satellite imagery, when running vector overlays, reprojection, or spatial statistics on shapefiles, GeoJSON, or GeoPackage data, when searching STAC catalogs and reading cloud-optimized GeoTIFFs, when classifying land cover or training ML models on Earth observation data, or when doing terrain, hydrology, or point cloud analysis. Not for general non-spatial data analysis or tabular ML.

geopandas

Guidance and local audit CLIs for Python workflows using GeoPandas 1.1.4 GeoSeries and GeoDataFrame for planar vector data: CRS handling, geometry validity and repair, sjoin, overlay, clip, dissolve, union, GeoPackage, GeoParquet and PostGIS I/O, and plotting. Use when reprojecting or assigning a CRS, auditing invalid or empty geometries, running spatial joins and checking cardinality, exporting vector data with a reproducible contract, or migrating code from GeoPandas 0.14 to 1.x. Use when checking privacy risks in coordinate data. Do not use for raster data or geodesic-only calculations.

get-available-resources

Detect host inventory and effective CPU, memory, disk, scheduler, container, and accelerator limits when a user asks for resource-aware planning or before a clearly resource-sensitive local workload. Produces a redacted JSON snapshot and conservative planning helpers without stress tests or assuming visible host hardware is usable.

gget

Queries 20+ bioinformatics databases and analysis services through the gget CLI and Python package, covering Ensembl gene search, info and sequences (ref, search, info, seq), BLAST, BLAT, MUSCLE, DIAMOND, PDB, AlphaFold, ELM, ARCHS4, CELLxGENE, Enrichr, Bgee, OpenTargets, cBioPortal, COSMIC, viral sequence downloads (virus), and 8cube mouse specificity and expression data. Use when looking up gene or transcript details, running a quick BLAST/BLAT search, fetching AlphaFold or PDB structures, running enrichment analysis on a gene list, downloading viral sequences with filters, or exploring disease and drug associations interactively. Not for batch processing or fine-grained BLAST control (use biopython) or multi-database Python pipelines (use bioservices).

gguf-quantization

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.

ginkgo-cloud-lab

Submit and manage protocols on Ginkgo Bioworks Cloud Lab (cloud.ginkgo.bio), a web-based interface for autonomous lab execution on Reconfigurable Automation Carts (RACs). Use when the user wants to run protein expression and purification (cell-free, E. coli, or Pichia), HiBiT or A280 or LabChip quantification, IVT mRNA/circRNA synthesis, thermal shift / developability assays, Echo-MS enzyme or analyte methods, SPR target onboarding, fluorescent pixel art, or otherwise interact with Ginkgo Cloud Lab services. Covers protocol selection, input preparation, pricing, and ordering workflows.

gptq

Quantizes LLMs to 4-bit (also 3-bit) with GPTQ using group-wise quantization (group size 128 by default), via AutoGPTQ and transformers. Covers loading pre-quantized GPTQ models, quantizing your own model, choosing group size, selecting ExLlamaV2, Marlin, or Triton kernels, and QLoRA fine-tuning with PEFT. Use when fitting 70B-class models onto limited or consumer GPUs, cutting memory about 4x versus FP16, speeding up inference, finding pre-quantized checkpoints on HuggingFace, or fine-tuning a quantized model with LoRA. For slightly better accuracy on newer GPUs use AWQ instead, and for simple 8-bit or on-the-fly quantization use bitsandbytes.

grpo-rl-training

Guides GRPO (Group Relative Policy Optimization) fine-tuning of language models with the TRL library, including GRPOTrainer configuration, composing multiple reward functions (correctness, format, length, style), dataset prep in chat format, Unsloth setup, LoRA merging, and monitoring reward, reward_std and KL. Use when training a model to follow a strict output format such as XML or JSON. Use when teaching verifiable tasks like math or code with objective correctness rewards. Use when improving chain-of-thought reasoning with custom reward functions. Use when debugging flat rewards, mode collapse, or OOM in GRPO runs. Do not use for plain supervised fine-tuning (use SFT) or when high-quality preference pairs exist (use DPO or PPO).

gtars

Inspects and plans work with Gtars, the Rust/Python/CLI toolkit for genomic intervals: BED RegionSet set algebra (reduce, setdiff, intersect, closest, cluster, gaps), overlap counts, consensus peak sets, coverage, tokenizers, fragment processing, and refget/BEDbase stores. Use when merging or comparing BED interval sets, building a consensus or universe from peaks, tokenizing regions for ML models, working with fragment files, or planning refget sequence-collection and BEDbase cache use. Use when checking BED coordinates, contigs, or strand before running gtars. Not for general BAM/VCF processing; use samtools or bedtools for that.

guidance

Constrains LLM output during generation with Guidance (Microsoft Research), using regex, select() choices, context-free grammars, token healing, and @guidance functions, with Anthropic, OpenAI, Transformers, and llama.cpp backends. Use when you need generated text to match a regex or fixed format such as dates, emails, or IDs. Use when you need guaranteed valid JSON, XML, or code from a model. Use when building multi-step generation workflows or ReAct-style agents with Python control flow. Use when classifying text into fixed categories with select(). Use when running local models and wanting grammar constraints. Not for Pydantic validation with automatic retries; use Instructor for that.

histolab

Extracts tiles and preprocesses H&E whole slide images with the histolab Python library (OpenSlide), covering slide inspection, tissue masks (TissueMask, BiggestTissueBoxMask), RandomTiler, GridTiler and ScoreTiler extraction, image and morphological filters, and Macenko or Reinhard stain normalization. Use when building a tile dataset from WSI files for deep learning. Use when detecting tissue and excluding background or pen annotations. Use when previewing tile locations and exporting tile CSV reports. Use when normalizing stain variation across slides. Not for spatial proteomics, multiplexed imaging, or deep learning pipelines; use pathml for those.

hqq-quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.

hugging-science

Use when the user is doing AI/ML work in a scientific domain such as biology, chemistry, physics, astronomy, climate, genomics, materials, medicine, ecology, energy, engineering, math, drug discovery, protein design, weather modeling, theorem proving, single-cell, or PDE solving. Hugging Science is a curated catalog of scientific datasets, models, blog posts, and interactive Spaces. This skill helps discover and use resources via datasets, transformers, the HF Inference API, gradio_client, and methodology citations.

huggingface-accelerate

Wraps existing PyTorch training scripts with HuggingFace Accelerate (Accelerator class, accelerate config, accelerate launch) so the same code runs on CPU, single GPU, multi-GPU, multi-node, TPU, or Apple MPS. Covers device placement, FP16/BF16/FP8 mixed precision, gradient accumulation, distributed checkpointing, and switching between DDP, DeepSpeed ZeRO, FSDP, and Megatron backends. Use when converting a single-GPU script to multi-GPU, enabling mixed precision, configuring DeepSpeed ZeRO or FSDP, or setting up gradient accumulation. Use when one script must run on different hardware. Not for callback-heavy training loops (use PyTorch Lightning) or multi-node orchestration with hyperparameter tuning (use Ray Train).

huggingface-tokenizers

Provides the HuggingFace Tokenizers library (Rust core with Python and Node.js bindings) for training and using BPE, WordPiece, and Unigram tokenizers. Covers the pipeline of normalizers, pre-tokenizers, models, post-processors and decoders, padding and truncation, batch encoding, alignment tracking, and conversion to transformers PreTrainedTokenizerFast. Use when training a custom tokenizer or vocabulary on a new corpus, tokenizing large text corpora quickly, mapping tokens back to character offsets for NER or question answering, or configuring normalization and special tokens. Use when wrapping a custom tokenizer for transformers. For SentencePiece models or tiktoken, use those tools instead; for loading a pretrained tokenizer only, AutoTokenizer is enough.

hypogenic

Plans and audits use of ChicagoHAI HypoGeniC/HypoRefine for LLM-assisted hypothesis generation from labeled text datasets. Use for the hypogenic package, its task configs, hypothesis banks, or HypoBench datasets—not for manual hypothesis formulation or scientific validation.

hypothesis-generation

Formulate evidence-bounded scientific questions, candidate hypotheses, rival explanations, causal or associational claims, discriminating predictions, measurements, and preregistration-ready analysis plans. Use when turning observations or preliminary findings into transparent, testable research plans without treating hypotheses as facts.

ieee-transactions

Format and submit papers to IEEE journals (Transactions, Journals, Letters, IEEE Access) and IEEE conferences using the IEEEtran LaTeX class or Word templates, covering journal vs conference modes, abstracts and index terms, numbered IEEE reference style, figures and equations, author biographies, page limits and overlength charges, IEEE PDF eXpress, copyright forms, double-blind options and ScholarOne/IEEE Author Portal submission. Use when writing for any IEEE venue, fixing IEEEtran layout issues, or checking IEEE submission compliance.

imaging-data-commons

Query and download public cancer imaging data from NCI Imaging Data Commons. Invoke for any question about IDC collections, cancer imaging datasets, DICOM data access, radiology (CT, MR, PET) or pathology AI training sets, metadata queries, visualization, or license checks — even when the user doesn't explicitly mention "IDC". No authentication required.

implementing-llms-litgpt

Implements and trains LLMs using Lightning AI's LitGPT with 20+ pretrained architectures (Llama, Gemma, Phi, Qwen, Mistral). Use when need clean model implementations, educational understanding of architectures, or production fine-tuning with LoRA/QLoRA. Single-file implementations, no abstraction layers.

infographics

Generates infographics from natural-language prompts using Nano Banana Pro image generation, with optional Perplexity Sonar research for facts and a Gemini 3.6 Flash quality review that regenerates only when the score is below a document-type threshold. Supports 10 types (statistical, timeline, process, comparison, list, geographic, hierarchical, anatomical, resume, social), 8 industry styles, and colorblind-safe palettes (wong, ibm, tol). Use when turning data or statistics into a visual summary. Use when building timelines, process guides, or side-by-side comparisons. Use when making geographic or organizational chart graphics. Use when producing social media or marketing visuals. Not for technical flowcharts, pathway diagrams, or CONSORT/PRISMA diagrams; use scientific-schematics instead.

instructor

Extracts structured, validated data from LLM responses using the Instructor Python library with Pydantic response models, including nested models, enums, custom validators, automatic retries with validation error feedback, and streaming of partial objects or iterables. Works with Anthropic, OpenAI, and local Ollama models. Use when pulling typed fields or entities out of free text, when classifying text into fixed categories, when an LLM must return JSON that passes schema validation, when failed extractions need automatic retry, or when streaming partial structured results. Not for prompt optimization (use DSPy) or building multi-step chains (use LangChain).

iso-standards-readiness

Prepares and structurally reviews readiness evidence for ISO management-system and laboratory-competence standards - ISO 13485 medical device QMS, ISO 14971 device risk management, ISO/IEC 17025 testing and calibration laboratories, and ISO 15189 medical laboratories. Use when organizing declared scope, controlled documents, risk-management files, scope of accreditation, traceability, CAPA, external-provider controls, or bounded local evidence manifests, and when separating ISO certification from laboratory accreditation, FDA QMSR inspection, CLIA certification, MDSAP, and EU MDR/IVDR evidence boundaries. Not for legal applicability, compliance, certification, or accreditation decisions; contains no clause text.

knowledge-distillation

Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller models with retained performance, transferring GPT-4 capabilities to open-source models, or reducing inference costs. Covers temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies.

lab-hardware-cad

Design custom laboratory hardware as parametric build123d models and export fabrication-ready STEP, STL, and DXF files - microfluidic chips and molds, optomechanical mounts and breadboard adapters, cuvette and microplate holders, tube racks, animal-behavior rigs, and 3D-printed instrument fixtures. Use when a research task needs a physical part that must mate with standardized labware, an optical table, a cage system, or a printer, CNC, or laser process.

labarchive-integration

Securely integrate with the official LabArchives ELN REST-like API and Inventory API v1. Use for regional endpoint selection, signed-request construction, user authorization and UID flows, local LA container validation, and verified LabArchives integration workflows.

lambda-labs-gpu-cloud

Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.

lamindb

Use when working with LaminDB, the open-source lineage-native lakehouse for biological datasets and models. Covers setup, artifact registration, query/search, lineage tracking, validation, ontology-backed annotation with Bionty, collections, branches, storage, and workflow integrations.

langchain

Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments.

langsmith-observability

LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.

latchbio-integration

Build, register, debug, and operate bioinformatics workflows on Latch using the Python SDK, CLI, Latch Data and Registry, Nextflow, Snakemake, programmatic execution, and Latch MCP. Use when authoring or deploying Latch workflows, configuring resources or interfaces, moving data, integrating Registry, or launching and monitoring runs.

latex-posters

Creates research posters in LaTeX with beamerposter, tikzposter, or baposter, covering page sizes (A0, A1, 36x48 inch), multi-column layouts, color schemes, figure and QR code placement, and pdflatex/overfull-box checks before printing. Also guides generating poster graphics with AI schematic tools under strict element and word limits. Use when preparing a conference or symposium poster, converting a paper into a poster, building a department poster template, or fitting a conference's size rules. Use when fixing overflow or unreadable text on a compiled poster. For slide decks, use a presentation skill instead.

liteparse

Local document and PDF parsing that returns spatial text with bounding boxes. Use for extracting text from PDFs, DOCX, Office files, and images; running OCR on scans; producing layout-preserved JSON for RAG; batch-ingesting folders of papers; or rendering pages to PNG for multimodal agents. Distinguishing capabilities are per-token bounding boxes, page raster output, and fully local processing with no cloud API.

literature-review

Runs systematic literature reviews by searching PubMed, arXiv, bioRxiv, and Semantic Scholar (plus web search via parallel-cli), screening studies, extracting data, synthesizing themes, and verifying citations with verify_citations.py. Produces a markdown and PDF review with a PRISMA-style flow diagram and a bibliography in APA, Nature, or Vancouver style. Use when conducting a systematic or scoping review, synthesizing research on a topic, writing the literature review section of a paper or thesis, identifying research gaps, or checking citations in a review. Not for looking up a single paper or database record; use the specific database skill instead.

llama-cpp

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

llama-factory

Guides fine-tuning of large language models with LLaMA-Factory, covering the WebUI no-code interface, training across 100+ supported models, quantized QLoRA at 2/3/4/5/6/8-bit precision, and multimodal model support. Use when setting up or configuring a LLaMA-Factory fine-tuning run, choosing a quantization level for QLoRA training, using the WebUI instead of writing training code, fine-tuning a multimodal model, or debugging a LLaMA-Factory training error. Not for general-purpose inference serving or for other training frameworks; use the matching skill for those.

llamaguard

Classifies LLM prompts and responses as safe or unsafe using Meta's LlamaGuard (7B v1, 8B v2 and v3) across six categories: violence and hate, sexual content, guns and illegal weapons, regulated substances, suicide and self-harm, and criminal planning. Covers HuggingFace Transformers and vLLM inference, a FastAPI moderation endpoint, Sagemaker deployment, and NeMo Guardrails integration. Use when filtering user prompts before an LLM call, moderating model outputs before display, serving a self-hosted moderation API, or reducing latency and GPU memory of a moderation model with vLLM or quantization. Use when tuning thresholds to cut false positives. For a simpler hosted option, use the OpenAI Moderation API instead.

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications.

llava

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

long-context

Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs.

mamba-architecture

Explains how to use Mamba selective state-space models (state-spaces/mamba package, Mamba-1 with d_state=16 and Mamba-2 with multi-head structure and d_state=128) for linear-time sequence modeling. Covers installation, the Mamba block, building a language model with MambaLMHeadModel, loading pretrained state-spaces checkpoints (130M to 2.8B) from HuggingFace, and benchmarking against Transformers. Use when implementing or loading Mamba models, processing very long sequences without a KV cache, building streaming applications, choosing between Mamba-1 and Mamba-2, or fixing install and CUDA memory problems. Requires Linux and an NVIDIA GPU. Do not use for standard Transformer models, or for RWKV, RetNet, or Hyena architectures.

markdown-mermaid-writing

Writes scientific documents and documentation as markdown with embedded Mermaid diagrams as the canonical, git-diffable source format. Covers a markdown style guide, a Mermaid style guide, 24 Mermaid diagram type references (flowchart, sequence, ER, Gantt, state, mindmap, radar, sankey, XY chart, and others), and 9 templates (research paper, status report, decision record, how-to, pull request, issue, kanban, presentation, project documentation). Use when writing a report, manuscript, README, or analysis in markdown. Use when adding a workflow, pipeline, timeline, or architecture diagram to a document. Use when choosing the right Mermaid diagram type or fixing Mermaid syntax such as radar-beta or xychart-beta. Use when starting from a document template or checking citations and heading rules. Not for scatter plots of real data or photorealistic images; use Python plotting or scientific-schematics for those.

market-research-reports

Build evidence-traceable market research reports and assumption-driven market sizing or forecast scenarios. Use for market definition, industry and customer evidence, competitive landscapes, TAM/SAM/SOM reconciliation, forecast sensitivity, and auditable report scaffolds.

markitdown

Converts documents to Markdown with Microsoft MarkItDown (Python API, markitdown CLI, markitdown-ocr plugin, markitdown-mcp server), covering PDF, Word, PowerPoint, Excel, HTML, CSV, EPUB, ZIP, and streams, plus Azure extraction. Use when preparing papers or reports for LLM/RAG ingestion; when batch-converting a literature folder to Markdown with provenance; when converting uploaded bytes or file streams; when OCR of scanned PDFs or image text is needed; when exposing conversion to a local agent through MCP. Do not use for bounding boxes or page coordinates (use LiteParse) or PDF merge/split/forms (use the pdf skill).

matchms

Process, clean, compare, and search tandem mass spectra with matchms. Use for MS/MS file I/O, metadata harmonization, peak filtering, spectral similarity, library matching, score matrices, and molecular-similarity networks. Use pyopenms instead for LC-MS feature detection or proteomics pipelines.

matlab

Designs, reviews, and migrates MATLAB R2026a and GNU Octave numerical code, covering functions with arguments blocks, arrays and indexing, tables and timetables, matlab.unittest tests, MATLAB Projects, exportgraphics figures, MAT file inventory, and MATLAB-Python interoperability. Includes static helper scripts that scan .m files, plan argv commands, and hash artifacts without executing anything. Use when writing or reviewing .m code, migrating code between MATLAB releases or to Octave, setting up reproducible tests and projects, inspecting an untrusted MAT or .mlx file safely, or checking Python Engine compatibility. Do not use for pure Python numerical work with NumPy or SciPy.

matplotlib

Low-level plotting library for full customization. Use when you need fine-grained control over every plot element, creating novel plot types, or integrating with specific scientific workflows. Export to PNG/PDF/SVG for publication. For quick statistical plots use seaborn; for interactive plots use plotly; for publication-ready multi-panel figures with journal styling, use scientific-visualization.

medchem

Filters and triages small-molecule libraries with the Python medchem library (datamol-io, v2.0.5) on top of RDKit and datamol. Covers drug-likeness rules (Lipinski, Veber, CNS, lead-like) via RuleFilters, structural alerts (ChEMBL-derived sets, NIBR, PAINS, Brenk), chemical group detection, ZINC-based complexity thresholds, scaffold constraints, and the medchem query language (QueryFilter). Use when screening a compound library for drug-likeness, removing PAINS or other structural-alert compounds, prioritizing hits for hit-to-lead or lead optimization, detecting functional groups, or combining criteria in one query. Not for computing general descriptors or fingerprints; use RDKit directly.

miles-rl-training

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.

ml-paper-writing

Write publication-ready ML/AI papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Use when drafting papers from research repos, structuring arguments, verifying citations, or preparing camera-ready submissions. For systems venues (OSDI, NSDI, ASPLOS, SOSP), use systems-paper-writing instead.

ml-training-recipes

Battle-tested PyTorch training recipes for all domains — LLMs, vision, diffusion, medical imaging, protein/drug discovery, spatial omics, genomics. Covers training loops, optimizer selection (AdamW, Muon), LR scheduling, mixed precision, debugging, and systematic experimentation. Use when training or fine-tuning neural networks, debugging loss spikes or OOM, choosing architectures, or optimizing GPU throughput.

mlflow

Tracks machine learning experiments and manages model lifecycles with MLflow, covering mlflow.log_param, log_metric and log_artifact, autologging for scikit-learn, PyTorch Lightning, XGBoost and HuggingFace Transformers, the Model Registry with versions and stage transitions, run searching, and local or cloud model serving. Use when logging parameters, metrics and artifacts for training runs, comparing runs across experiments, registering and promoting model versions from Staging to Production, serving a logged model for inference, or reproducing an experiment from an MLflow project. Not for hyperparameter-sweep dashboards or general data versioning; use a dedicated tool for those.

modal

Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. Use when deploying or serving AI/ML models, running GPU-accelerated workloads (training, fine-tuning, inference), serving web endpoints, scheduling batch jobs, or scaling Python code to cloud containers with the Modal SDK.

modal-serverless-gpu

Serverless GPU cloud platform for running ML workloads. Use when you need on-demand GPU access without infrastructure management, deploying ML models as APIs, or running batch jobs with automatic scaling.

model-merging

Merge multiple fine-tuned models using mergekit to combine capabilities without retraining. Use when creating specialized models by blending domain-specific expertise (math + coding + chat), improving performance beyond single models, or experimenting rapidly with model variants. Covers SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies.

model-pruning

Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods.

moe-training

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures like Mixtral 8x7B or DeepSeek-V3, or scaling model capacity without proportional compute increase. Covers MoE architectures, routing mechanisms, load balancing, expert parallelism, and inference optimization.

molecular-dynamics

Runs and analyzes molecular dynamics simulations using OpenMM and MDAnalysis. Covers system preparation with PDBFixer and OpenFF/GAFF2, force field choice (AMBER14, CHARMM36m, ff19SB), energy minimization, NVT/NPT equilibration, production MD, and trajectory analysis (RMSD, RMSF, protein-ligand contacts). Use when simulating a protein or protein-ligand system on GPU, assessing how a mutation affects protein dynamics, characterizing ligand binding mode, quantifying per-residue flexibility, or modeling membrane proteins and disordered proteins. Not for GROMACS or NAMD workflows, which the skill lists only as alternatives.

molfeat

Converts SMILES strings or RDKit/datamol molecules into numerical features using molfeat (0.11.0), which provides calculators, scikit-learn compatible transformers, and pretrained embedding models. Covers fingerprints (ECFP, MACCS, MAP4), RDKit and Mordred descriptors, pharmacophore descriptors, and pretrained models such as ChemBERTa and GIN, with parallel processing and caching. Use when building QSAR or QSPR models from SMILES. Use when choosing among molecular featurizers for a property prediction task. Use when generating embeddings for virtual screening, similarity search, or chemical space clustering. Use when adding a featurizer to a scikit-learn pipeline. Not for general cheminformatics tasks such as reading or editing structures, which belong to RDKit itself.

nanogpt

Provides nanoGPT, Karpathy's minimal PyTorch GPT implementation (model.py and train.py), with workflows for training character-level Shakespeare on CPU, reproducing GPT-2 124M on OpenWebText with multi-GPU torchrun, fine-tuning pretrained GPT-2 checkpoints, and training on custom text. Use when learning how GPT and transformer blocks work from scratch. Use when running a small training experiment on CPU or a single GPU. Use when modifying a transformer variant in plain PyTorch. Use when preparing character-level or BPE binary datasets. Use when troubleshooting out-of-memory or slow training in nanoGPT. Not for production or large-scale distributed training; use HuggingFace Transformers or Megatron-LM instead.

nature-portfolio

Prepare manuscripts for Nature and Nature Portfolio journals (Nature, Nature Communications, Nature Methods, Nature Biotechnology, Scientific Reports and other Nature-branded titles), covering article types, summary paragraph vs abstract, main-text and display-item limits, Methods and Extended Data, reporting summaries, data/code availability, figure preparation, references and presubmission enquiries. Use when targeting any Nature Portfolio journal, reformatting a paper for one, or checking a submission against its guidelines.

ncats-arax

Queries the NCATS Translator ARAX production API for bounded, typed, provenance-rich one-hop and endpoint-pinned two-hop biomedical knowledge-graph relationships. Use for Biolink-constrained RTX-KG2 lookup, explicit selected-provider ARAX federation, separate entity normalization, qualifier-aware graph traversal, and inspection of TRAPI edge bindings, publications, and knowledge-source provenance. Do not use for inference, ranking, open-ended pathfinding, clinical guidance, or sensitive queries.

nemo-curator

GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with RAPIDS. Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora.

nemo-evaluator-sdk

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

nemo-guardrails

Adds runtime safety rails to LLM applications with NVIDIA NeMo Guardrails, configured through Colang 2.0 flows. Covers jailbreak and prompt-injection detection, self-check input/output validation, retrieval-based fact-checking, hallucination detection, PII filtering via Presidio, toxicity detection via ActiveFence, and LlamaGuard integration. Use when adding programmable safety rules to a production LLM app, blocking jailbreaks or prompt injection, filtering PII from model inputs and outputs, verifying generated claims against retrieved sources, or tuning false positives and latency of guardrail checks. For standalone moderation only, use LlamaGuard or the OpenAI Moderation API instead.

networkx

Create, analyze, and visualize complex networks and graphs in Python with NetworkX. Use when working with network/graph data structures, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks (random, scale-free, small-world), reading/writing graph file formats, or drawing network topologies. Common applications include social, biological, transportation, and citation networks.

neurokit2

Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.

neuropixels-analysis

Analyze Neuropixels extracellular recordings end-to-end with SpikeInterface. Covers loading SpikeGLX/Open Ephys/NWB data, preprocessing, drift/motion correction, Kilosort4 (and CPU) spike sorting, quality metrics, and unit curation (threshold-based, model-based UnitRefine, and AI-assisted visual review). Use when working with Neuropixels 1.0/2.0 recordings, spike sorting, or extracellular electrophysiology analysis.

nextflow

Build, run, and debug Nextflow data pipelines and nf-core workflows end to end. Use whenever the user mentions Nextflow, nf-core, .nf files, nextflow.config, DSL2, processes/channels/operators, samplesheets, or wants to run a community pipeline (e.g. nf-core/rnaseq, nf-core/sarek), write or test a module/subworkflow with nf-test, configure executors/containers (Docker, Singularity/Apptainer, Conda, Wave), scale a workflow to HPC/SLURM or cloud (AWS Batch, Google Batch, Azure, Kubernetes), or debug a failed/-resume run. Make sure to use this skill for any reproducible scientific/bioinformatics workflow work even if the user does not say the word "Nextflow", and for authoring nf-core-compliant pipelines, modules, configs, and linting.

nnsight-remote-interpretability

Provides guidance for interpreting and manipulating neural network internals using nnsight with optional NDIF remote execution. Use when needing to run interpretability experiments on massive models (70B+) without local GPU resources, or when working with any PyTorch architecture.

omero-integration

Securely inspect and automate microscopy data workflows against OMERO.server with omero-py, BlitzGateway, OMERO CLI, tables, annotations, ROIs, rendering, and documented OMERO.web APIs. Use for scoped OMERO inventory, metadata export, import/export planning, or reviewed write workflows.

onekgpd

Query the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Use when a question is about individuals or variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene or region, which individuals are homozygous-reference at a position, which variants exist in the dataset or carried by specified individuals in a gene or region, the relatedness between two specified individuals. Variants are returned with 1000 Genomes allele frequencies (AF), gnomAD v4.1 exome and genome AF, AlphaMissense score, and HGVSp annotations.

ontology-term-resolution

Resolve free-text scientific labels to ontology term IDs and validate existing CURIEs against the EBI Ontology Lookup Service (OLS4). Also look up prefixes in Bioregistry, resolve compact identifiers via Identifiers.org, map lab shorthand with ZOOMA, and build Ontobee term pages. Use whenever an ontology identifier must be produced or checked - annotating tissue, cell type, disease, phenotype, assay, chemical, organism, sex, or developmental stage fields; preparing metadata for GEO, ENA, BioSamples, CELLxGENE, HCA, or ISA-Tab submission; auditing a metadata table of term IDs; checking whether a term is obsolete and what replaced it; or deciding HPO vs HP. Triggers include "ontology term", "ontology ID", "CURIE", "controlled vocabulary", "UBERON", "CL:", "MONDO", "HPO", "EFO", "ChEBI", "NCBITaxon", "GO term", "PATO", "Zooma", "Bioregistry", "Identifiers.org", "Ontobee", "annotate this tissue/cell type/disease", and any request to emit or verify an identifier shaped like PREFIX:0001234.

open-notebook

Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. Use when organizing research materials into notebooks, ingesting diverse content sources (PDFs, videos, audio, web pages, Office documents), generating AI-powered notes and summaries, creating multi-speaker podcasts from research, chatting with documents using context-aware AI, searching across materials with full-text and vector search, or running custom content transformations. Supports 16+ AI providers including OpenAI, Anthropic, Google, Ollama, Groq, and Mistral with complete data privacy through self-hosting.

openpiv

Particle Image Velocimetry (PIV) analysis with OpenPIV. Use when extracting velocity fields from PIV image pairs, analyzing fluid dynamics or flow visualization experiments, cross-correlating interrogation windows, validating and replacing spurious PIV vectors, or computing vorticity, strain rate, and turbulence statistics from measured velocity fields.

openrlhf-training

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

opentrons-integration

Author, review, migrate, simulate, and troubleshoot official Opentrons Python Protocol API v2 protocols for Flex and OT-2 robots. Use for robot-specific liquid handling, deck and labware setup, pipettes, modules, runtime parameters, liquid classes, and Opentrons App analysis. Use pylabrobot instead when one workflow must support multiple robot vendors.

optimize-for-gpu

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. Use for CUDA/GPU optimization; CPU-bound NumPy, SciPy, pandas, scikit-learn, NetworkX, scikit-image, vector-search, image-processing, graph, simulation, or file-I/O workloads; CuPy, cuDF, cuML, cuGraph, cuVS, cuCIM, KvikIO, Warp, Newton, Numba-CUDA, or RAFT questions; and profiling, memory-transfer, kernel, or multi-GPU bottlenecks. Also use when large data-parallel Python code is slow and GPU acceleration is a plausible option, even if the user does not name CUDA.

optimizing-attention-flash

Enables Flash Attention for transformer models using PyTorch native scaled_dot_product_attention (PyTorch 2.2+) or the flash-attn library, including multi-query attention, sliding window attention, and FP8 on H100 (FlashAttention-3). Covers profiling speedup, checking accuracy against a baseline, and troubleshooting install and GPU support errors. Use when training or running transformers on long sequences (over 512 tokens), when standard attention runs out of GPU memory, when attention is the inference bottleneck, when switching a PyTorch model to the flash backend, or when tuning attention on H100 GPUs. Not for CPU inference, V100 GPUs, or sequences under 256 tokens; consider xFormers for other attention variants.

outlines

Generates guaranteed-valid structured output from LLMs with Outlines (dottxt.ai), constraining token sampling via finite state machines for JSON schemas, Pydantic models, regex, choice lists, and integer/float types, on Transformers, llama.cpp, vLLM, and limited OpenAI backends. Use when you need output that always parses as valid JSON or matches a regex. Use when extracting typed data into Pydantic models from a local model. Use when classifying text into a fixed set of categories. Use when running batch structured generation on vLLM or Transformers. Use when controlling token sampling at the grammar level. Not for API models needing automatic retries (use Instructor).

pacsomatic

Operator toolkit for nf-core/pacsomatic matched tumor-normal workflows from BAM inputs. Use this skill when the user needs to validate run inputs, generate pacsomatic-compliant samplesheets, prepare reproducible Nextflow launch artifacts, run locally or submit to schedulers (LSF/Slurm/PBS/SGE), and triage execution failures. Triggers on requests to run pacsomatic, prepare launch commands/scripts, perform dry-run checks, or troubleshoot pipeline startup and scheduler submission errors.

paper-corpus-rag

Build grounded question answering and retrieval-augmented generation (RAG) over your own collection of research papers, with answers that cite the exact paper and passage. Use when a user wants to "chat with" or search a folder of PDFs, synthesize evidence across a literature corpus, find which paper says X, or build a vector/hybrid index with SQLite FTS5, pgvector (Postgres), Chroma, Qdrant or FAISS. Includes a zero-dependency local full-text index with citable hits.

paper-lookup

Search 18 scholarly APIs for papers, preprints, citations, open-access full text, repository records, and journal OA status, and return results with reproducible provenance. Covers PubMed, PMC, Europe PMC, bioRxiv, medRxiv, arXiv, OpenAlex, Crossref, Semantic Scholar, CORE, Unpaywall, OpenCitations, PubTator3, Zenodo, Figshare, ROR, BioStudies, and DOAJ. Use when searching for papers, citations, DOI/PMID/arXiv lookups, abstracts, full text, open-access PDFs, preprints, citation graphs, author publications, biomedical entity annotations, deposited records (Zenodo, Figshare, BioStudies), institution ROR IDs, or any scholarly literature query. Triggers on mentions of any supported database or requests like "find papers on X", "look up this DOI", "who cites this paper", or "get me the PDF".

paperclip

Search and read full-text biomedical papers, FDA/PMDA/EMA regulatory documents, clinical trial registries, and UniProt/PDB/ChEMBL entries with the Paperclip CLI from GXL. Covers installing and authenticating the paperclip binary with a PAPERCLIP_API_KEY, the read-only virtual filesystem under /papers, /fda, /trials, /proteins and /clipboard, source-scoped semantic search, corpus-wide grep, metadata lookup and SQL, map/reduce reading across many papers, figure vision analysis, opt-in paper repositories with claim verification, and line-pinned citations. Use when asked to install paperclip, run paperclip search/grep/map/reduce/sql/repo, find or read biomedical literature, regulatory filings or clinical trials through paperclip, or produce citations with line numbers.

paperzilla

Chat with your agent about projects, recommendations, and canonical papers in Paperzilla. Use when users ask for recent project recommendations, canonical paper details, markdown-based summaries, recommendation feedback, feed export, or Atom feed URLs.

parallel-web

Runs the parallel-cli tool for web workflows: web search, URL and PDF extraction, deep research reports, structured data enrichment of supplied rows, FindAll entity discovery, and recurring web monitors. Prefers primary literature and institutional sources for scientific queries. Use when looking up current web evidence or a bounded research question. Use when fetching content from a known URL, PDF, or JavaScript-rendered page. Use when adding web-sourced fields to a list of companies, people, or products. Use when discovering entities that match natural-language criteria. Use when an explicitly exhaustive multi-source report is requested. Use when tracking web changes on a recurring schedule. Not for one-time checks of a known page that need no CLI, or for offline literature databases.

pathml

Covers local, research-only computational pathology with PathML 3.0.5: loading and tiling whole-slide images (OpenSlide, Bio-Formats), preprocessing and QC pipelines run via SlideData.run(), .h5path data management, multiplex image quantification, spatial graph construction (KNN, RAG, HACT), and bounded local ONNX model inference planning. Use when loading or tiling slides, building tissue-mask or stain pipelines, managing .h5path files and patient-level splits, quantifying CODEX or Vectra multiplex images, or building cell and tissue graphs. Not for clinical diagnosis or patient care decisions.

pathogen-variant-surveillance

Query live pathogen genomic surveillance data through the GenSpectrum LAPIS API to find which viral lineages are circulating now, how fast they are growing, and what mutations they carry. Use whenever a question depends on the current state of a pathogen population rather than on remembered facts - which SARS-CoV-2 variant is dominant, whether a Pango lineage is still designated or has been withdrawn, what clade or genotype of H5N1 is in a host or region, whether a PCR primer or assay target still matches circulating sequence, or how a lineage's prevalence has moved week to week. Triggers include "variant surveillance", "genomic surveillance", "what variant is circulating", "dominant variant", "Pango lineage", "lineage prevalence", "growth advantage", "SARS-CoV-2 variant", "XFG", "clade 2.3.4.4b", "H5N1 genotype", "influenza clade", "RSV/mpox/measles/dengue lineage", "CoV-Spectrum", "LAPIS", "Nextclade", "pango-designation", and any request to report what a pathogen population looks like today.

pathway-enrichment

Run pathway and gene-set enrichment analysis on gene lists or ranked gene data, then interpret the results. Use whenever the user has a set of genes (differentially expressed genes from PyDESeq2/Scanpy, CRISPR-screen hits, cluster marker genes, proteomics hits) and wants to know which biological pathways, GO terms, or gene sets are over-represented or enriched. Covers over-representation analysis (ORA / Enrichr / Fisher / hypergeometric), ranked Gene Set Enrichment Analysis (GSEA / preranked), single-sample scoring (ssGSEA/GSVA), and functional profiling via gseapy, g:Profiler, Enrichr libraries, MSigDB, GO, KEGG, Reactome, and WikiPathways — plus gene-ID mapping, choosing the right background universe, multiple-testing correction, redundancy reduction, dotplots/enrichment maps, and publication-ready tables. Use this for "pathway analysis", "enrichment analysis", "GO enrichment", "KEGG/Reactome pathways", "GSEA", "over-representation", "functional annotation", or "what pathways are my genes in".

peer-review

Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments. Use for authorized review of scientific manuscripts, protocols, preprints, or research proposals; reporting-guideline selection; claim–evidence checks; methods, statistics, reproducibility, ethics, figure/table, and citation critique; or revision-response planning.

peft-fine-tuning

Fine-tunes LLMs with Hugging Face PEFT, using LoRA, QLoRA, IA3, AdaLoRA, prefix tuning, and prompt tuning so that under 1% of parameters are trained. Covers rank, alpha, and target module selection, loading and merging adapters, multi-adapter serving, and integration with TRL SFTTrainer, Axolotl, and vLLM. Use when fine-tuning 7B-70B models on limited GPU memory, when running QLoRA on a single 24GB GPU, when choosing LoRA rank and alpha, when merging or swapping adapters on one base model, or when debugging CUDA OOM or an adapter that does not apply. Not for full fine-tuning of small models under 1B parameters or cases needing all weights updated.

pennylane

Hardware-agnostic quantum ML framework with automatic differentiation. Use when training quantum circuits via gradients, building hybrid quantum-classical models, or needing device portability across IBM/Google/Rigetti/IonQ. Best for variational algorithms (VQE, QAOA), quantum neural networks, and integration with PyTorch or JAX. For hardware-specific optimizations use qiskit (IBM) or cirq (Google); for open quantum systems use qutip.

phoenix-observability

Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.

pi-agent

Build with and use Pi, the minimal terminal coding harness. Use for installing Pi, configuring providers/models/settings/environment variables, creating Pi skills/extensions/packages/themes/prompt templates, embedding Pi through the SDK, integrating over RPC or JSON event streams, parsing sessions, running local models through the llama.cpp router, developing custom Pi providers and TUI components, or using ecosystem packages such as pi-subagents (delegation/orchestration), pi-mcp-adapter (MCP servers), pi-interview (interactive forms), and pi-web-access (web search, fetching, video understanding).

pinecone

Guides use of Pinecone, a managed serverless vector database, through its Python client and the LangChain and LlamaIndex integrations. Covers creating indexes, upserting and querying vectors, metadata filtering, namespaces, hybrid dense and sparse search, index management, and deleting vectors. Use when building a production RAG system on a hosted vector store, adding semantic search or recommendations without running infrastructure, isolating per-user or per-tenant data with namespaces, combining dense and sparse vectors in one query, or filtering results by metadata. Do not use for self-hosted or local stores (Chroma, Weaviate) or offline similarity search (FAISS).

pkpd-modeling

Pharmacokinetic and pharmacodynamic modelling and simulation - non-compartmental analysis, compartmental and population PK, PK/PD and exposure-response, TMDD, PBPK orientation, bioequivalence, allometric scaling and first-in-human dose, drug interaction prediction, and Bayesian therapeutic drug monitoring. Use when analysing concentration-time data, deriving exposure metrics, fitting PK or PD models, or evaluating dosing regimens. Triggers include "pharmacokinetics", "pharmacodynamics", "PK/PD", "NCA", "non-compartmental", "AUC", "Cmax", "lambda z", "half-life", "clearance", "volume of distribution", "compartmental model", "population PK", "popPK", "NONMEM", "nlmixr2", "Pharmpy", "Monolix", "exposure-response", "Emax", "EC50", "indirect response", "effect compartment", "TMDD", "PBPK", "bioequivalence", "RSABE", "ABEL", "allometric scaling", "first-in-human", "MABEL", "drug-drug interaction", "DDI", "ICH M12", "concentration-QTc", "therapeutic drug monitoring", "MIPD", and "dosing regimen".

plos

Prepare manuscripts for PLOS journals (PLOS ONE, PLOS Biology, PLOS Computational Biology, PLOS Genetics, PLOS Medicine, PLOS Pathogens, PLOS Neglected Tropical Diseases, PLOS Global Public Health and others), covering structure, abstract and author summary, Vancouver-style references, the mandatory data availability policy, ethics and financial disclosure statements, figure requirements (PACE), the PLOS LaTeX template, reporting guidelines and PLOS ONE's publication criteria. Use when targeting any PLOS journal or checking compliance with PLOS policies.

polars

High-performance DataFrame library for Python ETL, analytics, and pandas migration. Use for expression-based data manipulation with lazy query optimization, parallel execution, streaming out-of-core processing, Arrow interoperability, and optional GPU execution.

polars-bio

Python library polars-bio for genomic interval operations and bioinformatics file I/O on Polars DataFrames, built on Arrow and DataFusion. Covers overlap, nearest, merge, cluster, coverage, complement, subtract, count_overlaps, per-base pileup depth, and read/scan/write of BED, VCF, BAM, CRAM, SAM, GFF/GTF, FASTA, FASTQ, plus SQL queries over those files. Use when intersecting or merging genomic intervals in Polars. Use when reading or streaming large BED, VCF, or BAM files, including from S3, GCS, or Azure. Use when computing read depth from BAM or CRAM. Use when running SQL on genomic files. Use when migrating bioframe code to a faster alternative. Not for pandas-only workflows that do not use Polars or DataFusion.

pptx-posters

Create and audit editable scientific posters in macro-free PowerPoint (.pptx) from author-approved local content and assets. Use when the requested deliverable is a PowerPoint research/conference poster and exact physical, printer, accessibility, provenance, and package-security checks are required.

presenting-conference-talks

Generates conference presentation slides (Beamer LaTeX PDF and editable PPTX) from a compiled paper with speaker notes and talk script. Use when preparing oral talks, spotlight presentations, or invited talks for ML and systems conferences.

prompt-guard

Classifies text with Meta's Prompt Guard, an 86M-parameter model loaded from HuggingFace, into BENIGN, INJECTION or JAILBREAK labels to detect prompt injections and jailbreak attempts in LLM applications. Covers user input filtering, third-party data and RAG document filtering, batch processing, threshold tuning, and sliding-window handling of texts over 512 tokens. Use when screening user prompts for jailbreaks before they reach an LLM. Use when filtering API responses or retrieved RAG documents for embedded instructions. Use when batch-scanning documents for injection. Use when tuning detection thresholds or false positives on security-related queries. Use when running a lightweight CPU or GPU classifier for 8 languages. Not for content moderation such as violence or hate; use LlamaGuard for that.

protocolsio-integration

Reads, validates, and exports protocols.io data using the documented REST v3/v4 endpoints and the official MCP endpoint, and builds non-executing mutation plans. The bundled client makes bounded GET requests to official hosts only with --execute, and also validates saved protocol JSON offline for step order, version, DOI, and attribution metadata. Use when fetching a protocol, its steps, materials, or PDF from protocols.io by exact version. Use when validating a saved protocol snapshot for provenance before reuse. Use when planning a protocol create, update, publish, upload, or organization export without running it. Use when checking protocols.io rate limits, pagination, or authentication. Do not use for other protocol repositories or general literature search.

pufferlib

Version-aware guidance for PufferLib reinforcement-learning environments, vectorization, policies, PuffeRL training, evaluation, and safe checkpoint review. Use when adapting Gymnasium/PettingZoo environments to published PufferLib 3.0.0 or working with the redesigned native 4.0 source line.

pydeseq2

Runs differential expression analysis on bulk RNA-seq count data with PyDESeq2, the Python port of DESeq2. Covers formulaic single- and multi-factor designs, contrasts, Wald tests, Benjamini-Hochberg FDR correction, optional apeGLM LFC shrinkage, pandas and AnnData (H5AD) integration, CSV export, volcano and MA plots, and a command-line script. Use when comparing gene expression between conditions such as treated vs control, adjusting for batch or covariates, porting an R DESeq2 workflow to Python, or building a Python pipeline for differential expression from raw integer counts. Not for single-cell data or for R-based DESeq2 itself.

pydicom

Reads, inspects, writes, and transforms local DICOM files with pydicom 3.x (dcmread, dcmwrite, pydicom.pixels), including metadata, transfer syntaxes, compressed pixel data plugins, frame decoding, private elements, DICOM JSON, and bounded pseudonymization review using bundled helper scripts. Use when extracting aggregate metadata from DICOM datasets without printing PHI, checking which transfer syntaxes and codec plugins a deployment needs, planning frame or memory limits before decoding pixel data, rendering one non-diagnostic frame, compressing or decompressing pixel data, or building and auditing a pseudonymized derivative. Not for diagnostic viewing, or for claiming DICOM PS3.15, HIPAA, or GDPR compliance.

pyhealth

Builds clinical deep-learning pipelines with PyHealth using its Dataset → Task → Model → Trainer → Metrics pattern. Covers loading MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14 and EHRShot data, defining prediction tasks (mortality, readmission, length of stay, drug recommendation, sleep staging, ICD coding, EEG events), models such as Transformer, RETAIN, GAMENet, SafeDrug and StageNet, training with Trainer, and ICD/ATC/NDC/RxNorm code lookup and cross-mapping. Use when building a clinical prediction model on EHR data. Use when working with MIMIC or eICU datasets. Use when recommending drugs or staging sleep from signals. Use when mapping medical codes between ICD, ATC, NDC or RxNorm. Not for generic PyTorch on tabular data.

pylabrobot

Develop and review PyLabRobot lab-automation resources, liquid-handling plans, offline simulations, and supported-device integrations. Use for PyLabRobot protocols or API questions; keep physical execution behind an explicit operator safety gate.

pymatgen

Analyzes, validates, converts, and transforms crystal structures and molecules with pymatgen. It covers CIF/POSCAR/XYZ/JSON conversion, symmetry-tolerance sweeps, local phase diagrams from computed entries, VASP and Q-Chem I/O, band structure and DOS parsing, and bounded Materials Project (mp-api) queries. Use when validating a CIF or checking disorder and oxidation states. Use when assigning space groups and testing symprec sensitivity. Use when converting structure files and checking for representation loss. Use when building a convex hull from total energies. Use when querying Materials Project with explicit fields and limits. Not for general molecular dynamics or non-materials cheminformatics.

pymc

Builds, fits, checks, and compares Bayesian models in Python with PyMC and ArviZ. Covers hierarchical (multilevel) models, NUTS MCMC sampling, variational inference (ADVI), prior and posterior predictive checks, convergence diagnostics (R-hat, ESS, divergences), and LOO/WAIC model comparison. Use when writing a PyMC model for regression, count, or binary data. Use when fitting a hierarchical model with partial pooling. Use when diagnosing divergences, low ESS, or high R-hat. Use when comparing candidate models with LOO. Use when choosing priors or running prior predictive checks. Not for non-Bayesian regression; use statsmodels or scikit-learn instead.

pymoo

Solves single- and multi-objective optimization problems in Python with pymoo, using NSGA-II, NSGA-III, MOEA/D, SPEA2, RVEA, GA, DE and PSO. Covers custom problems (Problem, ElementwiseProblem, FunctionalProblem), constraint handling, mixed-variable problems, ZDT/DTLZ/WFG benchmarks, genetic operators, parallel evaluation, Pareto front visualization and multi-criteria decision making. Use when finding Pareto-optimal trade-offs between conflicting objectives. Use when defining a constrained or mixed-variable problem and choosing an evolutionary algorithm. Use when benchmarking algorithms on standard test problems. Use when customizing crossover or mutation operators. Use when picking one solution from a Pareto front. Not for gradient-based or convex solvers.

pyopenms

Complete mass spectrometry analysis platform. Use for proteomics and metabolomics workflows—feature detection, peptide/protein identification, label-free and isobaric quantification, adduct/accurate-mass annotation, and complex LC-MS/MS pipelines. Supports extensive file formats and algorithms. For simple spectral comparison and small-molecule library matching use matchms.

pysam

Python/HTSlib workflows for genomic files. Use when reading, querying, filtering, or writing SAM/BAM/CRAM, VCF/BCF, FASTA/FASTQ, or tabix data with pysam, including pileup, coverage, indexing, and CRAM references.

pytdc

Uses the PyTDC package (import tdc, Therapeutics Data Commons) to discover therapeutic ML tasks from tdc.metadata, plan and load approved datasets, apply task-aware splits (random, scaffold, cold_split, combination, time), run evaluator metrics, evaluate benchmark groups such as admet_group, and run bounded molecular-oracle scoring. Use when selecting a TDC task or dataset, planning a download-free split, scoring predictions with TDC evaluators, running a benchmark group evaluation, or checking dataset licenses and cache effects before downloading. Use when scoring molecules with TDC oracles like QED. Not for generic RDKit cheminformatics or training molecule generators.

pytorch-fsdp2

Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.

pytorch-lightning

Organizes PyTorch training code with the lightning package (PyTorch Lightning): LightningModule, LightningDataModule, Trainer, callbacks such as ModelCheckpoint and EarlyStopping, loggers (TensorBoard, W&B, MLflow, Comet, CSV), and multi-GPU/TPU strategies (DDP, FSDP, DeepSpeed). Use when structuring a PyTorch model into training, validation and test steps. Use when configuring a Trainer for multi-GPU or TPU runs. Use when writing a LightningDataModule for data loading. Use when adding checkpointing, early stopping or experiment logging to a training run. Use when choosing between DDP, FSDP and DeepSpeed for a model size. Not for plain PyTorch loops without Lightning.

pytorch-lightning-distributed

High-level PyTorch framework with Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks system, and minimal boilerplate. Scales from laptop to supercomputer with same code. Use when you want clean training loops with built-in best practices.

pyvene-interventions

Provides guidance for performing causal interventions on PyTorch models using pyvene's declarative intervention framework. Use when conducting causal tracing, activation patching, interchange intervention training, or testing causal hypotheses about model behavior.

pyzotero

Reads and writes Zotero libraries from Python with pyzotero 1.13.0 and the Zotero Web API v3: items, collections, tags, attachments, saved searches, full-text content, and BibTeX, CSL-JSON, and bibliography export. Also covers the pyzotero CLI and MCP server for a local Zotero 7. Use when fetching or searching library items programmatically; when creating, updating, or deleting references, collections, or tags; when uploading PDF attachments or downloading files; when exporting citations as BibTeX or CSL-JSON; when building research automation around a Zotero user or group library. Not for formatting citations in a manuscript or verifying references against external databases.

qdrant-vector-search

High-performance vector similarity search engine for RAG and semantic search. Use when building production RAG systems requiring fast nearest neighbor search, hybrid search with filtering, or scalable vector storage with Rust-powered performance.

qiskit

Build, simulate, transpile, and execute quantum circuits with Qiskit and IBM Quantum Runtime. Use for Qiskit 2.x circuits and operators, V2 Sampler or Estimator primitives, target-aware transpilation, local or noisy simulation, IBM QPU execution, Runtime sessions or batches, error mitigation, and Qiskit ecosystem packages.

quantizing-models-bitsandbytes

Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.

qutip

Simulate and audit closed and open quantum-system models with QuTiP 5, including deterministic, trajectory, steady-state, spectral, and phase-space workflows. Use for local quantum-dynamics work where physical assumptions, dimensions, and numerical convergence must be explicit.

ray-data

Scalable data processing for ML workloads. Streaming execution across CPU/GPU, supports Parquet/CSV/JSON/images. Integrates with Ray Train, PyTorch, TensorFlow. Scales from single machine to 100s of nodes. Use for batch inference, data preprocessing, multi-modal data loading, or distributed ETL pipelines.

ray-train

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.

rdkit

Guides use of RDKit (Python) for reading and writing SMILES, MOL/SDF, and InChI, computing descriptors (MW, LogP, TPSA), generating Morgan/MACCS/atom-pair fingerprints, running SMARTS substructure searches, applying reaction SMARTS, and building 2D/3D coordinates with ETKDG. Use when parsing or sanitizing molecules that fail default sanitization, calculating Tanimoto similarity or clustering compounds, filtering libraries by substructure, embedding and optimizing conformers, or computing Murcko scaffolds and molecule hashes. Use when fine-grained control over sanitization or algorithms is needed. For simpler standard workflows, use datamol instead, which wraps RDKit.

rebuttal-and-response-to-reviewers

Plan and write responses to peer review, including journal "response to reviewers" letters for revise-and-resubmit, conference rebuttals under strict length limits (OpenReview/ICLR, NeurIPS, ICML, ACL ARR, CVPR), author responses to meta-reviews, and appeals. Use when a user receives reviews, must triage reviewer comments, draft point-by-point replies, decide what new experiments to run, disagree respectfully with a reviewer, or track manuscript changes for a revision. For authors answering reviews of their own manuscript; to write a review of someone else's paper, use peer-review instead.

reference-manager-interop

Move and sync reference libraries between Zotero, Mendeley, EndNote, JabRef, Paperpile and writing tools (LaTeX/BibTeX, Word, Google Docs, Pandoc, Quarto, Overleaf). Use when converting RIS, BibTeX, BibLaTeX or CSL-JSON files, migrating a library, setting up auto-exported .bib files, choosing a citation style (CSL), or fixing citations lost between tools. Includes a zero-dependency BibTeX/RIS/CSL-JSON converter.

relsa-severity-assessment

Multivariate severity assessment and humane endpoint prediction for laboratory animal studies using the RELSA (RELative Severity Assessment) score and ARIMA-based foRcast forecasting. Use when combining welfare readouts — body weight or weight loss, body temperature, clinical or nesting scores, biomarkers, activity, heart rate, burrowing, wheel running — into one severity score per animal per day, when asking which animals are at risk of reaching a humane endpoint or when one will be reached, when defining attention/danger zones or thresholds on a severity scale by kernel density estimation, or when reporting severity for a 3Rs, refinement, animal-welfare, or EU Directive 2010/63/EU severity-assessment context. Covers directionality ("turned" variables), baseline normalization, reference sets, RELSA weights, ARIMA prediction intervals, and RMSE/PICP/MPIW evaluation.

reproducibility-statement

Prepare the reproducibility, transparency and open-science parts of a paper, including data and code availability statements, reproducibility checklists (NeurIPS, ICML, ICLR, ACL Responsible NLP, Nature reporting summaries), research artifact packaging (Zenodo DOI, CITATION.cff, environment lockfiles, seeds), ethics/broader-impact statements, CRediT author contributions, competing interests and AI-use disclosures. Use when a venue requires any of these statements or checklists, or when preparing code/data for release alongside a paper.

research-agent-skills

Navigate the Research Agent Skills collection by Kalaris Labs for academia across AI, machine learning, biology, chemistry, medicine, physics, and academic writing. Use when a researcher needs to choose a field of study, identify relevant specialist SKILL.md files, or coordinate a cross-disciplinary research workflow. For a narrow task with a matching specialist skill already available, use that skill directly.

research-grants

Guides writing of competitive research grant proposals for NSF, NIH, DOE, DARPA, and Taiwan NSTC. Covers agency-specific formatting and review criteria, specific aims, project descriptions, significance and innovation narratives, broader impacts, budgets and justifications, timelines, biosketches, and the NSTC CM03 form. Use when drafting a proposal for one of these agencies. Use when writing specific aims or broader impacts statements. Use when preparing a budget justification or milestone plan. Use when responding to a solicitation, BAA, or reviewer comments in a resubmission. Use when checking a draft against page limits and submission requirements. For general manuscript writing, use a scientific writing skill instead.

research-knowledge-graph

Turn a bibliography or literature corpus into a knowledge graph of papers, authors, venues, topics and citation links, then analyze it (citation clusters, key papers, bridging work, research gaps) or export it to Neo4j, Gephi, NetworkX, Obsidian or Markdown-graph tools such as graphify. Use when mapping a research field, building a citation network, finding influential or bridging papers, or visualizing how a literature connects. Zero-dependency builder with optional OpenAlex enrichment.

research-lookup

Compile current scholarly evidence for a scientific manuscript or research brief. Use when the user explicitly asks to gather literature, references, background evidence, competing findings, or a manuscript research packet. Uses Parallel Search by default, Parallel Extract for source verification, Parallel Research for explicitly deep/exhaustive work, optional explicit Parallel Chat, and optional Perplexity only when requested or allowed as a failure fallback.

research-skill-creator

Create, improve and test agent skills for research workflows (paper writing, lab protocols, analysis pipelines, domain databases) that meet the Agent Skills specification and this repository's quality and security bar. Use when turning a repeated research task into a reusable skill, contributing a new skill to Research Agent Skills, rewriting an existing skill, writing trigger-accurate descriptions, or adding evals. Works alongside Anthropic's official skill-creator.

rowan

Rowan is a cloud-native molecular modeling and medicinal-chemistry workflow platform with a Python API. Use for pKa and macropKa prediction, conformer and tautomer ensembles, docking and analogue docking, protein-ligand cofolding, MSA generation, molecular dynamics, permeability, descriptor workflows, and related small-molecule or protein modeling tasks. Ideal for programmatic batch screening, multi-step chemistry pipelines, and workflows that would otherwise require maintaining local HPC/GPU infrastructure.

rwkv-architecture

Covers the RWKV (Receptance Weighted Key Value) architecture, an RNN/Transformer hybrid with O(n) inference and no KV cache, including RWKV-7, its parallel GPT-mode training and sequential RNN-mode inference, state passing, fine-tuning with DeepSpeed, and CUDA kernel setup. Use when generating text token by token with constant memory, processing very long contexts of 100K+ tokens, fine-tuning an RWKV model, comparing RWKV memory and speed against Transformers, or debugging RWKV state handling, loading, or out-of-memory errors. Prefer a standard Transformer when peak accuracy matters more than memory, and Mamba for state-space models.

scanpy

Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for exploratory scRNA-seq analysis with established workflows. For deep learning models use scvi-tools; for data format questions use anndata.

scholar-evaluation

Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.

science-aaas

Prepare manuscripts for Science and the Science family of journals (Science, Science Advances, Science Translational Medicine, Science Robotics, Science Immunology, Science Signaling), covering Research Article vs Report formats, abstracts and one-sentence summaries, reference and notes style, Supplementary Materials, data/code policies, figure requirements and the initial-submission vs revision workflow. Use when targeting a Science journal, converting a manuscript to Science style, or checking a submission against Science author instructions.

scientific-brainstorming

Facilitates evidence-aware scientific ideation with independent generation, structured discussion, explicit assumptions, transparent evaluation, adversarial review, and decision logs. Use for early-stage research brainstorming or prioritizing candidate directions; hand off empirical validation, study design, ethics or regulatory review, and clinical questions to appropriate experts or skills.

scientific-critical-thinking

Evaluate scientific claims and evidence quality. Use for assessing experimental design validity, identifying biases and confounders, applying evidence grading frameworks (GRADE, Cochrane Risk of Bias), or teaching critical analysis. Best for understanding evidence quality, identifying flaws. For formal peer review writing use peer-review.

scientific-schematics

Generates publication-style scientific diagrams as raster PNG images from a natural-language prompt, using Nano Banana 2 via OpenRouter, then scores each image with Gemini 3.6 Flash against a document-type threshold (journal, conference, thesis, grant, preprint, poster, etc.) and regenerates up to twice. Writes versioned PNGs and a review_log.json. Use when drawing neural network architectures, CONSORT or PRISMA flowcharts, biological signaling pathways, system or block diagrams, or circuit schematics. Use when a figure needs a quality score and critique recorded for a paper, poster, or grant. Not for vector, PDF, SVG, or EPS output, or for data plots; use a plotting skill for those. Prompts leave the machine, so avoid unpublished or patient data.

scientific-slides

Build slide decks and presentations for research talks. Use this for making PowerPoint slides, conference presentations, seminar talks, research presentations, thesis defense slides, or any scientific talk. Provides slide structure, design templates, timing guidance, and visual validation. Works with PowerPoint and LaTeX Beamer.

scientific-visualization

Create and audit truthful, accessible, publication-ready scientific figures with Matplotlib, Seaborn, or Plotly. Use for figure design, multi-panel layouts, uncertainty and missing-data displays, color/contrast review, image metadata validation, and journal export planning.

scientific-writing

Draft, revise, and audit scientific manuscripts or reports with explicit evidence provenance, reporting-guideline coverage, authorship accountability, confidentiality controls, and local consistency checks. Use for manuscript sections, references, declarations, tables, figures, or submission preparation when scientific accuracy and traceability matter.

scikit-bio

Python library scikit-bio for biological sequence and community-ecology analysis: DNA/RNA/protein sequences, pair_align alignment, phylogenetic trees (NJ, UPGMA, GME/BME, Newick), alpha/beta diversity including Faith's PD and UniFrac, PCoA/CCA/RDA ordination, PERMANOVA/ANOSIM/Mantel tests, ancom and dirmult differential abundance, BIOM tables, and FASTA/FASTQ/GenBank I/O. Use when computing microbiome diversity from a BIOM or feature table. Use when running PCoA and PERMANOVA on a distance matrix. Use when building or comparing phylogenetic trees from sequences. Use when reading, converting, or aligning FASTA/FASTQ/GenBank/Newick data. Use when testing differential abundance in compositional count data. Not for general-purpose sequence scripting where Biopython is enough.

scikit-learn

Covers classical machine learning in Python with scikit-learn (sklearn): classification and regression estimators, clustering and dimensionality reduction, preprocessing, Pipeline and ColumnTransformer, cross-validation, metrics, and GridSearchCV hyperparameter tuning. Use when training or comparing classifiers and regressors on tabular or text data, clustering data and choosing the cluster count, building leakage-free preprocessing pipelines, evaluating models with cross-validation and metrics, or tuning hyperparameters. Not for deep learning; use a dedicated neural network framework instead.

scikit-survival

Builds, evaluates, and audits right-censored survival analysis workflows with scikit-survival (sksurv): Cox PH, Coxnet, IPC ridge, survival trees, forests, boosting, and SVMs, plus nonparametric cumulative incidence for competing risks. Covers leakage-safe scikit-learn pipelines, nested CV, and censoring-aware metrics. Use when fitting survival models on time-to-event data, building structured outcome arrays, computing IPCW concordance, dynamic AUC, or Brier scores, estimating cause-specific cumulative incidence, or tuning models without leakage. Not for Fine-Gray regression, which scikit-survival does not provide.

scvelo

Performs RNA velocity analysis with scVelo on single-cell RNA-seq AnnData objects that have spliced and unspliced layers (from velocyto, STARsolo, kallisto|bustools, or alevin-fry). Covers stochastic and dynamical velocity models, velocity graphs and embedding arrows, latent time, PAGA trajectory graphs, and driver gene ranking. Use when inferring differentiation direction from snapshot data, estimating latent time from splicing kinetics, finding driver genes of a trajectory, or adding velocity arrows to a Scanpy UMAP. Not for fate probability modeling (use CellRank) or datasets without unspliced counts.

scvi-tools

Trains and applies scvi-tools probabilistic deep generative models (scVI, scANVI, totalVI, MultiVI, PeakVI, DestVI, Solo, CellAssign, MrVI and others) on AnnData or MuData single-cell data using PyTorch. Covers batch correction, integration, cell type annotation, probabilistic differential expression, and scRNA-seq, ATAC-seq, CITE-seq, spatial and methylation data. Use when integrating batches or datasets with scVI, annotating cells with scANVI, jointly modeling RNA and protein or ATAC, deconvolving spatial spots, or testing differential expression with uncertainty. For standard preprocessing and clustering pipelines, use scanpy instead.

seaborn

Statistical visualization with pandas integration. Use for quick exploration of distributions, relationships, and categorical comparisons with attractive defaults. Best for box plots, violin plots, pair plots, heatmaps. Built on matplotlib. For interactive plots use plotly; for publication styling use scientific-visualization.

segment-anything-model

Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks in an image.

sentence-transformers

Generates sentence, text, and image embeddings locally with the Python sentence-transformers (SBERT) library, using pre-trained Hugging Face models such as all-MiniLM-L6-v2, all-mpnet-base-v2, and multilingual variants. Covers encoding, cosine similarity, semantic search, batch encoding, fine-tuning, and LangChain/LlamaIndex integration. Use when building embeddings for RAG, running semantic search or similarity scoring, clustering or classifying text, embedding multilingual text without an API, or fine-tuning an embedding model on domain data. For API-based or managed embeddings, use OpenAI or Cohere Embed instead.

sentencepiece

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

sglang

Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.

shap

Explain and audit machine-learning predictions with SHAP. Use for selecting SHAP explainers and maskers, computing and validating feature attributions, handling multi-output explanations, and producing local or global SHAP visualizations.

simpo-training

Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more efficient than DPO. Use for preference alignment when want simpler, faster training than DPO/PPO.

simpy

Builds, tests, and analyzes bounded process-based discrete-event simulations in Python with SimPy 4.1.2: Environment, Timeout, Process, AnyOf/AllOf conditions, interrupts, Resource, PriorityResource, PreemptiveResource, Container, Store, time-weighted monitoring, and replication-based output analysis, plus bundled safe CLIs. Use when modeling queues, production lines, logistics, or service operations with generator processes; when contending entities need shared resources or preemption; when instrumenting queue length or utilization correctly; when running independent replications with seeds and confidence intervals; when debugging event ordering or run(until=) boundaries. Not for continuous-time ODE or agent-based modeling frameworks.

skypilot-multi-cloud-orchestration

Multi-cloud orchestration for ML workloads with automatic cost optimization. Use when you need to run training or batch jobs across multiple clouds, leverage spot instances with auto-recovery, or optimize GPU costs across providers.

slime-rl-training

Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.

sparse-autoencoder-training

Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features. Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations in language models.

speculative-decoding

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

springer-lncs

Format papers for Springer Lecture Notes in Computer Science (LNCS) and related proceedings series (LNAI, LNBI, CCIS) and Springer Nature journals using the llncs class or Springer Nature's sn-jnl template, covering page limits, abstract and keywords, splncs04 references, ORCID, running heads, camera-ready packages and the consent-to-publish form. Use when writing for conferences published in LNCS (e.g. ECCV, MICCAI, ESWC, many workshops), preparing an LNCS camera-ready, or targeting Springer Nature journals.

stable-baselines3

Production-ready reinforcement learning algorithms (PPO, SAC, DQN, TD3, DDPG, A2C) with scikit-learn-like API. Use for standard RL experiments, quick prototyping, and well-documented algorithm implementations. Best for single-agent RL with Gymnasium environments. For high-performance parallel training, multi-agent systems, or custom vectorized environments, use pufferlib instead.

stable-diffusion-image-generation

Generates images with Stable Diffusion models (SD 1.5, SDXL, SD 3.0, Flux) through the HuggingFace Diffusers library, covering text-to-image, image-to-image, inpainting, outpainting, ControlNet conditioning, LoRA adapters, scheduler swapping, and GPU memory optimization. Use when generating images from text prompts, transforming or restyling an existing image, filling masked regions of an image, adding spatial control from edges, poses, or depth maps, loading LoRA style adapters, or fixing out-of-memory and black-image errors in a Diffusers pipeline. Not for API-only generation without a GPU, where DALL-E 3 fits better.

statistical-analysis

Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting. Use whenever a user wants to compare groups, test a hypothesis, analyze experimental or survey data, check statistical assumptions, compute required sample sizes, or write up results - even if they never name a specific test. Covers t-tests, ANOVA, chi-square, correlation, regression, non-parametric and Bayesian methods. For low-level model APIs, see the statsmodels and pymc skills.

statistical-power

Sample-size and statistical power calculations for planning studies. Use whenever someone asks "how many subjects/samples/replicates do I need", wants an a priori power analysis, a minimum detectable effect (MDE), a power curve, or needs to justify a sample size for a grant, IRB protocol, or pre-registration. Covers closed-form power for t-tests, ANOVA, proportions, correlations, chi-square, and regression, plus simulation-based (Monte Carlo) power for designs with no formula — logistic/Poisson regression, mixed models, cluster-randomized trials, survival, and interactions. Use this skill even when the request only mentions an effect size, alpha, or "80% power" without saying "power analysis" explicitly. For laying out the study (randomization, blocking, factorial/DOE, crossover, sequential designs) use experimental-design; for analyzing data already collected and reporting it use statistical-analysis.

statsmodels

Statistical models library for Python. Use when you need specific model classes (OLS, GLM, mixed models, ARIMA) with detailed diagnostics, residuals, and inference. Best for econometrics, time series, rigorous inference with coefficient tables. For guided statistical test selection with APA reporting use statistical-analysis.

sympy

Use when you need exact symbolic math in Python — algebra, calculus, equation solving, symbolic linear algebra, or code generation via lambdify/LaTeX. Prefer NumPy or SciPy when floating-point approximations are sufficient.

systematic-review-prisma

Plan, run and report systematic reviews and meta-analyses to PRISMA 2020 standards. Use when writing a review protocol (PROSPERO/OSF), building reproducible database search strings (PubMed, Embase, Scopus, Web of Science), deduplicating exports, organizing title/abstract and full-text screening, assessing risk of bias, drawing the PRISMA flow diagram, or completing the PRISMA 2020 checklist. Includes a dedup tool and a checked PRISMA diagram generator.

systems-paper-writing

Provides paragraph-level structural blueprints for 10-12 page systems papers targeting OSDI, SOSP, ASPLOS, NSDI, and EuroSys. Covers page budgets per section, writing patterns (gap analysis, observation-driven, contribution list, thesis formula), evaluation structure, venue checklists, reviewer guidelines, deadlines, and LaTeX templates for OSDI, NSDI, ASPLOS, and SOSP. Use when planning the section and page layout of a systems paper. Use when drafting an introduction, motivation, design, or evaluation section. Use when checking a draft against venue page limits and a pre-submission checklist. Use when anticipating how systems reviewers will assess a paper. Use when looking up conference formats or deadlines. For NeurIPS/ICML/ICLR papers or citation verification, use ml-paper-writing instead.

tamarind

Access a collection of open-source molecular design and structural biology tools on the Tamarind Bio platform, via its REST API or MCP server — no local GPUs required. Tamarind bundles popular open-source models for structure prediction (AlphaFold, Boltz, Chai, ESMFold), protein, binder, and de novo design (RFdiffusion, ProteinMPNN, BoltzGen), antibody and nanobody design and developability, protein-ligand docking (DiffDock, Autodock Vina), binding-affinity prediction, MSA generation, and molecular dynamics. Use when the user mentions Tamarind or tamarind.bio, wants to run any of these open-source tools in the cloud, references app.tamarind.bio/api or the x-api-key header, or needs to submit batches of sequences for structural or biophysical characterization.

tensorboard

Logs and views ML training data in TensorBoard using PyTorch SummaryWriter and TensorFlow/Keras callbacks: scalars, images, text, histograms, model graphs, embedding projector, hyperparameter tables, PR curves, and TensorFlow or PyTorch profiler traces. Use when plotting loss and accuracy curves during training, comparing multiple runs in one dashboard, inspecting weight and gradient distributions, projecting embeddings with PCA or t-SNE, tracking hyperparameter experiments, or finding performance bottlenecks in a training loop. Not for hosted experiment tracking with team collaboration; use a dedicated tracking service for that.

tensorrt-llm

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

tiledbvcf

Stores and queries genomic variant data in TileDB-VCF datasets using the tiledbvcf Python API and CLI (create, store, export, list, stat). Covers ingesting single-sample VCF/BCF files with .csi or .tbi indexes, adding samples incrementally, querying regions and samples in parallel, and exporting to VCF/BCF or TSV, on local disk or S3, Azure, and GCS. Use when building a variant database for a cohort, adding new samples to an existing dataset, querying specific regions across many samples, exporting subsets of a large VCF collection, or preparing population genomics data such as allele frequency or GWAS inputs. Not for multi-sample VCFs, which are unsupported, or for one-off parsing of a single VCF file.

timesfm-forecasting

Zero-shot time series forecasting with Google's TimesFM foundation model. Use for any univariate time series (sales, sensors, energy, vitals, weather) without training a custom model. Supports CSV/DataFrame/array inputs with point forecasts and prediction intervals. Includes a preflight system checker script to verify RAM/GPU before first use.

torch-geometric

PyTorch Geometric (PyG) for graph neural networks — node/link/graph classification, message passing (GCN, GAT, GraphSAGE, GIN), heterogeneous graphs, neighbor sampling, and custom datasets. Use when working with torch_geometric, not for general NetworkX analytics or non-graph PyTorch models.

torchdrug

Build and troubleshoot TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and knowledge graph reasoning. Use when code imports torchdrug or needs its datasets, models, tasks, or Engine.

torchforge-rl-training

Provides guidance for PyTorch-native agentic RL using torchforge, Meta's library separating infra from algorithms. Use when you want clean RL abstractions, easy algorithm experimentation, or scalable training with Monarch and TorchTitan.

training-llms-megatron

Trains large language models (2B-462B parameters) with NVIDIA Megatron-Core using tensor, pipeline, sequence, context, and expert parallelism, plus FP8 on H100 and MoE configuration for Mixtral-style models. Covers choosing TP/PP/DP/CP sizes, launching distributed training, tuning micro-batch size, and fixing low MFU, out-of-memory errors, and diverging loss. Use when training models above 10B parameters on NVIDIA A100/H100 GPUs, when setting up 3D parallelism for a LLaMA-style model, when configuring expert parallelism for MoE training, or when trying to raise MFU toward 40-47%. Use PyTorch FSDP, DeepSpeed, or HuggingFace Accelerate instead for models under 70B or simpler setups.

transformer-lens-interpretability

Provides guidance for mechanistic interpretability research using TransformerLens to inspect and manipulate transformer internals via HookPoints and activation caching. Use when reverse-engineering model algorithms, studying attention patterns, or performing activation patching experiments.

transformers

Hugging Face Transformers for loading Hub models, running pipeline inference, text generation, and Trainer fine-tuning on NLP, vision, audio, and multimodal tasks. Use when working with AutoModel, pipelines, tokenizers, or TrainingArguments—not for general ML outside the Transformers library.

treatment-plans

Format and structurally validate local treatment-plan documentation after clinical decisions have already been supplied and verified by authorized licensed professionals. Use for source traceability, clinician-authored intervention records, goals and checkpoints, shared-decision records, reconciliation handoffs, and release gates—not for clinical decision-making.

umap-learn

Reduces and embeds high-dimensional data with umap-learn (UMAP) in Python, including 2D/3D visualization, supervised and semi-supervised UMAP, DensMAP, AlignedUMAP, Parametric UMAP (Keras), transform() on new data, and inverse transforms. Use when visualizing high-dimensional data as a 2D or 3D embedding, preprocessing features for HDBSCAN clustering, using partial labels to guide an embedding, aligning embeddings across time points or batches, or projecting unseen samples into a trained embedding. Tune n_neighbors, min_dist, n_components, and metric. Not for linear PCA or t-SNE-specific workflows.

uncertainty-and-units

Track physical units and propagate measurement uncertainty in scientific calculations using pint and uncertainties. Use for unit conversion and dimensional checking, GUM uncertainty budgets, Type A and Type B evaluation, coverage factors and expanded uncertainty, Monte Carlo propagation, significant-figure and plus-minus reporting, error propagation through curve fits, CODATA constants, auditing Python code for stripped units or broken uncertainty propagation, and order-of-magnitude plausibility checks using dimensionless groups (Reynolds, Peclet, Damkohler, Knudsen, Biot, Womersley), characteristic scales such as diffusion time or Debye length, and observed magnitude ranges. Trigger on "is this number physically reasonable", "sanity check these units", "what regime is this flow in", or a result that looks off by orders of magnitude.

unslop-academic-writing

Remove AI slop from research writing so papers, theses, grant proposals, reviews and rebuttals read as written by a careful human expert. Covers stock vocabulary (delve, tapestry, pivotal, underscores), empty emphasis, hedge stacks, formulaic signposting, "not only X but also Y" constructions, em-dash pileups, monotone rhythm, and claims vaguer than the data. Use when drafting or revising academic text with an AI assistant, when a draft "sounds like ChatGPT", before submission, or to match an author's own voice. Includes a zero-dependency slop linter for Markdown, LaTeX, text and Word files.

unsloth

Provides guidance on fine-tuning large language models with Unsloth, a library for faster, lower-memory training using LoRA and QLoRA, based on its official documentation (references/llms-txt.md). Use when setting up LoRA or QLoRA fine-tuning of an LLM with Unsloth, reducing GPU memory use during training, looking up Unsloth features or APIs, debugging Unsloth training code, or learning Unsloth best practices. Not for general model training with plain Hugging Face Transformers or PEFT without Unsloth.

usfiscaldata

Query the U.S. Treasury Fiscal Data REST API for federal financial data. No API key required. Use for national debt (Debt to the Penny), Daily Treasury Statements, Monthly Treasury Statements, Treasury securities auctions, interest rates, foreign exchange rates, savings bonds, or U.S. government revenue and spending statistics.

vaex

Processes and analyzes tabular datasets too large for RAM using Vaex, a Python library for lazy, out-of-core DataFrames over memory-mapped HDF5 and Arrow files, with CSV and Parquet import/export. Covers virtual columns, filtering, groupby aggregations, large-data heatmaps and histograms, and vaex-ml transformers, PCA, and K-means. Use when opening or converting multi-gigabyte CSV/HDF5/Arrow/Parquet files, computing fast statistics on billions of rows, visualizing massive datasets, building ML pipelines that do not fit in memory, or speeding up slow aggregations with lazy evaluation and delay=True. Prefer polars when data fits in RAM, or dask for cluster-distributed work.

venue-templates

Prepare journal manuscripts, conference papers, research posters, and grant documents using venue-specific formatting guidance and bundled LaTeX scaffolds. Use when selecting an official template, checking current page or anonymity rules, adapting academic writing to a venue, or inspecting a submission PDF.

verl-rl-training

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

waypoint-bio

Use when working with Outpost Bio's open microbiome foundation models - the Waypoint checkpoints (Waypoint-6m, Waypoint-45m, Waypoint-170m), the Atlas pretraining corpus, the Compass eight-task benchmark, or the waypoint CLI from the waypoint-bio package. Covers embedding microbiome samples, fine-tuning on taxonomic abundance data, benchmarking a checkpoint on Compass, pretraining a GPT-2 model on taxonomic abundance profiles, and converting MetaPhlAn, Kraken2, QIIME 2, or MGnify abundance tables into waypoint format.

weights-and-biases

Logs and tracks machine learning experiments with Weights & Biases (W&B, wandb): metrics, hyperparameters, checkpoints, sweeps, artifacts with lineage, model registry, custom charts and shareable reports. Includes integrations for PyTorch, HuggingFace Transformers, PyTorch Lightning and Keras/TensorFlow. Use when logging training runs and comparing them across configurations, running automated hyperparameter sweeps, versioning datasets and models as artifacts, registering model versions, or sharing results with a team workspace. Use when setting up wandb.init for offline or unstable-connection training. Not for general data versioning outside ML runs.

whisper

Transcribes and translates audio with OpenAI's Whisper (openai-whisper Python package and whisper CLI), covering model sizes from tiny to large plus turbo, language specification, initial prompts, word and segment timestamps, temperature fallback, batch processing, and subtitle generation. Use when transcribing speech, podcasts, meetings, or video audio to text. Use when translating non-English speech to English. Use when working with noisy or multilingual audio in any of 99 languages. Use when choosing a Whisper model size for available VRAM. Use when generating subtitles from audio. Not for speaker diarization or live captioning; use AssemblyAI or Deepgram instead.

zarr-python

Guides use of Zarr-Python 3 for storing chunked, compressed N-dimensional arrays and groups, with local, in-memory, ZIP, and fsspec-backed S3/GCS/HTTP stores, plus NumPy, Dask, and Xarray integration. Covers array creation, resizing and appending, attributes, chunk and shard sizing, codecs, consolidated metadata, and v2-to-v3 migration. Use when creating or opening Zarr arrays or groups, choosing chunk sizes or compression for large datasets, reading or writing arrays on S3 or GCS, appending to time-series arrays, or migrating code from Zarr-Python 2 to 3. Use when setting up parallel I/O with Dask or Xarray. For labeled multi-dimensional datasets, use Xarray instead.