Skip to content

cometweb-io/cometweb-agent-skills

v2.0.1MIT

CometWeb Labs specialist skills for research, product operations, quality workflows, adversarial review, QA, release readiness, and evidence-based decisions.

skill-evaluator

Design and evaluate Agent Skill experiments that measure whether a skill improves model behavior, discovery, task success, reliability, cost, or latency relative to a no-skill/prior-version baseline, including host/model comparisons, judge agreement, quality-cost Pareto trade-offs, and runtime drift under frozen measurement identity. Use when the user asks to benchmark, A/B test, evaluate, compare, prove, regress-test, or measure a skill across supported harnesses, or to prepare executable real-host suites when execution is unavailable. Do not use for static package/routing audits (use skill-auditor), to create/edit a skill (use skill-creator), to fabricate real-host results, or to claim universal superiority from one configuration.

Read SKILL.md at the source

Pinned to revision 6626bb65beb0, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.