cometweb-io/cometweb-agent-skills
CometWeb Labs specialist skills for research, product operations, quality workflows, adversarial review, QA, release readiness, and evidence-based decisions.
Run an always-current, evidence-governed AI decision council for material, high-stakes, multi-domain decisions where options, risk gates, forecasts, and living memory matter. Use when the user explicitly asks to "przepuść przez Radę", "zapytaj Radę", compare consequential options, challenge a major plan, decide GO/NO-GO/TEST/DEFER on strategic choices, verify whether a prior decision is still current, inspect Council health, or improve the Council itself. Do not use for routine weekly product prioritization, whole-repo roadmapping, release-candidate gates, customer support triage, or single-domain specialist work when a dedicated skill exists. Orchestrate blind advisers, conditional specialists, binding risk gates, live/fresh evidence, temporal truth, contradiction testing, forecasts, living Decision Memory, watch dependencies, human escalation, and champion/challenger evaluation.
Naturalize, edit, or substantially rewrite English and Polish prose while preserving meaning, factual constraints, and the author's intentional voice. Use specifically when the user asks to humanize text, sound less like generic AI writing, remove AI tells, perform a strong/deep/robust rewrite, clean suspicious invisible Unicode in prose or Markdown, or re-express text with provenance-aware caution. Supports blogs, articles, LinkedIn posts, emails, proposals, documentation, release notes, and READMEs. Do not use merely for generic proofreading, unrelated copywriting, or AI-authorship detection. Never claim human authorship, detector defeat, or watermark removal without an appropriate supported test.
Run a final evidence-backed acceptance gate on a content, research, documentation, sales, or other knowledge artifact against an explicit brief, required evidence, unresolved findings, and format/QA criteria. Use when the user asks whether an artifact is actually ready, complete, publishable as a deliverable, or has passed its defined acceptance contract. Do not use as the final production-release gate for software, to choose among consequential strategic options, to perform the initial review, or to invent acceptance criteria after seeing the result; use release-readiness, ai-council, or the relevant reviewer instead.
Design, curate, version, and quality-control benchmark or evaluation corpora for Agent Skills and quality workflows, including task taxonomies, discovery/forced/negative controls, adversarial and regression cases, difficulty strata, holdout isolation, provenance, duplication, contamination risk, coverage balance, and immutable benchmark hashes. Use when the user asks to build or maintain an eval dataset, golden set, benchmark suite, regression corpus, holdout, challenge set, or representative test cases. Do not use to execute the model experiment or claim lift (use skill-evaluator), to design the grading rubric (use rubric-designer), to edit the candidate skill (use skill-creator), or to treat leaked/known cases as a clean holdout.
Turn an ambiguous request for a content, research, sales, documentation, or knowledge artifact into an explicit execution contract with audience, objective, evidence policy, constraints, acceptance criteria, risks, and handoff fields. Use when the task is underspecified, expensive to redo, spans multiple specialists, or needs a durable brief before writing/research. Do not use to write the final artifact, run broad product discovery, make a consequential decision, or replace product-operator, evidence-researcher, content-writer, or ai-council.
Builds a fresh, provenance-aware context snapshot for CometWeb and its active projects before product, portfolio, strategy, GTM, pricing, outreach, meeting, content, or advisory work. Use when the request depends on current state across GitHub/repositories, Notion, CometWeb Insight, CRM/communications, files, live CometWeb surfaces, or multiple CometWeb projects such as Insight, CometBase, CometPen, CometLens/Extensions, cometweb.io, agent-skills, RankProof, and research. Also use for "what changed", full/portfolio refreshes, meeting prep, and as a preflight for Product Operator, AI Council, Repo to Roadmap, or Skill Orchestrator. Include First Principles governance for material product/GTM/pricing/portfolio decisions. Do not use for simple conceptual questions or when one direct authoritative source is sufficient.
Continuous competitive intelligence and competitor change detection. Use when the user asks to monitor competitors over time, refresh existing competitor profiles, detect what changed since a prior scan, track pricing/product/positioning/SEO/ads/reviews/company changes, maintain a competitor watchlist, produce recurring competitor digests, verify competitor claims, analyze cross-competitor trends, or turn observed deltas into product/GTM/sales implications. Prefer competitor-profiling for a one-time initial deep profile; use this skill when temporal state, snapshots, deltas, alerts, freshness, evidence provenance, or recurring intelligence operations matter.
Review an existing piece of content constructively against its brief, intended reader, factual/evidence requirements, structure, clarity, specificity, usefulness, and internal consistency. Use when the user wants editorial QA, actionable findings, a pre-publication review, or a reasoned assessment of what must change. Do not use for adversarial roasting, full rewrites, evidence collection, humanization-only editing, or final release acceptance; route those to content-roaster, content-writer/ai-humanize, evidence-researcher, or artifact-acceptance.
Run evidence-anchored adversarial reviews of marketing, sales, product, editorial, and long-form content such as landing pages, pricing pages, offers, emails, posts, articles, case studies, reports, ebooks, lead magnets, documentation, product pages, demo/trial flows, enterprise procurement or trust content, release announcements, and sales decks. Use when the user asks to roast, tear apart, red-team, stress-test, compare revisions of, or brutally critique content and wants each material criticism tied to exact source evidence, decision impact, proof burden, counterevidence search, a concrete repair, and an observable acceptance check. Do not use for scientific peer review, repository/code critique, rewrite-only work, one-off fact checking, or full SEO/GEO/AEO audits; route those to science-roaster, repo-roaster, ai-humanize, evidence-researcher, or seo-geo-aeo-maxxing.
Create evidence-aware informational or editorial prose such as articles, guides, reports, documentation, research-backed explainers, and other bounded knowledge content from an explicit brief or sufficiently clear request. Use when the user wants the actual written artifact and factual integrity, reader utility, structure, and claim discipline matter. Do not use primarily for persuasion-first landing-page copy, ads, lifecycle or cold email, generic copy-editing, humanization-only rewrites, long-form publication/release orchestration, hostile critique, or primary research collection; route those to the relevant copy, email, ai-humanize, longform-publisher, evidence-researcher, content-reviewer, or content-roaster specialist when available.
Run the operational customer-to-resolution loop across support conversations, customer cases, incidents, account risk, commitments, internal handoffs, feedback clusters, and GitHub engineering work. Use when asked to triage an inbox or ticket queue, investigate a customer problem, detect or coordinate an incident, build an operational account 360, watch churn/non-renewal signals, find overdue promises or stalled escalations, dedupe customer-reported bugs into GitHub, produce customer-ops briefs, or verify that a fix actually resolved the customer-visible symptom. Composes with Gmail/support tools, HubSpot/CRM, billing, product analytics, GitHub, Notion/incident records, and connected files. Do not use for broad VOC/persona research, retention-program design, general CRM architecture, analytics implementation, product roadmap prioritization, or security exploitation; hand those workflows to the specialist skill.
Find, verify, qualify, compare, and manage companies as design partners, co-development partners, beta partners, paid pilots, lighthouse customers, or early adopters. Use for design-partner discovery, early-adopter shortlists, qualification of an existing candidate list, cohort selection, partner-readiness validation, active-partner review, or revalidation of prior candidates. Go deeper than generic prospecting by optimizing for learning value, urgency, representativeness, implementation feasibility, real user/champion access, mutual commitment, evidence quality, transferability, cohort coverage, and cost-to-learn. Separate desk-research fit from live readiness, preserve evidence lineage/freshness, and prevent prestige or contract size from replacing product-learning quality. Hand off broad lead generation to prospecting, research synthesis to customer-research, and outreach copy to cold-email.
Create, research, validate, edit, redesign, and prepare evidence-backed ebooks, white papers, and practical PDF workbooks, especially the CometWeb ebook series. Use for requests such as zrob ebook, napisz poradnik PDF, deep research do ebooka, zweryfikuj ebook, popraw okladke, or prepare a publication from research through reviewed PDF. Own publication scope, research coverage, claim-to-source mapping, editorial structure, meaning-preserving Polish/English editing, series design, revision-aware QA, and publication handoff. Reuse evidence-researcher and ai-humanize when available. Do not trigger for a standalone blog post, a generic website audit, simple PDF extraction, or merely a request to install a skill. Research-only and redesign-only requests must stay within their requested scope. Never imply that automated checks prove factual accuracy or that a created file is installed or publicly published.
Build auditable Evidence Packs for consequential research, fact-checking, due diligence, verification, and cross-skill evidence handoff. Use when ChatGPT must decompose a question into material claims, inspect primary or system-of-record sources, verify currentness/effective dates/versions, distinguish source artifacts from claim-specific evidence, search for falsifiers and negative evidence, resolve contradictions, detect derivative or non-independent sources, identify evidence gaps, compare evidence deltas, or prepare reusable evidence for AI Council, product, technical, audit, SEO/GEO/AEO, sales, or customer workflows. Do not use for casual single-fact lookup or as the final decision-maker when a dedicated decision skill exists.
Aggregate repeated findings, acceptance failures, eval results, user corrections, and workflow incidents across skill runs to identify recurring failure patterns and propose evidence-backed improvements to skill instructions, routing, references, kernels, evals, or host integration. Use when the user wants the skill system to learn from repeated mistakes, run a retrospective, improve reliability over time, or convert observed failures into regression tests. Do not use to silently self-modify skills, learn from a single weak signal, replace product/research analytics, or apply repository changes without explicit authorization.
Use when ChatGPT must build, refresh, reconcile, or release a long-form publication such as an ebook, report, playbook, white paper, guide, handbook, or research-backed article across manuscript and derived document formats.
Evidence-grounded cross-domain portfolio control plane for deciding what a user or small team should focus on across products, client work, research/academic commitments, growth, operations, and side projects. Use when the user asks what to do this week/next 7-30 days across multiple projects, how to allocate limited capacity, which commitments conflict, what to pause/drop/delegate, how to reconcile client deadlines with product/research work, or for a portfolio-wide review. Do not use for deep prioritization inside one product/repo, release GO/NO-GO, or executing a multi-skill workflow when a specialist/control-plane skill owns that task.
Evidence-governed product operating system that reconciles GitHub implementation and release state, Notion roadmap/tasks/decision docs, product context, and available outcome signals to answer what the team should do next. Use for questions such as what remains, what to build/fix/verify next, how to finish or unstick a product, what is actually done, whether roadmap and repo agree, how to plan the next product cycle, what changed since the last review, or what should wait/stop. Produces bounded BLOCKER/VERIFY NOW/DECISION NOW/NOW/NEXT/LATER/STOP actions, dependency-aware sequencing, state drift, readiness, immutable snapshots/deltas, confidence, done conditions, and specialist handoffs. Read-only by default; delegate deep audits and consequential decisions instead of duplicating specialist skills.
Analyze external products, websites, apps, APIs, documentation, and source repositories to extract evidence-backed, transferable, implementable product and engineering patterns rather than generic competitor summaries. Use when the user asks to teardown, reverse-engineer at a product/architecture level, benchmark, study, or learn from another product/repo; asks what workflows, UX mechanics, architecture, onboarding, monetization, developer experience, operations, reliability, or implementation ideas are worth adapting; asks "what can we borrow/learn/implement from X" or Polish equivalents such as "przeanalizuj produkt/repo", "wyciagnij wzorce", or "co warto wdrozyc". Also use for multi-product pattern synthesis and source-to-target adaptation. Do not use as the primary skill for ongoing competitor monitoring/delta analysis, broad competitor dossiers, external-facing comparison pages, or a roadmap of the user's own repo with no external/reference target.
Orchestrate the CometWeb artifact/repository/skill quality loop across briefing, creation, review, adversarial testing, repair, acceptance, measurement, selective revalidation, and skill runtime lifecycle while preserving candidate/contract identity, frozen policy/rubric locks, finding lineage, rollout state, rollback readiness, disagreement state, and quality debt. Use when the user wants one entry point to run or resume the full quality workflow, coordinate quality specialists, process a batch/campaign, reconcile findings, or govern a measured skill candidate through canary/staged rollout. Do not use for general multi-skill orchestration outside quality workflows, to replace specialist analysis, to make consequential strategy decisions, or to declare software production readiness; use skill-orchestrator, ai-council, or release-readiness for those cases.
Assess whether a specific release candidate, build, artifact, application, service, mobile/desktop build, or API is ready for production in a named environment and issue an evidence-backed GO / GO_WITH_CONTROLS / NO_GO / DEFER verdict across product acceptance, QA, security, operations/reliability, documentation, billing/entitlements, and support/incident readiness. Use when the user names or implies a concrete candidate (version, build ID, branch/tag, artifact digest, deploy target) for launch/release gates, pre-deploy audits, "is this build ready to ship?", hotfix readiness, post-incident releases, and repeated delta/revalidation reviews. Do not use for first-time whole-project roadmap baselines or "analyze the entire repo" — use Repo to Roadmap. Do not use for ongoing weekly prioritization on an existing roadmap — use Product Operator. Orchestrate specialist evidence without pretending to replace security scans, live-app QA, legal/privacy review, or deployment authorization.
Convert findings from reviewers, roasters, audits, tests, or acceptance gates into a minimal dependency-aware repair set, optionally apply authorized edits, and verify that fixes close root causes without introducing regressions. Use when the user wants issues actually fixed rather than merely analyzed, especially after content-reviewer, content-roaster, science-roaster, repo-roaster, or artifact-acceptance. Do not use to invent new requirements, perform the initial broad audit, make consequential strategy decisions, or claim a finding is fixed without fresh verification evidence.
Run evidence-anchored adversarial reviews of software repositories, codebases, monorepos, pull requests, branches, modules, architecture, tests, configuration, CI/CD, migrations, data pipelines, infrastructure, and engineering hygiene. Use when the user asks to roast, tear apart, red-team, stress-test, re-check a repaired repo, or brutally critique a repo/codebase and wants concrete file/symbol evidence, critical-invariant reasoning, trust-boundary and state-transition analysis, execution-path reachability, blast radius, false-positive checks, repairs, and executable verification steps. Do not use as the primary skill for whole-project roadmapping, external-product pattern extraction, runtime web QA, or final release GO/NO_GO; route those to repo-to-roadmap, product-teardown, web-app-auditor, or release-readiness.
Analyze an entire software/product project and turn verified project truth into an evidence-based, dependency-aware, reusable roadmap. Use for first-time or delta whole-project baselines such as "analyze the whole repo/project", "what is left to build before target state", "create/update a roadmap from GitHub/Notion", "przeanalizuj całe repo", or "update the roadmap after recent changes". Inventory topology before searching, separate intent/presence/behavior/release/outcome truth, prove material absence instead of inferring it from search misses, expose coverage and contradictions, model capabilities/critical journeys, prioritize target blockers and gates, validate hard dependencies, and create immutable baseline/delta roadmap snapshots. Do not use for weekly sprint control, release-candidate GO/NO_GO gates, or "is build v1.2.3 ready to ship?" — use Product Operator or Release Readiness instead. Do not use as an implementation agent, narrow code review, or specialist security/SEO/CRO audit.
Design, normalize, lint, and freeze evaluation rubrics before substantive review or benchmarking, with observable criteria, explicit pass/fail semantics, evidence floors, materiality, blocker rules, scope boundaries, and immutable rubric hashes. Use when the user asks to create grading criteria, an evaluation rubric, acceptance scoring framework, reviewer scorecard, red-team rubric, or repeatable assessment contract. Do not use to judge the candidate itself, to write the artifact, to run empirical skill benchmarks, or to move criteria after seeing the result; use the relevant reviewer, artifact-acceptance, skill-evaluator, or quality-loop-operator for those tasks.
Run an adversarial Reviewer #2-style critique of scientific manuscripts, papers, protocols, theses, methods, analyses, reviewer responses, validation studies, and research drafts. Use when the user asks to roast, peer-review, red-team, stress-test, re-review a revision of, or challenge scientific work and wants every material criticism anchored to exact source evidence, inferential type, validity domain, counterevidence search, minimum repair burden, and an observable verification condition. Do not use for generic content critique, repository/code review, one-off claim verification, or research-program planning; route those to content-roaster, repo-roaster, evidence-researcher, or research-program-operator.
Run evidence-governed, multi-pillar website visibility audits across technical SEO, relevance, authority/trust, GEO/AI citation readiness, and AEO/answer extraction. Use when the user explicitly asks for SEO+GEO+AEO or "maxxing", a broad end-to-end search/AI visibility audit, cross-pillar diagnosis, same-rubric competitor comparison, or repeat/delta audit. Use PILLAR only when this skill is explicitly requested for one pillar. Do not use for isolated schema/meta tasks, ongoing competitor monitoring, whole-product roadmapping, or release-candidate GO/NO-GO; hand accepted findings to the corresponding specialist. Diagnosis only; never mutate live sites.
Audit an Agent Skill or a repository of skills for trigger quality, scope overlap, instruction conflicts, progressive-disclosure cost, broken references/dependencies, eval blind spots, false-green paths, semantic-version and public-contract compatibility, host-support overclaims, package hygiene, supply-chain risks, migration/deprecation gaps, and drift between registry, docs, tests, and shipped archives. Use when the user asks to audit, review, roast, harden, compare, or quality-check a skill or skill library itself. Do not use to create/edit the skill (use skill-creator), to empirically benchmark model behavior with-vs-without it (use skill-evaluator), to audit ordinary software (use repo-roaster), or to issue a production release verdict.
Design and evaluate Agent Skill experiments that measure whether a skill improves model behavior, discovery, task success, reliability, cost, or latency relative to a no-skill/prior-version baseline, including host/model comparisons, judge agreement, quality-cost Pareto trade-offs, and runtime drift under frozen measurement identity. Use when the user asks to benchmark, A/B test, evaluate, compare, prove, regress-test, or measure a skill across supported harnesses, or to prepare executable real-host suites when execution is unavailable. Do not use for static package/routing audits (use skill-auditor), to create/edit a skill (use skill-creator), to fabricate real-host results, or to claim universal superiority from one configuration.
Plan and execute multi-skill CometWeb workflows with CW-AIP handoffs when the user wants one entry point instead of tagging each specialist. Use when the user asks to orchestrate, sequence, or run end-to-end flows (e.g. evidence then Council, audit then release, research then weekly ops), run everything needed for a goal, or chain Evidence Researcher with ai-council or other skills in order. Supports execution_mode auto|single_thread|isolated_subagents. Loads each step's SKILL.md, runs it fully, and passes envelopes between steps. Do not use for a single-domain task when one specialist skill is enough, for routing-only questions without execution, or to bypass AI Council decision policy or Release Readiness gate rules.
Alias for skill-orchestrator with execution_mode=isolated_subagents. Use when the user asks for multiagent orchestration, separate agents per skill, Task/subagent per step, or strict isolation between Evidence Researcher, AI Council, auditors, and other skills. Do not use when a single specialist suffices, when the host has no subagent/Task API (use skill-orchestrator single_thread), for routing-only questions, or to bypass AI Council or Release Readiness policy.
Evidence-driven QA and product audit for user-facing websites and web applications. Use when asked to audit, review, inspect, QA, verify, click through, test, or "roast" a site/app/page/dashboard/ checkout/form, including UI/UX, accessibility, regression, data integrity, and critical flows. Supports browser, browser+source, screenshot-only, source-only, and fetch-only environments. Do not use for backend-only review, greenfield implementation, whole-repo roadmapping, or as the final production-release gate; provide candidate-bound QA evidence to Release Readiness when relevant. Do not perform penetration testing/exploitation with no product-audit goal.