Experimental evaluation and improvement of agent skills and plugins, including improvement methods.
Use when evaluating agent skills or plugins, comparing versions, or running evidence-based improvement experiments, including changes to the improvement method itself. Not for ordinary task execution, routine skill authoring, one-off edits, or general design reviews.