Skip to content

madhukar04012/agentforge

v1.5.0Apache-2.0

AI Agent Development, Evaluation & Deployment Platform

agentforge-eval

This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes. Applies to any agentforge project, whatever framework the agent is written in. Do NOT use for agent API code patterns (ADK: use agentforge-adk-code), deployment (use agentforge-deploy), or project scaffolding (use agentforge-scaffold).

Version
1.5.0
Read SKILL.md at the source

Pinned to revision 27d427027fdd, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.