testing-llm
LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Use when testing AI features, validating LLM outputs, or building evaluation pipelines.
- License
- MIT
- Compatibility
- Claude Code 2.1.220+.
Pinned to revision 1ff988bd66da, so it is the text this page describes rather than whatever the author pushed since.
Pre-approved tools experimental
Experimental field. Support varies between clients, so this list is what the author declared, not what your client will enforce.
- Read
- Glob
- Grep
- WebFetch
- WebSearch
Files
- skills/testing-llm/SKILL.md
- skills/testing-llm/checklists/llm-test-checklist.md
- skills/testing-llm/references/healer-agent.md
- skills/testing-llm/references/ork-delta.md
- skills/testing-llm/rules/_sections.md
- skills/testing-llm/rules/llm-evaluation.md
- skills/testing-llm/rules/llm-mocking.md
- skills/testing-llm/test-cases.json
Every link opens the file at its source, pinned to the revision this page describes.