Skip to content

neerajcodz/ml-stack

v0.2.0MIT

Requirements-driven ML research, experiments, validation, and evidence audit for coding agents.

ml-stack-evaluation

Evaluate ML models or systems when a user asks for benchmarks, metrics, slices, calibration, robustness, LLM evaluation, smoke tests, pass/fail thresholds, or go/no-go decisions; preserve revisions and uncertainty.

Read SKILL.md at the source

Pinned to revision 6829bde67fb6, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.