nemo-evaluator-sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/nemo-evaluator-sdk/SKILL.md
- skills/nemo-evaluator-sdk/references/adapter-system.md
- skills/nemo-evaluator-sdk/references/configuration.md
- skills/nemo-evaluator-sdk/references/custom-benchmarks.md
- skills/nemo-evaluator-sdk/references/execution-backends.md
- skills/nemo-evaluator-sdk/references/python-api-usage.md
Every link opens the file at its source, pinned to the revision this page describes.