hermes-labs-ai/hermes-rubric
Assess named artifacts with evidence-cited Hermes Rubric results and explicit coverage limits.
Changelog
[1.2.3] — 2026-09-15
Fixed
- Stop the published quickstart recipe when
mktemp -dfails, instead of continuing with an empty$workdirthat made every quoted path expand to/post.mdand/result.json(#32)._documented_argv()also now screens the published shell tokens for leftover variables before substitution, so a--basetempcontaining$is not misread as undocumented.
[1.2.2] — 2026-09-12
Added
- Add a native OpenAI Agents SDK adapter (
hermes-rubric[openai-agents]) that renders a completedRunResultinto cited Hermes evidence and grades it without re-running the agent or calling a model itself. Tool guardrail rejections, SDK-resolved names for hosted tool calls, and reasoning text emitted without a summary all reach the rendered run.
Fixed
- Prevent malformed or incomplete scoring responses from being converted into fallback scores that could enter an aggregate. Batch mode retains its per-dimension retry before surfacing a score-stage failure.
[1.2.1] — 2026-09-02
Added
- Add a native Inspect AI scorer that preserves Hermes evidence, coverage, and receipts in Inspect score metadata and supports post-hoc log re-scoring.
[1.2.0] — 2026-09-02
Added
- Add
--pin-rubric <path>for comparable re-grades against an unchanged rubric from a prior JSON result, with pin provenance recorded in the receipt.
[1.1.1] — 2026-08-17 — Apache-2.0 successor release
- Release the portable assessment core under Apache-2.0.
[1.1.0] — 2026-08-14 — portable assessment core
- Add one-call in-memory and path APIs:
assess,assess_path, and async wrappers. - Add typed top-level results with stable serialization and result schema
1.0. - Support caller-provided frozen rubrics alongside synthesis and artifact classes.
- Report complete/partial evidence coverage, byte/source facts, and limitations.
- Normalize public errors by pipeline stage while preserving exception causes.
- Add caller-policy feedback with quality, evidence, and coverage gap kinds.
- Delegate the CLI to the shared public orchestrator while preserving existing flags, exit behavior, and output keys.
- Lead documentation with the agent/application integration path and state the current prefix-window limitation explicitly.
- Add release notes, an adapter contract, and a portable agent-output example.
[1.0.2] — 2026-08-04 — scoring identity and integration clarity
- Pin Stage-3 dimension IDs and names to the synthesized rubric across per-dimension, batched, missing-result, and parse-fallback paths.
- Publish dedicated Documentation and Changelog links in package metadata.
- State the backend requirement before the first README command.
- Align quickstart commands and Python imports with the current CLI and API.
- Clarify which paths are local, which cloud backends are optional, and which stages remain non-deterministic.
- Remove unsupported recomputation and retired research claims from current public documentation.
- Add an OIDC Trusted Publishing workflow for PyPI releases.
1.0.1 — 2026-07-26 — correctness patch
- Honor the configured target window throughout both Stage-2 evidence paths instead of silently applying a second 6,000-character cap.
- Preserve the caller-bound rubric
target_typerather than allowing backend output to reclassify it. - Implement
hermes-rubric --versionand derive receipttool_versionfrom the same runtime version surface.
1.0.0 — 2026-04-28 — first official release
The repo has been public since 2026-04-24 as a 0.9-era preview. v1.0.0 is the first tagged release on PyPI + GitHub Releases, headlined by class-aware rubric templates.
The 0.1.x and 0.2.0 entries below were internal-only iterations toward v1.0.
What v1.0 includes (vs the 0.9-era preview)
Added --artifact-class <name> flag. When set, Stage-1 LLM rubric synthesis is bypassed and a deterministic dim set is loaded from a YAML template. The same input + same class produces the same dim set across runs — addressing the non-determinism observed in v0.1.x where Stage-1 synthesized different dims on every run.
New classes (4)
social-post— X / Twitter / Bluesky. Voice + platform-fit dominate.show-hn-post— Hacker News launch posts. Substance + fab-block dominate.linkedin-post— LinkedIn announcements. Procurement-voice + defensibility dominate.outreach-email— Cold sales emails. Quote-first opener + voice-match dominate.
Each class template includes 7-9 fixed dimensions with weights and evidence instructions, a class-specific slop-signature list (injected into llm_fool dim), and voice priors (injected into voice_match dim). outreach-email adds banned_subject_patterns for the subject_neutrality dim.
CLI
--artifact-classflag added; back-compat preserved (omitted = v0.1 behavior)--intentand--contextoptional when--artifact-classis set
Internal
- New
hermes_rubric.classesmodule - 9 new tests in
tests/test_classes.py; full suite 109 passed, 2 skipped - pyyaml added as runtime dependency
Why
Observed in v0.1.x: scoring the same artifact 3 times produced 3 different rubric hashes because the LLM invented dims fresh each run. Aggregate scores varied 5.4–6.2 across runs, with no way to compare them dim-by-dim. Class-aware preloading determinizes Stage-1: the dim set comes from YAML, not from the LLM.
Back-compat
All existing CLI calls work unchanged. Receipts include a new rubric.rubric_source field ("class-template" when applicable; absent otherwise).
0.1.3 — 2026-04-25
- Refactor: Bias-compensation preambles (intent-debias, scope-class)
moved out of
hermes_rubric.preamblesand into the upstreamhermes-blindpackage (which is the bias-compensation domain). - New runtime dependency:
hermes-blind>=0.1.1. hermes_rubric.preamblesis now a thin re-export shim for thehermes_blind.preamblesmodule. External callers importing from the old path keep working for one migration cycle. Will be removed in v0.2.0.- CLI surface is unchanged:
hermes-rubric --scope-class {gate-plan,sweep-plan,results-bundle} --intent-debiasparses and behaves identically to v0.1.2. - Behavior byte-identical to v0.1.2 — verified by 6-combo
frozen-prompt regression test in
tests/test_preambles_shim.py. - 73 tests green (was 62 + 11 new shim/byte-identity tests).
0.1.2 — 2026-04-25
--batchflag: one LLM call per stage (evidence + score), reducing 2N+1 calls per run to 3- Prompt-layer isolation via per-
<DIM>blocks with explicit "score only within your block" invariant - dim_id-keyed reassembly; rubric dim order re-imposed in
compute_aggregateregardless of LLM return order - Auto-fallback to per-dim mode on JSON parse failure or oversize prompt; mode logged in receipt
- All clamps preserved byte-for-byte (hedge [3,7], no-evidence cap 3, self-marketing cap 6); consolidated in
_apply_clamps - 8 new tests in
tests/test_batch.pycovering reassembly, missing-dim fallback, parse-failure fallback, clamp suffix preservation, and per-dim-vs-batched golden equivalence on a frozen rubric fixture - Default behavior byte-identical to 0.1.x;
--batchis opt-in
0.1.0 — 2026-04-23
Initial release.
- Three-stage pipeline: rubric synthesis, evidence collection, scoring
- Auto-backend detection: claude-cli (preferred), ollama-local (fallback)
- Hedge enforcement: low-confidence dimensions clamped to [3,7]; no-evidence dimensions capped at ≤3
- Reproducibility receipt in every output
- 14 tests (5 backends, 3 synthesize, 4 score, 2 adversarial)
- Adversarial tests confirm fluency-vs-substance resistance and no-evidence cap
- Calibration dataset: 15 cases, 5 domains
- META-RUBRIC: 7 dimensions from 24 corpus-derived failure modes
- Applied: 4-paper scoring (langquant, cogito-ergo, 2 Zenodo papers)
- Handbook entry:
evidence-first-rubric-synthesis.md