Healthcare Agents
Healthcare administration support, from a messy problem to a reviewable artifact.
This is the v2.0.0-beta.1 development candidate. It adds six deep workflows, compact skills, a custom workflow builder and repeatable aggregate calculations. Publication and customer cloud-host qualification are pending; the prepared POSIX Codex profile has bounded native qualification recorded below; npx healthcare-agents from the public registry does not yet deliver this candidate.
Synthetic denial investigation: 10% baseline, 18% current, 8 percentage points higher
The picture is generated from the committed synthetic case and the actual local calculation. These are invented demonstration data, not patient records or measured customer outcomes.
Start with a task
Use the Healthcare Agents skill in an approved agent environment:
Investigate the denial-rate increase. Use the supplied aggregate reports, identify evidence gaps, and draft an owner-led recovery plan.
The host model develops the administrative artifact. The local CLI routes tasks, validates workflow contracts and calculates supported aggregate quantities; it does not call a model or execute an operational workflow.
For this checked-out candidate:
npm ci --ignore-scripts
node bin/cli.js workup "Commercial payer denial rate jumped" --workflow denial-spike-workup --json
node bin/cli.js admin run examples/admin-v2/denial-spike-workup.json
Use synthetic or approved aggregate evidence. This package is not a PHI-processing environment and does not make final clinical, legal, coding, billing, audit or compliance authority.
Six deep workflows
| Task | Useful artifact | Local calculation |
|---|---|---|
| Denial spike | Cohort comparison, hypotheses, evidence pulls and recovery owners | Baseline/current rates, percentage points and relative change |
| Ambulatory access | Demand/capacity scenario and monitored access plan | Shortfall, net weekly capacity and backlog clearance scenario |
| Survey evidence | Standard-to-evidence gap register and retest plan | Missing-evidence counts; no compliance certification |
| Prior authorization appeal | Source-linked administrative packet and document gap map | Supplied calendar window and missing documents |
| Underpayment | Contract-linked variance register | USD line/net variance, with overpayments retained |
| Discharge barriers | Aggregate barrier owners and dependencies | Supplied barrier counts and reported avoidable days |
Each contract contains domain checks, evidence requirements, decisions and an observable completion gate. See the catalog and case contract. All six have committed synthetic examples and independently specified expected arithmetic in the regression suite.
The existing 51 specialists and 16 workups remain available. Compact role briefs load only the chosen role; detailed original prompts remain optional references. Free-text choose/workup calls offer discovery candidates without selecting a primary specialist or executable workup. The user or host interprets intent, negation, completed work and multiple goals, then supplies validated IDs or a structured selection. Explicit multi-selections retain every workstream. Routing authority and migration documents the deliberate change from lexical auto-selection; routing scores remain discovery heuristics.
Specialist coverage
51 specialists across 10 administrative domains remain available on demand.
| Domain | Specialists |
|---|---|
| Clinical Operations | 8 |
| Emergency Preparedness | 1 |
| Health IT & Informatics | 6 |
| Operations & Administration | 7 |
| Payer & Managed Care | 6 |
| Pharmacy Programs | 2 |
| Population Health & Community Health | 3 |
| Quality, Safety & Compliance | 7 |
| Revenue Cycle & Finance | 6 |
| Strategy & Advisory | 5 |
Build your own workflow
Define the outcome, specialist, human owner, evidence fields and acceptance criteria. Use the example or the builder skill.
node bin/cli.js admin validate workflows/admin-v2/custom-example.json
node bin/cli.js admin build workflows/admin-v2/custom-example.json --output ./community-workflow
The new directory contains SKILL.md, a detailed reference, a source-attributed case template and a hash manifest. Existing output directories are refused. Custom workflows are instruction packs; the builder does not invent executable formulas, install services or grant external-action authority. A valid contract still needs task-level behavioral evaluation.
Customer deployment surfaces
node bin/cli.js admin export codex denial-spike-workup --output ./codex-denials
node bin/cli.js admin export azure denial-spike-workup --output ./azure-denials
| Host | Delivered surface | Qualification |
|---|---|---|
| Codex | Skills, portable plugin manifests and local stdio MCP | Actual packaged-server client tests; pinned native MCP campaign receipt separate |
| Claude | Skills and correlated Messages client-tool callback | Local callback/contract tests; live Claude pending |
| ChatGPT | Plugin manifests and local MCP transport | Local protocol/packaging tested; remote web connection pending |
| Azure / Microsoft Foundry | Instructions and Responses function callback | Local bridge tested; SDK/model/tenant deployment pending |
| Databricks | Instructions and Chat Completions callback | Local bridge tested; hosted ResponsesAgent and tenant deployment pending |
node bin/mcp-server.js --stdio
node bin/host-tool-bridge.js --list azure
Eight read-only MCP tools expose catalog discovery, the six validated calculations and an in-memory workflow draft. Strict schemas, provenance, cancellation, deadlines and structured errors apply. MCP and customer callback guide covers plugin launchers, loopback-only HTTP and actual host requirements. The local Python bridge requires Node and installed dependencies. No cloud SDK, credentials or customer-host service is installed by this package.
Existing installation options and platform exports remain available. Generated text and installation simulations do not prove acceptance by every current live client.
Import source receipts
The optional Data MCP receipt importer maps a pinned public-evidence bundle into one of the six case contracts. It keeps full receipts, lineage, missingness and conflicts in a custody sidecar. Operators supply exact field bindings; absent or contradictory evidence blocks case creation.
node bin/cli.js admin import-evidence examples/admin-v2/data-mcp/denial-spike-workup.bundle.json \
--mapping examples/admin-v2/data-mcp/denial-spike-workup.mapping.json \
--output ./imported-denials --python /path/to/python-with-pydantic
This optional preprocessing step requires Python 3.11+ and the pinned validator dependency. The six examples are wholly synthetic and were generated through the Data MCP producer CLI; they do not establish public metric availability or customer-host deployment. The existing eight MCP tools remain unchanged.
Evidence and boundaries
Every v2 numeric input requires an explicit evidence record with source, origin and as-of date. Synthetic and aggregate modes cannot be silently mixed. Missingness and unknown recoverable cash remain explicit. Unexpected fields, invalid dates and impossible denominators fail validation.
Source-family cards in the existing evidence packs are leads to verify, not automatically verified citations. The bounded primary source review records the new workflows' reviewed context and exact applicability limits. Payer terms, standards, deadlines and clinical facts need the applicable exact source and an accountable owner. External submissions, outreach and production changes require scoped authority. A draft or calculated metric does not demonstrate operational completion.
flowchart LR
Task[Administrative task] --> Route[Workflow or specialist]
Route --> Skill[Compact skill and selected reference]
Sources[Approved evidence and dates] --> Draft[Host-produced artifact]
Skill --> Draft
Sources --> Calc[Optional local aggregate calculation]
Calc --> Draft
Draft --> Review[Qualified human review and authorized action]
Trust and safety covers approved environments, minimum necessary handling, source freshness and human ownership. Existing review protocols and strict frozen-input contracts remain an optional specialist seam.
Verify the candidate
npm run test:admin-v2
# Optional receipt importer, with the pinned Python validator installed:
HAG_EVIDENCE_PYTHON=/path/to/python-with-pydantic npm run test:public-evidence
npm run test:admin-adapter
npm run test:admin-consumer
npm run test:admin-mcp
npm run test:plugin-consumer
npm run test:routing-intent
npm run test:result-boundary
npm test
The new suite tests observable arithmetic, source attribution, missing/contradictory inputs, routing abstention, real CLI file generation and all 30 target/workflow payload combinations. The adapter test exercises all six cases through Python and the actual Node CLI. A separate clean-consumer test installs the npm tarball offline into a disposable project, runs six independent external cases and six provenance failures, resolves payload references from host-style folders, and tests the builder and routing. Release checks retain existing schema, safety, review, installer and package gates.
Offline checks do not measure model quality, customer usefulness or healthcare accuracy. A prior-snapshot native Codex campaign completed twelve synthetic tasks with actual MCP calls and bounded factual/receipt grading; it was not rerun for the revised routing-authority/result-limit contract and it does not establish clinical certification, general routing reliability, customer outcomes or token savings. Candidate status separates local consistency from publication and model/host qualification. Public-channel checks must fail closed when they cannot verify the actual versions.
Eval Status
Historical records report 51/51 evaluated, an average 94.18, and 51/51 tracked improved under the old rubric. These are internal prompt-rubric results, not certification or outcome validation. They are retained in eval/results.tsv; local replay evidence is incomplete. The remaining eval backlog includes repeated independent task-level, domain and customer-host evaluation beyond the bounded native Codex campaign.
Self-Improvement Kit
The historical rubric and role baselines remain immutable. The existing evaluation procedure is retained for explicitly requested historical-loop work. New outcome evidence should record task inputs, independent expected facts, exact model/runtime identity, prompts and source hashes, critical errors, human judgments, cost and latency. No superiority claim is made for the compact instructions.
Maintenance and contributions
Use one versioned release across package, lockfile, plugin and installer. Verify registry and repository payload identity before publication. Keep source-review dates honest; the retained specialist review metadata is not current regulatory verification.
Preserve specialist domain identity, meaningful contracts and owner boundaries. Add realistic negative and conflicting-evidence cases when extending a workflow. Use the workflow contribution guide and report reproducible issues through the repository.
Apache-2.0 · License · Release publishing
For the tested BuilderBob Codex host, the local package profile prepares an independent marketplace and a plugin-owned Node MCP registration using existing Node/npm. Python is optional through --stdio-bridge. Both profiles pass native install, eight-tool/six-workflow calls and exact restoration on Ubuntu/Codex CLI 0.160.0; other hosts need separate acceptance. No cache-path lookup, per-thread server override or security-setting change is required. Run npm run test:codex-profile and npm run test:stdio-bridge; model quality remains a separate exact-configuration check.