graph-agents-cli
Graphs in, agents out.
Build, evaluate and deploy LangGraph agents on your own Kubernetes, with one CLI and six skills for your coding agent.
PyPI Python CI License Docs Skills tuned with SkillOpt Demo video
Get started · Guides · Reference · Skills benchmark · Changelog · Issues
graph-agents-cli is a port of Google's agents-cli to LangGraph on self-hosted Kubernetes: it keeps the lifecycle and the scaffold engine, and replaces ADK on Google Cloud with LangGraph, Helm and any cluster (how the two compare).
Works with your coding agent: Claude Code · Codex · Gemini CLI · Cursor · Antigravity · and others. Or run each command yourself: the CLI works without a coding agent.
What's new
- 2026-09-29 · 0.3.1 is on PyPI. The first release published there:
uv tool install graph-agents-cli(release notes). - 2026-09-29 · 0.3.0: agents calling agents. Agents ask other agents over A2A for the
user they serve, a person's approvals stay with that person, and
systemchecks, wires and deploys several projects as one (guide). Also new: structured final answers in a JSON shape you declare, and skills tuned with Microsoft's SkillOpt (results). Measured while building it (0.3.0 notes):- wiring an agent to five other agents takes 5
peer addcommands instead of 40apicommands and 418 hand-written lines; - replies that were not a bare JSON document (gpt-5-mini, 24-case triage task) went from 28 of 48 to 0 of 96 with structured answers.
- wiring an agent to five other agents takes 5
What you get
| Feature | What it does |
|---|---|
| Scaffold a real service | create renders a LangGraph project: streaming chat API, A2A endpoint, eval harness, hardened Helm chart, GitHub Actions workflows. |
| Run it locally, keyless | run sends one prompt, playground serves a dev chat page; a deterministic fake model needs no key. |
| An eval gate CI can enforce | Deterministic checks and model judges; the exit code of eval run is the gate. |
| Outbound calls under a policy | api-policy.yaml declares the APIs tools may call; anything else is refused, and chosen calls wait for a human. |
| Agents that ask other agents | peer add declares the agents one asks over A2A; system checks, wires and deploys several projects as one. |
| Secure by default | One auth policy on every surface: a shared bearer key, OIDC/JWT or your own. |
| Deploy to any cluster | deploy --env dev, staging or prod with Helm, from your machine, GitHub Actions or Argo CD. |
| Skills for your coding agent | Six skills teach it the same lifecycle. All six are measured on a 104-task benchmark, and SkillOpt tuned three of them: workflow, scaffold and observability. |
Nothing is specific to one domain or one company: a project picks its auth policy, declares its APIs and configures the rest through environment variables and chart values.
Demo video
https://github.com/user-attachments/assets/ebc36951-e765-490e-8c45-d412f1cecba7
Get started
You need Python 3.12 or 3.13 and uv.
setup uses Node.js (npx skills) when it is there; deploying also needs helm, kubectl,
git and a docker that builds with BuildKit.
1. Install the CLI and the skills
uv tool install graph-agents-cli
graph-agents-cli setup # optional: add the six skills to the coding agents it finds
setup also reinstalls the same release from its git tag
(Installation & setup
explains why).
Other ways to install: pipx, the release tag from GitHub, extras
pipx install graph-agents-cli
uv tool install git+https://github.com/ss7172/graph-agents-cli@v0.3.1 # the same release, from its git tag
uv tool install 'graph-agents-cli[a2a,langsmith]' # adds run --mode a2a and eval submit
Mirrors and disconnected installs: Installation & setup.
2. Run your first agent, no model key needed
graph-agents-cli create my-agent && cd my-agent
cp .env.example .env
export MODEL_PROVIDER=fake # the keyless, deterministic test model
graph-agents-cli login --write-env # checks the setup, generates API_KEY
graph-agents-cli install
graph-agents-cli run "What's the weather in San Francisco?"
graph-agents-cli eval run # the exit code is the gate
For a real model, leave out the export line and login --write-env asks for the key. The
provider is one setting: openai, anthropic, gemini, or openai-compatible for Ollama,
vLLM, TGI and others. The Quickstart
walks through each step.
3. Build with your coding agent, then ship it
Once setup has installed the skills, open your coding agent in an empty directory and ask,
for example:
"Use graph-agents-cli to build an agent that answers questions about orders from our orders API (
GET /orders) and can put an order on hold (PATCH /orders/{order_id}). Holding an order must wait for my approval. Use the fake model locally; when the evals pass, deploy it to my local kind cluster."
The skills lead each step and stop for your decision at every gate
(coding-agent tutorial).
Or do it by hand: add an API tool under a policy, pass the eval run gate, then
deploy --env dev to a local kind cluster
(manual tutorial).
Skills, measured
setup installs six skills, all measured on gac-bench: 104 realistic graph-agents-cli tasks
with deterministic verifiers. SkillOpt proposed edits to the workflow, scaffold and
observability skills, and every edit was reviewed by hand before it shipped.
| Skill | What your coding agent learns |
|---|---|
graph-agents-cli-workflow | The lifecycle (understand, scaffold, build, evaluate, deploy, observe), approval before deploy |
graph-agents-cli-langgraph-code | LangGraph patterns the template uses: tools with API_CALLS, checkpointers, streaming, interrupts |
graph-agents-cli-scaffold | create, scaffold enhance, scaffold upgrade and every flag |
graph-agents-cli-eval | The eval gate, dataset schema, checks, judges and quality metrics |
graph-agents-cli-deploy | Environments, rollouts and rollback, secrets, the Helm chart, Argo CD |
graph-agents-cli-observability | Opt-in tracing (LangSmith or OTLP), JSON logs, /metrics, /health and /ready |
0.3 skills against 0.2 skills on gac-bench (hard pass rate, same CLI build, only the skill text differs):
| Harness | Tasks | 0.2 skills | 0.3 skills | Change |
|---|---|---|---|---|
| Claude Code (claude-sonnet-5-5) | 104 | 0.84 | 0.97 | +0.12 (p = 0.0006) |
| Codex (gpt-5.6-terra) | 30 | 0.77 | 0.97 | +0.20 (p = 0.031) |
Codex ran the test split of all six skills plus val of workflow and observability. Change is the mean of per-task differences, so it need not equal the difference of the two scores.
Where the gain is, per skill (Claude Code, val and test tasks)
| Skill | 0.2 skills | 0.3 skills | Change |
|---|---|---|---|
graph-agents-cli-workflow | 0.36 | 1.00 | +0.64 |
graph-agents-cli-scaffold | 0.61 | 1.00 | +0.39 |
The other four skills scored 0.92 or higher before and moved by 0.00 to +0.06.
No test task went down on either harness. There is one real regression on Claude Code (1 of 2 reps of a new deploy task in the val split), and the test-split gains alone are not significant (p = 0.25 and 0.5). Full tables: final-v0.3.md.
Commands
| Command | What it does |
|---|---|
graph-agents-cli setup | Install the CLI and the skills into detected coding agents |
graph-agents-cli create <name> | Create a LangGraph agent project from a template |
graph-agents-cli run "prompt" | Run the agent with a single prompt |
graph-agents-cli eval run | Run the agent over the eval dataset and grade it |
graph-agents-cli deploy --env dev | Deploy to the current Kubernetes context |
All commands
| Command | What it does |
|---|---|
login · info · install | Check provider keys, LangSmith and kubeconfig (optionally write .env) · show configuration and version · install project dependencies |
create · scaffold | Create a project · scaffold, enhance and upgrade projects |
run · playground · lint | One prompt · the app locally with reload and a dev chat page · code checks plus the API-policy check |
api · approvals · auth | Declare the outbound APIs tools may call · list and decide gated calls · dev-only JWTs |
peer · system | Declare the agents this agent asks over A2A · check, wire and deploy agents that call each other |
eval | Evaluate agents and compare results |
build · deploy · secrets · infra | Build the image · deploy to Kubernetes · the app's Secret per environment · check prerequisites (read-only) |
setup · update · extension | Install skills · force-reinstall them · manage extensions (experimental) |
graph-agents-cli <command> --help prints the same reference as the site's
CLI page.
Documentation
| Section | Pages |
|---|---|
| Get started | Installation · Quickstart · Build with a coding agent · Manual workflow · The lifecycle |
| Build | Develop · Authentication · API policy · Human approval · Agents calling agents · Evaluation · Extensions |
| Operate | Deploy · CI/CD · Secrets · Observability · Upgrading · Offline · Security |
| Reference | CLI · Environment · HTTP API · api-policy.yaml · System file · Manifest · Exit codes · Skills · Compared with google-agents-cli |
Status
Version 0.3.1, alpha, on PyPI and as tags on GitHub (release notes). Interfaces may still change between minor versions; the changelog lists every breaking change with its migration steps.
0.3's acceptance tests covered a system of six agents built with peer add and system apply
on a local cluster, security probes, a probe of 20 agents with 2 replicas each, an issuer that
hangs, and an upgrade from 0.2.0 with data. Blocker and major issues are fixed before a
release; the rest are parked.
- Known issues: parked issues, each with its impact and a workaround where one exists.
- Changelog: every release and its migration steps.
- Contributing: development setup, tests, templates, releases and the upstream-sync process.
Credits
graph-agents-cli stands on two projects:
- google-agents-cli, Copyright 2026 Google LLC, Apache-2.0. graph-agents-cli is a fork of it, started from 1.6.1, with the Google Cloud specific parts removed; files kept from it keep their Google LLC copyright headers. NOTICE lists the modifications.
- SkillOpt from Microsoft, which found the skill rules that 0.3 adopted after review.
License
Apache-2.0: see LICENSE.