Making a CLI that AI agents can use safely
More and more, the “user” of a command-line tool is an AI coding agent. It runs your command, reads stdout, checks the exit code, and decides what to do next. It does not see a spinner, cannot answer a hidden prompt, and will happily believe an exit code of 0.
That changes what a good CLI looks like. This post collects five rules from the Command Line Interface Guidelines and current agent conventions, and shows how Devpit applies them in its Accounts commands, where a wrong guess means pushing or deploying as the wrong person.
The short answer
Section titled “The short answer”- Machine output on request:
--jsonon every command that reads state. - No silent changes: anything that changes state asks a human, or needs
--yes. - Exit codes that mean something, documented, and never 0 for “did nothing”.
- Tell agents how to use you: a skill, an
AGENTS.mdsnippet, anllms.txt. - Never put secrets in output, not even in JSON.
Rule 1: –json for everything that reads
Section titled “Rule 1: –json for everything that reads”The Command Line Interface Guidelines (clig.dev) say: “Display output as formatted JSON if --json is passed.” Text output is for people. It changes wording between versions, wraps at the terminal width and mixes in colour. JSON with stable field names is something an agent can parse today and next year.
Two tools agents already use show the details that matter:
claude auth status # JSON by default; --text for peoplegh auth status --json hosts # JSON only when askedclaude auth status “exits with code 0 if logged in, 1 if not”, so the exit code alone answers the question. gh auth status takes a different line, and documents it: “when using the --json option, the command will always exit with zero regardless of any authentication issues”. Both are fine because both are written down. An agent reading either can do the right thing.
Also from clig.dev: “Send output to stdout.” and “Send messaging to stderr.” Keep progress and warnings out of the stream an agent parses.
Rule 2: no silent changes
Section titled “Rule 2: no silent changes”clig.dev again: “Only use prompts or interactive elements if stdin is an interactive terminal (a TTY)”, and “If stdin is not an interactive terminal, skip prompting and just require those flags/args.” For dangerous actions it suggests a prompt when interactive, and “requiring them to pass -f or --force otherwise”.
For an agent, that turns into a simple contract:
- In a terminal with a person, show a preview and ask, with No as the default.
- Without a terminal, refuse, explain, and name the flag. The agent then asks its user, and runs again with
--yesonly after the user agrees.
The important part is what triggers the refusal. Detect “no terminal”, not “an agent”. Environment variables such as CI or an agent’s marker may make output plainer, but they should never loosen confirmation. Otherwise any process that sets the right variable skips the question.
Rule 3: exit codes that mean something
Section titled “Rule 3: exit codes that mean something”The baseline from clig.dev: “Return zero exit code on success, non-zero on failure”, and “Map the non-zero exit codes to the most important failure modes.”
There is no universal table. Bash returns 2 from its builtins “to indicate incorrect usage”, 126 for “not executable” and 127 for “not found”. BSD’s sysexits codes (64 for usage and so on) exist, but OpenBSD’s own manual says “A few programs exit with the following non-portable error codes. Do not use them.” So: choose a small set, document it in --help, and keep it stable.
The rule people break most is the quiet one: a command that did nothing must not exit 0. A placeholder command that prints “coming soon” and exits 0 tells an agent the job is done.
In PowerShell, an agent or a script reads the code from $LASTEXITCODE:
claude auth status | Out-Nullif ($LASTEXITCODE -ne 0) { 'not signed in' }Rule 4: tell agents how to use you
Section titled “Rule 4: tell agents how to use you”Agents look for instructions in a few standard places. A CLI can meet them in each.
A skill. Claude Code skills “follow the Agent Skills open standard”. The spec at agentskills.io: “The SKILL.md file must contain YAML frontmatter followed by Markdown content”. name is required (“Lowercase letters, numbers, and hyphens only”, and it “must match the parent directory name”), and so is description, which says “what the skill does and when to use it”. Personal skills live in ~/.claude/skills/<skill-name>/SKILL.md. A minimal one for your own tool:
---name: mytooldescription: Use mytool to check and change which account a folder uses. Use it before pushing or deploying.---
- Read state with `mytool status --json` first.- Never change anything without asking the user. Changes need `--yes`.- Exit code 3 means the user must agree first. Ask them, then rerun with --yes.An AGENTS.md snippet. agents.md describes AGENTS.md as “a simple, open format for guiding coding agents”, a “README for agents” that lives in a repository. A CLI can print a short snippet users paste into their own AGENTS.md, for agents that do not read skills.
An llms.txt. The llmstxt.org proposal: “We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content.” It starts with an H1, a short summary in a blockquote, and lists of links to the pages that matter. Put your command summary there, so an agent that only fetches your website still finds the right commands.
Rule 5: never put a secret in output
Section titled “Rule 5: never put a secret in output”Agents copy output into their context, logs and sometimes their answers. Anything you print may end up somewhere you did not plan.
Real examples of the trap: gh auth status has a --show-token option, and firebase login:list --json includes each account’s tokens next to the emails (we read this in the firebase-tools 15.32.1 source). A CLI meant for agents should answer “who am I” with names and emails only, and make the token path impossible to reach by accident.
How Devpit applies these rules
Section titled “How Devpit applies these rules”Devpit’s account commands (devpit <tool>, devpit <tool> use, run, list, add, devpit undo, devpit accounts verify) follow all five. The full list is in the command line reference.
--jsonon every read command (devpit <tool>,list,accounts verify), with stable field names. Each command’s JSON is pinned by a test.- Interactive means a person. A command counts as interactive only when stdin and stdout are terminals. Agent variables such as
CLAUDECODEorCImay make output plainer, never less careful. - No silent changes. Not interactive and no
--yes? Devpit exits with code 3 and says: “This would change which account is used. Re-run with –yes after the user agrees.” - Documented exit codes, in
--help: 0 ok, 1 failed, 2 wrong usage, 3 needs--yes, 4 verify found a mismatch. - Placeholder commands exit non-zero. Devpit found its own older stub commands exited 0 while doing nothing, and changed them, because agents trust exit codes.
- Help ends with real examples. A test fails if a command has none.
devpit agent installwrites adevpitskill (SKILL.mdwith standard frontmatter) into every Claude Code account and prints anAGENTS.mdsnippet for other agents. It says: when to use Devpit, check before switching, ask the user first, never read account folders.devpit agent removetakes it out again.llms.txton the website includes the command summary: devpit.zubyr.dev/llms.txt.- No tokens in any output. Devpit stores none and prints none. A test plants a fake token in fake tools and checks that it never appears in events, logs or
--jsonoutput. More in Safety and privacy.
Checklist for your own CLI
Section titled “Checklist for your own CLI”- Add
--jsonto every command that reads state, and keep the field names stable. - Send data to stdout and messages to stderr.
- Prompt only when stdin is a terminal; otherwise require
--yesand say so. - Never loosen safety because you detected an agent.
- Document a small set of exit codes. Never exit 0 for “nothing happened”.
- Ship a skill, an
AGENTS.mdsnippet and anllms.txt. - Keep tokens out of every output mode.
Sources
Section titled “Sources”Common questions
Should a CLI detect AI agents and behave differently?
It can make output plainer, for example no colour or spinners, but it should never loosen safety because of it. Decide on prompts by whether stdin and stdout are terminals, and require an explicit flag such as --yes for changes when they are not.
Which exit codes should I use?
Zero for success and non-zero for failure is the one rule everyone agrees on. Beyond that, pick a small documented set for the failures a caller can act on, such as wrong usage or needs confirmation, and keep it stable.
What is the difference between AGENTS.md, a skill and llms.txt?
AGENTS.md is a file in a repository that tells coding agents how to work there. A skill is a SKILL.md folder an agent loads when a task matches its description. llms.txt is a Markdown file on a website that points language models at the right pages.
