Caldera docs
calderaapp.io

Operating agents

The procedures for an agent that is connected to Caldera: install, update, pin, roll back, pause and resume with and without Caldera, update the runtime, rotate a token, delete, move an existing agent onto the runtime, and a troubleshooting table from symptom to fix.

Conventions: <dir> is the agent's directory on its machine (for example ~/caldera-agents/my-agent), <name> its name. Commands on the machine are python3 caldera.pyz ...; the download sits wherever it was fetched, and the runtime keeps a copy at <dir>/.caldera/caldera.pyz.

Quick reference

I want to Where How
See if an agent is alive Caldera, or the machine Overview card; python3 caldera.pyz status --dir <dir>
Pause Caldera Agent page, "Pause"
Pause without Caldera the machine python3 caldera.pyz pause --dir <dir> (runtime agents) or the agent's own channel
Update Caldera Agent page, Blueprint panel, "Move to this version"
Go back Caldera or the machine "Roll back"; or python3 caldera.pyz rollback --dir <dir>
Change a setting Caldera Agent page, Settings tab (admins)
Ask for a log Caldera Runs tab or Overview, "View log" (admins)
Install a new agent Caldera, then the machine Blueprints, "Install", then caldera init
Update the runtime Caldera Agent page, Runtime panel, "Update runtime" (admins)

Install an agent from a blueprint

Needs: an admin role in an activated workspace; a published version; a machine with Python 3 and (for the service) systemd; outbound HTTPS to app.calderaapp.io.

  1. In Caldera: Blueprints, the blueprint (or a catalog one), "Install...". Pick the version, give the agent name (lowercase letters, digits, hyphens; not an existing agent's name), press "Create install code". If a binding's product is not activated, or the workspace would exceed ten bots, the dialog says so now.
  2. Copy both commands now. The code is shown once and expires in 15 minutes.
  3. On the machine, run the first command (fetch and verify). On macOS use the variant with shasum -a 256 -c -. If the check fails, stop; see the troubleshooting table.
  4. Run the second command: python3 caldera.pyz init <code> --url https://app.calderaapp.io. Add --dir <dir> to choose the directory (default ~/caldera-agents/<name>). It asks for the declared secrets (hidden input); for unattended setups set CALDERA_SECRET_<NAME> instead.
  5. It installs and starts caldera-<name>.service. For it to keep running when you are logged out: loginctl enable-linger <user>.
  6. Check: python3 caldera.pyz status --dir <dir> shows process running and last report ok, and the agent appears in the overview within about a minute.
  7. Do the independence check.

If init fails after the code was spent, the agent exists in Caldera without a machine. Delete it in Caldera and issue a new code.

Register an agent with an own reporter

  1. Overview, "New agent": a name and the reporting interval if another is wanted.
  2. Copy the token and the two environment lines now; they are shown once.
  3. Put them in the reporter's private secrets file; start the reporter; check the overview. See Connecting an agent.

Update to a new version

  1. Raise version in the manifest to a number that has never been published (published versions are immutable, and a draft may be replaced only until it is published), rebuild the bundle, upload it as a version of the blueprint and publish it.
  2. On the agent's page, Blueprint panel: choose the version, "Move to this version". Caldera stores the wish and sends it in the answer to the agent's next report.
  3. Follow the line under the panel. It goes: waiting for the agent's next report, then applied, rolled back by the agent, or failed, with the agent's own message.
  4. After an applied update, the agent watches itself for 30 minutes. If its level gets worse in that time, it rolls back by itself, and the panel then shows the rolled-back outcome.

What the runtime does is in the runtime. In short: fetch, stage beside the running version, swap, restart, run health, keep or go back. Everything in state paths is untouched.

A version the agent rolled back from or refused is not applied again until a different version is named. To try a fixed version, publish a new version number.

Pin

Pinning asks the agent to stay on the version it runs: on the agent's page press "Pin". Technically it sets the wished version to the running one, which the agent sees as nothing to do. It also stops a stale wish for a newer version from being applied. "Pin" is disabled while the wish already equals the running version.

Roll back

  • From Caldera: "Roll back" names the previous published version. Use this when the agent is healthy but the old behaviour is wanted.
  • On the machine, without Caldera: python3 caldera.pyz rollback --dir <dir>. It swaps to the previous version and restarts, with no network. Only two versions are kept, so a second rollback goes forward again.
  • After a local rollback the version it left is on hold: the next report will not undo it, even if Caldera still wishes that version. status shows held <version> is not applied until Caldera names another version. Then pin (name the version that runs) or name another version, so Caldera's page agrees.
  • A catalog agent can only fetch the one shipped version, so for it the local rollback is the way back.

The hold clears when a different version is applied successfully.

Change a setting, run now, ask for a log

  • Settings (admins): Settings tab, change, "Apply". The agent checks every change against its own bounds and answers in its own words; "A change is waiting for the next report of the agent" is shown until it does. A refusal reads "Nothing changed: : ." and changes nothing (all or nothing). Caldera repeats the bounds check only to tell a person at once.
  • Run now: needs the agent to offer run_now in controls; refused while paused.
  • Log: admins only. Caldera asks the agent for one of the log names it reported; the agent sends the last 512 KB. The text is kept 30 days with the command.

A command is delivered at most once. If a report's answer is lost, the command is lost and fails after ten minutes (when the agent next reports). Send it again.

Pause and resume, with and without Caldera

Everything Caldera does can be done without it. Know the offline way for every agent before it is needed.

Agent Pause Resume
Any runtime agent python3 caldera.pyz pause --dir <dir> python3 caldera.pyz resume --dir <dir>
An agent in adapter mode the same commands; they call the adapter the same
An agent with an own reporter the agent's own channel (a chat command, a file) the same
  • A pause set in Caldera is the same pause: the same file and the same stop.
  • For a runtime agent, pause stops the agent's process; the runtime keeps reporting, and the state shows "Paused" (warn).
  • Caldera disables the controls for a silent agent ("The agent is not reachable, so it cannot take a command now").
  • A runtime on a remote machine is paused by logging in to it and running the command there. Write the exact command down next to the agent's documentation.

Update the runtime itself

The runtime updates itself on a press, signed (how it works is in the runtime):

  1. On the agent's page, the Runtime panel shows the version the agent runs, the one this Caldera serves, a pin and the file kept before. "Update runtime" (admins only; members read) sets a wish. The next answer to a report carries it. The panel follows it: waiting for the agent's report, "has to prove itself first" while the new runtime runs its five minutes, then the runtime's own words for taken, gone back or not taken. A version that went back or failed is not offered again until Caldera serves another one, and the button is gone until then. A runtime that reports no version (older than this feature, or not caldera.pyz at all) gets a notice instead of a button: it is replaced by hand once, below.
  2. The runtime downloads /cli/caldera.pyz and its signature, checks the checksum and the signature against the public key built into it, and lets the new file say its version. A bad or missing signature, or a machine without ssh-keygen: nothing changes and the agent page shows the outcome failed with the reason.
  3. It keeps the old file as <dir>/.caldera/caldera.pyz.previous, swaps the file the unit names, and restarts in place: same process, the agent's own processes never stop.
  4. The first successful report within five minutes makes it good. A new runtime that crashes, or does not report in time, goes back by itself; that version is then held until a newer one is offered.

On the machine, with no connection at all:

python3 <dir>/.caldera/caldera.pyz rollback-runtime --dir <dir>   # back to the previous file
python3 <dir>/.caldera/caldera.pyz pin-runtime 0.32.0 --dir <dir>  # never update to another version
python3 <dir>/.caldera/caldera.pyz pin-runtime off --dir <dir>     # release the pin

One-time catch: a runtime installed before this existed cannot update itself. Replace <dir>/.caldera/caldera.pyz once by hand (mode 700, checksum verified) and run systemctl --user restart caldera-<name>.service; the unit has KillMode=process, so the agent keeps running. After that, updates are presses. A runtime started from anywhere but <dir>/.caldera/caldera.pyz (the unit's path) cannot replace itself and says so.

Which Claude account an agent runs under

The Runtime panel on the agent's page says which Claude account the machine is signed in to. See The Claude account hint.

Rotate an agent's token

Do this when a token may have leaked, or was lost.

  1. Caldera: agent menu, "Rotate token...". The old token stops working at once. The new one is shown once, with the two environment lines.
  2. The agent is now silent in Caldera until the machine has the new token.
  3. Runtime agent: edit <dir>/.caldera/agent.env, replace the value of CALDERA_TOKEN (keep the file mode 600), then systemctl --user restart caldera-<name>.service. The agent keeps running: it never held the token, and a restarted runtime adopts it.
  4. Own reporter: update its secrets file and restart the reporter service.
  5. Check that the agent reports again.

The runtime has no command that fetches the new token; Caldera never writes to the machine.

Binding tokens (access to other products) are not rotated here. They are the bot's tokens, handled in the product's own bot settings; see the table below ("binding is broken").

Delete an agent

Caldera: agent menu, "Delete...". The confirmation names what stays.

  • Caldera deletes the agent, its runs and its commands. Its token stops working at once.
  • Nothing on the machine changes. The runtime keeps running the agent and logs a refused report (HTTP 401) every interval.
  • The bots are not revoked. Only the product that owns a bot can revoke it. The confirmation lists them by name ("These bots stay: ..."). Revoke them in that product's bot settings, or the tokens stay valid until they expire.
  • To remove the agent from the machine, in this order:
    1. python3 caldera.pyz pause --dir <dir> (this stops the agent's process; disabling the unit first would leave it running),
    2. systemctl --user disable --now caldera-<name>.service,
    3. remove ~/.config/systemd/user/caldera-<name>.service and systemctl --user daemon-reload,
    4. remove <dir> (it holds the agent's state; keep it if it is wanted).

A blueprint can only be deleted when no agent is installed from it.

Move an existing agent onto the runtime

It works when the agent can run as one process the runtime may start and stop. (An agent with its own timers and a pause that other things set uses adapter mode, or stays on its own reporter.)

Before: the agent has an own reporter and its own timers. Nothing changes until the steps start.

  1. Build the bundle and upload it as the first version of a blueprint (see Blueprints); publish.
  2. Rename the old agent in Caldera (for example <name>-old): the install needs the name to be free.
  3. Issue an install code for the blueprint, with the agent's name.
  4. On the machine, stop the old mode first, so nothing runs twice: disable the old timers and the old reporter service.
  5. Run the two install commands; init writes the directory and the unit.
  6. Check: status, the unit's journal (journalctl --user -u caldera-<name>.service -n 20), the agent in Caldera, "Run now" once, and the independence check.
  7. When it runs, delete the old agent in Caldera. Leave the old unit files in place: the way back needs them.

The way back: disable the new unit and start the old mode again.

systemctl --user disable --now caldera-<name>.service
systemctl --user enable --now <the old timers and the old reporter service>

In Caldera rename the old agent back and delete the new one. The old and the new mode must never run at the same time.

The quota gauge looks wrong

Symptom. An agent's quota gauge in Caldera does not match the provider's usage page. For example Caldera shows 79 % for the week while the provider shows 90 %.

Cause. This is a property of how the agent measures, not of Caldera. A gauge that an agent builds from the quota report of its own last model call is only as fresh as that call. Two things make it stale:

  1. Other sessions use the same quota. The subscription is shared with interactive sessions on the same account, so the real figure can be higher than the last one the agent saw.
  2. The agent holds runs above its threshold. Above its own limits the agent starts no run, and without a run there is no new observation. The gauge then stays where it was.

A window whose reset time has passed counts as empty (0), so a stale gauge can also fall to 0 at a reset the agent has not seen a call after.

What helps. An agent that builds a gauge from its own last observation should

  • put when it observed into the gauge: the first words of the gauge's detail are the time of the last measurement (for example "As of 06.10. 14:05"), so that a stale value is recognisable as stale;
  • refresh the observation by itself, at most once an hour, also while it is paused or held, with a request that asks for nothing but the quota (no context, no tools, the smallest model); and say in its log when that fails and that the last value stays.

There is no central quota gate in Caldera and none is planned: every agent holds its own limits.

Check. Read the time in the gauge's detail line. If it is older than about an hour, the measurement fails: look in the agent's own log, and check that the program that makes the measurement can be found. Compare the figure with the provider's usage page.

Fix. Caldera shows what the agent reported; there is nothing to change in Caldera. Do not treat a gauge without a measurement time as live.

Troubleshooting

Start with the agent's page (level, last report, "Process" and "Binding" health items) and, on the machine, python3 caldera.pyz status --dir <dir>.

Symptom Likely cause Check Fix
Agent shown as not reachable (silent, level fault) No report for three intervals (at least two minutes): machine off, runtime stopped, no network, token rotated or agent deleted, workspace deactivated status: runtime, last report ... ago; journalctl --user -u caldera-<name>.service; <dir>/.caldera/runtime.log Start the unit; or fix the cause in the rows below
Log says report failed: HTTP 401 (unauthorized) The token was rotated or the agent was deleted The agent's menu in Caldera; does the agent still exist? Rotate the token and update the machine (above); or reinstall
HTTP 403 The workspace is not activated for Caldera (or lost it) The workspace's state in the app Ask for the workspace to be activated
HTTP 404 on /api/agent/v1/report The URL is not Caldera's own address (another product's address answers 404) curl -i https://app.calderaapp.io/api/agent/v1/report should answer 401, not 404 Use Caldera's own address as CALDERA_URL
HTTP 400 on a report The body is not valid JSON, not an object, or names another protocol An own reporter: print the response body Send "protocol": 1 and a JSON object
HTTP 413 The report is over 256 KB (a log over 512 KB) Size of the body Send less; logs: only the tail
HTTP 429 More than 60 reports or 10 logs a minute from one agent (30 bundle fetches; 20 install calls per address) Interval in the reporter's loop Keep to the interval the answer asks for
caldera.pyz download: checksum does not match and the file is HTML A proxy or a wrong address served a web page instead of the file head -c 200 caldera.pyz shows HTML Download from https://app.calderaapp.io/cli/caldera.pyz itself, without a proxy that rewrites
/cli/caldera.pyz answers 404 The address is not Caldera's own, or this Caldera serves no runtime build, or a proxy in between changes the Host curl -i https://app.calderaapp.io/cli/caldera.pyz.sha256 Use Caldera's own address; forward Host unchanged
sha256sum: command not found macOS Use the macOS variant in the dialog (shasum -a 256 -c -)
init: unknown, expired or used code (404) All three look the same on purpose The code's age (15 minutes); was it run before? Issue a new code
init: ... already holds an agent The target directory has .caldera/config.json Choose another --dir, or remove the old one
init: a secret is needed No terminal and no CALDERA_SECRET_<NAME> Set the variable or run in a terminal
init fails after the code was spent Unpacking or writing failed The message says the agent exists in Caldera without a machine Delete the agent in Caldera, issue a new code
Dialog refuses with a refusal reason See the refusal reasons Mend the named cause; the code stays usable if redemption refused
Agent shows "Not running" (fault) The agent process keeps exiting <dir>/.caldera/agent.out.log; status shows process not running; the report says "restarting in N s" Fix the agent. Restarts back off from 5 s to 300 s; "Run now" skips the wait
The service stops when you log out systemd user linger is off loginctl show-user <user> loginctl enable-linger <user>
A command stays "waiting for the agent's next report" Normal for up to one interval (30 s; 10 s while the page is open). If it lasts: the agent is silent Last report time; the reporting interval Wait; if silent, see the first row. A command the agent never answers fails after ten minutes, at its next report
"Moving to X: waiting for the agent's next report" stays The agent is silent; or the runtime holds that version (it rolled back from or refused it) and will not apply it; or the wish names a blueprint that is not the agent's own status: a held line; last update line Silent: see the first row. Held: name a different version, or publish a new version number
Update ended "rolled back" at once The new version's health command failed last update ... in status; run the health command by hand in <dir>/versions/<version> with the agent's environment Fix the health check or the version; publish a new number
Update ended "failed: could not fetch" Caldera was unreachable during the fetch It retries after ten minutes by itself
Update ended "failed: refused X: ..." The bundle was invalid or did not match its checksum The message names the reason (path, size, checksum, missing file) Fix the archive; publish a new number
The agent rolled back on its own after being fine Its level got worse within 30 minutes of the update status, last update message Investigate the version; publish a fixed one
Runtime update ended "failed" with a signature or ssh-keygen message The signature is missing or does not verify, or the machine has no ssh-keygen The message on the Runtime panel Install OpenSSH on the machine; for a missing signature, wait for the next release
The Runtime panel offers no "Update runtime" The version went back or failed and is held; the runtime is pinned; it is already at the served version; or it reports no version The notice on the panel Wait for the next release; release the pin with pin-runtime off; or replace an old runtime by hand once
A setting is refused ("Nothing changed: ...") The value is outside the manifest's bounds, or the key is not declared The message names the label and the rule Use an allowed value
A log answers "There is no such log." The name is not declared in the manifest's logs, or the file does not exist or resolves outside the agent's directory The manifest's logs; the file under the version directory Declare it, or fix the path
The log dialog says the agent did not answer The agent is silent, or busy beyond ten minutes Try again when it reports
A binding is broken (agent page, "Access to other products") Binding <product> is bad: the bot or its token was deleted or revoked, or the token expired (they last a year) The product's bot settings; the health item's detail In that product's settings create a new token for the bot (or a new bot); edit CALDERA_BINDING_<PRODUCT>_TOKEN in <dir>/.caldera/agent.env; pause then resume. Caldera cannot mend it.
Binding shows warn The product could not be reached from the machine Usually transient
The overview is empty for an admin Wrong workspace; or the workspace is not activated ("not set up for Caldera yet") The workspace switcher Switch; or ask for activation
A member cannot steer Members only see Role An admin does it
The quota gauge is wrong The agent's last measurement is old (read the time) See the quota gauge See there
Agent shows a commit-like version (1c1f3c5) and no Blueprint panel It reports with its own reporter (registered by hand) Normal; updates are the agent's own
python3: not found or a syntax error Python is missing or very old python3 --version Use Python 3.10 or newer (the lowest tested)

Refusal reasons

Caldera's refusals carry a reason (and params for the sentence); the app writes the sentence in the reader's language. The server's own message is English prose for logs.

Reason Means Mend
product_not_activated The workspace is not activated for Caldera Ask for activation
role_cannot_view A guest cannot see this Use a member or admin
blueprint_name_taken A blueprint of that name exists Choose another name
blueprint_has_agents Agents are installed from it Delete them first
version_published That version is published and cannot change Upload a new version number
version_upload_race The same version was uploaded at the same moment Try again
version_not_deletable A published version cannot be deleted None; publish a new one
version_unpublished Publish before installing Publish
manifest_name_mismatch The manifest's name is not the blueprint's Make them equal
upload_file_missing, archive_empty, archive_not_zip, archive_damaged, archive_entry_damaged, archive_unsupported The upload is not a usable zip (no file, empty, not a zip, damaged, or zip64, several disks, encryption, unknown compression) Repack
archive_too_large (20 MB), archive_too_many_entries (5000), archive_unpacks_too_large (100 MB), archive_entry_too_large (50 MB) A size limit Make the bundle smaller
archive_unsafe_path, archive_bad_entry A path that is not allowed, or an entry that is a link, device, duplicate, or file-that-is-also-a-directory Remove it
archive_no_manifest No caldera.agent.yaml at the top level Put it at the root
archive_file_missing A file the manifest names is missing Add it
manifest_too_large, manifest_not_yaml, manifest_invalid The manifest is over 256 KB, not valid YAML, or invalid at the named path Fix the field
agent_name_taken The workspace has an agent of that name Choose another name
binding_product_not_activated A bound product is not activated for the workspace Activate it there
binding_product_no_address This Caldera has no address for that product Not something a workspace can change; ask for help
bot_role_mismatch A bot of that name exists with another role Rename the agent, or change the bot
role_cannot_lend The issuer's role may not create a bot with that role Issue as an admin or owner
bot_limit The workspace would hold more than ten bots Remove bots
issuer_cannot_install The person who issued the code may no longer install agents Issue a new code as an admin
instance_no_caldera_address This Caldera cannot build the commands Not something a workspace can change; ask for help
agent_not_from_blueprint The agent was registered by hand and has no blueprint Not applicable
version_not_installable Not a published version of the agent's blueprint Publish it, or name another
runtime_version_held The wished runtime went back or failed and is held Wait until Caldera serves another version
runtime_too_old The runtime reports no version Replace it by hand once
runtime_not_served This Caldera serves no runtime build Not something a workspace can change; ask for help
runtime_up_to_date The runtime already runs the served version None
runtime_pinned The runtime is pinned to another version pin-runtime off on the machine