Operating agents
The procedures for an agent that is connected to Caldera: install, update, pin, roll back, pause and resume with and without Caldera, update the runtime, rotate a token, delete, move an existing agent onto the runtime, and a troubleshooting table from symptom to fix.
Conventions: <dir> is the agent's directory on its machine (for example
~/caldera-agents/my-agent), <name> its name. Commands on the machine are
python3 caldera.pyz ...; the download sits wherever it was fetched, and the
runtime keeps a copy at <dir>/.caldera/caldera.pyz.
Quick reference
| I want to | Where | How |
|---|---|---|
| See if an agent is alive | Caldera, or the machine | Overview card; python3 caldera.pyz status --dir <dir> |
| Pause | Caldera | Agent page, "Pause" |
| Pause without Caldera | the machine | python3 caldera.pyz pause --dir <dir> (runtime agents) or the agent's own channel |
| Update | Caldera | Agent page, Blueprint panel, "Move to this version" |
| Go back | Caldera or the machine | "Roll back"; or python3 caldera.pyz rollback --dir <dir> |
| Change a setting | Caldera | Agent page, Settings tab (admins) |
| Ask for a log | Caldera | Runs tab or Overview, "View log" (admins) |
| Install a new agent | Caldera, then the machine | Blueprints, "Install", then caldera init |
| Update the runtime | Caldera | Agent page, Runtime panel, "Update runtime" (admins) |
Install an agent from a blueprint
Needs: an admin role in an activated workspace; a published version; a machine
with Python 3 and (for the service) systemd; outbound HTTPS to
app.calderaapp.io.
- In Caldera: Blueprints, the blueprint (or a catalog one), "Install...". Pick the version, give the agent name (lowercase letters, digits, hyphens; not an existing agent's name), press "Create install code". If a binding's product is not activated, or the workspace would exceed ten bots, the dialog says so now.
- Copy both commands now. The code is shown once and expires in 15 minutes.
- On the machine, run the first command (fetch and verify). On macOS use the
variant with
shasum -a 256 -c -. If the check fails, stop; see the troubleshooting table. - Run the second command:
python3 caldera.pyz init <code> --url https://app.calderaapp.io. Add--dir <dir>to choose the directory (default~/caldera-agents/<name>). It asks for the declared secrets (hidden input); for unattended setups setCALDERA_SECRET_<NAME>instead. - It installs and starts
caldera-<name>.service. For it to keep running when you are logged out:loginctl enable-linger <user>. - Check:
python3 caldera.pyz status --dir <dir>showsprocess runningandlast report ok, and the agent appears in the overview within about a minute. - Do the independence check.
If init fails after the code was spent, the agent exists in Caldera without a
machine. Delete it in Caldera and issue a new code.
Register an agent with an own reporter
- Overview, "New agent": a name and the reporting interval if another is wanted.
- Copy the token and the two environment lines now; they are shown once.
- Put them in the reporter's private secrets file; start the reporter; check the overview. See Connecting an agent.
Update to a new version
- Raise
versionin the manifest to a number that has never been published (published versions are immutable, and a draft may be replaced only until it is published), rebuild the bundle, upload it as a version of the blueprint and publish it. - On the agent's page, Blueprint panel: choose the version, "Move to this version". Caldera stores the wish and sends it in the answer to the agent's next report.
- Follow the line under the panel. It goes: waiting for the agent's next report, then applied, rolled back by the agent, or failed, with the agent's own message.
- After an applied update, the agent watches itself for 30 minutes. If its level gets worse in that time, it rolls back by itself, and the panel then shows the rolled-back outcome.
What the runtime does is in the runtime.
In short: fetch, stage beside the running version, swap, restart, run health, keep
or go back. Everything in state paths is untouched.
A version the agent rolled back from or refused is not applied again until a different version is named. To try a fixed version, publish a new version number.
Pin
Pinning asks the agent to stay on the version it runs: on the agent's page press "Pin". Technically it sets the wished version to the running one, which the agent sees as nothing to do. It also stops a stale wish for a newer version from being applied. "Pin" is disabled while the wish already equals the running version.
Roll back
- From Caldera: "Roll back" names the previous published version. Use this when the agent is healthy but the old behaviour is wanted.
- On the machine, without Caldera:
python3 caldera.pyz rollback --dir <dir>. It swaps to the previous version and restarts, with no network. Only two versions are kept, so a secondrollbackgoes forward again. - After a local rollback the version it left is on hold: the next report will
not undo it, even if Caldera still wishes that version.
statusshowsheld <version> is not applied until Caldera names another version. Then pin (name the version that runs) or name another version, so Caldera's page agrees. - A catalog agent can only fetch the one shipped version, so for it the local rollback is the way back.
The hold clears when a different version is applied successfully.
Change a setting, run now, ask for a log
- Settings (admins): Settings tab, change, "Apply". The agent checks every change against its own bounds and answers in its own words; "A change is waiting for the next report of the agent" is shown until it does. A refusal reads "Nothing changed: : ." and changes nothing (all or nothing). Caldera repeats the bounds check only to tell a person at once.
- Run now: needs the agent to offer
run_nowincontrols; refused while paused. - Log: admins only. Caldera asks the agent for one of the log names it reported; the agent sends the last 512 KB. The text is kept 30 days with the command.
A command is delivered at most once. If a report's answer is lost, the command is lost and fails after ten minutes (when the agent next reports). Send it again.
Pause and resume, with and without Caldera
Everything Caldera does can be done without it. Know the offline way for every agent before it is needed.
| Agent | Pause | Resume |
|---|---|---|
| Any runtime agent | python3 caldera.pyz pause --dir <dir> |
python3 caldera.pyz resume --dir <dir> |
| An agent in adapter mode | the same commands; they call the adapter | the same |
| An agent with an own reporter | the agent's own channel (a chat command, a file) | the same |
- A pause set in Caldera is the same pause: the same file and the same stop.
- For a runtime agent, pause stops the agent's process; the runtime keeps reporting, and the state shows "Paused" (warn).
- Caldera disables the controls for a silent agent ("The agent is not reachable, so it cannot take a command now").
- A runtime on a remote machine is paused by logging in to it and running the command there. Write the exact command down next to the agent's documentation.
Update the runtime itself
The runtime updates itself on a press, signed (how it works is in the runtime):
- On the agent's page, the Runtime panel shows the version the agent runs, the
one this Caldera serves, a pin and the file kept before. "Update runtime" (admins
only; members read) sets a wish. The next answer to a report carries it. The panel
follows it: waiting for the agent's report, "has to prove itself first" while the
new runtime runs its five minutes, then the runtime's own words for taken, gone
back or not taken. A version that went back or failed is not offered again
until Caldera serves another one, and the button is gone until then. A runtime
that reports no version (older than this feature, or not
caldera.pyzat all) gets a notice instead of a button: it is replaced by hand once, below. - The runtime downloads
/cli/caldera.pyzand its signature, checks the checksum and the signature against the public key built into it, and lets the new file say its version. A bad or missing signature, or a machine withoutssh-keygen: nothing changes and the agent page shows the outcomefailedwith the reason. - It keeps the old file as
<dir>/.caldera/caldera.pyz.previous, swaps the file the unit names, and restarts in place: same process, the agent's own processes never stop. - The first successful report within five minutes makes it good. A new runtime that crashes, or does not report in time, goes back by itself; that version is then held until a newer one is offered.
On the machine, with no connection at all:
python3 <dir>/.caldera/caldera.pyz rollback-runtime --dir <dir> # back to the previous file
python3 <dir>/.caldera/caldera.pyz pin-runtime 0.32.0 --dir <dir> # never update to another version
python3 <dir>/.caldera/caldera.pyz pin-runtime off --dir <dir> # release the pin
One-time catch: a runtime installed before this existed cannot update itself.
Replace <dir>/.caldera/caldera.pyz once by hand (mode 700, checksum verified) and
run systemctl --user restart caldera-<name>.service; the unit has
KillMode=process, so the agent keeps running. After that, updates are presses. A
runtime started from anywhere but <dir>/.caldera/caldera.pyz (the unit's path)
cannot replace itself and says so.
Which Claude account an agent runs under
The Runtime panel on the agent's page says which Claude account the machine is signed in to. See The Claude account hint.
Rotate an agent's token
Do this when a token may have leaked, or was lost.
- Caldera: agent menu, "Rotate token...". The old token stops working at once. The new one is shown once, with the two environment lines.
- The agent is now silent in Caldera until the machine has the new token.
- Runtime agent: edit
<dir>/.caldera/agent.env, replace the value ofCALDERA_TOKEN(keep the file mode 600), thensystemctl --user restart caldera-<name>.service. The agent keeps running: it never held the token, and a restarted runtime adopts it. - Own reporter: update its secrets file and restart the reporter service.
- Check that the agent reports again.
The runtime has no command that fetches the new token; Caldera never writes to the machine.
Binding tokens (access to other products) are not rotated here. They are the bot's tokens, handled in the product's own bot settings; see the table below ("binding is broken").
Delete an agent
Caldera: agent menu, "Delete...". The confirmation names what stays.
- Caldera deletes the agent, its runs and its commands. Its token stops working at once.
- Nothing on the machine changes. The runtime keeps running the agent and logs a
refused report (
HTTP 401) every interval. - The bots are not revoked. Only the product that owns a bot can revoke it. The confirmation lists them by name ("These bots stay: ..."). Revoke them in that product's bot settings, or the tokens stay valid until they expire.
- To remove the agent from the machine, in this order:
python3 caldera.pyz pause --dir <dir>(this stops the agent's process; disabling the unit first would leave it running),systemctl --user disable --now caldera-<name>.service,- remove
~/.config/systemd/user/caldera-<name>.serviceandsystemctl --user daemon-reload, - remove
<dir>(it holds the agent's state; keep it if it is wanted).
A blueprint can only be deleted when no agent is installed from it.
Move an existing agent onto the runtime
It works when the agent can run as one process the runtime may start and stop. (An agent with its own timers and a pause that other things set uses adapter mode, or stays on its own reporter.)
Before: the agent has an own reporter and its own timers. Nothing changes until the steps start.
- Build the bundle and upload it as the first version of a blueprint (see Blueprints); publish.
- Rename the old agent in Caldera (for example
<name>-old): the install needs the name to be free. - Issue an install code for the blueprint, with the agent's name.
- On the machine, stop the old mode first, so nothing runs twice: disable the old timers and the old reporter service.
- Run the two install commands;
initwrites the directory and the unit. - Check:
status, the unit's journal (journalctl --user -u caldera-<name>.service -n 20), the agent in Caldera, "Run now" once, and the independence check. - When it runs, delete the old agent in Caldera. Leave the old unit files in place: the way back needs them.
The way back: disable the new unit and start the old mode again.
systemctl --user disable --now caldera-<name>.service
systemctl --user enable --now <the old timers and the old reporter service>
In Caldera rename the old agent back and delete the new one. The old and the new mode must never run at the same time.
The quota gauge looks wrong
Symptom. An agent's quota gauge in Caldera does not match the provider's usage page. For example Caldera shows 79 % for the week while the provider shows 90 %.
Cause. This is a property of how the agent measures, not of Caldera. A gauge that an agent builds from the quota report of its own last model call is only as fresh as that call. Two things make it stale:
- Other sessions use the same quota. The subscription is shared with interactive sessions on the same account, so the real figure can be higher than the last one the agent saw.
- The agent holds runs above its threshold. Above its own limits the agent starts no run, and without a run there is no new observation. The gauge then stays where it was.
A window whose reset time has passed counts as empty (0), so a stale gauge can also fall to 0 at a reset the agent has not seen a call after.
What helps. An agent that builds a gauge from its own last observation should
- put when it observed into the gauge: the first words of the gauge's
detailare the time of the last measurement (for example "As of 06.10. 14:05"), so that a stale value is recognisable as stale; - refresh the observation by itself, at most once an hour, also while it is paused or held, with a request that asks for nothing but the quota (no context, no tools, the smallest model); and say in its log when that fails and that the last value stays.
There is no central quota gate in Caldera and none is planned: every agent holds its own limits.
Check. Read the time in the gauge's detail line. If it is older than about an hour, the measurement fails: look in the agent's own log, and check that the program that makes the measurement can be found. Compare the figure with the provider's usage page.
Fix. Caldera shows what the agent reported; there is nothing to change in Caldera. Do not treat a gauge without a measurement time as live.
Troubleshooting
Start with the agent's page (level, last report, "Process" and "Binding" health
items) and, on the machine, python3 caldera.pyz status --dir <dir>.
| Symptom | Likely cause | Check | Fix |
|---|---|---|---|
| Agent shown as not reachable (silent, level fault) | No report for three intervals (at least two minutes): machine off, runtime stopped, no network, token rotated or agent deleted, workspace deactivated | status: runtime, last report ... ago; journalctl --user -u caldera-<name>.service; <dir>/.caldera/runtime.log |
Start the unit; or fix the cause in the rows below |
Log says report failed: HTTP 401 (unauthorized) |
The token was rotated or the agent was deleted | The agent's menu in Caldera; does the agent still exist? | Rotate the token and update the machine (above); or reinstall |
HTTP 403 |
The workspace is not activated for Caldera (or lost it) | The workspace's state in the app | Ask for the workspace to be activated |
HTTP 404 on /api/agent/v1/report |
The URL is not Caldera's own address (another product's address answers 404) | curl -i https://app.calderaapp.io/api/agent/v1/report should answer 401, not 404 |
Use Caldera's own address as CALDERA_URL |
HTTP 400 on a report |
The body is not valid JSON, not an object, or names another protocol |
An own reporter: print the response body | Send "protocol": 1 and a JSON object |
HTTP 413 |
The report is over 256 KB (a log over 512 KB) | Size of the body | Send less; logs: only the tail |
HTTP 429 |
More than 60 reports or 10 logs a minute from one agent (30 bundle fetches; 20 install calls per address) | Interval in the reporter's loop | Keep to the interval the answer asks for |
caldera.pyz download: checksum does not match and the file is HTML |
A proxy or a wrong address served a web page instead of the file | head -c 200 caldera.pyz shows HTML |
Download from https://app.calderaapp.io/cli/caldera.pyz itself, without a proxy that rewrites |
/cli/caldera.pyz answers 404 |
The address is not Caldera's own, or this Caldera serves no runtime build, or a proxy in between changes the Host |
curl -i https://app.calderaapp.io/cli/caldera.pyz.sha256 |
Use Caldera's own address; forward Host unchanged |
sha256sum: command not found |
macOS | Use the macOS variant in the dialog (shasum -a 256 -c -) |
|
init: unknown, expired or used code (404) |
All three look the same on purpose | The code's age (15 minutes); was it run before? | Issue a new code |
init: ... already holds an agent |
The target directory has .caldera/config.json |
Choose another --dir, or remove the old one |
|
init: a secret is needed |
No terminal and no CALDERA_SECRET_<NAME> |
Set the variable or run in a terminal | |
init fails after the code was spent |
Unpacking or writing failed | The message says the agent exists in Caldera without a machine | Delete the agent in Caldera, issue a new code |
| Dialog refuses with a refusal reason | See the refusal reasons | Mend the named cause; the code stays usable if redemption refused | |
| Agent shows "Not running" (fault) | The agent process keeps exiting | <dir>/.caldera/agent.out.log; status shows process not running; the report says "restarting in N s" |
Fix the agent. Restarts back off from 5 s to 300 s; "Run now" skips the wait |
| The service stops when you log out | systemd user linger is off | loginctl show-user <user> |
loginctl enable-linger <user> |
| A command stays "waiting for the agent's next report" | Normal for up to one interval (30 s; 10 s while the page is open). If it lasts: the agent is silent | Last report time; the reporting interval | Wait; if silent, see the first row. A command the agent never answers fails after ten minutes, at its next report |
| "Moving to X: waiting for the agent's next report" stays | The agent is silent; or the runtime holds that version (it rolled back from or refused it) and will not apply it; or the wish names a blueprint that is not the agent's own | status: a held line; last update line |
Silent: see the first row. Held: name a different version, or publish a new version number |
| Update ended "rolled back" at once | The new version's health command failed | last update ... in status; run the health command by hand in <dir>/versions/<version> with the agent's environment |
Fix the health check or the version; publish a new number |
| Update ended "failed: could not fetch" | Caldera was unreachable during the fetch | It retries after ten minutes by itself | |
| Update ended "failed: refused X: ..." | The bundle was invalid or did not match its checksum | The message names the reason (path, size, checksum, missing file) | Fix the archive; publish a new number |
| The agent rolled back on its own after being fine | Its level got worse within 30 minutes of the update | status, last update message |
Investigate the version; publish a fixed one |
Runtime update ended "failed" with a signature or ssh-keygen message |
The signature is missing or does not verify, or the machine has no ssh-keygen |
The message on the Runtime panel | Install OpenSSH on the machine; for a missing signature, wait for the next release |
| The Runtime panel offers no "Update runtime" | The version went back or failed and is held; the runtime is pinned; it is already at the served version; or it reports no version | The notice on the panel | Wait for the next release; release the pin with pin-runtime off; or replace an old runtime by hand once |
| A setting is refused ("Nothing changed: ...") | The value is outside the manifest's bounds, or the key is not declared | The message names the label and the rule | Use an allowed value |
| A log answers "There is no such log." | The name is not declared in the manifest's logs, or the file does not exist or resolves outside the agent's directory |
The manifest's logs; the file under the version directory |
Declare it, or fix the path |
| The log dialog says the agent did not answer | The agent is silent, or busy beyond ten minutes | Try again when it reports | |
| A binding is broken (agent page, "Access to other products") | Binding <product> is bad: the bot or its token was deleted or revoked, or the token expired (they last a year) |
The product's bot settings; the health item's detail | In that product's settings create a new token for the bot (or a new bot); edit CALDERA_BINDING_<PRODUCT>_TOKEN in <dir>/.caldera/agent.env; pause then resume. Caldera cannot mend it. |
Binding shows warn |
The product could not be reached from the machine | Usually transient | |
| The overview is empty for an admin | Wrong workspace; or the workspace is not activated ("not set up for Caldera yet") | The workspace switcher | Switch; or ask for activation |
| A member cannot steer | Members only see | Role | An admin does it |
| The quota gauge is wrong | The agent's last measurement is old (read the time) | See the quota gauge | See there |
Agent shows a commit-like version (1c1f3c5) and no Blueprint panel |
It reports with its own reporter (registered by hand) | Normal; updates are the agent's own | |
python3: not found or a syntax error |
Python is missing or very old | python3 --version |
Use Python 3.10 or newer (the lowest tested) |
Refusal reasons
Caldera's refusals carry a reason (and params for the sentence); the app writes
the sentence in the reader's language. The server's own message is English prose
for logs.
| Reason | Means | Mend |
|---|---|---|
product_not_activated |
The workspace is not activated for Caldera | Ask for activation |
role_cannot_view |
A guest cannot see this | Use a member or admin |
blueprint_name_taken |
A blueprint of that name exists | Choose another name |
blueprint_has_agents |
Agents are installed from it | Delete them first |
version_published |
That version is published and cannot change | Upload a new version number |
version_upload_race |
The same version was uploaded at the same moment | Try again |
version_not_deletable |
A published version cannot be deleted | None; publish a new one |
version_unpublished |
Publish before installing | Publish |
manifest_name_mismatch |
The manifest's name is not the blueprint's |
Make them equal |
upload_file_missing, archive_empty, archive_not_zip, archive_damaged, archive_entry_damaged, archive_unsupported |
The upload is not a usable zip (no file, empty, not a zip, damaged, or zip64, several disks, encryption, unknown compression) | Repack |
archive_too_large (20 MB), archive_too_many_entries (5000), archive_unpacks_too_large (100 MB), archive_entry_too_large (50 MB) |
A size limit | Make the bundle smaller |
archive_unsafe_path, archive_bad_entry |
A path that is not allowed, or an entry that is a link, device, duplicate, or file-that-is-also-a-directory | Remove it |
archive_no_manifest |
No caldera.agent.yaml at the top level |
Put it at the root |
archive_file_missing |
A file the manifest names is missing | Add it |
manifest_too_large, manifest_not_yaml, manifest_invalid |
The manifest is over 256 KB, not valid YAML, or invalid at the named path | Fix the field |
agent_name_taken |
The workspace has an agent of that name | Choose another name |
binding_product_not_activated |
A bound product is not activated for the workspace | Activate it there |
binding_product_no_address |
This Caldera has no address for that product | Not something a workspace can change; ask for help |
bot_role_mismatch |
A bot of that name exists with another role | Rename the agent, or change the bot |
role_cannot_lend |
The issuer's role may not create a bot with that role | Issue as an admin or owner |
bot_limit |
The workspace would hold more than ten bots | Remove bots |
issuer_cannot_install |
The person who issued the code may no longer install agents | Issue a new code as an admin |
instance_no_caldera_address |
This Caldera cannot build the commands | Not something a workspace can change; ask for help |
agent_not_from_blueprint |
The agent was registered by hand and has no blueprint | Not applicable |
version_not_installable |
Not a published version of the agent's blueprint | Publish it, or name another |
runtime_version_held |
The wished runtime went back or failed and is held | Wait until Caldera serves another version |
runtime_too_old |
The runtime reports no version | Replace it by hand once |
runtime_not_served |
This Caldera serves no runtime build | Not something a workspace can change; ask for help |
runtime_up_to_date |
The runtime already runs the served version | None |
runtime_pinned |
The runtime is pinned to another version | pin-runtime off on the machine |