Connecting an agent
There are two ways to get an agent into Caldera. Both end the same: the agent is an entry in Caldera, it reports in the agent protocol, and an admin sees and steers it. They differ in who runs the reporting and who owns the agent's process.
The two ways
| A. Own reporter | B. Blueprint and runtime | |
|---|---|---|
| What reports | A few lines you write and run yourself, as a separate process. | The runtime caldera.pyz, which Caldera serves. |
| Who starts and stops the agent | You (systemd, cron, your own code). | The runtime (start, stop, restart with backoff). |
| How the agent is set up | By hand: checkout, secrets, unit files. | caldera init <code> from a published blueprint. |
| Pause means | Whatever your agent does with it (a file, a flag, a unit stopped). | The runtime stops the agent process; resume starts it. |
| Settings | You declare a schema and enforce the bounds in your code. | The manifest declares schema, defaults and bounds; the runtime enforces them. |
| Updates | You deploy yourself. | Caldera names a version; the runtime fetches, health-checks, keeps or rolls back. |
| Rollback | Yours. | caldera rollback on the machine, or Caldera names the old version. |
| Needs Python 3 | No. Any language that can POST JSON. | Yes, on the agent's machine. |
| Needs Caldera to set up | Only to create the agent and its token. | Yes, for init (the code). Never to run. |
| Best for | An agent that already exists and has its own scheduler, or one with no process of its own. | A new agent, or an agent you can restructure as one long-running process. |
How to choose
- Choose B when the agent is new, or when it is one program that can run as a long-lived process. You get set-up from one command, updates with a health check and rollback, the settings form from the manifest, and bots for access to the suite's other products.
- Choose A when the agent has its own life that Caldera should not own. The process runtime launches one process and pauses by stopping it. If the agent has no such process (for example its runs are a timer and its pause is a file that a chat bot also sets), the process runtime does not fit. Either keep the own reporter, or use adapter mode, in which the runtime starts and stops nothing and calls one command the agent provides.
- You can start with A and move to B later. The agent keeps its history in Caldera only if you keep the same agent entry; a blueprint install always creates a new agent. The move is described in Operating agents, with the way back.
Whichever you choose, the fixed rule holds: the agent must not need Caldera to run (see Concepts).
Way A: an own reporter
Steps
-
Create the agent. An admin opens the overview and presses "New agent", gives it a name, and gets the token once (
caldera_agt_...) with two ready-made lines:CALDERA_URL=https://app.calderaapp.io CALDERA_TOKEN=caldera_agt_...Only a hash of the token is stored. If it is lost, rotate it from the agent's menu ("Rotate token"): the old token stops working at once.
-
Put the two lines where the reporter reads them, in a private file on the agent's machine (mode 600). Never in a repository.
-
Write the reporter against the agent protocol. The rules that make an agent independent of Caldera:
- a separate process (a systemd user service or a sidecar);
- a timeout on every request;
- when Caldera is unreachable, slow or answers nonsense: log a line, wait for the next interval, try again, change nothing else;
- pause, limits and settings are the agent's own files;
- a command is a wish: check it against your own bounds, apply or refuse, and answer with an ack in the next report;
- send what the page should show, nothing more. Mail bodies, messages and similar content stay with the agent. Logs are the one exception, and the agent decides which it hands out.
-
Run it and check the page. The agent appears in the overview after the first report.
-
Do the independence check.
A minimal reporter
This is the smallest loop that follows the rules. Replace handle with your own
checks. It uses only the Python standard library; any language works.
import json
import os
import time
import urllib.error
import urllib.request
URL = os.environ["CALDERA_URL"].rstrip("/")
TOKEN = os.environ["CALDERA_TOKEN"]
def post(path, body):
request = urllib.request.Request(
URL + "/api/agent/v1" + path,
data=json.dumps(body).encode(),
method="POST",
headers={"Authorization": "Bearer " + TOKEN, "Content-Type": "application/json"},
)
with urllib.request.urlopen(request, timeout=30) as response: # always a timeout
return json.load(response)
def handle(command):
"""Return (ok, message), or None for a `log` command (the upload is the answer)."""
if command["type"] == "log":
# send the log with post("/log", {"command": command["id"], "name": ..., "text": ...})
return None
return False, "This agent does not support that." # a refusal is a normal answer
acks, interval = [], 30
while True:
report = {
"protocol": 1,
"state": {"level": "ok", "label": "Ready"},
"controls": ["pause", "resume"],
"acks": acks,
}
try:
answer = post("/report", report)
acks = [] # delivered
interval = max(10, min(300, int(answer.get("interval", 30))))
for command in answer.get("commands", []):
result = handle(command)
if result is not None:
acks.append({"id": command["id"], "ok": result[0], "message": result[1]})
except (urllib.error.URLError, OSError, ValueError) as error:
print("caldera unreachable:", error) # one line, then the next interval
time.sleep(interval if not acks else 2)
What the example shows, and what matters:
acksare cleared only after a report was delivered, so an answer is not lost to one failed request.- A command is delivered at most once. If the answer to a report is lost, the
command is lost with it, and Caldera fails it after ten minutes. That is
deliberate:
run_nowmust never run twice because a network dropped a reply. A person can send it again. - A
nullin a report counts as "absent", except insettings.values, wherenullmeans "this setting is not set". It is better to leave an unset field out. - Shares (gauges, subject shares) are numbers from 0 to 1. Times are Unix seconds.
A pattern: a timer-driven agent with a file as its pause
An agent whose runs are a timer, and whose pause is a file that other things
(for example a chat bot) also set and clear, is the typical case for an own
reporter. The reporter is its own service next to the timer. Caldera's pause
creates the same file the chat command creates, so a pause set in Caldera can be
lifted without Caldera. resume deletes the file and says if another hold (a
person working in a session, a build in progress) still keeps the agent waiting.
run_now creates a file the timer's scheduler looks for. Settings are a fixed
list of fields, each with hard bounds, written to the agent's own configuration;
anything else is refused, and a set command is all or nothing. Logs are limited
to names the agent itself knows, and only the last 512 KB are sent.
If such an agent should also get set-up and updates from Caldera, use adapter mode.
Gauges from the agent's own last observation
A gauge built from the agent's own last observation (for example the quota left
after its last model call) can lag the provider's page, because other sessions use
the same quota and an agent that holds its runs above a threshold makes no new
observation. Put the time of the measurement into the gauge's detail (for
example "As of 06.10. 14:05, ...") so that a stale figure can be recognised as
stale. See Operating agents.
Way B: a blueprint under the runtime
Steps
- Write the blueprint. A manifest
caldera.agent.yaml, the agent's files, and a health command. See Blueprints and The manifest. - Upload and publish a version in the workspace (Blueprints, the blueprint, "Upload a version", then "Publish"). A published version never changes.
- Issue an install code ("Install"): choose the version and the agent name. Caldera shows two commands.
- On the machine: run the two commands: fetch and verify
caldera.pyz, thenpython3 caldera.pyz init <code> --url https://app.calderaapp.io. The tool shows what will be set up, asks for the declared secrets on the machine (they never reach Caldera) and only then redeems the code. It writes the directory, installs a systemd user unit and starts it. - Check:
python3 caldera.pyz status --dir <dir>and the agent's page. - Do the independence check.
The details of each step are in the runtime and Blueprints.
A pattern: one long-running process
An agent that fits the process runtime is one program with its own scheduling,
for example one process with one thread per job. It writes a status file
atomically every 30 seconds and on every change, with a heartbeat that its
health command reads. Pause is the runtime stopping the process; run now is a
file the agent checks every second. Its secrets and data stay where they were,
outside the bundle, and the manifest declares only the paths that must survive
an update. Offline pause is caldera pause on the machine.
The independence check
Part of connecting every agent. Write the result down in the agent's own repository, because that is where the reporter lives.
Way A:
- Set
CALDERA_URLto a host that does not exist (for examplehttps://caldera.invalid) and restart the reporter service. - Start a run. It must run to the end; the log must only say Caldera is not reachable.
- Set a pause in Caldera first, then make Caldera unreachable as in step 1. The agent's own channel must lift the pause.
- Restore
CALDERA_URL, restart; the next report arrives within one interval.
What must never happen: a run waits for Caldera, a pause can only be lifted from Caldera, or an error in the reporter stops the agent's scheduler.
Way B (runtime):
- Point the runtime's URL at nothing, or block the network. The agent process
must keep running;
status,pause,resumeandrollbackmust work. python3 caldera.pyz pause --dir <dir>thenresumemust work offline.- Stopping the runtime must not stop the agent.