CrashPilotX
Sign in

Install CrashPilotX

Start with Ubuntu Linux or Ubuntu on WSL. We are keeping the first platform solid before adding more.

Ubuntu 22.04 and 24.04 on Linux machines

Part 1 - Install the agent

1

Install the agent

curl -fsSL https://crashpilotx.com/install.sh \
  | sudo bash

The installer supports Ubuntu only for now. It installs the Python agent, systemd services, heartbeat and hourly verified update-check timers, and /etc/crashpilot/.env.

2

Verify and run an analysis

sudo crashpilot doctor
sudo crashpilot analyze

Connect the machine in Part 2 before expecting live dashboard heartbeats.

Part 2 - Connect to the dashboard

Outbound-only push mode

1

Create a system on the dashboard

Sign in, open Systems, click Add system, and create a push-mode system.

Go to Systems
2

Connect the machine

# New install or update existing install and connect in one command:
curl -fsSL https://crashpilotx.com/install.sh \
  | sudo bash -s -- --connect cpilot_<your-connection-string>

# Already installed? Connect only:
sudo crashpilot configure cpilot_<your-connection-string>
sudo crashpilot heartbeat

Copy the real command from the dashboard. Heartbeats send CPU, memory, disk, GPU, dmesg, and agent-health data outbound. No inbound port or public URL is needed.

Automated enrollment

For fleets that add and remove machines on their own

1

Create a join token

On the Systems page, create a join token. Mark it ephemeral for autoscaling groups, spot instances, CI runners, and Kubernetes nodes, so machines that go away are retired automatically. Use one token per Kubernetes cluster.

2

Enroll at first boot with cloud-init

#cloud-config
runcmd:
  - curl -fsSL https://crashpilotx.com/install.sh | CRASHPILOT_ENROLL='cpjoin_...' bash
3

Or enroll every Kubernetes node

kubectl create namespace crashpilot
kubectl -n crashpilot create secret generic crashpilot-agent \
  --from-literal=CRASHPILOT_ENROLL_TOKEN='cpjoin_...'
kubectl apply -f https://raw.githubusercontent.com/CrashPilotX/crashpilot-agent/main/k8s/daemonset.yaml

Create the Secret before the DaemonSet, because pods read it only when they start. Each node enrolls under its own node name and signs off when its pod stops.

Automation API

Add and retire systems from Terraform, lifecycle hooks, or scripts

1

Create an API key

On the Systems page, create an API key. It is shown once, with your endpoint and public key filled in:

export CRASHPILOT_URL='https://<project>.supabase.co'
export CRASHPILOT_ANON_KEY='<public anon key>'
export CRASHPILOT_API_KEY='cpk_...'

When you create a key, choose what it may do: add and retire systems, read fleet health, or both. A key made only to read health, for a status page or a monitoring check, cannot change anything. A key that adds systems receives the agent credentials of every system it adds, which is enough to report as that system, so keep it in a secret store. No key can read reports, alert text, or settings. What a key may do is fixed when it is created; to change it, revoke the key and create another. Revoking stops it at once.

2

Add a system

jq -n \
  '{p_api_key: env.CRASHPILOT_API_KEY, p_name: "web-7", p_external_id: "i-0abc123", p_labels: ["web", "prod"]}' |
  curl --fail-with-body -sS "$CRASHPILOT_URL/rest/v1/rpc/api_create_system" \
    -H "apikey: $CRASHPILOT_ANON_KEY" -H "Content-Type: application/json" -d @- > system.json

The response carries the new system's credentials, with HTTP 201:

{
  "system_id": "6f1c0d9e-...",
  "agent_token": "...",
  "name": "web-7",
  "created": true,
  "restored": false
}

Only p_name is required. Pass p_external_id, such as an instance ID, to make the call safe to retry: the same identity always gets the same system back with fresh credentials and HTTP 200, and a retired one returns with "restored": true. The name and up to 12 labels apply only when the system is first created.

3

Connect the machine

CONN="cpilot_$(jq -c --arg url "$CRASHPILOT_URL" --arg key "$CRASHPILOT_ANON_KEY" \
  '{url: $url, key: $key, system_id: .system_id, token: .agent_token}' system.json | base64 -w0)"

# Hand $CONN to the machine, for example in its cloud-init user-data, and run
# there as root. Through the environment it stays out of process arguments.
curl -fsSL https://crashpilotx.com/install.sh | CRASHPILOT_CONNECT="$CONN" bash
4

Retire it when it goes away

jq -n '{p_api_key: env.CRASHPILOT_API_KEY, p_external_id: "i-0abc123"}' |
  curl --fail-with-body -sS "$CRASHPILOT_URL/rest/v1/rpc/api_retire_system" \
    -H "apikey: $CRASHPILOT_ANON_KEY" -H "Content-Type: application/json" -d @-

Pass p_external_id or p_system_id. It works like Retire on the Systems page: the system leaves the fleet, its credentials stop working, and its history is kept until retention ends. It also retires machines that enrolled themselves, by the identity they enrolled under, such as aws:i-0abc123 or k8s:node-name, so a termination hook can retire an instance the moment it goes. Retiring twice is not an error.

5

Read fleet health, for a status page or a monitoring check

jq -n '{p_api_key: env.CRASHPILOT_API_KEY}' |
  curl --fail-with-body -sS "$CRASHPILOT_URL/rest/v1/rpc/api_system_health" \
    -H "apikey: $CRASHPILOT_ANON_KEY" -H "Content-Type: application/json" -d @-

# One system, by the identity it enrolled or was added under:
jq -n '{p_api_key: env.CRASHPILOT_API_KEY, p_external_id: "i-0abc123"}' |
  curl --fail-with-body -sS "$CRASHPILOT_URL/rest/v1/rpc/api_system_health" \
    -H "apikey: $CRASHPILOT_ANON_KEY" -H "Content-Type: application/json" -d @-

It answers with the whole fleet, or the one system you asked for:

{
  "fleet": { "total": 3, "online": 2, "offline": 1, "signed_off": 0, "retired": 0 },
  "systems": [
    {
      "system_id": "6f1c0d9e-...",
      "name": "web-7",
      "external_id": "i-0abc123",
      "labels": ["web", "prod"],
      "ephemeral": false,
      "status": "online",
      "last_heartbeat": "2026-09-12T12:30:30Z",
      "seconds_since_heartbeat": 42,
      "hostname": "web-7",
      "agent_version": "0.2.0",
      "maintenance_until": null,
      "signed_off_at": null,
      "retired_at": null,
      "retired_reason": null,
      "open_alerts": { "warning": 1 },
      "last_crash_at": "2026-09-01T03:12:00Z"
    }
  ]
}

Only a key created with health reads may call this; any other gets 403 scope_required, just as a key made only to read health gets it from the add and retire calls. status is one of online, signed_off (shut down cleanly), offline and retired, worked out the same way the dashboard does. Retired systems are left out unless you pass p_include_retired or ask for one by name.

Errors

Every error is JSON with a stable code and a readable message. Branch on the code, not the status: a misspelled parameter name also returns 404, with the code PGRST202.

HTTPcodeMeaning
401invalid_api_keyThe key is wrong, revoked, or expired.
403system_capThe account already has 50 active systems. Retire one first.
404system_not_foundNo system with that system_id or external_id in this account.
409ambiguous_external_idSeveral active systems share the external_id. Retire by system_id.
422invalid_name, invalid_external_id, invalid_labels, invalid_targetThe request needs fixing; the message says how.
429rate_limitedMore than 60 calls in a minute with this key. Wait for Retry-After seconds.
403scope_requiredThe key was not made for this call: it cannot read health, or it can only read health. What a key may do is fixed when it is created.

Optional Ubuntu packages

>
smartmontools
SMART disk health
>
lm-sensors
CPU/GPU/board temperatures
>
rasdaemon
Machine Check Exceptions
>
nvidia-smi
NVIDIA GPU Xid errors

First value checks

After connecting, run these to verify setup, send a heartbeat, and produce a first report.

sudo crashpilot doctor
sudo crashpilot heartbeat
sudo crashpilot analyze --force

Remove CrashPilotX

Stop services, remove files, and delete the dashboard system

1

Stop CrashPilotX services

sudo systemctl disable --now \
  crashpilot-heartbeat.timer crashpilot-heartbeat.service \
  crashpilot-update.timer crashpilot-update.service \
  crashpilot-snapshot.timer crashpilot-snapshot.service 2>/dev/null || true
sudo systemctl disable --now crashpilot.service crashpilot-api@root.service \
  crashpilot-signoff.service 2>/dev/null || true
2

Remove installed files

sudo rm -f /usr/local/bin/crashpilot
sudo rm -f /etc/systemd/system/crashpilot.service
sudo rm -f /etc/systemd/system/crashpilot-api@.service
sudo rm -f /etc/systemd/system/crashpilot-signoff.service
sudo rm -f /etc/systemd/system/crashpilot-heartbeat.service
sudo rm -f /etc/systemd/system/crashpilot-heartbeat.timer
sudo rm -f /etc/systemd/system/crashpilot-update.service
sudo rm -f /etc/systemd/system/crashpilot-update.timer
sudo rm -f /etc/systemd/system/crashpilot-snapshot.service
sudo rm -f /etc/systemd/system/crashpilot-snapshot.timer
sudo systemctl daemon-reload
sudo rm -rf /opt/crashpilot /etc/crashpilot
3

Remove a user install

rm -f ~/.local/bin/crashpilot
rm -rf ~/.local/share/crashpilot ~/.config/crashpilot

In the dashboard, open Systems and delete the system record if you no longer want historical reports or heartbeats for that machine.