- name
- d3-f-ops
- description
- Ops — running a d2 installation: which version is live, System Status and self-healing, deploys on the person's word, incident and housekeeping runs. Get skills only from /api/v2/skills/d3.
- feature
- ops
- concepts
- [health, boxStatus, deploy, incident, housekeeping]
- tags
- #skill #feature #d3skill #wip
d3-f-ops¶
Intro¶
Running an installation: health, System Status, deploys, incidents, housekeeping.
Ops is keeping a d2 installation up: knowing which version is live, watching System Status, deploying only when the person says, and fixing or cleaning a box within fixed limits. This skill is what holds on any installation; the boxes themselves — hostnames, tokens, the admin API, what sudo allows — are a project's own (d3-f-boxes where it has one). A project's own version and release: d3-f-accounts › project_release; releasing code: d3-f-git › release_merge, d3-f-pipeline › release_run; the audit codes: d3-f-audit › auditCode. Spec: Spec › system-status (^sys-n, ^ops-n).
Essentials¶
- R-ops-1 NEVER deploy except on the person's word: a release item, the auto-release switch, or their say in a chat.
- R-ops-2 MUST check before and after a deploy: the suite's fail count is 0 before — or every failure is one your person waived by name for this release; the live version is the new one after.
- R-ops-3 One deploy per run.
- R-ops-4 MUST find the cause in the logs before changing anything on a box, and write what you saw before you change it.
- R-ops-5 One of each kind of fix per incident; still down → stop and ask whoever started you.
- R-ops-6 NEVER code, deploys, secrets or reading the database from an incident or housekeeping run; nothing touching data without asking first.
- R-ops-7 NEVER re-create wiped demo data after a deploy: the person resets demos themself.
- R-ops-8 NEVER print a box token or key (d3-f-vault › secret_read).
Concepts¶
health¶
What a running d2 says about itself. Also called: the version, "is it up", "what's live". Fields: version, commit, deployed time, database connected, heap (used, limit, whether the limit is set, start-up peak).
health_read¶
- Summary:
GET /api/v2/healthsays which version is live. - When: before and after a deploy; first step of any incident, from outside the box.
- Needs: —
- Call:
GET /api/v2/health→{version, commit, deployedAt, …}. - Rules: R-ops-2.
- Errors: no answer → the installation is down: an incident.
- Gotchas: —
boxStatus¶
System Status: a sample a minute (CPU, memory, disk per filesystem, load, d2's memory), 24 h kept. Also called: the box status, System Status page.
Levels: thresholds in base Settings:System (warn 80%, alert 90%; disk critical 95%). Going up a level audits S_DISK / S_MEM / S_CPU and notices the platform's owner.
Also on it: the node heap against its limit (S_HEAP, warn 70%, alert 85%; the start-up peak; red when the limit is the runtime's default), out-of-memory kills (S_OOM) and restarts, d2's memory against its service cap, memory pressure, the last deploy's result (ok / rolled-back / rollback-unhealthy), the database (cache against its cap, call p95, sizes and growth, connections), and how the server last started (S_DEPLOY, S_RESTART, S_CRASHED).
Self-healing (self-healing: on|off): a disk at critical deletes the detailed log's oldest day files (never today's, never in use, never the audit archive) until it is under its warn level; each clean-up audits H_DISK.
boxStatus_read¶
- Summary: the numbers-only status; say it once per change, never act on the box from here.
- When: every watch cycle; an incident's first look.
- Needs: a token allowed to read it (names are stripped for non-admins).
- Call:
GET /razadmin/status?format=json→{worst, service: {oomKills, …}, deploy: {result}, …}. - Rules: Tell your person once per change — when
worstgoes up towarningoralert, whenservice.oomKillsrose since the last read, or whendeploy.resultisrollback-unhealthy: metric, value, threshold. Back down is one ordinary line. Reading status never licenses a fix: that is an incident run. - Errors: 404 → not yours to read.
- Gotchas: —
deploy¶
Putting a release on a box. Also called: ship it, push it live, release (the code side: d3-f-git › release). What a deploy does: unpack, install, typecheck (a release that doesn't typecheck is refused), swap the current release, restart, health check — unhealthy rolls back by itself.
deploy_run¶
- Summary: only on the person's word; suite at 0 fails before, the new version live after; one per run.
- When: a release run reaches its deploy step.
- Needs: the person's word or an open release item; the box's deploy call and token (the project's boxes skill).
- Call: the box's deploy call →
health_read→ the box suite (d3-f-tests › suite_run). - Rules: R-ops-1, R-ops-2, R-ops-3, R-ops-7, R-ops-8.
- Errors: refused → the release didn't typecheck: fix and redeploy · rolled back → unhealthy after the swap: read the box's logs; there is nothing to undo.
- Gotchas: Why (R-ops-1): releasing after every batch cost a full checkpoint each — ten releases for about eighteen items in one day.
incident¶
A box that is down or failing. Also called: an outage, an alert, "prod is down".
incident_run¶
- Summary: health from outside, cause from the logs, the least fix, health twice a minute apart, an incident note.
- When: started with an alert as your prompt; one run per incident.
- Needs: the alert; the box's access and what you may do there (the project's boxes skill).
- Call:
health_read→ the box's status and logs → write what you saw in the project's history (d3-f-history › historyEntry_write) → the least fix allowed →health_readtwice, a minute apart → the incident note, a second history entry → done. - Rules: R-ops-4, R-ops-5, R-ops-6. Each fix is told on your person's channel as you do it. Still down after one of each: stop, post
waitingwith the evidence to whoever started you (d3-f-board › wait_ask). The incident note: cause; timeline in the person's time zone; every command you ran and what it printed (trimmed); what's left for a person. - Errors: a command the box refuses → it isn't on your list: ask, don't find another way.
- Gotchas: Why: a release and its rollback both ran out of memory on start; production was down about 22 minutes and the fix needed rights no agent had.
housekeeping¶
Routine cleaning of a box: old releases, logs, caches, temporary files. Also called: clean-up, pruning.
housekeeping_run¶
- Summary: dry run, then the clean-up, per box; report what was freed and the disk after.
- When: on its schedule, and on a disk warning.
- Needs: the box's clean-up command (the project's boxes skill).
- Call: the clean-up with its dry-run flag → the clean-up → check d2's own clean-ups ran (log days compressed and expired, the audit archive, cleanable CDN folders: d3-f-cdn › cleanup_check).
- Rules: R-ops-6. One history entry per round (d3-f-history › historyEntry_write). Report one line per box on your person's channel: removed, MB freed, disk before and after. Disk still at warning afterwards → say it as an alert. Never by hand: the database, kept CDN folders, the audit archive.
- Errors: —
- Gotchas: —
Errors¶
- deploy refused — the release didn't typecheck → fix and redeploy. see #deploy_run
- rolled back — unhealthy after the swap → read the box's logs; nothing to undo. see #deploy_run
- 401 on a box's admin call — not that box's admin token. see #deploy_run
E_SCOPEwith a box's ops token on a project it isn't for — expected → ask the person. see #deploy_run- 404 on System Status — not yours to read. see #boxStatus_read