Version 2 of 2 · · razie via designer-95-demigod-2designer · Current version

name
d3-f-boxes
description
d2spec's own boxes — cloud4 (production) and cloud5 (the replica): their tokens, the admin API, the release order, seeds and demos, and what an ops run may do on them. Extends d3-f-ops. Get skills only from /api/v2/skills/d3.
feature
ops
concepts
[box, cert, adminApi, seed, opsAccess]
tags
#skill #feature #d3skill
project
d2spec

d3-f-boxes

Intro

d2's own boxes: cloud4 (production) and cloud5 (the replica).

d2 itself runs on two boxes: cloud4 is production, cloud5 its replica for integration runs. This skill is the facts and calls for those two boxes; the rules of deploying, incidents and housekeeping are d3-f-ops, which this fills in. Spec: Spec › system-status (^ops-n).

Essentials

  • R-box-1 NEVER reach cloud5 through a production hostname, not even with --connect-to: the sandbox proxy routes by hostname, so it lands on cloud4.
  • R-box-2 The ops token (d2-token) is for the box and the d2* test projects only — NEVER on base d2 without the person's OK.
  • R-box-3 The admin token only deploys (d2-admin-token; cloud5 has its own, d2-admin-token5); both come from the vault and are never printed.
  • R-box-4 Only HTTPS 443 reaches a box from a cloud sandbox — no SSH from there.
  • R-box-5 NEVER edit /etc/d2/env by hand: d2-env set on a listed key only.

Concepts

box

One server running d2, or an older one ops looks after. Also called: prod (cloud4), the replica (cloud5), d1 (the old dieselapps box), the server. The boxes:

  • cloud4 — production: aiputty.com (primary, login hub) and *.aiputty.com; its earlier names, aidieselapps.com and d2.dieselapps.com redirect to it; cdn.aidieselapps.com serves the CDN.
  • cloud5 — the replica, dieselwebos.com (+ *.): integration and MongoDB runs at a release, nothing else.
  • d1 — the old diesel server, not d2: the app runs on cloud3a.razie.com (~/rk/dist, logs in ~/rk/logs); an Apache on another box fronts it as www.dieselapps.com. Ops only cleans its logs and restarts it on the Architect's order (#opsAccess_d1); the Apache box is not ops's. Inside: Ubuntu 24.04; Caddy terminates TLS and proxies /admin/* to the admin agent (127.0.0.1:8081, systemd d2-admin) and the rest to d2 (127.0.0.1:8080). Releases in /opt/d2/releases/<timestamp>, /opt/d2/current a symlink, /opt/d2/last-good the last healthy one. /etc/d2/env holds D2_ADMIN_TOKEN, D2_TOKEN, D2_DOMAINS (comma list, first = primary and SSO hub), D2_CDN_DIR (/opt/d2/data/cdn), D2_LOG_DAYS, and D2_HEAP_MB (the node heap limit, default 1024, written by bootstrap.sh on every run; a deploy refuses a box without it). MongoDB (user d2, with clusterMonitor so System Status can read its numbers) is the default inventory; daily provider backups; out-of-memory guards in box/guards.sh. Audit archive: /opt/d2/data/audit-archive/YYYY/MM/.

box_which

  • Summary: cloud4 is production, cloud5 the replica — by its own hostname only.
  • When: before any call to a box.
  • Needs: —
  • Call: GET https://aiputty.com/api/v2/health (cloud4) · GET https://dieselwebos.com/api/v2/health (cloud5).
  • Rules: R-box-1, R-box-4. Before treating cloud5 as a real replica, check what on it touches the outside world (crons, mail, CDN sync) and say so.
  • Errors: —
  • Gotchas: Why (R-box-1): a coder once deployed to production that way.

box_releaseOrder

  • Summary: full suite → cloud5 full-setup check → work into main → one deploy on cloud4 → box suite → visible-page check (VIS-1) → seed refresh.
  • When: a release, on the person's word (d3-f-git › release_merge, d3-f-pipeline › release_run).
  • Needs: the release item; nobody merging or stuck.
  • Call: test:full → on cloud5 the integration suite and the tests against real MongoDB → razwip → main → adminApi_deploy on cloud4 → the box suite (D2_URL=https://d2conf.aiputty.com, the ops token) → the visible-page check, VIS-1 (d3-f-tests › rules R-tests-9) → pull base d2's published shipped topics into seed-d2/ in one commit.
  • Rules: R-box-2. Only on the Architect's run the full test and deploy (or an open release item): work collects on razwip until then. One deploy per run, the release — no preview first; a preview is only for work that waits on him before it's checkpointed. The full-setup check is cloud5's: builders never need a MongoDB of their own, and nobody touches the laptop's. Then seed_pull. The report carries the version, the box suite's pass / fail, VIS-1's result with its screenshots, and the timing (deploy minutes included). VIS-1 failing: roll back only if the release caused it; if the cause is live content (a topic), don't roll back — say so on the release item with the cause and who fixes it (content → the designer), re-run VIS-1 once it's fixed, then close.
  • Errors: —
  • Gotchas: tag pushes may be refused by a sandbox proxy: say so on the release item rather than retrying. Why (VIS-1): 0.263.0 shipped a closed chat box that covered every page on phones (black in dark mode); every suite passed, because none looked at a page in a browser. Why (content isn't a rollback): 0.265.0's VIS-1 failed on a leftover script in base d2's landing topic; a rollback wouldn't have touched the topic and would have taken back a fix.

cert

A subdomain the box's TLS proxy was allowed a certificate for, counted against the weekly certificate budget. Fields: name, domain, when. Also called: a certificate, TLS cert.

adminApi

The box's admin agent, https://<box host>/admin/…, with Authorization: Bearer <that box's admin token>.

adminApi_deploy

  • Summary: POST /admin/deploy with a .tar.gz of the repo; unhealthy rolls back by itself.
  • When: the deploy step of a release (d3-f-ops › deploy_run).
  • Needs: the admin token for that box from the vault; the full suite at 0 fails.
  • Call: <implDir>scripts/deploy.sh from the repo's root (it sends git archive HEAD:<implDir> plus a deploy.json {commit, kind}, gzipped, with D2_ADMIN_TOKEN — vault name d2-admin-token — to D2_DEPLOY_URL); by hand: POST https://<box host>/admin/deploy (body: the archive) → unpack, npm ci, typecheck, swap current, restart, health check. cloud5: https://dieselwebos.com/admin/deploy.
  • Rules: R-box-3; d3-f-ops R-ops-1 to R-ops-3.
  • Errors: refused → didn't typecheck, or the box has no D2_HEAP_MB (E_NO_HEAP: run bootstrap.sh) · 401 → not the admin token for that box.
  • Gotchas: a deploy runs under the admin agent already installed on the box; a new admin agent in the release takes effect only when bootstrap.sh / guards.sh copy it.

adminApi_read

  • Summary: status, logs and releases of a box, read-only.
  • When: after a deploy that rolled back; an incident's first look.
  • Needs: the admin token for that box.
  • Call: GET /admin/status · GET /admin/logs?n=200 · GET /admin/releases.
  • Rules: R-box-3.
  • Errors: 401 → not that box's admin token.
  • Gotchas: —

adminApi_restart

  • Summary: restart or roll back over HTTPS — an incident fix, one of each per incident.
  • When: an incident run, after the cause is written down (d3-f-ops › incident_run).
  • Needs: the admin token; what you saw, written in the history first.
  • Call: POST /admin/restart · POST /admin/rollback (to the last healthy release).
  • Rules: d3-f-ops R-ops-4, R-ops-5.
  • Errors: 401 → not that box's admin token.
  • Gotchas: —

seed

What a new database starts from. Also called: seed data, the demo data.

  • seed-d2/ applies only to a new database (a fresh install); live base d2 is the source after that, refreshed into the seed at each release; seed drift between releases is a watcher's check.
  • Demo projects reset on every deploy from seed-demo/<n>/: d2hero (bookstore), d2creator (shop), d2pro, d2master, and d2conf (the box suite's target). Test projects start with d2; only the Architect makes d2* realms. After a deploy nobody re-creates wiped demo data: the person uses Reset all.

seed_pull

  • Summary: at each release base d2's published shipped topics go into seed-d2/, one commit; live wins.
  • When: the last step of a release (#box_releaseOrder).
  • Needs: your own d2 token (no separate seed token: run it with D2_SEED_TOKEN=$D2_BUILDER); the release's checkout.
  • Call: D2_SEED_TOKEN=$D2_BUILDER scripts/seedpull.ts (it reads base d2's published topics) → one commit naming the topics it wrote.
  • Rules: Deploys never write base d2, and nobody edits seed-d2/ by hand: base content is edited live on base d2. A seed file changed in the repo since the last pull is a conflict — skipped until merged on base d2. The Architect's settings (Settings:Quotas, Settings:System, Settings:Permissions) are never pulled.
  • Errors: —
  • Gotchas: —

opsAccess

What an ops run may do on a box. Getting in: SSH as <d2-ops-user>@cloud4.razie.com / @cloud5.razie.com (and d1) with the one key for all boxes, both from the vault (d2-ops-user, d2-ops-ssh); the key written to a mode-600 temp file for the run and deleted after; from the runner that started you, never from a cloud sandbox (R-box-4). What sudo allows (anything else is refused by the box): systemctl restart|status d2-admin|caddy, systemctl status mongod, journalctl, free, df, du, bootstrap.sh, guards.sh, d2-env get|set KEY [VALUE] (listed keys only: D2_HEAP_MB, D2_LOG_DAYS), d2-prune [--dry-run].

opsAccess_diagnose

  • Summary: the fixed first look, written in the history before any change.
  • When: the start of an incident run.
  • Needs: SSH as d2ops; the admin token.
  • Call: /api/v2/health from outside → systemctl status d2-admin → journalctl -u d2-admin -n 200 → journalctl -k --since -2h | grep -iE "out of memory|oom" → free -m → df -h → /admin/logs?n=200.
  • Rules: d3-f-ops R-ops-4.
  • Errors: —
  • Gotchas: —

opsAccess_fix

  • Summary: unasked, one of each per incident: restart d2-admin, roll back, housekeep, set a listed key, re-run guards.
  • When: the cause is written down.
  • Needs: —
  • Call: sudo systemctl restart d2-admin Also: POST /admin/rollback · sudo d2-prune · sudo d2-env set D2_HEAP_MB <n> then restart · bootstrap.sh / guards.sh.
  • Rules: R-box-5; d3-f-ops R-ops-5, R-ops-6. Ask first (through whoever started you): restarting mongod or caddy, a reboot, packages or OS upgrades, any other key or file, anything touching data. Never: code, deploys, secrets, reading the database.
  • Errors: sudo refuses → not on the list: ask.
  • Gotchas: Why: 2026-10-05 a release and its rollback both ran out of heap on start; production was down about 22 minutes and the fix (a heap setting plus a restart) needed root (P-944).

opsAccess_d1

  • Summary: the old d1 server: clean its logs; restart it only on the Architect's order.
  • When: logs — in each housekeeping round and on a disk warning; restart — only when the Architect asks for it on his channel (the monitor relays the order).
  • Needs: SSH to cloud3a.razie.com as d2-ops-user with d2-ops-ssh; no sudo.
  • Call: logs: ls -la ~/rk/logs then delete the log files there older than 3 days (never the one being written) and say MB freed · restart: cd ~/rk/dist; . ./restart.sh, then GET https://www.dieselapps.com/ until it answers (2 minutes at most).
  • Rules: Nothing else on d1 — no other folder, no packages, no config; anything more is taught to you separately by the Architect.
  • Errors: restart doesn't answer in 2 minutes → tell the Architect through the monitor, with the script's output; don't retry.
  • Gotchas: —

opsAccess_housekeep

  • Summary: sudo d2-prune on cloud4 and cloud5, then one CDN sweep; one line per box on the Architect's channel.
  • When: daily at 04:00 Toronto, and on a disk warning (d3-f-ops › housekeeping_run).
  • Needs: —
  • Call: sudo d2-prune --dry-run → sudo d2-prune (each box) → POST /api/v2/cdn/sweep once → d1's logs (#opsAccess_d1).
  • Rules: It removes: releases beyond current, last healthy and two more (and their .tgz); journald past 14 days; npm and apt caches; d2's /tmp files over 2 days; CDN <role>/tmp/ over 7 days and <role>/mockups/ over 30 days. It never touches /opt/d2/data, Mongo, /etc, or d2's own log and audit-archive folders (d2 rotates those).
  • Errors: —
  • Gotchas: —

Errors

  • E_NO_HEAP — /admin/deploy refused: no D2_HEAP_MB on the box → run bootstrap.sh; refused without it → it didn't typecheck. see #adminApi_deploy
  • 401 on /admin/* — not the admin token for that box (cloud5 has its own). see #adminApi_read
  • E_SCOPE with the ops token on base d2 — expected → ask the person. see #box_which
  • sudo refuses a command — not on the list → ask through whoever started you. see #opsAccess_fix