deckhearth/.convoys/migrate-ci-to-self-hosted.md
varutasu e933140560
convoy: forbidden-pattern gate + AGENTS.md docs (briefs 3+4) (#133)
* convoy: forbidden-pattern gate + docs (briefs 3+4)

Closes out the migrate-ci-to-self-hosted convoy with the two
defensive follow-ups Brief 1+2 (PR #132) intentionally deferred.

Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps
`.github/workflows/` for `runs-on: ubuntu-latest` and fails unless
the match is in the documented allowlist (currently
`agent-context-drift.yml` only, per Decision D4). Self-tested
locally against the post-migration tree: 0 violations. Renames the
job from "Forbidden patterns (7 checks)" → "(8 checks)" and
normalizes the older "Check N/6" labels to "N/8" for consistency
(the inherited mix of `/6` and `/7` was a known cosmetic from the
unify-glass-panel-surfaces convoy).

Brief 4 — AGENTS.md § 6 and § 7 updates:
- § 6 "CI behavior": Playwright smoke runtime range updated to
  cover post-migration cold vs. warm cache (was a stale 59s figure
  from pre-migration ubuntu-latest).
- § 6 new top-level bullet "Self-hosted runner pool" alongside
  "CI minute optimizations" — covers where runners live, where
  caches are bind-mounted on CT 111, the Postgres rewire on
  CT 102, and the agent-context-drift.yml exemption + how Check 8
  enforces it.
- § 7 new bullet for the operational story: PAT rotation cadence
  + the D5 one-line `sed` revert path for when axiom is offline
  mid-PR-storm. Cross-references the axiom-server CT 111 README
  and the Beszel down alert.

Convoy doc — status flipped queued → shipped, shipped_in lists
both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron +
seed-visual-baselines-on-linux items machine-greppable.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop placeholder comment now PR #133 number is known

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:18:21 -05:00

11 KiB
Raw Blame History


slug: migrate-ci-to-self-hosted status: shipped opened: 2026-06-05 shipped: 2026-06-05 owner: rstillw shipped_in:

  • PR #132 (Briefs 1+2 — workflow migration + migrate-job rewire)
  • PR #133 (Briefs 3+4 — forbidden-pattern gate + AGENTS.md docs) follow_ups:
  • cleanup-stale-ci-runs-cron (weekly GC on CT 102 for ci_run_* DBs older than 7d; Risk #4 defensive)
  • seed-visual-baselines-on-linux (now easier with axiom; see queued-follow-ups below) prerequisites:
  • CT 111 (ci-runner) provisioned and online on axiom (192.168.68.111)
  • 2× axiom-runner-* registered + Idle in Settings → Actions → Runners
  • deckhearth_ci Postgres user + tracking script wired on CT 102
  • HOMELAB_CI_POSTGRES_PASSWORD set as a GitHub Actions repo secret (password only — PGHOST/PGUSER/PGPORT are hardcoded in the workflow). Earlier draft of this convoy used a single HOMELAB_CI_POSTGRES_BASE_URL URL secret, but during Brief 1+2 validation psql "$URL/postgres" failed with invalid option -- '/' — the URL-parse path was fragile. Split secret + standard PG* env vars sidesteps it.
  • Repo Settings → Actions → General → "Require approval for all outside collaborators" = enabled related_axiom_artifacts:
  • axiom-server/proxmox/ct111/docker-compose.yml
  • axiom-server/proxmox/ct111/.env.example
  • axiom-server/proxmox/ct111/README.md
  • axiom-server/.cursor/rules/ct111-ci-runner.mdc

migrate-ci-to-self-hosted

Problem

GitHub Actions billing on the free tier blocks CI on a private repo when the monthly minute budget runs out (hit during the unify-glass-panel-surfaces convoy, 2026-06-04; PRs #124 + #125 had to admin-merge without CI). The slash-ci-minutes convoy (PR #126) reduced consumption by ~60% via paths-ignore, grep consolidation, and caching, but a busy week of implementation work still trips the limit.

This convoy migrates all 4 GitHub Actions workflows off ubuntu-latest (GitHub-hosted, billed) onto the axiom homelab runner (CT 111, self-hosted, free). It also rewires the migrate job to use CT 102's shared Postgres instead of spinning up an ephemeral container per run — saving ~30s/PR and eliminating the Docker-in-runner pull cost.

Non-goals

  • Migrating to a different CI provider (CircleCI, Buildkite, etc.) — overkill.
  • Hosting the production app on axiom — Vercel keeps the deploy story simple and the homelab is already at ~95% RAM allocation. Out of scope.
  • Replacing the Vercel Preview deployments — Vercel still builds previews; Playwright smoke + visual-diff still run against those previews from the self-hosted runner.

Workflows to migrate

File Jobs Notes
.github/workflows/ci.yml lint, schema-map-fresh, forbidden-patterns, migrate, test migrate needs the Postgres rewire (see below)
.github/workflows/preview-smoke.yml gate, smoke smoke job needs Chromium — first run will npx playwright install and cache it
.github/workflows/visual-diff.yml (visual) same Chromium cache benefit; baselines still TBD
.github/workflows/pr-health-rollup.yml rollup trivial — single gh pr comment job
.github/workflows/agent-context-drift.yml (weekly cron) Optional: leave on ubuntu-latest so the cron runs even when axiom is down. Decision deferred — see Risk #3

Change shape per job:

# Before
jobs:
  lint:
    runs-on: ubuntu-latest

# After
jobs:
  lint:
    runs-on: [self-hosted, axiom]

Migrate-job Postgres rewire

Today (ci.yml § migrate):

services:
  postgres:
    image: postgres:16
    env: { POSTGRES_USER: postgres, POSTGRES_PASSWORD: postgres, POSTGRES_DB: deckhearth_test }
    ports: ["5432:5432"]
env:
  POSTGRES_URL: postgres://postgres:postgres@localhost:5432/deckhearth_test

Shipped (post-validation revision):

env:
  PGHOST: 192.168.68.102
  PGPORT: '5432'
  PGUSER: deckhearth_ci
  PGPASSWORD: ${{ secrets.HOMELAB_CI_POSTGRES_PASSWORD }}
  DBNAME: ci_run_${{ github.run_id }}_${{ github.run_attempt }}
steps:
  - uses: actions/checkout@v4
  - uses: actions/setup-node@v4
    with: { node-version: '20', cache: npm }
  - name: Install postgresql-client
    run: sudo apt-get update -qq && sudo apt-get install -y -qq postgresql-client
  - name: Cache node_modules
    uses: actions/cache@v4
    with:
      path: node_modules
      key: node-modules-${{ runner.os }}-node20-${{ hashFiles('package-lock.json') }}
  - run: npm ci
  - name: Create per-run database
    run: |
      psql -d postgres -c "CREATE DATABASE \"$DBNAME\";"
      echo "POSTGRES_URL=postgres://$PGUSER:$PGPASSWORD@$PGHOST:$PGPORT/$DBNAME" >> .env.local      
  - run: npm run migrate up
  - name: Drop per-run database (always)
    if: always()
    run: psql -d postgres -c "DROP DATABASE IF EXISTS \"$DBNAME\";"

psql reads PG* env vars automatically so we never need to assemble a connection URL on the psql command line (which was the failure mode in the first validation attempt). node-pg-migrate still wants a POSTGRES_URL, hence the inline URL written to .env.local. The runner is ephemeral so leaking the password into .env.local is bounded to that single job.

Why per-run DB:

  • Two PRs migrating in parallel don't collide.
  • A failed migration leaves a dirty DB behind — if: always() ensures cleanup.
  • Naming with run_id + run_attempt is collision-free even with re-runs.
  • psql is available on Debian 13 — add postgresql-client to CT 111's install if not already present (see axiom-server CT 111 README).

Architecture decisions to ratify

  • D1. Use runs-on: [self-hosted, axiom] (not [self-hosted] alone) so that if another runner is ever added with a different axiom-* label, these workflows still match correctly.
  • D2. Keep ubuntu-latest as the literal string in ZERO workflow files after migration — make CT 111 a hard dependency rather than a soft one. Rationale: dual-mode workflows hide drift (e.g. cache-key OS mismatch when swapping between the two). Single mode is simpler to reason about; if axiom is down, the operator runs the 1-line revert (D5).
  • D3. Rewire migrate to per-run DB on CT 102 (see above), NOT keep the ephemeral services.postgres block. The shared Postgres is already there and underutilized; the services: block on a self-hosted runner requires Docker-in-runner which adds 30s of pull/start time per run.
  • D4. Leave agent-context-drift.yml on ubuntu-latest. It's a weekly cron, costs ~2 min/month, and runs independently of axiom uptime. Trading $0 for resilience is a good trade here.
  • D5. Document a 1-line revert path in AGENTS.md § 7 Deployment: sed -i 's/\[self-hosted, axiom\]/ubuntu-latest/g' .github/workflows/*.yml for the case where axiom is offline mid-PR-storm and we need GitHub-hosted fallback fast. Operator manually re-enables billing or accepts the consumption for that day.

Risks

# Risk Likelihood Mitigation
1 CT 111 down → PRs queue indefinitely medium Beszel alerts on CT 111 down; D5 revert path documented
2 PAT expires silently → new jobs fail registration medium Calendar reminder at PAT mint time (90 days); runner logs surface failure on next restart
3 Malicious PR exfiltrates from runner low (repo-scoped, "require approval" enabled) Repo Settings gate; runner has no creds beyond secrets.HOMELAB_CI_POSTGRES_PASSWORD (scoped to deckhearth_ci, CREATEDB but no superuser, no access to other apps' databases)
4 Per-run DB litter on CT 102 if if: always() cleanup itself fails low Add a weekly cron on CT 102: psql ... -c "DROP DATABASE IF EXISTS …" FOREACH ci_run_* older than 7d
5 Cache poisoning across runs (shared ~/.npm between runner-1 and runner-2) low npm ci validates against package-lock.json checksum; corrupt cache is self-healing
6 Two ephemeral runners insufficient for peak load (5+ jobs per PR) medium Add runner-3: block in CT 111 compose; ~200 MB RAM per slot
7 Workflow file regressions (drift back to ubuntu-latest) low Add a forbidden-patterns check (8th gate): grep -rE "^\s*runs-on:\s*ubuntu-latest" .github/workflows/ should match only the agent-context-drift cron

Decomposition (proposed briefs)

  1. Brief 1 — Workflow migration. Single PR. Find/replace runs-on: ubuntu-latestruns-on: [self-hosted, axiom] in ci.yml, preview-smoke.yml, visual-diff.yml, pr-health-rollup.yml. Leave agent-context-drift.yml untouched (D4). Add a HOMELAB_CI_POSTGRES_BASE_URL secret to the repo before the PR opens (otherwise migrate job will fail on first run).
  2. Brief 2 — Migrate-job rewire. Same PR or split? Recommend SAME PR — the migrate job is part of ci.yml, splitting introduces a window where migrate runs on the self-hosted runner with no Postgres. Keep coupled.
  3. Brief 3 — Forbidden-pattern gate (Risk #7). Add 8th check to the forbidden-patterns job: runs-on: ubuntu-latest is only allowed in agent-context-drift.yml. Cheap insurance against drift.
  4. Brief 4 — Documentation. Update AGENTS.md § 6 Testing + § 7 Deployment with the self-hosted runner story + the D5 revert path. Add axiom-server/proxmox/ct111/README.md as a cross-reference.

Briefs 1 + 2 are tightly coupled — recommend single combined PR. Briefs 3 + 4 are independent and can dispatch in parallel after Brief 1+2 lands.

Validation

After Brief 1+2 merges:

  1. Open a no-op PR (e.g. add a trailing newline to README.md — wait, that hits paths-ignore. Use a 1-line code comment in pages/index.js instead).
  2. Confirm in the PR's Checks tab: all jobs report In progress on axiom-runner-1 or axiom-runner-2 within 10s of dispatch.
  3. Confirm GitHub Actions billing page shows zero minute consumption for that PR.
  4. SSH axiom and verify the per-run DB was created + dropped: ./proxmox/scripts/sync.sh exec 102 "docker exec postgres psql -U postgres -c '\l'" — no ci_run_* databases should linger.

Acceptance

  • 4/4 workflows green on self-hosted runner.
  • GitHub Actions minute consumption drops to ~0 (only the weekly agent-context-drift cron remains on ubuntu-latest).
  • Forbidden-pattern gate (8th check) catches accidental runs-on: ubuntu-latest reintroduction.
  • AGENTS.md § 6/7 updated; CT 111 README cross-referenced.

Queued follow-ups

  • seed-visual-baselines-on-linux — already queued (see ship-readiness.md). With axiom now in the loop, this gets easier: baselines can generate on CT 111 directly via npm run test:visual:update against a preview deploy, producing Linux-compatible PNGs the CI runner will match exactly.
  • cleanup-stale-ci-runs-cron — weekly cron on CT 102 to drop ci_run_* DBs older than 7d (Risk #4 mitigation, defensive).