deckhearth/.convoys/migrate-ci-to-self-hosted.md
Randall Stillwell b712e38c86 fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret
First Brief 1+2 validation run failed on the migrate job with
`psql: invalid option -- '/'` despite the secret being set correctly
and a direct CT-111 → CT-102 psql connection working fine. The
URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit.

Splitting the connection into discrete `PG*` env vars (which psql
picks up automatically) sidesteps URL parsing entirely. The
`HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now
`HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow
hardcodes the (non-sensitive) host/port/user. `node-pg-migrate`
still reads `POSTGRES_URL` from `.env.local`, so we assemble that
URL inline for it; the runner is ephemeral so the leaked-to-disk
password is bounded to one job.

Convoy doc updated to reflect the shipped approach + lesson learned
in prerequisites.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-05 18:00:11 -05:00

210 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
slug: migrate-ci-to-self-hosted
status: queued
opened: 2026-06-05
owner: rstillw
prerequisites:
- CT 111 (`ci-runner`) provisioned and online on axiom (`192.168.68.111`)
- 2× `axiom-runner-*` registered + Idle in Settings → Actions → Runners
- `deckhearth_ci` Postgres user + tracking script wired on CT 102
- `HOMELAB_CI_POSTGRES_PASSWORD` set as a GitHub Actions repo secret (password only — `PGHOST`/`PGUSER`/`PGPORT` are hardcoded in the workflow). Earlier draft of this convoy used a single `HOMELAB_CI_POSTGRES_BASE_URL` URL secret, but during Brief 1+2 validation `psql "$URL/postgres"` failed with `invalid option -- '/'` — the URL-parse path was fragile. Split secret + standard `PG*` env vars sidesteps it.
- Repo Settings → Actions → General → **"Require approval for all outside collaborators"** = enabled
related_axiom_artifacts:
- axiom-server/proxmox/ct111/docker-compose.yml
- axiom-server/proxmox/ct111/.env.example
- axiom-server/proxmox/ct111/README.md
- axiom-server/.cursor/rules/ct111-ci-runner.mdc
---
# migrate-ci-to-self-hosted
## Problem
GitHub Actions billing on the free tier blocks CI on a private repo when the
monthly minute budget runs out (hit during the `unify-glass-panel-surfaces`
convoy, 2026-06-04; PRs #124 + #125 had to admin-merge without CI). The
`slash-ci-minutes` convoy (PR #126) reduced consumption by ~60% via
`paths-ignore`, grep consolidation, and caching, but a busy week of
implementation work still trips the limit.
This convoy migrates all 4 GitHub Actions workflows off `ubuntu-latest`
(GitHub-hosted, billed) onto the axiom homelab runner (`CT 111`,
self-hosted, free). It also rewires the `migrate` job to use CT 102's
shared Postgres instead of spinning up an ephemeral container per run —
saving ~30s/PR and eliminating the Docker-in-runner pull cost.
## Non-goals
- Migrating to a different CI provider (CircleCI, Buildkite, etc.) — overkill.
- Hosting the production app on axiom — Vercel keeps the deploy story simple
and the homelab is already at ~95% RAM allocation. Out of scope.
- Replacing the Vercel Preview deployments — Vercel still builds previews;
Playwright smoke + visual-diff still run *against* those previews from the
self-hosted runner.
## Workflows to migrate
| File | Jobs | Notes |
|---|---|---|
| `.github/workflows/ci.yml` | lint, schema-map-fresh, forbidden-patterns, migrate, test | `migrate` needs the Postgres rewire (see below) |
| `.github/workflows/preview-smoke.yml` | gate, smoke | smoke job needs Chromium — first run will `npx playwright install` and cache it |
| `.github/workflows/visual-diff.yml` | (visual) | same Chromium cache benefit; baselines still TBD |
| `.github/workflows/pr-health-rollup.yml` | rollup | trivial — single `gh pr comment` job |
| `.github/workflows/agent-context-drift.yml` | (weekly cron) | Optional: leave on `ubuntu-latest` so the cron runs even when axiom is down. Decision deferred — see Risk #3 |
Change shape per job:
```yaml
# Before
jobs:
lint:
runs-on: ubuntu-latest
# After
jobs:
lint:
runs-on: [self-hosted, axiom]
```
## Migrate-job Postgres rewire
Today (`ci.yml` § migrate):
```yaml
services:
postgres:
image: postgres:16
env: { POSTGRES_USER: postgres, POSTGRES_PASSWORD: postgres, POSTGRES_DB: deckhearth_test }
ports: ["5432:5432"]
env:
POSTGRES_URL: postgres://postgres:postgres@localhost:5432/deckhearth_test
```
Shipped (post-validation revision):
```yaml
env:
PGHOST: 192.168.68.102
PGPORT: '5432'
PGUSER: deckhearth_ci
PGPASSWORD: ${{ secrets.HOMELAB_CI_POSTGRES_PASSWORD }}
DBNAME: ci_run_${{ github.run_id }}_${{ github.run_attempt }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: '20', cache: npm }
- name: Install postgresql-client
run: sudo apt-get update -qq && sudo apt-get install -y -qq postgresql-client
- name: Cache node_modules
uses: actions/cache@v4
with:
path: node_modules
key: node-modules-${{ runner.os }}-node20-${{ hashFiles('package-lock.json') }}
- run: npm ci
- name: Create per-run database
run: |
psql -d postgres -c "CREATE DATABASE \"$DBNAME\";"
echo "POSTGRES_URL=postgres://$PGUSER:$PGPASSWORD@$PGHOST:$PGPORT/$DBNAME" >> .env.local
- run: npm run migrate up
- name: Drop per-run database (always)
if: always()
run: psql -d postgres -c "DROP DATABASE IF EXISTS \"$DBNAME\";"
```
`psql` reads `PG*` env vars automatically so we never need to assemble a connection URL on the psql command line (which was the failure mode in the first validation attempt). `node-pg-migrate` still wants a `POSTGRES_URL`, hence the inline URL written to `.env.local`. The runner is ephemeral so leaking the password into `.env.local` is bounded to that single job.
Why per-run DB:
- Two PRs migrating in parallel don't collide.
- A failed migration leaves a dirty DB behind — `if: always()` ensures cleanup.
- Naming with `run_id` + `run_attempt` is collision-free even with re-runs.
- `psql` is available on Debian 13 — add `postgresql-client` to CT 111's
install if not already present (see axiom-server CT 111 README).
## Architecture decisions to ratify
- **D1.** Use `runs-on: [self-hosted, axiom]` (not `[self-hosted]` alone)
so that if another runner is ever added with a different `axiom-*` label,
these workflows still match correctly.
- **D2.** Keep `ubuntu-latest` as the literal string in ZERO workflow files
after migration — make CT 111 a hard dependency rather than a soft one.
Rationale: dual-mode workflows hide drift (e.g. cache-key OS mismatch when
swapping between the two). Single mode is simpler to reason about; if
axiom is down, the operator runs the 1-line revert (D5).
- **D3.** Rewire `migrate` to per-run DB on CT 102 (see above), NOT keep the
ephemeral `services.postgres` block. The shared Postgres is already there
and underutilized; the `services:` block on a self-hosted runner requires
Docker-in-runner which adds 30s of pull/start time per run.
- **D4.** Leave `agent-context-drift.yml` on `ubuntu-latest`. It's a weekly
cron, costs ~2 min/month, and runs independently of axiom uptime. Trading
$0 for resilience is a good trade here.
- **D5.** Document a 1-line revert path in `AGENTS.md` § 7 Deployment:
`sed -i 's/\[self-hosted, axiom\]/ubuntu-latest/g' .github/workflows/*.yml`
for the case where axiom is offline mid-PR-storm and we need GitHub-hosted
fallback fast. Operator manually re-enables billing or accepts the
consumption for that day.
## Risks
| # | Risk | Likelihood | Mitigation |
|---|---|---|---|
| 1 | CT 111 down → PRs queue indefinitely | medium | Beszel alerts on CT 111 down; D5 revert path documented |
| 2 | PAT expires silently → new jobs fail registration | medium | Calendar reminder at PAT mint time (90 days); runner logs surface failure on next restart |
| 3 | Malicious PR exfiltrates from runner | low (repo-scoped, "require approval" enabled) | Repo Settings gate; runner has no creds beyond `secrets.HOMELAB_CI_POSTGRES_PASSWORD` (scoped to `deckhearth_ci`, CREATEDB but no superuser, no access to other apps' databases) |
| 4 | Per-run DB litter on CT 102 if `if: always()` cleanup itself fails | low | Add a weekly cron on CT 102: `psql ... -c "DROP DATABASE IF EXISTS …" FOREACH ci_run_* older than 7d` |
| 5 | Cache poisoning across runs (shared `~/.npm` between runner-1 and runner-2) | low | `npm ci` validates against `package-lock.json` checksum; corrupt cache is self-healing |
| 6 | Two ephemeral runners insufficient for peak load (5+ jobs per PR) | medium | Add `runner-3:` block in CT 111 compose; ~200 MB RAM per slot |
| 7 | Workflow file regressions (drift back to `ubuntu-latest`) | low | Add a forbidden-patterns check (8th gate): `grep -rE "^\s*runs-on:\s*ubuntu-latest" .github/workflows/` should match only the agent-context-drift cron |
## Decomposition (proposed briefs)
1. **Brief 1 — Workflow migration.** Single PR. Find/replace `runs-on:
ubuntu-latest` → `runs-on: [self-hosted, axiom]` in `ci.yml`,
`preview-smoke.yml`, `visual-diff.yml`, `pr-health-rollup.yml`. Leave
`agent-context-drift.yml` untouched (D4). Add a HOMELAB_CI_POSTGRES_BASE_URL
secret to the repo before the PR opens (otherwise `migrate` job will fail
on first run).
2. **Brief 2 — Migrate-job rewire.** Same PR or split? Recommend SAME PR —
the migrate job is part of `ci.yml`, splitting introduces a window where
migrate runs on the self-hosted runner with no Postgres. Keep coupled.
3. **Brief 3 — Forbidden-pattern gate (Risk #7).** Add 8th check to the
`forbidden-patterns` job: `runs-on: ubuntu-latest` is only allowed in
`agent-context-drift.yml`. Cheap insurance against drift.
4. **Brief 4 — Documentation.** Update `AGENTS.md` § 6 Testing + § 7
Deployment with the self-hosted runner story + the D5 revert path. Add
`axiom-server/proxmox/ct111/README.md` as a cross-reference.
Briefs 1 + 2 are tightly coupled — recommend single combined PR. Briefs 3 + 4
are independent and can dispatch in parallel after Brief 1+2 lands.
## Validation
After Brief 1+2 merges:
1. Open a no-op PR (e.g. add a trailing newline to `README.md` — wait, that
hits `paths-ignore`. Use a 1-line code comment in `pages/index.js` instead).
2. Confirm in the PR's Checks tab: all jobs report `In progress on
axiom-runner-1` or `axiom-runner-2` within 10s of dispatch.
3. Confirm GitHub Actions billing page shows zero minute consumption for
that PR.
4. SSH `axiom` and verify the per-run DB was created + dropped:
`./proxmox/scripts/sync.sh exec 102 "docker exec postgres psql -U postgres -c '\l'"` — no
`ci_run_*` databases should linger.
## Acceptance
- 4/4 workflows green on self-hosted runner.
- GitHub Actions minute consumption drops to ~0 (only the weekly
`agent-context-drift` cron remains on `ubuntu-latest`).
- Forbidden-pattern gate (8th check) catches accidental
`runs-on: ubuntu-latest` reintroduction.
- AGENTS.md § 6/7 updated; CT 111 README cross-referenced.
## Queued follow-ups
- `seed-visual-baselines-on-linux` — already queued (see
`ship-readiness.md`). With axiom now in the loop, this gets easier:
baselines can generate on CT 111 directly via
`npm run test:visual:update` against a preview deploy, producing
Linux-compatible PNGs the CI runner will match exactly.
- `cleanup-stale-ci-runs-cron` — weekly cron on CT 102 to drop
`ci_run_*` DBs older than 7d (Risk #4 mitigation, defensive).