deckhearth/.convoys/migrate-ci-to-self-hosted.md

218 lines
11 KiB
Markdown
Raw Normal View History

convoy: migrate CI to self-hosted axiom runners (briefs 1+2) (#132) * convoy: migrate CI to self-hosted axiom runners (briefs 1+2) Moves 4 of 5 GitHub Actions workflows from `ubuntu-latest` to the new `stwl-labs` org-level self-hosted pool (CT 111 axiom-runner-1..4) and rewires the `migrate` job to use CT 102's shared Postgres via per-run databases. Changes: - ci.yml: lint, schema-map-fresh, forbidden-patterns, migrate, test -> `[self-hosted, axiom]`. migrate job drops `services.postgres` (saved ~30s/run of image pull) and switches to `HOMELAB_CI_POSTGRES_BASE_URL` secret + per-run DB (`ci_run_<run_id>_<run_attempt>`) with `always()` cleanup so failed migrations don't leak DBs. - preview-smoke.yml: gate + smoke -> self-hosted. Playwright browser cache lives under /opt/appdata/gha-runner/shared-cache/playwright on the host bind mount; first PR primes it, subsequent runs reuse. - visual-diff.yml: gate + visual -> self-hosted (same Playwright cache). - pr-health-rollup.yml: rollup -> self-hosted. - agent-context-drift.yml: deliberately LEFT on ubuntu-latest (D4 in convoy doc). Weekly cron stays GitHub-hosted so it runs even when axiom is down. Why on this side and not the runner side: - migrate adds an explicit `sudo apt-get install -y postgresql-client` step (~10s, amortized via apt-cache survival). The runner image doesn't ship psql; baking it in would require a custom image and doesn't earn its keep for one job. Repo prereqs (set before this PR opens): - `HOMELAB_CI_POSTGRES_BASE_URL` repo secret set (value pattern: `postgres://deckhearth_ci:<pw>@192.168.68.102:5432`) - `deckhearth_ci` Postgres user created on CT 102 with CREATEDB, no superuser - stwl-labs org Actions settings: "Require approval for all outside collaborators" + runner group rejects public repos - 4 runners online: `axiom-runner-1..4`, status Idle Follow-ups (per convoy): - Brief 3: forbidden-pattern gate to catch `runs-on: ubuntu-latest` re-introduction outside the agent-context-drift allowlist - Brief 4: AGENTS.md updates + 1-line revert path (D5) - Weekly cron on CT 102 to GC any `ci_run_*` DBs older than 7d (Risk #4 mitigation) Co-authored-by: Cursor <cursoragent@cursor.com> * fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret First Brief 1+2 validation run failed on the migrate job with `psql: invalid option -- '/'` despite the secret being set correctly and a direct CT-111 → CT-102 psql connection working fine. The URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit. Splitting the connection into discrete `PG*` env vars (which psql picks up automatically) sidesteps URL parsing entirely. The `HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now `HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow hardcodes the (non-sensitive) host/port/user. `node-pg-migrate` still reads `POSTGRES_URL` from `.env.local`, so we assemble that URL inline for it; the runner is ephemeral so the leaked-to-disk password is bounded to one job. Convoy doc updated to reflect the shipped approach + lesson learned in prerequisites. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: trigger vercel preview after stwl-labs reauth Co-authored-by: Cursor <cursoragent@cursor.com> * ci: re-test playwright after vercel project rebind Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:46:31 -04:00
---
slug: migrate-ci-to-self-hosted
convoy: forbidden-pattern gate + docs (briefs 3+4) Closes out the migrate-ci-to-self-hosted convoy with the two defensive follow-ups Brief 1+2 (PR #132) intentionally deferred. Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps `.github/workflows/` for `runs-on: ubuntu-latest` and fails unless the match is in the documented allowlist (currently `agent-context-drift.yml` only, per Decision D4). Self-tested locally against the post-migration tree: 0 violations. Renames the job from "Forbidden patterns (7 checks)" → "(8 checks)" and normalizes the older "Check N/6" labels to "N/8" for consistency (the inherited mix of `/6` and `/7` was a known cosmetic from the unify-glass-panel-surfaces convoy). Brief 4 — AGENTS.md § 6 and § 7 updates: - § 6 "CI behavior": Playwright smoke runtime range updated to cover post-migration cold vs. warm cache (was a stale 59s figure from pre-migration ubuntu-latest). - § 6 new top-level bullet "Self-hosted runner pool" alongside "CI minute optimizations" — covers where runners live, where caches are bind-mounted on CT 111, the Postgres rewire on CT 102, and the agent-context-drift.yml exemption + how Check 8 enforces it. - § 7 new bullet for the operational story: PAT rotation cadence + the D5 one-line `sed` revert path for when axiom is offline mid-PR-storm. Cross-references the axiom-server CT 111 README and the Beszel down alert. Convoy doc — status flipped queued → shipped, shipped_in lists both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron + seed-visual-baselines-on-linux items machine-greppable. Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:49:07 -04:00
status: shipped
convoy: migrate CI to self-hosted axiom runners (briefs 1+2) (#132) * convoy: migrate CI to self-hosted axiom runners (briefs 1+2) Moves 4 of 5 GitHub Actions workflows from `ubuntu-latest` to the new `stwl-labs` org-level self-hosted pool (CT 111 axiom-runner-1..4) and rewires the `migrate` job to use CT 102's shared Postgres via per-run databases. Changes: - ci.yml: lint, schema-map-fresh, forbidden-patterns, migrate, test -> `[self-hosted, axiom]`. migrate job drops `services.postgres` (saved ~30s/run of image pull) and switches to `HOMELAB_CI_POSTGRES_BASE_URL` secret + per-run DB (`ci_run_<run_id>_<run_attempt>`) with `always()` cleanup so failed migrations don't leak DBs. - preview-smoke.yml: gate + smoke -> self-hosted. Playwright browser cache lives under /opt/appdata/gha-runner/shared-cache/playwright on the host bind mount; first PR primes it, subsequent runs reuse. - visual-diff.yml: gate + visual -> self-hosted (same Playwright cache). - pr-health-rollup.yml: rollup -> self-hosted. - agent-context-drift.yml: deliberately LEFT on ubuntu-latest (D4 in convoy doc). Weekly cron stays GitHub-hosted so it runs even when axiom is down. Why on this side and not the runner side: - migrate adds an explicit `sudo apt-get install -y postgresql-client` step (~10s, amortized via apt-cache survival). The runner image doesn't ship psql; baking it in would require a custom image and doesn't earn its keep for one job. Repo prereqs (set before this PR opens): - `HOMELAB_CI_POSTGRES_BASE_URL` repo secret set (value pattern: `postgres://deckhearth_ci:<pw>@192.168.68.102:5432`) - `deckhearth_ci` Postgres user created on CT 102 with CREATEDB, no superuser - stwl-labs org Actions settings: "Require approval for all outside collaborators" + runner group rejects public repos - 4 runners online: `axiom-runner-1..4`, status Idle Follow-ups (per convoy): - Brief 3: forbidden-pattern gate to catch `runs-on: ubuntu-latest` re-introduction outside the agent-context-drift allowlist - Brief 4: AGENTS.md updates + 1-line revert path (D5) - Weekly cron on CT 102 to GC any `ci_run_*` DBs older than 7d (Risk #4 mitigation) Co-authored-by: Cursor <cursoragent@cursor.com> * fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret First Brief 1+2 validation run failed on the migrate job with `psql: invalid option -- '/'` despite the secret being set correctly and a direct CT-111 → CT-102 psql connection working fine. The URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit. Splitting the connection into discrete `PG*` env vars (which psql picks up automatically) sidesteps URL parsing entirely. The `HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now `HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow hardcodes the (non-sensitive) host/port/user. `node-pg-migrate` still reads `POSTGRES_URL` from `.env.local`, so we assemble that URL inline for it; the runner is ephemeral so the leaked-to-disk password is bounded to one job. Convoy doc updated to reflect the shipped approach + lesson learned in prerequisites. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: trigger vercel preview after stwl-labs reauth Co-authored-by: Cursor <cursoragent@cursor.com> * ci: re-test playwright after vercel project rebind Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:46:31 -04:00
opened: 2026-06-05
convoy: forbidden-pattern gate + docs (briefs 3+4) Closes out the migrate-ci-to-self-hosted convoy with the two defensive follow-ups Brief 1+2 (PR #132) intentionally deferred. Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps `.github/workflows/` for `runs-on: ubuntu-latest` and fails unless the match is in the documented allowlist (currently `agent-context-drift.yml` only, per Decision D4). Self-tested locally against the post-migration tree: 0 violations. Renames the job from "Forbidden patterns (7 checks)" → "(8 checks)" and normalizes the older "Check N/6" labels to "N/8" for consistency (the inherited mix of `/6` and `/7` was a known cosmetic from the unify-glass-panel-surfaces convoy). Brief 4 — AGENTS.md § 6 and § 7 updates: - § 6 "CI behavior": Playwright smoke runtime range updated to cover post-migration cold vs. warm cache (was a stale 59s figure from pre-migration ubuntu-latest). - § 6 new top-level bullet "Self-hosted runner pool" alongside "CI minute optimizations" — covers where runners live, where caches are bind-mounted on CT 111, the Postgres rewire on CT 102, and the agent-context-drift.yml exemption + how Check 8 enforces it. - § 7 new bullet for the operational story: PAT rotation cadence + the D5 one-line `sed` revert path for when axiom is offline mid-PR-storm. Cross-references the axiom-server CT 111 README and the Beszel down alert. Convoy doc — status flipped queued → shipped, shipped_in lists both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron + seed-visual-baselines-on-linux items machine-greppable. Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:49:07 -04:00
shipped: 2026-06-05
convoy: migrate CI to self-hosted axiom runners (briefs 1+2) (#132) * convoy: migrate CI to self-hosted axiom runners (briefs 1+2) Moves 4 of 5 GitHub Actions workflows from `ubuntu-latest` to the new `stwl-labs` org-level self-hosted pool (CT 111 axiom-runner-1..4) and rewires the `migrate` job to use CT 102's shared Postgres via per-run databases. Changes: - ci.yml: lint, schema-map-fresh, forbidden-patterns, migrate, test -> `[self-hosted, axiom]`. migrate job drops `services.postgres` (saved ~30s/run of image pull) and switches to `HOMELAB_CI_POSTGRES_BASE_URL` secret + per-run DB (`ci_run_<run_id>_<run_attempt>`) with `always()` cleanup so failed migrations don't leak DBs. - preview-smoke.yml: gate + smoke -> self-hosted. Playwright browser cache lives under /opt/appdata/gha-runner/shared-cache/playwright on the host bind mount; first PR primes it, subsequent runs reuse. - visual-diff.yml: gate + visual -> self-hosted (same Playwright cache). - pr-health-rollup.yml: rollup -> self-hosted. - agent-context-drift.yml: deliberately LEFT on ubuntu-latest (D4 in convoy doc). Weekly cron stays GitHub-hosted so it runs even when axiom is down. Why on this side and not the runner side: - migrate adds an explicit `sudo apt-get install -y postgresql-client` step (~10s, amortized via apt-cache survival). The runner image doesn't ship psql; baking it in would require a custom image and doesn't earn its keep for one job. Repo prereqs (set before this PR opens): - `HOMELAB_CI_POSTGRES_BASE_URL` repo secret set (value pattern: `postgres://deckhearth_ci:<pw>@192.168.68.102:5432`) - `deckhearth_ci` Postgres user created on CT 102 with CREATEDB, no superuser - stwl-labs org Actions settings: "Require approval for all outside collaborators" + runner group rejects public repos - 4 runners online: `axiom-runner-1..4`, status Idle Follow-ups (per convoy): - Brief 3: forbidden-pattern gate to catch `runs-on: ubuntu-latest` re-introduction outside the agent-context-drift allowlist - Brief 4: AGENTS.md updates + 1-line revert path (D5) - Weekly cron on CT 102 to GC any `ci_run_*` DBs older than 7d (Risk #4 mitigation) Co-authored-by: Cursor <cursoragent@cursor.com> * fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret First Brief 1+2 validation run failed on the migrate job with `psql: invalid option -- '/'` despite the secret being set correctly and a direct CT-111 → CT-102 psql connection working fine. The URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit. Splitting the connection into discrete `PG*` env vars (which psql picks up automatically) sidesteps URL parsing entirely. The `HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now `HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow hardcodes the (non-sensitive) host/port/user. `node-pg-migrate` still reads `POSTGRES_URL` from `.env.local`, so we assemble that URL inline for it; the runner is ephemeral so the leaked-to-disk password is bounded to one job. Convoy doc updated to reflect the shipped approach + lesson learned in prerequisites. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: trigger vercel preview after stwl-labs reauth Co-authored-by: Cursor <cursoragent@cursor.com> * ci: re-test playwright after vercel project rebind Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:46:31 -04:00
owner: rstillw
convoy: forbidden-pattern gate + docs (briefs 3+4) Closes out the migrate-ci-to-self-hosted convoy with the two defensive follow-ups Brief 1+2 (PR #132) intentionally deferred. Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps `.github/workflows/` for `runs-on: ubuntu-latest` and fails unless the match is in the documented allowlist (currently `agent-context-drift.yml` only, per Decision D4). Self-tested locally against the post-migration tree: 0 violations. Renames the job from "Forbidden patterns (7 checks)" → "(8 checks)" and normalizes the older "Check N/6" labels to "N/8" for consistency (the inherited mix of `/6` and `/7` was a known cosmetic from the unify-glass-panel-surfaces convoy). Brief 4 — AGENTS.md § 6 and § 7 updates: - § 6 "CI behavior": Playwright smoke runtime range updated to cover post-migration cold vs. warm cache (was a stale 59s figure from pre-migration ubuntu-latest). - § 6 new top-level bullet "Self-hosted runner pool" alongside "CI minute optimizations" — covers where runners live, where caches are bind-mounted on CT 111, the Postgres rewire on CT 102, and the agent-context-drift.yml exemption + how Check 8 enforces it. - § 7 new bullet for the operational story: PAT rotation cadence + the D5 one-line `sed` revert path for when axiom is offline mid-PR-storm. Cross-references the axiom-server CT 111 README and the Beszel down alert. Convoy doc — status flipped queued → shipped, shipped_in lists both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron + seed-visual-baselines-on-linux items machine-greppable. Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:49:07 -04:00
shipped_in:
- PR #132 (Briefs 1+2 — workflow migration + migrate-job rewire)
- PR #133 (Briefs 3+4 — forbidden-pattern gate + AGENTS.md docs)
convoy: forbidden-pattern gate + docs (briefs 3+4) Closes out the migrate-ci-to-self-hosted convoy with the two defensive follow-ups Brief 1+2 (PR #132) intentionally deferred. Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps `.github/workflows/` for `runs-on: ubuntu-latest` and fails unless the match is in the documented allowlist (currently `agent-context-drift.yml` only, per Decision D4). Self-tested locally against the post-migration tree: 0 violations. Renames the job from "Forbidden patterns (7 checks)" → "(8 checks)" and normalizes the older "Check N/6" labels to "N/8" for consistency (the inherited mix of `/6` and `/7` was a known cosmetic from the unify-glass-panel-surfaces convoy). Brief 4 — AGENTS.md § 6 and § 7 updates: - § 6 "CI behavior": Playwright smoke runtime range updated to cover post-migration cold vs. warm cache (was a stale 59s figure from pre-migration ubuntu-latest). - § 6 new top-level bullet "Self-hosted runner pool" alongside "CI minute optimizations" — covers where runners live, where caches are bind-mounted on CT 111, the Postgres rewire on CT 102, and the agent-context-drift.yml exemption + how Check 8 enforces it. - § 7 new bullet for the operational story: PAT rotation cadence + the D5 one-line `sed` revert path for when axiom is offline mid-PR-storm. Cross-references the axiom-server CT 111 README and the Beszel down alert. Convoy doc — status flipped queued → shipped, shipped_in lists both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron + seed-visual-baselines-on-linux items machine-greppable. Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:49:07 -04:00
follow_ups:
- cleanup-stale-ci-runs-cron (weekly GC on CT 102 for ci_run_* DBs older than 7d; Risk #4 defensive)
- seed-visual-baselines-on-linux (now easier with axiom; see queued-follow-ups below)
convoy: migrate CI to self-hosted axiom runners (briefs 1+2) (#132) * convoy: migrate CI to self-hosted axiom runners (briefs 1+2) Moves 4 of 5 GitHub Actions workflows from `ubuntu-latest` to the new `stwl-labs` org-level self-hosted pool (CT 111 axiom-runner-1..4) and rewires the `migrate` job to use CT 102's shared Postgres via per-run databases. Changes: - ci.yml: lint, schema-map-fresh, forbidden-patterns, migrate, test -> `[self-hosted, axiom]`. migrate job drops `services.postgres` (saved ~30s/run of image pull) and switches to `HOMELAB_CI_POSTGRES_BASE_URL` secret + per-run DB (`ci_run_<run_id>_<run_attempt>`) with `always()` cleanup so failed migrations don't leak DBs. - preview-smoke.yml: gate + smoke -> self-hosted. Playwright browser cache lives under /opt/appdata/gha-runner/shared-cache/playwright on the host bind mount; first PR primes it, subsequent runs reuse. - visual-diff.yml: gate + visual -> self-hosted (same Playwright cache). - pr-health-rollup.yml: rollup -> self-hosted. - agent-context-drift.yml: deliberately LEFT on ubuntu-latest (D4 in convoy doc). Weekly cron stays GitHub-hosted so it runs even when axiom is down. Why on this side and not the runner side: - migrate adds an explicit `sudo apt-get install -y postgresql-client` step (~10s, amortized via apt-cache survival). The runner image doesn't ship psql; baking it in would require a custom image and doesn't earn its keep for one job. Repo prereqs (set before this PR opens): - `HOMELAB_CI_POSTGRES_BASE_URL` repo secret set (value pattern: `postgres://deckhearth_ci:<pw>@192.168.68.102:5432`) - `deckhearth_ci` Postgres user created on CT 102 with CREATEDB, no superuser - stwl-labs org Actions settings: "Require approval for all outside collaborators" + runner group rejects public repos - 4 runners online: `axiom-runner-1..4`, status Idle Follow-ups (per convoy): - Brief 3: forbidden-pattern gate to catch `runs-on: ubuntu-latest` re-introduction outside the agent-context-drift allowlist - Brief 4: AGENTS.md updates + 1-line revert path (D5) - Weekly cron on CT 102 to GC any `ci_run_*` DBs older than 7d (Risk #4 mitigation) Co-authored-by: Cursor <cursoragent@cursor.com> * fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret First Brief 1+2 validation run failed on the migrate job with `psql: invalid option -- '/'` despite the secret being set correctly and a direct CT-111 → CT-102 psql connection working fine. The URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit. Splitting the connection into discrete `PG*` env vars (which psql picks up automatically) sidesteps URL parsing entirely. The `HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now `HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow hardcodes the (non-sensitive) host/port/user. `node-pg-migrate` still reads `POSTGRES_URL` from `.env.local`, so we assemble that URL inline for it; the runner is ephemeral so the leaked-to-disk password is bounded to one job. Convoy doc updated to reflect the shipped approach + lesson learned in prerequisites. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: trigger vercel preview after stwl-labs reauth Co-authored-by: Cursor <cursoragent@cursor.com> * ci: re-test playwright after vercel project rebind Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 00:46:31 -04:00
prerequisites:
- CT 111 (`ci-runner`) provisioned and online on axiom (`192.168.68.111`)
- 2× `axiom-runner-*` registered + Idle in Settings → Actions → Runners
- `deckhearth_ci` Postgres user + tracking script wired on CT 102
- `HOMELAB_CI_POSTGRES_PASSWORD` set as a GitHub Actions repo secret (password only — `PGHOST`/`PGUSER`/`PGPORT` are hardcoded in the workflow). Earlier draft of this convoy used a single `HOMELAB_CI_POSTGRES_BASE_URL` URL secret, but during Brief 1+2 validation `psql "$URL/postgres"` failed with `invalid option -- '/'` — the URL-parse path was fragile. Split secret + standard `PG*` env vars sidesteps it.
- Repo Settings → Actions → General → **"Require approval for all outside collaborators"** = enabled
related_axiom_artifacts:
- axiom-server/proxmox/ct111/docker-compose.yml
- axiom-server/proxmox/ct111/.env.example
- axiom-server/proxmox/ct111/README.md
- axiom-server/.cursor/rules/ct111-ci-runner.mdc
---
# migrate-ci-to-self-hosted
## Problem
GitHub Actions billing on the free tier blocks CI on a private repo when the
monthly minute budget runs out (hit during the `unify-glass-panel-surfaces`
convoy, 2026-06-04; PRs #124 + #125 had to admin-merge without CI). The
`slash-ci-minutes` convoy (PR #126) reduced consumption by ~60% via
`paths-ignore`, grep consolidation, and caching, but a busy week of
implementation work still trips the limit.
This convoy migrates all 4 GitHub Actions workflows off `ubuntu-latest`
(GitHub-hosted, billed) onto the axiom homelab runner (`CT 111`,
self-hosted, free). It also rewires the `migrate` job to use CT 102's
shared Postgres instead of spinning up an ephemeral container per run —
saving ~30s/PR and eliminating the Docker-in-runner pull cost.
## Non-goals
- Migrating to a different CI provider (CircleCI, Buildkite, etc.) — overkill.
- Hosting the production app on axiom — Vercel keeps the deploy story simple
and the homelab is already at ~95% RAM allocation. Out of scope.
- Replacing the Vercel Preview deployments — Vercel still builds previews;
Playwright smoke + visual-diff still run *against* those previews from the
self-hosted runner.
## Workflows to migrate
| File | Jobs | Notes |
|---|---|---|
| `.github/workflows/ci.yml` | lint, schema-map-fresh, forbidden-patterns, migrate, test | `migrate` needs the Postgres rewire (see below) |
| `.github/workflows/preview-smoke.yml` | gate, smoke | smoke job needs Chromium — first run will `npx playwright install` and cache it |
| `.github/workflows/visual-diff.yml` | (visual) | same Chromium cache benefit; baselines still TBD |
| `.github/workflows/pr-health-rollup.yml` | rollup | trivial — single `gh pr comment` job |
| `.github/workflows/agent-context-drift.yml` | (weekly cron) | Optional: leave on `ubuntu-latest` so the cron runs even when axiom is down. Decision deferred — see Risk #3 |
Change shape per job:
```yaml
# Before
jobs:
lint:
runs-on: ubuntu-latest
# After
jobs:
lint:
runs-on: [self-hosted, axiom]
```
## Migrate-job Postgres rewire
Today (`ci.yml` § migrate):
```yaml
services:
postgres:
image: postgres:16
env: { POSTGRES_USER: postgres, POSTGRES_PASSWORD: postgres, POSTGRES_DB: deckhearth_test }
ports: ["5432:5432"]
env:
POSTGRES_URL: postgres://postgres:postgres@localhost:5432/deckhearth_test
```
Shipped (post-validation revision):
```yaml
env:
PGHOST: 192.168.68.102
PGPORT: '5432'
PGUSER: deckhearth_ci
PGPASSWORD: ${{ secrets.HOMELAB_CI_POSTGRES_PASSWORD }}
DBNAME: ci_run_${{ github.run_id }}_${{ github.run_attempt }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: '20', cache: npm }
- name: Install postgresql-client
run: sudo apt-get update -qq && sudo apt-get install -y -qq postgresql-client
- name: Cache node_modules
uses: actions/cache@v4
with:
path: node_modules
key: node-modules-${{ runner.os }}-node20-${{ hashFiles('package-lock.json') }}
- run: npm ci
- name: Create per-run database
run: |
psql -d postgres -c "CREATE DATABASE \"$DBNAME\";"
echo "POSTGRES_URL=postgres://$PGUSER:$PGPASSWORD@$PGHOST:$PGPORT/$DBNAME" >> .env.local
- run: npm run migrate up
- name: Drop per-run database (always)
if: always()
run: psql -d postgres -c "DROP DATABASE IF EXISTS \"$DBNAME\";"
```
`psql` reads `PG*` env vars automatically so we never need to assemble a connection URL on the psql command line (which was the failure mode in the first validation attempt). `node-pg-migrate` still wants a `POSTGRES_URL`, hence the inline URL written to `.env.local`. The runner is ephemeral so leaking the password into `.env.local` is bounded to that single job.
Why per-run DB:
- Two PRs migrating in parallel don't collide.
- A failed migration leaves a dirty DB behind — `if: always()` ensures cleanup.
- Naming with `run_id` + `run_attempt` is collision-free even with re-runs.
- `psql` is available on Debian 13 — add `postgresql-client` to CT 111's
install if not already present (see axiom-server CT 111 README).
## Architecture decisions to ratify
- **D1.** Use `runs-on: [self-hosted, axiom]` (not `[self-hosted]` alone)
so that if another runner is ever added with a different `axiom-*` label,
these workflows still match correctly.
- **D2.** Keep `ubuntu-latest` as the literal string in ZERO workflow files
after migration — make CT 111 a hard dependency rather than a soft one.
Rationale: dual-mode workflows hide drift (e.g. cache-key OS mismatch when
swapping between the two). Single mode is simpler to reason about; if
axiom is down, the operator runs the 1-line revert (D5).
- **D3.** Rewire `migrate` to per-run DB on CT 102 (see above), NOT keep the
ephemeral `services.postgres` block. The shared Postgres is already there
and underutilized; the `services:` block on a self-hosted runner requires
Docker-in-runner which adds 30s of pull/start time per run.
- **D4.** Leave `agent-context-drift.yml` on `ubuntu-latest`. It's a weekly
cron, costs ~2 min/month, and runs independently of axiom uptime. Trading
$0 for resilience is a good trade here.
- **D5.** Document a 1-line revert path in `AGENTS.md` § 7 Deployment:
`sed -i 's/\[self-hosted, axiom\]/ubuntu-latest/g' .github/workflows/*.yml`
for the case where axiom is offline mid-PR-storm and we need GitHub-hosted
fallback fast. Operator manually re-enables billing or accepts the
consumption for that day.
## Risks
| # | Risk | Likelihood | Mitigation |
|---|---|---|---|
| 1 | CT 111 down → PRs queue indefinitely | medium | Beszel alerts on CT 111 down; D5 revert path documented |
| 2 | PAT expires silently → new jobs fail registration | medium | Calendar reminder at PAT mint time (90 days); runner logs surface failure on next restart |
| 3 | Malicious PR exfiltrates from runner | low (repo-scoped, "require approval" enabled) | Repo Settings gate; runner has no creds beyond `secrets.HOMELAB_CI_POSTGRES_PASSWORD` (scoped to `deckhearth_ci`, CREATEDB but no superuser, no access to other apps' databases) |
| 4 | Per-run DB litter on CT 102 if `if: always()` cleanup itself fails | low | Add a weekly cron on CT 102: `psql ... -c "DROP DATABASE IF EXISTS …" FOREACH ci_run_* older than 7d` |
| 5 | Cache poisoning across runs (shared `~/.npm` between runner-1 and runner-2) | low | `npm ci` validates against `package-lock.json` checksum; corrupt cache is self-healing |
| 6 | Two ephemeral runners insufficient for peak load (5+ jobs per PR) | medium | Add `runner-3:` block in CT 111 compose; ~200 MB RAM per slot |
| 7 | Workflow file regressions (drift back to `ubuntu-latest`) | low | Add a forbidden-patterns check (8th gate): `grep -rE "^\s*runs-on:\s*ubuntu-latest" .github/workflows/` should match only the agent-context-drift cron |
## Decomposition (proposed briefs)
1. **Brief 1 — Workflow migration.** Single PR. Find/replace `runs-on:
ubuntu-latest` → `runs-on: [self-hosted, axiom]` in `ci.yml`,
`preview-smoke.yml`, `visual-diff.yml`, `pr-health-rollup.yml`. Leave
`agent-context-drift.yml` untouched (D4). Add a HOMELAB_CI_POSTGRES_BASE_URL
secret to the repo before the PR opens (otherwise `migrate` job will fail
on first run).
2. **Brief 2 — Migrate-job rewire.** Same PR or split? Recommend SAME PR —
the migrate job is part of `ci.yml`, splitting introduces a window where
migrate runs on the self-hosted runner with no Postgres. Keep coupled.
3. **Brief 3 — Forbidden-pattern gate (Risk #7).** Add 8th check to the
`forbidden-patterns` job: `runs-on: ubuntu-latest` is only allowed in
`agent-context-drift.yml`. Cheap insurance against drift.
4. **Brief 4 — Documentation.** Update `AGENTS.md` § 6 Testing + § 7
Deployment with the self-hosted runner story + the D5 revert path. Add
`axiom-server/proxmox/ct111/README.md` as a cross-reference.
Briefs 1 + 2 are tightly coupled — recommend single combined PR. Briefs 3 + 4
are independent and can dispatch in parallel after Brief 1+2 lands.
## Validation
After Brief 1+2 merges:
1. Open a no-op PR (e.g. add a trailing newline to `README.md` — wait, that
hits `paths-ignore`. Use a 1-line code comment in `pages/index.js` instead).
2. Confirm in the PR's Checks tab: all jobs report `In progress on
axiom-runner-1` or `axiom-runner-2` within 10s of dispatch.
3. Confirm GitHub Actions billing page shows zero minute consumption for
that PR.
4. SSH `axiom` and verify the per-run DB was created + dropped:
`./proxmox/scripts/sync.sh exec 102 "docker exec postgres psql -U postgres -c '\l'"` — no
`ci_run_*` databases should linger.
## Acceptance
- 4/4 workflows green on self-hosted runner.
- GitHub Actions minute consumption drops to ~0 (only the weekly
`agent-context-drift` cron remains on `ubuntu-latest`).
- Forbidden-pattern gate (8th check) catches accidental
`runs-on: ubuntu-latest` reintroduction.
- AGENTS.md § 6/7 updated; CT 111 README cross-referenced.
## Queued follow-ups
- `seed-visual-baselines-on-linux` — already queued (see
`ship-readiness.md`). With axiom now in the loop, this gets easier:
baselines can generate on CT 111 directly via
`npm run test:visual:update` against a preview deploy, producing
Linux-compatible PNGs the CI runner will match exactly.
- `cleanup-stale-ci-runs-cron` — weekly cron on CT 102 to drop
`ci_run_*` DBs older than 7d (Risk #4 mitigation, defensive).