* convoy: forbidden-pattern gate + docs (briefs 3+4)
Closes out the migrate-ci-to-self-hosted convoy with the two
defensive follow-ups Brief 1+2 (PR #132) intentionally deferred.
Brief 3 — Check 8 of `forbidden-patterns` in ci.yml. Greps
`.github/workflows/` for `runs-on: ubuntu-latest` and fails unless
the match is in the documented allowlist (currently
`agent-context-drift.yml` only, per Decision D4). Self-tested
locally against the post-migration tree: 0 violations. Renames the
job from "Forbidden patterns (7 checks)" → "(8 checks)" and
normalizes the older "Check N/6" labels to "N/8" for consistency
(the inherited mix of `/6` and `/7` was a known cosmetic from the
unify-glass-panel-surfaces convoy).
Brief 4 — AGENTS.md § 6 and § 7 updates:
- § 6 "CI behavior": Playwright smoke runtime range updated to
cover post-migration cold vs. warm cache (was a stale 59s figure
from pre-migration ubuntu-latest).
- § 6 new top-level bullet "Self-hosted runner pool" alongside
"CI minute optimizations" — covers where runners live, where
caches are bind-mounted on CT 111, the Postgres rewire on
CT 102, and the agent-context-drift.yml exemption + how Check 8
enforces it.
- § 7 new bullet for the operational story: PAT rotation cadence
+ the D5 one-line `sed` revert path for when axiom is offline
mid-PR-storm. Cross-references the axiom-server CT 111 README
and the Beszel down alert.
Convoy doc — status flipped queued → shipped, shipped_in lists
both PRs, and follow_ups makes the cleanup-stale-ci-runs-cron +
seed-visual-baselines-on-linux items machine-greppable.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore: drop placeholder comment now PR #133 number is known
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* convoy: migrate CI to self-hosted axiom runners (briefs 1+2)
Moves 4 of 5 GitHub Actions workflows from `ubuntu-latest` to the new
`stwl-labs` org-level self-hosted pool (CT 111 axiom-runner-1..4) and
rewires the `migrate` job to use CT 102's shared Postgres via per-run
databases.
Changes:
- ci.yml: lint, schema-map-fresh, forbidden-patterns, migrate, test ->
`[self-hosted, axiom]`. migrate job drops `services.postgres` (saved
~30s/run of image pull) and switches to `HOMELAB_CI_POSTGRES_BASE_URL`
secret + per-run DB (`ci_run_<run_id>_<run_attempt>`) with `always()`
cleanup so failed migrations don't leak DBs.
- preview-smoke.yml: gate + smoke -> self-hosted. Playwright browser
cache lives under /opt/appdata/gha-runner/shared-cache/playwright on
the host bind mount; first PR primes it, subsequent runs reuse.
- visual-diff.yml: gate + visual -> self-hosted (same Playwright cache).
- pr-health-rollup.yml: rollup -> self-hosted.
- agent-context-drift.yml: deliberately LEFT on ubuntu-latest (D4 in
convoy doc). Weekly cron stays GitHub-hosted so it runs even when
axiom is down.
Why on this side and not the runner side:
- migrate adds an explicit `sudo apt-get install -y postgresql-client`
step (~10s, amortized via apt-cache survival). The runner image
doesn't ship psql; baking it in would require a custom image and
doesn't earn its keep for one job.
Repo prereqs (set before this PR opens):
- `HOMELAB_CI_POSTGRES_BASE_URL` repo secret set (value pattern:
`postgres://deckhearth_ci:<pw>@192.168.68.102:5432`)
- `deckhearth_ci` Postgres user created on CT 102 with CREATEDB,
no superuser
- stwl-labs org Actions settings: "Require approval for all outside
collaborators" + runner group rejects public repos
- 4 runners online: `axiom-runner-1..4`, status Idle
Follow-ups (per convoy):
- Brief 3: forbidden-pattern gate to catch `runs-on: ubuntu-latest`
re-introduction outside the agent-context-drift allowlist
- Brief 4: AGENTS.md updates + 1-line revert path (D5)
- Weekly cron on CT 102 to GC any `ci_run_*` DBs older than 7d
(Risk #4 mitigation)
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(migrate): use PGHOST/PGUSER/PGPASSWORD instead of URL secret
First Brief 1+2 validation run failed on the migrate job with
`psql: invalid option -- '/'` despite the secret being set correctly
and a direct CT-111 → CT-102 psql connection working fine. The
URL-parse path in `psql "$PGBASE/postgres"` was the fragile bit.
Splitting the connection into discrete `PG*` env vars (which psql
picks up automatically) sidesteps URL parsing entirely. The
`HOMELAB_CI_POSTGRES_BASE_URL` repo secret is now
`HOMELAB_CI_POSTGRES_PASSWORD` — password only — and the workflow
hardcodes the (non-sensitive) host/port/user. `node-pg-migrate`
still reads `POSTGRES_URL` from `.env.local`, so we assemble that
URL inline for it; the runner is ephemeral so the leaked-to-disk
password is bounded to one job.
Convoy doc updated to reflect the shipped approach + lesson learned
in prerequisites.
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: trigger vercel preview after stwl-labs reauth
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: re-test playwright after vercel project rebind
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>