--- name: scan-visual-catalog-search classification: feature success_metric: | After catalog embeddings exist and the new kNN path is on, ≥50% of legitimate card scans (scan_attempts excluding not_a_card) auto-match a printing with no picker — the original add-real-ocr-layer 70% target, measured honestly. Gemini L2 share of attempts falls vs the Phase 1 post-ship baseline. skip: - ia - ui-design - visual - a11y - design - flag status: open created: 2026-08-14 depends_on: - improve-scan-card-detection umbrella: scanner-identify-upgrade model_policy: default_session: auto roles: role-conductor: composer-2.5-fast role-architect: composer-2.5 role-ia-architect: composer-2.5-fast role-ux-reviewer: composer-2.5-fast role-ui-designer: composer-2.5-fast role-implementer: composer-2.5-fast role-reviewer: cursor-grok-4.5-high role-security-auditor: gpt-5.6-terra-medium role-design-system-auditor: cursor-grok-4.5-high role-a11y-auditor: cursor-grok-4.5-high role-doc-writer: auto escalate_to: claude-sonnet-5-thinking-medium escalate_to_premium: claude-4.6-opus-high-thinking never_premium: - role-reviewer - role-security-auditor - role-design-system-auditor - role-a11y-auditor - role-ui-designer - role-doc-writer --- # Convoy: scan-visual-catalog-search Phase 3 of `scanner-identify-upgrade`. Identify cards by **what they look like**, not by reading the name. This is how Manabox / Delver-style scanners get printing-accurate matches on foil and alt-art. Blocked on `improve-scan-card-detection`: embeddings of unwarped phone crops will not match catalog `image_url`s reliably. ## Why Name-only `pg_trgm` cannot distinguish printings. Tesseract fails on foil and stylized type even when it emits 8–20 characters (76.7% L1 escalate with non-empty text). Gemini can read a name but still returns a picker when set/number are missing, and costs 5/min. The catalog already stores `cards.image_url` for every imported printing. Precompute an embedding per row; at scan time embed the warped crop and take top-k cosine. OCR / Gemini become hints and unknown-card fallback. ## Scope ### In scope - **Migration:** `pgvector` (confirm Neon availability) + `cards.embedding` (or a side table `card_embeddings`) + ANN index. Update `docs/SCHEMA_MAP.md`. - **Offline embed job:** new dated script or `npm run` task that reads `image_url`, writes vectors. Idempotent. Rate-limit the embedding provider. Do **not** edit historical `scripts/add-*.js`. - **Architect picks the embedder** (one): 1. Gateway embedding model (same `AI_GATEWAY_API_KEY`, server-only). 2. In-browser MobileCLIP-S2 / SigLIP ONNX for the *query* crop, with catalog vectors baked or fetched — only if weight license + size are acceptable. - **Identify path:** new Layer-0 (or replace L1) — kNN then auto-match / disambiguate / escalate to existing L2. Log `scan_attempts.layer = 0` (or Architect-ratified value). - **Similarity thresholds** analogous to 0.85 / 0.60, tuned on a held-out set of scan captures if any exist in Blob. - Auth + rate-limit on any new route. No client-side API keys (`forbidden-client-side-llm-keys` must stay green). ### Out of scope - Replacing Gemini entirely on day one — keep L2 for catalog misses and low similarity. - Training a custom card CNN. - Python GPU service. - Changing scanner chrome / cart. - Auto-approving the 27 pending `card_submissions`. ## Roles invoked 1. `role-architect` — embedder, schema, layer numbering, thresholds, brief split (migration / backfill / route). Security-sensitive: escalate to Sonnet if schema + new route land together. 2. `role-implementer`. 3. Audit: `role-reviewer` + `role-security-auditor` (required). ## Todos - [ ] Architect: confirm `pgvector` on prod Neon tier - [ ] Brief 1 — migration + SCHEMA_MAP - [ ] Brief 2 — catalog backfill job (idempotent) - [ ] Brief 3 — identify kNN route + client escalate order - [ ] Threshold bake-off on real crops - [ ] Re-measure auto-match % excluding `not_a_card` ## Likely file ownership | Area | Files | | --- | --- | | Schema | `migrations/*_card-embeddings.js`, `docs/SCHEMA_MAP.md` | | Backfill | new `scripts/` job (dated) or `lib/card-embed-backfill.js` | | Query | new `lib/card-visual-match.js`, `pages/api/scan/identify-by-image.js` **or** fold into existing identify | | Client | `lib/scanner-card-identify.js` (`tryLayer1TextIdentify` sibling) | Do not rewrite `lib/scanner-card-detection.js` here. ## Multitask dispatch Brief 1 first. Brief 2 after 1. Brief 3 after 1 (can overlap 2 if the route degrades to escalate-when-empty-index). ``` /multitask role-reviewer + role-security-auditor ``` Group id: `audit-scan-visual-catalog-search-`. ## CI impact | Workflow / job | Behavior | | --- | --- | | `schema-map-fresh` | **Fires** — migration + SCHEMA_MAP | | `ci.yml` migrate | Must apply `pgvector` on CT 102 CI Postgres — Architect must verify the extension is available there or gate the migration | | `forbidden-client-side-llm-keys` | Blocking | ## Operator action - Confirm Neon `pgvector` (or Neon’s equivalent) on the prod project. - Budget: one embedding per catalog image, plus one per live scan if the query embed is server-side. Architect publishes a cost note before Brief 2 runs against prod images. - No new browser secrets. ## Conductor notes This is the accuracy leap. Do not start it to "try CLIP" before Phase 2 crops are rectified — that wastes the backfill. If Architect finds `pgvector` unavailable on CI Postgres, stop and write a fallback (external index vs skip-CI-extension plan) rather than shipping an untestable migration.