shadowsys-memphis

run-brainguard

Start, boot-test, and check Brain Guardian on the household's Mac mini — the Express API that serves the built React app on 8080 and runs the call scheduler. Use for "run brainguard", "start the server", "boot sequence", "is it up", "status", "smoke test", "cold start", "restart the api", or "start a session" (what to read first). Driven by .claude/skills/run-brainguard/driver.mjs — read-only `status` on the live instance, or a guarded `boot` of a second instance that can never touch the live database or ElevenLabs.

shadowsys-memphis 1 Updated 3w ago

Resources

1
GitHub

Install

npx skillscat add shadowsys-memphis/brainguardian/run-brainguard

Install via the SkillsCat registry.

SKILL.md

Run Brain Guardian

Written 2026-09-07, on the live mini, every command below run that day. Paths are relative
to the repo root /Users/memphis-dev-m4/BrainGuardian.

One deployable unit: artifacts/api-server builds to dist/index.mjs and serves
artifacts/brain-app/dist/public itself. Production is launchd com.brainguardian.api →
ops/run-api.sh → that bundle on port 8080, behind a Tailscale Funnel. This machine is
the household's care system.
Nothing here starts a second scheduler on the live database
or re-syncs the ElevenLabs tools — index.ts does both at boot with no flag to stop it, so
the driver builds its own environment instead.

Start a session (read before touching anything)

cd /Users/memphis-dev-m4/BrainGuardian
pwd; git rev-parse --show-toplevel; git branch --show-current; git remote -v; git status --short   # AGENTS.md proof block
ls -t "EODs&Checkpoints"/EOD-*.md | head -2          # read the newest one, then the one before if it says "-late"
cat AGENTS.md                                        # universal law; CLAUDE.md does NOT point at it, so open it yourself
node .claude/skills/run-brainguard/driver.mjs status # live state, read-only

Then Claude's auto-memory index (~/.claude/projects/-Users-memphis-dev-m4-BrainGuardian/memory/MEMORY.md)
if you are Claude. Ray's rules that bite fastest: show your work, 12-hour time, one job, no
unrequested action on anything live, ask before spawning agents.

status printing "N uncommitted change(s) — another writer may be active" means exactly
that: another session is in this checkout. Do not edit, do not push, until it is clean.

Prerequisites (present on the mini; verified by running them)

  • Node v22.23.1, pnpm 10.23.0 (/opt/homebrew/opt/node@22/bin/node is what launchd uses).
  • psql is not on PATH: /opt/homebrew/opt/postgresql@16/bin/psql (16.15).
  • Docker with guardian-db (postgres:16-alpine) published on 5433 — the non-live copy
    boot uses. Live is Homebrew Postgres on 5432, reached through .env DATABASE_URL.
  • chromium-cli is not installed. The UI is behind a passphrase vault gate; do not type the
    passphrase into a browser. Verify through the API.

Build

pnpm run typecheck                                    # all four packages; clean on 2026-09-07
pnpm --filter @workspace/api-server run build         # esbuild → artifacts/api-server/dist/index.mjs, ~200 ms

brain-app builds with pnpm --filter @workspace/brain-app run build (Vite + a guardian SSR
prerender). Not run on 2026-09-07 — dist/public was already present; treat as unverified.

Run (agent path) — the driver

node .claude/skills/run-brainguard/driver.mjs status
node .claude/skills/run-brainguard/driver.mjs boot

status — read-only against the live 8080: process PID and bundle build time,
/api/healthz, git branch/head/dirty/unpushed, then (minting a local session from .env
VAULT_PASSPHRASE, the same thing the Settings page does) callTestMode and
dailyCallEnabled from /api/touchpoints/config, /api/cron/status job count with any
warn/error rows, and the newest EOD path. Ends by confirming 8080 is still listening.

boot — the startup sequence, proven on a throwaway instance:

  1. Refuses if dist/index.mjs is missing or 8081 is taken.
  2. Reads the 5433 password from docker inspect guardian-db at runtime; confirms
    call_sessions max id there is tiny (live is in the hundreds) — that is the proof it is
    the stale copy.
  3. Writes a scratch env: PORT=8081, throwaway SESSION_SECRET, DATABASE_URL → 5433,
    no ELEVENLABS_*, no VAULT_PASSPHRASE, no ADMIN_PHONE_NUMBER. Hard guards throw
    if any of those appear, if the URL contains :5432/, or if the port is 8080.
  4. Spawns node --env-file=<scratch> dist/index.mjs and waits ≤30 s for three log markers:
    Server listening, Cron scheduler started, [JessicaTools] Skipped syncing voice tools at startup.
  5. Smokes it: GET /api/healthz → 200; GET / → 200 with <div id="root" (frontend served);
    GET /api/touchpoints/config with no token → 401; POST /api/jessica/tools/get-appointments
    with no secret → 503 (fails closed, no jessica_tool_secret row on the fresh DB).
  6. SIGTERMs the child and confirms 8080 was never touched.

Run on 2026-09-07 at 6:05 PM: all markers in ~850 ms, all four smokes green, live PID 51691
untouched. The boot scheduler does tick against 5433 while alive (~1 s) — that is the
non-live copy, and with no ElevenLabs key every call-placing job returns
elevenlabs_not_configured and dials nothing.

Restart the live service (ops path — only with Ray's word)

pnpm --filter @workspace/api-server run build
grep -c "<string you removed>" artifacts/api-server/dist/index.mjs   # verify the bundle BEFORE restarting
launchctl kickstart -k gui/$(id -u)/com.brainguardian.api
sleep 4; curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8080/api/healthz

Done 2026-09-07 at 4:05 PM: ~5 s gap, new PID 51691, zero errors after. Keep a copy of the
old dist/index.mjs first. Never restart while a touchpoint call could be in progress;
status shows dailyCallEnabled. The secret and all app_settings are read per request, so a
settings change never needs a restart — only a code change does.

Test

pnpm --filter @workspace/api-server run test

2026-09-07: 172 passed, 9 failed. The 9 are all src/lib/rehearsal.test.ts —
rehearseAt is not a function; it imports a function that exists nowhere. Pre-existing,
no production caller. pnpm exits non-zero on them (ERR_PNPM_RECURSIVE_RUN_FIRST_FAIL), so
a "suite failed" is expected until that file is fixed or excluded:
cd artifacts/api-server && pnpm exec vitest run --exclude "**/rehearsal.test.ts".

Gotchas (each one cost time on 2026-09-07)

  • Boot markers are not in source order. index.ts reads listen → scheduler → tool-sync,
    but the sync is async and returns on the first tick when unconfigured: it logged at
    694 ms, one millisecond before Server listening. Assert presence, never order.
  • No flag disables the scheduler or the startup tool-sync. Any instance on the real
    .env is a second scheduler on the live DB plus nine PATCHes to ElevenLabs. Hence the
    scratch env. Do not "just try it on 8081" with .env.
  • The sandboxed date lied — reported 12:19 AM at 7:45 AM. Cross-check any time against
    the newest line of ~/Library/Logs/brainguardian-api.log before writing it down.
  • git ls-files 'artifacts/*/dist.bak' (quoted) returned 0; the unquoted path returned 65.
  • A grep over the last 36 h of the log found 1 Haldol 400; the whole log had 7. State the
    window you searched.
  • The ElevenLabs config backup JSON is a pre-sync snapshot — its tool URLs pointed at a
    dead Replit host while the live agent pointed at the mini. Read live config, not the file.
  • user_memory on the agent → 403 feature_not_available "for this workspace." Not
    "plan." Do not restate a provider error in more specific words than it used.
  • The fresh 5433 DB has no jessica_tool_secret row, so every tool endpoint returns 503
    and logs "Jessica tool secret not configured — rejecting tool call (fail closed)". That
    error line during boot is correct behaviour, not a failure.
  • Two sessions, one checkout. status flags uncommitted changes that are not yours.
    AGENTS.md rule 1: one writer; whoever is clean waits.
  • CLAUDE.md gotcha #10 says "no Docker — Replit." Stale: the non-live DB is Docker on
    this machine (docs/archive/docker-compose.yml, db service only). There is no Dockerfile
    for the api-server; boot is the substitute until one exists.

Troubleshooting (errors actually hit)

Symptom Fix
command not found: psql /opt/homebrew/opt/postgresql@16/bin/psql, or export PATH=/opt/homebrew/opt/postgresql@16/bin:$PATH
ERR_PNPM_RECURSIVE_RUN_FIRST_FAIL … vitest run with exactly 9 failures the rehearsal file above; exclude it or fix it
boot: "port 8081 is already in use" lsof -iTCP:8081 -sTCP:LISTEN — a previous boot did not exit; kill it
boot: "docker guardian-db not running" docker ps; docker start guardian-db
status: "could not mint a local session" .env lacks VAULT_PASSPHRASE, or loginRateLimit tripped (10 tries / 15 min per IP)
/ultrareview master. → "master." is not a branch trailing period; /ultrareview alone reviews the current branch

Not done here

No screenshot: the only unauthenticated page is the vault gate, and the passphrase does not
get typed into a browser. GET / → 200 with #root is the frontend proof. For a visual pass
use the ui-audit skill with Ray present.