Use when the user says: "doesn't work", "still broken", "hung/frozen", "disappeared after update", "used to work", "metric is zero but UI OK" — or when you need to investigate an incident, find the hang root cause, verify a process was actually restarted. Covers: facts before theories (storage/logs/PID), symptom vs root cause, silent failure by metrics, single consumer, restart ritual, hang localization via progress marker, timeouts, cache masking. Do not use for test writing (testing-discipline) and refactoring (agent-refactor-safety).
Resources
1Install
npx skillscat add youlianvr/oper-share/debug-incident-protocol Install via the SkillsCat registry.
SKILL.md
Debug & incident protocol: facts before theories, hangs
Distillation of debug and incident sessions. Source:
the GLAV-PATTERNS document, blocks E (38-44), L.9-15, M.34/56;UNIVERSAL-PATTERNS.md — "Hang Protection."
1. Core: facts, not opinions
- FACTS BEFORE THEORIES — order: 1) config flags, 2) DB row, 3) log lines, 4) only then code-hypothesis. First incident report contains a fact, not an opinion.
- SYMPTOM ≠ ROOT CAUSE — "minutes not deducting" — symptom; silent except + wrong parser — root cause. Trace call path to side-effect; fix root ONCE. One fix closes all surfaces.
- IF METRIC FLAT WHILE FEATURE "WORKS" — SILENT FAILURE — UX success + zero metric = error swallowing. Look for except/early return on the metric path.
- SINGLE CONSUMER FOR EXCLUSIVE STREAMS — long-poll/queue/lock file — one owner. Kill duplicates before start; health = exactly one PID.
- RESTART RITUAL IS PART OF THE FIX — code on disk ≠ code in memory. After fix: stop all → start one → verify log. Check: PID creation time > edit time.
- ENCODING OF CONSOLE ≠ ENCODING OF PRODUCT — mojibake in Windows console doesn't mean corrupted data. Check UTF-8 in client/file; don't "fix" data because of cp1251.
- INCIDENT CHECKLIST TEMPLATE — recurring incidents → checklist, not heroics: flags? storage? logs? single instance? money path except? duration source? Checklist lives next to runbook.
2. Hang localization (UNIVERSAL-PATTERNS)
- "Can't hang" — not an argument. Anything can hang: network without timeout, Read-Host, interactive prompt, infinite loop, GUI wrapper. Only proven by running with a timeout.
- Localize via progress marker, not last line — marker ("Describing X", "=== stage N ===") printed at START of block → hung INSIDE last block with marker. No markers — add them.
- Progress — to FILE, not pipe —
cmd 2>&1 | Out-File prog.txt; pipe to tail can itself hang or lose tail on timeout kill. - Isolate suspect BEFORE full run — one block with small timeout (30-60s), not the full set with 900s.
- Tool API first —
--help/Get-Help one check; two failed attempts with guessed params = minus two timeouts. - Chain A && B && C masks hang location — separate stages: each command alone, its own timeout, its own marker.
- First suspect — infrastructure, not logic — BeforeAll/dot-source/modules/network/interactive are more guilty than "instant" tests.
- Compare with last successful run — git diff/file list: culprit is usually in the changes.
- Timeout — always and progressive — small for suspect, increase only if operation is legitimately long.
3. Processes (Windows/Linux)
- PID TRACE CHAIN — tree: ParentProcessId → launcher → working directory. Don't look at PID without understanding who launched whom.
- LOG TIMELINE CORRELATION — log time vs file LastWriteTime: log older than edit = process not restarted.
- STOP-ALL-THEN-START-ONE — kill all old processes before starting new (two processes with one token = Conflict).
- PROCESS CREATION TIME CHECK — CreationDate < edit time = process on old code.
- POST-RESTART LOG TIMESTAMP FILTER — after restart, read only lines after start time (old Conflict in log ≠ current problem).
4. MCP failed at startup (disk-mount race) — new in 2.7
- Symptom: at harness start ALL MCP servers with paths on an external disk fail AT ONCE: python servers — "MCP error -32000: Connection closed" (process died immediately: script file not found), scripts — "ENOENT posix_spawn"; only servers independent of the disk survive (binary on the system disk, root in an env variable).
- Root: the harness started before the OS mounted the external disk (udisks2 mounts lazily at login; harness autostart wins the race).
- Check: harness process start time (
ps -o lstart) vs mount time (journalctl -k | grep mounted— "EXT4-fs (sdb1): mounted"). The harness log may be in UTC — account for the offset when correlating with local time. - Before diagnosing: run the server BY HAND — if it answers MCP initialize, the server is alive; the problem is the harness start environment, not the server. All servers failing at once + working individually = infrastructure, not server code.
- Fix now: restart the harness service — the disk is already mounted, everything comes up (opencode:
opencode2 service restart, or close/reopen the TUI; the harness has no MCP retry — the failed status sticks). - Prevention (industry): an fstab entry with
nofail— the external disk mounts at boot, BEFORE login (udisks2 not involved); or systemdRequiresMountsFor=<path>on the harness unit. Gotcha (verified when installing the fix): a mount unit from fstab enters local-fs.target, andsystemd-tmpfiles-setup(which creates/run/media) runsAfter=local-fs.target— the per-user subdirectory (/run/media/<user>) is created by udisks2 only at login, and systemd creates only the last path component. Fix: a drop-in with[Mount] ExecStartPre=/usr/bin/mkdir -p /run/media/<user>.
Checklist
- facts collected (config flags → DB row → log lines → code hypothesis)
- process really restarted (PID creation > edit time; fresh log lines after restart)
- cache excluded (bytes actually served)
- disk-mount race excluded (harness start time vs disk mount time)
- output: fact + root cause + one fix + smoke