pedrohcgs

qa-quarto

Adversarial Quarto-vs-Beamer parity QA. A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixer agent applies fixes; loops until APPROVED (max 5 rounds). Use when user says "qa the quarto", "check parity", "does the html match the pdf?", "quarto matches beamer?", or after a translate-to-quarto run. Requires both the `.qmd` rendered and a `.pdf` benchmark.

pedrohcgs 1,544 2,983 Updated 2w ago
GitHub

Install

npx skillscat add pedrohcgs/claude-code-my-workflow/qa-quarto

Install via the SkillsCat registry.

About this skill

This skill performs adversarial quality assurance between Quarto HTML slides and a Beamer PDF benchmark by looping a critic agent that audits parity against a fixer agent that applies corrections, repeating until approved or five rounds complete. It solves the problem of catching content loss, visual regression, and notation drift when translating Beamer lectures to Quarto. Use it whenever a Quarto render needs verification against its original PDF.

SKILL.md

Adversarial Quarto vs Beamer QA Workflow

Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.

Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.


Workflow

Phase 0: Pre-flight → Phase 1: Critic audit → Phase 2: Fixer → Phase 3: Re-audit → Loop until APPROVED (max 5 rounds)

Hard Gates (Non-Negotiable)

Gate Condition
Overflow NO content cut off
Plot Quality Interactive charts >= static plots
Content Parity No missing slides/equations/text
Visual Regression Quarto >= Beamer in all dimensions
Slide Centering Content centered, no jumping
Notation Fidelity All math verbatim from Beamer

Phase 0: Pre-flight

  1. Locate Beamer (.tex/.pdf) and Quarto (.qmd/.html) files
  2. Check freshness (re-render if QMD newer than HTML)
  3. Verify TikZ SVGs if applicable

Phase 1: Initial Audit

Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to quality_reports/[Lecture]_qa_critic_round1.md.

Phase 2: Fix Cycle

If not APPROVED, launch quarto-fixer agent to apply fixes (Critical → Major → Minor), re-render, and verify.

Phase 3: Re-Audit

Re-launch critic to verify fixes. Loop back to Phase 2 if needed.

Iteration Limits — loop-until-dry

This is the loop-until-dry primitive from `orchestrator-protocol.md`: the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per `orchestration-schemas.md`); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.

  • Fallback cap: 5 rounds bounds a non-converging loop, then escalate to the user with remaining issues.
  • Two-strikes: the same gate failing in rounds N and N+2 is flagged for the user, not patched again (`summary-parity.md`).
  • APPROVED iff every hard gate passes (zero CRITICAL).

Final Report

Save to quality_reports/[Lecture]_qa_final.md with hard gate status, iteration summary, and remaining issues.

Findings are validated, not just written (v2.5)

This skill's reviewers emit findings under the machine-checked contract in
`finding-schema.json`. Reports are JSON arrays.

Smoke-test the harness before spending review effort — a run that fans out reviewers and
then cannot write a valid report has wasted the whole pass:

echo '[]' | python3 scripts/validate-findings.py

Then, before presenting any summary:

python3 scripts/validate-findings.py <report>.json   # exit 0 required

What the contract forces, and why:

  • rule — the documented rule or standard violated. A finding citing no rule is an
    opinion, and opinions do not gate a commit.
  • failing_case — a concrete configuration under which the claim breaks, or the exact
    missing hypothesis. "This could be clearer" does not validate.
  • id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds is
    exact and the two-strikes rule is checkable rather than eyeballed.
  • mechanicaltrue only for fixes that cannot change a result (typo, cross-reference,
    formatting, label). Never for an estimand, assumption, specification, inference
    procedure, sample definition, or reporting language: those return to the researcher.

Apply the per-lens evidence burdens and the "does NOT count" filters in
`orchestration-schemas.md` §7 before
verification, so known false alarms never reach the judge. The verifier pass is
refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is
dropped, not downgraded to a warning.