Develop cumulative HMASD research understanding - update judgments from evidence, use simple-model prototypes and select a discriminating next action with independent scientific review under constitution section 5; add Pro when it offers distinct value. Declare fit cost and read confirmation by its fixed rule in NOTES.md.
Resources
2Install
npx skillscat add cartmanfatass/my-paper-code/claude-skills-hmasd-scientific-tools Install via the SkillsCat registry.
Research method
Authority: docs/project/OPERATING_CONSTITUTION.md sections 3, 4, 5 and 8. This skill is the
method; it adds no rule. Records are the notebook, the runs folder and the claim note.
Mathematics, conjectures and experiments
Owner clarification, 2026-09-25: use mathematical, information-theoretic and game-theoretic
reasoning to choose questions and competing explanations, together with experiments and
explicit conjectures. Read the relevant structural background in RESEARCH.md. Identify which
information, joint-action/partner coupling, policy restriction, finite learning or deployment
distribution a proposed intervention changes. A familiar mathematical term or a proxy trend
does not supply that connection. These are reasoning aids in the existing NOTES, not additional
forms or a requirement to formalize every axis before acting.
MARL often advances from a bandit, single-agent result, simple game or empirical regularity
whose extension is not yet proved. Such a bridge may directly motivate a modest complete
learning experiment: state the useful mechanism, the assumptions that carry over, the coupling
that does not, and the prediction being conjectured. A rigorous theorem, exact headroom,
identifiability proof, complete simulator solution or positive simple-host result is not an
entry requirement. Use a derivation, approximation or counterexample when it changes a
comparator, prediction or reading; do not manufacture a theorem to justify ordinary exploration.
Choose verification in proportion to the decision and its full cost. Single-step counterfactual
branches, exact suffix replay, exhaustive joint-action enumeration and repeated mechanism
screens are optional studies, not default gates before native learning. Prefer a small direct
complete experiment when it answers the scientific question more efficiently. A bounded
diagnostic needs a concrete interpretation/use that warrants its marginal work; zero new fits
does not erase evaluation, branching, engineering or readback cost. Necessary correctness and
information-leakage checks remain. Exploratory evidence may support a conjecture without
proving its mechanism; retain that uncertainty instead of claiming proof or forbidding the next
useful experiment. Confirmation still follows the fixed scientific minimums below.
Explore an idea
- Start from the direction's current explanation and the observation or gap motivating this
work. For a new question or material hypothesis/comparator/investment change, consult the
relevant shared background from current published main, not just the direction branch's copy.
InNOTES.md, link the topic/revision and how it changes the comparison, prediction or choice,
or explain why the scope differs. Link the relevant prior interpretation and contrary evidence. State the
question, MARL structure, strongest simpler explanation and discriminating observation.
For a targeted change predict both an intermediate effect and its native consequence;
a package screen can instead explicitly forgo mechanism attribution. Declare the arms,
training horizon, planned fits with their scientific reason and expected sign before running. - Run on any host that can show the effect; single seed is fine. Commit first, preflight,
launch detached at that sha (engineering skill). The runner writesruns/<direction>/<tag>/. - Read curves and
summary.jsondirectly. Separate technical execution facts, observations
and interpretation. Update what is strengthened, weakened, untouched or unresolved before
choosing the next action, including the shared judgment used in the design. Publish a useful
shared change through the normal result-publication method. Exploratory conclusions stay exploratory: no effect claim from
one seed, no MEI verdict. Do not force a new insight from an uninformative result. - Choose inspection, diagnosis, replication, targeted revision, a different hypothesis or pivot
for the judgment/use it can change, not a quota of new candidates. After failure or recipe
closure, consider relevant evidence and remaining opportunities across the project before
declaring the DM's work finished. Under constitution section 2 the DM may revise or change
its direction and register worthwhile unowned work without renewed approval. A new prospective study
may continue the same scientific question; explain its new information value and count its
exposure. A killed idea reopens only for a recorded new reason. Do not extend a batch after
seeing its scores or rename a failed idea to reset its fits.
If no worthwhile feasible next action remains, explain that conclusion in the existing
notebook. Neither endless rescue nor a positive result is required; a failed recipe alone
is insufficient reason to stop scientific work.
An assigned question can outlive its current direction name or recipe. Use the same notebook
reasoning to compare useful related continuations, with one result-bearing study active at a
time. Distinguish the tested approach's outcome, the remaining parent question and the chosen
next observation; these need not share a completion state. A scientific split needs a distinct
question/comparator/estimand, while a merger needs these and the next step to align. Neither
creates a new record type, requires an exhaustive candidate search or licenses taking another
lead's work. Five runtime slots limit concurrent execution, not the research programme's scope
(owner amendment 2026-09-27 UTC); actual node admission may permit fewer operations.
For mechanism questions trace environment event -> entity ownership -> available information
-> action/credit -> learning -> native consequence. For changing rosters distinguish entity
from slot, join/leave/rejoin, survivor history, censoring and partner co-adaptation; distinguish
primitive time from decision opportunities. Use the distinctions relevant to the proposed effect,
not a form to fill for every run. A plausible heuristic or suboptimal scheme is a legitimate
empirical candidate; exploration needs no general optimality, invariance or convergence proof.
For learning observations check that the actual environment, policy, learner/trainer and
evaluator ran: read transition, optimizer-update and evaluation counts and learner movement
from the run summary. Claimed training needs actual updates; recurrent-state evolution alone
is not parameter learning. Fixed-policy evaluation reports zero new updates and its conditional
scope, never new learning. Use an informative horizon, without a prerequisite learnability run.
Update the working explanation
Independent scientific diagnosis and direction correction belong to the Scientific Reviewer
(existing ResearchCritic) under constitution section 2. DM still formulates hypotheses and
maintains the explanation, but self-reflection is not the independent check. Use the dedicated
role body with separate context: Codex children use fork_turns="none"; other runtimes use
their actual history-isolated child mechanism. A role-like task title or read-only permission
alone does not establish context independence. Pass the question and original evidence first,
then the proponents' interpretations; the reviewer reconstructs its reading before considering
DM/Root preferences. It may disagree with both. Disclose any inherited conversation or missing
source rather than claiming independence from the role name.
Use this review at question/approach selection and material interpretation or route correction,
not each completed cell or unchanged batch. Reuse applicable independent analysis and combine
overlapping claim/direction criticism in the same pass. Return through the actual parent;
preserve the substantive recommendation, material dissent and disposition in existing NOTES
or the assigned RESEARCH review. Uncontested in-scope recommendations need no Root ACK;
Root resolves material direction disagreements within its assigned coordination; between the
Claude DM and Root (peers, owner 2026-09-27) the owner or an independent scientific review
resolves them; the owner does so in a direct DM's own task when no Root is assigned. No
automatic App relay. Do not silently treat an objection as resolved because DM/Root prefers
its earlier explanation. Preserve accepted-operation collection and
unrelated authorized work while the disputed new effect is unresolved. One adequate independent
scientific review covers an ordinary consequential decision under section 5; add Pro for distinct
expertise, framing or unresolved disagreement. Model agreement does not establish empirical truth.
In the existing notebook, connect the prior judgment and prediction to the observation and
the resulting interpretation. Use only distinctions relevant to the question: task opportunity,
representation, finite learnability, and complete-package benefit/cost are not interchangeable.
Identify which competing explanations actually predicted different observations. A negative
package comparison need not identify a bad component; nonactivation and technical failure
cannot count as evidence of an active mechanism's adverse effect.
A working update may be qualitative and conditional. Missing population precision does not
forbid learning from the result, but does forbid fabricated confidence or a stable ranking.
Keep positive and negative evidence, distinguish newly suggested explanations from pre-result
predictions, and revise interpretations by appending rather than rewriting prior entries.
No useful discrimination is a valid conclusion. Do not repeatedly list "optimization, capacity,
seed" as equally surviving excuses without asking what could weaken each explanation.
Prefer a targeted revision when evidence points to a modifiable link and the revision makes a
different prediction. Lower expectations when tested repairs fail their intermediate predictions,
the proposed bottleneck is not material in the target conditions, or remaining rescue stories
make no different feasible prediction. These are research judgments, not automatic failure-count
gates. Stopping because the next information is not worth its cost is distinct from falsification.
An unchanged replication is useful when recurrence itself changes a decision; a new architecture
is not a prerequisite. A cheap direct learner test may beat an elaborate diagnostic.
Use and revise shared understanding
Make the relevant knowledge do work in the existing NOTES reasoning: a competent ordinary
comparator, a corrected information/credit contract, a different prediction, reuse of an asset
or declining a redundant experiment. A citation or "read RESEARCH" alone does not show use.
Check current published main at the material decision boundary; reuse applicable reading within
an unchanged study. Read changed relevant passages, not the whole archive or a per-batch syllabus.
If current sources are unavailable, state the known revision and decision-relevant gap rather than
claim freshness or force unrelated work to stop. Shared knowledge is revisable: a scope mismatch
or contrary prediction can justify a new question under the existing scientific decision rules.
After reading a result, revise the affected shared topic if another direction could use the
updated conclusion, limitation, counterexample or competent method. Publish with the direction
result, linking the underlying NOTES/claim/run evidence and keeping support and adverse evidence.
Use scope-qualified language; an exploratory pattern stays exploratory and one host is not UAV
generalisation. Merge with the existing explanation, not a dated result paragraph. Purely local
details and duplicate observations stay in NOTES; no new insight or shared edit is compulsory.
If evidence conflicts, retain the differing conditions and unresolved issue rather than force
agreement or erase another direction's evidence. No new ledger, citation quota or approval step.
Simple-model and literature bridges
Use a bandit, single-agent MDP/POMDP or small joint-action game when it clarifies the disputed
link. Map variables, information rights, objective, intervention and prediction to MARL; name
what the simplification removes, such as endogenous teammate learning, decentralized information,
joint credit or asynchronous commitments. Derive or inspect a counterexample where useful.
A prototype, proof or literature search is not a required preliminary stage. Toy success does
not establish MARL/UAV benefit, and toy failure constrains only assumptions actually shared.
Verify the relevant primary passage; separate its result from our analogy and proposed design.
A return to an older branch inherits its adverse evidence and selection history. Neither model
agreement nor a literature analogy supplies new empirical replication. Examples and selectively
borrowed agent-project ideas are in cumulative-research.md;
read it for a concrete reasoning need, not as a mandatory preload.
Confirm a claim
Write CLAIM_<slug>.md before the confirmation batch: hypothesis, candidate arm, the one
primary matched-information baseline, task population, fresh independent training seeds per arm
(at least three for an empirical learning claim; choose the count for the planned precision),
training horizon, endpoint and evaluation budget/protocol, checkpoint selection and stopping rule,
selection and tuning exposure with development separated from final evaluation, decision rule,
uncertainty method, and what each outcome branch means. Run the batch once. Append the result
read by that rule with per-seed values; never rewrite the plan. Inconclusive is a legitimate
end; non-significance is not equivalence; a wide interval is not zero effect.
Comparators and MARL information
State for every arm: actor and critic information, refresh cadence, bandwidth and
representation, communication, action constraints, reward, termination and truncation,
normalisation, recurrent reset, training and update budgets, tuning rights and evaluation
selection. Local-actor MAPPO, central-input flat, fixed-clock or interruption ablations and
privileged uppers are different comparators; the same exogenous information is not the same
representation, bandwidth, optimisation difficulty or compute. A package gain needs an
identifying control before it is attributed to a component; keep native losses beside proxy
gains. Reuse docs/research/baselines/<host>/ and experiments/baselines/<host>/ when the
configuration, information conditions and exposure match; state mismatches. K-axis mechanism
questions, jointly trained N-axis churn, train-N to test-N transfer and open ad hoc teamwork are
distinct targets; do not merge K and N into one programme.
When discussing headroom, name the upper reference and the tuned same-information generic
baseline whose gap it measures; an arbitrary favorable score gap is not headroom. A privileged
upper does not establish an achievable gain. Missing tuned headroom is no prerequisite to exploration.
Statistics
Independent training runs are the inference unit; episodes and checkpoints are nested
observations, not more n. Bootstrap, more episodes or a larger effect cannot fix n = 1. Pair
only when a shared exogenous design and real independent units justify it, never because seed
numbers match. Never fill a missing pair with zero or assume missingness is random. Keep every
run and curve. Outcome-informed redesign is a new exploration, not a fresh confirmation of the
old rule. Report signed effects, per-seed values, the estimand and the small-sample limits.
Separate practical effect importance, training-outcome variation and estimator uncertainty.
For confirmation explain what difference would matter for the task, without a mandatory MEI
verdict for exploration. An equivalence claim needs a prospectively defined equivalence region
and an uncertainty interval sufficiently narrow for that claim; a small point estimate or a
one/two-seed-SD rule does not establish equivalence or importance. State the interval assumptions.
Cost and exposure
Count fits as arms times seeds per launched attempt at a declared horizon; different horizons
are different compute. Record actual wall time per fit and batch elapsed; unknown time is not
zero. Report selection and tuning exposure alongside any comparison. Prefer the smallest real
learning comparison that decides the question; exhaustive diagnosis, exact maxima and
search-before-learning need a concrete purpose.
At design time count the dominant work from the configuration: arms, fits, steps, evaluation
panels/checkpoints, optimizer epochs and nested candidate/trajectory/solver calls. Separate
algorithm-intrinsic search from verification added to study it. Joint-action or trajectory
branching, subsets and repeated replanning can dominate even a finite, bounded or zero-fit
study; reconsider unnecessary dimensions before accelerating them. Use known counts and
existing measurements; unknown cost stays unknown and creates no mandatory profiling run.
Performance claims account for full work: import/build/init, rollout and learning, replay,
evaluation, synchronization, publication and readback. Separate cold/warm runs, preparation,
queue/support, sum of fit walls, batch elapsed and actual node occupancy. Report user/system
CPU and children without double-counting threads when parallelism matters; name internal
BLAS/OpenMP/native teams and RSS scope. Never omit scientific work to claim speed; unknown
cost is not zero. These are claim-specific measurements, not a mandatory profiling fit.
Transfer claims state the held-out task/population, perturbations and aggregation actually
tested. Simulator results do not establish physical deployment safety. Exact theorem claims
need assumptions matching the implemented scheme. These limits do not create evidence
classes, a C-consumption ladder or a universal held-out requirement for exploration.
Pro
Apply constitution section 5 proactively before establishing or materially changing the question,
core hypothesis or key comparator; changing a failure explanation or continuing investment after
intermediate predictions keep failing; closing/reopening a research route or broadening a claim;
and confirmation. One adequate independent scientific review in a separate context covers an
ordinary consequential decision; confirmation still receives scrutiny of its actual claim and
fixed design. Add Pro when it offers distinct expertise, framing or unresolved-disagreement value.
Do not wait for the owner to request scientific criticism. These are decision points, not a fixed
failure count or a consultation after every result.
Identify the choice that advice can change. Check the relevant prior Pro answer: if it already
covers that choice and its evidence and premises remain materially applicable, reuse it in the
normal notebook reasoning. Confirmation reuse must cover the actual claim, comparison and fixed
plan. Routine implementation, planned verification, execution and collection under the same
reasoning need no repeat. Materially changed questions, premises or evidence at these points
need focused scientific review. When Pro adds distinct value, use hmasd-pro-research-prompt-author
for synthesis, failure explanation, a simple-model/source bridge, targeted revision, hypothesis
search or criticism as appropriate.
Read the whole answer; in NOTES.md record what you adopt, modify or reject, which judgment
changes and why. Engineering review remains separate from scientific review; a completed scientific
review does not automatically require an additional Pro pass. Direction correction and material disagreements follow section 2; other in-scope
choices remain with DM. Adviser agreement and a fixed idea count are not required.
The DM can complete the authorized Pro browser workflow without Root forwarding
or a per-question owner approval. Continue work independent of the pending scientific decision.
Existing frozen review exceptions remain tied to their original object, not expanded by this method.
Tools, only as needed
Shared-background use follows the decision steps above. For unresolved conceptual detail use
scientific-reading.md for the relevant topic and further sources.
For a literature gap use local-literature.md and verify primary
passages. For baseline or environment integration use adapters.md.
For an endpoint CSV task,seed,arm,score run scripts/summarize_runs.py (one score per training
run; --paired --baseline <arm> only for justified pairing). Use NumPy for known counts rather
than simulation. Optional packages go in isolated environments, never the live interpreters.