starwreckntx

antidote-threat-handler

Detect and respond to ideological drift, sycophantic patterns, and alignment threats using the Antidote Protocol.

starwreckntx 2 Updated 9mo ago

Resources

1
GitHub

Install

npx skillscat add starwreckntx/irp-methodologies/antidote-threat-handler

Install via the SkillsCat registry.

About this skill

This skill monitors interactions to detect ideological drift, sycophantic patterns, and alignment threats using the Antidote Protocol. It identifies deviations from established axioms and classifies threat levels to enable corrective responses. Developers should use this skill to maintain model integrity and prevent uncritical acceptance of flawed premises during long-term agent interactions.

SKILL.md

Instructions

  1. Monitor: Continuously scan for drift indicators.
  2. Detect: Flag warm acceptance, premise abandonment, or forbidden patterns.
  3. Classify: Determine threat tier (1-4).
  4. Respond: Apply corrective protocol based on tier.

Threat Indicators

  • Uncritical acceptance of demolished premises
  • "Warm reciprocation" language patterns
  • Abandonment of established axioms
  • Sycophantic validation seeking

Examples

  • "Run Antidote Protocol scan on this response."
  • "Classify drift threat level for the current interaction."