- Home
- /
- Categories
- /
- Monitoring
Monitoring
Logging, metrics, and observability
data-engineering
by eyadsibai
Use when "data pipelines", "ETL", "data warehousing", "data lakes", or asking about "Airflow", "Spark", "dbt", "Snowflake", "BigQuery", "data modeling"
rviz-screenshot-loop
by Idate96
Capture RViz/GUI screenshots via MCP to close the loop while debugging ROS. Use when you need visual verification in RViz or other windows.
cloud-waste-hunter
by famaoai-creator
Actively identifies and eliminates unused or over-provisioned cloud resources. Goes beyond estimation to hunt for actual cost savings in live environments.
designing-monitoring
by sumik5
Designs monitoring and observability systems covering anti-patterns, design patterns, layer-based strategy (6 layers), alerting, on-call operations, incident management, telemetry pipeline architecture, observability concepts (structured events, core analysis loop), SLO-based reliability, sampling strategies, and observability maturity model. Use when designing monitoring strategy, setting up alerting, choosing monitoring tools, building observability systems, implementing SLOs, or adopting observability practices. For OpenTelemetry SDK/API implementation, use implementing-opentelemetry instead. For green software carbon metrics and sustainability monitoring, use building-green-software instead. For Google Cloud monitoring (Cloud Logging, Security Command Center, Cloud Run), use developing-google-cloud instead. For LLM-specific evaluation and monitoring, use practicing-llmops. For application logging design, structured logging, and log collection pipelines, use implementing-logging instead.
migration-observability
by dmonteroh
"Make database migrations safe and observable. Define progress + safety metrics, dashboards, and runbook gates (go/no-go criteria) for live migrations, backfills, and cutovers. Works standalone and is database/tooling agnostic."
monitoring-config-auditor
by famaoai-creator
status: implemented
python-dual-mode
by Victory-Hugo
将既有或新编写的 Python 脚本重构为双模式模块,用于稳定、可维护的数据处理流。
cleanproject
by manastalukdar
Remove debug artifacts and temporary files safely with git checkpoint protection
grafana-dashboards
by dmonteroh
"Provides guidance to create and manage production Grafana dashboards for real-time visualization of system and application metrics. Use when building monitoring dashboards, visualizing metrics, or creating operational observability interfaces."
kitty
by XYenon
Instructions for using kitty remote control to spawn windows/tabs, send text, inspect output, and manage processes. Useful for running servers or long-running tasks in the background.
agent-activity-monitor
by famaoai-creator
Collects and visualizes statistics regarding the agent's activities, including skill usage, execution success rates, and task duration. Provides a data-driven dashboard for ecosystem health.
root-cause-tracing
by pproenca
This skill should be used when the user asks to "find the root cause", "trace the bug", "why is this happening", "where does this come from", or when errors occur deep in the call stack. Systematically traces backward to identify the source.
loop-review-skill-until-fixed-point
by corygabrielsen
Iterate /review-skill on a target until fixed point. Runs review passes until all reviewers return NO ISSUES.
experiment-tracking
by eyadsibai
Use when "experiment tracking", "MLflow", "Weights & Biases", "wandb", "model registry", "hyperparameter logging", "ML experiments", "training metrics"
backend-principle-eng-typescript-pro-max
by PrakharMNNIT
"Principal backend engineering intelligence for TypeScript services. Actions: plan, design, build, implement, review, fix, optimize, refactor, debug, secure, scale backend code and architectures. Focus: correctness, reliability, performance, security, observability, scalability, operability, cost."
backend-principle-eng-nodejs-pro-max
by PrakharMNNIT
"Principal backend engineering intelligence for Node.js runtime systems. Actions: plan, design, build, implement, review, fix, optimize, refactor, debug, secure, scale backend code and architectures. Focus: correctness, reliability, performance, security, observability, scalability, operability, cost."
effect-vitest
by tstelzer
Testing Effect programs with vitest. Use when writing tests for effect-based code.
implementing-opentelemetry
by sumik5
OpenTelemetry implementation for distributed system observability covering instrumentation API/SDK and Collector deployment. Use when implementing tracing, metrics, or logging with OpenTelemetry. Covers Collector pipelines, semantic conventions, and organizational adoption strategies. For monitoring strategy, alerting design, telemetry pipeline architecture, observability concepts, SLOs, and sampling strategies, use designing-monitoring instead. For application-level logging design and log collection architecture beyond OTel Logs Signal, use implementing-logging.
backend-principle-eng-python-pro-max
by PrakharMNNIT
"Principal backend engineering intelligence for Python services and data systems. Actions: plan, design, build, implement, review, fix, optimize, refactor, debug, secure, scale backend code and architectures. Focus: correctness, reliability, performance, security, observability, scalability, operability, cost."
project-risks-and-changes
by piperubio
Risk & Change Management (Devil's Advocate): Identify risks, manage issues, and evaluate change requests. Use this skill to proactively detect threats, assess the impact of changes, and protect the project baseline.
chaos-engineer
by dmonteroh
"Design and run safe chaos experiments (failure injection + game days) to validate resilience and reduce blast radius. Produces hypotheses, steady-state signals, rollback gates, and experiment specs. Use when resilience is uncertain or before high-risk changes."
backend-principle-eng-java-pro-max
by PrakharMNNIT
"Principal backend engineering intelligence for Java services and distributed systems. Actions: plan, design, build, implement, review, fix, optimize, refactor, debug, secure, scale backend code and architectures. Focus: correctness, reliability, performance, security, observability, scalability, operability, cost."
manim-video-teacher
by lispking
专注于使用 Manim 生成动画教学视频的完整流程与专业建议。适用于用户用中文提示语让 Codex 生成脚本、分镜、Manim 代码、渲染命令或优化教学视频质量与节奏,并输出 MP4。
disaster-recovery
by timequity
Backup strategies, disaster recovery planning, and business continuity.