kuhrezumaya17-spec

visual-storyboarding

Use when a user has a visual creative idea (video, short film, animation, comic storyboard) and needs to refine it into frame-level visual design before any implementation or AI generation

kuhrezumaya17-spec 0 Updated 1mo ago

Resources

1
GitHub

Install

npx skillscat add kuhrezumaya17-spec/wordstocamera

Install via the SkillsCat registry.

SKILL.md

Visual Storyboarding

Overview

Transform rough creative ideas into precise frame-level visual design documents. One question per dimension, multiple-choice when possible, lock each segment before moving forward.

When to Use

  • User has a visual creative idea that needs structural refinement
  • User needs help breaking "I have a vision" into specific frame descriptions
  • Before any AI image/video generation attempt
  • Before any actual filming or production

When NOT to use:

  • Pure text writing (novels, articles)
  • Non-visual projects (code, data analysis)
  • User already has precise, shot-by-shot specs

Core Process

Creative summary → Global setup → Act/scene breakdown → Finalize & save → Multi-tool prompts
       ↑                ↑               ↑                   ↑                  ↑
   Understand     A/B line contrast  One question      Full output      Adapt per platform
   intent first   Space + light      per message       locked & saved    + failure warnings
                  Sound principles   Multiple-choice   Ready for tools

Key Principles

One Question, One Dimension

Never stack questions. Each message asks about exactly ONE visual dimension. Wait for answer before next question.

Multiple-Choice Preferred

Present 2-3 concrete options with trade-offs. Always include a "You describe it" escape hatch. Lead with recommendation and explain why.

Explain What the User Doesn't Know

When a user says "I don't understand lighting," don't skip it — explain each option in human terms:

  • What it looks like
  • What the audience will feel
  • Which serves their story best
  • Why you recommend one

Anti-AI-Aesthetic Early Warning

If the user's visual style is deliberately "imperfect" (low-res / grainy / mockumentary / rough), warn early: most AI image tools default to clean and beautiful, and fighting that bias is expensive. Don't wait until prompt-writing to discover this.

Lock Then Move

Each segment confirmed → locked. Never revisit locked segments unless user explicitly asks. Prevents infinite circling.

Don't Re-Ask What's Already Said

If the user's original description already implies a dimension's answer, don't re-ask. Instead, confirm in one sentence:

"你说了'像要拥抱'——毛巾横向展开像伸开双臂,对吧?"

Only go into detailed follow-up questions when the original description isn't sufficient to generate a good prompt, or when the dimension has multiple valid interpretations that even the user may not have considered (e.g., lighting direction, color tone).

Per-Scene Checklist

For every scene, work through these dimensions in order (one at a time):

  1. Space — What kind of room? How big? How many people? What's the atmosphere?
  2. Camera — Where is it placed? Does it move? What device type? Why that choice?
  3. Foreground / Background — What blocks the lens? What's in the distance? What's out of focus?
  4. Subject — Who? What are they doing? What expression? Is this about action or stillness?
  5. Light — Where from? Hard or soft? What gets lit, what stays dark? Why that direction?
  6. Color & Texture — Color temperature? Image quality (clean vs grainy vs distorted)? Any special treatment?
  7. Sound — Ambient noise? Action sounds? Breathing? Lines? What's absent (no music, no voiceover)?

Global Setup Before Scene Work

Before asking about individual scenes, establish:

Item Why
Two-line structure (A/B, inside/outside, past/present) Prevents everything blending into same texture
Space contrast between lines "Dark grid office vs quiet room" — locks visual contrast early
Light source per line Single source per line = realistic, controllable
Sound rules No music? No voiceover? All rough? — sets texture baseline

Design Document Format

Save completed design to docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md with:

  • Global settings table
  • Per-act/scene tables with all dimensions
  • All confirmed dialogue/quotes
  • Status markers (✅ confirmed / ⏳ in progress)

Common Mistakes

Mistake Fix
Writing camera angles as prompt language ("low angle, dutch tilt") Write what the VIEWER sees, not what the camera does
Asking lighting preferences without explaining what each option means Explain visually in human terms, then ask
Letting user rush through multiple dimensions at once Hold the line: one at a time
Not warning about AI limitations until prompt-writing stage Flag early when visual style conflicts with AI defaults
Revisiting locked decisions without user permission Only reopen when user explicitly asks

Red Flags — Slow Down

  • User jumps ahead: "Let's just do all four scenes at once"
  • User skips a dimension: "Light doesn't matter, let's move on"
  • User wants to write prompts before design is complete
  • "This is simple enough, we don't need to go through every dimension"

Each of these means: Hold position. Complete current dimension first.

Final Step: Multi-Tool Prompt Generation

After the design is fully confirmed and saved, ask:

"设计定稿了。要不要我分别生成不同 AI 工具的提示词版本?比如小云雀、LiblibAI、Midjourney、DALL·E 等——每个工具写法不同,我可以针对各平台写适配版。"

Tool-Specific Adaptation Rules

Tool Style Notes
小云雀/国产 自然中文段落描述 中文零损耗,可写长段落
Midjourney 英文 + 关键短语,用 -- 参数控制风格 需英文,用逗号分隔关键词而非完整句子
DALL·E 英文自然语言,精炼 对"监控""偷拍"等词敏感,可能触发审核
LiblibAI / SD 中文 + 可选模型/LoRA/ControlNet 可指定纪实/胶片模型,控图最强

Anti-Failure Warnings

When writing prompts, explicitly warn about known failure modes of each tool based on the visual style:

  • If the design uses low-res / grainy / mockumentary visual: warn that AI defaults toward clean/beautiful
  • If the design uses unusual camera angle (chest-forward not selfie): warn about common perspective errors
  • If the design involves specific facial expressions: warn about consistency across multiple generations

Categories