Use when a user has a visual creative idea (video, short film, animation, comic storyboard) and needs to refine it into frame-level visual design before any implementation or AI generation
Resources
1Install
npx skillscat add kuhrezumaya17-spec/wordstocamera Install via the SkillsCat registry.
Visual Storyboarding
Overview
Transform rough creative ideas into precise frame-level visual design documents. One question per dimension, multiple-choice when possible, lock each segment before moving forward.
When to Use
- User has a visual creative idea that needs structural refinement
- User needs help breaking "I have a vision" into specific frame descriptions
- Before any AI image/video generation attempt
- Before any actual filming or production
When NOT to use:
- Pure text writing (novels, articles)
- Non-visual projects (code, data analysis)
- User already has precise, shot-by-shot specs
Core Process
Creative summary → Global setup → Act/scene breakdown → Finalize & save → Multi-tool prompts
↑ ↑ ↑ ↑ ↑
Understand A/B line contrast One question Full output Adapt per platform
intent first Space + light per message locked & saved + failure warnings
Sound principles Multiple-choice Ready for toolsKey Principles
One Question, One Dimension
Never stack questions. Each message asks about exactly ONE visual dimension. Wait for answer before next question.
Multiple-Choice Preferred
Present 2-3 concrete options with trade-offs. Always include a "You describe it" escape hatch. Lead with recommendation and explain why.
Explain What the User Doesn't Know
When a user says "I don't understand lighting," don't skip it — explain each option in human terms:
- What it looks like
- What the audience will feel
- Which serves their story best
- Why you recommend one
Anti-AI-Aesthetic Early Warning
If the user's visual style is deliberately "imperfect" (low-res / grainy / mockumentary / rough), warn early: most AI image tools default to clean and beautiful, and fighting that bias is expensive. Don't wait until prompt-writing to discover this.
Lock Then Move
Each segment confirmed → locked. Never revisit locked segments unless user explicitly asks. Prevents infinite circling.
Don't Re-Ask What's Already Said
If the user's original description already implies a dimension's answer, don't re-ask. Instead, confirm in one sentence:
"你说了'像要拥抱'——毛巾横向展开像伸开双臂,对吧?"
Only go into detailed follow-up questions when the original description isn't sufficient to generate a good prompt, or when the dimension has multiple valid interpretations that even the user may not have considered (e.g., lighting direction, color tone).
Per-Scene Checklist
For every scene, work through these dimensions in order (one at a time):
- Space — What kind of room? How big? How many people? What's the atmosphere?
- Camera — Where is it placed? Does it move? What device type? Why that choice?
- Foreground / Background — What blocks the lens? What's in the distance? What's out of focus?
- Subject — Who? What are they doing? What expression? Is this about action or stillness?
- Light — Where from? Hard or soft? What gets lit, what stays dark? Why that direction?
- Color & Texture — Color temperature? Image quality (clean vs grainy vs distorted)? Any special treatment?
- Sound — Ambient noise? Action sounds? Breathing? Lines? What's absent (no music, no voiceover)?
Global Setup Before Scene Work
Before asking about individual scenes, establish:
| Item | Why |
|---|---|
| Two-line structure (A/B, inside/outside, past/present) | Prevents everything blending into same texture |
| Space contrast between lines | "Dark grid office vs quiet room" — locks visual contrast early |
| Light source per line | Single source per line = realistic, controllable |
| Sound rules | No music? No voiceover? All rough? — sets texture baseline |
Design Document Format
Save completed design to docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md with:
- Global settings table
- Per-act/scene tables with all dimensions
- All confirmed dialogue/quotes
- Status markers (✅ confirmed / ⏳ in progress)
Common Mistakes
| Mistake | Fix |
|---|---|
| Writing camera angles as prompt language ("low angle, dutch tilt") | Write what the VIEWER sees, not what the camera does |
| Asking lighting preferences without explaining what each option means | Explain visually in human terms, then ask |
| Letting user rush through multiple dimensions at once | Hold the line: one at a time |
| Not warning about AI limitations until prompt-writing stage | Flag early when visual style conflicts with AI defaults |
| Revisiting locked decisions without user permission | Only reopen when user explicitly asks |
Red Flags — Slow Down
- User jumps ahead: "Let's just do all four scenes at once"
- User skips a dimension: "Light doesn't matter, let's move on"
- User wants to write prompts before design is complete
- "This is simple enough, we don't need to go through every dimension"
Each of these means: Hold position. Complete current dimension first.
Final Step: Multi-Tool Prompt Generation
After the design is fully confirmed and saved, ask:
"设计定稿了。要不要我分别生成不同 AI 工具的提示词版本?比如小云雀、LiblibAI、Midjourney、DALL·E 等——每个工具写法不同,我可以针对各平台写适配版。"
Tool-Specific Adaptation Rules
| Tool | Style | Notes |
|---|---|---|
| 小云雀/国产 | 自然中文段落描述 | 中文零损耗,可写长段落 |
| Midjourney | 英文 + 关键短语,用 -- 参数控制风格 |
需英文,用逗号分隔关键词而非完整句子 |
| DALL·E | 英文自然语言,精炼 | 对"监控""偷拍"等词敏感,可能触发审核 |
| LiblibAI / SD | 中文 + 可选模型/LoRA/ControlNet | 可指定纪实/胶片模型,控图最强 |
Anti-Failure Warnings
When writing prompts, explicitly warn about known failure modes of each tool based on the visual style:
- If the design uses low-res / grainy / mockumentary visual: warn that AI defaults toward clean/beautiful
- If the design uses unusual camera angle (chest-forward not selfie): warn about common perspective errors
- If the design involves specific facial expressions: warn about consistency across multiple generations