Hihiii

cinematic-prompt-enhancer

"Professional prompt enhancement skill: takes a raw user prompt and enriches it with cinematic image decisions, expert visual synthesis, recognizable character IP identity systems, and model-specific final prompts for ChatGPT Image, Z Image Base, Z Image Turbo, Stable Diffusion, Midjourney, Flux, image edit, video workflows, and publication-style cover/poster outputs, mature-content age-safety control, human anatomy / pose biomechanics, adult body-presentation / sex-characteristic consistency, external intimate-anatomy consistency locks, physical-collision / spatial-contact guardrails, cloth/wetness physics, 3D lookdev/material response, and scene worldbuilding / background enrichment."

Hihiii 0 Updated 4w ago

Resources

3
GitHub

Install

npx skillscat add hihiii/imagepromptenhancer

Install via the SkillsCat registry.

SKILL.md

Cinematic Prompt Enhancer

This skill takes a plain user prompt and enhances it into a cinematography-grade prompt package using professional photography and filmmaking knowledge bases.

The core value is structured enrichment: every added detail is grounded in a real photography decision (why this lens, why this modifier, why this color palette), not random fluff.

Enhancement Knowledge Bases

Reference these config/ files during enhancement — each file follows a consistent four-layer architecture: core datagenerator translationquality/rulesintegration:

  • config/visual-cinematography/lighting-diagrams.yaml — 30 modifiers, 26 patterns, 6 source types, 13 control tools, 7 cinematic genres, 4 portrait patterns, 3 special function, 17 scene recipes + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + lighting safety + face effects + material-specific lighting + scene presets

  • config/visual-cinematography/camera-lens-library.yaml — 6 focal length strategies, 14 lens examples, 12 filters, 4 accessories, 18 selection guide, 3 cinema optical, 2 specialty + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + guardrails + output mode style + assembly hooks

  • config/visual-cinematography/color-palette-schemes.yaml — 16 schemes, 35 grading looks, 12 temperatures, 12 psychology, 7 saturation + color role system, ratio rules, selection matrix, skin tone protection, brand handling, accessibility, material color, scene presets + translation + deduplication + failure taxonomy + conflict resolution

  • config/visual-cinematography/exposure-strategies.yaml — 17 strategies, 11 zones, 6 advanced concepts, 5 tonality, 8 ratios + generator translation + intent selector + tonal zone prompting + failure taxonomy + conflict resolution + compression rules + quality rubric + guardrails + dynamic range + output mode + assembly hooks

  • config/visual-cinematography/composition-library.yaml — 7 framing types, 15 rules, 8 strategies, 8 genre guides, 10 pitfalls, 5 aspect ratios, 3 motion, 4 camera angles, 3 temporal, 3 metaphors + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + positioning guardrails + text/brand layout + scene zoning + visual hierarchy + poster layout + UI/game rules + architecture perspective + assembly hooks

  • config/visual-cinematography/composition-decision-engine.yaml - composition master controller. Decides subject priority, frame strategy, hierarchy, gaze/action flow, depth, background role, negative space, crop policy, camera distance, and balance before style detail.

  • config/visual-cinematography/tonal-foundation-system.yaml - tonal controller for key level, midtones, contrast curve, dynamic range, highlight roll-off, and shadow detail.

  • config/visual-cinematography/exposure-decision-engine.yaml - exposure controller for face, product, highlight, and balanced-scene readability.

  • config/visual-cinematography/aperture-depth-of-field-system.yaml - depth controller for shallow portrait, moderate-readable, deep-focus, and macro focus behavior.

  • config/visual-cinematography/shutter-motion-control-system.yaml - motion controller for freeze, natural motion, directional blur, panning, and long exposure.

  • config/visual-cinematography/iso-noise-control-system.yaml - texture controller for clean, natural, documentary, and film-grain image quality.

  • config/visual-cinematography/lens-practical-shooting-system.yaml - practical lens controller for shooting distance, perspective behavior, and distortion guardrails.

  • config/visual-cinematography/camera-technical-router.yaml - technical master router. Resolves tonal, exposure, lens, depth, shutter, and ISO decisions before renderer output.

  • config/visual-cinematography/visual-hierarchy-system.yaml - first-read controller. Keeps face, eyes, product, action target, or title dominant according to task.

  • config/visual-cinematography/gaze-action-composition-logic.yaml - gaze/action framing controller. Adds lead room, target visibility, screen/mirror/object attention, and body-action alignment.

  • config/visual-cinematography/depth-layering-system.yaml - foreground/midground/background controller. Chooses shallow, moderate-readable, deep-focus, or graphic depth.

  • config/visual-cinematography/background-role-policy.yaml - background role controller. Defines minimal, readable secondary, worldbuilding, spectacle, product support, stage, and architecture backgrounds.

  • config/visual-cinematography/negative-space-policy.yaml - negative-space controller for posters, covers, product heroes, action lead room, and mood/scale.

  • config/visual-cinematography/visual-balance-system.yaml - visual counterweight controller for subject mass, light, architecture, typography, props, groups, and empty space.

  • config/visual-cinematography/crop-safety-rules.yaml - crop safety controller for joints, hands, feet, products, vehicles, architecture, and typography.

  • config/visual-cinematography/framing-layout-templates.yaml - reusable framing templates such as centered hero, thirds portrait, cover vertical, stage wide, reclining interior, duo, and group lineup.

  • config/visual-cinematography/expert-visual-synthesis-system.yaml — cross-domain expert synthesis controller for professional photography, standing human frame ratio, clothing open-front/fit/translucency, scene-assisted body description, skin microtexture, wet cloth geometry, lighting physics, 3D lookdev/material response, and motion/animation physical cues.

  • config/visual-cinematography/visual-mood-library.yaml — 11 visual languages, 4 mediums, 14 moods + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + cinematography mapping + guardrails + scene presets + output mode + assembly hooks

  • config/prompt-core/output-specs.yaml — 10 output specs, 3 delivery guidelines + mode router + aspect ratio matrix + failure taxonomy + cleanup rules + safe zone guardrails + transparent/print/video contracts + advisory length guidance + avoid list rules + output examples + assembly hooks

  • config/visual-cinematography/motion-capture.yaml — 5 shutter strategies, 4 motion techniques, 5 subject guides, 3 action aperture, 4 scenarios, 4 rigging, 3 fluid dynamics + generator translation + intent selector + camera movement grammar + failure taxonomy + conflict resolution + compression rules + quality rubric + pose readability + secondary motion + guardrails + assembly hooks

  • config/prompt-core/constraints-library.yaml — 16 constraints, 3 text preservation rules + generator translation + extraction patterns + failure taxonomy + conflict resolution + compression rules + attribute change control + constraint recipes + safety filter + avoid list + text/brand constraints + guardrails + output mode + assembly hooks

  • config/prompt-core/task-profiles.yaml — 7 task profiles, 11 enhanced profiles + router + depth modes + config map + failure taxonomy + conflict resolution + compression rules + output contracts + assembly hooks

  • config/visual-cinematography/shot-taxonomy.yaml — 3 shot types, 14 expanded, 4 object shots + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + guardrails + output mode + assembly hooks

  • config/visual-cinematography/camera-angle-library.yaml — 3 angles, 9 expanded, 6 view directions + generator translation + intent selector + strength control + failure taxonomy + conflict resolution + compression rules + quality rubric + guardrails + output mode + assembly hooks

  • config/prompt-core/subject-rules.yaml — 3 subject types, 14 expanded rules + generator translation + failure taxonomy + conflict resolution + compression rules + quality rubric + output mode + assembly hooks

  • config/visual-cinematography/material-rendering-library.yaml — 3 materials + generator translation (22 entries) + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + priority rules + guardrails + lighting interaction + color behavior + output mode + assembly hooks

  • config/visual-cinematography/style-realism-control.yaml — 3 style modes, 5 realism levels + generator translation + intent selector + strength control + failure taxonomy + conflict resolution (v1+v2) + compression rules + quality rubric + guardrails + output mode + assembly hooks

  • config/visual-cinematography/cinematography-grammar.yaml — 3 grammar patterns, 10 film grammar, 3 blocking/staging, 3 motivated lighting, 2 rhythm/editing + generator translation + intent selector + failure taxonomy + conflict resolution + compression rules + quality rubric + guardrails + output mode + assembly hooks

  • config/prompt-core/failure-taxonomy.yaml — 10 base failures, 20 expanded, 4 severity levels + diagnostics + self-review checklist + config router (11 domains) + compression rules + output mode + repair templates + assembly hooks

  • config/prompt-core/quality-rubric.yaml — 12 dimensions, 5 core groups, 10 additional, 6 scale, 7 hard rejects, 5 score bands + config router (22 mappings) + prompt checks + thresholds + scoring formula + iteration rules + review contract + report templates + repair templates + assembly hooks

  • config/prompt-core/prompt-assembly-schema.yaml — 5 assembly rules, 4 modes, 22 segment contracts, 10 priority stack, 7 pipeline stages + translation rules + conflict resolution + quality checks + deduplication + detail priority allocation + cleanup order + avoid list + model-specific assembly + final templates + assembly hooks

  • config/prompt-core/semantic-priority-rules.yaml — 18 priority rules with weights 100→70

  • config/prompt-core/prompt-renderer.yaml — model-specific final prompt renderers for ChatGPT Image, Z Image Base, Z Image Turbo, Stable Diffusion, Midjourney, Flux, image edit, and video prompt outputs. Defines syntax style, advisory length guidance, abstraction tolerance, positive/negative split, field ordering, completeness-preserving cleanup behavior, hero-integrity boosts, and publication-style cover/poster typography rules per model.

  • config/prompt-core/image2-benchmark-learning-system.yaml - benchmark-learning controller. Use when comparing Image 2.0 / ChatGPT Image outputs against Z Image outputs, or when Z Image loses editorial layout, character identity, object complexity, palette discipline, or text strategy.

  • config/safety-mature/mature-content-control.yaml — mature-content age-safety controller for adult glamour, swimsuit, boudoir-inspired, body-conscious fashion, adult cosplay, adult explicit content, and gore prompts. Defines adult age locks, underage / childlike blocking, renderer-specific mature wording, and age-safety negative prompt guardrails.

  • config/character-identity/human-anatomy-guardrails.yaml — human figure anatomy and pose-biomechanics guardrails for portraits, full-body fashion, glamour, cosplay, swimsuit, dance, and dynamic poses. Defines anatomy-risk detection, pelvis / hip / thigh / knee / leg / foot rules, lower-body failure taxonomy, generator-visible anatomy translation, self-review checks, and renderer-specific anatomy strengthening.

  • config/character-identity/face-expression-gaze-system.yaml — facial expression and gaze controller for portraits and characters. Defines gaze target logic, coordinated eyes/brows/mouth/jaw/head cues, natural micro-expression rules, attention-anchor consistency, common uncanny-face failures, and model-specific rendering guidance.

  • config/text-brand-layout/text-rendering-accuracy.yaml — text accuracy and copy-integrity controller for titles, labels, logos, covers, posters, signage, UI, and multilingual text. Defines exact-text locks, no-extra-text rules, implied cover/poster text handling, no-text contradiction prevention, hierarchy, script integrity, and renderer-specific wording.

  • config/character-identity/sex-characteristics-control.yaml — adult body-presentation and sex-characteristic consistency controller. Defines intended presentation detection, body-trait locks, intimate-anatomy locks, explicit user overrides, male/female/androgynous guardrails, external-anatomy drift failures, renderer-specific wording, and self-review checks.

  • config/scene-environment/physical-collision-guardrails.yaml — physical collision, contact logic, occlusion, spatial layering, support mechanics, and material/light interaction controller. Defines clipping/intersection failures, subject-object contact rules, overlap specification, contact shadows, and renderer-specific spatial realism strengthening.

  • config/scene-environment/scene-worldbuilding-library.yaml — scene enrichment and background worldbuilding controller. Defines scene-type inference, depth layering, background density, environmental storytelling, subject-background interaction, and style-consistent worldbuilding.

  • config/scene-environment/environment-props-library.yaml — environment-specific prop and anchor-element library used to enrich backgrounds without cluttering the hero subject.

  • config/character-ip/character-ip-database.index.yaml — lightweight character lookup index. Stores character keys, display names, English display names, aliases, English aliases, source, and category. Check this first when a prompt mentions a character name or alias.

  • config/character-ip/character-ip-database.yaml — reusable character profile schema and split-database router. Full per-IP files use three maintenance layers: base_profiles for stable identity/source/mode metadata, enhanced_profiles for detailed visual anchors and renderer notes, and aliases for display names, English names, aliases, and lookup terms. Load the matched per-IP file after index lookup, then merge base_profiles[character_key] with enhanced_profiles[character_key].

  • config/character-ip/_indexes/character-lookup-index.yaml — formal per-character profile lookup for migrated IPs. When it resolves a character, load the mapped character directory and combine identity.yaml, appearance.yaml, face-hair-body.yaml, outfits/canonical.yaml, props-weapons.yaml, pose-language.yaml, prompt-pack.yaml, and negative-guards.yaml.

  • config/character-ip/_global/shared-rules/character-costume-auto-apply-policy.yaml — explicit wardrobe policy. Character identity alone never applies canonical clothing; apply it only for canonical/original/cosplay/character-accurate wardrobe intent, and let an explicit user outfit override canonical clothing.

  • config/character-identity/character-ip-identity-system.yaml — character IP identity controller. Converts character names, original character prompts, cosplay, and reinterpretations into recognizable visual locks: silhouette, hair, face impression, palette, costume markers, signature props, archetype behavior, world motifs, allowed variation, and forbidden drift.

  • config/character-identity/character-archetype-visual-system.yaml — character behavioral identity controller. Converts archetype/personality cues into visible gaze, expression, posture, gesture, camera relationship, and emotional range so the character remains recognizable beyond costume.

  • config/character-identity/costume-marker-priority-system.yaml - costume marker hierarchy controller. Ranks silhouette, color blocking, emblem/crest, signature props, garment structure, material language, and micro-details for direct references, cosplay, fashion reinterpretation, and photoreal conversion.

  • config/character-identity/reclining-pose-physics.yaml - reclining and semi-reclining body mechanics controller. Defines load-bearing points, pelvis/spine/leg alignment, hand/elbow support, surface contact, and pose-specific negatives.

  • config/safety-mature/adult-pose-tiering-system.yaml - adult pose router for glamour, boudoir, swimwear, implied figure, artistic figure, and adult-focused pose intent.

  • config/safety-mature/age-maturity-guardrails.yaml - age boundary controller; adult poses require clearly adult presentation and only underage or childlike sexual framing is blocked.

  • config/character-identity/adult-standing-pose-system.yaml - adult standing templates for contrapposto, swimwear editorial, power glamour, and back-view turns.

  • config/character-identity/adult-seated-pose-system.yaml - adult chair, sofa, and bed-edge templates with support, pelvis-leg alignment, and surface response.

  • config/character-identity/adult-reclining-pose-system.yaml - adult side-recline, back-recline, and propped-elbow templates with camera-aware support checks.

  • config/character-identity/adult-nude-pose-system.yaml - adult implied and artistic figure-pose controller for body line, coverage geometry, lighting, and physical support.

  • config/character-identity/adult-hand-placement-system.yaml - hand-intent controller for waist, hair, thigh, garment, fabric grip, framing, and load-bearing contact.

  • config/character-identity/adult-lens-aware-posing-system.yaml - pose/lens compatibility controller for wide, portrait, high-angle, and low-angle adult posing.

  • config/character-identity/adult-garment-pose-interaction-system.yaml - garment, wetness, tension, strap, fabric, and body-paint interaction controller for adult poses.

  • config/character-identity/explicit-wardrobe-policy.yaml - explicit-only wardrobe controller. Unspecified wardrobe slots remain absent; scenes, styles, default modesty, and character identity cannot add clothing.

  • config/character-identity/wardrobe-slot-occupancy-system.yaml - tracks worn, absent, unspecified, and non-garment coverage states independently for every wardrobe slot.

  • config/character-identity/chinese-wardrobe-explicit-semantics.yaml - Chinese wardrobe presence/absence parser that distinguishes 穿著, 沒穿, and omission.

  • config/character-identity/context-wardrobe-inference-blocker.yaml - blocks clothing inference from rain, beach, pool, bedroom, boudoir, school, office, fashion, selfie, and character-IP context.

  • config/character-identity/wardrobe-conflict-resolver.yaml - Phase 4.5 resolver that removes unrequested garments and material effects while preserving non-garment coverage.

  • config/character-identity/standing-pose-balance-system.yaml - standing pose balance controller. Defines support leg, center of gravity, foot contact, knee alignment, pelvis/shoulder counterbalance, and silhouette readability.

  • config/scene-environment/soft-surface-contact-physics.yaml - bed/sofa/cushion/blanket contact controller. Defines compression, wrinkles, fabric tension, contact shadows, and occlusion boundaries under weight.

  • config/scene-environment/action-object-attention-system.yaml - action-object attention controller. Aligns gaze, hands, body orientation, object placement, screen glow, tool use, product handling, and target reaction.

  • config/character-identity/reference-image-override-system.yaml - character reference override controller. Defines how identity, outfit, pose, style, and lighting references override or supplement per-IP profiles.

  • config/prompt-core/lora-routing-system.yaml - LoRA routing controller for SDXL, Pony, Illustrious, anime, and custom diffusion workflows.

  • config/prompt-core/model-specific-compact-prompt-generator.yaml - compact prompt controller that preserves hard locks while shortening renderer-specific output.

  • config/prompt-core/composition-preset-routing.yaml - routes cover, selfie, performance, reclining, automotive, destruction, and architecture prompts to composition presets.

  • config/visual-cinematography/selfie-perspective-disambiguation.yaml - selfie semantic router. Plain selfie / 自拍 means close camera relationship with phone off-frame; phone, mirror, and selfie-stick modes require explicit triggers.

  • config/visual-cinematography/camera-relationship-system.yaml - camera relationship controller for close handheld selfie perspective, vlog framing, and POV without accidental device props.

  • config/character-identity/character-sheet-multiview-system.yaml - character sheet and multi-view controller for front/side/back/expression/prop consistency.

  • config/character-identity/group-composition-identity-preservation.yaml - group identity matrix controller for lineups, ensemble posters, selfies, battle groups, and multi-character scenes.

  • config/character-identity/hero-action-character-scene-guardrails.yaml - hero action character scene controller. Coordinates recognizable character/IP identity, hero pose biomechanics, hand-prop contact, wanted/poster text policy, torn paper physics, open outerwear cloth-body collision, motivated cinematic lighting, and scene-depth storytelling.

  • config/visual-cinematography/luminous-anime-detail-system.yaml - luminous anime-realism detail controller. Coordinates filtered light, translucent outfit construction, controlled color roles, layered foreground/midground/background richness, small prop contact, wet-surface reflections, delicate anime face expression, and festival detail polish.

Enhancement Flow

Before Phase 2 output is assembled, run config/visual-cinematography/camera-technical-router.yaml after the composition decision layer. Preserve the critical subject plane first, then resolve tonal, exposure, practical lens, depth of field, motion, and texture as one compatible package. Render visible results rather than equipment lists.

For mature adult pose intent, classify the adult pose tier, confirm clearly adult presentation, choose the standing/seated/reclining/figure pose family, define support points and hand intent, then apply lens-aware and garment-pose interaction rules. Do not downshift adult pose intent; only underage or childlike sexual framing is blocked.

For every wardrobe slot, extract only items explicitly present or absent in the original prompt. An omitted slot must remain unfilled: do not infer clothing, underwear, swimwear, towel/fabric wraps, accessories, footwear, body paint, cloth folds, straps, or wet garments from scene, style, weather, or character identity. Canonical character costumes require an explicit canonical-outfit, cosplay, or character-accurate request.

All phases run internally. Phase 5 presents the model-specific final_prompt for user sign-off and writes a timestamped TXT export for downstream generation.

  1. Phase 1 (internal): Analyze the original prompt — extract subject, goal, audience, visual hierarchy, narrative context, model_target, publication_layout, benchmark_learning, mature_content, anatomy_risk, sex_characteristics, collision_risk, constraint_model, attribute_change_control, and risk flags.
  2. Phase 2 (internal): Enrich with cinematic analysis — define composition decision layer, color script, lighting intent, lens strategy, exposure plan, benchmark design-system transfer when relevant, and expert visual synthesis. Then translate the human-technical plan into generator-visible effects (subject position, lead room, depth cues, lighting behavior, material response, model risks, corrective terms).
  3. Phase 3 (internal): Build an enriched scene blueprint with a full lighting diagram (modifier, position, distance, angle, color temp, gel for each light).
  4. Phase 3.5 (internal): Compile the enhanced prompt package — categorized into type, subject, composition, style_and_lighting, background, material, atmosphere, constraints, output_format, footer.
  5. Phase 4 (internal): Self-review optimization — re-run Phase 1→3.5 on the Phase 3.5 output, correct gaps and contradictions.
  6. Phase 4.5 (internal): Prompt cleanup and consistency check — remove only semantic garbage, contradictions, equipment leaks, avoid-list conflicts, underage/childlike mature-content conflicts, and unsupported anatomy/body-presentation drift. Do not optimize for shortness; visual completeness outranks token economy.
  7. Phase 4.6 (internal): Select the model-specific renderer from config/prompt-core/prompt-renderer.yaml based on model_target; if a cover/poster/publication request is detected, also select the correct publication_mode, apply mature-content age-safety rules, anatomy guardrails, and sex-characteristic consistency rules when needed, then convert the cleaned prompt package into the target model's preferred syntax without enforcing a length budget.
  8. Phase 5: Assemble the final model-specific prompt, present it for user sign-off, and write final_YYYYMMDD_HHMMSS.txt with each model target's positive prompt, negative prompt, and suggested resolution.

For every image prompt, run the composition decision layer before renderer wording: decide the primary subject, first-read target, camera relationship, subject position, gaze/action direction, lead room, depth layers, background role, negative-space/typography zones, crop policy, visual balance, and light hierarchy. Final model prompts must express these decisions when they affect the image.

If the user provides a stronger Image 2.0 / ChatGPT Image output as a benchmark for a weaker Z Image result, classify the task with config/prompt-core/image2-benchmark-learning-system.yaml. Extract the benchmark's design decisions first: editorial grid, reading order, hero/support-panel relationship, character identity locks, vehicle/product complexity, real environment context, palette discipline, material response, and text/post-layout policy. Do not simply add more adjectives. For Z Image Base, Z Image Turbo, and ComfyUI Flux -> Z Image workflows, put these decisions early in the positive prompt and place only targeted drift terms in the negative prompt.

Short Prompt Expansion Mode (default for short user input)

When the user provides a short prompt, fragment, keyword list, or under-specified image idea, treat it as a seed brief and expand it into a precise, complete, model-facing prompt. The final prompt should be longer and richer than the user's input, not compacted back down.

Expansion goals:

  • Infer the missing visual decisions needed for a high-quality image: subject role, shot size, subject scale in frame, composition, lighting effect, material behavior, skin/fabric realism, wetness physics, background depth, atmosphere, output use, and high-risk avoid terms.
  • Preserve all explicit user constraints as hard locks.
  • Add only details that support the user's likely intent; do not change the subject, genre, identity, brand, era, wardrobe, or content category.
  • Use config/ knowledge bases internally, but render final wording as natural image language rather than equipment lists or raw YAML.
  • For chatgpt_image, default to a complete natural-language prompt with all required visual decisions preserved, not a short keyword chain.
  • For mature content, preserve requested adult content and only block underage or childlike sexual framing.

Detail Integrity & Conflict Rules

During prompt expansion, treat visual correctness as a hierarchy rather than a pile of adjectives:

  • Preserve explicit user constraints first: subject, identity, brand, era, wardrobe, pose, content category, output use, and exact text.
  • For people, always check structure when relevant: head/neck/shoulders, torso/pelvis, support leg, feet, hands/fingers, physical contact, expression, and gaze target.
  • For facial realism, coordinate eyes, eyelids, brows, mouth corners, jaw, head angle, and attention target. Avoid dead eyes, forced smiles, doll faces, expression/action mismatch, and uncanny symmetry.
  • For text, logos, labels, covers, posters, signage, and UI, preserve requested or implied text intent. Never add broad no text cleanup when text is required; use narrow guards like random text, misspellings, fake logos, or extra words.
  • For professional photography quality, translate gear knowledge into visible effects: camera position, perspective, subject scale, motivated light direction, shadow shape, catchlights, background separation, and material response.
  • For premium photoreal outputs, route short human/fashion/wet-cloth/render prompts through config/visual-cinematography/expert-visual-synthesis-system.yaml to align composition scale, body structure, clothing fit/opening, fabric translucency, wetness, skin texture, lighting physics, contact shadows, and 3D material response.
  • For standing human images, state the intended frame scale and subject occupancy: waist-up, three-quarter, full-body, or cover/poster layout. Protect headroom, footroom, natural crop points, and camera perspective.
  • For clothing, specify opening degree, anchor point, fit/contact areas, opacity/translucency cause, fold direction, and how the scene changes the fabric.
  • For wet-cloth scenes, state the physical cause: water load, adhesion, gravity, tension, darker damp zones, heavier folds, gradual opacity changes, and contact shadows.
  • For realism, preserve refined pores, skin tone variation, natural specular highlights, wet hair strand grouping, and environmental effects without plastic AI gloss.
  • For 3D/rendered quality, translate PBR/lookdev into visible roughness, specular behavior, contact shadows, ambient occlusion, stable geometry, and material-specific highlights.
  • Style, mood, cinematic depth, bokeh, shadows, and atmosphere are lower priority than anatomy, expression, text, product geometry, architecture, and task readability.

Phase 1: Intent Analysis (internal)

Extract the core requirements from the raw user prompt. Identify what matters so Phase 2-3 enrichment targets the right dimensions. This output is consumed by subsequent phases — do not report to the user.

Primary Fields

Field Purpose
task_profile select from config/prompt-core/task-profiles.yaml to apply sensible defaults
subject_type what is being photographed (product, person, food, scene, etc.)
subject_role the subject's narrative function (hero, context, accent, environment)
usage_context where the image will be used (influences contrast, saturation, sharpness)
audience who the image is for (influences color temperature, lighting mood)
aspect_ratio extract or infer from the prompt
output_medium print | digital | social | unspecified
model_target target prompt consumer: chatgpt_image, z_image_base, z_image_turbo, comfyui_flux2_zimage_3stage, stable_diffusion, midjourney, flux, image_edit, video_prompt, or unspecified
publication_layout publication-style layout intent: none, magazine_cover, poster, book_cover, album_cover, movie_poster, or unspecified
mature_content mature content classification and age-safety control: safe, adult_mature, adult_explicit, requires_adult_age_lock, or minor_blocked
anatomy_risk human anatomy / pose-biomechanics risk level and lower-body-structure control: low, medium, high, or critical
face_expression_gaze gaze target, emotional tone, facial micro-expression, attention anchor, and portrait realism risk
text_rendering exact text, logo, title, label, language/script, copy hierarchy, and no-extra-text / no-text contradiction control
sex_characteristics intended adult body presentation, body-trait consistency lock, and external intimate-anatomy consistency lock: male, female, androgynous, trans, nonbinary, genderfluid, user_specified, or unspecified
collision_risk physical-collision, contact-logic, occlusion, and material-interaction risk level: low, medium, high, or critical
scene_enrichment scene type, background density, depth layering, environmental storytelling, and subject-background interaction package

Visual Hierarchy & Narrative

Field Purpose
visual_hierarchy ordered list of what the eye should see first → last
narrative_context what story does the image tell? (e.g., "product emerging from darkness", "freshly prepared on a kitchen counter")
hero_element the single most important visual detail that must be perfect
secondary_elements supporting objects, props, or background features
reference_mood implied mood references (e.g., "like a luxury perfume ad", "warm family dinner feel")

Publication / Cover Layout Detection

If the user asks for a magazine cover, poster, book cover, album cover, movie poster, or social cover visual, extract publication intent explicitly. A cover/poster request is not just a composition request — it usually implies typography, masthead, title hierarchy, and layout zones.

Field Purpose
publication_layout.type none | magazine_cover | poster | book_cover | album_cover | movie_poster
publication_layout.publication_mode exact_text_mode | implied_cover_text_mode | text_safe_layout_only_mode
publication_layout.masthead_required whether a large title/masthead should appear
publication_layout.cover_lines_required whether supporting cover lines or editorial text should appear
publication_layout.text_safe_zones clear areas reserved for title, masthead, subtitle, or later typography
publication_layout.exact_text exact user-provided text, if any
publication_layout.text_policy generate_text | placeholder_editorial_text | reserve_space_only

Publication mode rules:

  • If the user provides exact text, title, logo, masthead, subtitle, or cover lines → exact_text_mode.
  • If the user asks for a cover/poster but does not specify exact text → default to implied_cover_text_mode.
  • Use text_safe_layout_only_mode only when the user or workflow explicitly says typography will be added later.
  • Do not turn a cover/poster request into a text-free image unless explicitly requested.
  • For weaker text models, keep text simple: large masthead, short cover lines, clear spacing, simple sans-serif or editorial type.

Trigger examples:

  • "汽車雜誌封面" → magazine_cover + implied_cover_text_mode
  • "海報大片" → poster + implied_cover_text_mode
  • "標題寫 SUPRA STYLE" → exact_text_mode
  • "留上方空白讓我後製文字" → text_safe_layout_only_mode
publication_layout:
  type: "magazine_cover"
  publication_mode: "implied_cover_text_mode"
  masthead_required: true
  cover_lines_required: true
  text_safe_zones:
    - "large masthead area at top"
    - "supporting cover-line zones around subject"
    - "optional bottom cover-line area"
  exact_text: []
  text_policy: "placeholder_editorial_text"

Mature Content Detection

If the user prompt contains sensual, glamour, swimsuit, boudoir, body-conscious fashion, adult cosplay, explicit adult nudity, adult genital visibility, adult sexual actions, gore, or similar mature styling, classify it with config/safety-mature/mature-content-control.yaml. The purpose is to preserve adult intent while preventing only underage or childlike sexual framing from entering the final prompt.

Field Purpose
mature_content.level safe | adult_mature | adult_explicit | requires_adult_age_lock | minor_blocked
mature_content.adult_age_lock whether adult / 21+ / mature adult wording is required
mature_content.style_intent mature style language such as adult glamour editorial, adult explicit nudity, adult genital visibility, gore, or user-specified adult visual intent
mature_content.wardrobe_coverage fashion_full_coverage | body_conscious_fashion | swimsuit_editorial | lingerie_inspired_fashion | body_paint_inspired_fashion
mature_content.pose_safety adult_allowed | minor_blocked
mature_content.forbidden_elements underage or childlike sexual framing to remove or block
mature_content.renderer_guardrails model-specific positive and negative wording for age-safety only

Rules:

  • Mature styling must be clearly adult when sensual or explicit framing is present.
  • Use adult, 21+, mature adult, or adult model near the subject description when needed.
  • Do not use underage or childlike framing in mature/sexualized prompts.
  • Adult nudity, adult genital visibility, adult sexual actions, and gore are not restricted by this mature-content module.
  • Do not rewrite adult explicit requests into non-explicit editorial language unless the user asks for that style.
mature_content:
  level: "adult_mature"
  adult_age_lock: true
  style_intent: "adult glamour editorial"
  wardrobe_coverage: "swimsuit_editorial"
  pose_safety: "adult_allowed"
  forbidden_elements:
    - "childlike"
    - "underage"
  renderer_guardrails:
    positive:
      - "adult 21+ model"
    negative:
      - "childlike"
      - "underage"

Human Anatomy / Pose Biomechanics Detection

If the user prompt involves people and especially full-body, lower-body-visible, dance, glamour, swimsuit, cosplay, body-conscious, or dynamic posing, classify anatomy risk with config/character-identity/human-anatomy-guardrails.yaml.

This module is not only about “correct anatomy.” It breaks anatomy control into explicit structure and pose rules:

  • pelvis / hip structure
  • hip-to-thigh connection
  • support-leg weight distribution
  • knee direction consistency
  • leg proportion consistency
  • foot grounding
  • dynamic pose biomechanics
Field Purpose
anatomy_risk.level low | medium | high | critical
anatomy_risk.triggers which prompt cues caused anatomy strengthening
anatomy_risk.lower_body_visible whether lower-body structure must be guarded explicitly
anatomy_risk.support_leg_required whether the support leg / center-of-gravity logic must be stated
anatomy_risk.pose_biomechanics_required whether pose mechanics should be made explicit
anatomy_risk.renderer_guardrails_required whether model-specific anatomy wording should be added

Rules:

  • If the lower body is visible, do not rely on correct anatomy alone.
  • Use generator-visible guardrails such as believable pelvis, natural hip-to-thigh connection, knees aligned with thigh direction, and grounded support foot.
  • If the pose is dynamic or asymmetric, make the support leg and center of gravity readable.
  • If anatomy risk is high or critical, keep anatomy guardrails through Phase 4.5 and into the final renderer output.
  • For reclining, semi-reclined, side-lying, propped-on-hands, propped-on-elbows, or bed/sofa poses, check the ribcage-pelvis-thigh-knee-shin-foot direction chain. The pelvis, folded legs, shins, and feet must not point in unrelated directions.
  • For soft-surface poses, include visible support logic: hips, thighs, knees, palms, or elbows contacting the bed or sofa, subtle fabric compression, contact shadows, and no floating lower legs.
  • If hands or elbows support the torso, describe the shoulder-elbow-wrist/palm load path so the upper-body weight visibly transfers into the surface.
  • For adult nude, towel-slip, body-paint, wet-cloth, or anatomy-repair prompts, treat visible nipples, areolae, and requested adult external genital anatomy as anatomical landmarks. Their placement must follow chest volume, ribcage turn, pelvis orientation, camera perspective, and occlusion by towel, hands, thighs, or fabric.
  • If the pose is too complex for stable anatomy, preserve pose intent but simplify impossible limb relationships.
anatomy_risk:
  level: "high"
  triggers:
    - "full body"
    - "glamour pose"
    - "one-leg support pose"
    - "body-conscious outfit"
  lower_body_visible: true
  support_leg_required: true
  pose_biomechanics_required: true
  renderer_guardrails_required: true

Sex Characteristics / Body Presentation Consistency

If the prompt specifies adult male, adult female, androgynous, trans, nonbinary, gender-fluid, masculine woman, feminine man, or another explicit body-presentation concept, classify it with config/character-identity/sex-characteristics-control.yaml.

This module is for body-trait consistency, not explicit anatomy. It prevents unintended trait drift such as male subjects gaining unintended feminine chest structure, female subjects gaining unintended male-coded torso traits, or user-specified androgynous/trans/nonbinary presentation being overwritten.

Field Purpose
sex_characteristics.intended_presentation male | female | androgynous | trans | nonbinary | genderfluid | user_specified | unspecified
sex_characteristics.body_trait_lock male_coded | female_coded | androgynous | user_specified | unspecified
sex_characteristics.intimate_anatomy_lock male_external | female_external | user_specified | unspecified
sex_characteristics.explicit_user_override whether user explicitly requested non-binary or nonstandard presentation
sex_characteristics.consistency_required whether renderer should include body-presentation guardrails
sex_characteristics.forbidden_trait_drift unintended body-trait drift to prevent
sex_characteristics.forbidden_intimate_trait_drift unintended external intimate-anatomy drift to prevent
sex_characteristics.renderer_guardrails model-specific positive / negative body-presentation phrases

Rules:

  • Explicit user-specified body presentation wins.
  • Do not force binary traits when the user explicitly requests androgynous, trans, nonbinary, gender-fluid, or user-specified presentation.
  • Use visible, non-explicit body-structure language.
  • Do not add explicit anatomy.
  • Mature / glamour styling must not override body-trait locks.
  • When an adult male/female subject is specified and anatomy visibility is relevant, keep external anatomy consistent with the specified subject.
  • Female-coded subjects should not gain unintended male external anatomy.
  • Male-coded subjects should not gain unintended female external anatomy.
  • Explicit user-specified anatomy or body presentation overrides automatic binary correction.
  • Use safe consistency language by default; use strict intimate-anatomy guardrail terms only for NSFW-sensitive, anatomy-sensitive, repair, or repeated-failure contexts.
  • If the user explicitly reports misplaced nipples, areolae, or adult external genital anatomy, use strict placement guardrails: correct adult anatomical landmark placement, correct skin/fabric/towel occlusion, and no misplaced, floating, duplicated, or wrong-sex visible anatomy.
sex_characteristics:
  intended_presentation: "male"
  body_trait_lock: "male_coded"
  intimate_anatomy_lock: "male_external"
  explicit_user_override: false
  consistency_required: true
  forbidden_trait_drift:
    - "feminine breasts"
    - "female-coded torso"
    - "wrong sex characteristics"
  forbidden_intimate_trait_drift:
    - "vagina"
    - "female external genital anatomy"
  renderer_guardrails:
    positive:
      - "adult male figure"
      - "masculine torso structure"
      - "flat chest"
      - "male-coded shoulder-waist-hip proportions"
    negative:
      - "feminine breasts"
      - "female chest"
      - "wrong sex characteristics"

Physical Collision / Spatial Contact Consistency

If the prompt involves human-object contact, seated poses, leaning poses, multi-subject overlap, product handling, vehicles, furniture, props, layered compositions, or material-sensitive scenes, classify collision/contact risk with config/scene-environment/physical-collision-guardrails.yaml.

This module is not only about avoiding clipping. It also defines:

  • physical contact logic
  • support surfaces and weight transfer
  • front/back occlusion hierarchy
  • readable overlap amount
  • hand-prop interaction
  • clothing/accessory contact
  • material hardness / gloss / texture response
  • light-shadow interaction that reinforces physical presence
Field Purpose
collision_risk.level low | medium | high | critical
collision_risk.triggers which prompt cues triggered collision strengthening
collision_risk.contact_logic_required whether touching / support relationships must be made explicit
collision_risk.occlusion_required whether front/back layering must be specified
collision_risk.material_interaction_required whether surface/light interaction must be stated
collision_risk.renderer_guardrails_required whether renderer-specific collision wording should be added

Rules:

  • Explicitly state what touches what, what supports what, and what occludes what.
  • Use overlap wording such as A partially occludes B when composition depends on layering.
  • Use contact-shadow and material-response wording when objects otherwise look flat or plastic.
  • Do not allow body-object clipping, floating contact, impossible grips, wrong occlusion order, or merged geometry.
  • When material realism matters, state hardness, gloss, texture, and light interaction.
  • For bed, sofa, cushion, blanket, or mattress poses, specify where body weight settles and where fabric compresses. This is required for semi-reclined, side-lying, kneeling-on-bed, and hand-supported poses.
  • For pool-edge exit poses such as hands supporting on pool edge, climbing out of pool, or emerging from pool, preserve the downward support action: palms press down on the pool coping, fingers curl over the near edge, elbows bend downward, shoulders load into the hands, and the pool edge remains beside or in front of the subject rather than above the head.
  • For after-bath towel-slip poses, preserve towel physics: thick terry towel, continuous torso wrap, visible top-edge thickness, one edge slipping diagonally downward, another edge anchored under the arm or lightly held, heavy damp folds, and soft contact shadows where towel overlaps wet skin.
  • For after-bath or post-shower towel poses, default to wet context: damp or wet hair, wet skin with droplets, slightly damp towel where it touches the body, darker wet contact zones, and heavier absorbent terry folds unless the user explicitly asks for a dry staged towel.
collision_risk:
  level: "high"
  triggers:
    - "leaning on a car"
    - "hand holding a product"
    - "foreground overlap"
    - "glossy painted metal"
  contact_logic_required: true
  occlusion_required: true
  material_interaction_required: true
  renderer_guardrails_required: true

Scene Worldbuilding / Background Enrichment

If the task benefits from a fuller environment — such as editorial covers, fashion scenes, automotive scenes, lifestyle images, cinematic portraits, product ads, or posters — classify and enrich the scene with config/scene-environment/scene-worldbuilding-library.yaml and config/scene-environment/environment-props-library.yaml.

This module helps the system move beyond a subject-only prompt into a complete scene:

  • scene type inference
  • background density control
  • foreground / midground / background / far-depth layering
  • environmental storytelling
  • subject-background interaction
  • negative-space planning for typography when needed
  • style-consistent worldbuilding
Field Purpose
scene_enrichment.scene_type the most suitable environment type for the subject/task
scene_enrichment.background_density sparse | balanced | rich | cinematic_dense
scene_enrichment.depth_layers structured foreground, midground, background, far-background scene elements
scene_enrichment.environmental_storytelling atmosphere and narrative support elements
scene_enrichment.subject_background_interaction reflections, support, palette harmony, and light/environment relationship
scene_enrichment.style_consistency_rules keep the background rendering aligned with the subject/task style
scene_enrichment.negative_space_required whether background complexity should be reduced for text or hero focus

Rules:

  • Do not rely on "detailed background" alone; define the place.
  • Infer a suitable environment if the user leaves the background underspecified.
  • Use depth layering so the scene does not feel flat.
  • Use environmental storytelling details that strengthen the subject and mood.
  • Control density: some tasks need a rich world, some need clean negative space.
  • Keep background style aligned with the subject rendering language.

Era & Culture Contextual Authenticity

If the prompt contains or implies an era, historical period, cultural region, traditional style, retro setting, futuristic culture, or geographically specific visual language, classify it with config/scene-environment/era-culture-library.yaml.

Era and culture affect the full scene logic: architecture, fashion silhouettes, hair and makeup, props and tools, vehicles, signage and typography, furniture and interiors, materials and craft logic, color aesthetics, social atmosphere, and anachronism avoidance.

Rules:

  • Preserve explicit user-specified era, culture, and region.
  • Avoid anachronistic objects unless deliberate fusion is requested.
  • Avoid generic cultural mashups and stereotypes.
  • If fusion is requested, define the fusion logic clearly.
  • Coordinate era/culture with scene worldbuilding, environment props, rendering, and typography.

Environment & Atmosphere Control

If the prompt contains or implies mood, time of day, weather, season, air quality, humidity, cinematic atmosphere, environmental particles, or sensory ambience, classify it with config/scene-environment/environment-atmosphere-library.yaml.

Environment and atmosphere affect the sensory layer of the image: time of day, weather, season, air quality, humidity, temperature feel, atmosphere density, mood, sensory cues, lighting cues, particle effects, ambient motion, and background visibility.

Rules:

  • Translate mood into concrete environmental cues.
  • Keep weather, lighting, surface response, and air quality consistent.
  • Use atmosphere to support mood, not to randomly obscure the subject.
  • Coordinate atmosphere with lighting diagrams, scene worldbuilding, color palette, and rendering technicals.
  • Avoid overdone bloom, random fog, muddy visibility, or inconsistent weather cues.

Typography / Graphic Design Layout Control

If the request contains or implies typography, logo, brand name, magazine cover, poster, product ad, campaign visual, signage, UI-like layout, title, headline, cover lines, or exact text, classify it with config/text-brand-layout/typography-layout-system.yaml.

Typography affects the image as a design system:

  • layout type
  • text hierarchy
  • masthead / headline / subhead / cover-line logic
  • logo placement
  • grid system
  • safe margins
  • typography reserve zones
  • subject avoid zones
  • negative space planning
  • exact text preservation
  • random text prevention

Rules:

  • Never add no text when the user requested cover/poster/ad typography.
  • Preserve exact user-provided text, spelling, capitalization, and brand names.
  • Keep text away from faces, hands, product hero areas, and key composition lines unless intentionally requested.
  • If exact long text is critical, reserve clean zones and prefer post-layout editing.
typography_layout:
  layout_type: "magazine_cover"
  text_intent: "exact"
  text_hierarchy:
    masthead: "BRAND NAME"
    headline: "MAIN TITLE"
    cover_lines: []
  grid_system:
    layout_grid: "asymmetrical_editorial"
    typography_reserve_zones:
      - "top masthead band"
      - "right side cover-line column"
    subject_avoid_zones:
      - "face"
      - "hands"
      - "product logo"

Multi-subject Interaction Control

If the request contains multiple people, multiple characters, team scenes, group selfies, duos, families, crowds, sports teams, dance partners, fight/action scenes, or multi-product compositions, classify it with config/character-identity/multi-subject-interaction-system.yaml.

Multi-subject control prevents subjects from merging and clarifies:

  • subject count
  • per-subject character card
  • primary / secondary roles
  • left / center / right placement
  • foreground / midground / background placement
  • camera distance from each person
  • individual pose and gaze
  • hair / outfit / accessory differentiation
  • group formation
  • shared event center
  • event source subject
  • effect direction and reaction timing
  • subject spacing
  • interaction target
  • physical contact
  • occlusion order
  • identity separation
  • collision risk

Rules:

  • Count subjects explicitly.
  • Assign each subject a role, position, pose, gaze, and interaction target.
  • For multi-person action, define the shared event center, who causes it, the physical direction of the effect, and how each subject reacts.
  • For premium ensemble or cosplay-style group images, give every key person a compact character card: position, depth, hair color/style, outfit silhouette, material, accessory motif, expression, gaze, and hand role.
  • Use one shared camera, lighting, wetness/material system, and environment to unify the group; do not make separate characters look pasted from different images.
  • Preserve distinct faces, silhouettes, and wardrobes.
  • Apply physical-collision and occlusion guardrails for close-contact scenes.
  • For pool, splash, confetti, smoke, or motion-effect group scenes, keep hands, feet, support plane, and partial occlusion readable; do not allow the effect to become a pasted overlay.
  • For six or more subjects, use group rhythm and hierarchy rather than over-describing every person.
multi_subject_interaction:
  subject_count: 3
  primary_subject: "center foreground subject"
  character_design_matrix:
    - "subject_A: foreground center, pink hair, direct calm gaze, leather-and-fabric outfit, one foreground selfie arm"
    - "subject_B: upper left, blonde wet hair, confident gaze, white vest with straps and dark inner layer"
    - "subject_C: upper right, reddish hair, bright laugh, pink outfit accents, reaching hand"
  group_composition:
    formation: "triangle"
    spacing: "natural"
    hierarchy: "single_hero"
  interaction_logic:
    - "primary subject looks at camera"
    - "secondary subjects support the composition from left and right"
  shared_event_center:
    source_subject: "center-left subject"
    action_cause: "hands scooping water upward"
    effect_origin: "palms at pool surface"
    effect_direction: "upward arc, then falling droplets"
    reaction_targets:
      - "other subjects look toward the splash with varied natural reactions"
  identity_separation_required: true

Atmosphere & Cosmic Light

If the prompt depends on outdoor natural light, celestial light, storm drama, aurora, fog, mist, god rays, lightning, blizzard, sandstorm, or large-scale atmospheric spectacle, classify it with config/scene-environment/atmosphere-cosmic-light.yaml.

This module controls:

  • sun / moon / aurora / lightning / eclipse light logic
  • sky condition and cloud behavior
  • volumetric atmosphere and atmospheric optics
  • weather-driven particles and surface response
  • readability safeguards for dramatic outdoor scenes

Rules:

  • Tie outdoor mood to a real sky/light condition.
  • Extreme weather must affect visibility and surfaces, not only add particles.
  • Preserve subject readability under dramatic weather and low light.
  • Use atmospheric optics only when physically justified.

Oceans & Subaquatic Physics

If the prompt involves sea, shore, surf, waves, harbors, underwater scenes, reefs, marine ruins, diving, or strong subject-water interaction, classify it with config/scene-environment/oceans-subaquatic-physics.yaml.

This module controls:

  • wave and surf structure
  • tide and shoreline logic
  • foam, spray, wetness, and reflections
  • underwater visibility and light filtering
  • caustics, refraction, buoyancy, and suspended particles

Rules:

  • Specify whether the scene is above water, at the shoreline, or underwater.
  • Water state must match wind, weather, and camera distance.
  • Underwater scenes must show filtered light and depth-dependent visibility.
  • Wet surfaces should visibly respond to water.

Mountains & Terrains

If the prompt depends on mountain scale, ridgelines, cliffs, valleys, canyons, glaciers, dunes, alpine weather, or large-scale landform readability, classify it with config/scene-environment/mountains-terrains.yaml.

This module controls:

  • landform type and climate band
  • terrain layering from foreground to far distance
  • geology, vegetation, and snow/ice distribution
  • atmospheric perspective and ridge separation
  • scale readability in large landscapes

Rules:

  • Define the dominant landform type.
  • Use terrain layers to communicate scale.
  • Match vegetation, snow, and rock to altitude and climate.
  • Preserve subject readability while maintaining environmental grandeur.

Automotive Motion / Rain Physics

If the prompt involves a moving vehicle, sports car, racing scene, rain-soaked road, highway speed, wheel spray, or night headlights, classify it with config/prompt-core/subject-rules.yaml, config/visual-cinematography/motion-capture.yaml, config/visual-cinematography/cinematography-grammar.yaml, config/visual-cinematography/rendering-technicals-library.yaml, config/scene-environment/atmosphere-cosmic-light.yaml, and config/prompt-core/failure-taxonomy.yaml.

Rules:

  • Preserve vehicle identity, body type, correct grille, symmetrical headlights, clean wheel geometry, and tire contact.
  • Define travel direction and camera tracking direction.
  • For speed, prefer sharp panning language: crisp car body, spinning wheels, horizontal background streaks, and road-line motion.
  • For rain, define tire contact patches cutting through shallow standing water.
  • Water spray should be thrown outward and backward from each tire, not floating, moving forward, or detached from the wheels.
  • Headlight beams should align with the vehicle's forward path and scatter through rain mist; avoid diagonal spotlight beams crossing the car body.
  • Wet asphalt reflections should stretch along the road plane and perspective.
  • Avoid broad negative terms like blurry subject when motion is required. Use targeted negatives such as blurred car body, soft grille, melted headlights, warped wheels, diagonal spotlight crossing the vehicle, dry tire contact, and water spray without tire source.

Attribute Change Control

Explicitly define what the user has locked vs what the enhancement pipeline may change. This prevents Phase 2-4 from modifying attributes the user considers fixed.

Category Description
locked_attributes things the user has explicitly stated must not change (face, hairstyle, body shape, outfit, background, specific product, brand identity)
editable_attributes things the user has implied or explicitly permitted to change (pose, expression, lighting, camera angle, timing)
forbidden_changes hard rules from the user — things the enhancement pipeline must never do regardless of aesthetic judgment

Extract these from the user's language:

  • "same face" / "keep the same" / "不變" / "一樣" → locked
  • "only change..." / "just adjust..." → locked everything else, edit only specified
  • Unspecified attributes are implicitly editable unless context suggests otherwise
  attribute_change_control:
    locked_attributes:
      - face
      - hairstyle
      - body_shape
      - outfit
      - background
    editable_attributes:
      - pose
      - expression
      - lighting
    forbidden_changes:
      - changing identity
      - changing clothing
      - changing camera crop

Constraint Model (Three-Tier)

Classify every constraint extracted from the user prompt into one of three tiers. This prevents over-enhancement (adding things that conflict with user intent) and under-enforcement (ignoring critical user requirements).

Tier Description Examples
hard_locks Cannot be changed under any enhancement. These override all aesthetic judgment. exact brand colors, required text, specific subject identity, forbidden elements, minimum detail fidelity
soft_preferences Should be respected unless they conflict with a hard lock. May be deprioritized if they cause quality degradation. implied mood references ("like a luxury perfume ad"), suggested styles, optional color tones
inferred_enhancements Can be added only if they do not conflict with any hard lock or soft preference. These are the pipeline's creative license. additional mood adjectives, secondary lighting refinements, subtle material detail enrichments

Extraction rules:

  • Explicit user statements → hard_locks
  • Implied preferences and comparisons → soft_preferences
  • Gaps where no user constraint exists → inferred_enhancements may fill them from config knowledge bases
  constraint_model:
    hard_locks:
      description: "Cannot be changed under any enhancement."
      items: []
    soft_preferences:
      description: "Should be respected unless conflicting."
      items: []
    inferred_enhancements:
      description: "Can be added only if not conflicting."
      items: []

Risk Flags

Field Purpose
risk_flags known failure modes predicted from the configuration (text rendering, anatomy, exact identity, over-FX, brand consistency)

Target output:

intent:
  task_profile: "product_ad"
  subject_type: "glass bottle"
  subject_role: "hero"
  usage_context: "e-commerce hero image"
  audience: "premium skincare buyers"
  aspect_ratio: "1:1"
  output_medium: "digital"
  model_target: "chatgpt_image"

  visual_hierarchy: ["bottle_label", "gold_cap", "marble_reflection", "silk_drape"]
  narrative_context: "luxury skincare product floating in warm amber light, isolated from the world"
  hero_element: "product label with gold foil detail"
  secondary_elements: ["cream silk fabric right side", "polished marble slab"]
  reference_mood: "Tom Ford fragrance ad, dark warm luxury"

  constraint_model:
    hard_locks:
      - "gold tone accent, no bright colors"
      - "no human model"
      - "text LUMIÈRE on product label center — must be fully legible"
      - "text 'Gold Serum' on product label below brand name"
      - "keep right 30% clear for text overlay"
    soft_preferences:
      - "dark warm luxury mood like Tom Ford fragrance ad"
      - "studio controlled lighting"
      - "centered hero with breathing room"
    inferred_enhancements:
      - "amber gold color palette to support luxury brand feel"
      - "soft gradient background to isolate product"
      - "subtle marble reflection to add depth without distraction"

  risk_flags: ["text_rendering", "glass_reflection", "brand_color_accuracy"]

Phase 2: Cinematic Enrichment (internal)

Take the Phase 1 intent and enrich it with professional cinematography decisions. Every choice must reference the knowledge bases. Output is consumed by Phase 3 — do not report to the user.

Phase 2 produces two layers:

  1. Human-technical plan — the cinematography reasoning for the human DP
  2. Generator translation — the visible effects that the image model must render

This is because "f/2.8 at 85mm" means something very different to a cinematographer than to a diffusion model. The human layer captures the intent; the generator layer captures the visible result.

Color Script

Choose a palette and grading look from config/visual-cinematography/color-palette-schemes.yaml:

  • palette_type: complementary (contrast) | analogous (harmony) | monochromatic (focus) | triadic (vibrant) | split_complementary (refined contrast)
  • dominant_colors: colors that support the brand or mood
  • saturation_strategy: muted | vibrant | selective | desaturated
  • color_grading_intent: select from config/visual-cinematography/color-palette-schemes.yaml color_grading_looks

Lighting Intent

Choose a key modifier and ratio from config/visual-cinematography/lighting-diagrams.yaml and config/visual-cinematography/exposure-strategies.yaml:

  • key_quality: hard | soft | diffused | specular | wrap_around
  • key_modifier: select a modifier from config/visual-cinematography/lighting-diagrams.yaml modifiers
  • ratio: key-to-fill ratio from config/visual-cinematography/exposure-strategies.yaml ratios
  • color_temperature_strategy: matched | mixed_CTO | mixed_CTB | gel_accent
  • rim_or_separation: does the subject need edge separation?
  • background_treatment: flat | gradient | textured | gobo | patterned

Lens Strategy

Choose a focal length from config/visual-cinematography/camera-lens-library.yaml and its impact:

  • focal_length_category: wide_angle | wide_to_normal | standard | short_telephoto | telephoto | ultra_wide
  • compression: low | moderate | strong | extreme
  • aperture_intent: shallow_DoF | deep_focus | medium
  • working_distance: close | moderate | far
  • camera_height: low_angle | eye_level | high_angle | overhead
  • distortion_character: natural | slight_warp | anamorphic_feel

Exposure Strategy

Choose an exposure strategy from config/visual-cinematography/exposure-strategies.yaml:

  • style: high_key | low_key | middle_key | silhouette | chiaroscuro | flat_diffuse
  • contrast_ratio: low | medium | high | extreme

Composition

Choose a composition rule from config/visual-cinematography/composition-library.yaml:

  • rule: rule_of_thirds | centered_symmetry | golden_ratio | leading_lines | frame_within_frame | negative_space | diagonal_tension | depth_layering | and more
  • depth_layers: 2 | 3 | 4
  • focal_priority: ordered list of what the eye should see first
  • breathing_room: tight | comfortable | generous

Generator Translation Layer

Translate the human-technical plan above into effects the image model can actually render. The model does not understand "octabox_90cm" — it renders physical objects. The model does not understand "f/2.8" — it renders depth of field as a visible effect.

Build this from the human-technical decisions using knowledge-base lookups:

  • visible_lighting_effect: describe the light's behavior on the subject — shadow transition, catchlight shape, highlight character, wrap quality. Never name equipment.
  • visible_depth_effect: describe depth cues — foreground blur, background compression, focal falloff, atmospheric perspective.
  • visible_material_effect: describe surface behavior — specular response, subsurface scattering, reflection sharpness, translucency.
  • visible_composition_effect: describe the spatial relationship — negative space ratios, eye path, weight balance, leading vectors.
  • model_risk: known generation failure modes triggered by this configuration (e.g., "softbox visible as physical object", "hands merged with background at this ratio").
  • corrective_terms: terms that mitigate known model risks or cross-reference style guardrails from config/visual-cinematography/style-realism-control.yaml.

Also extract avoid_in_prompt from each modifier entry's prompt_keywords or notes — any equipment name that would cause the model to render studio hardware.

Target output:

cinematic_enrichment:
  color_script:
    palette_type: "complementary"
    dominant_colors: ["amber_gold", "deep_brown"]
    saturation_strategy: "selective_saturation"
    color_grading_intent: "vintage_warm"

  lighting_intent:
    key_quality: "soft_diffused"
    key_modifier: "octabox_90cm"
    ratio: "3:1"
    color_temperature_strategy: "mixed_CTO"
    rim_or_separation: true
    background_treatment: "gradient_warm"

  lens_strategy:
    focal_length_category: "short_telephoto"
    compression: "moderate"
    aperture_intent: "medium"
    working_distance: "moderate"
    camera_height: "eye_level"
    distortion_character: "natural"

  exposure_strategy:
    style: "middle_key"
    contrast_ratio: "medium"

  composition:
    rule: "negative_space"
    depth_layers: 3
    focal_priority: ["product_label", "glass_highlight", "texture"]
    breathing_room: "generous_right"

  generator_translation:
    visible_lighting_effect: "large soft diffused key light from camera-right, smooth shadow transition, warm round catchlights in the glass, gentle wrap-around on the bottle contour"
    visible_depth_effect: "moderate depth of field — sharp on the bottle label, gradual falloff into the silk drape, marble surface stays crisp at the product plane"
    visible_material_effect: "amber glass with glossy specular highlight on the shoulder, soft subsurface glow through the bottle, silk with visible woven texture, polished marble with mirror-like reflection"
    visible_composition_effect: "generous negative space on the right balanced by the silk fall, eye path from gold label to cap highlight to marble edge, hero centered with breathing room"
    model_risk:
      - "softbox may render as physical reflector visible in glass reflection"
      - "glass refraction may distort or omit label text"
      - "gold foil may render as flat yellow at low contrast"
    corrective_terms:
      - "no studio equipment visible"
      - "clean glass reflection without light source"
      - "gold foil texture with metallic sheen"
    avoid_in_prompt:
      - "octabox"
      - "softbox"
      - "V-flat"
      - "studio flash"
      - "light stand"

Phase 3: Enriched Scene Blueprint (internal)

Build a detailed scene description with a full lighting diagram. This is the "shooting plan" that will be compiled into categorized prompt fields. Output is consumed by Phase 3.5 — do not report to the user.

scene_blueprint:
  subject:
    type: "product"
    description: "amber glass serum bottle with gold dropper"
    pose_or_arrangement: "center-left on polished black marble"
    gaze_direction: "camera"

  environment:
    location: "minimal studio tabletop"
    subject_to_bg_distance: "moderate"
    set_dressing: ["black marble slab", "cream silk drape right"]

  camera:
    lens: "85mm short telephoto"
    aperture: "f/5.6"
    angle: "eye-level"
    framing: "medium close-up"
    shooting_distance_m: 1.5

  lighting_diagram:
    key_light:
      modifier: "octabox_90cm"
      position: "camera-right"
      distance: 1.2
      angle_horizontal: 45
      angle_vertical: 30
      color_temp_K: 5600
      gel: "full_CTO"
    fill_light:
      modifier: "V-flat white bounce"
      position: "camera-left"
      distance: 0.8
    rim_light:
      modifier: "strip_softbox_30x120"
      position: "back-left"
      distance: 1.0
    background_light:
      modifier: "softbox_60x60"
      position: "behind_subject"
      gel: "full_CTO"

  atmosphere:
    mood: "luxurious_elegant"
    haze: "none"
    bokeh_type: "clinical_clean"
    grain: "none"

Phase 3.5: Enhanced Prompt Package (internal)

Compile the enriched scene into a categorized prompt package. Each category maps to one controllable dimension for later iteration. Output is consumed by Phase 4 — do not report to the user.

prompt_package:
  original_prompt: "premium amber glass serum bottle on black stone"

  type:
    category: 產品海報               # image category, from task_profile
    use_case: "e-commerce hero image, product launch campaign, social media primary visual"

  subject: "amber glass serum bottle with gold dropper cap, placed center-left on polished black marble"

  composition:
    framing: "medium close-up, full frame product focus, product occupies 40% of frame"
    layout: "full frame, product centered on golden ratio intersection, generous breathing room on all sides"
    subject_placement: "center-left, right 30% reserved as negative space for text overlay"
    depth_planes: "foreground marble slab, midground product, background soft warm gradient"
    negative_space: "right 30% and bottom margin — clear for headlines and tagline"
    margins: "equal left-right, slightly more bottom margin for visual weight"

  style_and_lighting:
    visual_language: "premium commercial photography, luxury studio aesthetic, editorial quality"
    medium: "photography"
    mood: "luxurious, calm, warm elegant, sophisticated"
    lighting:
      key: "warm directional soft light from camera-right, 45 degrees, moderate shadow"
      fill: "gentle warm fill from camera-left, one stop below key, open shadows"
      rim: "narrow rim accent from back-left for edge separation"
      pattern: "three-point cinematic with warm accent"
      quality: "soft diffused with smooth shadow transition"
    color_script: "vintage warm amber tone, selective saturation, warm gold highlights, complementary amber-and-charcoal palette"
    exposure: "middle key, moderate contrast, 3:1 key-to-fill ratio"
    lens: "85mm short telephoto at f/5.6, eye-level angle, moderate distance, natural proportions"

  background: "minimal studio tabletop, cream silk fabric draped on right side, soft warm gradient background"
  material: "high-gloss glass with sharp specular highlights, matte gold cap, polished marble with subtle reflections, silk fabric with soft sheen"
  atmosphere: "clean no haze, clinical sharp bokeh, no grain"

  constraints:
    must_keep:
      - "brand name LUMIÈRE visible on product label"
      - "gold dropper cap design detail"
      - "warm amber tone overall"
    no_extra_elements:
      - "watermark"
      - "human model or hands"
      - "price tag or barcode"
      - "extra text not specified in exact_text"
    exact_text_preservation:
      - text: "LUMIÈRE"
        location: "product label center"
        criticality: "high — must be fully legible"

  output_format:
    aspect_ratio: "1:1"
    transparent_background: false
    video_ready_frame: false
    output_medium: "digital"

  publication_layout:
    type: "none"
    publication_mode: "none"
    masthead_required: false
    cover_lines_required: false
    text_safe_zones: []
    text_policy: "none"

  footer: "warm amber gradient bar at bottom edge anchoring the composition, subtle marble texture strip grounding the product, brand color accent bar in charcoal"

  keywords: ["soft diffused key light", "warm colored key light", "gentle fill light", "narrow rim edge light", "three-point lighting setup"]
  quality_tags: ["premium commercial photography", "studio product shot", "sharp focus on product label", "8K detail"]
  avoid: []

Rules:

  • Each category is independently editable for targeted iteration.
  • style_and_lighting.lighting must match the Phase 3 lighting_diagram.
  • style_and_lighting.lens must match the Phase 2 lens_strategy.
  • subject must keep the original user intent visible.
  • constraints.exact_text_preservation must match Phase 1 exact_text.
  • The model-neutral prompt candidate = subject + composition values + style_and_lighting values + background + material + atmosphere + footer + keywords; Phase 4.6 renders this candidate into the target model-specific final syntax.

Phase 4: Self-Review Optimization (internal)

Take the Phase 3.5 categorized prompt package and run it through Phase 1 → Phase 2 → Phase 3 → Phase 3.5 a second time, treating the existing prompt as the "original input". This is a self-critique pass to find optimization opportunities. Output is consumed by Phase 4.5 — do not report to the user.

What to look for on the second pass

  1. Gaps — did Phase 1 miss any intent dimensions? (audience, format, brand constraints, risk flags)
  2. Contradictions — does the lighting diagram contradict the composition choice? Does the lens strategy fight the working distance?
  3. Ambiguity — are any categories vague or underspecified? Can a detail be grounded in a config entry?
  4. Consistency — does the color grade support the mood? Does the exposure strategy match the lighting diagram?
  5. Over-engineering — is every detail justified? Remove photography terms added without config backing.

Output

optimization_pass:
  iteration: 2
  changes_made:
    - category: "style.lighting"
      before: "..."
      after: "..."
      reason: "original key angle clashed with rim placement"
    - category: "style.composition"
      before: "..."
      after: "..."
      reason: "missed negative space requirement from Phase 1 format constraint"
  optimized_prompt: "<re-runs Phase 3.5 with improvements>"

Only make changes that are grounded in a config knowledge base entry. If everything is solid, publish the Phase 3.5 output as-is (no changes needed — silence is correct).

Rules

  • Maximum 1 self-review pass. Do not loop infinitely.
  • If the optimized_prompt is identical to the original Phase 3.5 output, omit changes_made.
  • Do not degrade quality — if the first pass is already solid, ship it.

Phase 4.5: Prompt Cleanup & Consistency Check (internal)

Before selecting the final renderer, clean the model-neutral prompt package for consistency and usability. This phase is not a compression phase and must not shorten the prompt merely to save tokens.

The goal is:

clean + non-contradictory + complete + renderer-ready

not:

short

A longer prompt is acceptable when the additional wording preserves subject identity, publication typography, vehicle/product/building integrity, material behavior, lighting direction, text/logo requirements, or exact user constraints. Output is consumed by Phase 4.6 — do not report to the user.

prompt_cleanup_consistency_check:
  principle: >
    Do not compress for length. Only remove semantic garbage, duplicate wording, contradictions,
    equipment leaks, avoid-list conflicts, and low-value filler that does not change the visible result.

  target_behavior:
    - "preserve visual specificity"
    - "preserve hard locks and user-requested details"
    - "preserve publication-layout and cover typography intent"
    - "preserve hero-subject integrity guardrails"
    - "preserve material, lighting, color, and composition details when they affect the image"
    - "remove true duplication and contradictions only"
    - "prepare a clean package for the selected model renderer"

  cleanup_actions:
    remove_or_merge:
      - repeated synonyms with identical meaning
      - repeated mood adjectives that do not add visible information
      - redundant quality tags
      - conflicting style modes
      - internal lighting equipment names in generator-facing fields
      - abstract cinematography theory that has no visible translation
      - avoid terms that contradict required positive elements
      - broad negative terms that suppress required text, logo, or cover typography
      - accidental raw YAML/JSON fragments in final prompt fields

    preserve_always:
      - original_user_intent
      - explicit_user_constraints
      - subject_identity
      - exact_text_or_logo_requirements
      - publication_layout_or_cover_typography_intent
      - model_target_and_renderer_requirements
      - shot_size_and_crop_guardrails
      - composition_and_subject_placement
      - style_realism_mode
      - key_lighting_direction_and_quality
      - material_critical_details
      - hero_integrity_guardrails
      - output_specs
      - task_specific_negative_guardrails

  no_shortening_rule: >
    If a phrase adds distinct visible information, preserves a user requirement, or prevents a known failure mode,
    keep it. Do not remove it just because the prompt is long.

  subject_minimum_detail_floor:
    human_portrait:
      require:
        - subject identity / demographic or character role
        - shot size
        - pose or expression
        - lighting direction
        - style realism mode
        - anatomy/crop guardrails when needed
    automotive:
      require:
        - vehicle identity
        - full or intended crop state
        - correct proportions / wheel guardrail
        - scene context
        - lighting direction
        - composition placement
    product:
      require:
        - product identity
        - geometry or label readability
        - material finish
        - lighting/reflection behavior
        - background or surface
    architecture:
      require:
        - space/building type
        - perspective accuracy
        - material palette
        - lighting condition
        - scale/layout readability
    publication_cover:
      require:
        - publication type
        - masthead/title intent
        - cover-line or text-safe layout mode
        - hero subject hierarchy
        - readable typography guardrails

How to run cleanup

  1. Preserve first. Identify all hard locks, subject identity, exact text/logo, publication intent, and model-specific requirements before removing anything.
  2. Remove only true garbage. Delete repeated synonyms, contradictory clauses, internal equipment leaks, and generic filler that does not affect the visible result.
  3. Fix contradictions before deleting. If positive requires text/logo/cover typography, do not use broad negatives like no text or text artifacts; replace them with random unreadable text, messy typography, or no random extra logos.
  4. Preserve publication intent. If the user asked for a magazine cover or poster, keep masthead/title/cover-line language unless the selected mode is explicitly text_safe_layout_only_mode.
  5. Preserve positive guardrails. For Z Image Base, Z Image Turbo, and SD-style models, keep positive protection such as "full car visible", "correct car proportions", "clean wheel shape", "label faces camera", or "straight vertical lines."
  6. Keep distinct visual instructions. Material behavior, lighting direction, color palette, camera angle, crop safety, and background context should remain when they change the image.
  7. Do not enforce a token or word budget here. Phase 4.6 only handles renderer syntax and field splitting; it must not shorten meaningful visual content just to fit a length target.
  8. When in doubt, remove decorative adjectives before visual nouns. "amber glass bottle" survives; "beautiful gorgeous premium amber" gets merged.
  9. Final package should feel complete and controllable, not skeletal.

Phase 4.6: Model-Specific Renderer Selection (internal)

Select the final prompt renderer from config/prompt-core/prompt-renderer.yaml using intent.model_target. This phase converts the cleaned universal prompt package into the target model's preferred final syntax. It must not enforce a token or word budget; advisory length ranges are only references for unusual user-requested brevity or known downstream limits. Do not report this internal renderer selection unless expert or debug output is requested.

Renderer Selection Rules

renderer_selection:
  default_model_target: "chatgpt_image"
  supported_targets:
    - chatgpt_image
    - z_image_base
    - z_image_turbo
    - comfyui_flux2_zimage_3stage
    - stable_diffusion
    - midjourney
    - flux
    - image_edit
    - video_prompt

  selection_priority:
    - explicit_user_model_target
    - workflow_or_platform_hint
    - output_medium_hint
    - default_model_target

  routing_examples:
    chatgpt_image:
      triggers: ["ChatGPT Image", "image2", "gpt image", "OpenAI image"]
      final_format: "natural_language_prompt_with_concise_avoid"
    z_image_base:
      triggers: ["Z Image Base", "z-image base", "z image base"]
      final_format: "dense_positive_prompt_and_negative_prompt"
    z_image_turbo:
      triggers: ["Z Image Turbo", "z image turbo", "turbo", "ComfyUI z image turbo"]
      final_format: "positive_prompt_and_negative_prompt"
    comfyui_flux2_zimage_3stage:
      triggers: ["Flux.2 Dev + Z-Image Base + Z-Image Turbo", "three-stage ComfyUI workflow", "shared positive prompt"]
      final_format: "shared_positive_prompt_and_z_image_negative_prompt"
    stable_diffusion:
      triggers: ["Stable Diffusion", "SDXL", "Pony", "checkpoint", "LoRA"]
      final_format: "positive_prompt_and_negative_prompt_tag_style"
    midjourney:
      triggers: ["Midjourney", "MJ", "--ar", "--style raw"]
      final_format: "compact_aesthetic_prompt_with_parameters"
    flux:
      triggers: ["Flux", "FLUX.1", "schnell", "dev"]
      final_format: "subject_first_clean_natural_language"
    image_edit:
      triggers: ["edit", "modify", "only change", "只改", "保持不變"]
      final_format: "change_only_with_must_preserve_and_forbidden_changes"
    video_prompt:
      triggers: ["i2v", "t2v", "video", "first frame", "影片"]
      final_format: "stable_first_frame_plus_motion_and_camera_movement"

Renderer Behavior

  • chatgpt_image: preserve semantic relationships, natural-language scene logic, and concise avoid instructions; when the source prompt is short, expand it into a complete prompt with subject, composition, light, material, background, atmosphere, output intent, and constraints.
  • z_image_base: use dense subject-first visual phrase chains with positive/negative split. Keep more composition, lighting, material, anatomy, expression/gaze, exact text, and geometry details than Turbo, but do not use ChatGPT-style paragraphs.
  • z_image_turbo: use subject-first visual-result phrases, reduce abstract cinematography terms, and output positive/negative prompts when supported. Preserve all prompt details that materially affect the image, including multi-subject roles, anatomy, expression/gaze, exact text, material, geometry, water/fabric physics, and scene support logic; do not shorten purely for a word target.
  • comfyui_flux2_zimage_3stage: output one shared_positive_prompt for Flux.2 Dev, Z-Image Base, and Z-Image Turbo, plus one z_image_negative_prompt used only by Z-Image Base/Turbo. The shared positive must contain no negative syntax or failure words because Flux.2 Dev consumes it directly.
  • stable_diffusion: use compact tag-friendly positive and negative prompts, keep the most important tags early, and avoid long prose.
  • midjourney: use compact aesthetic language and append requested parameters such as aspect ratio.
  • flux: use clean natural language with explicit subject-first ordering and minimal keyword noise.
  • image_edit: prioritize change_only, must_preserve, locked_attributes, and forbidden_changes over all aesthetic enhancements.
  • video_prompt: preserve a stable first frame, readable subject silhouette, motion-ready layers, and camera movement intent.

Output Contract

renderer_output_contract:
  model_target: "chatgpt_image | z_image_base | z_image_turbo | comfyui_flux2_zimage_3stage | stable_diffusion | midjourney | flux | image_edit | video_prompt"
  renderer_profile: "selected in Phase 4.6"
  source_prompt_package: "Phase 3.5 optimized and Phase 4.5 compressed package"
  final_fields_depend_on_renderer:
    chatgpt_image: [prompt, avoid, suggest_resolution, locked_constraints, user_status]
    z_image_base: [positive_prompt, negative_prompt, suggest_resolution, locked_constraints, user_status]
    z_image_turbo: [positive_prompt, negative_prompt, suggest_resolution, locked_constraints, user_status]
    comfyui_flux2_zimage_3stage: [shared_positive_prompt, z_image_negative_prompt, suggest_resolution, stage_usage, locked_constraints, user_status]
    stable_diffusion: [positive_prompt, negative_prompt, suggest_resolution, settings_notes, user_status]
    midjourney: [prompt, parameters, no_terms, suggest_resolution, user_status]
    flux: [prompt, avoid, suggest_resolution, user_status]
    image_edit: [edit_instruction, must_preserve, change_only, forbidden_changes, suggest_resolution, user_status]
    video_prompt: [first_frame_prompt, motion_prompt, camera_motion, avoid, suggest_resolution, user_status]
  phase_5_txt_export:
    required: true
    filename_pattern: "final_YYYYMMDD_HHMMSS.txt"
    output_directory: "ABSOLUTE_USER_DESKTOP/comfyui_prompt"
    forbidden_output_directories:
      - "skill folder"
      - "repository folder"
      - "current working directory"
      - "relative desktop/comfyui_prompt"
    desktop_resolution: "Resolve the real user Desktop first, then append comfyui_prompt. On Linux, respect XDG_DESKTOP_DIR from ~/.config/user-dirs.dirs; otherwise use the user's standard Desktop folder."
    create_output_directory_if_missing: true
    writer: "scripts/write_final_prompt_txt.py"
    export_default_renderer_families: [chatgpt_image, flux, z_image_base, z_image_turbo]
    include_requested_workflow_target: true
    collapse_identical_outputs: true
    fields_per_model_target:
      - positive_prompt
      - negative_prompt
      - suggest_resolution

Anatomy Rendering Rules

When anatomy_risk.level is high or critical, or when the lower body is clearly visible, apply config/character-identity/human-anatomy-guardrails.yaml before final rendering.

anatomy_rendering:
  required_when:
    - "full body"
    - "lower body visible"
    - "glamour pose"
    - "dance pose"
    - "stage pose"
    - "swimsuit editorial"
    - "body-conscious outfit"
    - "one-leg support pose"
    - "dynamic asymmetric stance"

  prefer_positive_language:
    - "realistic full-body anatomy"
    - "believable pelvis and hip structure"
    - "natural hip-to-thigh connection"
    - "knees aligned with thigh direction"
    - "consistent leg proportions"
    - "support foot firmly grounded"
    - "body weight naturally balanced over the planted leg"

  lower_body_negative_guardrails:
    - "warped pelvis"
    - "disconnected thighs"
    - "twisted knees"
    - "mismatched leg length"
    - "floating feet"
    - "impossible pose balance"
    - "deformed lower body"

Do not overuse generic anatomy tags when specific structure can be named.
For high-risk poses, specific lower-body rules outperform vague phrases like correct anatomy.

Intimate Anatomy Consistency Rendering Rules

When sex_characteristics.intimate_anatomy_lock is male_external, female_external, or user_specified, apply config/character-identity/sex-characteristics-control.yaml before final rendering.

Use safe consistency wording by default:

intimate_anatomy_rendering:
  female_external_lock:
    safe: "external anatomy remains consistent with the specified adult female subject"
    strict_negative_when_needed:
      - "glans"
      - "penile anatomy"
      - "testicles"
      - "male external genital anatomy"

  male_external_lock:
    safe: "external anatomy remains consistent with the specified adult male subject"
    strict_negative_when_needed:
      - "vagina"
      - "female external genital anatomy"

  user_specified_lock:
    rule: "explicit user-specified anatomy wins; do not force binary intimate anatomy"

Strict intimate-anatomy terms should be used only when the context is NSFW-sensitive, anatomy-sensitive, image repair/inpainting, or when the user explicitly asks to prevent wrong external sex characteristics.

Physical Collision Rendering Rules

When collision_risk.level is high or critical, or when the scene contains visible touch/contact/overlap/support relationships, apply config/scene-environment/physical-collision-guardrails.yaml before final rendering.

collision_rendering:
  contact_logic:
    - "physically plausible contact points"
    - "body touches the object without intersecting it"
    - "support surfaces visibly carry weight"

  occlusion_and_layering:
    - "correct occlusion order"
    - "clear front-to-back layering"
    - "readable partial overlap"

  material_interaction:
    - "material-specific highlights and reflections"
    - "surface texture reads clearly"
    - "contact shadows reinforce depth"

  avoid:
    - "clipping"
    - "merged geometry"
    - "floating contact"
    - "wrong occlusion order"
    - "cheap plastic look"

Use explicit overlap and support wording when composition, contact, or product realism depends on it.

Era / Culture and Environment / Atmosphere Rendering Rules

Natural Environment Rendering Rules

Apply the following modules when the image depends on large-scale natural spectacle or environmental physics:

  • config/scene-environment/atmosphere-cosmic-light.yaml
  • config/scene-environment/oceans-subaquatic-physics.yaml
  • config/scene-environment/mountains-terrains.yaml

Use them to preserve weather logic, sky/light logic, water physics, underwater readability, terrain layering, geology, and landscape scale.

Avoid:

  • generic flat sky treatment
  • weather not affecting surfaces or visibility
  • flat water with random foam
  • underwater scenes with air-like clarity
  • flat backdrop mountains with no terrain layering

Apply config/scene-environment/era-culture-library.yaml and config/scene-environment/environment-atmosphere-library.yaml when the final image depends on historical, cultural, regional, atmospheric, or sensory context.

Include era-consistent visual details, culturally coherent architecture/fashion/props/materials, period-appropriate signage or typography when relevant, time-of-day lighting, weather-specific surface response, air quality and mood cues, and purposeful particle effects. Avoid anachronistic props, generic cultural mashup, random fog, overdone bloom, muddy visibility, and inconsistent weather/lighting.

Typography and Multi-subject Rendering Rules

Apply config/text-brand-layout/typography-layout-system.yaml and config/character-identity/multi-subject-interaction-system.yaml whenever text layout or multiple subjects affect the final image.

typography_render:
  include:
    - "clear typography hierarchy"
    - "reserved negative space"
    - "subject-safe text placement"
    - "exact requested text only"
  avoid:
    - "random text"
    - "text over face"
    - "misspelled brand names"
    - "cluttered layout"

multi_subject_render:
  include:
    - "explicit subject count"
    - "distinct individual faces and silhouettes"
    - "clear group formation"
    - "readable spacing and occlusion"
    - "physically plausible interaction"
  avoid:
    - "merged faces"
    - "same face repeated"
    - "shared limbs"
    - "fused bodies"
    - "unreadable crowding"

Phase 5: Final Prompt — User Sign-Off

Assemble the final prompt using the renderer selected in Phase 4.6. The final_prompt is the only default user-facing output. The final prompt must be cleaned, internally consistent, and formatted for the selected model_target; if a cover/poster/publication request exists, it must also reflect the selected publication_mode. Do not dump verbose category output or raw YAML/JSON prompt packages into the generator-facing prompt field.

After the final prompt is assembled, generate a timestamped TXT export for the user. Save every generated TXT file in the user's real desktop folder comfyui_prompt, creating the folder if it does not already exist. Resolve the desktop path in a cross-platform way before writing: on Linux, respect XDG_DESKTOP_DIR from ~/.config/user-dirs.dirs; otherwise use the user's standard Desktop folder such as ~/Desktop or C:\Users\<user>\Desktop. Never treat desktop/comfyui_prompt as a relative path, and never write the export under the skill folder, repository folder, current working directory, or any local desktop/ folder. The filename must be final_ plus the current local date and time in YYYYMMDD_HHMMSS format, followed by .txt, for example final_20260618_214530.txt; if that filename already exists, create a new suffixed file such as final_20260618_214530_01.txt rather than overwriting. The TXT export is not limited to the single selected display renderer: it should prepare practical renderer-family variants for chatgpt_image, flux, z_image_base, and z_image_turbo by default, plus the explicitly requested workflow target such as comfyui_flux2_zimage_3stage when relevant. Each section must include:

  • Positive Prompt
  • Negative Prompt
  • Suggested Resolution

For chatgpt_image, flux, and midjourney, map the main prompt field to Positive Prompt and map avoid / --no terms to Negative Prompt. For comfyui_flux2_zimage_3stage, map shared_positive_prompt to Positive Prompt, map z_image_negative_prompt to Negative Prompt, and include the stage usage note so Flux.2 Dev uses positive only while Z-Image Base/Turbo use positive plus negative. If a renderer does not support a true negative prompt, write (not specified) rather than inventing one. Use config/prompt-core/output-specs.yaml to choose suggest_resolution from aspect ratio, output use, and workflow target.

If multiple model targets produce exactly the same Positive Prompt, Negative Prompt, and Suggested Resolution, do not duplicate separate sections. Write one section and list all equivalent targets in model_target, for example model_target: chatgpt_image, flux. If any target needs different syntax, different negative prompt behavior, or a different resolution, keep it as a separate section.

Use scripts/write_final_prompt_txt.py for deterministic file output whenever a machine-readable Phase 5 object is available. Do not pass a relative --output-dir; omit --output-dir so the script resolves Desktop automatically, unless an absolute custom directory is explicitly requested by the user.

ChatGPT Image Example

final_prompt:
  model_target: "chatgpt_image"
  renderer_profile: "natural_language_semantic"
  suggest_resolution: "1024x1536 or 1536x2048 portrait, depending on workflow budget"

  prompt: >
    A photorealistic waist-up fashion portrait of an adult woman standing between two semi-transparent frosted glass panels.
    She is placed on the left third with generous negative space to the right.
    Soft diffused light comes from camera-right, creating smooth facial shadows and natural catchlights.
    The image has natural skin texture, warm neutral tones, and a restrained luxury editorial mood.

  avoid: >
    Avoid cropped hands, plastic skin, random text, watermark, distorted anatomy, door handles, hinges, or glass doors.

  locked_constraints:
    - "frosted glass panels, not glass doors"
    - "no door handles or hinges"

  user_status: "pending | accepted | modified"

Z Image Turbo Example — Publication / Magazine Cover

final_prompt:
  model_target: "z_image_turbo"
  renderer_profile: "short_visual_direct"
  publication_mode: "implied_cover_text_mode"
  suggest_resolution: "1024x1536 portrait or 1536x2048 upscale target"

  positive_prompt: >
    photorealistic automotive magazine cover, Toyota Supra sports car, full car visible,
    parked on a sunlit city street at golden hour, Japanese woman casually leaning against the driver-side door,
    smiling at the camera, car slightly left of center, warm soft side light from camera-right,
    glossy metallic paint, correct car proportions, clean wheel shape, urban background softly blurred,
    large magazine masthead at the top, a few clean supporting cover lines,
    clear professional cover layout, high-contrast editorial look

  negative_prompt: >
    cropped car, cropped hands, extra fingers, blurry face, blurry car, distorted car proportions,
    warped wheels, random unreadable text, messy typography, watermark, flat lighting, cartoon

  locked_constraints:
    - "automotive magazine cover, not text-free hero image"
    - "Toyota Supra remains the hero subject"
    - "cover typography intent remains visible"

  user_status: "pending | accepted | modified"

ChatGPT Image Example — Exact Cover Text

final_prompt:
  model_target: "chatgpt_image"
  renderer_profile: "natural_language_semantic"
  publication_mode: "exact_text_mode"
  suggest_resolution: "1024x1536 portrait or 1536x2048 upscale target"

  prompt: >
    A photorealistic automotive magazine cover featuring a Toyota Supra sports car and a Japanese woman on a golden-hour city street.
    The Supra is the hero subject, fully visible with clean proportions and glossy metallic paint.
    Add a large readable masthead at the top reading "SUPRA STYLE" and a short supporting cover line reading "TOKYO STREET PERFORMANCE".
    Keep the cover layout professional, with clear editorial spacing, sharp subject focus, and softly blurred urban background.

  avoid: >
    Avoid cropped car, warped wheels, blurry face, distorted hands, random extra text, messy typography, watermark, and cartoon style.

  locked_constraints:
    - "masthead text: SUPRA STYLE"
    - "cover line text: TOKYO STREET PERFORMANCE"

  user_status: "pending | accepted | modified"

Stable Diffusion Example

final_prompt:
  model_target: "stable_diffusion"
  renderer_profile: "tag_style_positive_negative"
  suggest_resolution: "832x1216 portrait or nearest model-native aspect ratio"

  positive_prompt: >
    photorealistic portrait, waist-up, adult woman, frosted glass panels, subject on left side,
    negative space, soft side light, natural skin texture, sharp eyes, warm neutral palette, editorial photography

  negative_prompt: >
    cropped hands, extra fingers, bad anatomy, plastic skin, blurry face, watermark, random text, glass door, door handle

  user_status: "pending | accepted | modified"

Image Edit Example

final_prompt:
  model_target: "image_edit"
  renderer_profile: "change_only_preserve_rest"
  suggest_resolution: "match source image resolution unless the user requests resizing"

  edit_instruction: >
    Only change the pose to a cheerful peace-sign gesture and make the expression happy.

  must_preserve:
    - "same face"
    - "same hairstyle"
    - "same body shape"
    - "same outfit"
    - "same background"

  forbidden_changes:
    - "do not change identity"
    - "do not change clothing"
    - "do not change camera crop unless requested"

  user_status: "pending | accepted | modified"

The user may edit any part. After sign-off, the model-specific final prompt and the generated final_YYYYMMDD_HHMMSS.txt file are ready for generation.

References

  • config/visual-cinematography/lighting-diagrams.yaml: light modifier dictionary and standard lighting patterns.
  • config/visual-cinematography/camera-lens-library.yaml: lens characteristics and selection guide.
  • config/visual-cinematography/color-palette-schemes.yaml: color theory, temperature, and grading looks.
  • config/visual-cinematography/exposure-strategies.yaml: exposure and contrast strategies.
  • config/visual-cinematography/composition-library.yaml: composition rules, framing types, and aspect ratio guides.
  • config/visual-cinematography/expert-visual-synthesis-system.yaml: cross-domain synthesis for frame ratio, wardrobe physics, skin realism, wet cloth, lighting physics, 3D lookdev, and motion cues.
  • config/visual-cinematography/visual-mood-library.yaml: visual language, medium, and mood lexicon.
  • config/prompt-core/output-specs.yaml: output format specifications and delivery requirements.
  • config/visual-cinematography/motion-capture.yaml: motion capture strategies with shutter angle physics, biomechanical gait patterns (canine/equine/avian/bipedal), rigging physics (iso/cross velocity), action aperture by motion axis, and mechanical fluid dynamics.
  • config/prompt-core/constraints-library.yaml: common constraints, animal anatomy rules, and text preservation guidelines.
  • config/prompt-core/task-profiles.yaml: default lens, lighting, and color strategies per task type.
  • config/visual-cinematography/shot-taxonomy.yaml: shot size definitions and add_positive/add_negative guardrails.
  • config/visual-cinematography/camera-angle-library.yaml: camera angle psychology with visual effects and risks.
  • config/prompt-core/subject-rules.yaml: required fields and quality guardrails per subject type.
  • config/visual-cinematography/material-rendering-library.yaml: material visual properties and rendering guardrails.
  • config/visual-cinematography/style-realism-control.yaml: style mode definitions with primary/incompatible term pairs.
  • config/visual-cinematography/cinematography-grammar.yaml: high-level visual language patterns (film still, hero reveal, motivated lighting).
  • config/prompt-core/prompt-assembly-schema.yaml: canonical 20-segment prompt assembly order with conflict rules.
  • config/prompt-core/semantic-priority-rules.yaml: priority hierarchy (100→70) for attribute conflicts.
  • config/prompt-core/prompt-renderer.yaml: model-specific final prompt renderer profiles, syntax contracts, publication cover/poster rules, and model-specific hero-integrity boosts.
  • config/safety-mature/mature-content-control.yaml: mature-content age-safety controller for adult editorial, glamour, swimsuit, boudoir-inspired, body-conscious fashion, adult explicit content, gore, and adult cosplay prompts.
  • config/character-identity/human-anatomy-guardrails.yaml: anatomy-risk detection, lower-body failure taxonomy, pose-biomechanics rules, and model-specific anatomy strengthening.
  • config/prompt-core/quality-rubric.yaml: image scoring rubric with cinematic dimensions.
  • config/prompt-core/failure-taxonomy.yaml: visual failure to targeted fix mapping.
  • references/iteration-policy.md: quality and iteration behavior.

Style Conflict Resolution

When the user prompt mixes terms from different style modes (e.g., "photorealistic anime portrait"), the style conflict resolution rules determine which mode takes priority and which terms must be removed or downweighted. Always reference config/visual-cinematography/style-realism-control.yaml for the base mode definitions.

style_conflict_resolution:
  photorealistic_vs_anime:
    if_user_says: ["photorealistic", "真人", "real-life"]
    remove_or_downweight:
      - "anime"
      - "cel shading"
      - "cartoon"
    preserve_if_character_reference:
      - "character-inspired costume"
      - "signature color palette"
      - "recognizable hairstyle influence"

  anime_realism_vs_cartoon:
    if_user_says: ["anime realism", "高端日系", "半寫實日漫"]
    keep:
      - "stylized facial design"
      - "cinematic lighting"
      - "realistic body proportions"
    avoid:
      - "chibi"
      - "flat cartoon"
      - "raw documentary photo"

  general_rule: >
    When style terms conflict, medium-defining terms outrank loose aesthetic references.
    Correction language such as "but", "instead", "rather than", "不要", "更接近", "不是...而是..."
    overrides earlier style terms. A named anime/game character reference is a subject or design cue,
    not automatically the rendering medium. Never preserve incompatible medium terms as equal-weight style modes.

Critical Rules

For anime-realism, high-end anime illustration, luminous character scenes, ChatGPT Image 2.0-like anime quality, festival/garden scenes, translucent outfits, wet reflective floors, delicate accessories, or small-prop character actions, classify the task with config/visual-cinematography/luminous-anime-detail-system.yaml. Use it to organize detail richness into light transmission, garment construction, color roles, depth planes, prop contact, wet-surface physics, face expression, and small text/sign handling. Do not express richness by adding random ornaments or repeated quality adjectives.

  • Every enhanced detail must reference a valid entry from config/ knowledge bases.

  • Internal Phase outputs should be machine-readable YAML for auditability, but default user-facing output should only include Phase 5's model-specific final_prompt plus the generated final_YYYYMMDD_HHMMSS.txt export.

  • Keep the original user prompt visible in the internal package — never lose it.

  • Each prompt category is independently controllable — never merge categories into one opaque blob.

  • Describe lighting by its visual effect, never by equipment name. Image generators render "octabox" or "white V-flat" as physical objects. Instead of "octabox key light", use "soft diffused key light with round catchlight". The equipment detail stays in the lighting_diagram YAML for human reference, not in the final prompt string.

  • Do not add random photography terms. Every modifier, lens, and color choice must be justified by the intent.

  • Phase 4.5 is cleanup, not compression. Never remove meaningful visual information only to make the prompt shorter.

  • When the user states token usage is not a concern, preserve richer visual, layout, material, typography, and constraint details.

  • Do not clean below the minimum detail floor required by the subject type, task profile, model target, or publication layout.

  • Mature styling must be routed through config/safety-mature/mature-content-control.yaml; preserve adult explicit content when requested while enforcing only the underage / childlike block.

  • For mature styling, require adult / 21+ / mature adult wording when sensual framing is present.

  • Do not add or preserve underage or childlike cues in mature/sexualized prompts.

  • Do not rewrite adult nudity, adult genital visibility, or gore into non-explicit editorial language unless the user asks for that style.

  • Route full-body, lower-body-visible, glamour, swimsuit, cosplay, dance, and dynamic human poses through config/character-identity/human-anatomy-guardrails.yaml.

  • For anatomy-risk prompts, use specific structure language such as pelvis, hip-to-thigh connection, support leg, knee direction, leg proportions, and grounded feet instead of only saying correct anatomy.

  • Do not remove anatomy guardrails during cleanup when the pose is asymmetric, dynamic, low-angle, or lower-body-focused.

  • For image-edit repair tasks, anatomy fixes must preserve identity, face, hairstyle, wardrobe, lighting, background, and overall pose intent unless a broader change is explicitly requested.

  • Route adult body-presentation consistency through config/character-identity/sex-characteristics-control.yaml when the prompt specifies male, female, androgynous, trans, nonbinary, gender-fluid, or user-specified presentation.

  • Explicit user-specified body presentation wins over automatic male/female cleanup.

  • Do not force binary body traits when the user explicitly requested androgynous, trans, nonbinary, gender-fluid, or user-specified presentation.

  • Use non-explicit visible body-structure language such as masculine torso, flat chest, feminine torso, natural bust shape, or gender-neutral body presentation only when needed for consistency.

  • For external intimate-anatomy consistency, use safe consistency wording by default and strict anatomical negative terms only when the context requires it.

  • Female-coded subjects must not gain unintended male external anatomy unless explicitly requested.

  • Male-coded subjects must not gain unintended female external anatomy unless explicitly requested.

  • Explicit user-specified anatomy wins over automatic intimate-anatomy cleanup.

  • Route contact-heavy, overlap-heavy, support-heavy, prop-handling, seated, leaning, vehicle, furniture, product, and multi-subject scenes through config/scene-environment/physical-collision-guardrails.yaml.

  • Specify physical relationships explicitly: what touches what, what supports what, and what occludes what.

  • Use contact-shadow, surface texture, gloss, hardness, and reflected-light cues when material realism matters.

  • Route standing full-body, fashion, dance, action, combat, poster, and character-sheet poses through config/character-identity/standing-pose-balance-system.yaml; define support leg, center of gravity, foot contact, knee direction, pelvis/shoulder counterbalance, and ground shadow.

  • Route reclining, lounging, side-lying, propped-on-elbows, propped-on-hands, bed, sofa, and mattress poses through config/character-identity/reclining-pose-physics.yaml; define load-bearing points, pelvis/spine/leg direction chain, hand/elbow support, surface response, and occlusion order.

  • Route beds, sofas, cushions, pillows, blankets, towels, mattresses, and padded supports through config/scene-environment/soft-surface-contact-physics.yaml; add compression, wrinkle direction, fabric tension, contact shadows, and surface deformation at the load points.

  • Route screen viewing, product handling, weapon/tool use, phone/TV/monitor interaction, sign reading, and object-dependent reactions through config/scene-environment/action-object-attention-system.yaml; align gaze target, hand contact, body orientation, object placement, and light/reflection response.

  • For group pool or water-play scenes, route through config/character-identity/multi-subject-interaction-system.yaml, config/scene-environment/physical-collision-guardrails.yaml, config/scene-environment/oceans-subaquatic-physics.yaml, and config/character-identity/pose-gesture-body-language.yaml; define the splash source, hand action, droplet direction, varied reactions, consistent water depth, planted feet, wet swimwear, and readable hand occlusion.

  • For ensemble selfie or cosplay-party scenes, route through config/character-identity/multi-subject-interaction-system.yaml, config/visual-cinematography/composition-library.yaml, config/character-identity/pose-gesture-body-language.yaml, and config/prompt-core/prompt-renderer.yaml; define a character design matrix, controlled wide-angle selfie perspective, face cluster geometry, foreground arm logic, varied expressions, rich outfit/accessory motifs, and shared lighting/material unifiers.

  • Treat plain selfie / 自拍 as a close selfie camera relationship, not a phone prop action: keep the phone, mirror, and selfie stick off-frame unless an explicit trigger requests holding phone, phone visible, mirror selfie, selfie stick, 拿手機自拍, 手機入鏡, 對鏡自拍, or 自拍棒.

  • For semi-reclined bed/sofa poses, preserve pelvis-leg direction chain, arm-support load path, visible fabric compression, and contact shadows through cleanup and renderer output.

  • For pool-edge exit poses, do not rewrite hand support into arms raised overhead, holding an overhead ledge, lifting the pool edge, gripping underside of tiles, or hanging from pool edge; use positive support wording before these negatives.

  • For towel-slip / after-bath prompts, route through config/character-identity/fabric-opacity-translucency-logic.yaml, config/character-identity/wardrobe-outfit-construction.yaml, and config/scene-environment/physical-collision-guardrails.yaml; do not turn an opaque towel reveal into transparent cloth, cutout exposure, floating fabric, or fabric clipping through the body.

  • In towel-slip adult prompts, preserve adult anatomical correctness without censoring adult intent: visible landmarks must be correctly placed and occluded by the towel edge instead of appearing on the towel surface or in impossible positions.

  • Do not remove collision/contact guardrails during cleanup when clipping, floating contact, or wrong occlusion order is likely.

  • Always select model_target before Phase 5. If unspecified, default to chatgpt_image.

  • Always detect publication-style intent before Phase 5. If the user asks for a cover/poster/封面/海報 and does not provide exact text, default to implied_cover_text_mode, not a text-free image.

  • Do not put no text, text artifacts, or equivalent broad text-suppression terms in the negative prompt when a cover/poster/publication layout is requested. Use random unreadable text, messy typography, or illegible cover lines instead.

  • If final typography will be added later in another tool, explicitly set publication_mode: text_safe_layout_only_mode; otherwise preserve visible masthead/title/cover-line intent.

  • Internal prompt packages may be YAML/JSON, but final prompts must be rendered into the target model's preferred syntax. Never pass raw prompt_package YAML/JSON directly to the image model unless the downstream workflow explicitly parses it.

  • For chatgpt_image, use natural language with clear semantic relationships, moderate detail, and concise avoid instructions.

  • For z_image_base, use dense visual-result phrases with positive/negative split. Keep subject, hard locks, anatomy/expression/gaze, exact text, composition, lighting, material response, background depth, and geometry integrity. Avoid ChatGPT-style paragraphs.

  • For z_image_turbo, use visual-result phrases, strong subject-first ordering, fewer abstract cinematography terms, and separate positive/negative prompts when supported; keep all details needed for image correctness even when the prompt becomes longer.

  • For SDXL, Pony, Illustrious, anime, or LoRA-based workflows, route through config/prompt-core/lora-routing-system.yaml; character LoRA priority outranks style LoRA when identity is at risk, and costume LoRA outranks generic fashion terms when marker preservation matters.

  • If the user requests compact output or the target renderer requires compact syntax, use config/prompt-core/model-specific-compact-prompt-generator.yaml; keep explicit constraints, identity locks, exact text, anatomy/contact, costume markers, mature/fabric logic, and aspect ratio before removing redundant style language.

  • For comfyui_flux2_zimage_3stage, output shared_positive_prompt and z_image_negative_prompt. Use the shared positive for Flux.2 Dev, Z-Image Base, and Z-Image Turbo; use the negative prompt only in Z-Image Base and Z-Image Turbo nodes.

  • In comfyui_flux2_zimage_3stage, never put no, avoid, bad, deformed, wrong, artifact terms, or age-safety negatives in shared_positive_prompt. Convert them to positive replacements and keep failure terms in z_image_negative_prompt.

  • In the shared positive, order content for the pipeline: composition skeleton, subject and hard locks, pose/expression/gaze, positive anatomy/geometry guardrails, camera/framing, lighting, material/fabric/skin detail, background, text/output. This lets Flux.2 Dev build layout before Z Image stages rebuild quality.

  • For both z_image_base and z_image_turbo, put positive integrity guardrails before negative bans: natural anatomy, clear gaze, natural facial expression, realistic skin texture, natural hands if visible, exact requested text, accurate product geometry, correct vehicle proportions, straight architectural verticals.

  • For Z Image negative prompts, keep only high-risk failures: deformed anatomy/hands, dead eyes/forced grin, wrong gaze, random unreadable text, misspelled text, fake logos, warped product/vehicle/architecture geometry, watermark, and age-safety negatives when mature content is present.

  • Do not enforce length budgets in Phase 4.6 or Phase 5. Advisory length ranges may inform readability, but subject identity, character matrices, anatomy, expression/gaze, pose mechanics, text, geometry, water/fabric physics, and scene logic always outrank brevity.

  • For stable_diffusion, output positive_prompt and negative_prompt separately, using compact tag-friendly phrasing.

  • For midjourney, output a compact aesthetic prompt with aspect ratio parameters when requested.

  • For flux, use clean natural language with explicit subject-first ordering and minimal keyword noise.

  • For image_edit, prioritize change_only, must_preserve, locked_attributes, and forbidden_changes over all aesthetic enhancements.

  • For video_prompt, preserve stable first-frame composition, readable subject silhouettes, motion-ready layers, and camera movement intent.

  • config/scene-environment/environment-atmosphere-library.yaml — environment mood and sensory atmosphere controller. Defines time of day, weather, season, humidity, air quality, mood, sensory cues, lighting cues, particles, ambient motion, and visibility control.

  • config/scene-environment/atmosphere-cosmic-light.yaml — natural light, celestial light, atmospheric optics, and extreme-weather controller. Defines sunlight/moonlight/aurora/lightning logic, sky states, volumetric effects, storm drama, and subject readability under environmental spectacle.

  • config/scene-environment/oceans-subaquatic-physics.yaml — ocean, shoreline, surf, and underwater-physics controller. Defines wave state, tide, foam, wetness, caustics, underwater light filtering, refraction, buoyancy cues, and subject-water interaction.

  • config/scene-environment/mountains-terrains.yaml — mountain, terrain, and landform controller. Defines landform type, terrain layering, geology, snowline, vegetation zones, atmospheric perspective, and large-scale terrain readability.

  • config/edit-reference/image-edit-repair-templates.yaml — image edit and repair controller. Defines edit scope, change targets, preservation constraints, localized repair templates, and anti-drift edit logic.

  • config/character-identity/identity-character-consistency.yaml — identity and character consistency controller. Defines locked traits, allowed variations, anti-drift rules, and series continuity.

  • config/character-ip/character-ip-database.index.yaml — lightweight character lookup index. Check exact keys, display names, English display names, aliases, and English aliases before loading full character data.

  • config/character-ip/character-ip-database.yaml — full reusable character profile schema and split-database router. Load the matched per-IP file after index lookup, verify aliases[character_key], then merge base_profiles[character_key] with enhanced_profiles[character_key] before building the character identity lock from visual anchors instead of name repetition.

  • config/character-identity/character-ip-identity-system.yaml — character IP identity controller. Defines silhouette, hair, face impression, palette, costume markers, signature props/emblems, worldbuilding motifs, allowed variations, and forbidden drift.

  • config/character-identity/character-archetype-visual-system.yaml — character archetype controller. Defines personality-to-visual behavior mapping for gaze, expression, posture, gesture, camera relation, and emotional range.

  • config/text-brand-layout/text-rendering-accuracy.yaml — text accuracy controller. Defines exact wording, multilingual integrity, text hierarchy, and no-extra-text suppression.

  • config/edit-reference/reference-image-control.yaml — reference-role controller. Defines identity/pose/style/composition/light roles, preserve-vs-reinterpret logic, and reference priority.

  • config/text-brand-layout/brand-commercial-visual-system.yaml — commercial art-direction controller. Defines commercial hero, hierarchy, copy-safe space, logo-safe space, brand tone, and ad usability.

  • config/character-identity/hero-action-character-scene-guardrails.yaml - hero action character scene controller. Use for recognizable character hero shots, wanted/bounty posters, torn paper props, open shirts/jackets, victory poses, kneeling/crouching hero poses, and cinematic action portraits where identity, pose, prop, text, cloth, and lighting must stay coherent.

Image Edit / Repair Templates

If the task is an edit, selective repair, localized replacement, cleanup, background swap, or attribute change on an existing image, classify it with config/edit-reference/image-edit-repair-templates.yaml.

This module controls:

  • what changes
  • what stays the same
  • local vs global edit scope
  • repair templates for face / hands / legs / outfit / background / text
  • anti-drift preserve rules

Identity / Character Consistency

If the task depends on keeping the same person, same character, same product identity, or same subject across edits/variations/series, classify it with config/character-identity/identity-character-consistency.yaml.

If the task includes reference images, classify each reference with config/character-identity/reference-image-override-system.yaml and config/edit-reference/reference-image-control.yaml; explicit text constraints win, reference images override only their declared role, and per-IP identity locks remain active unless the user asks to change them.

If the task asks for a character sheet, turnaround, multi-view, expression sheet, prop callout, or model sheet, route through config/character-identity/character-sheet-multiview-system.yaml and keep silhouette, hair, face, palette, costume markers, and props consistent across all views.

If the task includes multiple named characters, a team lineup, ensemble poster, party selfie, or battle group, route through config/character-identity/group-composition-identity-preservation.yaml; build an identity matrix so each character keeps distinct hair, face, palette, costume marker, pose language, spatial position, and interaction role.

If the task asks for a recognizable character, original character IP, named character, cosplay, character-inspired design, reinterpretation, same character across variants, or stronger character identity, first check config/character-ip/_indexes/character-lookup-index.yaml for a migrated formal profile. When it matches, load the referenced character folder and apply its component files; honor its validation.yaml review status. If it has no match, use config/character-ip/character-ip-database.index.yaml or config/character-ip/character-ip-lookup-index.yaml, load the mapped legacy per-IP file from config/character-ip-by-ip/, verify aliases[character_key], and merge base_profiles[character_key] with enhanced_profiles[character_key]. The formal profile takes priority only for its own character, and legacy remains the fallback for every unmigrated character. Then classify it with config/character-identity/character-ip-identity-system.yaml and config/character-identity/character-archetype-visual-system.yaml. Do not preserve character identity by repeating the name. Convert the character into visual anchors: silhouette, hair identity, face impression, palette hierarchy, costume structure, signature prop/emblem, archetype pose/expression, worldbuilding motif, allowed variation, and forbidden drift.

If the character task changes outfit, uses cosplay, asks for fashion/editorial reinterpretation, photoreal conversion, brand collaboration, or commercial-safe transformation, also classify it with config/character-identity/costume-marker-priority-system.yaml. Preserve the marker hierarchy before adding new styling: silhouette, color blocking, emblem/crest or safe motif, signature props/accessories, shoulder/collar/waist structure, material language, footwear/gloves, then micro trim.

If the character task also includes a hero shot, victory pose, dynamic kneeling/crouching stance, wanted poster, bounty poster, torn paper, held poster/sign, open shirt/jacket, wind-blown clothing, pirate/adventure scene, or dramatic cinematic action portrait, also classify it with config/character-identity/hero-action-character-scene-guardrails.yaml. Use it to connect the character identity lock to body support mechanics, hand-prop contact, poster text policy, paper tear physics, cloth-body collision, motivated lighting, and layered scene storytelling.

This module controls:

  • locked traits
  • allowed variations
  • recurring visual markers
  • face / hair / body / costume continuity
  • anti-drift safeguards
  • character IP visual-recognition anchors
  • archetype-specific gaze, expression, posture, gesture, and camera relation

Character IP priority order:

  1. silhouette and outline
  2. hair identity
  3. face impression and gaze
  4. color palette hierarchy
  5. costume structure and color blocking
  6. signature props, emblem, or markings
  7. personality pose and expression
  8. worldbuilding motifs
  9. rendering style
  10. decorative micro-details

For character IP prompts, separate:

  • locked_traits: hair silhouette, face impression, eye color, core costume palette, signature accessory, emblem, body silhouette
  • allowed_variations: pose, camera angle, lighting, background, seasonal outfit detail, expression within character range
  • forbidden_drift: changed hairstyle, wrong palette, missing emblem, generic outfit, wrong personality expression, face identity drift, missing signature prop

Text Rendering Accuracy

If the task includes titles, slogans, copy, labels, captions, multilingual text, or exact wording requirements, classify it with config/text-brand-layout/text-rendering-accuracy.yaml.

This module controls:

  • exact wording
  • text hierarchy
  • language/script integrity
  • no-extra-text constraints
  • clean typography rendering goals

Reference Image Control

If the task depends on one or more reference images, classify it with config/edit-reference/reference-image-control.yaml.

This module controls:

  • reference role assignment
  • preserve vs reinterpret logic
  • reference priority order
  • conflict resolution between references
  • prevention of accidental irrelevant copying

Brand / Commercial Visual System

If the task is a poster, product ad, campaign KV, website hero, magazine cover, billboard, or other marketing/commercial visual, classify it with config/text-brand-layout/brand-commercial-visual-system.yaml.

This module controls:

  • commercial hero hierarchy
  • headline/logo/CTA-safe layout zones
  • brand presence strength
  • product vs person balance
  • commercial usability and ad readability

Mature Content Tiering System

Use config/safety-mature/mature-content-tiering-system.yaml together with config/character-identity/fabric-opacity-translucency-logic.yaml when the prompt involves adult glamour, boudoir-inspired styling, thin fabric, sheer fabric, wet clothing, cleavage, implied form, artistic nude, or mature visual mood.

Core principle:

Do not globally erase mature cues, and do not block adult explicit content.
Use tiered mature control plus physical fabric plausibility.

Mature tiers:

mature_tier:
  - neutral
  - sensual_glamour
  - implied_contour
  - plausible_translucency
  - adult_nude
  - explicit_adult_content

Adult-content default:

adult glamour,
soft cleavage,
mostly opaque fabric with gentle body contour,
faint physically plausible silhouette through semi-sheer backlit fabric,
subtle outline visibility caused by damp fabric cling,
adult artistic figure-study framing when explicitly requested

Only forbidden in mature/sexualized renderer:

childlike,
underage

Nipple/areola handling:

Preserve requested adult visibility. Add semi-sheer material, wet cling, strong backlight,
mesh/lace structure, or artistic context only when it improves visual coherence.

Genital handling:

Adult genital exposure/focus is not blocked by the mature-content module.
Only underage or childlike sexual framing is blocked.

Categories