Seedream V5 — Complete Human Prompting Guide
Welcome to Seedream V5
There's a moment every image-generator user knows: you type a careful prompt, hit generate, and get something that looks like your description was run through a blender. Two objects merged. The lighting is wrong. The text is garbled. You spend twenty minutes trying to fix it by adding more adjectives.
Seedream V5 is different — and the difference runs deep.
Most image models work like this: your text gets encoded into numbers, and those numbers steer a diffusion process toward pixels. It's essentially a very sophisticated pattern-matcher. Seedream V5 adds a step before that: reasoning. The model reads your prompt the way a thoughtful art director would — parsing what you want, how the elements relate, what the lighting physics should look like, what the cultural context implies — and then it generates.
This reasoning-first architecture (described by ByteDance Seed) is why Seedream V5 can handle genuinely complex prompts without falling apart. A prompt describing five objects with specific positions and per-element colors doesn't just blend into visual noise — the model works out the spatial logic first.
On top of that reasoning core, Seedream V5 ships four features you won't find combined anywhere else:
- Real-time web search — the model can look things up mid-generation, pulling in current events, brand visuals, and live data.
- HEX color codes — drop
#FF006Edirectly into your prompt and get that exact hot pink, not "something pinkish." - JSON structured prompting — describe multi-subject scenes as structured data with per-element positions, colors, and actions.
- Cultural multi-language prompting — write in French and get Parisian light; write in Korean and get Hanok architecture. The language genuinely shifts the visual DNA of the output.
This guide walks you through all of it — from a quick-start copy-paste prompt to building production-ready commercial art direction from scratch.
Table of Contents
- Model Identity & Specs
- Core Architecture: Why Reasoning Matters
- Quick Start — 3 Copy-Paste Prompts
- Exclusive Features Deep Dive
- 4.1 Real-Time Web Search
- 4.2 HEX Color Code Precision
- 4.3 JSON Structured Prompting
- 4.4 Multi-Language Cultural Prompting
- 4.5 Example-Based Controllable Editing
- The Prompt Formula
- Prompting Best Practices
- Domain-Specific Prompting
- Image Editing
- Model Family: Choosing the Right Version
- Common Mistakes → Better Version
- Prompt Building Workshop
- Cheat Sheet
1. Model Identity & Specs
Identity
| Field | Value |
|---|---|
| Product Name | Seedream 5.0 (also referenced as Seedream V5) |
| Lite Variant | Seedream 5.0 Lite |
| Developer | ByteDance Seed AI |
| Launch Date | February 2026 (Preview: early February; Lite: late February) |
| Category | Image generation and editing |
| Access | Dreamina (Jimeng AI), Together AI API, fal.ai, WaveSpeed AI, Atlas Cloud, various third-party platforms |
| Architecture | DiT (Diffusion Transformer) with built-in Chain-of-Thought reasoning |
Technical Specifications
| Parameter | Seedream 5.0 | Seedream 5.0 Lite |
|---|---|---|
| Default Resolution | 2048×2048 (2K) | 2048×2048 (2K) |
| Max Resolution | 4096×4096 (4K) | 2K |
| Inference Speed | ~1.8s (2K image) | ~50 seconds (via API) |
| Parameters | ~20B | ~20B |
| Aspect Ratios | 1:1, 3:2, 4:3, 16:9, 21:9, 9:16, and custom | 1:1, 16:9, 9:16, and custom |
| Max Output Images | Up to 4 per request | Up to 4 per request |
| Output Formats | PNG, JPEG | PNG, JPEG |
| Seed Support | Yes (reproducibility) | Yes |
| Web Search | Yes (toggleable) | Yes |
| Text Rendering | Very good; 99%+ accuracy claimed | Very good |
| Multi-Image Consistency | Yes (character identity across batches) | Yes |
| Image Editing | Yes (natural language edits, style/color/lens transfer) | Yes |
Where to access: fal.ai, Together AI, WaveSpeed AI, Dreamina/Jimeng AI
2. Core Architecture: Why Reasoning Matters
Most image models are fancy pattern-matchers. You describe something, and the model finds the visual pattern that best fits your description — amazing for straightforward prompts, unreliable for complex ones.
Seedream V5 takes a different approach. It's a reasoning-first image model, built on a DiT (Diffusion Transformer) backbone with integrated Chain-of-Thought (CoT) reasoning. Before a single pixel is computed, the model runs a four-stage process:
- Parse — reads the prompt and decomposes it into components: subject, materials, lighting, spatial layout, intent, and context.
- Reason — works through spatial relationships, physics, and lighting consistency. Where does that shadow fall if the light is upper-left? How does candlelight interact with hammered copper? How do five objects arranged in a flat-lay relate to each other in perspective?
- Resolve — synthesizes those inferences into a coherent visual specification.
- Generate — only now does pixel generation begin, working from a well-formed internal model of the scene.
What this means for you practically:
- Complex multi-element prompts work. Describe five products with specific positions and individual colors — you'll get a coherent scene, not a fusion blob.
- The model understands intent, not just nouns. Saying "a product hero image" tells it something about hierarchy and composition, not just subject matter.
- Physics is respected. Glass refracts correctly. Metals scatter light realistically. Fabric shows proper drape and sheen. You don't need to describe physics laws — the model reasons them out.
- Long, layered prompts don't degrade. Where other models start to lose coherence around the 100-word mark, Seedream V5's reasoning step keeps everything organized — though you should still keep prompts under 200 words to avoid internal contradictions.
This architecture is also what makes web search work so well (Section 4.1): the model reasons about what information it needs to look up before starting generation.
3. Quick Start — 3 Copy-Paste Prompts
No preamble, no setup. These three prompts work as-is. Copy, paste, and generate.
Prompt 1 — Product Photography
A premium glass perfume bottle with a faceted stopper, product placed on a slab of #F5F0E8 warm ivory marble. The bottle contains a liquid in #D4A853 warm amber gold. Three-point studio lighting with a key light from the upper right, a soft fill from the left, and a subtle rim light creating a highlight edge. Background fades from #FAFAFA off-white to #E8E0D5 warm greige. Commercial product photography, minimalist, clean shadows. 3:2 aspect ratio.Prompt 2 — Editorial Portrait
A 30-year-old woman with natural curly dark hair, wearing a rust-orange linen blazer over a white tee, sitting at a light wood cafe table. Late afternoon window light from the left casts warm golden tones across her face and creates a long shadow on the table. She's mid-laugh, looking slightly off-camera. An untouched cappuccino sits in the foreground, slightly out of focus. Editorial portrait photography, Kinfolk magazine aesthetic. Shot on Sony A7R V, 85mm f/1.8. 4:3 aspect ratio.Prompt 3 — Architecture
The main reading hall of a grand 19th-century public library, shot from the center of the hall looking toward the far end. Vaulted ceiling with cast-iron arches painted forest green and gold. Rows of solid oak reading tables with green-shaded brass lamps, all lit. Tall arched windows on both sides let in cool afternoon light. A few lone scholars visible in the middle distance. Shot on a tilt-shift lens for perfect verticals. Architectural photography, Richard Pare style. 16:9 aspect ratio.4. Exclusive Features Deep Dive
4.1 Real-Time Web Search
Seedream V5 is the only image model that can search the web during generation. When your prompt references something that exists in the real world — a brand, a landmark, a current event, a trending topic — you can turn on web search and the model will retrieve live information to ground its output accurately.
When to turn it ON:
- Current events, news, trending topics
- Real brand identities (logos, color systems, visual language)
- Specific landmarks, buildings, natural wonders
- Real products you want visually accurate representations of
- Time-sensitive content ("today's weather," "this week's top story")
When to turn it OFF:
- Fictional content where you don't want real-world anchoring
- Consistent/reproducible results where you use the same seed across a batch
- Stylized art where creative interpretation is the goal
Step-by-Step Tutorial: Using Web Search
Step 1: Enable the web search toggle in your interface (Dreamina/fal.ai/WaveSpeed AI — look for a "Web Search" or "Search" toggle near the prompt box).
Step 2: Write your prompt referencing something real and current. Be specific — the more precise your reference, the better the retrieval.
Step 3: Generate. The model will pull context before generating.
Step 4: If the output doesn't look right, try adding "verified" or "accurate" to your prompt, or narrow your reference (e.g., instead of "the Eiffel Tower," try "the Eiffel Tower as of 2025, current nighttime illumination").
Web Search Example Prompts:
Search for the current Aurora Borealis forecast for Tromsø, Norway. Generate a dramatic landscape photograph showing the actual predicted activity level tonight — auroras filling the sky above a dark fjord with snow-covered spruce trees in the foreground. Long exposure photography style, National Geographic quality.Generate a poster for today's top trending news headline, styled as a breaking news graphic with bold red and white typography. The headline should be displayed prominently as a rendered text element.Search for the current visual identity and flagship product of Bang & Olufsen. Create a minimalist product shot in their documented brand style — dark backgrounds, clean geometric forms, premium materials.Search for the architectural style of the Neue Nationalgalerie in Berlin. Create a new architectural rendering of a fictional building designed in that same vocabulary — flat roof, glass curtain walls, elevated on a plinth — in a forested setting at dusk.4.2 HEX Color Code Precision
This is one of the most practical features in Seedream V5, and one that immediately earns its keep in any commercial workflow. Instead of describing "a kind of dusty rose-ish pinkish beige," you drop in #C9A0A0 and get exactly that.
The basic rule: Drop the HEX code inline in your prompt, immediately followed by the color name. This redundancy is intentional — the code gives precision, the name gives context.
Works best on:
- Large surfaces: backgrounds, walls, product bodies, clothing
- Graphic elements: shapes, borders, text
- Gradients (from one HEX to another)
- Brand color systems
Step-by-Step Tutorial: HEX Color Prompting
Step 1: Identify the exact colors you need. Use a brand color guide, a tool like Coolors, or a color picker from your design system.
Step 2: Write your prompt as normal, but replace any color description with #HEXCODE color-name.
Instead of: "a dark teal background"
Write: "a #006D77 dark teal background"Step 3: For multi-element scenes with separate color control on each element, consider switching to JSON format (see Section 4.3).
Step 4: For gradients, specify both ends:
Background transitions from #1A1A2E deep navy at the bottom to #3A86FF electric blue at the top.Step 5: Generate and check. If the color isn't right, try adding a second descriptor:
#FF5733 vivid coral-orange → "a #FF5733 vivid coral-orange, almost tomato red"HEX Color Prompting Examples:
A reusable water bottle in #3D405B slate blue with a #F4F1DE cream silicone sleeve and a #81B29A sage green lid. Placed on a weathered concrete surface with moss. Product photography, natural diffused light from above, subtle long shadow.Minimalist business card design on #F8F4EF warm white cardstock. The company name "MORRO" appears in #2C2C2C near-black, using a thin serif typeface. A single horizontal rule in #C9A060 warm gold separates the name from the contact details below. Clean, Swiss design aesthetic.An abstract gradient poster: background sweeps from #0D1117 near-black at the left to #6E40C9 deep violet at the right. Overlapping translucent geometric circles in #3FB4D4 cyan and #F7A04A amber. Minimal Bauhaus-influenced graphic design.Mini Color Palette Reference
Here are some useful pre-built palettes for common use cases, with HEX codes ready to drop in:
Tech / Modern Dark
- Background:
#0D1117near-black - Primary text:
#F0F6FCoff-white - Accent:
#58A6FFelectric blue - Highlight:
#3FB950green
Luxury / Warm Cream
- Background:
#FAF7F2warm ivory - Primary text:
#1A1A18near-black - Gold accent:
#C9A060warm gold - Shadow tone:
#D4C5A9warm greige
Brand — Common Reference Colors
| Brand | HEX | Name |
|---|---|---|
| Coca-Cola Red | #F40009 | vivid red |
| Facebook Blue | #1877F2 | medium blue |
| Instagram Gradient Start | #F58529 | warm orange |
| Spotify Green | #1DB954 | vivid green |
| Tiffany Blue | #0ABAB5 | teal blue |
| Hermès Orange | #FF6600 | vivid orange |
| Chanel Black | #1A1A1A | near-black |
| Apple Silver | #A2AAAD | cool silver |
Nature / Earth Tones
- Sky:
#87CEEBlight blue - Forest:
#2D5016deep green - Earth:
#8B6914warm brown - Sand:
#F4D03Fwarm yellow - Stone:
#9E9E9Emedium grey
Pastel / Soft Lifestyle
- Blush:
#FFB3C1soft pink - Lavender:
#C8A2C8lilac - Butter:
#FFF3CDpale yellow - Mint:
#B2DFDBsoft teal - Peach:
#FFCCB3soft orange
4.3 JSON Structured Prompting
When you have a scene with multiple subjects and you care exactly where each one sits, what color it is, and how it behaves — a plain text prompt becomes a negotiation with uncertainty. JSON prompting removes that uncertainty.
Instead of describing a five-product flat-lay in one dense paragraph (and hoping the model figures out the layout), you declare each subject as a structured object with explicit attributes: position, color, action, size. The model reads your data structure the way a set designer reads a blocking sheet.
When to use JSON:
- Multi-subject scenes requiring precise placement (flat-lays, product lineups, food spreads)
- Per-element color control across several objects
- Commercial art direction with exact specifications
- Any time you need to control more than three distinct elements
When to stick to plain text:
- Single-subject images
- Creative/artistic prompts where model freedom is welcome
- Quick iterations and explorations
Step-by-Step Tutorial: Building a JSON Prompt
Step 1: List all the subjects in your scene. Don't worry about order yet — just inventory everything that needs to appear.
Step 2: Assign each subject a position. Use compass/clock positions: center, upper-left, lower-right, center-left, or descriptive: foreground-center, background-right.
Step 3: Assign each subject a color. Use HEX codes, color names, or both.
Step 4: Add any per-subject actions or states ("slightly tilted", "lid open", "steam rising", "casting a shadow").
Step 5: Set your scene-level properties: overall style, lighting, camera angle and specs.
Step 6: Paste the full JSON as your prompt. No quotation marks around the whole thing — just the raw JSON object { ... }.
Original Guide Example — Breakfast Flat-Lay
{
"scene": "flat-lay breakfast spread on a rustic wooden table",
"subjects": [
{"description": "avocado toast with poached egg", "position": "center-left", "color": "green, golden yolk"},
{"description": "ceramic coffee cup with latte art", "position": "upper-right", "color": "#8B4513 brown, white foam"},
{"description": "small glass of fresh orange juice", "position": "lower-right", "color": "#FFA500 orange"},
{"description": "folded linen napkin", "position": "lower-left", "color": "#F5F5DC cream"},
{"description": "silver fork and knife", "position": "beside center-left", "color": "metallic silver"},
{"description": "small vase with wildflowers", "position": "upper-left", "color": "purple, yellow, green"}
],
"style": "editorial food photography, overhead shot",
"color_palette": ["warm earth tones", "#F5F5DC", "#8B4513", "#2D5016"],
"lighting": "soft morning light from the upper right, gentle shadows",
"camera": {"angle": "top-down", "lens": "50mm", "depth_of_field": "moderate"}
}New Example 1 — Fashion Flat-Lay
{
"scene": "fashion editorial flat-lay on a pale oak wood surface",
"subjects": [
{"description": "oversized camel wool coat, folded and laid flat", "position": "center", "color": "#C4942A camel"},
{"description": "cream ribbed turtleneck sweater, partially tucked under coat", "position": "center-left", "color": "#FAF0E6 cream"},
{"description": "dark navy wide-leg trousers, folded in thirds", "position": "lower-center", "color": "#1A1F36 dark navy"},
{"description": "tan leather crossbody bag with gold hardware", "position": "upper-right", "color": "#A0785A tan leather, #D4AF37 gold"},
{"description": "white ankle socks with minimal ribbing", "position": "lower-right", "color": "#FAFAFA white"},
{"description": "chunky white leather sneakers", "position": "lower-left", "color": "#F8F8F0 off-white"},
{"description": "small sprig of dried pampas grass", "position": "upper-left", "color": "#D4C5A9 warm beige"}
],
"style": "fashion editorial flat-lay, @leefromamerica aesthetic, clean and aspirational",
"color_palette": ["#C4942A", "#FAF0E6", "#1A1F36", "#D4C5A9"],
"lighting": "bright even diffused daylight, no harsh shadows, shot near a large north-facing window",
"camera": {"angle": "top-down 90°", "lens": "35mm", "depth_of_field": "deep, everything sharp"}
}New Example 2 — Tech Product Lineup
{
"scene": "premium tech accessories product lineup on a #1C1C1E near-black matte surface",
"subjects": [
{"description": "over-ear noise-canceling headphones", "position": "far-left", "color": "#2C2C2E space grey", "action": "standing upright on headband"},
{"description": "slim laptop, lid closed", "position": "center-left", "color": "#3A3A3C graphite silver", "action": "closed, slight angle toward viewer"},
{"description": "wireless mechanical keyboard", "position": "center", "color": "#1C1C1E black with #F5F5F5 white keycap legends"},
{"description": "wireless trackpad", "position": "center-right", "color": "#E5E5EA light silver"},
{"description": "USB-C hub with visible ports", "position": "far-right", "color": "#2C2C2E space grey"},
{"description": "braided USB-C cable, loosely coiled", "position": "foreground-center", "color": "#5E5CE6 indigo purple"}
],
"style": "premium tech product photography, MKBHD / Marques Brownlee studio aesthetic",
"color_palette": ["#1C1C1E", "#2C2C2E", "#5E5CE6", "#F5F5F5"],
"lighting": "dramatic side lighting from the left with a single warm key light, subtle rim light from the right, deep shadows",
"camera": {"angle": "slightly elevated 20°", "lens": "85mm", "depth_of_field": "slight — closest item sharpest"}
}New Example 3 — Food Spread
{
"scene": "abundant mezze spread on a weathered terracotta-tiled Mediterranean table, outdoor setting",
"subjects": [
{"description": "white ceramic bowl of hummus topped with olive oil and paprika", "position": "center", "color": "#F5DEB3 wheat beige, #FF4500 paprika red"},
{"description": "wooden board with sliced pita bread, slightly charred", "position": "upper-center", "color": "#D2A679 golden brown"},
{"description": "small plate of marinated olives, mix of green and black", "position": "upper-left", "color": "green and black"},
{"description": "vine leaves stuffed with rice (dolmades), stacked", "position": "left", "color": "#556B2F dark olive green"},
{"description": "white plate of sliced grilled halloumi with char marks", "position": "right", "color": "#F5F5DC cream with golden grill marks"},
{"description": "small ceramic ramekin of tzatziki with dill", "position": "lower-right", "color": "white with green flecks"},
{"description": "scattered fresh herbs — mint, parsley, dill", "position": "throughout", "color": "vivid green"},
{"description": "small glass carafe of iced water with lemon slices", "position": "upper-right", "color": "clear glass, #FFF44F yellow lemon"}
],
"style": "Mediterranean food photography, Ottolenghi cookbook aesthetic, abundant and generous",
"color_palette": ["terracotta", "#F5DEB3", "#556B2F", "#D2A679"],
"lighting": "warm afternoon Mediterranean sunlight from the upper left, dappled through an overhead vine canopy",
"camera": {"angle": "overhead at 45°", "lens": "50mm", "depth_of_field": "moderate — edges slight blur"}
}4.4 Multi-Language Cultural Prompting
This is Seedream V5's most quietly powerful feature. Write a prompt in French and you get more than translated French words — you get Parisian light, ochre-tinted plaster walls, narrow Haussmanian windows. Write in Japanese and you get the visual vocabulary of East Asian aesthetics: specific material textures, a particular quality of negative space, subtle seasonal symbolism.
The model has learned to associate visual culture with language, so the language you write in acts as a cultural style anchor that goes deeper than any written instruction.
Language → Visual Style Map
| Language | Effect on Output |
|---|---|
| French | Parisian architecture, Mediterranean light quality, café culture |
| Japanese | East Asian aesthetic sensibilities, wabi-sabi textures, seasonal motifs |
| Korean | Hanok architecture, contemporary K-aesthetic, Korean visual culture |
| Arabic | Middle Eastern design patterns, zellige tiles, mashrabiya screens |
| Hindi | South Asian visual culture, Diwali/festival aesthetics, vibrant color systems |
| Italian | Mediterranean color palette, classical architectural style, Renaissance visual references |
| Russian | Bold constructivist influences, birch forests, specific Slavic folk motifs |
| Spanish | Moorish architectural echoes, terracotta and azulejo, golden-hour saturation |
Tips for Multi-Language Prompting:
- Match your language to the scene's cultural context for maximum authenticity.
- Mix languages: write the scene description in the native language, then add English technical terms (camera specs, aspect ratio) at the end.
- RTL scripts (Arabic), Devanagari (Hindi), Hangul (Korean), Cyrillic (Russian) all render correctly.
- 12+ languages have been tested with documented cultural visual style shifts (fal.ai Blog).
Step-by-Step Tutorial: Multi-Language Prompting
Step 1: Write the scene description in the language that matches the cultural context. If you're creating a Moroccan riad, write the scene in Arabic or Darija. If it's a Tokyo convenience store at night, write in Japanese.
Step 2: Keep technical/photographic terms in English at the end — they work better that way.
Step 3: If you want to mix authenticity with precision (e.g., Japanese aesthetic but also specific HEX colors), write the scene in Japanese, then switch to English for colors, camera, and style specs.
Multi-Language Example Prompts:
Korean — Hanok Village (from original guide)
북촌한옥마을의 좁은 골목길. 전통 기와지붕 사이로 겨울 오후의 부드러운 햇살이 비추고, 감나무에 남은 마지막 감이 매달려 있다. 한복을 입은 젊은 여성이 돌담 옆을 걷고 있다. Shot on Fujifilm X-T5, 23mm f/1.4, natural film simulation. 3:2 aspect ratio.Japanese — Autumn Temple
京都の古い禅寺の庭。秋の紅葉が境内を覆い、石畳の参道に落ち葉が積もっている。早朝の霞がかかった光の中、一人の僧侶が箒で葉を掃いている。Shot on Phase One XF, 110mm, f/4. Quiet, contemplative mood. 3:2 aspect ratio.(Kyoto ancient Zen temple garden. Autumn leaves cover the grounds, fallen leaves piled on the stone-paved approach. In early morning hazy light, a lone monk sweeps leaves with a broom.)
French — Parisian Market
Le marché du dimanche matin à Belleville, Paris. Des étals de légumes colorés s'alignent sous des parasols rayés bleu et blanc. Une vieille femme inspecte des artichauts, son cabas à carreaux rempli de provisions. Lumière douce d'octobre, légèrement brumeuse. Photographie de rue, Cartier-Bresson style. 4:3 format.(The Sunday morning market in Belleville, Paris. Colorful vegetable stalls lined under blue-and-white striped parasols. An old woman inspects artichokes, her checkered shopping bag full of provisions. Soft October light, slightly misty.)
Arabic — Moroccan Riad
فناء رياض مراكشي تقليدي في المساء. نافورة مركزية من الرخام الأبيض تعكس ضوء القناديل النحاسية. أزلاج ملونة تغطي الجدران. رائحة ورود وبرتقال في الهواء الدافئ. Shot on Leica M11, 28mm. Warm evening light. 1:1 aspect ratio.(Traditional Marrakchi riad courtyard in the evening. A central white marble fountain reflects the light of brass lanterns. Colorful zellige tiles cover the walls. The scent of roses and orange in the warm air.)
Italian — Amalfi Coast
Una mattina di luglio sulla costiera amalfitana. Barche da pesca colorate — azzurro, bianco, rosso — ormeggiate nel piccolo porto di Positano. Case pastello si arrampicano sulla scogliera. Luce obliqua delle dieci del mattino. Fotografia paesaggistica, colori vividi e saturi. Shot on Nikon Z9, 24-70mm. 16:9 aspect ratio.(A July morning on the Amalfi Coast. Colorful fishing boats — blue, white, red — moored in the small harbor of Positano. Pastel houses climb the cliff. Ten-o'clock oblique morning light. Landscape photography, vivid saturated colors.)
4.5 Example-Based Controllable Editing
One of the more sophisticated editing modes: you provide a before-and-after image pair to show the model a transformation, and it learns the pattern and applies the same transformation to new images. This is more consistent than describing the transformation in words, because you're showing it rather than telling it.
What it handles well:
- Color grading transfer (show it a film look, it applies that exact grade)
- Aging effects (show a face aging, it applies to other faces)
- Season changes (show summer-to-autumn transition, it replicates)
- Style conversions (show photo-to-illustration, it learns that style)
Best practices:
- Make your example before-and-after pair as clean and unambiguous as possible.
- The more distinctly different the before and after are in only one dimension (just color, just season, just style), the more predictable the result.
5. The Prompt Formula
Every good Seedream V5 prompt follows this structure (not all parts are always needed):
Subject > Setting > Style > Lighting > Technical| Component | What It Controls | Example |
|---|---|---|
| Subject | Main element; specify materials, textures, actions | "A weathered brass compass with a cracked crystal face" |
| Setting | Environment, time of day, background details | "resting on a hand-drawn nautical chart in a dimly lit captain's cabin" |
| Style | Photography style, art movement, film reference, publication aesthetic | "shot on Hasselblad 500C, Kodak Portra 400" |
| Lighting | Direction, quality, color temperature | "warm candlelight from the left, deep amber tones" |
| Technical | Camera specs, lens, depth of field, aspect ratio | "macro lens, f/2.8, shallow depth of field, 3:2 aspect ratio" |
Key principles:
- Not all five components are always needed — Subject + Style is often enough.
- More specifics = fewer model decisions = more predictable output.
- Keep prompts under 200 words. Beyond that, internal contradictions creep in.
- The formula is a guide, not a law. Some of the best prompts are two sentences.
6. Prompting Best Practices
6.1 The Three Levels of Prompt Detail
You can think of prompt complexity in three levels. Start at Level 1, then add layers when you need more control.
Level 1 — Simple: Basic subject + setting
A red ceramic coffee mug on a wooden table.Level 2 — Adding Context: Details, props, light source
A red ceramic coffee mug with steam rising, next to an open book and reading glasses, warm morning light from a nearby window.Level 3 — Full Control: Materials, actions, precise lighting, camera/film specs
A handmade red ceramic mug with uneven glaze and a chip on the rim, steam curling upward. It sits on a weathered oak table next to a dog-eared paperback and tortoiseshell reading glasses. Warm directional light streams through a mullioned window from the upper left, casting soft shadows. Shot on Mamiya RB67, Kodak Portra 160, f/4, shallow depth of field. Quiet, contemplative mood.The Level 3 version doesn't just add more words — it replaces generic terms with specific ones. "Wooden table" becomes "weathered oak." "Reading glasses" becomes "tortoiseshell reading glasses." These specifics are what turn a generic image into a distinctive one.
6.2 Text Rendering
Seedream V5 has claimed 99%+ text accuracy — and it shows. When you need legible text in your image, follow these rules:
- Always put text in quotation marks. This is non-negotiable.
"Today's Special"renders;Today's Specialmay not. - Short text works best. Single words and short phrases are near-perfect. Multiple long sentences start to degrade.
- Describe the text's visual style.
"neon lettering","hand-painted sign","engraved brass plate","chalk lettering". - Give the text a surface. Floating text in space looks wrong — anchor it to a billboard, a mug, a window, a book cover.
Example:
A vintage-style coffee shop chalkboard menu displaying "Today's Special: Lavender Latte $5.50" in hand-drawn chalk lettering with small decorative flourishes. Dark green chalkboard surface, warm interior lighting.6.3 Writing with Intent
Tell the model what the image is for. This shapes composition, hierarchy, and format in ways that "make it look good" never will.
Use intent markers like:
"poster"— expects a headline, visual hierarchy, negative space for text"product hero image"— clean background, subject-first composition"editorial photo"— candid feel, interesting crop, narrative tension"infographic"— data-forward, legible text, visual structure"UI mockup"— flat, clean, functional-looking"book cover"— title space at top, author space at bottom, atmosphere-forward
6.4 Layout Control
Don't leave layout to chance on important work. Call it out explicitly:
"Centered headline, subtitle beneath, clean margins, three-panel grid""Rule of thirds composition, subject in the left third, negative space right""Symmetric framing, centered subject, mirror reflection in surface below""Full-bleed image, no safe zone, extreme edges included""Lower third: title and byline. Upper two-thirds: image."
6.5 Style Anchors
Always include at least one style anchor — a reference that tells the model the visual DNA of the image. Without a style anchor, the model makes its own choice, which may not align with your vision.
Strong style anchors:
"minimal Swiss graphic design""editorial photography, New York Times Magazine""commercial product photography, Apple.com style""hand-drawn botanical illustration, 19th century""Wes Anderson color palette and symmetry""film noir, 1940s Hollywood""Bauhaus poster design""Kinfolk magazine aesthetic""National Geographic photography""Bon Appétit editorial food photography"
6.6 Spatial Language for Complex Scenes
When multiple subjects appear in one scene, vague spatial language creates ambiguity. Be explicit about where things are.
Positional vocabulary:
- Cardinal:
"on the left","in the upper right corner","at the bottom center" - Depth:
"in the foreground","mid-ground","background","on the horizon" - Relational:
"between them","behind","towering over","dwarfed by","leaning against" - Camera-relative:
"filling the frame","dead center","in the lower right third"
Avoid:
Two people at a café table.Use:
On the left side, a woman in a blue coat reads a newspaper. On the right side, a man in a grey sweater sips coffee. Between them, a small marble table with a single red rose in a bud vase.6.7 Negative Prompts
Negative prompts in Seedream V5 work — but don't use them preemptively. Generate first. Look at what's wrong. Then add targeted negatives to fix specific issues.
General safety net (add when you want a clean result):
Avoid: blurry, low quality, distorted, deformed, watermark, text overlay, cropped, out of frameTargeted fixes based on output issues:
Avoid: blurry hands, extra fingers, merged faces
Avoid: distorted architecture, bent verticals
Avoid: oversaturated colors, blown highlights
Avoid: anachronistic elements, modern objects in historical scenes7. Domain-Specific Prompting
Product Photography
Product photography rewards specificity about materials above anything else. The model's physics reasoning means that if you say "matte ceramic," it will render non-reflective diffuse surfaces correctly. If you say "anodized aluminum," you get that distinctive brushed-metal look. Be specific about materials, and the lighting will follow logically.
Original example:
Matte white ceramic, anodized aluminum base, brushed gold accents. Three-point studio lighting, clean white sweep background.Additional examples:
A glass gin bottle with a hand-stamped wax seal in #8B0000 deep red, filled with #E8D5B7 pale amber liquid. Label reads "BOTANICA No. 7" in letterpress serif type on cream stock. Placed on a rough slate surface with scattered juniper berries and dried orange peel. Studio lighting with a single overhead strip softbox. Warm, slightly amber color grade. 3:2 aspect ratio.A stainless steel vacuum travel mug in #2C3E50 charcoal, with a matte powder-coated finish. The surface shows subtle fingerprint smudges. Placed on a granite countertop next to car keys and a crumpled receipt. Hard directional morning light from a skylight overhead. Commercial product photography, utilitarian aesthetic, slight grain. 1:1 aspect ratio.A handmade leather wallet in #8B6914 cognac, opened to show interior card slots in natural tan. Placed on a dark walnut wood surface with dramatic single-source side lighting from the left, casting a long shadow to the right. Background is #1A1A1A near-black. The edge stitching is visible and deliberate. Product detail photography, premium craftsmanship aesthetic. 4:3 aspect ratio.Flat lay of a complete men's grooming set: safety razor in chrome, shaving brush with badger bristles, ceramic shaving bowl in matte white, a bar of soap with the text "OAKMOSS" embossed on it, and a glass bottle of aftershave. All arranged on a #1C1C1C near-black stone surface. Overhead diffused studio light. Grooming editorial, GQ magazine aesthetic. 1:1 aspect ratio.Editorial Photography
A 25-year-old Korean woman with short black hair and minimal makeup, wearing an oversized camel wool coat, walking through Bukchon Hanok Village in Seoul. Late afternoon light, long shadows. Shot on Fujifilm GFX 100S, 80mm, f/2. Gentle color grade, slight warmth. (Original example)An elderly fisherman in his late 70s, deep tan, weathered hands, mending a fishing net on a dock in a small Portuguese coastal village. He wears a dark navy wool sweater and a worn yellow rain hat. Overcast sky diffuses the light evenly. Documentary photography style, intimate and unhurried. Shot on Leica Q3, 28mm. 3:2 aspect ratio.A teenage boy in a Chicago public park, golden hour, holding a basketball under his arm and looking at his phone. He wears an oversized white tee and basketball shorts. His shadow stretches long across dry summer grass. Shot from a distance, candid feel. Telephoto, 200mm, f/4. Street photography, Frank Paulin influence. 3:2 aspect ratio.Two elderly women playing chess at a weathered outdoor table in Trastevere, Rome. They're deep in concentration. Around them: pigeons, cobblestones, a Vespa leaning against a wall. Dappled afternoon light through plane trees. Documentary photography, intimate and unposed. Shot on Nikon FM2, Ilford HP5 black-and-white film. 4:3 aspect ratio.Architecture
A brutalist concrete apartment building in post-Soviet Tbilisi, shot from a low angle. Dramatic cumulus clouds behind. Late afternoon raking light emphasizes the textured concrete. Architectural photography, Hélène Binet style. 4:3 aspect ratio. (Original example)Interior of a Nordic sauna in a contemporary Finnish lakehouse. Thermory ash wood walls and ceiling with a low, warm glow from hidden LED strips behind the benches. Steam rising from a large steatite stove. Through a small square window, a snow-covered pine forest and frozen lake are visible. Architectural photography, warm tones, Dezeen magazine quality. 3:2 aspect ratio.The underside of a curved concrete freeway overpass, shot looking up at the intersection of two ramps. The geometry creates an unexpected abstract composition of arcs, columns, and shadow. Shot in late afternoon when sun cuts through at a low angle, creating dramatic light bars across the concrete. Architectural photography, Julius Shulman-influenced geometry. 1:1 aspect ratio.A rooftop terrace garden on a contemporary Barcelona apartment building. Lush grasses, olive trees, and drought-tolerant planting in long Corten steel planters. A hammered copper water feature in the center. The terrace looks out over terracotta rooftops toward the Sagrada Família in the middle distance at dusk, still being lit by the last warm light. Architectural photography, RCR Arquitectes aesthetic. 16:9 aspect ratio.Food Photography
Overhead shot of a deconstructed sushi platter. Each piece placed precisely on a black slate surface. Garnished with microgreens and a thin line of wasabi. Soft box lighting from upper left. Steam visible from the miso soup bowl in the upper right corner. Editorial food photography, Bon Appétit magazine style. (Original example)A single slice of tres leches cake on a white ceramic plate, soaked through and glistening, topped with fresh mango slices fanned out and a single hibiscus flower. Shot from a 30-degree angle, close-up. The fork is in frame, mid-bite, with a small piece of cake on the tines. Dramatic key light from the left casts a long shadow. Food photography, Food52 editorial style. 3:2 aspect ratio.A cast iron skillet with a freshly made shakshuka — the tomato sauce is vivid #CC2200 deep red with whole blistered cherry tomatoes, four poached eggs nestled in the sauce with bright golden yolks, topped with white feta crumbles, fresh green herbs, and a swirl of olive oil. The skillet is on a rough stone surface with a torn piece of sourdough beside it. Overhead editorial food photography, warm tones. 1:1 aspect ratio.Close-up, macro shot of a cinnamon roll being pulled apart, showing layered flaky dough, dark brown cinnamon-sugar filling, and thick vanilla cream cheese frosting dripping off the sides. Shot on a black slate board. Steam rising. Key light from above and slightly behind, making the frosting glisten. Moody, dramatic food photography. 1:1 aspect ratio.Cinematic Scenes
Shot on ARRI Alexa, anamorphic lens, 2.39:1 widescreen. Kubrick symmetry. Kodak 5219 500T tungsten film stock. (Original example)A lone astronaut sits on the edge of a crater on an alien moon, looking out over an endless rocky plain. Two moons hang in a violet sky. The astronaut's helmet visor reflects the landscape — a miniature version of the scene behind us. Shot on ARRI Alexa 65, wide anamorphic lens, 2.39:1. Teal and orange color grade. Denis Villeneuve sci-fi aesthetic.An empty double-decker bus moving through a deserted London street at 4am, all lights on inside. Rain-slicked streets reflect the red bus and yellow streetlights. A fox sits on the pavement watching the bus pass. Long exposure, motion blur on the bus wheels. Cinematic, Christopher Doyle DOP style, Kodak Vision3 film look. 16:9 aspect ratio.Landscape Photography
The salt flats of Salar de Uyuni, Bolivia, at sunrise. Perfectly still water covers the flats to a shallow depth, creating a mirror reflection of the sky — pink and gold clouds above matching perfectly below. A lone cactus in the far left reflects in the water. Shot on Phase One XT, 35mm, f/11. Long exposure. Landscape photography, Peter Lik style. 16:9 aspect ratio.A Japanese cedar forest (sugi) in Yakushima Island, early morning. Ancient moss-covered trees with trunks three meters wide. Volumetric beams of light breaking through the canopy, illuminating drifting mist. A small stone torii gate almost entirely consumed by moss is visible in the middle distance. Shot on Hasselblad H6D, 50mm. 4:3 aspect ratio.The Faroe Islands coastline in winter — dramatic basalt sea stacks rising from a grey North Atlantic. Overcast sky merges with the water at the horizon. Waves crash at the base of the stacks in a burst of white foam. No human presence. Shot with a wide 16mm lens, slight lens distortion at the edges adds drama. Landscape photography, moody, desaturated Scandinavian palette. 16:9 aspect ratio.Golden hour at Antelope Canyon, Arizona. Narrow slot canyon walls glow in flowing layers of orange, amber, and deep crimson. A single shaft of sunlight enters from above and hits the sandy floor, creating a perfect cone of light. A lone figure stands in the beam, silhouetted. 10-second exposure. Landscape photography, maximally saturated warm tones. 4:3 aspect ratio.Data Visualization
The reasoning-first architecture makes Seedream V5 particularly strong at data visualization — it can reason through what a chart should look like before drawing it.
An elegant dark-theme data visualization poster showing global average temperatures from 1880 to 2024 as a horizontal warming stripe chart. Each year is a vertical stripe, color-coded from #0A3060 deep blue (coldest) through #FFFFFF white (average) to #8B0000 dark red (hottest). The stripes fill the full width. Title: "Global Warming, 1880–2024" in white, clean sans-serif at the top. Source credit at the bottom in small grey text. 16:9 aspect ratio.An infographic showing the water cycle — evaporation, condensation, precipitation, runoff, groundwater — as a cross-section diagram. Background transitions from #87CEEB sky blue at the top to #8B6914 earth brown at the bottom. Each stage has a clean icon and short label in white. Arrows show the flow direction. Scientific illustration style, textbook quality. 4:3 aspect ratio.A circular radial bar chart showing the world's 12 most spoken languages by native speakers, displayed as bars radiating outward from a center point. Each language has a distinct accent color. Clean white background, Inter typeface, labels at the end of each bar. Modern data visualization aesthetic, Nadieh Bremer inspired. 1:1 aspect ratio.8. Image Editing
Feed Seedream V5 a reference image plus an edit prompt. The model infers what you want to keep and what to change — you don't need to mask regions manually.
What works well:
| Edit Type | Example Prompt |
|---|---|
| Season/weather changes | "Transform this summer scene into a snowy winter landscape" |
| Time shifts | "Make it nighttime with city lights reflecting on the wet street" |
| Add objects | "Add a coffee cup to the table on the right" |
| Remove objects | "Remove the car from the background" |
| Material changes | "Change the wooden fence to wrought iron" |
| Style transfers | "Recreate this photograph as a watercolor painting" |
| Clothing changes | "Change the outfit to a red evening gown" |
| Background replacement | "Replace the background with a tropical beach at sunset" |
| Color grading | "Apply a moody desaturated film look, crush the blacks" |
| Aging/de-aging | "Age this portrait by 30 years, keeping the bone structure and eyes the same" |
Critical tips:
- Specify what to keep:
"Keep the face, pose, and lighting; change only the background"— this is one of the most important habits in editing. - Describe the end state, not the action:
"The sky is now a dramatic sunset with orange and purple clouds"rather than"Change the sky to a sunset". - Use before/after reference pairs (Section 4.5) when you need a consistent transformation style applied across multiple images.
9. Model Family: Choosing the Right Version
| Version | Best For | Avoid When |
|---|---|---|
| 5.0 | Real-time web search, current events, complex reasoning tasks, anatomical diagrams, architectural blueprints, data visualization, HEX color precision, JSON multi-element scenes, multi-language cultural output | You just need a quick aesthetic image iteration |
| 5.0 Lite | Rapid iteration, cost-sensitive workflows, prototyping, 2K output is sufficient | You need 4K output or the most complex reasoning |
| 4.5 | Portraits, visual beauty, highest aesthetic quality, deep editing, multi-image character consistency | Web search grounding matters, or complex logical reasoning is needed |
| 4.0 | Fast iteration at the lowest cost, agile production pipelines | Maximum image quality is the priority |
Choose 5.0 over alternatives when:
- Prompt references current events, brands, or real-world entities
- You need precise color via HEX codes
- You have complex multi-element scenes requiring JSON structure
- Multi-language cultural authenticity matters
- Logical reasoning tasks: anatomical diagrams, architectural blueprints, data visualization
- Character consistency across a batch (WaveSpeed AI)
10. Common Mistakes → Better Version
Mistake 1: Vague Adjectives
Before:
A beautiful, amazing coffee shop interior.After:
A narrow two-story coffee shop interior in Portland. Ground floor with exposed brick walls painted forest green, vintage brass fixtures, and a hand-painted menu above an espresso bar. A worn wooden spiral staircase leads to a mezzanine with mismatched armchairs. Morning light through a large street-facing window. Warm, inviting. Editorial photography, Kinfolk aesthetic. 4:3 aspect ratio.Why it's better: Every adjective is replaced by a concrete visual fact.
Mistake 2: Keyword Soup
Before:
8K ultra realistic hyper detailed dramatic cinematic amazing beautiful bokeh trending artstationAfter:
A dramatic wide-angle shot of a volcanic landscape at dusk. Solidified lava flows in the foreground, still steaming from the last eruption. Towering cinder cone in the mid-ground. A sky bruised with purple, orange, and deep crimson. Shot on Sony A7R V, 24mm, f/8. National Geographic landscape photography.Why it's better: Natural sentences with intent beat comma-separated adjective lists every time.
Mistake 3: Forgetting Quotation Marks for Text
Before:
A product label that says Botanica Gin No. 7 in serif type.After:
A product label displaying "Botanica Gin No. 7" in hand-set serif letterpress type on cream paper stock.Why it's better: Text in quotes is treated as a text-rendering instruction, not a description.
Mistake 4: No Style Anchor
Before:
A woman walking through a city.After:
A woman in a white linen dress walking through a quiet Athens street at midday, dappled shade from an overhead bougainvillea. She's mid-stride, looking slightly away. Documentary street photography, Magnum Photos aesthetic. Shot on Leica M10, 35mm f/2. 3:2 aspect ratio.Why it's better: The style anchor ("documentary street photography, Magnum Photos") sets the visual DNA.
Mistake 5: Vague Spatial Descriptions
Before:
A kitchen counter with a bunch of fruit, some cooking things, and a plant.After:
A white marble kitchen counter. On the left, a wooden bowl holds a pyramid of oranges and a bunch of green bananas. In the center, a cast iron skillet and a wooden spatula rest on a trivet. On the right, a small terracotta pot with a trailing pothos plant. Morning light from an unseen window to the right. Editorial food/lifestyle photography. 16:9 aspect ratio.Why it's better: Each element gets an explicit position.
Mistake 6: HEX Codes Without Color Names
Before:
A logo on a background of #3D2B1F with text in #F5E6D3.After:
A circular logo embossed in #F5E6D3 warm cream against a #3D2B1F dark espresso brown background. The logo mark is a minimalist coffee plant illustration above the wordmark "CULTIVAR" in a thin extended serif. Luxury branding aesthetic.Why it's better: The color name contextualizes the HEX code and helps the model place it correctly.
Mistake 7: Vague Edit Instructions
Before:
Make this photo look better and more professional.After:
Keep the subject, pose, and foreground objects identical. Change the background to a seamless studio backdrop in #F0EDE8 warm off-white. Relight the scene with soft three-point studio lighting: key light from upper left, fill from the right, subtle rim light from behind. Remove any distracting elements from the original background. Result should look like a professional commercial product shot.Why it's better: "Better" is undefined. The revision specifies every change and every preservation.
Mistake 8: Too-Long Prompt with Contradictions
Before:
A dark moody atmospheric underwater scene in an ancient temple at midnight with bright sunlight streaming through the windows and a vibrant colorful tropical reef with neon bright fish and also a realistic black-and-white photograph from the 1930s and an oil painting in vivid saturated colors with a minimalist color palette of just black and white showing a crowded baroque detailed scene with lots of negative space.After:
An ancient stone temple submerged underwater, colonized by a coral reef. Fish school through a collapsed archway. Shafts of filtered green-blue light break through gaps in the ceiling from the surface above. The mood is eerie and beautiful — civilization reclaimed by nature. Underwater photography, wide angle, 16mm, slightly desaturated, teal color cast. 16:9 aspect ratio.Why it's better: The original contradicts itself (midnight + bright sunlight, dark moody + neon bright, oil painting + photograph). The revision picks a coherent single vision.
Mistake 9: Listing Negatives in the Main Prompt
Before:
A mountain landscape — no people, no buildings, no roads, no powerlines, no fences, no cars, no aircraft, no litter, no modern elements, no artificial objects.After:
A completely untouched mountain wilderness — rocky peak rising above the treeline, glacial lake in the foreground, dense old-growth forest on the slopes. No signs of human presence anywhere. Landscape photography, pristine wilderness aesthetic. Shot on Horseman SW617, 90mm. 16:9 aspect ratio.
Avoid: any human-made structures, powerlines, roads, or modern objects.Why it's better: Use the main prompt to describe what you want. Use the Avoid/negative prompt to list specific exclusions.
Mistake 10: Ignoring Aspect Ratio
Before:
Create a professional LinkedIn banner for a tech company.After:
A professional LinkedIn banner image for a B2B cloud infrastructure company. Clean dark background in #0D1117 near-black. On the left third, bold white sans-serif text: "Building the Future of Cloud". A subtle abstract geometric network visualization (nodes and connecting lines) in #3FB4D4 teal fills the right two-thirds, with the lines fading toward the left where the text sits. Modern, technical, authoritative. 8:1 aspect ratio (LinkedIn banner format).Why it's better: Intent (LinkedIn banner), use context, and aspect ratio together tell the model what this image is actually for.
11. Prompt Building Workshop
This section walks you through building a real prompt from scratch, explaining each decision. Think of it as watching over a designer's shoulder.
Workshop Project: A Coffee Brand Campaign Hero Shot
Brief: Create the hero image for a new specialty coffee brand called "Altitude." It should feel premium and adventurous — positioned for a young professional audience who wants to feel like their morning coffee connects them to something larger than a convenience purchase.
Step 1 — Start with just the subject.
A ceramic coffee cup.Too generic. Every image model can generate this. We haven't said anything distinctive.
Step 2 — Specify the subject's materials and state.
A matte black ceramic coffee cup, the surface showing slight kiln texture, filled with a flat white with minimal latte art.Better. Now we have material (matte black ceramic), texture (kiln texture), and contents (flat white, minimal latte art).
Step 3 — Add a setting that supports the brand's story.
A matte black ceramic coffee cup with kiln texture, filled with a flat white with minimal latte art. Placed on a rough granite ledge on a mountain summit, the Andes visible in the background — snow-capped peaks partially in cloud.Now the setting earns the brand name "Altitude." The contrast between the intimate (a cup of coffee) and the vast (mountain range) creates tension.
Step 4 — Add lighting that matches the mood.
A matte black ceramic coffee cup with kiln texture, filled with a flat white with minimal latte art. Placed on a rough granite ledge on a mountain summit, the Andes visible in the background — snow-capped peaks partially in cloud. Early morning light, the sun just cresting the mountains, sending long golden rays across the granite. Steam rises from the cup into cold air."Early morning light" and "steam into cold air" add time-of-day specificity and atmosphere. They also reinforce the adventurous narrative: someone was up before sunrise.
Step 5 — Add a style anchor and technical specs.
A matte black ceramic coffee cup with kiln texture, filled with a flat white with minimal latte art. Placed on a rough granite ledge on a mountain summit, the Andes visible in the background — snow-capped peaks partially in cloud. Early morning light, the sun just cresting the mountains, sending long golden rays across the granite. Steam rises from the cup into cold air. Shot on Hasselblad H6D, 85mm, f/2.8 — subject sharp, mountains softly out of focus. Campaign photography, premium brand, editorial quality. 16:9 aspect ratio.The camera specs (Hasselblad, 85mm, f/2.8) communicate a certain quality level and the depth-of-field intent. "Campaign photography, premium brand, editorial quality" is the style anchor. The aspect ratio (16:9) signals this is for a wide-format hero placement.
The Result: A complete, production-ready prompt in five steps.
This prompt can now go directly to Seedream V5. Notice what we didn't do: we didn't add "8K ultra-realistic hyper-detailed" anywhere. Every word does real work.
Workshop: Adding HEX Color Control
Now let's suppose the brand has a specific color system. The brief says: primary color is #1B3A2D (a deep forest green), accent is #C9A060 (warm gold).
Add the colors to the prompt:
A matte #1B3A2D deep forest green ceramic coffee cup with kiln texture, filled with a flat white with minimal latte art. The base of the cup sits on a small #C9A060 warm gold ceramic saucer. Placed on a rough granite ledge on a mountain summit, the Andes visible in the background. Early morning light, the sun just cresting the mountains. Steam rises into cold air. Shot on Hasselblad H6D, 85mm, f/2.8. Campaign photography, premium brand, editorial quality. 16:9 aspect ratio.Two HEX codes added, and the image is now brand-accurate without any post-processing.
Workshop: Escalating to JSON for Multiple Products
If the campaign needs a three-product lineup shot instead of a single hero:
{
"scene": "three coffee products in a dramatic mountain summit setting, Andes peaks in the background, early morning light",
"subjects": [
{
"description": "matte ceramic coffee cup with kiln texture and flat white latte art",
"position": "center",
"color": "#1B3A2D deep forest green cup, #C9A060 warm gold saucer",
"action": "steam rising into cold air"
},
{
"description": "200g kraft paper coffee bag with a sealed tin-tie, branded label",
"position": "left",
"color": "#1B3A2D forest green label with #C9A060 gold text",
"action": "standing upright, facing slightly toward center"
},
{
"description": "hand-held pour-over coffee dripper with glass carafe below",
"position": "right",
"color": "#1B3A2D dark ceramic dripper, clear glass carafe",
"action": "a thin stream of coffee visible mid-pour"
}
],
"style": "campaign product photography, premium outdoor brand, editorial quality",
"lighting": "early morning golden hour, sun cresting mountain peaks from the right, long shadows to the left, cold blue ambient fill from the sky",
"camera": {"angle": "slightly elevated 25°, looking slightly down at products", "lens": "85mm", "depth_of_field": "products sharp, mountains softly out of focus"}
}12. Cheat Sheet
The Prompt Formula
[Subject + Materials] > [Setting + Time] > [Style Anchor] > [Lighting] > [Camera + Aspect Ratio]Quick-Reference Rules
| Rule | Do | Don't |
|---|---|---|
| Text rendering | Put text in "quotes" | Write text without quotes |
| Colors | Use #HEX color-name | Use vague color descriptions for critical colors |
| Spatial layout | "on the left", "in the foreground" | "somewhere around there" |
| Style | Name a reference: "Bon Appétit style" | Leave style unspecified |
| Edits | Specify what to keep AND change | Just say "make it better" |
| Length | Under 200 words | Multi-paragraph prose |
| Negatives | Use Avoid/negative field after generating | Front-load negatives into main prompt |
| Multi-element scenes | Use JSON format | Cram 6 objects into one paragraph |
The Five Components
| Component | Keywords |
|---|---|
| Subject | material, texture, state, action, scale |
| Setting | location, time of day, season, weather, background detail |
| Style | publication name, photographer name, art movement, film look |
| Lighting | direction (left/right/top), quality (hard/soft), color temperature |
| Technical | camera brand, lens mm, f-stop, aspect ratio, output resolution |
Feature Quick-Reference
| Feature | How to Use |
|---|---|
| Web Search | Toggle ON in interface; reference real-world entities in prompt |
| HEX Colors | Drop inline: #FF006E hot pink — always follow code with color name |
| JSON Prompting | Replace plain text with JSON object {} for multi-element scenes |
| Multi-Language | Write scene description in the language of the culture; add English for tech specs |
| Editing | Feed reference image + edit prompt; specify what to keep |
Common Style Anchors by Domain
| Domain | Style Anchors |
|---|---|
| Product | "Apple.com product photography", "editorial, Wallpaper* magazine", "GQ grooming" |
| Food | "Bon Appétit editorial", "Ottolenghi cookbook", "Food52 lifestyle", "Nobu plating" |
| Architecture | "Hélène Binet", "Dezeen magazine", "Richard Pare", "Julius Shulman" |
| Portrait | "Magnum Photos documentary", "Kinfolk magazine", "Annie Leibovitz editorial" |
| Landscape | "National Geographic", "Peter Lik", "Ansel Adams B&W" |
| Cinematic | "Denis Villeneuve aesthetic", "ARRI Alexa, anamorphic", "Kodak Vision3" |
| Data Viz | "Nadieh Bremer", "Edward Tufte principles", "New York Times graphics" |
| Graphic | "Swiss International Style", "Bauhaus poster", "Wes Anderson palette" |
Common Negative Prompts
General: blurry, low quality, distorted, deformed, watermark, text overlay, cropped, out of frame
Portraits: blurry hands, extra fingers, merged faces, unnatural skin texture
Architecture: bent verticals, distorted perspective, anachronistic elements
Food: unappetizing textures, grey tones, artificial-looking garnishesAspect Ratio Quick Guide
| Use Case | Ratio |
|---|---|
| Social media post, product hero | 1:1 |
| LinkedIn banner, YouTube thumbnail | 16:9 |
| Portrait/story format | 9:16 |
| Standard photo print | 3:2 or 4:3 |
| Cinema / panoramic | 21:9 |
| Ultra-wide banner | 8:1 (custom) |
The 30-Second Prompt Check
Before hitting generate, ask:
- Does my subject have material/texture specifics? ✓
- Do I have at least one style anchor? ✓
- Is any required text in "quotation marks"? ✓
- Did I specify lighting direction? ✓
- Did I specify aspect ratio? ✓
- Is the prompt under 200 words? ✓
- If multiple subjects: did I specify positions? ✓
- If brand colors: are HEX codes in there? ✓
Guide compiled and expanded from: fal.ai Blog, ByteDance Seed, GenAIntel, WaveSpeed AI, Together AI, Reddit r/fal. Last updated: April 2026.
