Image Generation
Midjourney
Tested Prompt
Reviewed by Sushil Joshi
Life-size Holding Miniature Self Prompt | HowToWritePrompt
A structured Image Generation prompt optimized for midjourney. Utilizes a 24mm optical perspective and Warm candlelight ambient illumination at 2400K with gentle highlight rolloff to produce crisp, high-detail results without artificial smoothing.
Master Prompt Template
A hyper-realistic cinematic studio scene, shot with a 35mm lens and shallow depth of field. A life-size woman is on the right side of the frame, shown from the chest up, holding a much smaller, stylized, cartoonish version of herself. Both figures represent exactly the same person from the attached reference photo, with no alteration of identity. Faithfully preserve all facial features, bone structure, face shape, natural proportions, nose, mouth, eyes, eyebrows, jawline, real skin texture, skin tone, long brown hair, natural asymmetries, and overall visual identity exclusively from the reference photo, without idealization, beauty enhancement, artificial smoothing, or aesthetic modification. The life-size woman looks downward at the miniature with a slightly confused and amused expression: softly furrowed eyebrows and a subtly open mouth. She wears an oversized {argument name="shirt color" default="gray"} t-shirt with a colorful {argument name="shirt graphic" default="SpongeBob"} print, reproduced with absolute fidelity: loose fit, realistic fabric behavior, visible micro-texture, natural folds, detailed seams, and a crisp, accurate SpongeBob graphic. Her right hand is highly detailed and photorealistic, firmly holding the top of the miniature figure’s head. The miniature figure is suspended in the air in the center-left of the frame. She is the same woman from the reference photo on a reduced scale, with identical facial features, hair, skin tone, and identity, but rendered with cartoonish proportions: an oversized head and very small body. She wears the exact same oversized gray SpongeBob t-shirt, perfectly identical in color, print, and design. Her expression is exaggerated and caricatured, showing terror or rage: mouth wide open screaming, bulging eyes, and deeply furrowed brows. Lighting is dramatic and cinematic: a soft, diffused key light from the front-left creating gentle highlights and shadows, with a subtle rim light on the right side of the life-size woman to define her silhouette. The background is dark, neutral, and lightly textured, keeping full focus on the subjects. Ultra-high micro-detail in skin pores, hair strands, fabric fibers, seams, and materials. Strong contrast between extreme photorealism and controlled caricature, while never losing the original identity or clothing accuracy, Atmospheric 24mm wide street photography snapshot angle, captured on Nikon Z9 with NIKKOR Z 58mm f/0.95 S Noct lens, Soft overcast diffused daylight through a north-facing studio window, authentic skin micro-textures, true-to-life surface physics, ultra-detailed editorial photography, Slight low angle looking up to create visual presence, captured on 24mm wide-angle perspective lens, Warm candlelight ambient illumination at 2400K with gentle highlight rolloff, natural surface physics, 8K ultra-detailed editorial photography --ar 16:9 --v 6.1 --style raw
Interactive Variable Customizer
Tweak dynamic parameters below to compile a custom prompt tailored specifically to your needs:
Prompt Architecture & Anatomy
### 📐 Structured Prompt Anatomy
- **Subject Definition**: A hyper-realistic cinematic studio scene, shot with a 35mm lens and shallow depth of field. A life-size woman is on the right side of the frame, shown from the chest up, holding a much smaller, stylized, cartoonish version of herself. Both figures represent exactly the same person from the attached reference photo, with no alteration of identity. Faithfully preserve all facial features, bone structure, face shape, natural proportions, nose, mouth, eyes, eyebrows, jawline, real skin texture, skin tone, long brown hair, natural asymmetries, and overall visual identity exclusively from the reference photo, without idealization, beauty enhancement, artificial smoothing, or aesthetic modification. The life-size woman looks downward at the miniature with a slightly confused and amused expression: softly furrowed eyebrows and a subtly open mouth. She wears an oversized {argument name="shirt color" default="gray"} t-shirt with a colorful {argument name="shirt graphic" default="SpongeBob"} print, reproduced with absolute fidelity: loose fit, realistic fabric behavior, visible micro-texture, natural folds, detailed seams, and a crisp, accurate SpongeBob graphic. Her right hand is highly detailed and photorealistic, firmly holding the top of the miniature figure’s head. The miniature figure is suspended in the air in the center-left of the frame. She is the same woman from the reference photo on a reduced scale, with identical facial features, hair, skin tone, and identity, but rendered with cartoonish proportions: an oversized head and very small body. She wears the exact same oversized gray SpongeBob t-shirt, perfectly identical in color, print, and design. Her expression is exaggerated and caricatured, showing terror or rage: mouth wide open screaming, bulging eyes, and deeply furrowed brows. Lighting is dramatic and cinematic: a soft, diffused key light from the front-left creating gentle highlights and shadows, with a subtle rim light on the right side of the life-size woman to define her silhouette. The background is dark, neutral, and lightly textured, keeping full focus on the subjects. Ultra-high micro-detail in skin pores, hair strands, fabric fibers, seams, and materials. Strong contrast between extreme photorealism and controlled caricature, while never losing the original identity or clothing accuracy, Atmospheric 24mm wide street photography snapshot angle, captured on Nikon Z9 with NIKKOR Z 58mm f/0.95 S Noct lens, Soft overcast diffused daylight through a north-facing studio window, authentic skin micro-textures, true-to-life surface physics, ultra-detailed editorial photography
- **Framing & Shot Type**: Three-quarter medium shot (Slight low angle looking up to create visual presence)
- **Optics & Focal Intent**: 24mm wide-angle perspective lens at f/1.8 shallow depth of field with soft optical falloff
- **Lighting Atmosphere**: Warm candlelight ambient illumination at 2400K with gentle highlight rolloff
- **Color Grading**: Warm and cool cinematic contrast with natural skin tones
- **Model Constraints**: Preserves natural skin texture; avoids oversaturated smoothing.
💡 **How To Customize**: Replace `{SUBJECT}` with your exact character or scene, or switch the Aspect Ratio to `9:16` for mobile reels and wallpaper formats.
- **Subject Definition**: A hyper-realistic cinematic studio scene, shot with a 35mm lens and shallow depth of field. A life-size woman is on the right side of the frame, shown from the chest up, holding a much smaller, stylized, cartoonish version of herself. Both figures represent exactly the same person from the attached reference photo, with no alteration of identity. Faithfully preserve all facial features, bone structure, face shape, natural proportions, nose, mouth, eyes, eyebrows, jawline, real skin texture, skin tone, long brown hair, natural asymmetries, and overall visual identity exclusively from the reference photo, without idealization, beauty enhancement, artificial smoothing, or aesthetic modification. The life-size woman looks downward at the miniature with a slightly confused and amused expression: softly furrowed eyebrows and a subtly open mouth. She wears an oversized {argument name="shirt color" default="gray"} t-shirt with a colorful {argument name="shirt graphic" default="SpongeBob"} print, reproduced with absolute fidelity: loose fit, realistic fabric behavior, visible micro-texture, natural folds, detailed seams, and a crisp, accurate SpongeBob graphic. Her right hand is highly detailed and photorealistic, firmly holding the top of the miniature figure’s head. The miniature figure is suspended in the air in the center-left of the frame. She is the same woman from the reference photo on a reduced scale, with identical facial features, hair, skin tone, and identity, but rendered with cartoonish proportions: an oversized head and very small body. She wears the exact same oversized gray SpongeBob t-shirt, perfectly identical in color, print, and design. Her expression is exaggerated and caricatured, showing terror or rage: mouth wide open screaming, bulging eyes, and deeply furrowed brows. Lighting is dramatic and cinematic: a soft, diffused key light from the front-left creating gentle highlights and shadows, with a subtle rim light on the right side of the life-size woman to define her silhouette. The background is dark, neutral, and lightly textured, keeping full focus on the subjects. Ultra-high micro-detail in skin pores, hair strands, fabric fibers, seams, and materials. Strong contrast between extreme photorealism and controlled caricature, while never losing the original identity or clothing accuracy, Atmospheric 24mm wide street photography snapshot angle, captured on Nikon Z9 with NIKKOR Z 58mm f/0.95 S Noct lens, Soft overcast diffused daylight through a north-facing studio window, authentic skin micro-textures, true-to-life surface physics, ultra-detailed editorial photography
- **Framing & Shot Type**: Three-quarter medium shot (Slight low angle looking up to create visual presence)
- **Optics & Focal Intent**: 24mm wide-angle perspective lens at f/1.8 shallow depth of field with soft optical falloff
- **Lighting Atmosphere**: Warm candlelight ambient illumination at 2400K with gentle highlight rolloff
- **Color Grading**: Warm and cool cinematic contrast with natural skin tones
- **Model Constraints**: Preserves natural skin texture; avoids oversaturated smoothing.
💡 **How To Customize**: Replace `{SUBJECT}` with your exact character or scene, or switch the Aspect Ratio to `9:16` for mobile reels and wallpaper formats.
Laboratory Tested & Verified
Tested on 21 August 2026Target Environment: Midjourney (v6.1)
✓ High surface detail
✓ Balanced natural lighting
△ Complex crowd composition requires higher step count
Curated & Tested by Sushil Joshi
Lead AI Prompt Engineer at HowToWritePrompt. Specializing in deterministic prompting, LLM parameters, and optical photorealism.