Ep 8: Image-to-Image
Text-to-image vs. image-to-image
In text-to-image, the AI starts from noise and generates from scratch. In image-to-image, the AI starts from a photo or sketch you provide and transforms it based on your prompt.
We'll take the text-to-image workflow from Ep 4: Text-to-Image from scratchEpisode 4 and modify it. The change is small: swap one node group.
The swap
- Delete the Empty Latent Image node
- Add a Load Image node. Drop in the image you want to transform
- Add a VAE Encode node between Load Image and the KSampler
You can't connect Load Image (blue) directly to the KSampler (pink). VAE Encode converts the visible image into latent space so the AI can work on it. Connect your existing Load VAE to both the VAE Encode and VAE Decode nodes.
Resize your input
Diffusion models work at specific resolutions (usually around 1024x1024). Feeding a 4K image directly will fail or produce artifacts.
Add a Resize Image node between Load Image and VAE Encode. Set it to 1024 pixels. Make sure it keeps the aspect ratio so nothing gets stretched.
Add a Resize Image node between Load Image and VAE Encode. Set it to 1024 pixels. Make sure it keeps the aspect ratio so nothing gets stretched.
Denoise is everything here
In image-to-image, the KSampler's denoise setting controls how much changes:
Value | Result |
|---|---|
0.2-0.4 | Subtle. Structure and colors stay mostly the same |
0.5-0.7 | Noticeable changes. Style shifts but composition holds |
0.8-1.0 | Dramatic. The original is barely a suggestion |
Default is 1.0, which ignores your input entirely. Lower it.
The AI doesn't "see" objects in your image. It sees pixels and colors. A low denoise preserves those color patterns. The prompt controls what the AI thinks it's looking at. For actual structural control (poses, edges, composition), you need Ep 10: ControlNet basicsControlNet.
FAQ
Why does my output look nothing like my input?
Denoise is too high. At 1.0 the AI ignores your image entirely. Try 0.3-0.5.
What resolution should my input be?
Match your model's training resolution. For most modern models, that's around 1024x1024. Add a Resize Image node to handle this automatically.