Ep 4: Text-to-Image from scratch
What we're building
A complete text-to-image workflow. No template. You'll place every node and understand why each one is there.
A text-to-image workflow does five things:
- Loads an AI model
- Takes a text prompt
- Creates a blank canvas for the AI to start from
- Generates the image
- Converts the result into something you can see and save
Place the KSampler
Start here. The KSampler is where generation happens. Everything else feeds into it. You'll see four inputs on the left: model, positive, negative, and latent image.
Load a diffusion model
Add a Load Diffusion Model node. Connect its violet output to the KSampler's violet model input. This gives the KSampler the AI model it needs to generate.
Drag from an empty connector and drop it on the canvas. ComfyUI shows you a filtered list of nodes that are compatible with that connector.
Add your prompts
Add two CLIP Text Encode nodes.
- Connect one to the KSampler's positive input (what you want: "sunflower on a meadow")
- Connect one to the negative input (what to avoid: "blurry, distorted")
Both need a CLIP model. Add a Load CLIP node and connect it to both. Make sure the CLIP model matches your diffusion model.
Create an empty starting canvas
Add an Empty Latent Image node. Set width and height (1024x1024 for most modern models). Connect it to the KSampler's latent image input. This is the blank slate the AI starts refining.
Decode and save
The KSampler outputs latent data (pink). You need to convert it to a visible image.
Add a VAE Decode node. Connect the KSampler's output to it. Load a VAE model and connect that too. Then add a Save Image node and connect the VAE Decode output (blue) to it.
Click run. Your image appears.
Generated images embed the workflow data. Anyone can drag a generated image onto their ComfyUI canvas to load the exact workflow that created it.
FAQ
What is the KSampler?
The core generation node. It takes noise and gradually refines it into an image that matches your prompt, using the model, positive/negative conditioning, and a starting canvas.
How do I know which CLIP and VAE to use?
Check your diffusion model's documentation. Each model specifies compatible CLIP and VAE files. Mismatched models produce poor results or errors.
Can I share this workflow?
Yes. All Floyo workflows are live sharable links that can be sent to other creators directly. To share with the same settings or inputs, add folks to your private Floyo team and share with them. You can also download the workflow JSON from the canvas.
Can I control pose or style in a text-to-image workflow?
Not with text alone. For structural control like pose, depth, or edges, you need ControlNet. That's covered in Ep 10: ControlNet basicsEpisode 10
.
Ep 3: Nodes 101Nodes 101
Ep 5: KSampler settingsKsampler