Ep 7: Models explained
Why models matter
Without models, your workflow is just a set of empty nodes. The models are what actually know how to generate images, translate text, encode data, and control composition. Here's what each type does.
Diffusion models
The main engine. A diffusion model takes random noise and turns it into a coherent image matching your prompt. Every image generation workflow needs one.
These are large files. Older models were around 6GB. Modern ones are commonly 10-20GB, some over 40GB.
Two file formats to know:
- .safetensors is the standard. Use this.
- .gguf is a compressed format that uses less memory at some quality cost.
Avoid .ckpt files. This older format can contain hidden executable code. Stick with .safetensors.
LoRA (Low-Rank Adaptation)
A LoRA is a small model that modifies a diffusion model's behavior. It doesn't generate on its own. It steers.
Example: a diffusion model doesn't know what a specific person looks like unless they were in the training data. You can create a LoRA trained on photos of that person. Pair it with the diffusion model, and it generates accurate images of them in different poses, lighting, and styles.
LoRAs work for anything: art styles, product designs, characters, clothing, architecture. They're small, so you can collect and swap between many of them.
CLIP
CLIP translates your text prompt into numbers the diffusion model can work with. It sits between your words and the AI.
A good CLIP model understands nuance: "a dog sitting calmly" is different from "a dog jumping excitedly." Using the wrong CLIP model is like using a translator who speaks a different language than your model.
Each diffusion model requires a specific CLIP model. Check the documentation.
VAE (Variational Auto Encoder)
The VAE converts data between two spaces:
- Encode: visible image → latent space (so the AI can work on it)
- Decode: latent space → visible image (so you can see the result)
If your colors look off or your output looks washed out, it might be the wrong VAE.
ControlNet
ControlNet gives you structural control over composition. Text prompts describe what to generate. ControlNet controls how it's arranged: the pose, the depth, the edges.
More on this in Episode 10.
The model landscape
The open-source image generation space moves fast. Stable Diffusion 1.5 kicked things off. SDXL was a major step up. Now there's WAN, QWAN, Flux, ZImagine, and more. New models come out constantly.
The point of this course isn't to memorize which model is best right now. It's to understand how models work in ComfyUI so that when the next one arrives, you already know how to use it.
FAQ
What is a LoRA and when would I use one?
A LoRA is a small add-on model that teaches a diffusion model something specific: a face, an art style, a product design. Load it alongside your diffusion model using a Load LoRA node.
Why do my images look bad even with a good model?
Most likely a mismatch. Your CLIP, VAE, and diffusion model all need to be compatible with each other. Check your diffusion model's documentation for the recommended CLIP and VAE.
What file format should I use for models?
.safetensors. It's the current standard and it's safe. .gguf works too for lower memory usage. Avoid .ckpt.