There is no single best upscaler for ComfyUI. There are two families, and the right pick is decided by one question: do you need to enlarge what is already there, or do you need the model to invent detail that was never captured?
Everything below sorts into those two answers. If you only want the short version: for speed and no extra VRAM, use an ESRGAN-family model with tiling; when a face, a texture or a compressed video frame has lost real information, a diffusion model like SeedVR2 is the only one of the two that can plausibly reconstruct it.
The four options people actually install
Most "best upscaler for ComfyUI" lists mix two different kinds of tool. These are the four that come up in practice:
| Option | Family | What it really does | Main cost |
|---|---|---|---|
| ESRGAN-family model (Real-ESRGAN and friends) | GAN | Enlarges and sharpens in one pass | Cannot invent missing detail; artefacts on already-sharp input |
| 4x-UltraSharp and similar community checkpoints | GAN | Same node path, tuned for a look | Quality varies by content; needs tiling on large images |
| Ultimate SD Upscale | GAN + tiling | Runs a GAN model in overlapping tiles so big images fit | Tiles can seam; slow on very large canvases |
| SeedVR2 and other diffusion restorers | Diffusion | Reconstructs plausible detail, not just edges | Far heavier per pixel; wants a GPU or a hosted run |
Speed and VRAM separate them cleanly. A GAN upscale is a single forward pass through a small network; a diffusion restore is a multi-step denoise where every step costs roughly a pass of its own.
How ComfyUI's stock upscale actually works
The default path is a loader plus one node. You drop a .pth upscale model into ComfyUI/models/upscale_models, add Load Upscale Model, wire it into Image Upscale with Model, and set a target. That is the whole GAN workflow, and it is why people start here: no extra dependencies, no sampling, milliseconds per pass.
The catch is what the model is allowed to do. A GAN upscaler learns a mapping from low-resolution patches to high-resolution patches. It is very good at edges and texture statistics, and it has no mechanism for "the pixels here are gone, guess what was there". On a soft but honest photo that is fine. On a face blurred by compression, or a video frame where the encoder smeared detail across seconds of frames, there is nothing to map, and the result is a clean, sharp, slightly wrong image.
Tiling: the VRAM answer that also becomes the quality problem
Ultimate SD Upscale exists because a single pass at 4x on a large canvas does not fit in memory. It splits the image into tiles, runs the upscaler on each, then stitches with overlap and a seam fix.
Two things follow. First, you can upscale images that never fit before, which is why the node is so popular. Second, every tile is upscaled with no knowledge of its neighbours, so the model can decide on a different texture in tile one than it did in tile two. Overlap and seam-fix settings reduce that, not remove it.
If you are choosing settings rather than models, the knobs that matter most are tile size against available VRAM, tile overlap, and seam-fix strength. Our SeedVR2 ComfyUI settings guide walks through them next to the loader options, because those two groups interact.
Diffusion restorers: what actually changes
A diffusion upscaler does not map patches. It starts from noise and denoises toward an image that is consistent with the low-resolution input, using a model that has seen an enormous range of real images. That is why it can put plausible eyelashes on a blurred face: it is reconstructing a likely image, constrained by the input, rather than enlarging a stored mapping.
The trade is honest and symmetric:
- Upside: it is the only family that can recover believable detail on compressed, blurry or old sources.
- Downside: it costs far more compute per pixel, it can hallucinate detail that was never in the source, and it needs enough VRAM that many people never finish a 4K pass on a laptop.
SeedVR2 is ByteDance's entry in this family, published under Apache-2.0 as a 3B and a 7B checkpoint. Both sizes, the precision options and the exact .safetensors names are covered on our SeedVR2 model files page; the node-by-node graph is on the ComfyUI workflow page.
Picking one: four cases, four answers
Written as decisions, because that is how the question actually arrives:
- You need more pixels per second than per detail. Use an ESRGAN-family model. Start at 2x, check the result, then decide whether 4x adds anything real.
- Your image is large and your card is small. Use tiling, and expect to tune overlap. If seams show up in flat areas like skies, lower the tile size and raise overlap.
- The source lost information. Compressed video, an old scan, a small face in a group photo. Go to a diffusion restorer. A GAN will give you a bigger version of the same blur.
- You want a finished file today, not a graph. Run it hosted instead — no checkpoint to download, no tile settings, and the browser video upscaler or the image upscaler both quote the credit cost before the job starts.
The most common mistake is buying a bigger GAN checkpoint when the problem is missing information. The second most common is running a diffusion restore at 4x on material that never needed it, which burns hours to reach a result a GAN would have matched in seconds.
When "best" is actually "fastest honest answer"
If you are comparing models because you are stuck mid-project, the decision can be cheaper than the research. A single upscale on your own fixture settles it: run the same crop through a GAN model and through a diffusion model, then look at the area that failed. If the GAN output is a sharper version of the blur, and the diffusion output has believable texture, your source class is decided, and so is your model class.
For a wider view of the video side specifically — including how the hosted route compares with desktop licences — see best AI video upscaler.
FAQ
What is the best upscaler for ComfyUI overall? There is no single answer because the two families solve different problems. Use an ESRGAN-family model when you need speed and the source is already sharp. Use a diffusion model like SeedVR2 when the source has lost detail you need back.
Do I need a diffusion model for video? Usually yes, more than for photos. Video compression destroys detail across frames, and a GAN cannot restore what the encoder removed. This is why video upscaling is where diffusion restorers earn their cost.
Which upscale model folder does ComfyUI use?
ComfyUI/models/upscale_models for GAN .pth upscalers. Diffusion checkpoints are different files and belong elsewhere — the folder layout for SeedVR2 specifically is on the model files page.
How much VRAM does an upscaler need in ComfyUI? A GAN upscaler with tiling will run on modest cards. A diffusion restore is a different order of magnitude and is the reason tiling, precision and offloading settings exist at all.
Can I upscale without installing ComfyUI? Yes. The hosted path on this site runs the same model family in the browser: upload a supported source, choose a target resolution, and see the credit cost before you commit.
