File size: 20,233 Bytes
e440059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2043d66
e440059
 
ef7e1e9
e440059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2043d66
 
 
 
 
 
 
 
 
 
e440059
 
 
 
 
 
 
 
2043d66
e440059
 
 
 
 
 
2043d66
3cc48cb
 
 
e440059
 
 
2043d66
e440059
 
2043d66
e440059
2043d66
 
 
 
e440059
 
 
2043d66
e440059
 
 
2043d66
 
 
 
 
 
 
 
e440059
 
 
 
 
 
 
 
2043d66
 
 
 
 
 
 
e440059
 
2043d66
3cc48cb
 
 
 
 
2043d66
 
 
3cc48cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e440059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24cc92d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e440059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
---
license: unlicense
tags:
  - flux2
  - flux
  - image-to-image
  - image-upscaling
  - latent-upscaling
  - super-resolution
  - diffusion
  - flow-matching
  - rectified-flow
  - generative-models
  - image-generation
  - pytorch
library_name: diffusers
base_model:
  - TensorForger/FlowUpscaler
  - black-forest-labs/FLUX.2-klein-4B
pipeline_tag: image-to-image
---


# Flow Upscaler



**Flow Upscaler** is a fast latent upscaler model that works in the [Flux.2](https://bfl.ai/models/flux-2) latent space.

Under the hood, it is a lightweight **Rectified Flow** model with **59M** parameters that generates upscaled latents in a single denoising step.


## Inference pipeline (`flow_upscaler_pipeline.py`)

This repo is in **diffusers format**: `FlowUpscalerPipeline` is a
`DiffusionPipeline` whose components are the UNet (`unet/upscaler_unet.py`,
a diffusers `ModelMixin`, weights in
`unet/diffusion_pytorch_model.safetensors`), a shared Flux.2 VAE, and a
`FlowMatchEulerDiscreteScheduler` β€” the plain-Python equivalent of the
ComfyUI node, loadable with `DiffusionPipeline.from_pretrained`. The VAE is
intentionally not bundled (it is shared with any Flux.2 checkpoint, and
FLUX.2-dev is gated), so it is passed in at load time;
`trust_remote_code=True` is required because the pipeline and UNet classes
are loaded from this repo. The pipeline runs the full loop:

1. **Encode** (`encode_image`): the input image is resized to a multiple of
   16, VAE-encoded, and the 32-channel latents are normalized with the VAE's
   BatchNorm running stats β€” the same convention the model was trained with.
2. **Upscale** (`upscale_latents`): pure noise is denoised into the 2x latent
   in a **single FlowMatchEuler step**, conditioned on the small latent.
   Passes chain for 4x/8x because a pass's output lives in the same
   normalized latent space as its conditioning input.
3. **Decode + color fix** (`__call__`): latents are denormalized and
   VAE-decoded. With the default `color_fix=True`, each pass ends with a
   latent low-frequency transplant and the decoded image gets a low-band
   L/a/b transplant from the input β€” the two-stage drift correction detailed
   in the next section.

Memory is bounded for large outputs: the VAE switches to tiling above 2048px,
the UNet runs through a chunked inference executor (`unet/upscaler_unet.py`,
detailed in the next section), and outputs are capped at
`MAX_OUTPUT_SIDE = 16384` px per side (excess passes are dropped with a
warning). This keeps even 16K outputs within a 24GB-class GPU in bf16.

```python
import torch
from diffusers import AutoencoderKLFlux2, DiffusionPipeline
from PIL import Image

vae = AutoencoderKLFlux2.from_pretrained(  # any Flux.2 checkpoint's VAE
    "black-forest-labs/FLUX.2-dev", subfolder="vae", torch_dtype=torch.bfloat16
)
pipeline = DiffusionPipeline.from_pretrained(
    "MinhNH232331M/FlowUpscaler-diffusers",
    vae=vae, torch_dtype=torch.bfloat16, trust_remote_code=True,
).to("cuda")

image = Image.open("input.png")
result = pipeline(image, num_passes=2).images[0]  # 2 passes = 4x
result.save("output.png")
```

Two compatibility paths are kept. `flow_upscaler.safetensors` at the repo
root is the same flat checkpoint the ComfyUI node consumes, and
`FlowUpscalerPipeline.from_single_file("flow_upscaler.safetensors", vae,
dtype=...)` replicates the pre-diffusers constructor for local checkouts.
The seed-based `pipeline.upscale_image(image, num_passes, seed=42, ...)`
method also remains, and a given seed reproduces the pre-conversion outputs
bit-for-bit (`__call__` takes a `generator` instead, diffusers-style).

For the VAE, consider
[MageFlow-VAE-diffusers](https://huggingface.co/MinhNH232331M/MageFlow-VAE-diffusers)
β€” an efficient drop-in replacement for the Flux.2 VAE that operates in the
same latent space (~10x faster encode, ~6x faster decode, roughly half the
peak VRAM, with a chunked decoder that keeps memory near-constant at large
resolutions). Since upscaling cost here is dominated by the encode/decode
ends, it speeds up the whole loop.

Pipeline call knobs: `num_passes` (2x per pass), `generator` (RNG for the
per-pass noise; `upscale_image` takes a `seed` instead), `color_fix` (master
switch for both correction stages), the per-stage `transplant_strength` /
`color_fix_strength` sliders described below, and `output_type` (`"pil"`
default, `"np"`, `"pt"`). `transplant_strength` also acts as a
fidelity/realism control β€” lower values track the input more faithfully,
higher values give the model's crisper rendition.


## UNet inference optimizations (`unet/upscaler_unet.py`)

The network is attention-free, so compute scales linearly with area β€” but the
stock forward's *memory* did not scale gracefully: at large latents each
residual block allocated 4-8 full-resolution temporaries (GroupNorm outputs,
SiLU copies, FiLM scale/shift maps, conv workspaces), and the skip-concat
up-block briefly held several 768-channel full-res tensors.
`unet/upscaler_unet.py` now layers four inference-only optimizations on the
same weights. All of them
are gated on `torch.is_grad_enabled()`, so training behavior is bit-for-bit
untouched.

Measured on an NVIDIA L4 (23 GB), 3840x2160 -> 7680x4320 (one 2x pass at a
960x540x32 latent; one full-res 384-channel activation = 0.37 GiB in bf16):

| Configuration | UNet-pass peak | Output change vs previous row |
|---|---|---|
| fp32 forward (training dtype) | 8.53 GiB | β€” |
| bf16 forward | 4.74 GiB | 54.5 dB PSNR (visually identical) |
| + in-place ops, early frees | 4.28 GiB | bit-identical |
| + split-concat up-block | 3.51 GiB | fp reduction order only (43.6 dB) |
| + chunked executor | **2.50 GiB** | **bit-identical** |

1. **bf16 execution.** The model was trained in fp32, but bf16 matches it to
   ~54 dB PSNR while halving peak memory and roughly doubling forward speed.
   The sinusoidal timestep embedding stays fp32 and is cast to the projection
   weight dtype (`ConditioningEncoder.forward`), which is what makes reduced
   precision work at all. The pipeline default remains fp32; pass
   `dtype=torch.bfloat16`.
2. **In-place ops.** SiLU, FiLM scale/shift, and residual adds mutate their
   operands at inference (backward would need the pre-op values, hence the
   grad gate). Conditioning tensors are freed the moment their FiLM layer has
   consumed them instead of surviving the whole forward.
3. **Split-concat up-block.** The first block after each skip concat is the
   network's memory peak: GN/SiLU/conv on `cat([x, skip])` at double width.
   Because the GroupNorm group count (32 over 768 channels) aligns with the
   concat boundary (channel 384 = group 16), the block decomposes exactly:
   GN and SiLU apply per half, and the 3x3/1x1 convs over the concat are the
   sums of per-half convs. No 768-channel tensor is ever materialized; the
   only numeric change is floating-point summation order.
4. **Chunked inference executor** β€” the main mechanism, engaged when the
   latent is at/above `CHUNK_MIN_PIXELS` (512x512, i.e. a 4096px output).
   Ops execute one at a time over *resident* full tensors, with only
   strip-sized working memory (`CHUNK_STRIP_PIXELS`, ~128k px per strip):
   * 3x3/1x1 convs run in row strips with a 1px halo into preallocated
     outputs (bit-exact vs the full conv); each block's final conv
     accumulates directly into its residual buffer.
   * FiLM applies strip-wise in place, never materializing the full-res
     2x-channel scale/shift map (0.74 GiB at an 8K-UHD latent).
   * GroupNorm stays the *native full-tensor op*. Its statistics are global
     by construction β€” the seam/color-shift failure mode of naive spatial
     tiling cannot occur β€” and since its full-size output just fills a
     live-set slot the neighboring convs already require, keeping it native
     costs no memory. It also keeps the input pristine, so the input doubles
     as the residual for identity-skip blocks.
   The net effect: each block's live set is capped at 3-4 activation-sized
   tensors, and the chunked path's output is `array_equal` with the standard
   inference path β€” there is no quality trade-off to track.

**Execution paths.** Every module picks between three code paths at call
time, so no configuration is needed:

| Path | When | Ops |
|---|---|---|
| Training | grad enabled | original out-of-place ops, upstream-identical |
| Unchunked inference | grad off, latent < `CHUNK_MIN_PIXELS` | whole-tensor ops with optimizations 1-3 (in-place, split-concat) |
| Chunked inference | grad off, latent >= `CHUNK_MIN_PIXELS` (512x512, i.e. 4096px output) | optimization 4 |

The unchunked path exists because below the threshold the entire transient
set is small (one activation at a 2048px output is 48 MB), so strip loops
would add per-strip kernel-launch overhead with nothing to save. Since both
inference paths are bit-identical, the threshold is purely a memory/latency
trade-off β€” never a quality one. Measured at the crossover (2048 -> 4096px,
the first size that chunks): 1.65 GiB chunked vs 1.93 GiB forced-unchunked
at equal speed (~2.0 s end-to-end); a 1024 -> 2048px pass runs unchunked at
1.05 GiB total, dominated by the VAE ends rather than the UNet. At 8K the
unchunked path would peak at 3.51 GiB (the split-concat row of the table
above) versus 2.50 GiB chunked, with the gap widening quadratically at
larger outputs.

**Memory model.** Peak VRAM ~= 4 x one full-res activation + ~1 GiB of
non-scaling base (weights, low-res conditioning stages, noise latent, strip
temporaries), where one activation is `(H/8)(W/8) x 384 x 2` bytes in bf16.
Chunking bounds the per-op overhead, not the total: GroupNorm's global stats
force each op's full input and output to exist between sync points, so the
resident term still scales with latent area. Measured end-to-end with the
MageFlow VAE on the L4:

| Output | UNet pass | VAE decode | Wall (total) |
|---|---|---|---|
| 8K UHD (from 4K) | 2.50 GiB | 1.8 GiB | ~6 s |
| 16K UHD (from 8K) | 7.63 GiB | 5.5 GiB | ~25 s |

`MAX_OUTPUT_SIDE = 16384` is calibrated from these numbers. If the upscaler
shares a GPU with resident diffusion pipelines, setting
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` before torch initializes
CUDA reduces allocator fragmentation (~0.6 GiB reserved-memory savings
measured).

Two further tiers exist if flat, resolution-independent memory is ever
needed, both unimplemented: accumulating GN statistics strip-wise and
recomputing convs in a second pass (~2 resident tensors, ~2x conv compute),
or streaming the resident tensors through pinned CPU memory (peak becomes
tile-sized, ~2 s of PCIe traffic at 16K).


## Color drift and corrections (this repo's pipeline)

Chained single-step passes drift measurably. The local wrapper
(`flow_upscaler_pipeline.py`, used by `app_local.py`) measures and corrects
three drift modes, all enabled by the default `color_fix=True`:

| Drift mode | Magnitude (uncorrected) | Cause |
|---|---|---|
| Per-channel latent mean drift | ~0.1 sigma/pass -> 2-5/255 RGB cast after 2 passes | single-step Euler DC error (VAE round-trip is clean) |
| Chroma amplification | ~8%/pass mean chroma; p99 +14/255 after 3 passes, concentrated on already-colorful regions (lips, skin) | model bias, rides on high-band latent content through the nonlinear decoder |
| Tone-curve stretch | dL_std +1.1, shadows -2.75, highlights +2 after 3 passes | side effect of the latent transplant below (raw model output is slightly *flat*) |

Two-stage correction:

1. **Latent low-frequency transplant** (`_lowfreq_transplant`), after every
   pass: the output latent's low band (antialiased downsample by
   `LOWFREQ_FACTOR = 8`) is replaced with the conditioning's. Fixes the mean
   drift and keeps the next pass's conditioning on-manifold, which measurably
   sharpens results. Cost: ~1.5 ms / +33 MiB at an 8K-UHD output.
   `transplant_strength` interpolates this stage from off (softer, truest to input structure) to
   full (default, crisper); color is unaffected at any setting. Beyond drift
   correction, this knob doubles as a **fidelity/realism control**: lower
   values stay measurably closer to the input (PSNR to input 43.9 dB at 0.0
   vs 38.4 dB at 1.0 in the 2-pass test below), higher values favor the
   model's sharper, more confident rendition.
2. **Pixel-space low-band L/a/b transplant** (`_lab_stat_match`), after the
   final decode: Lab diff maps computed at the input's resolution
   (area-downsampled output vs input, fp32), bilinearly upsampled and added
   to the decoded image in fp16 row strips on GPU. Below the input's Nyquist
   the input is ground truth, so regional color *and* tone are pinned exactly
   while the model keeps all detail above it. Cost incl. uint8 download:
   13 / 54 / 235 ms and +51 / +128 / +404 MiB transient VRAM at 2K / 4K / 8K.
   `color_fix_strength` interpolates this stage linearly; residual drift scales with `1 - strength`.

### Measured results

Protocol: output area-downsampled (`INTER_AREA`) back to input size and
compared against the input in Lab; worst cases from a 512px-base photo set,
2-3 pass chains, seed 42, NVIDIA L4, bf16. Reproduce with
`test_flow_upscaler.ipynb` (one section per design choice, with plots).

| Metric (at input scale) | Uncorrected | Both fixes on |
|---|---|---|
| RGB mean cast, 2 passes | up to 5/255 per channel | <= 0.05/255 |
| Mean chroma (Lab C_mean) | +7...+15% | within +/-0.5% |
| Chroma p99 (worst case, 3 passes) | +13.8/255 | +0.6/255 |
| Tone contrast (L_std), 3 passes | +1.08 | +0.03 |
| Shadows p5 / highlights p95 | -2.75 / +1.96 | 0.00 / 0.00 |

Knob response (2-pass restore-photo test, other knob at its 1.0 default β€”
the two stages are orthogonal by construction and by measurement):

| `transplant_strength` | 1.0 | 0.5 | 0.0 |
|---|---|---|---|
| chroma error dC | -0.02 | -0.02 | -0.02 |
| sharpness (mean abs Laplacian) | 7.94 | 4.86 | 3.28 |
| PSNR to input (dB) | 38.4 | 41.7 | 43.9 |

| `color_fix_strength` | 1.0 | 0.5 | 0.0 |
|---|---|---|---|
| chroma error dC | -0.02 | +0.53 | +1.16 |
| sharpness | 7.94 | 7.98 | 8.07 |

Compounding is per pass: before the pixel-space fix, mean chroma ran +8.6%
after 1 pass and +14.8% after 2; the latent mean drift is ~0.1 sigma per pass on average with
0.35-0.48 sigma outlier channels. The VAE round-trip itself is clean (chroma
-0.02, RGB means <= 0.9/255) β€” every drift mode above is created by the flow
UNet or by the correction stack itself.

### Visual demo

One-image walkthrough of the two knobs (`demo_images/`): a 320x480 portrait
upscaled 2 passes (4x) at seed 42.

**Color fix** β€” both runs at `transplant_strength = 0` so the output
structure matches the original and only color differs. Uncorrected, the skin
drifts visibly pale/green (R βˆ’3.1, G βˆ’0.5, B βˆ’2.8 /255 mean cast; mean
chroma +3.5%); with the fix, means pin to within 0.2/255 and chroma to +0.1%:

![color fix comparison](demo_images/compare_color_fix.png)

**Latent transplant** β€” both runs at `color_fix_strength = 1`. Strength 1
nearly doubles measured sharpness (mean abs Laplacian 1.40 -> 2.71) at the
cost of input fidelity (PSNR to input 43.1 -> 40.8 dB) β€” the
fidelity/realism tradeoff, visible in the lashes and iris detail:

![latent transplant comparison](demo_images/compare_latent_transplant.png)

Inputs and full outputs for all three settings are in `demo_images/`. Color
metrics for this demo were measured through the gradio demo's lossy webp
output, which inflates the residual percentiles slightly; the tables above
use the lossless notebook protocol.

### Nyquist / frequency-band analysis

Band bookkeeping, with frequencies written as fractions of each grid's
Nyquist: one latent pixel decodes to 8 image pixels, so the latent grid's
Nyquist sits at 1/16 cycles/px in pixel space; a conditioning at half the
output resolution owns everything below 0.5x the output-latent Nyquist. Each
correction stage operates in the widest band where its reference is ground
truth *and* its domain assumption holds:

* **Stage 1** corrects below output-latent Nyquist / `LOWFREQ_FACTOR` (= /8).
  The limiting factor is not the conditioning (ground truth up to 0.5) but
  how faithfully latent-space bands map to pixel-space bands β€” see below.
* **Stage 2** corrects below the *input image's* Nyquist, in pixel space,
  where the low band is the pixel-space low band by definition β€” no mapping
  assumption needed. Everything above the input Nyquist is model-invented
  detail with no reference, and is deliberately left untouched.

The latent->pixel mapping was measured directly on the Mage decoder: inject a
band-limited 0.1-sigma perturbation into one latent spectral band, decode,
and integrate the luma response spectrum per pixel band (bands as fractions
of the latent Nyquist; pixel content above 1.0 is finer than the latent grid
can represent spatially):

| Latent band | -> same pixel band | -> finer than latent Nyquist |
|---|---|---|
| 0 - 1/16 | 87.7% | 7.8% |
| 1/16 - 1/8 | 90.6% | 4.2% |
| 1/8 - 1/4 | 77.1% | 15.0% |
| 1/4 - 1/2 | 64.7% | 26.2% |
| 1/2 - 1 | 38.0% | 58.1% |

Three design rules fall out of this table:

1. Latent band surgery is sound only in the top two rows (~88-91%
   correspondence) β€” exactly where `LOWFREQ_FACTOR` 8-16 operates. Below 77%
   the spliced content scatters across unrelated pixel frequencies and reads
   as ghosting (matches the failed wide-band experiments below).
2. The ~8% upward leakage from the lowest band is the latent transplant's
   sharpening / tone-stretch side effect, quantified β€” the energy it injects
   into finer-than-latent-Nyquist pixel detail.
3. Energy fractions understate *statistical* bias: high-band latent chroma
   texture rectifies through the nonlinear decoder into a mean-saturation
   shift carrying almost no low-band energy (the mechanism behind drift
   mode 2). Exact guarantees therefore only come from the pixel-space stage.





**[ComfyUI Node](https://github.com/TensorForger/comfyui-flow-upscaler)**

Features:

* Upscaling from **512x512** to **1024x1024** takes **8ms*** 
* The model is trained for **2X** upscaling, but multiple passes can be chained to reach up to **8K** resolution
* A full pipeline with Flux generation, upscaling to **8K**, and decoding runs in just **25 seconds** (on RTX 5090)
* The training process uses **Flow Distillation** with Flux.2 as a teacher, forcing the model to learn strong image semantics

*On RTX 5090, in latent space, without decoding, see benchmark [here](https://github.com/tensorforger/CTGMWorkshop).

Here is one **4X** upscaled image (two passes):

![comparison](https://raw.githubusercontent.com/tensorforger/tensorforger/main/assets/upscaler_comparison.png)

## How it works

Architecturally, Flow Upscaler is a U-Net with SDXL-style ResNet blocks. It takes a noisy sample as input and predicts velocity as output. The generation process happens directly in high-resolution latent space.

The low-resolution latents are passed through a separate conditioning encoder that produces control signals, which are injected into the main U-Net encoder using FiLM conditioning.

No attention layers are used, so compute scales linearly with image area. This makes generation at **8K** resolution possible.

![example](https://raw.githubusercontent.com/tensorforger/tensorforger/main/assets/flow_upscaler_architecture.PNG)

The model is trained using **Flow Distillation** with Flux.2-klein-4B as a teacher. We generated **20K** diverse images with Flux, storing the initial noise, generated latents, and downscaled latents used for conditioning.

The downscaled latents are created by decoding high-resolution latents, downscaling them in pixel space, and encoding them back into latents. Direct latent downscaling introduces artifacts and breaks latent patterns, resulting in blurry decoded images.

![example](https://raw.githubusercontent.com/tensorforger/tensorforger/main/assets/flow_upscaler_training_approach.PNG)


## Training code

If you want to explore the training code or use the model outside ComfyUI, see:

`notebooks/flow_upscaler` in [https://github.com/tensorforger/CTGMWorkshop](https://github.com/tensorforger/CTGMWorkshop)