← Back to the Lab

Qwen-Image-2.1 one stage at a time: a second pipeline for stageload

pytorchapple-siliconmemorybenchmark

The first stageload note measured one pipeline, Pixal3D. A text-to-image pipeline has the same shape: a text encoder, a denoiser and an image decoder run one after another. So on 6 October I measured one, Qwen-Image-2.1 in diffusers, on the same 48 GB Mac. Its text encoder takes 16.3 GB, its transformer 13.3 GB and its VAE 1.3 GB. Each run made one 1024 × 1024 image in 20 steps, in its own process under the memory guard.

Eager is the usual from_pretrained(...).to("mps"). For staged, the pipeline's two large models are stand-ins. They answer .config from the checkpoint, so the pipeline can read it before the denoising begins, and they load the real model on any other use. Three hooks mark where the stages begin, at encode_prompt, prepare_latents and _unpack_latents. The pipeline's code is unchanged.

Peak memory per stage of Qwen-Image-2.1 on a 48 GB Mac, eager and staged

Eager stayed at 31 to 33 GB through encoding and denoising. The VAE decode needs about 11 GB on top of the weights, which took it to 43.6 GB; swap grew by 8 GB, and the guard stopped it before it wrote an image. Staged released the transformer before the decode. It peaked at 19.0 GB while encoding and at 14.4 GB in the decode, and added no swap in either of its two runs.

Staged spent 13 s loading the two models. The denoising took 109 s in its first run, 79 s in its second and 77.5 s for eager; the first run's trace does not show why it was slower.

What I could not show

The two staged runs produced the same image bit for bit. Eager produced none on this Mac, so these runs do not show that staged loading gives the same image as eager. An earlier eager run did not even get through loading: it started a minute after a video job had ended, and swap grew by 10 GB within 18 s.

The code is in bench/qwen_image.py, and the traces and the summary are in the repository.