Viggle released a distilled model based on Qwen-Image-2.1, offering text-to-image with 4 transformer passes instead of 40. It includes two versions: a full transformer and a LoRA adapter.
The LoRA adapter (340 MB) loads at runtime and requires num_inference_steps=4 with true_cfg_scale=1.0. Text-to-image works well, but complex image editing remains less accurate than the base model.
Commercial use is prohibited under the Qwen RESEARCH LICENSE AGREEMENT. The model is a preview version, not a replacement for the base model.
Highlights
- transformer/ — full fine-tuned transformer (bf16, 14.2 GB). Replaces the base transformer; exact, no adapter.
- Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors — LoRA adapter (rank 64, 340 MB) loaded on top of the
- numinferencesteps=4, truecfgscale=1.0, no negative prompt. The adapter was distilled for exactly this;
- Use the shipped scheduler config (or FlowMatchEulerDiscreteScheduler.fromconfig(pipe.scheduler.config,
- LoRA flavour: leave the LoRA scale at 1.0 (alpha equals rank).
- Reference-image order determines which image image 1 / image 2 in the prompt refers to. Without
- Prompt rewriting is optional and was not used in training; the official
- peft users can load peft/ directly: