Supra2-IMG is a text-to-image model with 100M parameters trained from scratch on high-quality synthetic data.
It uses a diffusion transformer with 104.1M parameters, a frozen Flan-T5-Base encoder, and SD-VAE-FT-MSE.
The model was trained for 10 epochs on the LucasFang/FLUX-Reason-6M dataset with 5.6M images.