Previous Post

SECourses Musubi Tuner - 1-Click to Install App for LoRA Training and Full Fine Tuning Krea 2, Ideogram 4, Qwen Image, Qwen Image Edit, Wan 2.1 and Wan 2.2, FLUX Klein, FLUX 2, Z Image Base and Turbo Models with Musubi Tuner with Ready Presets

Next Post
SECourses Musubi Tuner - 1-Click to Install App for LoRA Training and Full Fine Tuning Krea 2, Ideogram 4, Qwen Image, Qwen Image Edit, Wan 2.1 and Wan 2.2, FLUX Klein, FLUX 2, Z Image Base and Turbo Models with Musubi Tuner with Ready Presets
1 / 36
DESCRIPTION

Patreon exclusive posts index to find our scripts easily, Patreon scripts updates history to see which updates arrived to which scripts and amazing Patreon special generative scripts list that you can use in any of your task.

Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord

Please also Star, Watch and Fork our Stable Diffusion & Generative AI  GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)

=======

Latest Zip File : SECourses_Musubi_Trainer_v30_2.zip

[click here to choose a membership and Join to download zip files]

Main training tutorial to learn training and how to use this app (mandatory to watch) : https://youtu.be/DPX3eBTuO_Y

Wan 2.2 training tutorial : https://youtu.be/ocEkhAsPOs4

SwarmUI : https://www.patreon.com/posts/114517862

ComfyUI : https://www.patreon.com/posts/105023709

=======

Currently 1-click to install on Windows, RunPod and Massed Compute with uv installation - ultra fast

We have model auto downloader that supports below models (run Windows_Download_Training_Model_Files.bat)

The following models have full Fine Tuning / DreamBooth and LoRA training with fully supported Torch Compile (really speeds up training up to 58% with no tradeoff)

All Qwen Image Models, FLUX 2 Dev, FLUX 2 Klein 4B and 9B, Z-Image Base, Z-Image Turbo, Ideogram 4, Krea 2 Raw, Krea 2 Turbo

Also has automatic Qwen VL captioning

The following models only supported with LoRA training with fully supported Torch Compile (really speeds up training up to 58% with no tradeoff)

All Wan 2.1/2.2 variants

Please use model downloader to not have any issues because your selected models may be wrong

Moreover, the Musubi Tuner automatically does FP8 and FP8 scaled conversion while loading BF16 model into RAM so we always use BF16 models for training

Our model downloader is ultra optimized and can reach 1GB per second on cloud machines

It also has SHA256 verification to ensure no corrupt model ever used

We are using our own fork of Musubi Trainer does it has many more features and improvements

The installers will generate a Python 3.12 VENV automatically and install everything inside there, thus your system or any other of your APPs will never be impacted

With our pre-compiled abi3 wheels for Windows and Linux, you can run on Python 3.10, 3.11, 3.12 and 3.13 but preferred version is 3.12

The following libraries compiled for Windows and Linux

archs : mslk, xformers, flash_attn, sageattention, torchao,

All of them are abi3 and Windows versions are compiled for Consumer GPUs, Linux versions compiled for Consumer + Cloud GPUs - so we cover all GPUs you have

Windows Requirements

For this auto installer to work you need to have installed Python 3.12.10 (may work with 3.10.x, 3.11.x, 3.13.x too), Git, FFmpeg, cuDNN 9.17+, CUDA 13.0, Visual Studio Community Edition with All c++ options

Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0

Follow its updated post with links and screenshots exactly : https://www.patreon.com/posts/windows-AI-requirements-tutorial-111553210

18 July 2026 V30.2 Update

We have added LTX 2.3 Int8 Row ConvRot HQ quantization preset to our Model Quantizer tab

The difference of Int8 Row ConvRot HQ presets is not very well known in the community

I have so far compiled Krea 2 and LTX 2.3 recently, each takes around 3-4 hours on RTX 5090

You can get them with our SwarmUI model downloader app, bundles are now also downloading them : https://www.patreon.com/posts/114517862

The difference is that, Int8 Row ConvRot HQ is better than GGUF Q8 quality and more than 100% faster on RTX 3000, 4000 and 5000 series (don't have 2000 or 1000 series so can't tell)

Krea 2 quality and speed test as below

Int8 Row ConvRot is 96.2% similar to BF16 meanwhile GGUF Q8 is only 90.0% and FP8 Scaled is 82.2% and NVFP4 is 63.7%

Moreover, Int8 Row ConvRot generates the output in 3.05 seconds, making it 1.82× faster than BF16, which takes 5.56 seconds.

NVFP4 takes 3.8 seconds and is 1.46× faster than BF16, whereas GGUF Q8 takes 6.06 seconds and is approximately 8.3% slower than BF16.

LTX 2.3 Int8 Row ConvRot Benchmarks are as Below

So Int8 Row ConvRot HQ is 100% faster than FP8 Quant Scaled and 50% faster than BF16 on RTX 5090

The quality is also excellent almost same as BF16

To update from V30, just run Windows_Install_and_Update.bat file

Make sure to have latest zip file always and overwrite older installer files

14 July 2026 V30.0 Update

This is a major update with lots of new features and massive performance improvements

Now we are finally able to catch Linux training speeds, I tested Krea 2 and verified at least for Krea 2, read entire update to see

Krea 2 - 36%, FLUX 2 Klein 9B - 46%, Z-Image Base / Turbo - 53%, Ideogram 4 - 14%, Qwen 2512 - 24.5%, Qwen 2511 - 40%, Wan 2.1 - 22%, Wan 2.2 - 15%, FLUX 2 - 15.6% faster

The following models are now fully supported for Full Fine Tuning / DreamBooth training

FLUX.2 dev, FLUX Klein 4B and 9B, Ideogram 4, Krea 2 Raw and Turbo

Demo presets updated as below

New shared full-model training engine

Full FP32 or BF16 DiT training and checkpoint export.

Exact training resume, including mid-epoch resumes and epoch boundaries.

Training-time sampling, checkpoint retention, metadata, Hugging Face upload, and memory-efficient saving.

Single-GPU block swapping and fused-backward memory optimizations.

Ordinary multi-GPU DDP support with automatic validation of incompatible options.

FLUX.1 Kontext full-DiT training is also available through the Musubi backend.

Three new experimental Automagic optimizers implemented from Ostris AI Toolkit (https://www.patreon.com/SECourses/posts/ostris-ai-1-for-140089077)

Automagic, Automagic2, and Automagic3 are available across the training tabs.

Adaptive learning rates with different per-element, per-tensor, and per-group strategies.

Support for LoRA and full Fine-Tuning / DreamBooth, block swapping, checkpoint resume, and low-precision training.

Automatic fused/non-fused selection where supported.

Unsafe combinations are rejected before training instead of silently producing a broken run.

The GUI now displays detailed optimizer-specific guidance.

Native PyTorch fused SDPA is preferred.

When unavailable, external FlashAttention is used only after passing real CUDA forward/backward compatibility tests.

Unsupported tensor types or layouts automatically fall back to working PyTorch SDPA.

Every GUI training tab now has a Use Legacy PyTorch SDPA compatibility switch.

Smarter torch.compile with block swapping

Make sure to enable for Wan 2.2 while Wan 2.1 retains its standard default.

Explicit user settings are preserved when switching tasks or loading configurations.

Improved handling of empty or missing Accelerate launch values.

Hardened Qwen and Z-Image full-model configuration previews and runtime exports.

Added correct full-fine-tuning metadata to Qwen and Z-Image checkpoints.

Maximum absolute gradient

All measured before clipping and sent to TensorBoard or Weights & Biases.

Added CUDA 13.2/PyTorch packaging support.

Improved Ninja and compiler discovery for virtual environments and Visual Studio.

The Massive Speed Improvements Are As Below

All trainings made at 1024x1024px and batch size 1, for LoRAs, LoRA rank is 128

Click images to see full sizes

All speed ups have absolutely 0 tradeoff except Torch Compile initial time requirement - so exactly same max quality

Krea 2 - 36% faster compared to default Musubi Trainer (default no Torch Compile)

We implemented special way of Torch Compile discovery so fully working on Windows too

FLUX 2 Klein 9B - 46% faster

Z-Image Base / Turbo - 53% faster

Ideogram 4 - 14% faster

Qwen 2512 - 24.5% faster

Qwen 2511 - 40% faster - Trained with edit images

Wan 2.1 - 22% faster

Wan 2.2 - 15% faster

FLUX 2 Dev - 15.6% faster

For updating from V29, extract zip file, overwrite files and just run install / update bat file

For updating older version, make a fresh new install

12 July 2026 V29.0 Update

This is a pretty massive update please read all

Please make a seperate new fresh install

We have fully moved to Torch 2.13.0 and CUDA 13.1 with pre-compiled libraries

Also now our preferred Python is 3.12.10, still should work with 3.10, 3.11, 3.13 but please have 3.12 for best

RunPod, SimplePod, Massed Compute installers auto installs 3.12 so you don't need to do anything

Linux users can use Massed Compute installers

For Windows, please use Windows_Install_and_Update.bat and Windows_Download_Training_Model_Files.bat

For Massed Compute and Local Linux : Massed_Compute_Instructions_READ.txt

For RunPod and Simple Pod : Runpod_SimplePod_Musubi_Trainer_Instructions.txt

I have started upgrading all of our apps into latest Torch 2.13 and CUDA 13

For this, I have pre-compiled the following wheels with all CUDA 13 features and with all CUDA archs : mslk, xformers, flash_attn, sageattention, torchao

All these libraries are properly compiled with abi3 thus works on Python 3.10, 3.11, 3.12 and 3.13

Full Krea 2 training added - currently only LoRA

Fully Ideogram 4 training added - currently only LoRA

Demo presets are available inside Demo_Training_Configs_FLUX-2_Z-Image_FLUX-Klein_WAN-21_Krea2_Ideogram4 folder

Hopefully fully researched parameters will come for Krea 2 and others soon

I really like Krea 2 so far, really great model

Model downloader updated for newer models and improved

Model downloader is now more robust and faster

Now when you load a config, it will check if config has valid model paths, if not, it will scan default model downloader downloaded model paths and auto set them

Works on all platforms Windows, Linux, Cloud

For example on RunPod and Massed Compute automatically filled model paths like below

Now you can download multiple models with coma seperation or ranges like 1,2,3 or 1-3,5

For JSON prompt generation we have new tutorial and tool that is useful for Ideogram 4 model training

Tutorial : https://youtu.be/TW3MRdd0MV4

Tool : Ultimate Image Captioner Pro https://www.patreon.com/SECourses/posts/ultimate-image-captioner-pro-162527725

Model Quantizer tab completely improved and now works much better with newer presets

Now our ComfyUI and SwarmUI fully supports Int8 Row ConvRot and this quantization is insane soon hopefully I will make a tutorial and show and also update SwarmUI and ComfyUI posts to show

I recommend to use INT8 ConvRot Learned (Best Quality / Slow) preset to generate quantization of base models it is slow but insane quality even better than GGUF Q8 and 2x faster on all GPUs

Quantization is only for base models not for LoRAs

Now warnings and errors will be displayed on Gradio too with notice bubles

Now Torch Compile works even better and C++ tools not needed, only Visual Studio Community Edition with C++ options please see updated requirements

If you get any errors follow below video and its source link

Make sure that your NVIDIA driver is updated (min 590+)

https://youtu.be/DrhUHnYfwC0

https://www.patreon.com/posts/windows-requirements-tutorial-written-post

Now displays accurate training step speed almost immediately after training started accurately

I fixed this issue on my maintained fork

Hopefully I will add LTX 2.3 training as well soon with audio support

Krea 2 Demo preset training VRAM usage and Speed as below for 1024x1024px

Krea 2 Demo preset with 0 Block Swap + Torch Compile training speed and VRAM usage as below at 1024x1024px

Ideogram 4 Demo preset training VRAM usage and Speed as below for 1024x1024px

Ideogram 4 Demo preset with 0 Block Swap + Torch Compile training speed and VRAM usage as below at 1024x1024px

Tests are made on Windows 11 with RTX 5090

Torch Compile do not affect model quality

4 June 2026 V28.3 Update

Since requested new feature implemented into Image Captioning

When overwrite is off, append new captions to existing text files instead of skipping them

Append Existing Captions

Just run Windows_Install_and_Update.bat to update

20 April 2026 V28.2

Torchaudio version bug fixed

Quantizer application updated to latest and now it uses prodigy to generate FP8 Scaled Quants or NVFP4

FLUX 1 DEV preset added

FLUX 2 Klein Models preset added

ERNIE preset added

All presets updated to use new prodigy and thus now faster and better quality

Hopefully I will add new Quant models and presets to our model downloader zip file soon : https://www.patreon.com/posts/swarmui-auto-and-114517862

Get latest zip file, extract and overwrite all and run Windows_Install_and_Update.bat to update

8 March 2026 V27.6

Model quantizer app significantly updated

Presets are updated and made much better

Zip file is still same just use Windows_Install_and_Update.bat to update

This FP8 Scaled generation is much better than default FP8 and here why

Currently generating LTX2.3 distilled FP8 Scaled max quality for you :)

10 February 2026 V27.5

Sample generation section and backend for Qwen, Wan, FLUX and Z-Image improved and fixed

7 February 2026 V27.2

This is a pretty big update please read all carefully

We have added new Model Quantizer page which supports FP8 Scaled, NVFP4 and more quantization - still in BETA

FLUX Training tab added

FLUX 2 Dev LoRA Training with Torch Compile - min 18 GB VRAM with 128 LoRA Rank and 1024px training (you can reduce LoRA rank and resolution to reduce VRAM or speed up training)

FLUX Klein 9B LoRA Training with Torch Compile - min 9.6 GB VRAM with 128 LoRA Rank and 1024px

FLUX Klein 4B LoRA Training with Torch Compile - min 5.6 GB VRAM with 128 LoRA rank and 1024px

Sadly no Fine Tuning of FLUX 2 or FLUX Klein yet, I notified Kohya and hoping soon

Z Image Training tab added

Z-Image Base Fine Tuning with Torch Compile - min 6 GB VRAM with 1024px

Z-Image Base LoRA Training with Torch Compile - min 6.1 GB VRAM with 128 LoRA rank and 1024px

Z-Image Turbo LoRA Training with Torch Compile - min 8.6 GB VRAM with 128 LoRA rank and 1536px

To update, extract latest zip file, overwrite older files and run Windows_Install_and_Update.bat

To download accurate training files use Windows_Download_Training_Model_Files.bat

New demo configs exists in Demo_Training_Configs_FLUX-2_Z-Image_FLUX-Klein_WAN-21 folder

Z-Image LoRA training config is set similar to AI Toolkit Z Image Turbo training tutorial config, the other ones are set approximately and full R&D will be made later hopefully and final official configs will be published later like Qwen_Training_Configs and Wan22_Training_Configs hopefully with a tutorial

Currently configs are set for lowest VRAM thus if you have higher VRAM, you can reduce block swap count and such

16 January 2026 V26.2 Zip File

RunPod template link updated and now we fully support SimplePod which is much faster and cheaper than RunPod

SIMPLEPOD CHEAPER AND FASTER THAN RUNPOD

Now we fully support SimplePod as well please use this link to register : https://simplepod.ai/ref?user=secourses

SimplePod is faster and cheaper than RunPod and works exactly same

E.g. RTX 5090 on RunPod is 0.89 USD per hour, on SimplePod it is 0.45$ per hour,

RTX PRO 6000 on RunPod is 1.84 USD per hour and on SimplePod it is 0.79 USD per hour

Please use this template on SimplePod : https://dash.simplepod.ai/account/explore/100/ref-secourses/

For permanent storage, generate it from Storage tab with any name and size you want and when selecting template with above link, click Edit and Use, select Persistence Volume and change mount point to /workspace

Up-to-date SimplePod tutorial starting from 21:51 : https://youtu.be/yOj9PYq3XYM?si=Z86wZZLBeYzWo1Qo&t=1311

As usual follow Massed_Compute_Instructions_READ.txt and Runpod_SimplePod_Musubi_Trainer_Instructions.txt to install and use and watch the tutorials

4 January 2026 V26.2 Update

For FP8 Model Converter tab for Improved (convert_to_quant learned rounding) method new feature Use pinned RAM for faster GPU transfers (page-locked memory) added

1 January 2026 V26 Update

Qwen Image 2512 BF16 added into downloader app you can train it exactly as Qwen Image 0 difference

New highly experimental and expert stuff Improved (convert_to_quant learned rounding) added to FP8 Converter tool

It is based on https://github.com/silveroxides/convert_to_quant repo

This is total expert stuff so i am not supporting normally

Just get and extract latest zip file to get new download and Windows_Install_and_Update.bat to have new FP8 Converter tool

31 December 2025 V25.2 Update

Now you can set Target frames on interface - this matters

Now there is Auto Normalize Target Frames which i recommend you to enable - auto enabled

Both options added to Wan Training Dataset preperation tab

29 December 2025 V25.1 Update

Installers upgraded to uv pip installation

Installation on RunPod is now like 100x faster literally and many times faster on Windows and RunPod

Long awaited Wan 2.2 training tutorial published : https://youtu.be/ocEkhAsPOs4

Qwen Image Edit 25-11 Model added to training models downloader ( Windows_Download_Training_Model_Files.bat )

Qwen Image Edit 25-11 Training support implemented just run Windows_Install_and_Update.bat to update

You can train it without control images just as regular text to image model

You will select your model version from dropdown now

17 December 2025 V23 Update

This is a massive update - Long waited Wan 2.2 official training configs published

I have fully researched Wan 2.2 with literally over 64 uniqute trainings and analyzed all of the results - I have used a 8x B200 cloud machine for this research

Hopefully will make a tutorial for Wan 2.2

After watching Qwen Tutorial, exactly same way you can train your Wan 2.2 Text to Image or Text to Video or Image to Video model

Just using static images datasets working perfect I have tested

We have configs for literally every GPU check the configs below

Hopefully will explain more in tutorial

Qwen Image Fine Tuning Tier1_84000_MB.toml config was broken due to Fused Backward pass Off - currently it is enabled - This is bug in Kohya Musubi

When training Wan 2.2 models, there were few serious bugs and all fixed and working perfect now

Extract latest zip file and get latest configs and run Windows_Install_and_Update.bat to update

29 November 2025 V21 Update

Big news, now we have famous Torch Compile support for Qwen and WAN and all QWEN Training configs we have are modified to have auto Torch Compile enabled

This brings around 7.5-15% speed up and lowers VRAM usage to some degree with absolutely 0% quality loss or anything so this is a total win for Speed and VRAM usage

Moreover, FP8 Model Converter now supports Z Image Turbo Model as well and that is how I published the FP8 Scaled model of Z Image Turbo Model first time in community

The quality is almost same as BF16 I have tested

Moreover, now if you have properly installed Visual Studio\2022\BuildTools\VC\Tools\MSVC it will auto start the training with cl.exe included enviroment

This allows you to use more complex Torch Compile options but this is not mandatory

Follow this tutorial to install BuildTools : https://youtu.be/DrhUHnYfwC0

To update overwrite older files and run Windows_Install_and_Update.bat file

Qwen Image Tutorial Video Instructions

Main training tutorial : https://youtu.be/DPX3eBTuO_Y

Ultra realism tutorial : https://youtu.be/XWzZ2wnzNuQ

This section is made for Qwen Image training comprehensive tutorial video

First of all follow below requirements tutorial and make sure all requirements properly installed or updated

https://youtu.be/DrhUHnYfwC0

This is mandatory to watch and apply to not have any issues

Download the latest zip file from either attachments (down below) or Latest Zip File (at the top of the post)

The follow the tutorial video every step

Auxiliary tools are as below

Ultimate Batch Image Preprocessing app

https://www.patreon.com/posts/120352012

Batch Image Caption Editor app

https://www.patreon.com/posts/108992085

How to install and use SwarmUI with ComfyUI backend (mandatory to watch - recorded few days ago fully up-to-date)

https://youtu.be/c3gEoAyL2IE

We use manually installed ComfyUI backend which supports all the latest libraries and all GPUs

Extremely fast Hugging Face upload / download notebook (needed for Cloud training but can be used in Windows too)

https://youtu.be/X5WVZ0NMaTg

Training used Style images dataset

GTA5_Style_Dataset.zip

Model : https://civitai.com/models/2084406?modelVersionId=2358426

Example training images dataset if you ever need

https://www.patreon.com/posts/114972274

14 November 2025 V20 Update

Main training tutorial: https://youtu.be/DPX3eBTuO_Y

Realism tutorial : https://youtu.be/XWzZ2wnzNuQ

Scroll down to get to Qwen Image Tutorial Video Instructions

Hopefully Torch Compile coming very soon

Totally free 15%+ speed up

Kohya adding that after I made this issue thread : https://github.com/kohya-ss/musubi-tuner/issues/716

Pull request is here : https://github.com/kohya-ss/musubi-tuner/pull/722

Once merged into main, hopefully I will add to our App as soon as possible

We have added 2 new feature tabs

1: LoRA Extractor

Supports extracting LoRA from a Fine-Tuned Qwen Model - works amazing

No other models supported yet by Musubi Tuner repo

2: LoRA Merger

Supports Merging LoRAs for Qwen, Skyreels, Wan

You can Merge LoRA to LoRA or Merge + Base Model to Base Model

I didn't find out how to use it best yet so try with different Multipliers

Only tested with Qwen LoRAs so far

Use latest zip file, extract and overwrite, run Windows_Install_and_Update.bat to update

30 October 2025 Update V19

We are finally fully ready to Qwen Image Tutorial with V19

All configs updated like below

Below ones are LoRA

Below ones are full Fine Tuning

New option Faster Model Loading (Uses more RAM but speeds up model loading speed - Enable for RunPod) added

Now when you enable Enable Qwen-Image-Edit-2509 Mode it will auto set metadata so SwarmUI will auto recognize checkpoints - this is needed for Fine Tuning

Now all Fine Tuning configs have extra args of setting metadata as 1328x1328 so SwarmUI will auto recognize resolution accurately

Zip file content below

28 October 2025 Update V18.7

Bug fixes and a new amazing feature called as Image Preprocessing and FP8 Model Converter

Scaled FP8 Model Converter is used to batch convert Qwen Image Fine-Tuned / DreamBooth models into scaled FP8 from BF16 weights

Extremely useful if you don't have over 60 GB GPUs and also reduces size to half

Almost same quality with using intelligent FP8 Scaled weight conversation

What this tool does is that, it preprocess your given folder image with Kohya Musubi tuner actual training code

So your pre-processed images is the images that are actually being trained by the trainer

Extremely useful to see your resolution, aspect ratios and what bucketing does

Moreover, currently Kohya doesn't use exif data image orientation so if you might get mis-oriented images surprise, use this tool to see and you can use fixed dataset as well

Detailed debug options added to latent caching section so you can enable debug and 1 by 1 look at the processed images

You can use pre-processed dataset as your training dataset if you wish saves time and resources

21 October 2025 Update V18.0

Qwen Image Edit Plus (2509) Fully supported now

Just load the DreamBooth or LoRA Config and then follow the below steps

Original files are included inside Qwen_Image_Training_Configs folder check each one

Extract latest zip file, overwrite your existing files and run Windows_Install_and_Update.bat to update

Model downloader now supports downloading Qwen Image Edit Plus (2509) model set download as well

Run Windows_Download_Training_Model_Files.bat to download desired model

Qwen Image Fine Tuning configs completed finally

Working on updating inference presets in SwarmUI

The logic of Qwen Image Edit Model training shown below

21 October 2025 Update V17.9

Qwen Image Full Fine Tuning / DreamBooth research finally completed

Hopefully full tutorial coming and more information will be shared soon

We have made configs starting from 5750_MB (6 GB GPUs) to 84000_MB (96 GB GPUs) GPUs

The lower VRAM requiring configs will use more RAM since we do block swapping and also they will be slower

I have prepared 3 different config set

200_epoch folder - best quality - lower learning rate - more epochs (more epochs takes more time linearly)

150_epoch - good quality - a little bit higher learning rate - medium epochs

75_epoch - lower quality - higher learning rate - lower epochs - faster training

All qwen configs moved into Qwen_Image_Training_Configs and seperated as LoRA_Training and Fine_Tuning_Training-DreamBooth sub folders

All configs updated including LoRA and below 20 GB configs modified like this

T5 text encoder caching will be made on CPU with BF16 - still really fast

After tutorial for this hopefully I will research Wan 2.2 training for ultra realistic image gen + video gen

20 October 2025 Update V17.8

Windows_Download_Training_Model_Files.bat updated and now you can download Wan 2.2 Text to Video and Image to Video models

I have converted official FP32 Wan weights into BF16 so use our downloader to download BF16 weights of Wan 2.2

BF16 yields much better quality at training and also slightly better or sometimes significantly better quality at inference

Example Wan 2.2 train configs added

Wan_2.2_Text_to_Video_Low_Noise_Test_LoRA_Config_10250_MB.toml

Uses 10250 MB VRAM

Wan_2.2_Text_to_Video_Both_Low_and_High_Noise_Test_LoRA_Config_12000_MB.toml

Uses 12000 MB VRAM

Example configs are using BF16 weights of Wan 2.2 now

Massed_Compute_Instructions_READ.txt and Runpod_Instructions_Trainer_READ.txt updated to have Wan 2.2 model download options

Wan training research results coming soon hopefully

Qwen Image fine tuning completed i am finalizing to publish hopefully

3 October 2025 Update V17.5

Please extract latest zip file, overwrite older files

Delete SECourses_Musubi_Trainer\musubi-tuner folder and run Windows_Install_and_Update.bat to update

Switch to low memory branch file removed since it is now merged

Bug fix on Linux systems - RunPod & Massed Compute for Wan training

Wan models training directory explanation improved : https://pasteboard.co/y5fI2Sf8Bc0L.png

29 September 2025 Status Update V17

Wan training support added

I think it is supporting all of the Wan trainings that Musubi Tuner supports

So far only Wan 2.1 Text to Video model LoRA training tested and a demo config put as : Wan_2.1_text_to_video_test_lora_8500_mb.toml

Windows_Download_Training_Model_Files.bat updated to download Wan 2.1 Text to Video models

The models will be downloaded into Training_Models_Qwen and Training_Models_Wan according to your selection

Model downloader updated for RunPod and Massed Compute as well

You can read full changelogs on app and below

Extract zip file, overwrite and run Windows_Install_and_Update.bat for update

27 September 2025 Status Update V16

Qwen Image LoRA trainings research and development completed fully

I have done over 50 full trainings to find out best parameters and workflow

You can see training checkpoints and files here : https://huggingface.co/MonsterMMORPG/Qwen_Image_LoRA_Training_Research/tree/main

To access files you can message me and i can let you with a special price but normally this is not needed by you

I have updated all the configs after recent experiments

I feel like 6e-05 is working slightly better than 5e-05 so updated configs to that

The latest experiments put into Final_Qwen_LoRA_Test_Results.zip file

I have trained different learning rates up to 1000 epochs on 28 images dataset and you can see grid results in the attached Final_Qwen_LoRA_Test_Results.zip file

So if you need faster training, you can possibly increase learning rate to like 8e-5 or 9e-5 and do like 100 or less but it is R&D depending on your needs

The app is now fully supporting Qwen Image full Fine Tuning / DreamBooth

I added a demo config file : dreambooth.toml

This demo config toml uses 5 GB VRAM only when you also run Windows_Switch_Low_RAM_Branch_Temporary.bat

There is a pull request of Kohya which reduces used shared VRAM signficiantly but it is not merged yet so we switch to that branch with that bat file

When you run Windows_Download_Training_Model_Files.bat again it will switch back to main branch

I am starting to fully research Qwen Image full Fine Tuning and also adding Wan 2.2 and Wan 2.1 training capability to the app hopefully as soon as possible

Installers updated to Torch 2.8, CUDA 12.9, Flash Attention 2.8.3, Sage Attention 2.2, xFormers 0.0.33

I have pre-compiled libraries for both Windows and Linux and working amazing

Windows Requirements

Python 3.10.11, FFmpeg, CUDA 12.9, cuDNN 9.12 or above, C++ tools, MSVC and Git

Don't worry CUDA 12.9 works with all GPUs

Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0

Follow its updated post with links and screenshots exactly : https://www.patreon.com/posts/click-to-open-post-used-in-tutorial-111553210

5 September 2025 Status Update V14

Some bug fixes and Configs_v1 added that you can use

New option added for sample generation Disable Automatic Prompt Enhancement

Thus you can provide fully customized prompt txt

Space character having paths will now work but still dont use space character in any path

.e.g my awesome images - is wrong but my_awesome_images is right

We support as low as 6 GB GPUs

Tier 1 is better than Tier 2 and Tier 2 better than Tier 3 and so on

There shouldn't be any big difference between Tier 1 and Tier 2

Tier 2 and Tier 3 also should be pretty close in terms of quality

To update download latest zip file, extract and overwrite older files and run Windows_Install_and_Update.bat

31 August 2025 Update V11

Now you can immediately and properly stop batch image captioning as well

Some Critical config save load bugs fixed please upgrade

e.g. : FIXED: Critical checkpoint removal bug - Checkpoints were being deleted immediately after saving when save_last_n_epochs=0

Now we give GPU ID = 0 as a default, if you have multiple GPUs like onboard GPU and external GPU, make sure the ID matches to your external GPU ID

It is set in Distributed GPUs and that may look like multiple GPU config but when no multiple GPU enabled, it still uses it

UI significantly improved please check newest screenshots given at the very top

The downloader app improved and now it will show the progress of each model on a single line not in multiple lines as it downloads

Added comprehensive parameter support for Qwen Image training with 100% coverage of Musubi Tuner parameters

NEW: Integrated search bar in Qwen Image Training tab - quickly find any setting without opening all panels

Tab renamed from "Qwen Image LoRA" to "Qwen Image Training" to reflect both LoRA and Fine-tuning support

Implemented Qwen-Image-Edit mode support for control image training (experimental - not fully tested)

Added control image resolution settings for Edit mode (dataset_qwen_image_edit_control_resolution_width/height)

Introduced dataset_qwen_image_edit_no_resize_control option for maintaining original control image sizes - this is for Qwen Image Edit model - not tested yet

Started implementation of Qwen Image Fine-Tuning mode (DreamBooth) - parameter infrastructure in place

Enhanced FP8 quantization descriptions with clearer GPU compatibility information

Improved timestep sampling with better documentation of qwen_shift vs standard shift methods

Added advanced flow matching parameters (logit_mean, logit_std, mode_scale) for fine-tuned control

Implemented complete VAE optimization settings (tiling, chunk_size, spatial_tile_sample_min_size)

Enhanced parameter descriptions throughout the GUI for better user understanding

All critical default values confirmed to match official Qwen Image documentation

Parameter accuracy validated at 100% for all implemented features

Some configs save and load were broken and all fixed hopefully

Please report if you notice error

e.g. one of the fixed one is GPU ID set

Please use latest zip file, overwrite previous files and just run Windows_Install_and_Update.bat

30 August 2025 Update V7

Dataset TOML file generate error fixed

Qwen2.5-VL image captioning turns out working perfect on Windows

It turns out my model file was corrupted even though it was same size

Therefore I have updated the model downloader and now it will check and verify SHA 256 of files therefore it will be 100% accurate

Prompt file selection folder icon issue fixed

Downloader file will use generated venv of installation

Make sure to run it after installation completed

Fixed skip existing captions functionality in Image Captioning with Qwen2.5-VL

Previously skipping was happening after caption generation which was destroying the skip logic

Now properly checks for existing captions before processing, significantly improving efficiency

Added full batch captioning status display in command line with progress tracking and ETA

Enhanced config save/load functionality for better reliability

Improved interface of Image Captioning with Qwen2.5-VL for better user experience

Various error fixes in the Qwen2.5-VL captioning pipeline

Fixed broken config save and load functionality for Optimizer Arguments and Scheduler Arguments

Improved Stop Training button responsiveness - now appears much earlier when Text Encoder caching starts

Enhanced training control for better user experience

A new full dedicated section for Sample generation implemented

It will automatically format your given sample txt file with the settings you set on GUI

So you just type prompts into txt file with new lines e.g.

ohwx man wearing a very nice amazing suit

ohwx man driving a luxury car

Please use latest zip file, overwrite previous files and just run Windows_Install_and_Update.bat

How To Give Accurate Dataset for Dataset Prepare

Example folder : E:\training_imgs_28_1328

So you give above path into UI

Inside this folder example sub folder for dataset images

E:\training_imgs_28_1328\1_ohwx man

Inside E:\training_imgs_28_1328\1_ohwx man

I have my images as below

How To Install and Use:

Windows:

Full tutorial coming soon hopefully

Use same folder logic of Kohya and use Generate Dataset Configuration button it will handle all

e.g. Parent Folder > sub folder like 1_ohwx man

Just use Windows_Install_and_Update.bat for install and update

Massed Compute (Recommend Cloud) :

Please register via this link : https://vm.massedcompute.com/signup?linkId=lp_034338&sourceId=secourses&tenantId=massed-compute

We have a special coupon for all GPUs : SECourses

If you want to learn more about GPUs and prices read this link : https://www.patreon.com/posts/126671823

Select RTX A6000 or Better GPU - like L40S or A6000 ADA or A100 or H100 or now RTX 6000 PRO

Then select our image SECourses from Creator dropdown

Then follow Massed_Compute_Instructions_READ.txt

Same as my any other Massed Compute installer script

Example tutorial for learn how to install and use Massed Compute

(Starts at 12:58) : https://youtu.be/KW-MHmoNcqo?si=G1WbG-Qw4ujWvOtG&t=778

RunPod (Cloud):

Please register via this link : https://get.runpod.io/955rkuppqv4h

Then follow Runpod_Instructions_READ.txt

Same as my any other RunPod installer script

Use the template written in Runpod_Instructions_READ.txt file

Example tutorial for learn how to install and use RunPod

(starts at 22:03) : https://youtu.be/KW-MHmoNcqo?si=QN8X8Sjn13ZYu-EU&t=1323

To See All Screenshots : https://www.reddit.com/r/SECourses/comments/1n4qq8y/massive_updates_and_improvements_to_secourses/

SECourses: FLUX, Tutorials, Guides, Resources, Training, Scripts PATREON 32 favs
VIEWS1
FILES38 files
POSTEDJul 14, 2026
ARCHIVEDJun 10, 2026