SECourses Musubi Tuner - 1-Click to Install App for LoRA Training and Full Fine Tuning Krea 2, Ideogram 4, Qwen Image, Qwen Image Edit, Wan 2.1 and Wan 2.2, FLUX Klein, FLUX 2, Z Image Base and Turbo Models with Musubi Tuner with Ready Presets
Patreon exclusive posts index to find our scripts easily, Patreon scripts updates history to see which updates arrived to which scripts and amazing Patreon special generative scripts list that you can use in any of your task.
Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord
Please also Star, Watch and Fork our Stable Diffusion & Generative AI GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)
=======
Latest Zip File : SECourses_Musubi_Trainer_v30_2.zip
[click here to choose a membership and Join to download zip files]
Main training tutorial to learn training and how to use this app (mandatory to watch) : https://youtu.be/DPX3eBTuO_Y
Wan 2.2 training tutorial : https://youtu.be/ocEkhAsPOs4
SwarmUI : https://www.patreon.com/posts/114517862
ComfyUI : https://www.patreon.com/posts/105023709
=======
Currently 1-click to install on Windows, RunPod and Massed Compute with uv installation - ultra fast
We have model auto downloader that supports below models (run Windows_Download_Training_Model_Files.bat)
The following models have full Fine Tuning / DreamBooth and LoRA training with fully supported Torch Compile (really speeds up training up to 58% with no tradeoff)
All Qwen Image Models, FLUX 2 Dev, FLUX 2 Klein 4B and 9B, Z-Image Base, Z-Image Turbo, Ideogram 4, Krea 2 Raw, Krea 2 Turbo
Also has automatic Qwen VL captioning
The following models only supported with LoRA training with fully supported Torch Compile (really speeds up training up to 58% with no tradeoff)
All Wan 2.1/2.2 variants
Please use model downloader to not have any issues because your selected models may be wrong
Moreover, the Musubi Tuner automatically does FP8 and FP8 scaled conversion while loading BF16 model into RAM so we always use BF16 models for training
Our model downloader is ultra optimized and can reach 1GB per second on cloud machines
It also has SHA256 verification to ensure no corrupt model ever used
We are using our own fork of Musubi Trainer does it has many more features and improvements
The installers will generate a Python 3.12 VENV automatically and install everything inside there, thus your system or any other of your APPs will never be impacted
With our pre-compiled abi3 wheels for Windows and Linux, you can run on Python 3.10, 3.11, 3.12 and 3.13 but preferred version is 3.12
The following libraries compiled for Windows and Linux
archs : mslk, xformers, flash_attn, sageattention, torchao,
All of them are abi3 and Windows versions are compiled for Consumer GPUs, Linux versions compiled for Consumer + Cloud GPUs - so we cover all GPUs you have
Windows Requirements
For this auto installer to work you need to have installed Python 3.12.10 (may work with 3.10.x, 3.11.x, 3.13.x too), Git, FFmpeg, cuDNN 9.17+, CUDA 13.0, Visual Studio Community Edition with All c++ options
Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0
Follow its updated post with links and screenshots exactly : https://www.patreon.com/posts/windows-AI-requirements-tutorial-111553210
18 July 2026 V30.2 Update
We have added LTX 2.3 Int8 Row ConvRot HQ quantization preset to our Model Quantizer tab
The difference of Int8 Row ConvRot HQ presets is not very well known in the community
I have so far compiled Krea 2 and LTX 2.3 recently, each takes around 3-4 hours on RTX 5090
You can get them with our SwarmUI model downloader app, bundles are now also downloading them : https://www.patreon.com/posts/114517862
The difference is that, Int8 Row ConvRot HQ is better than GGUF Q8 quality and more than 100% faster on RTX 3000, 4000 and 5000 series (don't have 2000 or 1000 series so can't tell)
Krea 2 quality and speed test as below
Int8 Row ConvRot is 96.2% similar to BF16 meanwhile GGUF Q8 is only 90.0% and FP8 Scaled is 82.2% and NVFP4 is 63.7%
Moreover, Int8 Row ConvRot generates the output in 3.05 seconds, making it 1.82× faster than BF16, which takes 5.56 seconds.
NVFP4 takes 3.8 seconds and is 1.46× faster than BF16, whereas GGUF Q8 takes 6.06 seconds and is approximately 8.3% slower than BF16.
LTX 2.3 Int8 Row ConvRot Benchmarks are as Below
So Int8 Row ConvRot HQ is 100% faster than FP8 Quant Scaled and 50% faster than BF16 on RTX 5090
The quality is also excellent almost same as BF16
To update from V30, just run Windows_Install_and_Update.bat file
Make sure to have latest zip file always and overwrite older installer files
14 July 2026 V30.0 Update
This is a major update with lots of new features and massive performance improvements
Now we are finally able to catch Linux training speeds, I tested Krea 2 and verified at least for Krea 2, read entire update to see
Krea 2 - 36%, FLUX 2 Klein 9B - 46%, Z-Image Base / Turbo - 53%, Ideogram 4 - 14%, Qwen 2512 - 24.5%, Qwen 2511 - 40%, Wan 2.1 - 22%, Wan 2.2 - 15%, FLUX 2 - 15.6% faster
The following models are now fully supported for Full Fine Tuning / DreamBooth training
FLUX.2 dev, FLUX Klein 4B and 9B, Ideogram 4, Krea 2 Raw and Turbo
Demo presets updated as below
New shared full-model training engine
Full FP32 or BF16 DiT training and checkpoint export.
Exact training resume, including mid-epoch resumes and epoch boundaries.
Training-time sampling, checkpoint retention, metadata, Hugging Face upload, and memory-efficient saving.
Single-GPU block swapping and fused-backward memory optimizations.
Ordinary multi-GPU DDP support with automatic validation of incompatible options.
FLUX.1 Kontext full-DiT training is also available through the Musubi backend.
Three new experimental Automagic optimizers implemented from Ostris AI Toolkit (https://www.patreon.com/SECourses/posts/ostris-ai-1-for-140089077)
Automagic, Automagic2, and Automagic3 are available across the training tabs.
Adaptive learning rates with different per-element, per-tensor, and per-group strategies.
Support for LoRA and full Fine-Tuning / DreamBooth, block swapping, checkpoint resume, and low-precision training.
Automatic fused/non-fused selection where supported.
Unsafe combinations are rejected before training instead of silently producing a broken run.
The GUI now displays detailed optimizer-specific guidance.
Native PyTorch fused SDPA is preferred.
When unavailable, external FlashAttention is used only after passing real CUDA forward/backward compatibility tests.
Unsupported tensor types or layouts automatically fall back to working PyTorch SDPA.
Every GUI training tab now has a Use Legacy PyTorch SDPA compatibility switch.
Smarter torch.compile with block swapping
Make sure to enable for Wan 2.2 while Wan 2.1 retains its standard default.
Explicit user settings are preserved when switching tasks or loading configurations.
Improved handling of empty or missing Accelerate launch values.
Hardened Qwen and Z-Image full-model configuration previews and runtime exports.
Added correct full-fine-tuning metadata to Qwen and Z-Image checkpoints.
Maximum absolute gradient
All measured before clipping and sent to TensorBoard or Weights & Biases.
Added CUDA 13.2/PyTorch packaging support.
Improved Ninja and compiler discovery for virtual environments and Visual Studio.
The Massive Speed Improvements Are As Below
All trainings made at 1024x1024px and batch size 1, for LoRAs, LoRA rank is 128
Click images to see full sizes
All speed ups have absolutely 0 tradeoff except Torch Compile initial time requirement - so exactly same max quality
Krea 2 - 36% faster compared to default Musubi Trainer (default no Torch Compile)
We implemented special way of Torch Compile discovery so fully working on Windows too
FLUX 2 Klein 9B - 46% faster
Z-Image Base / Turbo - 53% faster
Ideogram 4 - 14% faster
Qwen 2512 - 24.5% faster
Qwen 2511 - 40% faster - Trained with edit images
Wan 2.1 - 22% faster
Wan 2.2 - 15% faster
FLUX 2 Dev - 15.6% faster
For updating from V29, extract zip file, overwrite files and just run install / update bat file
For updating older version, make a fresh new install
12 July 2026 V29.0 Update
This is a pretty massive update please read all
Please make a seperate new fresh install
We have fully moved to Torch 2.13.0 and CUDA 13.1 with pre-compiled libraries
Also now our preferred Python is 3.12.10, still should work with 3.10, 3.11, 3.13 but please have 3.12 for best
RunPod, SimplePod, Massed Compute installers auto installs 3.12 so you don't need to do anything
Linux users can use Massed Compute installers
For Windows, please use Windows_Install_and_Update.bat and Windows_Download_Training_Model_Files.bat
For Massed Compute and Local Linux : Massed_Compute_Instructions_READ.txt
For RunPod and Simple Pod : Runpod_SimplePod_Musubi_Trainer_Instructions.txt
I have started upgrading all of our apps into latest Torch 2.13 and CUDA 13
For this, I have pre-compiled the following wheels with all CUDA 13 features and with all CUDA archs : mslk, xformers, flash_attn, sageattention, torchao
All these libraries are properly compiled with abi3 thus works on Python 3.10, 3.11, 3.12 and 3.13
Full Krea 2 training added - currently only LoRA
Fully Ideogram 4 training added - currently only LoRA
Demo presets are available inside Demo_Training_Configs_FLUX-2_Z-Image_FLUX-Klein_WAN-21_Krea2_Ideogram4 folder
Hopefully fully researched parameters will come for Krea 2 and others soon
I really like Krea 2 so far, really great model
Model downloader updated for newer models and improved
Model downloader is now more robust and faster
Now when you load a config, it will check if config has valid model paths, if not, it will scan default model downloader downloaded model paths and auto set them
Works on all platforms Windows, Linux, Cloud
For example on RunPod and Massed Compute automatically filled model paths like below
Now you can download multiple models with coma seperation or ranges like 1,2,3 or 1-3,5
For JSON prompt generation we have new tutorial and tool that is useful for Ideogram 4 model training
Tutorial : https://youtu.be/TW3MRdd0MV4
Tool : Ultimate Image Captioner Pro https://www.patreon.com/SECourses/posts/ultimate-image-captioner-pro-162527725
Model Quantizer tab completely improved and now works much better with newer presets
Now our ComfyUI and SwarmUI fully supports Int8 Row ConvRot and this quantization is insane soon hopefully I will make a tutorial and show and also update SwarmUI and ComfyUI posts to show
I recommend to use INT8 ConvRot Learned (Best Quality / Slow) preset to generate quantization of base models it is slow but insane quality even better than GGUF Q8 and 2x faster on all GPUs
Quantization is only for base models not for LoRAs
Now warnings and errors will be displayed on Gradio too with notice bubles
Now Torch Compile works even better and C++ tools not needed, only Visual Studio Community Edition with C++ options please see updated requirements
If you get any errors follow below video and its source link
Make sure that your NVIDIA driver is updated (min 590+)
https://www.patreon.com/posts/windows-requirements-tutorial-written-post
Now displays accurate training step speed almost immediately after training started accurately
I fixed this issue on my maintained fork
Hopefully I will add LTX 2.3 training as well soon with audio support
Krea 2 Demo preset training VRAM usage and Speed as below for 1024x1024px
Krea 2 Demo preset with 0 Block Swap + Torch Compile training speed and VRAM usage as below at 1024x1024px
Ideogram 4 Demo preset training VRAM usage and Speed as below for 1024x1024px
Ideogram 4 Demo preset with 0 Block Swap + Torch Compile training speed and VRAM usage as below at 1024x1024px
Tests are made on Windows 11 with RTX 5090
Torch Compile do not affect model quality
4 June 2026 V28.3 Update
Since requested new feature implemented into Image Captioning
When overwrite is off, append new captions to existing text files instead of skipping them
Append Existing Captions
Just run Windows_Install_and_Update.bat to update
20 April 2026 V28.2
Torchaudio version bug fixed
Quantizer application updated to latest and now it uses prodigy to generate FP8 Scaled Quants or NVFP4
FLUX 1 DEV preset added
FLUX 2 Klein Models preset added
ERNIE preset added
All presets updated to use new prodigy and thus now faster and better quality
Hopefully I will add new Quant models and presets to our model downloader zip file soon : https://www.patreon.com/posts/swarmui-auto-and-114517862
Get latest zip file, extract and overwrite all and run Windows_Install_and_Update.bat to update
8 March 2026 V27.6
Model quantizer app significantly updated
Presets are updated and made much better
Zip file is still same just use Windows_Install_and_Update.bat to update
This FP8 Scaled generation is much better than default FP8 and here why
Currently generating LTX2.3 distilled FP8 Scaled max quality for you :)
10 February 2026 V27.5
Sample generation section and backend for Qwen, Wan, FLUX and Z-Image improved and fixed
7 February 2026 V27.2
This is a pretty big update please read all carefully
We have added new Model Quantizer page which supports FP8 Scaled, NVFP4 and more quantization - still in BETA
FLUX Training tab added
FLUX 2 Dev LoRA Training with Torch Compile - min 18 GB VRAM with 128 LoRA Rank and 1024px training (you can reduce LoRA rank and resolution to reduce VRAM or speed up training)
FLUX Klein 9B LoRA Training with Torch Compile - min 9.6 GB VRAM with 128 LoRA Rank and 1024px
FLUX Klein 4B LoRA Training with Torch Compile - min 5.6 GB VRAM with 128 LoRA rank and 1024px
Sadly no Fine Tuning of FLUX 2 or FLUX Klein yet, I notified Kohya and hoping soon
Z Image Training tab added
Z-Image Base Fine Tuning with Torch Compile - min 6 GB VRAM with 1024px
Z-Image Base LoRA Training with Torch Compile - min 6.1 GB VRAM with 128 LoRA rank and 1024px
Z-Image Turbo LoRA Training with Torch Compile - min 8.6 GB VRAM with 128 LoRA rank and 1536px
To update, extract latest zip file, overwrite older files and run Windows_Install_and_Update.bat
To download accurate training files use Windows_Download_Training_Model_Files.bat
New demo configs exists in Demo_Training_Configs_FLUX-2_Z-Image_FLUX-Klein_WAN-21 folder
Z-Image LoRA training config is set similar to AI Toolkit Z Image Turbo training tutorial config, the other ones are set approximately and full R&D will be made later hopefully and final official configs will be published later like Qwen_Training_Configs and Wan22_Training_Configs hopefully with a tutorial
Currently configs are set for lowest VRAM thus if you have higher VRAM, you can reduce block swap count and such
16 January 2026 V26.2 Zip File
RunPod template link updated and now we fully support SimplePod which is much faster and cheaper than RunPod
SIMPLEPOD CHEAPER AND FASTER THAN RUNPOD
Now we fully support SimplePod as well please use this link to register : https://simplepod.ai/ref?user=secourses
SimplePod is faster and cheaper than RunPod and works exactly same
E.g. RTX 5090 on RunPod is 0.89 USD per hour, on SimplePod it is 0.45$ per hour,
RTX PRO 6000 on RunPod is 1.84 USD per hour and on SimplePod it is 0.79 USD per hour
Please use this template on SimplePod : https://dash.simplepod.ai/account/explore/100/ref-secourses/
For permanent storage, generate it from Storage tab with any name and size you want and when selecting template with above link, click Edit and Use, select Persistence Volume and change mount point to /workspace
Up-to-date SimplePod tutorial starting from 21:51 : https://youtu.be/yOj9PYq3XYM?si=Z86wZZLBeYzWo1Qo&t=1311
As usual follow Massed_Compute_Instructions_READ.txt and Runpod_SimplePod_Musubi_Trainer_Instructions.txt to install and use and watch the tutorials
4 January 2026 V26.2 Update
For FP8 Model Converter tab for Improved (convert_to_quant learned rounding) method new feature Use pinned RAM for faster GPU transfers (page-locked memory) added
1 January 2026 V26 Update
Qwen Image 2512 BF16 added into downloader app you can train it exactly as Qwen Image 0 difference
New highly experimental and expert stuff Improved (convert_to_quant learned rounding) added to FP8 Converter tool
It is based on https://github.com/silveroxides/convert_to_quant repo
This is total expert stuff so i am not supporting normally
Just get and extract latest zip file to get new download and Windows_Install_and_Update.bat to have new FP8 Converter tool
31 December 2025 V25.2 Update
Now you can set Target frames on interface - this matters
Now there is Auto Normalize Target Frames which i recommend you to enable - auto enabled
Both options added to Wan Training Dataset preperation tab
29 December 2025 V25.1 Update
Installers upgraded to uv pip installation
Installation on RunPod is now like 100x faster literally and many times faster on Windows and RunPod
Long awaited Wan 2.2 training tutorial published : https://youtu.be/ocEkhAsPOs4
Qwen Image Edit 25-11 Model added to training models downloader ( Windows_Download_Training_Model_Files.bat )
Qwen Image Edit 25-11 Training support implemented just run Windows_Install_and_Update.bat to update
You can train it without control images just as regular text to image model
You will select your model version from dropdown now
17 December 2025 V23 Update
This is a massive update - Long waited Wan 2.2 official training configs published
I have fully researched Wan 2.2 with literally over 64 uniqute trainings and analyzed all of the results - I have used a 8x B200 cloud machine for this research
Hopefully will make a tutorial for Wan 2.2
After watching Qwen Tutorial, exactly same way you can train your Wan 2.2 Text to Image or Text to Video or Image to Video model
Just using static images datasets working perfect I have tested
We have configs for literally every GPU check the configs below
Hopefully will explain more in tutorial
Qwen Image Fine Tuning Tier1_84000_MB.toml config was broken due to Fused Backward pass Off - currently it is enabled - This is bug in Kohya Musubi
When training Wan 2.2 models, there were few serious bugs and all fixed and working perfect now
Extract latest zip file and get latest configs and run Windows_Install_and_Update.bat to update
29 November 2025 V21 Update
Big news, now we have famous Torch Compile support for Qwen and WAN and all QWEN Training configs we have are modified to have auto Torch Compile enabled
This brings around 7.5-15% speed up and lowers VRAM usage to some degree with absolutely 0% quality loss or anything so this is a total win for Speed and VRAM usage
Moreover, FP8 Model Converter now supports Z Image Turbo Model as well and that is how I published the FP8 Scaled model of Z Image Turbo Model first time in community
The quality is almost same as BF16 I have tested
Moreover, now if you have properly installed Visual Studio\2022\BuildTools\VC\Tools\MSVC it will auto start the training with cl.exe included enviroment
This allows you to use more complex Torch Compile options but this is not mandatory
Follow this tutorial to install BuildTools : https://youtu.be/DrhUHnYfwC0
To update overwrite older files and run Windows_Install_and_Update.bat file
Qwen Image Tutorial Video Instructions
Main training tutorial : https://youtu.be/DPX3eBTuO_Y
Ultra realism tutorial : https://youtu.be/XWzZ2wnzNuQ
This section is made for Qwen Image training comprehensive tutorial video
First of all follow below requirements tutorial and make sure all requirements properly installed or updated
This is mandatory to watch and apply to not have any issues
Download the latest zip file from either attachments (down below) or Latest Zip File (at the top of the post)
The follow the tutorial video every step
Auxiliary tools are as below
Ultimate Batch Image Preprocessing app
https://www.patreon.com/posts/120352012
Batch Image Caption Editor app
https://www.patreon.com/posts/108992085
How to install and use SwarmUI with ComfyUI backend (mandatory to watch - recorded few days ago fully up-to-date)
We use manually installed ComfyUI backend which supports all the latest libraries and all GPUs
Extremely fast Hugging Face upload / download notebook (needed for Cloud training but can be used in Windows too)
Training used Style images dataset
Model : https://civitai.com/models/2084406?modelVersionId=2358426
Example training images dataset if you ever need
https://www.patreon.com/posts/114972274
14 November 2025 V20 Update
Main training tutorial: https://youtu.be/DPX3eBTuO_Y
Realism tutorial : https://youtu.be/XWzZ2wnzNuQ
Scroll down to get to Qwen Image Tutorial Video Instructions
Hopefully Torch Compile coming very soon
Totally free 15%+ speed up
Kohya adding that after I made this issue thread : https://github.com/kohya-ss/musubi-tuner/issues/716
Pull request is here : https://github.com/kohya-ss/musubi-tuner/pull/722
Once merged into main, hopefully I will add to our App as soon as possible
We have added 2 new feature tabs
1: LoRA Extractor
Supports extracting LoRA from a Fine-Tuned Qwen Model - works amazing
No other models supported yet by Musubi Tuner repo
2: LoRA Merger
Supports Merging LoRAs for Qwen, Skyreels, Wan
You can Merge LoRA to LoRA or Merge + Base Model to Base Model
I didn't find out how to use it best yet so try with different Multipliers
Only tested with Qwen LoRAs so far
Use latest zip file, extract and overwrite, run Windows_Install_and_Update.bat to update
30 October 2025 Update V19
We are finally fully ready to Qwen Image Tutorial with V19
All configs updated like below
Below ones are LoRA
Below ones are full Fine Tuning
New option Faster Model Loading (Uses more RAM but speeds up model loading speed - Enable for RunPod) added
Now when you enable Enable Qwen-Image-Edit-2509 Mode it will auto set metadata so SwarmUI will auto recognize checkpoints - this is needed for Fine Tuning
Now all Fine Tuning configs have extra args of setting metadata as 1328x1328 so SwarmUI will auto recognize resolution accurately
Zip file content below
28 October 2025 Update V18.7
Bug fixes and a new amazing feature called as Image Preprocessing and FP8 Model Converter
Scaled FP8 Model Converter is used to batch convert Qwen Image Fine-Tuned / DreamBooth models into scaled FP8 from BF16 weights
Extremely useful if you don't have over 60 GB GPUs and also reduces size to half
Almost same quality with using intelligent FP8 Scaled weight conversation
What this tool does is that, it preprocess your given folder image with Kohya Musubi tuner actual training code
So your pre-processed images is the images that are actually being trained by the trainer
Extremely useful to see your resolution, aspect ratios and what bucketing does
Moreover, currently Kohya doesn't use exif data image orientation so if you might get mis-oriented images surprise, use this tool to see and you can use fixed dataset as well
Detailed debug options added to latent caching section so you can enable debug and 1 by 1 look at the processed images
You can use pre-processed dataset as your training dataset if you wish saves time and resources
21 October 2025 Update V18.0
Qwen Image Edit Plus (2509) Fully supported now
Just load the DreamBooth or LoRA Config and then follow the below steps
Original files are included inside Qwen_Image_Training_Configs folder check each one
Extract latest zip file, overwrite your existing files and run Windows_Install_and_Update.bat to update
Model downloader now supports downloading Qwen Image Edit Plus (2509) model set download as well
Run Windows_Download_Training_Model_Files.bat to download desired model
Qwen Image Fine Tuning configs completed finally
Working on updating inference presets in SwarmUI
The logic of Qwen Image Edit Model training shown below
21 October 2025 Update V17.9
Qwen Image Full Fine Tuning / DreamBooth research finally completed
Hopefully full tutorial coming and more information will be shared soon
We have made configs starting from 5750_MB (6 GB GPUs) to 84000_MB (96 GB GPUs) GPUs
The lower VRAM requiring configs will use more RAM since we do block swapping and also they will be slower
I have prepared 3 different config set
200_epoch folder - best quality - lower learning rate - more epochs (more epochs takes more time linearly)
150_epoch - good quality - a little bit higher learning rate - medium epochs
75_epoch - lower quality - higher learning rate - lower epochs - faster training
All qwen configs moved into Qwen_Image_Training_Configs and seperated as LoRA_Training and Fine_Tuning_Training-DreamBooth sub folders
All configs updated including LoRA and below 20 GB configs modified like this
T5 text encoder caching will be made on CPU with BF16 - still really fast
After tutorial for this hopefully I will research Wan 2.2 training for ultra realistic image gen + video gen
20 October 2025 Update V17.8
Windows_Download_Training_Model_Files.bat updated and now you can download Wan 2.2 Text to Video and Image to Video models
I have converted official FP32 Wan weights into BF16 so use our downloader to download BF16 weights of Wan 2.2
BF16 yields much better quality at training and also slightly better or sometimes significantly better quality at inference
Example Wan 2.2 train configs added
Wan_2.2_Text_to_Video_Low_Noise_Test_LoRA_Config_10250_MB.toml
Uses 10250 MB VRAM
Wan_2.2_Text_to_Video_Both_Low_and_High_Noise_Test_LoRA_Config_12000_MB.toml
Uses 12000 MB VRAM
Example configs are using BF16 weights of Wan 2.2 now
Massed_Compute_Instructions_READ.txt and Runpod_Instructions_Trainer_READ.txt updated to have Wan 2.2 model download options
Wan training research results coming soon hopefully
Qwen Image fine tuning completed i am finalizing to publish hopefully
3 October 2025 Update V17.5
Please extract latest zip file, overwrite older files
Delete SECourses_Musubi_Trainer\musubi-tuner folder and run Windows_Install_and_Update.bat to update
Switch to low memory branch file removed since it is now merged
Bug fix on Linux systems - RunPod & Massed Compute for Wan training
Wan models training directory explanation improved : https://pasteboard.co/y5fI2Sf8Bc0L.png
29 September 2025 Status Update V17
Wan training support added
I think it is supporting all of the Wan trainings that Musubi Tuner supports
So far only Wan 2.1 Text to Video model LoRA training tested and a demo config put as : Wan_2.1_text_to_video_test_lora_8500_mb.toml
Windows_Download_Training_Model_Files.bat updated to download Wan 2.1 Text to Video models
The models will be downloaded into Training_Models_Qwen and Training_Models_Wan according to your selection
Model downloader updated for RunPod and Massed Compute as well
You can read full changelogs on app and below
Extract zip file, overwrite and run Windows_Install_and_Update.bat for update
27 September 2025 Status Update V16
Qwen Image LoRA trainings research and development completed fully
I have done over 50 full trainings to find out best parameters and workflow
You can see training checkpoints and files here : https://huggingface.co/MonsterMMORPG/Qwen_Image_LoRA_Training_Research/tree/main
To access files you can message me and i can let you with a special price but normally this is not needed by you
I have updated all the configs after recent experiments
I feel like 6e-05 is working slightly better than 5e-05 so updated configs to that
The latest experiments put into Final_Qwen_LoRA_Test_Results.zip file
I have trained different learning rates up to 1000 epochs on 28 images dataset and you can see grid results in the attached Final_Qwen_LoRA_Test_Results.zip file
So if you need faster training, you can possibly increase learning rate to like 8e-5 or 9e-5 and do like 100 or less but it is R&D depending on your needs
The app is now fully supporting Qwen Image full Fine Tuning / DreamBooth
I added a demo config file : dreambooth.toml
This demo config toml uses 5 GB VRAM only when you also run Windows_Switch_Low_RAM_Branch_Temporary.bat
There is a pull request of Kohya which reduces used shared VRAM signficiantly but it is not merged yet so we switch to that branch with that bat file
When you run Windows_Download_Training_Model_Files.bat again it will switch back to main branch
I am starting to fully research Qwen Image full Fine Tuning and also adding Wan 2.2 and Wan 2.1 training capability to the app hopefully as soon as possible
Installers updated to Torch 2.8, CUDA 12.9, Flash Attention 2.8.3, Sage Attention 2.2, xFormers 0.0.33
I have pre-compiled libraries for both Windows and Linux and working amazing
Windows Requirements
Python 3.10.11, FFmpeg, CUDA 12.9, cuDNN 9.12 or above, C++ tools, MSVC and Git
Don't worry CUDA 12.9 works with all GPUs
Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0
Follow its updated post with links and screenshots exactly : https://www.patreon.com/posts/click-to-open-post-used-in-tutorial-111553210
5 September 2025 Status Update V14
Some bug fixes and Configs_v1 added that you can use
New option added for sample generation Disable Automatic Prompt Enhancement
Thus you can provide fully customized prompt txt
Space character having paths will now work but still dont use space character in any path
.e.g my awesome images - is wrong but my_awesome_images is right
We support as low as 6 GB GPUs
Tier 1 is better than Tier 2 and Tier 2 better than Tier 3 and so on
There shouldn't be any big difference between Tier 1 and Tier 2
Tier 2 and Tier 3 also should be pretty close in terms of quality
To update download latest zip file, extract and overwrite older files and run Windows_Install_and_Update.bat
31 August 2025 Update V11
Now you can immediately and properly stop batch image captioning as well
Some Critical config save load bugs fixed please upgrade
e.g. : FIXED: Critical checkpoint removal bug - Checkpoints were being deleted immediately after saving when save_last_n_epochs=0
Now we give GPU ID = 0 as a default, if you have multiple GPUs like onboard GPU and external GPU, make sure the ID matches to your external GPU ID
It is set in Distributed GPUs and that may look like multiple GPU config but when no multiple GPU enabled, it still uses it
UI significantly improved please check newest screenshots given at the very top
The downloader app improved and now it will show the progress of each model on a single line not in multiple lines as it downloads
Added comprehensive parameter support for Qwen Image training with 100% coverage of Musubi Tuner parameters
NEW: Integrated search bar in Qwen Image Training tab - quickly find any setting without opening all panels
Tab renamed from "Qwen Image LoRA" to "Qwen Image Training" to reflect both LoRA and Fine-tuning support
Implemented Qwen-Image-Edit mode support for control image training (experimental - not fully tested)
Added control image resolution settings for Edit mode (dataset_qwen_image_edit_control_resolution_width/height)
Introduced dataset_qwen_image_edit_no_resize_control option for maintaining original control image sizes - this is for Qwen Image Edit model - not tested yet
Started implementation of Qwen Image Fine-Tuning mode (DreamBooth) - parameter infrastructure in place
Enhanced FP8 quantization descriptions with clearer GPU compatibility information
Improved timestep sampling with better documentation of qwen_shift vs standard shift methods
Added advanced flow matching parameters (logit_mean, logit_std, mode_scale) for fine-tuned control
Implemented complete VAE optimization settings (tiling, chunk_size, spatial_tile_sample_min_size)
Enhanced parameter descriptions throughout the GUI for better user understanding
All critical default values confirmed to match official Qwen Image documentation
Parameter accuracy validated at 100% for all implemented features
Some configs save and load were broken and all fixed hopefully
Please report if you notice error
e.g. one of the fixed one is GPU ID set
Please use latest zip file, overwrite previous files and just run Windows_Install_and_Update.bat
30 August 2025 Update V7
Dataset TOML file generate error fixed
Qwen2.5-VL image captioning turns out working perfect on Windows
It turns out my model file was corrupted even though it was same size
Therefore I have updated the model downloader and now it will check and verify SHA 256 of files therefore it will be 100% accurate
Prompt file selection folder icon issue fixed
Downloader file will use generated venv of installation
Make sure to run it after installation completed
Fixed skip existing captions functionality in Image Captioning with Qwen2.5-VL
Previously skipping was happening after caption generation which was destroying the skip logic
Now properly checks for existing captions before processing, significantly improving efficiency
Added full batch captioning status display in command line with progress tracking and ETA
Enhanced config save/load functionality for better reliability
Improved interface of Image Captioning with Qwen2.5-VL for better user experience
Various error fixes in the Qwen2.5-VL captioning pipeline
Fixed broken config save and load functionality for Optimizer Arguments and Scheduler Arguments
Improved Stop Training button responsiveness - now appears much earlier when Text Encoder caching starts
Enhanced training control for better user experience
A new full dedicated section for Sample generation implemented
It will automatically format your given sample txt file with the settings you set on GUI
So you just type prompts into txt file with new lines e.g.
ohwx man wearing a very nice amazing suit
ohwx man driving a luxury car
Please use latest zip file, overwrite previous files and just run Windows_Install_and_Update.bat
How To Give Accurate Dataset for Dataset Prepare
Example folder : E:\training_imgs_28_1328
So you give above path into UI
Inside this folder example sub folder for dataset images
E:\training_imgs_28_1328\1_ohwx man
Inside E:\training_imgs_28_1328\1_ohwx man
I have my images as below
How To Install and Use:
Windows:
Full tutorial coming soon hopefully
Use same folder logic of Kohya and use Generate Dataset Configuration button it will handle all
e.g. Parent Folder > sub folder like 1_ohwx man
Just use Windows_Install_and_Update.bat for install and update
Massed Compute (Recommend Cloud) :
Please register via this link : https://vm.massedcompute.com/signup?linkId=lp_034338&sourceId=secourses&tenantId=massed-compute
We have a special coupon for all GPUs : SECourses
If you want to learn more about GPUs and prices read this link : https://www.patreon.com/posts/126671823
Select RTX A6000 or Better GPU - like L40S or A6000 ADA or A100 or H100 or now RTX 6000 PRO
Then select our image SECourses from Creator dropdown
Then follow Massed_Compute_Instructions_READ.txt
Same as my any other Massed Compute installer script
Example tutorial for learn how to install and use Massed Compute
(Starts at 12:58) : https://youtu.be/KW-MHmoNcqo?si=G1WbG-Qw4ujWvOtG&t=778
RunPod (Cloud):
Please register via this link : https://get.runpod.io/955rkuppqv4h
Then follow Runpod_Instructions_READ.txt
Same as my any other RunPod installer script
Use the template written in Runpod_Instructions_READ.txt file
Example tutorial for learn how to install and use RunPod
(starts at 22:03) : https://youtu.be/KW-MHmoNcqo?si=QN8X8Sjn13ZYu-EU&t=1323