Ultimate Image Captioner Pro - Qwen3 VL 8B Instruct with Full JSON and Text Captioning - Joy Caption Beta 1, Alpha 1& 2, Pre Alpha, Fully Working JSON Prompt Builder, Saved Outputs, 1-Click to Install on Windows, RunPod, Massed Compute, SimplePod, Linux
Our Patreon exclusive posts index list all of the apps we have (over 100+ AI apps, scripts, trainers, presets and more). Use CTRL+F to find whatever app you looking on our index.
Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord
Please also Star, Watch and Fork our Stable Diffusion & Generative AI GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)
=======
Latest installer zip file : Ultimate_Image_Captioner_Pro_v2.zip
[click here to choose a membership and Join to download zip files]
With Ideogram 4 model as you know JSON prompting is now a thing and we need JSON prompts for both inference and training
Therefore, a new app was necessary to solve this issue and I built the very best local app out there for this task
Fully supported model list is as below with fully working robust Torch Compile
Qwen Vision Models
Qwen3-VL 8B Instruct (default)
Huihui Qwen3-VL 8B Instruct Abliterated
Qwen3-VL 4B Instruct
Qwen3-VL 2B Instruct
Qwen3-VL 30B-A3B Instruct
Qwen3.6 27B
Huihui Qwen3.6 27B Abliterated
Joy Caption Models
Joy Caption Beta 1
Joy Caption Alpha 2
Joy Caption Alpha 1
Joy Caption Pre Alpha
Full features of the app introduced below with screenshots so please read
The installer will auto download all the necessary models with 16 connections + SHA256 hash verification
Full tutorial video : https://youtu.be/TW3MRdd0MV4
SwarmUI and ComfyUI zip files updated for Ideogram 4 model, model downloads and presets and workflows already added
SwarmUI : https://www.patreon.com/SECourses/posts/114517862
ComfyUI : https://www.patreon.com/SECourses/posts/105023709
Windows Requirements
Python 3.12.10, FFmpeg, CUDA 13, cuDNN 9.17 or above, Visual Studio Community Edition with all C++ options selected
Don't worry CUDA 13 works with all GPUs - make sure you have updated NVIDIA driver
Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0
Follow its updated post with links and screenshots exactly : https://www.patreon.com/SECourses/posts/requirements-written-tutorial-111553210
For RunPod, SimplePod, Massed Compute and Linux please follow:
Massed_Compute_Instructions_READ.txt
Runpod_SimplePod_Ultimate_Caption_Instructions.txt
The application runs on Torch 2.13 with CUDA 13, supports literally every GPU out there including server GPUs
Moreover, we are using latest libraries that I compiled as below
19 July 2026 V2.1
Application upgraded to newest Torch 2.13 with above seen pre-compiled wheels that works perfect
So either do a fresh install or, extract latest zip file, overwrite installers, delete venv folder inside Ultimate_Image_Captioner_Pro and run installer again for update
The following models are fully supported now
Qwen3-VL 8B Instruct (default), Huihui Qwen3-VL 8B Instruct Abliterated, Qwen3-VL 4B Instruct, Qwen3-VL 2B Instruct, Qwen3-VL 30B-A3B Instruct, Qwen3.6 27B, Huihui Qwen3.6 27B Abliterated
We have implemented fully working torch compile feature that speeds up Qwen Vision Models 84% and works on all Joy Captions as well
Moreover, now we are using Python 3.12.10 so please install Python 3.12.10 if you don't have yet
For Torch Compile to work, you need to have installed Visual Studio Community Edition with All c++ options
C++ tools not needed anymore, only Visual Studio Community Edition
Requirements post is updated for this : https://www.patreon.com/SECourses/posts/requirements-written-tutorial-111553210
New Qwen Image models implemented that fully works
They will be downloaded fully automatically when you first time use them, only default Qwen3-VL 8B Instruct is auto downloaded with the initial installation
Joy Caption Alpha 2, Joy Caption Alpha 1, Joy Caption Pre Alpha models will not be auto downloaded anymore with initial installation
They will be auto downloaded when you first time use them
All Joy Caption outputs are now displayed with text-wrapping and has copy generated prompt button feature
3 July 2026 V1.2
Qwen image captioning made more robust
Such as in some cases it was adding imgur links and not anymore this bug exists
Apply Box Edits button is now Apply Box Edits & Save so every edit is automatically saved in the outputs folder
Overwrites generated json file and re-generates boxed image
The changes you made in JSON Box Preview or JSON Elements were not being saved in outputs folder and now they will be saved
JSON Prompt Builder significantly improved
Now it will accurately recognize selected file's accurate outputs folder path and all changes will be saved
If your file is not in outputs folder, use Browse File to load file
Now lets say you started empty design and saved, it will be saved in a new folder inside outputs folder and keep using that folder as long as you work on that json prompt
When you were switching between folders, it was not properly updating displayed values and this issue fixed
Now when you switch Saved Outputs tab it will auto refresh and show latest
Now Saved Outputs tab is auto sorted by latest but you can re-sort by clicking display headers
To update just run Windows_Install_Update_App.bat file
Zip file is still same
Ultimate Image Captioner Pro Features
Click on images to see them full resolution
1-Click to install on Windows, RunPod, Massed Compute (Linux users please use Massed Compute scripts) and SimplePod
Fully support JSON prompt generation
Full custom user preset save and load, after restart remembers last saved / used one
Fully edit generated json values content or boxes and reconstruct json and easy 1-click to copy
Hide / display boxes to easy work, fully drag to change position or resize and make bigger or smaller from interface
We have got 35 Ideogram 4 presets ready to select and use
Fully working automatically selected GPU VRAM presets for both Qwen and Joy Caption models
You can run as subprocess thus it will leave 0 VRAM or RAM usage after captioning
Moreover, you can set which GPU ID to run captioning on, or set multiple GPUs to distribute batch captioning
Auto save box drawn images
Add suffix, prefix or word replaces to generated captions automatically
All Joy Caption Models are supported with full features like Qwen (e.g. batch captioning, VRAM presets, various preset prompts and save options, etc.)
Fully working JSON prompt builder that you can build from scratch or load existing image
Add boxes, write info, modify boxes, move them, resize them, etc. all fully working
If you load existing image, if that image has JSON file, it will be auto loaded
This way, you can load your previous generations and modify and work further on them
Every generated output is saved in outputs folder with full metadata as well
Fully working view saved outputs screen that you can quickly navigate and find your previous generations and see them with full details, info, etc.
The page has full filtering and pagination features so use them as well
CMD screen shows full details of what is happening even token / second as well
Following CMD outputs is especially useful for batch folder processing
Qwen JSON prompt generator is so amazing that it has a specific field for written text on images, analyze below image to understand logic
Pay attention to JSON Elements table you will see text
You can fully edit the JSON Elements table and re-generate, auto generated JSON as you wish
When you click Apply Box Edits it will overwrite generated JSON file and boxes drawn image in respected output folder
Saved output json files are beautified and saved - can be disabled if you wish