Ultimate Image Captioner Pro - Qwen3 VL 8B Instruct with Full JSON and Text Captioning - Joy Caption Beta 1, Alpha 1& 2, Pre Alpha, Fully Working JSON Prompt Builder, Saved Outputs, 1-Click to Install on Windows, RunPod, Massed Compute, SimplePod, Linux

Next Post
Ultimate Image Captioner Pro - Qwen3 VL 8B Instruct with Full JSON and Text Captioning - Joy Caption Beta 1, Alpha 1& 2, Pre Alpha, Fully Working JSON Prompt Builder, Saved Outputs, 1-Click to Install on Windows, RunPod, Massed Compute, SimplePod, Linux
1 / 15
DESCRIPTION

Our Patreon exclusive posts index list all of the apps we have (over 100+ AI apps, scripts, trainers, presets and more). Use CTRL+F to find whatever app you looking on our index.

Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord

Please also Star, Watch and Fork our Stable Diffusion & Generative AI  GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)

=======

Latest installer zip file : Ultimate_Image_Captioner_Pro_v2.zip

[click here to choose a membership and Join to download zip files]

With Ideogram 4 model as you know JSON prompting is now a thing and we need JSON prompts for both inference and training

Therefore, a new app was necessary to solve this issue and I built the very best local app out there for this task

Fully supported model list is as below with fully working robust Torch Compile

Qwen Vision Models

Qwen3-VL 8B Instruct (default)

Huihui Qwen3-VL 8B Instruct Abliterated

Qwen3-VL 4B Instruct

Qwen3-VL 2B Instruct

Qwen3-VL 30B-A3B Instruct

Qwen3.6 27B

Huihui Qwen3.6 27B Abliterated

Joy Caption Models

Joy Caption Beta 1

Joy Caption Alpha 2

Joy Caption Alpha 1

Joy Caption Pre Alpha

Full features of the app introduced below with screenshots so please read

The installer will auto download all the necessary models with 16 connections + SHA256 hash verification

Full tutorial video : https://youtu.be/TW3MRdd0MV4

SwarmUI and ComfyUI zip files updated for Ideogram 4 model, model downloads and presets and workflows already added

SwarmUI : https://www.patreon.com/SECourses/posts/114517862

ComfyUI : https://www.patreon.com/SECourses/posts/105023709

Windows Requirements

Python 3.12.10, FFmpeg, CUDA 13, cuDNN 9.17 or above, Visual Studio Community Edition with all C++ options selected

Don't worry CUDA 13 works with all GPUs - make sure you have updated NVIDIA driver

Follow this requirements tutorial video exactly : https://youtu.be/DrhUHnYfwC0

Follow its updated post with links and screenshots exactly : https://www.patreon.com/SECourses/posts/requirements-written-tutorial-111553210

For RunPod, SimplePod, Massed Compute and Linux please follow:

Massed_Compute_Instructions_READ.txt

Runpod_SimplePod_Ultimate_Caption_Instructions.txt

The application runs on Torch 2.13 with CUDA 13, supports literally every GPU out there including server GPUs

Moreover, we are using latest libraries that I compiled as below

19 July 2026 V2.1

Application upgraded to newest Torch 2.13 with above seen pre-compiled wheels that works perfect

So either do a fresh install or, extract latest zip file, overwrite installers, delete venv folder inside Ultimate_Image_Captioner_Pro and run installer again for update

The following models are fully supported now

Qwen3-VL 8B Instruct (default), Huihui Qwen3-VL 8B Instruct Abliterated, Qwen3-VL 4B Instruct, Qwen3-VL 2B Instruct, Qwen3-VL 30B-A3B Instruct, Qwen3.6 27B, Huihui Qwen3.6 27B Abliterated

We have implemented fully working torch compile feature that speeds up Qwen Vision Models 84% and works on all Joy Captions as well

Moreover, now we are using Python 3.12.10 so please install Python 3.12.10 if you don't have yet

For Torch Compile to work, you need to have installed Visual Studio Community Edition with All c++ options

C++ tools not needed anymore, only Visual Studio Community Edition

Requirements post is updated for this : https://www.patreon.com/SECourses/posts/requirements-written-tutorial-111553210

New Qwen Image models implemented that fully works

They will be downloaded fully automatically when you first time use them, only default Qwen3-VL 8B Instruct is auto downloaded with the initial installation

Joy Caption Alpha 2, Joy Caption Alpha 1, Joy Caption Pre Alpha models will not be auto downloaded anymore with initial installation

They will be auto downloaded when you first time use them

All Joy Caption outputs are now displayed with text-wrapping and has copy generated prompt button feature

3 July 2026 V1.2

Qwen image captioning made more robust

Such as in some cases it was adding imgur links and not anymore this bug exists

Apply Box Edits button is now Apply Box Edits & Save so every edit is automatically saved in the outputs folder

Overwrites generated json file and re-generates boxed image

The changes you made in JSON Box Preview or JSON Elements were not being saved in outputs folder and now they will be saved

JSON Prompt Builder significantly improved

Now it will accurately recognize selected file's accurate outputs folder path and all changes will be saved

If your file is not in outputs folder, use Browse File to load file

Now lets say you started empty design and saved, it will be saved in a new folder inside outputs folder and keep using that folder as long as you work on that json prompt

When you were switching between folders, it was not properly updating displayed values and this issue fixed

Now when you switch Saved Outputs tab it will auto refresh and show latest

Now Saved Outputs tab is auto sorted by latest but you can re-sort by clicking display headers

To update just run Windows_Install_Update_App.bat file

Zip file is still same

Ultimate Image Captioner Pro Features

Click on images to see them full resolution

1-Click to install on Windows, RunPod, Massed Compute (Linux users please use Massed Compute scripts) and SimplePod

Fully support JSON prompt generation

Full custom user preset save and load, after restart remembers last saved / used one

Fully edit generated json values content or boxes and reconstruct json and easy 1-click to copy

Hide / display boxes to easy work, fully drag to change position or resize and make bigger or smaller from interface

We have got 35 Ideogram 4 presets ready to select and use

Fully working automatically selected GPU VRAM presets for both Qwen and Joy Caption models

You can run as subprocess thus it will leave 0 VRAM or RAM usage after captioning

Moreover, you can set which GPU ID to run captioning on, or set multiple GPUs to distribute batch captioning

Auto save box drawn images

Add suffix, prefix or word replaces to generated captions automatically

All Joy Caption Models are supported with full features like Qwen (e.g. batch captioning, VRAM presets, various preset prompts and save options, etc.)

Fully working JSON prompt builder that you can build from scratch or load existing image

Add boxes, write info, modify boxes, move them, resize them, etc. all fully working

If you load existing image, if that image has JSON file, it will be auto loaded

This way, you can load your previous generations and modify and work further on them

Every generated output is saved in outputs folder with full metadata as well

Fully working view saved outputs screen that you can quickly navigate and find your previous generations and see them with full details, info, etc.

The page has full filtering and pagination features so use them as well

CMD screen shows full details of what is happening even token / second as well

Following CMD outputs is especially useful for batch folder processing

Qwen JSON prompt generator is so amazing that it has a specific field for written text on images, analyze below image to understand logic

Pay attention to JSON Elements table you will see text

You can fully edit the JSON Elements table and re-generate, auto generated JSON as you wish

When you click Apply Box Edits it will overwrite generated JSON file and boxes drawn image in respected output folder

Saved output json files are beautified and saved - can be disabled if you wish

SECourses: FLUX, Tutorials, Guides, Resources, Training, Scripts PATREON 32 favs
VIEWS2
FILES16 files
POSTEDJul 19, 2026
ARCHIVEDJul 3, 2026