Previous Post

ACE-Step 1.5 XL Premium - Better Music & Song Generator Than SUNO 5.0 - Remix and Repaint Features, SAM Audio Processing - Windows, RunPod, Massed Compute, Linux 1-click Installers

Next Post
ACE-Step 1.5 XL Premium - Better Music & Song Generator Than SUNO 5.0 - Remix and Repaint Features, SAM Audio Processing - Windows, RunPod, Massed Compute, Linux 1-click Installers
1 / 142
DESCRIPTION

Our Patreon exclusive posts index list all of the apps we have (over 100+ AI apps, scripts, trainers, presets and more). Use CTRL+F to find whatever app you looking on our index.

Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord

Please also Star, Watch and Fork our Stable Diffusion & Generative AI  GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)

=======

Latest installer zip file : ACE_Step_v7.zip

[click here to choose a membership and Join to download zip files]

Quick Info

This app has the following repos perfectly combined into our premium app with additional improvements and features such as optimized model loading, VRAM, quality, accuracy and performance optimizations, batch folder processing and many more (all models automatically downloaded and everything installed into a Python 3.11 VENV)

We got VRAM presets for every GPUs already set, read changelogs below to learn everything, slowly top to bottom read recommended

ACESTEP XL 1.5 (both inference + training) : https://github.com/Runware/ACE-Step-1.5-XL

https://deepwiki.com/ace-step/ACE-Step-1.5/5-generation-features

SAM-Audio Segment from Facebook / META : https://github.com/facebookresearch/sam-audio

Massive optimizations made for this model , it is working amazing

Auto-Editor : https://github.com/wyattblue/auto-editor

TrackAICleaner Post Processing : https://github.com/mikecastrodemaria/TrackAICleaner

DiffPitcher : https://github.com/haidog-yaqub/DiffPitcher

ACE Step 1.5 XL is the newest State Of The Art (SOTA) Music and Song generator model. It has 3 variants and we support all 3 variants (Turbo, SFT, Base) with fully automatic setup, models download, VRAM presets for all GPUs starting from 4 GB and with all best researched generation values / settings / configurations.

Windows Requirements

Python 3.12.10, FFmpeg, CUDA 13 cuDNN 9.17, Git, Visual Studio Community Edition with Desktop Development with C++ (all checkboxes checked)

Make sure that your NVIDIA driver is updated (min 590+)

If you get any errors follow below video and its source link

https://youtu.be/DrhUHnYfwC0

https://www.patreon.com/posts/windows-requirements-tutorial-written-post

The zip file contains installers for

Windows : Windows_Install_or_Update.bat

Please follow requirements video for Windows before starting installation : https://youtu.be/DrhUHnYfwC0

Requirements tutorial is 1 time mandatory for all of my applications

Windows installer will download only ACEStep 1.5 XL Turbo model

To download all models also run Windows_Download_All_Models.bat after installation

RunPod and SimplePod : Runpod_SimplePod_ACE_Step_Instructions.txt

Massed Compute / Local Linux : Massed_Compute_Instructions_READ.txt

RunPod, Massed Compute installers automatically downloads all 3 ACEStep 1.5 XL models, Turbo, SFT and Base

Zip file also has ACE_Step_Lyric_Generation_Instructions_For_LLMs.txt which you can use to better format your Music / Song lyrics and style by providing this file to your favorite LLM

The installers will generate a Python 3.12 VENV automatically and install everything inside there, thus your system or any other of your APPs will never be impacted

With our pre-compiled abi3 wheels for Windows and Linux, you can run on Python 3.10, 3.11, 3.12 and 3.13 but preferred version is 3.12

The following libraries compiled for Windows and Linux

archs : mslk, xformers, flash_attn, sageattention, torchao,

All of them are abi3 and Windows versions are compiled for Consumer GPUs, Linux versions compiled for Consumer + Cloud GPUs - so we cover all GPUs you have

12 July 2026 V7.0.1 Update

You need to delete your venv folder and run this update - existing installation may remain

I have started upgrading all of our apps into latest Torch 2.13 and CUDA 13

For this, I have pre-compiled the following wheels with all CUDA 13 features and with all CUDA archs : mslk, xformers, flash_attn, sageattention, torchao

All these libraries are properly compiled with abi3 thus works on Python 3.10, 3.11, 3.12 and 3.13

So I have updated the SECourses ACE-STEP XL 1.5 installers with newest libraries

It will automatically generate Python 3.12 venv and install everything there, if you don't have 3.12, it will use your default python

Make sure to have 3.12.10 it is best

Make sure to have nodejs 22+ installed - 22 preferred and tested

Follow requirements tutorial ( https://youtu.be/DrhUHnYfwC0 ) and latest ACE-Step tutorial : https://youtu.be/hzKSt5WUAm0

At the LoRA training tab, we had missed supported languages (50)

Now you can train with all Languages the model supports for inference

Arabic, Azerbaijani, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Persian, Finnish, French, Hebrew, Hindi, Croatian, Haitian Creole, Hungarian, Indonesian, Icelandic, Italian, Japanese, Korean, Latin, Lithuanian, Malay, Nepali, Dutch, Norwegian, Punjabi, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Serbian, Swedish, Swahili, Tamil, Telugu, Thai, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Cantonese, Chinese

We have massively improved Torch Compile with newest Torch 2.13 and our massive backend improvements

Now Torch Compile happens with 8-cpu threads by default

Moreover, there were unncessarily compiled parts and they are not compiled anymore

Moreover, we have updated some libraries and fixed some bugs

The result is that, previously Torch Compile was taking 240 seconds to compile, now taking like 10 seconds

Once Torch Compile is warm, it is taking around 20-30 seconds to generate full songs on RTX 5090 - click to result

A wider audit of the app has been made and following bugs fixed

Presets lost multi-select and explicit empty CheckboxGroup values.

Negative prompts did not reliably synchronize both ways.

GPU tiers did not synchronize from Advanced back to Generate Song.

Initialize Service returned fewer values than its event wiring expected.

Batch Folder ignored Generate Song duration and seed settings.

Grid Testing passed shifted positional LoRA arguments.

Saving a dataset could erase per-sample instrumental labels.

Stopped LoKr training incorrectly reported successful completion.

For updating, get the latest zip file, overwrite older files, delete venv folder inside ACE-Step_Premium and run installer

5 July 2026 V6.3.4 Update

Negative prompt disabled for Turbo models since they are CFG 1

Negative prompt synching with advanced tab - simple tab issue fixed and now works smooth

Just run install update bat file to update

29 June 2026 V6.3.3 Update

ACESTEP XL 1.5 VRAM presets Tier 3, 4, 5 and 6a significantly improved

When first time model is quantized into Int8 or FP8, it was taking too much time and this issue fixed now like 100x faster

Moreover, after first time caching, the cache files will be saved inside ACE-Step_Premium\.cache and reused next time so next time generation will be instant even after restarting the app

Just run installer / update bat file to upgrade

25 June 2026 V6.3.2 Update

How Remix Source Start and Remix Source End works upgraded

Now it regenerates entire remix and then only merges the selected part into remixed original song

With this approach, you can iteratively remix certain parts quickly and get the ultimate best song

Just run installer / update bat file to upgrade

23 June 2026 V6.3.1 Update

Same V6 zip file just run installer to update

We fixed an important generation quality regression

The issue was caused by the new Remix Melody Retention default value leaking into normal Text-to-Music generation

This made XL Turbo start diffusion almost at the end of the schedule, so an 8-step generation effectively ran only about 1 real DiT step.

That explained why some users saw extremely fast generations, for example around 20 seconds instead of the expected longer runtime, with poor quality results.

We also fixed Turbo inference step handling. Previously, Turbo requests above 8 steps were still being clamped back to 8.

Now Turbo supports up to 20 steps correctly, so setting 16 steps actually runs 16 steps.

Sampler mode added to the Generate Song tab and set heun as default since almost no speed difference and it is better than previous euler

Still you can test and compare both if you wish

I did a lot of testing with Remix presets and sadly it is hard to make best for every case so you better test for your own cases

Now there are 4 presets, Default, Same Lyrics Big Change, Same Lyrics Medium Change and Different Lyrics

Enable torch compile and generate subsequently until you get a good result really fast

Negative prompt feature implemented to every field

Works only with SFT and Base model since Turbo model has CFG 1

New button Use Generated Result as Source added so that you can quickly set generated song to remix, edit, whatever you want to do

Useful for iterative processing

Advanced section of ACESTEP XL 1.5 app completely revamped and made much better as below

21 June 2026 V6.0 Update

Gangam Style in English : https://x.com/SECourses/status/2068512611725975733

This is a massive improvements and fixes update

We have moved to the Gradio 6.19 and thus transformers library upgraded to 5.3

So i had to fix pipeline for newest transformers library

PyTorch 5Hz LM generation now uses modern Transformers forward-pass features such as cache_position and logits_to_keep when available.

Gradio interface event handling was heavily improved for Gradio 6.19.

Many UI sync events now run without queue overhead and without unnecessary progress overlays.

Stale Gradio status timers are automatically hidden, fixing stuck timer/progress artifacts.

Long-running actions now show progress only on relevant outputs instead of slowing down unrelated UI elements.

Library metadata display was changed from heavy JSON rendering to copy-friendly text, making the library page smoother.

Batch processing, batch extract, audio processing, generation, library, LoRA, LoKR, dataset, and SAM Audio UI wiring were updated for smoother behavior.

Model loading architecture updated for newer transformers library and now model loading faster

Due to newer transformers library, now torch compile is even faster than before

PyTorch LM loading now tries the faster SDPA attention path on CUDA.

Audio-code generation now has a compact valid-token sampling path, avoiding unnecessary full-vocabulary processing during constrained generation - No quality loss

VAE tiled decode was optimized by preventing pathological tiny-stride chunking - Nno quality loss

With VAE optimization + transformers library, now torch compile is able to generate full song in 30 seconds on RTX 5090

GPU VRAM presets are re-tested and updated as below

Wildcards special character was [] and now it is fixed and changed into {}

CoT Language Detection and Caption / Style Auto Improve was enabled in some presets and now they are all disabled - so you have to manually enable

They were causing unexpected issues and problems

If CoT language is enabled, the LM-detected language is used only when vocal language is set to auto/unknown.

Explicit user-selected vocal language is preserved and no longer unexpectedly overwritten.

In advanced tab now you can set explicit vocal language and this fixed so many issues

Advanced tab set Vocal Language will update Generate Song tab Vocal Language as well or vice-versa

Remix presets implemented and literally 1-click first test result you can see here

Gangam Style in English : https://x.com/SECourses/status/2068512611725975733

Hopefully will make a mini tutorial video so open bell on Youtube : https://www.youtube.com/SECourses

Batch audio processing now reports status immediately when scanning starts.

Batch Extract now normalizes itself to Extract mode internally, instead of requiring the user to manually switch generation mode first.

Batch progress display was improved.

Batch queue restore defaults now keep CoT caption/language disabled unless explicitly enabled.

For updating please get latest v6 zip file, overwrite previous files and run installer bat file

If you get any errors, please delete ACE-Step_Premium\venv and then run installer again

19 June 2026 V5.5 Update

Full tutorial video published finally for inference : https://youtu.be/9C_6qNKjgpA

I started working on LoRA training tutorial as well hopefully soon

With 5.5 optimizer specific parameters are now shown that you can set, I am also working on to make them auto default best hopefully

There was a visual bug that hidden Remix Melody Retention and Direct Source Latents (no_fsq) on Remix songs page and this bug fixed and app scanned entirely and all visuals verified

Default value set to 0.97 one of our expert remixer recommended that

Just run Windows_Install_or_Update.bat to update, the zip file not changed

18 June 2026 V5.4 Update

Now batch folder processing for ACESTEP XL 1.5 and SAM Audio has this extra option Save only output

This is useful to get only processed files and no other stuff like remaining part of the songs or metadata files, etc.

18 June 2026 V5.3 Update

Wildcard feature implemented

It works both for Style / Captions and Lyrics with syntax verification as well

It will work in batch folder processing as well so you can write that way in txt files

If you enable Auto improve lyrics or Auto improve style they may break your syntax so don't enable when using wildcards

Just run Windows_Install_or_Update.bat to update same zip file still

Also full inference tutorial published that covers every topic in details including how to install on Windows, RunPod, Massed Compute and SimplePod : https://youtu.be/9C_6qNKjgpA

16 June 2026 V5.2 Update

Default Remix value is now 0.95 instead of 1

Seed box and Random seed option moved to a much easier to use place

Last generation seed value will be auto set in seedbox so you can uncheck random seed and keep working with same seed now easier

14 June 2026 V5.1 Update

Use Repaint with lyrics added to the Repaint tab of ACESTEP XL 1.5

14 June 2026 V5.0.0 Update

I am still working on inference tutorial and as I used and as you made new feautre requests new features arrived

Auto-Editor trim output was not working properly and this bug fixed now should work much better when you use it in SAM Audio Segment or ACESTEP XL 1.5 Extract

This is really useful to get only vocals and trim empty / no vocal parts for training

Extract All stems feature implemented to ACESTEP XL 1.5 extract tab

Extracted stems will be saved in same folder with suffixes like brass, guitar, vocal, etc.

I noticed that extracting stems much better working on full songs rather than part of songs like 1 minute split part for some reason for ACESTEP XL 1.5 extract

Extract logic improved

Each different extract may yield different results so you can try multiple times to get better extract

Auto-Editor workflow export significantly improved

In Audi Processing tab enable Auto-Editor trim silent sections

Then Set Processing Preset = None

Then select your Auto-Editor workflow export like DaVinci Resolve

Then use Local Audio/Video Path with Browse File button or direct path

This way you will get almost instantly .fcpxml with accurate file path or whatever supported format you pick

SAM Audio Segmet now supports Batch Segment

You can use Batch Segment with 2 ways

First way is enable Batch Segment checkbox and type your stems / segments into Custom Prompt with ; seperation

Second way is select multiple Quick Prompt from dropdown and it will segment / extract every one of them

Custom prompt section overwrites Quick Prompt selections

Extracted stems / segments will be saved in same folder with suffixes like brass, guitar, vocal, etc.

Load Metadata feature implemented as a new tab

Select the generation_manifest.json and it will load every single configuration / parameter of that generation

Get the latest zip file, overwrite older files and run Windows_Install_or_Update.bat file for update or fresh install

To have all models (ACESTEP XL 1.5 Base and ACESTEP XL 1.5 SFT) run Windows_Download_All_Models.bat after installation

12 June 2026 V4.9.5 Update

In Audio Processing tab now there is None Processing Preset which unchecks all Audio Enhancement

Now there is Disable upload preview checkbox in Audio Processing tab

Use for very large videos or containers like multi-GB MKV files. When enabled, Gradio will not render the uploaded media preview, avoiding slow browser/Gradio post-processing such as MKV-to-MP4 preview conversion. Processing still uses the original uploaded file.

Gradio does post processing to every video file if not mp4 therefore other formats will take massive time to display if they are big : https://github.com/gradio-app/gradio/issues/13527

11 June 2026 V4.9.4 Update

Auto-Editor trim silent parts descriptions updated

Apply automatically to generated songs was mistakenly enabled by default and this issue fixed so you can enable if you wish

Analyze button won't overwrite your lyrics anymore

Auto-Editor trim silent parts feature in SAM Audio Segment and ACESTEP XL 1.5 Extract will now use the settings / parameters set in Audio Processing tab

In ACESTEP XL 1.5 extract mode when Auto-Editor trim was selected, it was not working accurately and now will work a bug fixed

Latest generated results sections labels fixed - for ACESTEP XL 1.5 advanced tab

For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat

11 June 2026 V4.9.3 Update

In the repaint task, if generated song is shorter than the selected repaint area, it will trim thus you won't have silent parts

The Repaint Strength description updated and fixed : When lyrics are provided, Repaint switches to text-to-music, so Repaint Strength has no impact when changing lyrics. To keep the same vocal audio, LoRA training and using that LoRA are mandatory.

Now output format can be selected in ACESTEP XL 1.5 modes

Default is set as mp3 since generated files were taking too much space

Now all generated files will obey the selected format e.g. like below

Lego mode was not working accurately and this issue fixed

Now in Lego mode, you will see only generated output as well such as you selected guitar so you will get the generated guitar song as well like below

For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat

10 June 2026 V4.9 Update

V4.9 is a pretty big and important update lots of fixes and improvements

Generation modes now explicity shows recommended models for ACESTEP XL 1.5

Previously, switching models without restarting the app was causing VRAM leak and OOM

This issue is fixed and now you can generate with Turbo model and then switch SFT or Base, and so on

To be 100% sure not have any RAM or VRAM leak, enable Use isolated subprocess generation checkbox

This option will slow subsequent generations and not mandatory, so enable if you are sure and needed

For Remix, Repaint, Lego and Complete, now you can set Instrument Start and End of source input and it will show live preview, really useful for Repaint

Instrument Start and End selection was not working accurately for Remix, Repaint, Lego and Complete but this bug fixed so now you can repaint just specific part of the model

Repaint was not using accurate methodologies and automatic inner prompt to repaint song accurately and now this issue also fixed

So now you can change specific part of the song and make it sing different vocal / lyrics etc perfectly working tested

Remix, Lego, Repaint and Complete mode errors fixed and they are made more robust

Optional Parameters, Batch Process, Settings will be closed by default now, so easier to read interface

Click them to open them again

Cluttering unrelated some information from Remix, Lego, Repaint and Complete modes removed such as Custom Guide from Remix

Generated results now will show followings

Latest Generated Result (Sample 1) : Is the full new repainted, remixed, etc song

Next to it Original Input, the original song for quickly listen both and compare

Latest Repainted Area, is the area of the song you repainted like between 30-40 seconds, this works for other modes too so you can listen only that particular section

Next to it, Latest Repainted Area Original, the original part of the song that was repainted, etc. to see before after quickly

For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat

10 June 2026 V4.8 Update

Torch compile feature implemented for ACESTEP 1.5 XL and SAM Audio processing

For ACESTEP XL 1.5, switch to advanced tab and enable, then you can switch back to Generate Song tab

ACESTEP XL 1.5 training also supports torch compile but not tested and verified yet

The initial torch compile may take some time but after that, repeated usage brings massive performance boost as shown as below

It won't recompile once compiled even if app is restarted, so it uses compile cache, if necessary it will recompile though

Initial compile may take time and may look like frozen but both inference and training tested and working

You have to have accurately setup CUDA, MSVC and C++ Tools for this to work since Torch compile depends on it

Therefore, follow requirements tutorial fully properly : https://youtu.be/DrhUHnYfwC0

The system is very robustly designed to automatically find accurate CUDA and C++ tools installation even if you have multiple installations

LoRA training speed with Torch Compile is 0.98 it / second and without Torch Compile is 0.78 it / second

25% faster

Use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat to update

ACESTEP XL 1.5 Inference Torch Compile

SAM Audio Inference Torch Compile

ACESTEP 1.5 XL LoRA Training Torch Compile

6 June 2026 V4.7.1 Update

Auto-Editor executable download now has alternative source if GitHub fails - now more robust

New feature DiffPitcher Pitch Fixer implemented into Audio Processing tab since requested

You can read more about it here : https://github.com/haidog-yaqub/DiffPitcher

The installer bat file will download necessary diffusion models automatically as safetensors files

Get latest zip file (4_7), overwrite previous files and run Windows_Install_or_Update.bat to update

6 June 2026 V4.6 Update

Audio processing tab significantly improved a lots of new features added

Now supports Run as subprocess and cancel button immediately

Now fully supports video inputs

Now supports Export Only Audio - very useful for getting audio from video if you don't need video

Now avoids reencoding of videos only if audio of video is processed - Auto-Editor triggers video processing

Now supports video re-encoding profiles

Now supports Auto-Editor workflow export for Davinci Resolve, Adobe Premiere Pro, Final Cut Pro, Shotcut and Kdenlive

Thus you can trim silent parts of your videos and continue editing in your favorite app, I use this literally to edit my tutorial videos

Now fully shows Audio Processing tab process progress in CMD and also on Gradio

Auto-Editor video processing may take quite time since it re-encodes video

Hopefully will make new tutorial soon

Zip file is same, just use Windows_Install_or_Update.bat to update

4 June 2026 V4.5 Update

SAM Audio model loading speed significantly improved like 2.5x faster than before

Unchecking Subprocess mode in SAM Audio was not working now works

So if you uncheck, after processing, it will keep model in VRAM thus instantly starts processing next task - in batch mode it doesn't unload model even if it is checked until batch process ends

New feature Predict spans added to the SAM Audio

Uses SAM-Audio's span predictor to estimate target time ranges from the text prompt when you did not provide anchors

This can improve quality of results depending on source file and the task so you can compare and see if improves

This can use slightly more VRAM and slightly slower

Advanced tab renamed into ACESTEP Advanced

Interface of following sections Custom, Remix, Repaint, Extract,Lego, Complete improved which are located in ACESTEP Advanced tab

Descriptions and buggy features of each section updated and improved as below:

Custom: Manual mode for precise control over caption, lyrics, BPM, key, duration, sampler settings, and advanced generation parameters. Use it when you already have a clear target and want to tune the result yourself. Switch to Generate Song main tab when you want to describe the idea in plain language and let AI fill in the details.

What it does: generates new music from your manual Caption, Lyrics, BPM, key, duration, and advanced settings.

How to use it: describe the target style and vocal delivery in Caption, write structured Lyrics with tags such as [Verse] and [Chorus], then set metadata only when you need tighter control. Leave Think on when you want the LM to plan; turn Think off only when using pasted LM Codes Hints.

Audio inputs: Reference Audio can guide timbre, mix, performance feel, and atmosphere, but it will not copy exact melody, rhythm, or lyrics. Source Audio is ignored in normal Custom generation and is only used by the Edit morph workflow.

Remix: Upload source audio and restyle it with your own caption and lyrics. The AI uses the original as a structural guide while applying your new style. Adjust Remix Strength to control how closely it follows the original (high = faithful cover, low = loose reinterpretation).

What it does: uses Source Audio as the structural guide for melody, rhythm, chords, arrangement, and timing while applying your new Caption and Lyrics.

How to use it: upload Source Audio, optionally trim it in Source Audio Preview, write the target style in Caption, provide replacement Lyrics if you want changed vocals, then adjust Remix Strength and Remix Melody Retention. Higher Remix Strength follows the source more closely; lower strength gives the model more room to reinterpret.

Audio inputs: Source Audio is the important input here. Reference Audio is only an extra global style cue. If the source is instrumental-only, Remix can follow the instrumental structure but still has to invent the vocal melody and phrasing for new lyrics.

Repaint: Upload Source Audio, choose a start/end range, and regenerate only that range. Caption/Lyrics describe the replacement section. Optional Reference Audio can guide style/timbre, but it is not the audio being edited.

Repaint: regenerate one time range of the source

What it does: keeps the Source Audio context and redraws only the selected start/end range. Use it to fix a bad section, replace a lyric phrase, change a solo, or smooth a transition without regenerating the whole song.

How to use it: upload Source Audio, set Repainting Start and End in seconds, then write Caption and Lyrics for the replacement section only. Use Conservative to protect boundaries, Balanced for normal edits, or Aggressive when the selected range should change more freely.

Audio inputs: Source Audio is the audio being edited. Reference Audio can nudge style/timbre for the replacement, but it is not the editable source and will not force exact melody or lyric timing.

Extract: Isolate a single track (vocals, drums, bass, etc.) from source audio using AI stem separation. Useful for creating instrumentals, acapellas, or isolating parts for remixing. Available on Base only.

Extract: isolate one stem from source audio

What it does: separates one selected track from Source Audio, such as vocals, drums, bass, guitar, keyboard, or other supported categories.

How to use it: upload Source Audio, choose Track Name, select the Extract output format if needed, then click Extract Stem. Use the extracted stem for acapellas, instrumentals, remix prep, cleanup, or analysis.

Audio inputs: Extract uses Source Audio only. Caption, Lyrics, Reference Audio, Think, BPM, and key are not creative controls for this mode.

Lego: Choose a predefined instrument category such as synth, bass, drums, or guitar. The AI generates that instrument and adds it over the existing source audio; you do not upload external stems. Upload the source track for context, trim it in Source Audio Preview if needed, choose the instrument to add, and describe only that new layer. Available on Base and SFT; Base generally gives the best results.

Lego: add one generated track over existing audio

What it does: creates the selected instrument category and layers it over Source Audio. This is for adding a new AI-generated part, not for uploading your own external stem.

How to use it: upload Source Audio, choose Track Name such as vocals, backing_vocals, drums, bass, guitar, or synth, optionally set the start/end range, then describe only the new layer in Caption. For vocals, provide Lyrics and describe the singer/delivery.

Audio inputs: Source Audio gives musical context for the new layer. Reference Audio can nudge global sound, but it will not act as a guide vocal. If you add vocals to an instrumental, the model must invent the sung melody and phrasing unless the source already contains that vocal structure.

Complete: Fill in selected missing tracks from source audio. Upload a partial arrangement or single stem, trim it in Source Audio Preview if needed, choose the tracks to add, and optionally set Complete Start/End to regenerate only that section while preserving the rest of the source. Available on Base and SFT; Base generally gives the best results.

Complete: fill missing tracks in a partial arrangement

What it does: listens to Source Audio and generates the selected missing track classes so the partial idea becomes a fuller arrangement.

How to use it: upload a partial track, single stem, or incomplete mix, choose the track classes to add, optionally set Complete Start and End to limit the generated section, then describe the desired finished arrangement in Caption. Use it for adding accompaniment around vocals, drums/bass under a sketch, or missing instruments in a section.

Audio inputs: Source Audio is the context that the new tracks must fit. Reference Audio can guide overall style, but it does not replace the source and does not force exact melodic or lyric timing.

Get latest zip file (4_3), overwrite previous files and run Windows_Install_or_Update.bat to update

4 June 2026 V4.4 Update

SAM Audio processing bug fixed

In ACESTEP XL 1.5 Advanced Extract tab, Analyze button was useless now it will show info message to use Track Name and click Extract Stem

Extract Stem will now show progress and status in Latest Result Status

Extract Stem limited to ACESTEP XL Base model since it works 100x better with Base than SFT

Now all advanced tab audio / video input fields will show preview

If preview doesn't show immediately, click X and re-select file this fixes Gradio bug

Now you can trim audio from Gradio preview as well

Audio previews visuality improved and trim feature visuality improved significantly for all upload audio fields and previews

Gradio version upgraded to 6.16.0

ACESTEP XL 1.5 Advanced Mode Lego and Complete features improved and bugs fixed

Description of how Lego mode works improved

Lego: Choose a predefined instrument category such as synth, bass, drums, or guitar. The AI generates that instrument and adds it over the existing source audio; you do not upload external stems. Upload the source track for context, trim it in Source Audio Preview if needed, choose the instrument to add, and describe only that new layer.

Get latest zip file (4_3), overwrite previous files and run Windows_Install_or_Update.bat to update

3 June 2026 V4.1 Update

V4.2 is a massive update so please carefully read all

For update please get latest v4_1 zip file, extract into install folder, overwrite and run Windows_Install_or_Update.bat file

If you get any errors for any reason, delete \ACE-Step_Premium\venv and then run installer bat file

When you select ACESTEP XL 1.5 SFT or Base model, in advaced tab, now all these options will be enabled and fully work

Simple, Custom, Remix, Repaint, Extract, Lego, Complete

Extract now fully works and you can pick what to extract from Track Name below

However I think new SAM Audio model is better still test and compare both

You can also use batch extract feature now if you want to batch process a folder of songs

Audio Processing tab improved and now we support extremely famous Auto-Editor

Auto-Editor is amazing library to trim silent - no spoken parts

I use this to trim out videos and can be very useful to trim vocal extraction

I use this to also cut silent parts of my tutorials, very useful to pre-process before editing

New tab SAM Audio Segment implemented

SAM Audio is state of the art prompt and mask based audio processing / seperation model from Facebook : https://ai.meta.com/research/samaudio/

It supports any custom text prompt and the below quick select presets

I have made massive amount of optimizations and programming to implement this model

BF16 pre-converted safetensors SAM-Audio and SAM-Audio Judge models will be automatically downloaded when you run installer or model downloader bat file

Normally released models were FP32 .pt models but I converted them to BF16 and safetensors format

SAM Audio models official pipeline was also first loading into RAM and then moving into GPU thus using extra RAM and slower

I made it directly to be loaded into GPU as BF16

Our implementation supports full sub-process running and auto trim feature - extremely useful to extract vocals for ACESTEP XL 1.5 LoRA vocal training

We already have VRAM presets for every GPU out there for SAM Audio model

It works with 20 seconds segmentation with 5 seconds overlap

20 seconds segmentation is what model authors recommend and used for training

Longer segmentation not improving quality but increases VRAM usage and reduces processing time

Only missing feature is Multi-diffusion text-only mode since authors didn't publish this but I opened an issue and expecting them to publish hopefully

We already have that mode coded by CODEX but I think it is not better due to our inaccurate implementation

We support batch folder processing to pre-process training songs as well or for any reason you want

Flash Attention were not working on Windows RTX 4000 series GPUs and this issue fixed

I have re-compiled Flash Attention 2.8.3 to fully support RTX 3000, 4000 and 5000 series GPUs with extra CUDA Arch a flag for SM120a

Linux Flash Attention with all GPUs (to include Cloud server GPUs too) SMs also recompiled and now will be used : 80;86;89;90;100;103;120

So no GPU should get any error with Flash Attention anymore

More information regarding CUDA archs : https://www.patreon.com/posts/159064759

ACESTEP XL 1.5 Advanced tab now supports uploading video files as well

They will be automatically converted into audio and used

If your video upload shows processing forever, click X icon and reupload

This is a Gradio bug I am trying to fix, refresh page also fixes

Audio Processing tab supports both Audio and Video uploading

SAM Audio Segment supports both Audio and Video uploading

Video upload previews are now capped to height 400px so they won't take entire web page space and look much better

31 May 2026 V3.9.1 Update

New full audio post-processing tab implemented to our premium app from TrackAICleaner repo

You can use this tab to both post-process your existing audio files as batch or as single file or automatically post process your generated songs

When it is enabled to auto post-process generated songs, it will save both original and post-processed songs in the outputs folder

You can use preview button to generate 60 second preview and compare quickly the effect impact

Use latest newer zip file, overwrite and run installer to update

28 May 2026 V3.9 Update

New feature LM Audio Codes added and enabled for all default presets

This is supposed to improve quality in all generations without any loss or VRAM increase

Updates made to fix below error that some users reported

Sadly I couldn't reproduce it yet to verify

Error: Generation produced NaN or Inf latents (shape=[1, 8261, 64], dtype=torch.bfloat16, device=cuda:0, nan=528704, inf=0).

Same zip file just run installer to update

26 May 2026 V3.8 Update

Version 3.8 is a very major upgrade for training

In Advanced tab when you click Analyze button now it will auto initialize model and won't throw error

Custom Preset System save and load issues fixed for some cases

Now we support DoRA for both training and inference (song generation)

DoRA is like LoRA but better quality for training close to full Fine Tuning of the entire model

Moreover now we have Target MLP feature for training

Also applies LoRA/DoRA to decoder MLP layers (gate_proj, up_proj, down_proj). This increases trainable capacity and VRAM use; leave off for the legacy attention-only path.

When MLP enabled, more parameters are trained thus it may be a little bit slower and may require more VRAM but it should improve quality I am still in research

Training Parameters screen significantly improved with lots of new features

Now we have Save best feature

It will save best loss having checkpoint and as new best loss having checkpoint reached, it will overwrite previous best

You can set Best smoothing window, Best min delta and Start saving best after epoch as you wish to make it as you wish

Now we support following Optimizers : adamw, adamw8bit, adafactor

Now we support following Schedulers : cosine, cosine_restarts, linear, constant, constant_with_warmup

I usually prefer constant, I am still in research of best hyper parameters and accurate way of training

Now we support following Timestep modes : continuous, discrete

Now we support Adaptive timestep

Now we have an amazing new feature to measure quality of training progress : Validation split %

Lets say you have 100 training songs, and you set 5%, so it will set 5 songs as validation and use 95 songs for training

Therefore, you can see actual quality improvement or degrade of the model therotically

You can see in below screenshot that as the training continued, the validation loss rate got worse and worse even though traidional loss got lower and lower

This is because the model got completely overtrained and cooked and memorized and lost its full generalization

With the newest features, training VRAM presets are also got updated as below

Auto LRC feature improved and now uses lesser VRAM moreover generated .lrc and .vtt files are now automatically saved inside generated output folder

Zip file is same just use Windows_Install_or_Update.bat to update

25 May 2026 V3.6 Update

Brings some massive improvements for training and LoRA usage

In Generate Song tab now Latest Song section will display status

Dataset tab completey remade and now when you load your generated dataset.json file, it will show caption statistics as word n-grams

You can select word-ngrams and it will list songs containing them

Then you can select songs to see their full details

Useful to see your dataset composition

Now there is Use Only Custom Trigger checkbox which will make all Caption / Style data to be only as your custom trigger

I am gonna test this approach to see if works better on training hopefully

We have a full new Grid Testing tab which lets you to generate songs with selected LoRAs to compare them very easy and quickly

So that you can compare your LoRA training checkpoints properly

You can select multiple LoRAs, any LoRAs you want with filter LoRAs extra feature to make selection easier

Grid results will have special naming and save to make them easier to compare e.g. like below

Just run installer to update

25 May 2026 V3.5 Update

When auto labeling, even though process finished it was still showing as processing at Raw Lyrics (from .txt file) and this bug fixed

The training speed display fixed and now it shows training speed accurately after first 20 steps completed like below

Auto labeling is taking massive time, therefore I have implemented true batch size to this process

User / Custom preset system now will save and load every field exists in training tab properly

Delete preset button now deletes the selected preset and loads the next custom preset

If all user custom presets get deleted, it will load default vram preset

Cancel auto labelling process and tensor generation process implemented

Now you can immediately cancel both processes when running in isolated subprocess

Now your lyric txt files can also contain style / caption

e.g. a.mp3 and a.txt in same folder

Format is like below

# Caption

Your caption

# Lyrics

Your lyrics

I tried to make it as much as possible robust so it should work fairly well see below example

Custom Trigger Tag was not working properly and this issue fixed

New feature Debug: save text prompts added so that you can see what is exactly used to generate Preprocessed tensor files which are actually used for training

Just run installer to update

24 May 2026 V3.0 Update

V3 is an important update for LoRA training

Some LoRA training bugs fixed and the training made more smoother

I also have compiled Flash Attention 2.8.4 latest version for Windows and Linux with massive GPU support for Torch 2.11 and CUDA 13

The reason is that we have upgraded our installer to Torch 2.11, CUDA 13 and Torchao 0.17.0 thus training is now even faster

This Flash Attention compile costed me like 50$ on RunPod for Linux and over 14 hours on Windows you can read more info here : https://www.patreon.com/posts/159064759

Download latest zip file, overwrite older files, delete ACE-Step_Premium\venv and run Windows_Install_or_Update.bat again to update or install

20 May 2026 V2.1 Update

Please read V2.0 update first

Browse Dataset JSON bug fixed which is needed to Preprocess files and generate training tensor files

LoRA refresh added to quick generation panel

In train LoRA tab we have improved the parameters you can set and default parameters are also improved

Now when you change training base model, it will auto update training parameters to best for each model specifics

Now Shift and Training Timestep Steps working accurately and set for each model : Base, SFT, Turbo

Now Resume Training State directly takes saved state file, state files are now directly saved read below to see

Unnecessary export LoRA and custom samples directory removed

How training generated files saved completely revamped and improved

Everything will be saved inside target folder with your training name like below

Much more organized and clean and ready to use after training

safetensors files are LoRA files ready to use and pt files are state files which you can use to continue training

Get latest zip file and run Windows_Install_or_Update.bat to update

Torchao upgraded to 0.16.0 for training

Zip file changes only when needed

20 May 2026 V2.0 Update

This is a massive upgrade and we have added so many new features so read carefully all

The deault models were all 32-bit however we were generating songs in BF16 or FP8 / Int8

Thus, the models were keeping double size on disk for no reason and taking more RAM and duration to load

Therefore, I have generated BF16 models and updated the app and model downloader

Thus, you can make a complete fresh install or, delete \ACE-Step_Premium\models folder and run Windows_Install_or_Update.bat / Windows_Download_All_Models.bat again to download new models

Make sure to get latest zip file and overwrite previous installer files

New all 3 models folder takes 41.9 GB, previously it was 69.8 GB

All default generation presets for all 3-models updated - the parameters are now more accurate

unlimited (>24GB) , tier6b (20-24GB), tier6a (16-20GB), tier5 (12-16GB), tier4 (8-12GB), tier3 (6-8GB), tier2 (4-6GB), tier1 (≤4GB)

ACEStep XL 1.5 Turbo, ACEStep XL 1.5 SFT, ACEStep XL 1.5 Base

Thus, you may expect better quality generation on ACEStep Turbo, SFT and Base models

LoRA selection added to Generate song tab as well since we now fully support LoRA training

Also before starting anything with LoRA training, select your Model from here

If you gonna train ACEStep 1.5 XL SFT, first select it from this screen to load all best SFT parameters then continue this is important

Use isolated subprocess generation was not working properly and this is fixed

Cancel Generation button added to both Generate Song and Advanced tab

For this button to work, Use isolated subprocess has to be enabled

🎓 LoRA Training tab completely remade

Now you can browse Dataset JSON Path and directly load

Now you can browse Browse Audio Folder and directly load

Dataset generation settings now fully working

Format Lyrics (LM) is not recommended

Transcribe Lyrics (LM) is recommended

If there are existing lyric files, Transcribe Lyrics (LM) will keep lyrics as it is and fill other data like Duration, Label, BPM, Key, Caption

Format is, audio_file_name.txt in same folder as audio

Also select your Dataset Model according to the model you gonna train and Dataset VRAM preset - it will be also auto set according to your GPU during initial start

If you get Out of Memory Error (OOM), move to 1 below VRAM Preset

Available auto-label and preprocess presets : 24 GB+ - quality, 12-16 GB, 10 GB+

Auto-Label All now fully working all bugs and errors fixed

Now use Browse Label Folder to pick auto labels saved folder

It will save labels after each labelling done

With using same folder, it will continue wherever left

Now you can navigate between each data item and manually change / fix and save

Selecting song from above listing will also update this Preview & Edit selected song

After auto labelling done, save your dataset into any desired location with any name

Now this auto labelling is fully VRAM optimized and auto unloads after completed and 100% free up VRAM and RAM

After dataset json saved, load dataset json and preprocess and generate training pt files into desired target folder

These pt tensor files will be used to train the model it contains every information needed to train

Now this preprocess is fully VRAM optimized and auto unloads after completed and 100% free up VRAM and RAM

Once you are done with Dataset builder move to Train LoRA tab

Train LoKr not tested yet but Train LoRA fully working and fully tested and optimized

Select folder of preprocessed tensors and load it

App will auto select VRAM Preset according to your GPU at initial load

If you get Out of Memory Error (OOM), move to 1 below VRAM Preset

Available presets : 24GB+, 16-24 GB, 12-16 GB, 10 GB+, 8-10 GB

48 GB or above GPUs can try turning off Gradient checkpointing - not tested yet

Select your LoRA Base Model again - same as auto label and tensor generation

Set your LoRA Training Name

Currently other parameters are set as best according to VRAM Preset

Learning rate and Training Epoch Count is in research

So far I tested below settings on 50 songs dataset

Set your Epoch count and Save Every N Epochs before starting training

Batch size 1 and Gradient Accumulation steps 1 are best quality

Set your Output directory as well where LoRA and training state files will be saved

This will be auto set to Loras folder so you can auto select from advanced tab Loras

Quick LoRA selection option added to fast generation tab as well

So you will be able to select any use your LoRAs immediately from this folder

You can also resume from training state file

Once all is set until this point you can start training

If you want to generate samples during training, we support it as well

So before starting training, make sure to enable sample and write your style and lyrics if you want

To update download latest zip file, overwrite older files and run Windows_Install_or_Update.bat file

ACEStep XL 1.5 SFT model training recommend but I am testing right now not concluded yet

Hopefully will make a full training tutorial soon

It takes around 100 minutes for RTX 5090 for 5000 steps with sample generation so pretty fast and best config uses 22.6 GB VRAM - so fits into RTX 4090 or 3090 as well

Model variants comparison as below

15 May 2026 V1.1 Update

The behaviour of Songs, Number of songs to generate sequentially is fixed

Will generate multiple songs in a loop

Just run Windows_Install_or_Update.bat to update

How To Use ACEStep 1.5 XL SECourses Premium App and its features

Main app interface screenshot

All 3 models SFT, Base and Turbo are fully automatically supported with VRAM presets

Each preset automatically updates best values when you change the selected ACEStep XL 1.5 model from selection

Each preset change also updates necessary config automatically according to VRAM tier selection

More VRAM tiers may have higher quality since using some different configs and models such as using acestep-5Hz-lm-4B vs 1.7B vs 0.6B, etc.

So for best quality, you may run the app on cloud services

3 minutes song generation takes around 40 seconds on RTX 5090 with Turbo model, other models are also very close to this and all are really fast locally

We have high quality FP8 Cache feature as well - custom implemented

On the first time, it will generate FP8 Scaled version of the used model, save it and use it when you next time use FP8 Scaled

You can set most needed parameters regarding song / music generation on our specially designed easy generation screen like below

Change language of your song

Change Male / Female

Set Instrumental

Set song duration, -1 means auto

Set number of songs you want to generate

You can also provide an image and if provided it will generate additional MP4 music video with generated audio file with desired video resolution while keeping your image aspect ratio

All generations are fully saved with full metadata into outputs folder as sub folders

You can use open outputs folder to open it quickly

For advanced users our advanced tab supports all the features these models have

They are Custom, Remix and Repaint

Every feature has a detailed instructions and description so read everything to understand how app works

You can set all the parameters individually but totally not needed since the presets we developed auto sets all of them and service is auto initialized so no need to click

We support LoRA folder feature as well that you can pick and use automatically

LoRA folder is ACE-Step_Premium\Loras - put your custom LoRAs here, app also supports LoRA training

More custom parameters that you can set if you want but all auto set to best with presets

You can make repaint as well

We have a custom Library page where you can see all your generations with their metadata

It is daily based filtered and very convenient to use

We have results page for some custom fast and easy operations

We have custom preset system where you can save your config and load them later if you want

Your last saved / used config will be auto remembered at next launch

Delete this folder to return back to defaults : ACE-Step_Premium\premium_user_presets

We have fully working LoRA training and I plan to make tutorial with it later hopefully

But currently you can use LLMs like Codex or Cursor or Claude or Gemini, etc. to get their help or look online sources and other sources or just try and learn yourself to train

We have fully working batch folder processing

Name your song txt files as you wish which will contain lyrics and also make their style files with suffix _style.txt in same folder

e.g. rap_song.txt and rap_song_style.txt, awesomesong.txt and awesomesong_style.txt and so on

SECourses: FLUX, Tutorials, Guides, Resources, Training, Scripts PATREON 32 favs
VIEWS1
FILES143 files
POSTEDJul 12, 2026
ARCHIVEDJun 22, 2026