ACE-Step 1.5 XL Premium - Better Music & Song Generator Than SUNO 5.0 - Remix and Repaint Features, SAM Audio Processing - Windows, RunPod, Massed Compute, Linux 1-click Installers
Our Patreon exclusive posts index list all of the apps we have (over 100+ AI apps, scripts, trainers, presets and more). Use CTRL+F to find whatever app you looking on our index.
Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord
Please also Star, Watch and Fork our Stable Diffusion & Generative AI GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)
=======
Latest installer zip file : ACE_Step_v7.zip
[click here to choose a membership and Join to download zip files]
Quick Info
This app has the following repos perfectly combined into our premium app with additional improvements and features such as optimized model loading, VRAM, quality, accuracy and performance optimizations, batch folder processing and many more (all models automatically downloaded and everything installed into a Python 3.11 VENV)
We got VRAM presets for every GPUs already set, read changelogs below to learn everything, slowly top to bottom read recommended
ACESTEP XL 1.5 (both inference + training) : https://github.com/Runware/ACE-Step-1.5-XL
https://deepwiki.com/ace-step/ACE-Step-1.5/5-generation-features
SAM-Audio Segment from Facebook / META : https://github.com/facebookresearch/sam-audio
Massive optimizations made for this model , it is working amazing
Auto-Editor : https://github.com/wyattblue/auto-editor
TrackAICleaner Post Processing : https://github.com/mikecastrodemaria/TrackAICleaner
DiffPitcher : https://github.com/haidog-yaqub/DiffPitcher
ACE Step 1.5 XL is the newest State Of The Art (SOTA) Music and Song generator model. It has 3 variants and we support all 3 variants (Turbo, SFT, Base) with fully automatic setup, models download, VRAM presets for all GPUs starting from 4 GB and with all best researched generation values / settings / configurations.
Windows Requirements
Python 3.12.10, FFmpeg, CUDA 13 cuDNN 9.17, Git, Visual Studio Community Edition with Desktop Development with C++ (all checkboxes checked)
Make sure that your NVIDIA driver is updated (min 590+)
If you get any errors follow below video and its source link
https://www.patreon.com/posts/windows-requirements-tutorial-written-post
The zip file contains installers for
Windows : Windows_Install_or_Update.bat
Please follow requirements video for Windows before starting installation : https://youtu.be/DrhUHnYfwC0
Requirements tutorial is 1 time mandatory for all of my applications
Windows installer will download only ACEStep 1.5 XL Turbo model
To download all models also run Windows_Download_All_Models.bat after installation
RunPod and SimplePod : Runpod_SimplePod_ACE_Step_Instructions.txt
Massed Compute / Local Linux : Massed_Compute_Instructions_READ.txt
RunPod, Massed Compute installers automatically downloads all 3 ACEStep 1.5 XL models, Turbo, SFT and Base
Zip file also has ACE_Step_Lyric_Generation_Instructions_For_LLMs.txt which you can use to better format your Music / Song lyrics and style by providing this file to your favorite LLM
The installers will generate a Python 3.12 VENV automatically and install everything inside there, thus your system or any other of your APPs will never be impacted
With our pre-compiled abi3 wheels for Windows and Linux, you can run on Python 3.10, 3.11, 3.12 and 3.13 but preferred version is 3.12
The following libraries compiled for Windows and Linux
archs : mslk, xformers, flash_attn, sageattention, torchao,
All of them are abi3 and Windows versions are compiled for Consumer GPUs, Linux versions compiled for Consumer + Cloud GPUs - so we cover all GPUs you have
12 July 2026 V7.0.1 Update
You need to delete your venv folder and run this update - existing installation may remain
I have started upgrading all of our apps into latest Torch 2.13 and CUDA 13
For this, I have pre-compiled the following wheels with all CUDA 13 features and with all CUDA archs : mslk, xformers, flash_attn, sageattention, torchao
All these libraries are properly compiled with abi3 thus works on Python 3.10, 3.11, 3.12 and 3.13
So I have updated the SECourses ACE-STEP XL 1.5 installers with newest libraries
It will automatically generate Python 3.12 venv and install everything there, if you don't have 3.12, it will use your default python
Make sure to have 3.12.10 it is best
Make sure to have nodejs 22+ installed - 22 preferred and tested
Follow requirements tutorial ( https://youtu.be/DrhUHnYfwC0 ) and latest ACE-Step tutorial : https://youtu.be/hzKSt5WUAm0
At the LoRA training tab, we had missed supported languages (50)
Now you can train with all Languages the model supports for inference
Arabic, Azerbaijani, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Persian, Finnish, French, Hebrew, Hindi, Croatian, Haitian Creole, Hungarian, Indonesian, Icelandic, Italian, Japanese, Korean, Latin, Lithuanian, Malay, Nepali, Dutch, Norwegian, Punjabi, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Serbian, Swedish, Swahili, Tamil, Telugu, Thai, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Cantonese, Chinese
We have massively improved Torch Compile with newest Torch 2.13 and our massive backend improvements
Now Torch Compile happens with 8-cpu threads by default
Moreover, there were unncessarily compiled parts and they are not compiled anymore
Moreover, we have updated some libraries and fixed some bugs
The result is that, previously Torch Compile was taking 240 seconds to compile, now taking like 10 seconds
Once Torch Compile is warm, it is taking around 20-30 seconds to generate full songs on RTX 5090 - click to result
A wider audit of the app has been made and following bugs fixed
Presets lost multi-select and explicit empty CheckboxGroup values.
Negative prompts did not reliably synchronize both ways.
GPU tiers did not synchronize from Advanced back to Generate Song.
Initialize Service returned fewer values than its event wiring expected.
Batch Folder ignored Generate Song duration and seed settings.
Grid Testing passed shifted positional LoRA arguments.
Saving a dataset could erase per-sample instrumental labels.
Stopped LoKr training incorrectly reported successful completion.
For updating, get the latest zip file, overwrite older files, delete venv folder inside ACE-Step_Premium and run installer
5 July 2026 V6.3.4 Update
Negative prompt disabled for Turbo models since they are CFG 1
Negative prompt synching with advanced tab - simple tab issue fixed and now works smooth
Just run install update bat file to update
29 June 2026 V6.3.3 Update
ACESTEP XL 1.5 VRAM presets Tier 3, 4, 5 and 6a significantly improved
When first time model is quantized into Int8 or FP8, it was taking too much time and this issue fixed now like 100x faster
Moreover, after first time caching, the cache files will be saved inside ACE-Step_Premium\.cache and reused next time so next time generation will be instant even after restarting the app
Just run installer / update bat file to upgrade
25 June 2026 V6.3.2 Update
How Remix Source Start and Remix Source End works upgraded
Now it regenerates entire remix and then only merges the selected part into remixed original song
With this approach, you can iteratively remix certain parts quickly and get the ultimate best song
Just run installer / update bat file to upgrade
23 June 2026 V6.3.1 Update
Same V6 zip file just run installer to update
We fixed an important generation quality regression
The issue was caused by the new Remix Melody Retention default value leaking into normal Text-to-Music generation
This made XL Turbo start diffusion almost at the end of the schedule, so an 8-step generation effectively ran only about 1 real DiT step.
That explained why some users saw extremely fast generations, for example around 20 seconds instead of the expected longer runtime, with poor quality results.
We also fixed Turbo inference step handling. Previously, Turbo requests above 8 steps were still being clamped back to 8.
Now Turbo supports up to 20 steps correctly, so setting 16 steps actually runs 16 steps.
Sampler mode added to the Generate Song tab and set heun as default since almost no speed difference and it is better than previous euler
Still you can test and compare both if you wish
I did a lot of testing with Remix presets and sadly it is hard to make best for every case so you better test for your own cases
Now there are 4 presets, Default, Same Lyrics Big Change, Same Lyrics Medium Change and Different Lyrics
Enable torch compile and generate subsequently until you get a good result really fast
Negative prompt feature implemented to every field
Works only with SFT and Base model since Turbo model has CFG 1
New button Use Generated Result as Source added so that you can quickly set generated song to remix, edit, whatever you want to do
Useful for iterative processing
Advanced section of ACESTEP XL 1.5 app completely revamped and made much better as below
21 June 2026 V6.0 Update
Gangam Style in English : https://x.com/SECourses/status/2068512611725975733
This is a massive improvements and fixes update
We have moved to the Gradio 6.19 and thus transformers library upgraded to 5.3
So i had to fix pipeline for newest transformers library
PyTorch 5Hz LM generation now uses modern Transformers forward-pass features such as cache_position and logits_to_keep when available.
Gradio interface event handling was heavily improved for Gradio 6.19.
Many UI sync events now run without queue overhead and without unnecessary progress overlays.
Stale Gradio status timers are automatically hidden, fixing stuck timer/progress artifacts.
Long-running actions now show progress only on relevant outputs instead of slowing down unrelated UI elements.
Library metadata display was changed from heavy JSON rendering to copy-friendly text, making the library page smoother.
Batch processing, batch extract, audio processing, generation, library, LoRA, LoKR, dataset, and SAM Audio UI wiring were updated for smoother behavior.
Model loading architecture updated for newer transformers library and now model loading faster
Due to newer transformers library, now torch compile is even faster than before
PyTorch LM loading now tries the faster SDPA attention path on CUDA.
Audio-code generation now has a compact valid-token sampling path, avoiding unnecessary full-vocabulary processing during constrained generation - No quality loss
VAE tiled decode was optimized by preventing pathological tiny-stride chunking - Nno quality loss
With VAE optimization + transformers library, now torch compile is able to generate full song in 30 seconds on RTX 5090
GPU VRAM presets are re-tested and updated as below
Wildcards special character was [] and now it is fixed and changed into {}
CoT Language Detection and Caption / Style Auto Improve was enabled in some presets and now they are all disabled - so you have to manually enable
They were causing unexpected issues and problems
If CoT language is enabled, the LM-detected language is used only when vocal language is set to auto/unknown.
Explicit user-selected vocal language is preserved and no longer unexpectedly overwritten.
In advanced tab now you can set explicit vocal language and this fixed so many issues
Advanced tab set Vocal Language will update Generate Song tab Vocal Language as well or vice-versa
Remix presets implemented and literally 1-click first test result you can see here
Gangam Style in English : https://x.com/SECourses/status/2068512611725975733
Hopefully will make a mini tutorial video so open bell on Youtube : https://www.youtube.com/SECourses
Batch audio processing now reports status immediately when scanning starts.
Batch Extract now normalizes itself to Extract mode internally, instead of requiring the user to manually switch generation mode first.
Batch progress display was improved.
Batch queue restore defaults now keep CoT caption/language disabled unless explicitly enabled.
For updating please get latest v6 zip file, overwrite previous files and run installer bat file
If you get any errors, please delete ACE-Step_Premium\venv and then run installer again
19 June 2026 V5.5 Update
Full tutorial video published finally for inference : https://youtu.be/9C_6qNKjgpA
I started working on LoRA training tutorial as well hopefully soon
With 5.5 optimizer specific parameters are now shown that you can set, I am also working on to make them auto default best hopefully
There was a visual bug that hidden Remix Melody Retention and Direct Source Latents (no_fsq) on Remix songs page and this bug fixed and app scanned entirely and all visuals verified
Default value set to 0.97 one of our expert remixer recommended that
Just run Windows_Install_or_Update.bat to update, the zip file not changed
18 June 2026 V5.4 Update
Now batch folder processing for ACESTEP XL 1.5 and SAM Audio has this extra option Save only output
This is useful to get only processed files and no other stuff like remaining part of the songs or metadata files, etc.
18 June 2026 V5.3 Update
Wildcard feature implemented
It works both for Style / Captions and Lyrics with syntax verification as well
It will work in batch folder processing as well so you can write that way in txt files
If you enable Auto improve lyrics or Auto improve style they may break your syntax so don't enable when using wildcards
Just run Windows_Install_or_Update.bat to update same zip file still
Also full inference tutorial published that covers every topic in details including how to install on Windows, RunPod, Massed Compute and SimplePod : https://youtu.be/9C_6qNKjgpA
16 June 2026 V5.2 Update
Default Remix value is now 0.95 instead of 1
Seed box and Random seed option moved to a much easier to use place
Last generation seed value will be auto set in seedbox so you can uncheck random seed and keep working with same seed now easier
14 June 2026 V5.1 Update
Use Repaint with lyrics added to the Repaint tab of ACESTEP XL 1.5
14 June 2026 V5.0.0 Update
I am still working on inference tutorial and as I used and as you made new feautre requests new features arrived
Auto-Editor trim output was not working properly and this bug fixed now should work much better when you use it in SAM Audio Segment or ACESTEP XL 1.5 Extract
This is really useful to get only vocals and trim empty / no vocal parts for training
Extract All stems feature implemented to ACESTEP XL 1.5 extract tab
Extracted stems will be saved in same folder with suffixes like brass, guitar, vocal, etc.
I noticed that extracting stems much better working on full songs rather than part of songs like 1 minute split part for some reason for ACESTEP XL 1.5 extract
Extract logic improved
Each different extract may yield different results so you can try multiple times to get better extract
Auto-Editor workflow export significantly improved
In Audi Processing tab enable Auto-Editor trim silent sections
Then Set Processing Preset = None
Then select your Auto-Editor workflow export like DaVinci Resolve
Then use Local Audio/Video Path with Browse File button or direct path
This way you will get almost instantly .fcpxml with accurate file path or whatever supported format you pick
SAM Audio Segmet now supports Batch Segment
You can use Batch Segment with 2 ways
First way is enable Batch Segment checkbox and type your stems / segments into Custom Prompt with ; seperation
Second way is select multiple Quick Prompt from dropdown and it will segment / extract every one of them
Custom prompt section overwrites Quick Prompt selections
Extracted stems / segments will be saved in same folder with suffixes like brass, guitar, vocal, etc.
Load Metadata feature implemented as a new tab
Select the generation_manifest.json and it will load every single configuration / parameter of that generation
Get the latest zip file, overwrite older files and run Windows_Install_or_Update.bat file for update or fresh install
To have all models (ACESTEP XL 1.5 Base and ACESTEP XL 1.5 SFT) run Windows_Download_All_Models.bat after installation
12 June 2026 V4.9.5 Update
In Audio Processing tab now there is None Processing Preset which unchecks all Audio Enhancement
Now there is Disable upload preview checkbox in Audio Processing tab
Use for very large videos or containers like multi-GB MKV files. When enabled, Gradio will not render the uploaded media preview, avoiding slow browser/Gradio post-processing such as MKV-to-MP4 preview conversion. Processing still uses the original uploaded file.
Gradio does post processing to every video file if not mp4 therefore other formats will take massive time to display if they are big : https://github.com/gradio-app/gradio/issues/13527
11 June 2026 V4.9.4 Update
Auto-Editor trim silent parts descriptions updated
Apply automatically to generated songs was mistakenly enabled by default and this issue fixed so you can enable if you wish
Analyze button won't overwrite your lyrics anymore
Auto-Editor trim silent parts feature in SAM Audio Segment and ACESTEP XL 1.5 Extract will now use the settings / parameters set in Audio Processing tab
In ACESTEP XL 1.5 extract mode when Auto-Editor trim was selected, it was not working accurately and now will work a bug fixed
Latest generated results sections labels fixed - for ACESTEP XL 1.5 advanced tab
For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat
11 June 2026 V4.9.3 Update
In the repaint task, if generated song is shorter than the selected repaint area, it will trim thus you won't have silent parts
The Repaint Strength description updated and fixed : When lyrics are provided, Repaint switches to text-to-music, so Repaint Strength has no impact when changing lyrics. To keep the same vocal audio, LoRA training and using that LoRA are mandatory.
Now output format can be selected in ACESTEP XL 1.5 modes
Default is set as mp3 since generated files were taking too much space
Now all generated files will obey the selected format e.g. like below
Lego mode was not working accurately and this issue fixed
Now in Lego mode, you will see only generated output as well such as you selected guitar so you will get the generated guitar song as well like below
For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat
10 June 2026 V4.9 Update
V4.9 is a pretty big and important update lots of fixes and improvements
Generation modes now explicity shows recommended models for ACESTEP XL 1.5
Previously, switching models without restarting the app was causing VRAM leak and OOM
This issue is fixed and now you can generate with Turbo model and then switch SFT or Base, and so on
To be 100% sure not have any RAM or VRAM leak, enable Use isolated subprocess generation checkbox
This option will slow subsequent generations and not mandatory, so enable if you are sure and needed
For Remix, Repaint, Lego and Complete, now you can set Instrument Start and End of source input and it will show live preview, really useful for Repaint
Instrument Start and End selection was not working accurately for Remix, Repaint, Lego and Complete but this bug fixed so now you can repaint just specific part of the model
Repaint was not using accurate methodologies and automatic inner prompt to repaint song accurately and now this issue also fixed
So now you can change specific part of the song and make it sing different vocal / lyrics etc perfectly working tested
Remix, Lego, Repaint and Complete mode errors fixed and they are made more robust
Optional Parameters, Batch Process, Settings will be closed by default now, so easier to read interface
Click them to open them again
Cluttering unrelated some information from Remix, Lego, Repaint and Complete modes removed such as Custom Guide from Remix
Generated results now will show followings
Latest Generated Result (Sample 1) : Is the full new repainted, remixed, etc song
Next to it Original Input, the original song for quickly listen both and compare
Latest Repainted Area, is the area of the song you repainted like between 30-40 seconds, this works for other modes too so you can listen only that particular section
Next to it, Latest Repainted Area Original, the original part of the song that was repainted, etc. to see before after quickly
For update / install use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat
10 June 2026 V4.8 Update
Torch compile feature implemented for ACESTEP 1.5 XL and SAM Audio processing
For ACESTEP XL 1.5, switch to advanced tab and enable, then you can switch back to Generate Song tab
ACESTEP XL 1.5 training also supports torch compile but not tested and verified yet
The initial torch compile may take some time but after that, repeated usage brings massive performance boost as shown as below
It won't recompile once compiled even if app is restarted, so it uses compile cache, if necessary it will recompile though
Initial compile may take time and may look like frozen but both inference and training tested and working
You have to have accurately setup CUDA, MSVC and C++ Tools for this to work since Torch compile depends on it
Therefore, follow requirements tutorial fully properly : https://youtu.be/DrhUHnYfwC0
The system is very robustly designed to automatically find accurate CUDA and C++ tools installation even if you have multiple installations
LoRA training speed with Torch Compile is 0.98 it / second and without Torch Compile is 0.78 it / second
25% faster
Use latest zip file (4_7), overwrite and run Windows_Install_or_Update.bat to update
ACESTEP XL 1.5 Inference Torch Compile
SAM Audio Inference Torch Compile
ACESTEP 1.5 XL LoRA Training Torch Compile
6 June 2026 V4.7.1 Update
Auto-Editor executable download now has alternative source if GitHub fails - now more robust
New feature DiffPitcher Pitch Fixer implemented into Audio Processing tab since requested
You can read more about it here : https://github.com/haidog-yaqub/DiffPitcher
The installer bat file will download necessary diffusion models automatically as safetensors files
Get latest zip file (4_7), overwrite previous files and run Windows_Install_or_Update.bat to update
6 June 2026 V4.6 Update
Audio processing tab significantly improved a lots of new features added
Now supports Run as subprocess and cancel button immediately
Now fully supports video inputs
Now supports Export Only Audio - very useful for getting audio from video if you don't need video
Now avoids reencoding of videos only if audio of video is processed - Auto-Editor triggers video processing
Now supports video re-encoding profiles
Now supports Auto-Editor workflow export for Davinci Resolve, Adobe Premiere Pro, Final Cut Pro, Shotcut and Kdenlive
Thus you can trim silent parts of your videos and continue editing in your favorite app, I use this literally to edit my tutorial videos
Now fully shows Audio Processing tab process progress in CMD and also on Gradio
Auto-Editor video processing may take quite time since it re-encodes video
Hopefully will make new tutorial soon
Zip file is same, just use Windows_Install_or_Update.bat to update
4 June 2026 V4.5 Update
SAM Audio model loading speed significantly improved like 2.5x faster than before
Unchecking Subprocess mode in SAM Audio was not working now works
So if you uncheck, after processing, it will keep model in VRAM thus instantly starts processing next task - in batch mode it doesn't unload model even if it is checked until batch process ends
New feature Predict spans added to the SAM Audio
Uses SAM-Audio's span predictor to estimate target time ranges from the text prompt when you did not provide anchors
This can improve quality of results depending on source file and the task so you can compare and see if improves
This can use slightly more VRAM and slightly slower
Advanced tab renamed into ACESTEP Advanced
Interface of following sections Custom, Remix, Repaint, Extract,Lego, Complete improved which are located in ACESTEP Advanced tab
Descriptions and buggy features of each section updated and improved as below:
Custom: Manual mode for precise control over caption, lyrics, BPM, key, duration, sampler settings, and advanced generation parameters. Use it when you already have a clear target and want to tune the result yourself. Switch to Generate Song main tab when you want to describe the idea in plain language and let AI fill in the details.
What it does: generates new music from your manual Caption, Lyrics, BPM, key, duration, and advanced settings.
How to use it: describe the target style and vocal delivery in Caption, write structured Lyrics with tags such as [Verse] and [Chorus], then set metadata only when you need tighter control. Leave Think on when you want the LM to plan; turn Think off only when using pasted LM Codes Hints.
Audio inputs: Reference Audio can guide timbre, mix, performance feel, and atmosphere, but it will not copy exact melody, rhythm, or lyrics. Source Audio is ignored in normal Custom generation and is only used by the Edit morph workflow.
Remix: Upload source audio and restyle it with your own caption and lyrics. The AI uses the original as a structural guide while applying your new style. Adjust Remix Strength to control how closely it follows the original (high = faithful cover, low = loose reinterpretation).
What it does: uses Source Audio as the structural guide for melody, rhythm, chords, arrangement, and timing while applying your new Caption and Lyrics.
How to use it: upload Source Audio, optionally trim it in Source Audio Preview, write the target style in Caption, provide replacement Lyrics if you want changed vocals, then adjust Remix Strength and Remix Melody Retention. Higher Remix Strength follows the source more closely; lower strength gives the model more room to reinterpret.
Audio inputs: Source Audio is the important input here. Reference Audio is only an extra global style cue. If the source is instrumental-only, Remix can follow the instrumental structure but still has to invent the vocal melody and phrasing for new lyrics.
Repaint: Upload Source Audio, choose a start/end range, and regenerate only that range. Caption/Lyrics describe the replacement section. Optional Reference Audio can guide style/timbre, but it is not the audio being edited.
Repaint: regenerate one time range of the source
What it does: keeps the Source Audio context and redraws only the selected start/end range. Use it to fix a bad section, replace a lyric phrase, change a solo, or smooth a transition without regenerating the whole song.
How to use it: upload Source Audio, set Repainting Start and End in seconds, then write Caption and Lyrics for the replacement section only. Use Conservative to protect boundaries, Balanced for normal edits, or Aggressive when the selected range should change more freely.
Audio inputs: Source Audio is the audio being edited. Reference Audio can nudge style/timbre for the replacement, but it is not the editable source and will not force exact melody or lyric timing.
Extract: Isolate a single track (vocals, drums, bass, etc.) from source audio using AI stem separation. Useful for creating instrumentals, acapellas, or isolating parts for remixing. Available on Base only.
Extract: isolate one stem from source audio
What it does: separates one selected track from Source Audio, such as vocals, drums, bass, guitar, keyboard, or other supported categories.
How to use it: upload Source Audio, choose Track Name, select the Extract output format if needed, then click Extract Stem. Use the extracted stem for acapellas, instrumentals, remix prep, cleanup, or analysis.
Audio inputs: Extract uses Source Audio only. Caption, Lyrics, Reference Audio, Think, BPM, and key are not creative controls for this mode.
Lego: Choose a predefined instrument category such as synth, bass, drums, or guitar. The AI generates that instrument and adds it over the existing source audio; you do not upload external stems. Upload the source track for context, trim it in Source Audio Preview if needed, choose the instrument to add, and describe only that new layer. Available on Base and SFT; Base generally gives the best results.
Lego: add one generated track over existing audio
What it does: creates the selected instrument category and layers it over Source Audio. This is for adding a new AI-generated part, not for uploading your own external stem.
How to use it: upload Source Audio, choose Track Name such as vocals, backing_vocals, drums, bass, guitar, or synth, optionally set the start/end range, then describe only the new layer in Caption. For vocals, provide Lyrics and describe the singer/delivery.
Audio inputs: Source Audio gives musical context for the new layer. Reference Audio can nudge global sound, but it will not act as a guide vocal. If you add vocals to an instrumental, the model must invent the sung melody and phrasing unless the source already contains that vocal structure.
Complete: Fill in selected missing tracks from source audio. Upload a partial arrangement or single stem, trim it in Source Audio Preview if needed, choose the tracks to add, and optionally set Complete Start/End to regenerate only that section while preserving the rest of the source. Available on Base and SFT; Base generally gives the best results.
Complete: fill missing tracks in a partial arrangement
What it does: listens to Source Audio and generates the selected missing track classes so the partial idea becomes a fuller arrangement.
How to use it: upload a partial track, single stem, or incomplete mix, choose the track classes to add, optionally set Complete Start and End to limit the generated section, then describe the desired finished arrangement in Caption. Use it for adding accompaniment around vocals, drums/bass under a sketch, or missing instruments in a section.
Audio inputs: Source Audio is the context that the new tracks must fit. Reference Audio can guide overall style, but it does not replace the source and does not force exact melodic or lyric timing.
Get latest zip file (4_3), overwrite previous files and run Windows_Install_or_Update.bat to update
4 June 2026 V4.4 Update
SAM Audio processing bug fixed
In ACESTEP XL 1.5 Advanced Extract tab, Analyze button was useless now it will show info message to use Track Name and click Extract Stem
Extract Stem will now show progress and status in Latest Result Status
Extract Stem limited to ACESTEP XL Base model since it works 100x better with Base than SFT
Now all advanced tab audio / video input fields will show preview
If preview doesn't show immediately, click X and re-select file this fixes Gradio bug
Now you can trim audio from Gradio preview as well
Audio previews visuality improved and trim feature visuality improved significantly for all upload audio fields and previews
Gradio version upgraded to 6.16.0
ACESTEP XL 1.5 Advanced Mode Lego and Complete features improved and bugs fixed
Description of how Lego mode works improved
Lego: Choose a predefined instrument category such as synth, bass, drums, or guitar. The AI generates that instrument and adds it over the existing source audio; you do not upload external stems. Upload the source track for context, trim it in Source Audio Preview if needed, choose the instrument to add, and describe only that new layer.
Get latest zip file (4_3), overwrite previous files and run Windows_Install_or_Update.bat to update
3 June 2026 V4.1 Update
V4.2 is a massive update so please carefully read all
For update please get latest v4_1 zip file, extract into install folder, overwrite and run Windows_Install_or_Update.bat file
If you get any errors for any reason, delete \ACE-Step_Premium\venv and then run installer bat file
When you select ACESTEP XL 1.5 SFT or Base model, in advaced tab, now all these options will be enabled and fully work
Simple, Custom, Remix, Repaint, Extract, Lego, Complete
Extract now fully works and you can pick what to extract from Track Name below
However I think new SAM Audio model is better still test and compare both
You can also use batch extract feature now if you want to batch process a folder of songs
Audio Processing tab improved and now we support extremely famous Auto-Editor
Auto-Editor is amazing library to trim silent - no spoken parts
I use this to trim out videos and can be very useful to trim vocal extraction
I use this to also cut silent parts of my tutorials, very useful to pre-process before editing
New tab SAM Audio Segment implemented
SAM Audio is state of the art prompt and mask based audio processing / seperation model from Facebook : https://ai.meta.com/research/samaudio/
It supports any custom text prompt and the below quick select presets
I have made massive amount of optimizations and programming to implement this model
BF16 pre-converted safetensors SAM-Audio and SAM-Audio Judge models will be automatically downloaded when you run installer or model downloader bat file
Normally released models were FP32 .pt models but I converted them to BF16 and safetensors format
SAM Audio models official pipeline was also first loading into RAM and then moving into GPU thus using extra RAM and slower
I made it directly to be loaded into GPU as BF16
Our implementation supports full sub-process running and auto trim feature - extremely useful to extract vocals for ACESTEP XL 1.5 LoRA vocal training
We already have VRAM presets for every GPU out there for SAM Audio model
It works with 20 seconds segmentation with 5 seconds overlap
20 seconds segmentation is what model authors recommend and used for training
Longer segmentation not improving quality but increases VRAM usage and reduces processing time
Only missing feature is Multi-diffusion text-only mode since authors didn't publish this but I opened an issue and expecting them to publish hopefully
We already have that mode coded by CODEX but I think it is not better due to our inaccurate implementation
We support batch folder processing to pre-process training songs as well or for any reason you want
Flash Attention were not working on Windows RTX 4000 series GPUs and this issue fixed
I have re-compiled Flash Attention 2.8.3 to fully support RTX 3000, 4000 and 5000 series GPUs with extra CUDA Arch a flag for SM120a
Linux Flash Attention with all GPUs (to include Cloud server GPUs too) SMs also recompiled and now will be used : 80;86;89;90;100;103;120
So no GPU should get any error with Flash Attention anymore
More information regarding CUDA archs : https://www.patreon.com/posts/159064759
ACESTEP XL 1.5 Advanced tab now supports uploading video files as well
They will be automatically converted into audio and used
If your video upload shows processing forever, click X icon and reupload
This is a Gradio bug I am trying to fix, refresh page also fixes
Audio Processing tab supports both Audio and Video uploading
SAM Audio Segment supports both Audio and Video uploading
Video upload previews are now capped to height 400px so they won't take entire web page space and look much better
31 May 2026 V3.9.1 Update
New full audio post-processing tab implemented to our premium app from TrackAICleaner repo
You can use this tab to both post-process your existing audio files as batch or as single file or automatically post process your generated songs
When it is enabled to auto post-process generated songs, it will save both original and post-processed songs in the outputs folder
You can use preview button to generate 60 second preview and compare quickly the effect impact
Use latest newer zip file, overwrite and run installer to update
28 May 2026 V3.9 Update
New feature LM Audio Codes added and enabled for all default presets
This is supposed to improve quality in all generations without any loss or VRAM increase
Updates made to fix below error that some users reported
Sadly I couldn't reproduce it yet to verify
Error: Generation produced NaN or Inf latents (shape=[1, 8261, 64], dtype=torch.bfloat16, device=cuda:0, nan=528704, inf=0).
Same zip file just run installer to update
26 May 2026 V3.8 Update
Version 3.8 is a very major upgrade for training
In Advanced tab when you click Analyze button now it will auto initialize model and won't throw error
Custom Preset System save and load issues fixed for some cases
Now we support DoRA for both training and inference (song generation)
DoRA is like LoRA but better quality for training close to full Fine Tuning of the entire model
Moreover now we have Target MLP feature for training
Also applies LoRA/DoRA to decoder MLP layers (gate_proj, up_proj, down_proj). This increases trainable capacity and VRAM use; leave off for the legacy attention-only path.
When MLP enabled, more parameters are trained thus it may be a little bit slower and may require more VRAM but it should improve quality I am still in research
Training Parameters screen significantly improved with lots of new features
Now we have Save best feature
It will save best loss having checkpoint and as new best loss having checkpoint reached, it will overwrite previous best
You can set Best smoothing window, Best min delta and Start saving best after epoch as you wish to make it as you wish
Now we support following Optimizers : adamw, adamw8bit, adafactor
Now we support following Schedulers : cosine, cosine_restarts, linear, constant, constant_with_warmup
I usually prefer constant, I am still in research of best hyper parameters and accurate way of training
Now we support following Timestep modes : continuous, discrete
Now we support Adaptive timestep
Now we have an amazing new feature to measure quality of training progress : Validation split %
Lets say you have 100 training songs, and you set 5%, so it will set 5 songs as validation and use 95 songs for training
Therefore, you can see actual quality improvement or degrade of the model therotically
You can see in below screenshot that as the training continued, the validation loss rate got worse and worse even though traidional loss got lower and lower
This is because the model got completely overtrained and cooked and memorized and lost its full generalization
With the newest features, training VRAM presets are also got updated as below
Auto LRC feature improved and now uses lesser VRAM moreover generated .lrc and .vtt files are now automatically saved inside generated output folder
Zip file is same just use Windows_Install_or_Update.bat to update
25 May 2026 V3.6 Update
Brings some massive improvements for training and LoRA usage
In Generate Song tab now Latest Song section will display status
Dataset tab completey remade and now when you load your generated dataset.json file, it will show caption statistics as word n-grams
You can select word-ngrams and it will list songs containing them
Then you can select songs to see their full details
Useful to see your dataset composition
Now there is Use Only Custom Trigger checkbox which will make all Caption / Style data to be only as your custom trigger
I am gonna test this approach to see if works better on training hopefully
We have a full new Grid Testing tab which lets you to generate songs with selected LoRAs to compare them very easy and quickly
So that you can compare your LoRA training checkpoints properly
You can select multiple LoRAs, any LoRAs you want with filter LoRAs extra feature to make selection easier
Grid results will have special naming and save to make them easier to compare e.g. like below
Just run installer to update
25 May 2026 V3.5 Update
When auto labeling, even though process finished it was still showing as processing at Raw Lyrics (from .txt file) and this bug fixed
The training speed display fixed and now it shows training speed accurately after first 20 steps completed like below
Auto labeling is taking massive time, therefore I have implemented true batch size to this process
User / Custom preset system now will save and load every field exists in training tab properly
Delete preset button now deletes the selected preset and loads the next custom preset
If all user custom presets get deleted, it will load default vram preset
Cancel auto labelling process and tensor generation process implemented
Now you can immediately cancel both processes when running in isolated subprocess
Now your lyric txt files can also contain style / caption
e.g. a.mp3 and a.txt in same folder
Format is like below
# Caption
Your caption
# Lyrics
Your lyrics
I tried to make it as much as possible robust so it should work fairly well see below example
Custom Trigger Tag was not working properly and this issue fixed
New feature Debug: save text prompts added so that you can see what is exactly used to generate Preprocessed tensor files which are actually used for training
Just run installer to update
24 May 2026 V3.0 Update
V3 is an important update for LoRA training
Some LoRA training bugs fixed and the training made more smoother
I also have compiled Flash Attention 2.8.4 latest version for Windows and Linux with massive GPU support for Torch 2.11 and CUDA 13
The reason is that we have upgraded our installer to Torch 2.11, CUDA 13 and Torchao 0.17.0 thus training is now even faster
This Flash Attention compile costed me like 50$ on RunPod for Linux and over 14 hours on Windows you can read more info here : https://www.patreon.com/posts/159064759
Download latest zip file, overwrite older files, delete ACE-Step_Premium\venv and run Windows_Install_or_Update.bat again to update or install
20 May 2026 V2.1 Update
Please read V2.0 update first
Browse Dataset JSON bug fixed which is needed to Preprocess files and generate training tensor files
LoRA refresh added to quick generation panel
In train LoRA tab we have improved the parameters you can set and default parameters are also improved
Now when you change training base model, it will auto update training parameters to best for each model specifics
Now Shift and Training Timestep Steps working accurately and set for each model : Base, SFT, Turbo
Now Resume Training State directly takes saved state file, state files are now directly saved read below to see
Unnecessary export LoRA and custom samples directory removed
How training generated files saved completely revamped and improved
Everything will be saved inside target folder with your training name like below
Much more organized and clean and ready to use after training
safetensors files are LoRA files ready to use and pt files are state files which you can use to continue training
Get latest zip file and run Windows_Install_or_Update.bat to update
Torchao upgraded to 0.16.0 for training
Zip file changes only when needed
20 May 2026 V2.0 Update
This is a massive upgrade and we have added so many new features so read carefully all
The deault models were all 32-bit however we were generating songs in BF16 or FP8 / Int8
Thus, the models were keeping double size on disk for no reason and taking more RAM and duration to load
Therefore, I have generated BF16 models and updated the app and model downloader
Thus, you can make a complete fresh install or, delete \ACE-Step_Premium\models folder and run Windows_Install_or_Update.bat / Windows_Download_All_Models.bat again to download new models
Make sure to get latest zip file and overwrite previous installer files
New all 3 models folder takes 41.9 GB, previously it was 69.8 GB
All default generation presets for all 3-models updated - the parameters are now more accurate
unlimited (>24GB) , tier6b (20-24GB), tier6a (16-20GB), tier5 (12-16GB), tier4 (8-12GB), tier3 (6-8GB), tier2 (4-6GB), tier1 (≤4GB)
ACEStep XL 1.5 Turbo, ACEStep XL 1.5 SFT, ACEStep XL 1.5 Base
Thus, you may expect better quality generation on ACEStep Turbo, SFT and Base models
LoRA selection added to Generate song tab as well since we now fully support LoRA training
Also before starting anything with LoRA training, select your Model from here
If you gonna train ACEStep 1.5 XL SFT, first select it from this screen to load all best SFT parameters then continue this is important
Use isolated subprocess generation was not working properly and this is fixed
Cancel Generation button added to both Generate Song and Advanced tab
For this button to work, Use isolated subprocess has to be enabled
🎓 LoRA Training tab completely remade
Now you can browse Dataset JSON Path and directly load
Now you can browse Browse Audio Folder and directly load
Dataset generation settings now fully working
Format Lyrics (LM) is not recommended
Transcribe Lyrics (LM) is recommended
If there are existing lyric files, Transcribe Lyrics (LM) will keep lyrics as it is and fill other data like Duration, Label, BPM, Key, Caption
Format is, audio_file_name.txt in same folder as audio
Also select your Dataset Model according to the model you gonna train and Dataset VRAM preset - it will be also auto set according to your GPU during initial start
If you get Out of Memory Error (OOM), move to 1 below VRAM Preset
Available auto-label and preprocess presets : 24 GB+ - quality, 12-16 GB, 10 GB+
Auto-Label All now fully working all bugs and errors fixed
Now use Browse Label Folder to pick auto labels saved folder
It will save labels after each labelling done
With using same folder, it will continue wherever left
Now you can navigate between each data item and manually change / fix and save
Selecting song from above listing will also update this Preview & Edit selected song
After auto labelling done, save your dataset into any desired location with any name
Now this auto labelling is fully VRAM optimized and auto unloads after completed and 100% free up VRAM and RAM
After dataset json saved, load dataset json and preprocess and generate training pt files into desired target folder
These pt tensor files will be used to train the model it contains every information needed to train
Now this preprocess is fully VRAM optimized and auto unloads after completed and 100% free up VRAM and RAM
Once you are done with Dataset builder move to Train LoRA tab
Train LoKr not tested yet but Train LoRA fully working and fully tested and optimized
Select folder of preprocessed tensors and load it
App will auto select VRAM Preset according to your GPU at initial load
If you get Out of Memory Error (OOM), move to 1 below VRAM Preset
Available presets : 24GB+, 16-24 GB, 12-16 GB, 10 GB+, 8-10 GB
48 GB or above GPUs can try turning off Gradient checkpointing - not tested yet
Select your LoRA Base Model again - same as auto label and tensor generation
Set your LoRA Training Name
Currently other parameters are set as best according to VRAM Preset
Learning rate and Training Epoch Count is in research
So far I tested below settings on 50 songs dataset
Set your Epoch count and Save Every N Epochs before starting training
Batch size 1 and Gradient Accumulation steps 1 are best quality
Set your Output directory as well where LoRA and training state files will be saved
This will be auto set to Loras folder so you can auto select from advanced tab Loras
Quick LoRA selection option added to fast generation tab as well
So you will be able to select any use your LoRAs immediately from this folder
You can also resume from training state file
Once all is set until this point you can start training
If you want to generate samples during training, we support it as well
So before starting training, make sure to enable sample and write your style and lyrics if you want
To update download latest zip file, overwrite older files and run Windows_Install_or_Update.bat file
ACEStep XL 1.5 SFT model training recommend but I am testing right now not concluded yet
Hopefully will make a full training tutorial soon
It takes around 100 minutes for RTX 5090 for 5000 steps with sample generation so pretty fast and best config uses 22.6 GB VRAM - so fits into RTX 4090 or 3090 as well
Model variants comparison as below
15 May 2026 V1.1 Update
The behaviour of Songs, Number of songs to generate sequentially is fixed
Will generate multiple songs in a loop
Just run Windows_Install_or_Update.bat to update
How To Use ACEStep 1.5 XL SECourses Premium App and its features
Main app interface screenshot
All 3 models SFT, Base and Turbo are fully automatically supported with VRAM presets
Each preset automatically updates best values when you change the selected ACEStep XL 1.5 model from selection
Each preset change also updates necessary config automatically according to VRAM tier selection
More VRAM tiers may have higher quality since using some different configs and models such as using acestep-5Hz-lm-4B vs 1.7B vs 0.6B, etc.
So for best quality, you may run the app on cloud services
3 minutes song generation takes around 40 seconds on RTX 5090 with Turbo model, other models are also very close to this and all are really fast locally
We have high quality FP8 Cache feature as well - custom implemented
On the first time, it will generate FP8 Scaled version of the used model, save it and use it when you next time use FP8 Scaled
You can set most needed parameters regarding song / music generation on our specially designed easy generation screen like below
Change language of your song
Change Male / Female
Set Instrumental
Set song duration, -1 means auto
Set number of songs you want to generate
You can also provide an image and if provided it will generate additional MP4 music video with generated audio file with desired video resolution while keeping your image aspect ratio
All generations are fully saved with full metadata into outputs folder as sub folders
You can use open outputs folder to open it quickly
For advanced users our advanced tab supports all the features these models have
They are Custom, Remix and Repaint
Every feature has a detailed instructions and description so read everything to understand how app works
You can set all the parameters individually but totally not needed since the presets we developed auto sets all of them and service is auto initialized so no need to click
We support LoRA folder feature as well that you can pick and use automatically
LoRA folder is ACE-Step_Premium\Loras - put your custom LoRAs here, app also supports LoRA training
More custom parameters that you can set if you want but all auto set to best with presets
You can make repaint as well
We have a custom Library page where you can see all your generations with their metadata
It is daily based filtered and very convenient to use
We have results page for some custom fast and easy operations
We have custom preset system where you can save your config and load them later if you want
Your last saved / used config will be auto remembered at next launch
Delete this folder to return back to defaults : ACE-Step_Premium\premium_user_presets
We have fully working LoRA training and I plan to make tutorial with it later hopefully
But currently you can use LLMs like Codex or Cursor or Claude or Gemini, etc. to get their help or look online sources and other sources or just try and learn yourself to train
We have fully working batch folder processing
Name your song txt files as you wish which will contain lyrics and also make their style files with suffix _style.txt in same folder
e.g. rap_song.txt and rap_song_style.txt, awesomesong.txt and awesomesong_style.txt and so on