VisoMaster Fusion and VisoMaster Automatic Installer - The Most Advanced 0-Shot Face Swap / Deep Fake APP - State of the Art - Windows and Massed Compute - Portable Installers
1-Click to install VisoMaster Fusion (newer much better) and classic VisoMaster on Windows and also on Massed Compute - SOTA Face Swap / Deep Fake APP - Newer interface and faster. Moreover portable 1-click Windows installers for All GPUs RTX 1000 series to 5000 series
Patreon exclusive posts index to find our scripts easily, Patreon scripts updates history to see which updates arrived to which scripts and amazing Patreon special generative scripts list that you can use in any of your task.
Join discord to get help, chat, discuss and also tell me your discord username to get your special rank : SECourses Discord
Please also Star, Watch and Fork our Stable Diffusion & Generative AI GitHub repository and join our Reddit subreddit and follow me on LinkedIn (my real profile)
=======
Follow the post slowly to learn how to use the app in full details and features
Latest installer zip file for VisoMaster Fusion (recommended most up-to-date) : VisoMaster_Fusion_v7.zip (recommended)
[click here to choose a membership and Join to download zip files]
Alternative Viso Master Fusion from Maxieto (Not works in Linux and Massed Compute) : VisoMaster_Fusion_Maxieto_v2.zip (not recommended)
Latest installer zip file for Viso Master Original : VisoMaster_v10.zip (not recommended)
Installers automatically installs with Torch 2.13 CUDA 13 and pre-compiled necessary libraries. Installer now also auto downloads necessary TensorRT and TensorRT Engine fully works.
If you don't want to follow requirements tutorial (https://youtu.be/DrhUHnYfwC0), you can right away use Windows_Start_Portable.bat
Windows_Start_Portable.bat automatically installs everything with all portable Python into a venv
So for Windows either Windows_Install.bat and then use Windows_Start.bat or directly and always use Windows_Start_Portable.bat
For Massed Compute and local linux machines : Massed_Compute_Instructions_READ.txt
14 July 2026 Massive Update
Deleting VENV folder is mandatory or make a fresh install
I have started upgrading all of our apps into latest Torch 2.13 and CUDA 13
For this, I have pre-compiled the following wheels with all CUDA 13 features and with all CUDA archs : mslk, xformers, flash_attn, sageattention, torchao
All these libraries are properly compiled with abi3 thus works on Python 3.10, 3.11, 3.12 and 3.13
Newly developed advanced face scan feature to find all faces and select the ones that are normally same but recognized by the models as different faces
Someone requested this and we added for you :)
Now play button directly shows processing FPS
Windows Requirements
For this auto installer to work you need to have installed Python 3.12.10 (may work with 3.10.x, 3.11.x, 3.13.x too), Git, FFmpeg, cuDNN 9.17+, CUDA 13.0, Visual Studio Community Edition with All c++ options
You can also use Windows_Start_Portable.bat - this doesn't require any requirements
Full updated requirements post with screenshots and links : https://www.patreon.com/posts/111553210
VisoMaster Fusion v7: The Complete AI Face-Swapping Studio
Installation, every feature, and real workflows
VisoMaster Fusion turns a folder full of specialist AI models into one coherent Windows production desk. It can replace and refine faces in still images, video, webcam feeds, virtual-camera output, and VR180 footage; preserve foreground obstacles; restore detail; transfer expressions and gaze; edit pose and makeup; organize repeatable presets; and render jobs in batches. The important part is not any one model. It is the control you get over the entire shot.
This is a complete, screenshot-led guide to the VisoMaster Fusion build audited on 13 July 2026. Every image comes from the installed application or its bundled, disclosed demo media. The final appendix is generated directly from the current interface definitions and covers all 396 declarative controls, while the main tutorial explains the surrounding menus, context actions, shortcuts, jobs, markers, and production workflow.
In this guide
Why VisoMaster Fusion feels different
Installation and launcher maintenance
Workspace tour
First complete face swap
Media, faces, and embeddings
Swap models, strength, and likeness
Masks, parsers, texture, color, and re-aging
Restoration and expression systems
Face Editor and denoiser
Detection, tracking, timeline, presets, and jobs
Output, live workflows, VR180, and performance
Real demo cases and quality control
Troubleshooting and production recipe
Complete 396-control reference
Why VisoMaster Fusion feels different
Many face tools stop at “choose two pictures and hope.” VisoMaster Fusion exposes the decisions that actually determine whether a difficult shot works: which identity model runs, how the face is aligned, how strongly identity is injected, which regions survive from the original, how occlusions are protected, where restoration happens, how expressions and gaze are driven, and how the final pixels are encoded.
That makes it valuable in two very different modes:
A beginner can load a target, add a source face, click the matching face card, enable swapping, and export.
An experienced operator can build a per-face pipeline with XSeg and parser masks, two restoration passes, expression retargeting, reference-guided denoising, tracking, frame enhancement, timeline markers, and batch jobs.
The application is especially compelling when a project contains several people. Each detected target face gets its own assignment and parameter set, so one person can use conservative settings while another uses a different model, mask, restorer, or expression treatment in the same shot.
The short version
The workspace at a glance
VisoMaster Fusion opens as a dense but logical studio. The center is the viewer. Media and identity libraries sit around it. The timeline and transport controls run underneath. A tall parameter stack holds the current face pipeline, while Settings, Jobs, Presets, output, and VRAM tools turn a preview into a repeatable render.
The full workspace at native resolution. Panels can be hidden from the View menu when the viewer or parameters need more room.
The major regions are:
Target Media: images, videos, and webcams you want to process. Filters switch among Images, Videos, and Webcams.
Input Faces: source portraits used to define replacement identities.
Faces / Embeddings: faces found in the target and reusable identities built from one or more source images.
Viewer: the current original or processed frame, with fit, 100%, pan, full-screen, theatre, compare, mask-debug, and save-image actions.
Parameters: Face Swap, Restorers/Expressions, and Face Editor are stored per selected target face. Denoiser is a global control whose reference data comes from the selected target’s assigned source face or embedding.
Timeline and transport: frame navigation, playback, markers, issue tools, recording, and audio.
Output: a compact destination field synchronized with the full output settings.
Job Manager: saved projects and unattended batch processing.
VRAM indicator and Clear VRAM: a live memory meter and explicit model-unload action.
The View menu can hide Target, Input, Job Manager, Faces/Embeddings, or Parameters. That is useful on smaller screens and when making a full-screen quality decision. Theatre mode can also be configured to force full-screen automatically.
The interface ships with 12 themes: True-Dark, OLED-Black, Windows11-Dark, Dark, Dark-Blue, Light, Solarized-Dark, Solarized-Light, Dracula, Nord, Gruvbox, and Monokai. Theme changes do not alter saved image/video pixels.
Your first complete face swap
The fastest way to understand the app is to finish one controlled project before exploring every slider. The bundled test_video.mp4 and face14.png make a high-contrast, reproducible pair: the source identity is visually distinct from both actors in the target shot.
1. Load target media
Use the Target Media file or folder action, or drag a supported image or video into the target list. Folder import can include subfolders when Target Media Include Subfolders is enabled in Settings. Select the new card to load it in the viewer.
For a webcam, switch the target filter to Webcams and add the desired device. Webcam backend, resolution, frame rate, and maximum stream count are configured in Settings.
2. Load a source identity
Add one or more clear portraits to Input Faces. A useful source is sharp, front-facing, evenly lit, and not heavily filtered. Multiple views of the same person can be combined into an embedding later; a single portrait is enough for the first pass.
The reproducible starting point: bundled target footage in Target Media and a bundled portrait in Input Faces.
3. Find the target face
On an image or the current video frame, choose Find Faces. The detected identity appears as a target-face card. On longer footage, open Scan Video and choose one of three strategies:
Quick: approximately two samples per second; best for a fast first inventory.
Smart: approximately six samples per second and the default balance for most footage.
Every Frame: exhaustive and slowest; valuable for very short, chaotic, or heavily cut material.
The scan opens a review grid instead of blindly adding every cluster. Inspect the candidates, select the people you want, and choose Add Selected. From a result you can also jump back to its source frame, which is extremely useful when a tiny or ambiguous face needs context.
Video scanning is review-first: choose sampling depth, inspect unique-face candidates, then add only the people the project needs.
4. Assign the source
Select the target-face card, then click the source-face card or embedding that should replace it. The assignment is per target identity. A normal click assigns one source; Ctrl-click adds or removes source faces and saved embeddings, and Shift-click can select a range of input faces. When several sources are assigned, the global Embedding Merge Method decides whether their identity vectors are combined by Mean or Median. In a group shot, repeat the process for each person and give each face its own settings.
Enable swapping with the main swap control or the S shortcut. The current frame updates through the selected pipeline. If a face is not being matched, lower Similarity Threshold cautiously; if the wrong person is being picked up, raise it.
Once a source is assigned, that target face owns its model, masks, restorers, expression treatment, and Face Editor state; global Denoiser and Settings controls remain shared.
5. Compare before refining
Press X to enable Face Compare. This is one of the best quality-control tools in the application because it makes identity, edge, color, eye, mouth, and occlusion differences immediately visible. Use the viewer’s fit command (Ctrl+0) for composition and 100% view (Ctrl+1) for texture and seam inspection. Middle-drag pans a zoomed frame.
The viewer can also display the swap mask, difference mask, or texture mask, depending on the active pipeline. These diagnostic views show why a composite behaves as it does. Right-click the viewer for Fit, 100% Zoom, Face Compare/mask choices, and Save Image; return to the normal result view before recording.
Start with the default Inswapper128 pipeline and resist the urge to turn on everything. Fix problems in this order:
Detection and assignment.
Alignment and identity strength.
Occlusions and borders.
Eyes, mouth, tongue, and parser regions.
Restoration and expression.
Color, texture, denoising, and final enhancement.
That ordering prevents a heavy restorer or color pass from hiding the real cause of an error.
6. Save or record
For a still, use the viewer’s save-image action. PNG is the default; Save Output Image in JPG Format changes both Save Image and image-batch output to JPG. For video, choose an output folder, move to the required start frame, and record with R. A direct recording is named automatically from the source plus a timestamp; a saved job can instead use its job name or an explicit output name. Timeline record markers can define several output ranges in one project, while standard markers store both per-face parameters and global controls.
The compact Output field above the target list and the Output field in Settings stay synchronized. Output to Target Location, Preserve Source Directory Structure, and Cluster Output by Source Name control where large batches land.
Media, faces, and embeddings
Target and input libraries
Both libraries accept individual files, folders, and drag-and-drop. Search narrows long lists. Thumbnail-size commands and clear/remove actions keep a large project manageable. The target list supports Images, Videos, and Webcams; source cards can open their file, be removed from the project, or be deleted from disk through their context menu, so read destructive labels carefully.
Settings can recurse through subfolders separately for target media and input faces. That is ideal for a prepared production tree, but disable it when a parent folder contains unrelated material.
Mean and Median embeddings
An embedding packages identity information so you can reuse it without reselecting raw portraits. Select several images of the same person and create:
A Mean embedding for a broad average across the chosen views.
A Median embedding that is more resistant to one unusual or poor source.
The denoiser can also use optional K/V maps associated with an embedding. Treat source selection as part of the creative process: a small set of clean, complementary angles usually beats a large folder containing blur, extreme expressions, occlusions, or inconsistent identities.
Advanced Embedding Editor
The Advanced Embedding Editor can load an identity, add another embedding into it, save under a new name, save selected components, and convert between VisoMaster and Rope formats. It includes search, select/deselect all, manual and alphabetical sorting, rename/copy/paste/delete context actions, and undo/redo.
Embeddings turn a one-off source selection into a reusable identity asset; the editor handles combination, conversion, organization, and revision.
Choosing a swap model
Open the Swapper section in Face Swap. The visible model menu contains ten choices, while the DFM option exposes its own downloaded-model selector.
Face Alignment Interpolation switches the face warp between Bilinear and Bicubic. Bilinear is the default; Bicubic can preserve more high-frequency edge and eye detail for a small cost. Pre Swap Sharpness sharpens the aligned original before the swap, while Similarity Threshold determines how closely a detected face must match the assigned target identity.
Model, resolution, alignment, and match threshold form the identity engine. Test them before piling on restoration.
Strength, anti-drift, and likeness
The Strength feature repeats or intensifies the swap up to 500%. Around 200% can help a weak identity, but repeated swaps can drift or flatten texture. Mode 2 (Anti-Drift & Texture) uses phase-correlation and frequency-separation logic to stabilize geometry and preserve skin texture during multiple passes. An amount of zero deliberately bypasses swapping while allowing the rest of the pipeline to process the original face.
Face Likeness applies a direct identity adjustment. Its Amount accepts negative as well as positive values, and Volumetric Strength controls the intensity of the likeness factor. Face Keypoints Replacer adjusts the source-face geometry toward the target’s keypoints by a chosen amount. These controls solve different problems: Strength reinforces the swap, Likeness adjusts identity features, and Keypoints changes geometry.
Identity tuning works best in small steps. Compare several representative frames after each change, especially at profiles and expressions.
Masks: the difference between a demo and a believable shot
A swap can have excellent identity and still fail because it paints over glasses, clips a cheek at profile, changes the inside of the mouth, or leaves a rectangular color seam. VisoMaster Fusion’s masking stack attacks those errors at several semantic levels.
Border and profile masks
Border Mask is on by default. It trims the bottom, left, right, and top of the swapped crop, then softens the transition with Border Blur. Increase only the side that shows a seam; overly aggressive borders can expose too much of the original face.
Profile Angle Mask fades the far side of a turned face. Angle Threshold determines when it begins, and Fade Strength controls the gradient. It is a fast solution for profile distortion without sacrificing the near side.
Occluder and XSeg
Occlusion Mask restores objects covering the face. Its Size expands or contracts the protected region, and Tongue/Mouth Priority prevents aggressive growth from erasing the inner mouth.
DFL XSeg Mask offers another learned obstacle/face segmentation route. It includes Size and shared blur controls plus two important protections:
Mouth/Lips Protection keeps an open mouth from being mistaken for an obstacle.
Face Obstacles Protection preserves genuine objects such as glasses, microphones, or a hand across the face when positive protection is required.
The optional XSeg Mouth sub-pass has its own size, blur, and mouth/upper-lip/lower-lip growth controls. Use it when the main XSeg mask is correct except around speech or teeth.
Text masking with CLIP
Enable Text Masking, enter comma-separated object descriptions, press Enter, and set Amount. This is useful when the protected item has a name—“microphone,” “glasses,” or “hand,” for example—but always verify several frames because text-guided segmentation can vary as an object moves or rotates.
The mask stack protects composition. Solve the narrowest problem first: border, profile, semantic obstacle, mouth sub-pass, then text guidance if needed.
Original-face parsers, eyes, mouth, tongue, and texture
The Original Face Parsers section selectively returns regions from the untouched target. Its FaceParser controls cover background, face, neck, hair, eyebrows, eyes, nose, mouth, lips, teeth, and related semantic areas, with amount and blur controls where appropriate. This is often the cleanest way to keep the target’s hairline, brows, teeth, or mouth cavity while replacing the identity around them.
Mouth Fit & Align
Mouth Fit & Align automatically anchors the restored original mouth inside the swapped face. Mouth Zoom scales that recovered or swapped mouth overlay around its anchor; it does not zoom the viewer. Use Original Mouth exposes cavity blur and shadow controls, while Smart Sharpen (USM) can restore perceived sharpness after compositing.
Restore Tongue preserves the target tongue, with a dedicated Tongue Edge Blur for a less cut-out transition. Paste After Restorer changes pipeline order when a face restorer would otherwise repaint the recovered mouth.
Restore Eyes and Restore Mouth
The dedicated Restore Eyes and Restore Mouth tools return those areas with individual blend and feather controls. They are ideal when identity is good but a model changes iris detail, eyelid shape, teeth, or speech articulation. Use the smallest blend that fixes the error; a full-strength patch can create a mismatched island.
Differencing and texture transfer
Differencing uses original-versus-swapped differences to limit where the result changes. Transfer Texture can reintroduce skin detail and offers a second mode, gamma, contrast, and CLAHE controls. Original and manipulated VGG feature masks, parser-based feature exclusions, and background exclusion let the texture stage focus on relevant facial information rather than spreading noise or background structure into the composite.
Mouth Fit, parser regions, and targeted eye, mouth, and tongue restoration recover only the areas that the swap actually damaged.
Color, artifacts, landmarks, blend, and re-aging
VisoMaster Fusion has both automatic and manual finishing layers. AutoColor estimates a target-consistent transfer, while Ending Color applies color treatment late in the pipeline. Manual RGB, brightness, contrast, saturation, sharpness, hue, gamma, and noise controls handle deliberate correction.
JPEG and MPEG tools can add compression characteristics, including block shifting, so a pristine generated face does not look pasted into heavily compressed footage. These are matching tools, not quality enhancers: add only enough degradation to meet the surrounding frame.
The landmark-correction section can scale the face and keypoint set, then offset the individual five-point landmarks. Use it for a persistent eye, nose, or mouth alignment error that masks cannot solve. Final Blend and Mask Blend then control how strongly the processed result and its mask integrate with the original.
Face Re-Aging exposes Source Age and Target Age. It can age or rejuvenate the identity while remaining inside the same per-face pipeline. After changing the age values, use Apply to recompute the source embedding and any K/V maps. Judge it on several expressions and lighting conditions; age effects that look persuasive on a frontal still may become too strong in motion.
Landmark correction, final blending, and optional re-aging provide precise late-stage control over geometry and age appearance.
Two-stage face restoration
The Restorers/Expressions tab contains Face Restorer 1 and Face Restorer 2. A two-stage design matters because restoration can serve different purposes at different moments: one pass can repair the swapped crop before expression or compositing work, while a restrained final pass can unify the finished face.
Both stages offer the current restorer family:
GFPGAN 1.4 and GFPGAN 1024.
CodeFormer.
GPEN at 256, 512, 1024, and 2048.
RestoreFormer++.
VQFR v2.
Alignment can use Original, Blend, or Reference behavior. Model-specific fidelity controls balance faithful reconstruction against stronger synthesis. Blend determines how much restored output survives. Auto Restore can adapt the amount, and its sharpness map targets areas that benefit from recovery instead of treating every pixel equally.
Restorer 2 can be placed at the end of the pipeline. That is useful after expression or denoising, but two strong passes can produce waxy skin, repeated-detail artifacts, or a face that is sharper than the surrounding footage. A disciplined recipe is one purposeful main restorer and, only when needed, a low-blend finishing pass.
Two restoration stages let repair and finishing happen at deliberate points instead of stacking two maximum-strength enhancers.
Expression restoration: Simple, Advanced, and Recast
Identity is only half of a face replacement. The target performance—eye openness, brows, lips, pose, gaze, and speech—must still feel connected to the body and shot. VisoMaster Fusion offers three expression modes and lets the expression stage run at the Beginning, After First Restorer, or After Second Restorer.
Simple mode
Simple mode exposes Neutral Factor, which controls how much target expression is added, and Expression Factor, which controls expression similarity and intensity. Animation Region limits the work to All, Eyes, or Lips, while Normalize Lips and its threshold stabilize mouth interpretation. Crop Scale and VY Ratio define the driving crop. Start here when the swap is structurally sound and needs only more faithful target expression.
Simple mode is the fast route: choose a region, set a restrained factor, and keep the target performance connected to the swap.
Advanced mode
Advanced mode separates eyes, eyebrows, lips, and general expression. Each family includes relative or neutral behavior and strength controls, with dedicated eye and lip retargeting and normalization. It also adds tools that are easy to miss in older documentation:
Camera Gaze Lock keeps the eyes oriented toward the camera, with Strength and Vertical Fine-Tune.
Micro-Expression Boost restores subtle movement that can disappear during a strong identity swap.
Relative Lids + Retargeted Gaze coordinates eyelid response with gaze changes so the eyes do not look mechanically rotated.
Treat eyes as a temporal problem, not a single-frame problem. A gaze setting can look perfect in a still but unnatural through a blink or head turn. Preview a short loop containing neutral eyes, a blink, and the largest gaze change.
Auto Mouth
Auto Mouth can trigger from the Expression Restorer or from the Face Parser alone. Confidence and EMA smoothing determine when and how steadily it engages; Strength and Region control the correction. Normalization, parser overrides, Exclude Upper Teeth, and a debug outline help diagnose speech frames without hiding the mechanism.
Use the debug outline briefly to confirm the selected region, then turn it off for final capture and output. Excluding upper teeth is helpful when the source correction would otherwise repaint stable target teeth.
Advanced mode separates expression systems and adds temporal gaze and mouth logic for shots that must survive motion, blinks, and speech.
Recast mode
Recast is a dedicated expression-driving path that uses the PerformRecast model with Enhancement and Replacement behavior. It includes expression strength, its own crop scale, animation region, eye and lip driving weights, temporal smoothing, and paste-back feathering.
Choose Enhancement when the existing performance is fundamentally correct and needs reinforcement. Choose Replacement when the driving expression should take a more dominant role. Crop scale and paste-back feather are as important as strength: they determine whether the driven face aligns and blends cleanly with the untouched frame.
Recast is a deliberate driving workflow, with independent crop, region, weight, smoothing, and paste-back controls.
Face Editor: pose, expression, gaze, and makeup
The Face Editor uses a LivePortrait-based editing path. To activate it, select a target face, enable Enable Face Pose/Expression Editor in the tab, and turn on the main Edit Faces button below the viewer; both switches must be on. It is not just a beauty panel: it can reposition and rescale the working crop, change pose, open or close eyes and lips, translate the face in three axes, and shape individual expressions.
The geometric controls include Position, Crop Scale, vertical crop offset, and crop blur. Pipeline Position can run the editor at the Beginning, After First Restorer, After Texture Transfer, or After Second Restorer. The performance controls cover eye and lip openness; pitch, yaw, and roll; X, Y, and Z movement; pout, purse, grin, smile, wink, brows, and gaze. Because several edits influence the same anatomy, make one family of adjustments at a time.
Makeup controls apply RGB color and blend values to the face, hair, eyebrows, and lips. They are useful for controlled stylization or shot matching, but high blend values can erase natural luminance variation. Sample changes at 100% view and in motion.
The Face Editor can reshape performance and styling after identity assignment without leaving the VisoMaster workspace.
ReF-LDM denoiser
Unlike Face Swap, Restorers/Expressions, and Face Editor, the Denoiser is global. Its reference K/V map comes from the selected target’s assigned source face or embedding; an embedding can store generated K/V maps for reuse. Exclusive Reference Path is a toggle that forces attention toward that reference K/V data, not a field for entering an image path. Base Seed makes comparisons deterministic. The denoiser can run in three places: Before Restorers, After First Restorer, or After All Restorers. Each pass can use Single Step or DDIM.
Single Step is the faster, simpler cleanup route and exposes a timestep plus latent sharpening.
DDIM adds iterative Steps and CFG guidance alongside timestep and latent sharpening, trading speed for a more involved diffusion pass.
Keep the same source assignment, reference toggle, and seed while comparing settings so changes come from the control you touched rather than a new random trajectory. Start with one pass. Multiple denoiser insertions can create a beautifully clean still that no longer tracks naturally through video.
The global denoiser is reproducible when its assigned reference K/V map and seed stay fixed. Place it where the pipeline needs cleanup, not everywhere at once.
Detection, landmarks, and ByteTrack
Face quality begins before the swap model. Face Detect Model provides RetinaFace, Yolov8, SCRFD, and Yunet, along with detection score, maximum face count, and rotation handling. A difficult face may need a different detector or rotated search rather than a lower global threshold.
Recognition Model is separate: it selects the Inswapper128ArcFace, SimSwapArcFace, GhostArcFace, or CSCSArcFace family used for identity similarity. If matching is unstable, verify this family as well as the detector and Similarity Threshold. Landmark choices include 5, 68, 3D-68, 98, 106, 203, and 478-point models. Associated controls set landmark confidence, extraction from points, mean-eye behavior, and keypoint smoothing. Show Landmarks, Show Bounding Boxes, and the related overlays make alignment failures visible instead of mysterious.
ByteTrack settings maintain identity continuity through motion and brief uncertainty. Tracking is especially helpful when the detector fluctuates across frames, but an incorrect track can carry a mistake forward. Inspect cuts, crossings, and re-entries; those are the moments where identity association is hardest.
Diagnostic overlays turn detection, alignment, and tracking into visible evidence. Disable them before final output.
Timeline, markers, issue scans, and frame control
The timeline is the production spine for video. Standard markers snapshot both per-face parameters and global controls. Playback and recording apply those states at marker frames, allowing a profile, lighting change, or short occlusion to receive a different pipeline. Scrubbing or jumping also updates the visible controls to the nearest marker at or before the destination when Track Markers on Video Seek is enabled.
Record Start and Record End markers define output ranges. Multiple marker pairs can render several selected segments without cutting the source in another editor. The transport shortcuts are designed for frame-accurate work:
ShortcutActionSpacePlay or pauseSToggle swappingRStart or stop recordingC / VStep one frame backward / forwardA / DJump by the configured skip amountZGo to the beginningFAdd a markerAlt+FRemove the current markerQ / WPrevious / next markerXToggle Face CompareTToggle Theatre modeCtrl+0 / Ctrl+1Fit view / 100% viewF11Full screenEscLeave Theatre mode or full screenMouse wheel over viewerZoom the viewer
Scan Tools does not judge aesthetics or general render quality. It walks eligible frame ranges and marks frames where each configured target identity cannot be matched under the current detector and recognizer settings; an unreadable frame marks every target. Record Start/End ranges constrain the scan, and saved settings markers are applied as it progresses. Start or abort the scan, move to the previous or next issue, drop or restore the current frame, drop all issue frames, clear the list, or restore dropped frames. Issue scanning is unavailable in VR180 mode. Document dropped-frame decisions when continuity matters because discarding frames can change timing.
Markers make per-face and global settings time-aware; Issue Scan flags target-identity matching failures inside eligible ranges before a long render.
Presets, workspaces, and context actions
Presets store reusable parameters. Apply one with the Apply button or a confirmed double-click. Face parameters and global controls are intentionally separate: enable Apply Settings when a preset should also change global Denoiser and Settings values. Right-click a preset to rename or delete it; deletion sends its file to the recycle bin.
Workspaces save the broader project: media paths, assignments, controls, markers, and other session state. On a clean exit, the app writes last_workspace.json and normally asks whether to reopen it next launch; Auto Load Last Workspace skips that prompt. Auto Save Workspace is separate and writes a workspace into the output folder after recording. Auto-load is convenient on a dedicated workstation; disable it when one machine moves frequently among unrelated or confidential projects.
Target-face context menus provide fast transfer between people or shots:
Copy and Paste Parameters.
Save Parameters and Settings.
Load Parameters.
Load Parameters and Settings.
Adjust thumbnail display.
Jump to a source frame.
Remove the face card.
This distinction is important: Face Swap, Restorers/Expressions, and Face Editor parameters belong to the selected target face, while Denoiser and Settings are global. Settings controls execution, detection, output, playback, recording, and device behavior.
The Presets panel stores reusable looks; Apply Settings decides whether global Denoiser and Settings values travel with the selected face parameters.
Job Manager and batch production
The Job Manager saves the current workspace as a job, loads or deletes saved jobs, refreshes the list, and processes all or selected jobs. Before processing, the application validates referenced files. Completed jobs move into jobs/completed, making it easier to distinguish queued work from finished work.
Two direct batch actions cover work that does not need a saved queue. Batch Process Selected Face (For Videos and Images) applies the current selected target-face configuration across every image and video in Target Media; videos use the full recording path and honor markers. Batch Process All Faces (For Images only) detects every face in each image and applies the current parameters and selected sources. Set an output folder and select at least one input face or embedding before starting either action.
A reliable batch workflow is:
Finish and preview one representative frame.
Inspect a profile, expression extreme, occlusion, and lighting change.
Add settings markers if one static pipeline is insufficient.
Confirm output destination, SDR/HDR path, encoder preset, quality/AQ, FPS limit, and source-audio expectations.
Save the workspace as a clearly named job.
Load the job once to prove its paths and assignments survived.
Queue the remaining targets, then process selected jobs before committing the entire list.
Jobs turn a configured workspace into a queueable unit and validate missing inputs before unattended processing.
Playback, recording, output, and encoding
Playback settings cover custom FPS, buffering, looping, full-screen theatre behavior, skip-step size, audio volume, and audio delay. Custom playback FPS, live-sound volume, and audio delay are preview controls; they do not set the recorded frame rate or rewrite source audio. The Processing FPS readout beneath Play reports current processing throughput rather than source or output FPS.
Recording can ask for confirmation before stopping, cap maximum output FPS, resize 16:9 footage to 1920x1080, keep controls active during capture, and open the destination after completion. FFMpeg options reveals preset, quality, automatic-quality, Spatial AQ, and Temporal AQ rows; it is not a codec selector. SDR recording uses 10-bit HEVC NVENC with presets p1 through p7. HDR Encoding - Use on HDR videos only (CPU) switches to 10-bit CPU libx265 with the HDR preset list. Recording automatically attempts to extract and mux the source audio; if that merge fails, the app can retain a video-only output.
Choose output settings before a long job:
Match source FPS unless the deliverable has a specific cap.
Use a short representative segment to test the fixed HEVC path in the actual destination player or editor.
Check source synchronization before a long run, remembering that preview audio volume/delay do not modify the automatically muxed output audio.
Use automatic quality for a safe start, then make an explicit quality decision for a final master.
Do not upscale to 1080p merely to hide a weak face crop; fix the face pipeline first.
Output to Target Location is convenient for one-off work, while a dedicated output root plus Preserve Source Directory Structure is safer for large batches. Cluster Output by Source Name groups results under the selected merged embedding’s name.
Global Settings determine how the application executes, previews, records, and organizes the per-face pipelines built elsewhere.
Frame enhancement, webcams, virtual camera, and VR180
The frame enhancer runs after the face pipeline and offers RealESRGAN x2/x4, RealESR-General x4v3, BSRGAN x2/x4, UltraSharp x4, UltraMix x4, DDColor Artistic/standard, and DeOldify Artistic/Stable/Video. Blend mixes enhanced output back with the original. Upscalers, colorizers, and deoldifiers solve different problems; select the family that matches the source rather than treating the list as a quality ladder.
Webcam Settings choose maximum stream count, backend, resolution, and FPS. Send Frames to Virtual Camera routes the processed result to an external application through OBS or Unity Capture backends. Test latency, resolution, and frame stability in the receiving application before a live session.
VR180 mode includes eye selection, tiled detection, field of view, and crop-resolution controls. It changes geometry and performance enough to deserve its own test project. Confirm left/right eye behavior and seams on the intended headset or player; ordinary flat-frame issue scanning is unavailable in this mode.
The remaining global swap settings automate repeat work. Auto Swap starts swapping with the selected source faces or embeddings when new image or video target media loads. Keep Selected Input Faces / Embeddings preserves those source selections while switching targets. Swap Input Face only once limits each source to its highest-scoring face match instead of every match above the threshold. Maximum DFM Models to use caps resident DFM models for the available VRAM, while Embedding Merge Method combines multi-source assignments by Mean or Median.
Enhancement and live/immersive modes share the workspace but have different quality checks: texture for upscaling, latency for webcam, and geometry for VR180.
Performance and VRAM tuning
Execution providers
The current package exposes CUDA, TensorRT, TensorRT-Engine, and CPU. The audited 3.6.1 package defaults to CUDA, even though some project documentation still describes TensorRT as the default.
CUDA is the dependable NVIDIA baseline and the right first choice.
TensorRT can accelerate compatible ONNX inference after engine preparation, with build time and model compatibility as trade-offs.
TensorRT-Engine targets prepared engine workflows where applicable.
CPU is a compatibility path, not the recommended production mode.
Settings also controls execution threads, worker delay, whether models remain in VRAM, and the input resize target from 540p through 2160p. Resize Input Source rescales media before AI processing while preserving aspect ratio; Frame Worker Delay postpones seek-triggered work slightly to reduce GPU overload and stutter. Keep Controls Active leaves the interface editable while recording, and Track Markers on Video Seek makes manual seeking load the applicable marker state. Keep Loaded Models in Memory improves repeated processing when VRAM is plentiful; the Clear VRAM button explicitly unloads cached models after settings or projects change.
Enable Mouse Wheel on Parameter Controls changes wheel behavior. When it is off, the wheel scrolls the parameter panel and holding Ctrl temporarily adjusts a hovered slider or dropdown. When it is on, the wheel adjusts hovered parameter controls directly, so disable it if accidental value changes are more costly than an extra modifier key.
The best tuning order is:
Prove the project with CUDA.
Choose a realistic input-resize ceiling.
Keep only repeatedly used models resident.
Measure a representative segment, not an easy still.
Try TensorRT only after the CUDA result is correct.
Compare output pixels as well as frames per second.
Current measured benchmark
The refreshed benchmark uses the bundled target video and source face on the audited RTX 5090 Laptop installation. It records environment, provider availability, warm-up, processed frame count, elapsed time, FPS, output hashes, and visual comparison so acceleration is not presented without reproducibility.
The benchmark shown here was rerun on the audited 3.6.1 installation; exact figures belong to this hardware, driver, models, footage, and configuration.
Measured result: across 360 measured frames (120 frames × three passes), TensorRT reached 25.61 FPS while CUDA reached 22.77 FPS—a 1.125× / 12.5% throughput gain on this RTX 5090 Laptop GPU. Median per-frame latency fell from 42.55 ms to 37.66 ms. The trade-off was memory: measured peak GPU-memory growth was 1,965 MB for TensorRT versus 1,517 MB for CUDA, an extra 448 MB. These numbers describe this machine, driver, model cache, and workload; benchmark your own system before locking a production provider.
Treat that number as a local reference, not a universal promise. Face count, detector, swap model, resolution, masks, restorers, expression logic, denoising, enhancement, source resolution, encoder path, driver, and engine cache can all change throughput.
Real demo cases
The following comparison uses only the disclosed example video and faces installed with the application. Each processed panel is linked to a saved configuration in the refreshed Tests_and_Configs package so the result can be reproduced instead of treated as marketing decoration.
One source frame, one disclosed identity, several intentional pipelines: baseline, compositing repair, restoration, and expression-aware finishing.
Case 1: fast, clean baseline
Use Inswapper128, Auto Resolution, default border mask, no denoiser, and one restrained restoration pass only if the source truly needs it. This is the right starting preset for clean frontal footage and a strong source portrait.
Case 2: glasses, hands, microphones, or hair across the face
Start with the baseline, then enable Occluder or DFL XSeg. Add Face Obstacles Protection when the object must survive, Mouth/Lips Protection when speech is being cut away, and a small shared blur. Use CLIP text masking only when the semantic mask still misses a named object.
Case 3: speech and expressive close-ups
Restore the target mouth or use parser regions, enable Tongue restoration when visible, and test Auto Mouth with EMA smoothing. Advanced expression mode can stabilize gaze and preserve micro-expressions. Preview a loop containing teeth, tongue, lip closure, a blink, and a head turn.
Case 4: low-resolution or compressed material
Match identity and mask first, then use one restorer with modest blend. Transfer Texture can return useful high-frequency character; JPEG/MPEG matching can stop an over-clean face from floating above the source. Apply frame enhancement only after the composite looks correct at native size.
Case 5: multi-face batch
Scan the video, add only relevant identities, assign a source to each, and tune each target card independently. Use parameters-only copy/paste for faces that share a pipeline, then correct individual similarity thresholds and masks. Save the workspace, validate it, and queue it as a job.
Quality-control checklist
Before exporting a still or a long render, inspect:
Identity at frontal, three-quarter, and profile angles.
Detection through cuts, crossings, blur, and re-entry.
Hairline, jaw, ears, neck, and border seams.
Glasses, hands, microphones, hair, and other occlusions.
Eyes at neutral gaze, extreme gaze, and blinks.
Lips, teeth, tongue, and full mouth closure during speech.
Color and exposure across lighting changes.
Texture at 100% view; watch for waxiness and repeated detail.
Temporal stability during head motion and expression changes.
Audio sync, output FPS, resolution, HEVC playback compatibility, and HDR/SDR behavior.
Marker transitions and every record start/end range.
A short encoded sample in the actual destination player or editor.
Theatre mode (T) is excellent for a final uninterrupted review. It can force full screen, while the normal full-screen command remains available with F11.
Theatre mode removes the control-room distraction. Watch the finished motion, not just a favorite frame.
Troubleshooting by symptom
No face is detected
Try another detector, lower Detect Score cautiously, enable relevant rotation search, increase the maximum-face count, and test a frame where the face is larger. Use Show Bounding Boxes and landmark overlays to see whether failure happens at detection or alignment.
The wrong person is swapped
Raise Similarity Threshold, verify the selected target card and source assignment, and confirm that Recognition Model matches the identity family you expect. Rescan the footage and review ByteTrack behavior at crossings or cuts. Swap Input Face only once may not suit a scene where identities enter and leave repeatedly.
The face looks right but the edges fail
Adjust the relevant Border side, use Profile Angle Mask for turned heads, and add a small mask blur. If an object crosses the face, use Occluder or XSeg rather than trying to hide it with a wider border.
Glasses, a hand, or a microphone disappears
Enable an occlusion-aware mask. With XSeg, try Face Obstacles Protection and tune Size/Blur. For a named object that remains missed, test CLIP Text Masking and verify it across motion.
Mouth or teeth look synthetic
Restore original mouth/parser regions, use Mouth Fit & Align, consider Paste After Restorer, and test Auto Mouth. Exclude upper teeth when the correction should not repaint them. Avoid a high-strength face restorer over the recovered mouth.
Eyes jitter or stare unnaturally
Reduce eye retargeting, use Relative Lids + Retargeted Gaze to couple eyelid motion to eye direction, and lower Camera Gaze Lock strength or Eyes Multiplier. Review a moving sequence rather than one frame.
The result is waxy
Lower restorer blend or disable the second restorer, reduce repeated swap Strength, return texture selectively, and avoid stacking denoising and frame enhancement. Compare at 100% rather than judging a scaled preview.
Processing is unexpectedly slow
Confirm the execution provider and GPU availability, lower input resize, disable expensive stages one at a time, unload unused cached models, and benchmark the same segment. A high-resolution swapper, two restorers, DDIM denoising, exhaustive detection, and frame enhancement are cumulative costs.
Models are missing or the app no longer starts after an update
Use the launcher’s Check/Update Models and Check/Update Dependencies actions. Repair Installation backs up changed tracked files before restoring the checkout. Rollback and branch switching can replace tracked local edits, so keep project workspaces and custom assets backed up outside the application tree.
A production recipe that scales
For repeatable work, separate creative approval from long rendering:
Ingest: organize targets, sources, and output roots; decide whether folder recursion is safe.
Identify: scan footage, review candidates, create or select embeddings, and assign every relevant person.
Baseline: choose the swap model and alignment; prove identity with all optional finishing disabled.
Composite: solve borders, profiles, obstacles, mouth, eyes, and parser regions.
Perform: add expression, gaze, or Face Editor changes only where the shot needs them.
Finish: restore, denoise, texture-match, color-match, and enhance in restrained stages.
Temporal QA: inspect representative motion and run issue scans where supported.
Encode test: render a short range with the final FPS limit, SDR/HDR path, preset, quality/AQ, and source audio.
Package: save a preset for the look, a workspace for the project, and a job for the queue.
Render and review: process selected jobs first, then inspect the encoded deliverable in its destination environment.
This method is slower than randomly moving sliders for the first five minutes and much faster than discovering a systematic mouth or tracking failure after a two-hour render.
Final verdict
VisoMaster Fusion succeeds because it treats face replacement as a production pipeline rather than a novelty effect. It gives newcomers a direct path from media to result, then keeps revealing depth: multiple identity engines, per-face configuration, semantic masking, two-stage restoration, three expression systems, pose and gaze editing, reference denoising, tracking, frame enhancement, timeline automation, jobs, live output, and VR180.
Its density is a strength once you adopt one rule: enable a tool only to solve a visible problem. Build identity first, protect composition second, preserve performance third, and finish the pixels last. The application’s compare modes, overlays, markers, saved state, and job system make that discipline practical.
If you want a local studio where a quick swap can grow into a controlled multi-shot workflow without exporting every problem to a different tool, VisoMaster Fusion v7 deserves a serious test.