AI Video Face Swap, Frame by Frame
Upload one source face and the clip you want it in. Every frame gets its own detection, landmark and blending pass, so the video face swap holds through motion instead of strobing.
A still image hands the model one problem. A video face swap hands it one problem per frame, then a harder problem on top: the frames have to agree with each other. Twenty four swaps that are each convincing alone will still read as fake if the jawline shifts a pixel between them, because the eye catches movement far better than it judges detail. That temporal demand is why a clip costs more, takes longer, and rewards preparation in a way photo work never does.
Under the hood each frame is decoded, scanned for a face, landmarked, swapped, blended back and written into a new file. FaceFusion does the work on RunPod GPUs, which is why a video face swap usually takes one to five minutes rather than the few seconds a still needs. Nothing carries between frames by default: every frame is detected on its own, which is why the pipeline survives hard cuts and why unstable detection surfaces as flicker.
There are no preset faces here and no library of real people to choose from. You upload the source face as an image and the target as an image or a video, so both halves are yours to answer for. A video face swap costs 30 credits flat, whatever the clip length, while an image costs 10. Free accounts get 10 welcome credits once and 10 more each day, so a clip needs a paid plan or several days of saved credits.
What the Video Face Swap Pipeline Actually Does
Detection on every single frame
Each frame gets its own detection, landmark and blending pass instead of a track interpolated forward from the first one. Hard cuts and occlusions do not derail the sequence, because frame 400 of a video face swap knows nothing about frame 399.
Controls that stop the strobing
Flicker is a detection problem, not a blending problem. Lower the detector score so borderline frames still register, raise the detector size for small faces, and set the selector to reference mode with a fixed reference frame number to lock one person across the whole video face swap.
Encoder level output control
The advanced panel exposes the output video encoder, preset, quality, scale and FPS, so a video face swap never leaves the pipeline more compressed than it arrived. Output audio encoder and volume sit in the same panel, because the sound passes straight through.
Trim before you spend
Trim frame start and trim frame end cut the material down to the shot you need. Fewer frames means less processing for the same 30 credits, and a tightly trimmed video face swap comes back in about a minute.
Enhancers that scale with frame count
Face enhance rebuilds texture inside the swapped region and frame enhance upscales the whole clip. Neither costs extra credits. Both cost extra time per frame, and on a video face swap that multiplies by the frame count, so switch them on deliberately.
How to Run a Video Face Swap in 3 Steps
Upload the source face
The source has to be a still image up to 10MB: one clear front facing face, both eyes visible. It sets the identity for every frame that follows, so sharpness here pays off across the whole clip.
Upload and trim the target
Drop in an MP4 or MOV up to 100MB. Set trim frame start and end before you launch the video face swap, since the charge is identical whether you process eight seconds or eighty.
Run it and wait
Most clips land in one to five minutes. The page polls every three seconds and stops waiting at fifteen minutes. Download the file at full resolution, with the original audio intact and no watermark.
A video face swap costs 30 credits against the 10 an image costs, so a free account needs three days of daily credits, or a paid plan.
Where a Clip Beats a Still
The format decides a great deal. These are the jobs where a video face swap earns the extra credits and the extra wait, all of them starting with footage you own.
Dialogue and reaction shots
A face that has to speak, blink and react gives away far more than a portrait does, which is why a video face swap is the more persuasive version when it lands. Keep the take short and the camera steady.
Look tests for production
Casting mockups and pitch reels need an actor moving, not standing still. A rough video face swap answers the question in five minutes and reads more honestly than a painted concept frame.
Fan edits and recuts
Recasting a scene, finishing a cosplay reel or assembling the crossover trailer nobody funded is the highest volume use of this format. Trim to the shot that carries the joke rather than processing a whole episode.
Short form social clips
Vertical video moves fast and forgives a lot. A nine second cut is a fraction of the frames a full scene demands, which keeps processing time down while the video face swap still carries the gag.
Your own face, somewhere else
The least glamorous and most common job: putting your own face into footage you were not standing in. Travel clips you missed, a group video you skipped, a take you wish you had been around for.
Working With Video, Not Against It
Trim and frame before you spend credits
Thirty credits buys one job, not one second, so the first decision is which frames you genuinely need. Trim frame start and trim frame end exist for this: pull the material down to the shot that carries the idea before you launch the video face swap. A ten second cut comes back in roughly a minute, while a two minute scene edges towards the fifteen minute polling limit and costs exactly the same. Trim first, then think about parameters.
Framing matters more here than on stills, because the face changes size as the subject moves. A face that fills the frame at the top of the shot and shrinks to forty pixels by the end swaps cleanly for half the clip and comes apart for the rest. Either trim to the stretch where the face stays a usable size, or raise the detector size so small faces still register. A wide establishing shot is the hardest material you can hand the model.
Why swaps flicker, and how to stop it
Flicker almost never originates in the blending stage. It comes from detection: on a handful of frames the face drops below the detector score threshold, that frame passes through unswapped, and the sequence strobes. Lower the detector score and those borderline frames come back into line. Motion blur, sharp profile turns and shadow across the face are the three usual causes, and all three answer to a lower threshold rather than a different swapper model.
The second source of instability is identity drift in a shot holding more than one person. With the selector left on many faces, the engine can hand the swap to somebody else the moment your subject leaves frame. Move the selector to reference mode, set the reference frame number to a frame where your subject is unmistakable, then tune reference position and distance so the match stays pinned for the length of the video face swap.
Do not let the encoder undo the work
Every video face swap re-encodes. The pipeline decodes the source, rewrites the pixels and writes out a new file, which means a second generation of lossy compression on top of whatever the original already carried. Output video encoder, preset and quality decide how much that second pass costs you. Leaving quality low on footage that started clean is the quickest way to make a technically perfect swap look worse than the material you began with.
Frame rate deserves the same care. Output FPS follows the source unless you change it, and changing it forces a resample, which can introduce judder that reads as a swap artefact even when the face work is flawless. Leave FPS alone without a concrete reason. If the source really is soft, frame enhance upscales the clip and face enhance rebuilds texture in the swapped region, paying in processing time per frame rather than credits.
Check This Before You Spend 30 Credits
A disappointing video face swap is nearly always an input problem. The model is the same one that runs on stills and it does not degrade between jobs. What changes is the footage: how far the face is turned, how much of the frame it fills, how heavy the motion blur gets, and how many people compete for the detector.
Because a clip costs three times what an image costs and takes a hundred times as long, a wasted run hurts. Spend fifteen seconds on the list below. It is cheaper to trim the shot or pick a different source photo now than to sit through the same video face swap twice.
- Source face is a sharp, front facing still under 10MB
- Target clip is MP4 or MOV and under 100MB
- Trim frame start and end cover only the shot you need
- The face stays a usable size across the trimmed range
- Selector is on reference mode if several people appear
- You hold the rights to the footage and consent for the face
Rules That Apply to Every Clip
There is no roster of real people on this page and there never will be. A one click video face swap built around a named public figure is how a tool like this becomes a harassment machine, so the source face is always something you supply and always something you answer for. Sexual content using a real person's likeness without consent is prohibited outright. So is output built to pass as a genuine, unedited recording of something that never happened, and moving footage is the highest risk format for exactly that, so it is worth stating plainly.
Parody, commentary, fan work, pitch and production material, and your own likeness all sit inside what this is for. Label edited media as edited when it leaves your hands. The law has moved fast: likeness and publicity rights, non consensual intimate imagery statutes and synthetic media rules around elections all carry real penalties today. Accounts used to target, sexualise or defraud a real person lose access, and we honour takedown requests from the people depicted in a video face swap.
Video Face Swap FAQ
Why does a video face swap cost 30 credits when an image costs 10?
Because the work is per frame. One second at 30fps is thirty detections, thirty landmark passes, thirty blends and a full re-encode, against one of each for a still. The 30 credits are flat for any clip length, so a long clip is better value per frame, and daily free credits alone will not cover one run.
Does the voice change?
No. A video face swap changes the face and nothing else. Audio passes through untouched, and the only sound settings are the output audio encoder and the volume. There is no voice cloning, no lip sync and no speech generation in this pipeline.
How long does a video face swap take?
One to five minutes for most material, driven by frame count rather than file size. The page polls every three seconds and gives up after fifteen minutes, so very long or high frame rate footage is better trimmed first.
What files can I upload?
The source face must be an image, up to 10MB. The target can be an image up to 10MB or a video up to 100MB in MP4 or MOV. Anything else is rejected before it costs a credit.
Why does my video face swap flicker?
The detector is losing the face on some frames. Lower the detector score first, then raise the detector size if the face sits small in frame. If the swap jumps to another person, move the selector to reference mode and pin it with a reference frame number.
Is there a watermark on the output?
No, on any plan. The file comes back at the resolution and scale you set, with no overlay. Uploads sit in presigned private storage scoped to your account, results stay in your job history, and nothing is published or used as training data.
Can I test with a still image first?
Yes, and you should. The same tool takes an image target for 10 credits, the cheap way to find out whether a source face works before you commit to a video face swap. Get the identity right on one frame, then run the clip.
Other Face Swap Tools
The same video face swap engine, pointed at a different kind of source material.
Start a Video Face Swap
Upload a face, trim your clip, and run it. Thirty credits covers one video face swap of any length.
Swap a Face Now