Skip to content
Get App

Dear Higgsfield, I reverse engineered you in 12 hours

Sirio reverse engineers Higgsfield's Genjutsu in 12 hours, showing how to bypass IP blocks and create AI face swaps safely.

Key Takeaways

  • AI face-swapping apps rely on workflows that hide identity but preserve motion to bypass IP filters.
  • Depth masks and vocal isolation are key techniques to avoid detection by platforms like Seedance 2.5.
  • Understanding the underlying AI workflow empowers users to create custom face-swap videos.
  • Open-source tools and APIs can be combined to replicate and improve commercial AI face-swap apps.
  • Ethical and legal considerations are important; this tutorial is for educational use only.

What the video covers

  • Sirio demonstrates how to reverse engineer Higgsfield's Genjutsu face-swapping AI in 12 hours.
  • The video explains how to bypass IP blocks in Seedance 2.5 to insert yourself into any video scene.
  • Sirio emphasizes understanding the workflow behind AI models rather than just using wrapper apps.
  • The tutorial covers masking techniques like depth masks to hide identity while preserving motion.
  • Audio processing is addressed by isolating vocals and removing background music to avoid detection.
  • Sirio uses open-source models and APIs such as Replicate, Foul, and Google MediaPipe for face tracking.
  • The importance of separating 'who' (identity) from 'how' (movement) in videos is explained.
  • The video includes practical steps with API key management and setting up environment files.
  • Sirio encourages viewers to build their own AI workflows and hints at future advanced tutorials.
  • The video is intended for educational purposes only and warns about IP and ethical considerations.

Answers

Questions about this video

What is the main goal of this video?

The video aims to teach viewers how to reverse engineer Higgsfield's Genjutsu AI to bypass IP blocks and create custom face-swapped videos using open-source tools.

How does the tutorial avoid IP blocking in face swap videos?

It hides the identity ('who') by using depth masks and isolating vocals while preserving motion ('how'), which prevents platforms like Seedance 2.5 from detecting copyrighted faces or music.

Is special software required to follow this tutorial?

No special proprietary software is needed; the tutorial uses open-source models, APIs like Replicate and Google MediaPipe, and explains the workflow so viewers can build their own solutions.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
Hi, welcome to this masterclass. My name is Sirio, and this is entirely AI. I'll show you. So that's me in a scene that I was never actually in, and I didn't even use any app to make it. And this is what actually happens when you try to do it.
00:15
Speaker A
the normal way. I'm uploading that same template video into Seedance 2.5 and asking you to swap me in. Blocked. But apps like Higgsfield's Genjutsu do this exact thing all day, every day with these exact videos.
00:30
Speaker A
The normal way. I'm uploading that same template video into Seedance 2.5 and asking you to swap me in. Blocked. But apps like Higgsfield's Genjutsu do this exact thing all day, every day with these exact videos.
00:53
Speaker A
any special software. A quick heads up before we start. This video is for educational purposes only.
00:58
Speaker A
So how the f*** are they getting around it? Hey friend, welcome back. Today we're going to reverse engineer Higgsfield's Genjutsu and we're going to get around the IP blocks in Seedance 2.5 so that you guys can put yourself into pretty much any movie scene or talking head you want without.
01:11
Speaker A
And by they, you know what I'm talking about. Okay, so no fluff. Here's how this is going to work.
01:16
Speaker A
Any special software. A quick heads up before we start. This video is for educational purposes only.
01:25
Speaker A
At every step along the way, I'm going to pause and tell you why it works and not just what to click. And then I'm going to give you the AI or the automatic version, the agent version, though the workflow so that you can turn it into a template for
01:39
Speaker A
I basically want you to understand how these AI models actually work and what they can really do. And this might end up being the first episode of a new series where I actually show you guys the stuff that they do not want you to know.
01:49
Speaker A
I'm Cereo and this is your time to build something of your own. Cut. Good job.
01:57
Speaker A
And by they, you know what I'm talking about. Okay, so no fluff. Here's how this is going to work.
02:01
Speaker A
So let's not waste any more time. You guys have all seen these videos that are basically everywhere right now.
02:07
Speaker A
We're going to start by doing the whole thing by hand. Okay, it's actually going to break on us twice and we're going to fix both times.
02:19
Speaker A
And most people get this from one of those wrapper apps that sells it as a feature.
02:24
Speaker A
At every step along the way, I'm going to pause and tell you why it works and not just what to click. And then I'm going to give you the AI or the automatic version, the agent version, though the workflow so that you can turn it into a template for.
02:34
Speaker A
It is not the AI model. It's the workflow. It's the thinking. They are selling you the thinking that they do for you and you just press a button. And so this tutorial is for those people who actually want to
02:46
Speaker A
Your own stuff. And in the next videos, we're going a lot further and we're going to make things like this. I'm Kim Kardashian and this is masterclass.
03:01
Speaker A
So let's go back to that blog generation from the start. This is the template. And all I did here was I head over to CDense 2.5.
03:10
Speaker A
I'm Cereo and this is your time to build something of your own. Cut. Good job.
03:20
Speaker A
But like I said, this exact video works totally fine in Genjutsu. But before we actually build anything, I think that it is important for us to understand why because if you get this part, the rest of the video is basically just
03:32
Speaker A
So what do you think? Was it any good? Man, you killed it. But first the basics.
03:39
Speaker A
And the way that I like to think about this is that the video is really carrying these two things at the same time.
03:45
Speaker A
So let's not waste any more time. You guys have all seen these videos that are basically everywhere right now.
03:59
Speaker A
movement. So the whole idea here is that you keep the how, the movement and you throw away the who, the IP and everything that we're going to do from here on out is really just hide one of those two things, obviously without losing the movement.
04:12
Speaker A
And the way that it works is that you upload a photo of yourself and then you pick a template video and then it just swaps you in for the person that is in that template video. It's the same moves, the same everything, the same lip syncing.
04:25
Speaker A
You can find it on hugging phase or replicate or foul. And all you have to do is that you upload your video, you click generate, and then you get something like this. So this is what's called the depth mask.
04:37
Speaker A
And most people get this from one of those wrapper apps that sells it as a feature.
04:47
Speaker A
Like it's the depth. If you've used Photoshop or or you've done retouching before, it's kind of very similar to dodge and burn.
04:55
Speaker A
And like I've said before, that's really the whole reason that the wrapper works and good for them. What I mean by that is that they are selling to you an API workflow.
05:03
Speaker A
This one is from foul. And this one is from replicate. I like the one from replicate. Just, I don't know.
05:10
Speaker A
It is not the AI model. It's the workflow. It's the thinking. They are selling you the thinking that they do for you and you just press a button. And so this tutorial is for those people who actually want to.
05:14
Speaker A
And there's one more thing, which is the audio, right? If we want the lips to match the song, then the audio has to go through as well. Now, the first thing that most people would say is just muted,
05:27
Speaker A
Build that button. I spend the last 12 to 24 hours on this going through open source models, masking, cutting people out and a whole bunch of other subs that didn't work. And I ended up with the one that I'm going to show you today.
05:36
Speaker A
So instead of muting it, we're going to keep the voice, but get rid of everything else. Keep the actual words, remove the soundtrack, remove the background music.
05:46
Speaker A
So let's go back to that blog generation from the start. This is the template. And all I did here was I head over to CDense 2.5.
05:53
Speaker A
And then I'm going to click on it. And over here on the right hand side, I'm going to go to audio, then voice changer, then voice filters, then I'm going to pick pitch up.
06:05
Speaker A
I uploaded this exact video and I typed replace the character in the video with the picture that I uploaded. And it got blocked for IP, which I mean, sure, it makes sense. These people are famous.
06:09
Speaker A
And then once that's done, I'm going back to the basics. I'm going to click on isolate voice, then keep only the vocals.
06:16
Speaker A
But like I said, this exact video works totally fine in Genjutsu. But before we actually build anything, I think that it is important for us to understand why because if you get this part, the rest of the video is basically just.
06:29
Speaker A
So that's both of them hidden. Right. C dense 2.5 isn't going to block this anymore because there are no faces.
06:37
Speaker A
Just details. There are really only two things that get you blocked here. The first one is the face. The second one is the music.
06:50
Speaker A
But there's one problem though, which is that the depth mask has no idea what the place actually looks like. It just either like a grayscale version of the depth mask like in FAL or like this other one here.
07:06
Speaker A
And the way that I like to think about this is that the video is really carrying these two things at the same time.
07:09
Speaker A
We need an API key. So I'm going to grab my enhancer API key for C dense 2.5 so we can test. And the link for that is in the description.
07:19
Speaker A
It's carrying who's in it. So the face, the voice, the music, and it's carrying how things move. So the motion, the timing and the lips and the filter, the IP filter only cares about the who. It does not care about the how and the how is the.
07:25
Speaker A
So your stuff is safe. I'm going to copy the docs right here. And then I'm going to paste them into codex or Claude.
07:34
Speaker A
Movement. So the whole idea here is that you keep the how, the movement and you throw away the who, the IP and everything that we're going to do from here on out is really just hide one of those two things, obviously without losing the movement.
07:46
Speaker A
And I'm going to head over to my enhancer API dashboard. I'm going to generate a key, copy it and I'm going to go back to the agent and I'm going to ask it to make an ENV file, which is just like a little private file in
08:00
Speaker A
So the first thing that we're going to have to do is turn this template into something where the AI cannot tell who the people are, but he can still see how they move. And for that, we're going to use something that is called depth anything.
08:09
Speaker A
Always put them in an ENV file. I'm going to click on here and I'm going to paste in my API key and we're all set.
08:16
Speaker A
You can find it on Hugging Face or Replicate or Foul. And all you have to do is that you upload your video, you click generate, and then you get something like this. So this is what's called the depth mask.
08:31
Speaker A
video first over and over until you like what you see. And then it finishes that exact same video in 1080p so you don't have to lose any detail. And it's not like an upscaler.
08:43
Speaker A
And what I mean by this is that it throws away what everything looks like and it keeps how close or how far things are from the camera.
08:49
Speaker A
And then when you're ready, you're like, cool, I'm going to pay for the paint once.
08:53
Speaker A
Like it's the depth. If you've used Photoshop or you've done retouching before, it's kind of very similar to dodge and burn.
09:00
Speaker A
What a lot of people will do is they're going to download it and then upload it right back as the reference video thinking that's the same thing it is not. And the reason why is that the draft is tied to a task ID and not
09:16
Speaker A
But in this case, it tracks all the movements, the movements are still there. But the faces, as you can see, they're not just a silhouette.
09:25
Speaker A
So there's no promises that you're going to get the same video back. So basically you do all of the cheap tries in draft mode and then you only pay for 1080p one time. Hopefully this makes sense.
09:37
Speaker A
This one is from Foul. And this one is from Replicate. I like the one from Replicate. Just, I don't know.
09:46
Speaker A
And if you guys haven't met teen yet, teen is this little creature that I've made with AI basically my companion.
09:53
Speaker A
I just like the colors of it. There's no difference. You can use the one from Foul too.
10:05
Speaker A
And then I'm just basically going to say, hey, use the draft mode for this generation.
10:11
Speaker A
And there's one more thing, which is the audio, right? If we want the lips to match the song, then the audio has to go through as well. Now, the first thing that most people would say is just muted,
10:19
Speaker A
It's going to host your inputs temporarily and then you're going to hit enter. And this usually takes about one minute.
10:24
Speaker A
but we cannot do that. And the reason why is that AI actually needs to hear the voice in order for it to move the mouth.
10:34
Speaker A
And it will be sitting right there with your request ID. That's also going to match your agents.
10:40
Speaker A
So instead of muting it, we're going to keep the voice, but get rid of everything else. Keep the actual words, remove the soundtrack, remove the background music.
10:50
Speaker A
So that's my face, but that it's very much the actor's body with all the long hair. And the second thing is that the whole thing is just gray. So one way to fix this is to have the agent describe
11:07
Speaker A
So what I'm going to do is that I'm going to take this depth mask and bring it into CapCut. And I'm just going to make a new project and drag the clip in.
11:21
Speaker A
And what I'm going to do is I'm going to take a screenshot of the original video and I'm going to drop it into the agent.
11:26
Speaker A
And then I'm going to click on it. And over here on the right-hand side, I'm going to go to audio, then voice changer, then voice filters, then I'm going to pick pitch up.
11:36
Speaker A
So that's the problem that I'm using. You can pause and screenshot this. Since we're doing this with the agent, the agent knows exactly what they send before and they can just update that request or update that payload, you don't have to do
11:47
Speaker A
This is a voice filter. And what it does, it just changes the pitch of the music.
11:52
Speaker A
Okay, now we have caller, which is nice. That's one of our two problems fixed.
11:59
Speaker A
And then once that's done, I'm going back to the basics. I'm going to click on isolate voice, then keep only the vocals.
12:12
Speaker A
I'm not entirely sure why. I think it's basically, I don't know, I think it just sees an outline and then it traces it.
12:18
Speaker A
And so now what's left is basically only the acapella without the music or soundtrack. So listen to it.
12:31
Speaker A
to put in the depth mask right on top of it. I'm going to mute the original clip.
12:36
Speaker A
So that's both of them hidden. Right. C Dense 2.5 isn't going to block this anymore because there are no faces.
12:50
Speaker A
Basically, what this is doing is that it's cutting out just the people and then it's laying them on top of the original video.
12:57
Speaker A
There's no music and we still have all of the movement, right? Pretty neat. Now, instead of giving it the original video, we're going to give it the depth mask video plus our own photos and then swap the people in.
13:08
Speaker A
So it was raising the outline before and now we are giving it something that that is totally different and it stops from tracing it.
13:14
Speaker A
But there's one problem though, which is that the depth mask has no idea what the place actually looks like. It just either like a grayscale version of the depth mask like in Foul or like this other one here.
13:19
Speaker A
Then I'm going to send it back to the agent and say, Hey, swap the old sorts video for this new one and make another draft.
13:26
Speaker A
And I think it's easier to show you guys. So let's run a quick test in order to do that.
13:34
Speaker A
We hit the face, we hit the music, we described the room, and then we laid the people over the real video. And that's way too much manual work to do every single time. But but you can do all of it with one click the same way Higgsfield does
13:49
Speaker A
We need an API key. So I'm going to grab my enhancer API key for C Dense 2.5 so we can test. And the link for that is in the description.
14:03
Speaker A
So really quickly it is 11 steps. One, we upload the original video. We keep the original video for also the final audio to demicus on replicate splits the vocals from the music.
14:16
Speaker A
And yes, it is cheaper than all of the other platforms that say that they are the cheapest and also doesn't train on your data.
14:31
Speaker A
Six, it puts those depth people over the original background. Seven, the pitched vocals go right inside of that video.
14:39
Speaker A
So your stuff is safe. I'm going to copy the docs right here. And then I'm going to paste them into Codex or Claude.
14:45
Speaker A
Ten, it throws away C dense audio and puts in the whole original soundtrack back.
14:50
Speaker A
I'm using Codex right now. That's really because I already had this project started in here. It does not mean that I'm endorsing OpenAI or anything. It's just easier for me right now. Hey, can you set this up for me?
15:00
Speaker A
Now you don't have to do any of that yourself. I'll show you all of it running as soon as it is set up.
15:07
Speaker A
And I'm going to head over to my enhancer API dashboard. I'm going to generate a key, copy it and I'm going to go back to the agent and I'm going to ask it to make an ENV file, which is just like a little private file in.
15:11
Speaker A
I built all of it for you guys. All you have to do is download the skill, give it to your agent and then plug in your API keys from enhancer and replicate so you can grab the skill in the
15:20
Speaker A
Your folder that holds your keys and saves the keys in there for safety. So please guys, never ever paste your keys in the chat.
15:33
Speaker A
questions along the way. And if you don't want to download the files for whatever reason, you can just ask your agent to build this workflow over here.
15:41
Speaker A
Always put them in an ENV file. I'm going to cli
15:52
Speaker A
The first one is you just drop in the video that you want to change and then the photos that you want to be placed in the video are like swapped and you say, Hey, run the workflow as simple as that.
16:07
Speaker A
The second one you do it yourself in the UI interface and that's what I'm going to do here. So you guys can see every step.
16:13
Speaker A
So the interface is really just four steps because you don't see everything that's happening in the back end. You have to upload, review, draft and export or save.
16:23
Speaker A
So I'm going to upload my video, which can be an MP4 or a MOV up to 30 seconds.
16:29
Speaker A
I'm going to leave it on advanced because that's the depth workflow we just built.
16:34
Speaker A
And then I'm going to click prepare. And so right now it's doing steps through one to seven all by itself.
16:40
Speaker A
You don't have to do anything. It's splitting the vocals from the music. It's pitching them up everything we did in cup cut.
16:47
Speaker A
And then it's making the depth video, cutting the people with Sam three laying them over the original background and cleaning up every little stray or edges that are off from that mask. This is step eight.
16:59
Speaker A
And it's not going to let me generate until I've watched the whole thing and I've approved it myself. And the reason for that is that might be a bad mask.
17:08
Speaker A
So we don't, I don't want you to basically just throw money down the drain. AI agents are unpredictable.
17:17
Speaker A
That's why we make sure that they check their work every time. So it's checking for holes in the body, the hair or hands.
17:24
Speaker A
This one to me looks very clean. So I'm going to prove this. And then I'm going to add my characters.
17:30
Speaker A
So character one, that's me. Character two is optional in this case, it's teen. And then the prompt is already written.
17:38
Speaker A
And you can edit that if you wish. And I'm going to generate my draft.
17:43
Speaker A
Cool. So here's our draft and listen to that. Take notes. That's the original audio on the pitched one.
17:55
Speaker A
And I'm pretty happy with this. So I'm going to approve that. And then it makes the 1080p and puts the original audio back one more time. And that one does cost extra.
18:08
Speaker A
And it's never going to do it without asking you first or approving. Now that workflow is the best one if you want to bulletproof your generations.
18:15
Speaker A
And it mostly uses the audio to match the words. So if your video has no audio, it will still generate for sure.
18:22
Speaker A
But there's one more workflow for when you want the face to move exactly like the original. And it's the same workflow, but we add what's called a face mash.
18:33
Speaker A
And what I mean by that, it's a little wireframe that gets drawn right on top of the face. It's kind of like those dots that they put on actor spaces in the behind the scene videos, they can track them, the dots can capture how the face moves and then
18:46
Speaker A
the computer can totally create a different character on top of that. So that same thing, same thing applies here.
18:53
Speaker A
It's going to track how the face moves and it's still going to hide the person behind it. And this is actually the easiest one because it only has two steps.
19:01
Speaker A
So you can change the pitch, you can add the face mash and there's no cutting anyone out like that. Like it just removes all those extra other steps.
19:10
Speaker A
So if you guys downloaded the project, you can just ask the agent to use this workflow instead, it's faster. And it's also right here in the interface.
19:17
Speaker A
It's just called fast, or you can ask it to install the face mash tracking and put it on top of the original video.
19:24
Speaker A
You send that in and you get something like this. That's dope. That's crazy. So what the face mash is doing is it's tracking 478 points on the face. So around the eyes, the lips, the nose, the whole face. And it is drawing that mash on every single frame.
19:42
Speaker A
And it saves all of those points with all the timing in a file that we have set up to be called face landmarks .json.
19:52
Speaker A
Everything is done and safe for you. You don't have to worry about it. And it uses Google's face landmark model from Google's official model storage.
20:02
Speaker A
One heads up though, at least right now, this works best with one person which is centered talking to the camera.
20:08
Speaker A
If you have two people moving around, I would just stick with a depth mask workflow.
20:13
Speaker A
And if you want to bullet prove this even more, you can ask the agent to put the face mash on top of the cut out depth video.
20:21
Speaker A
So on top of this one right here. And then you get something like this.
20:24
Speaker A
And then you send that in as a source video. And again, all of this works by just talking to it.
20:29
Speaker A
There's nothing for you to install yourself. You just ask the agent to set up the Google Media Pipe Face Landmarker or any other packages. Okay, this has been a long video.
20:41
Speaker A
So from here, you guys can keep making things better. You can add a transcriber. So it writes down what's said in the video and puts that into the prompt. You could have the agent watch the original video and describe the
20:53
Speaker A
movement in detail second by second. And then you can even use your own voice.
20:57
Speaker A
So you basically make your and I'm here to spill the secrets behind API's so that you can build million dollar workflows.
21:06
Speaker A
That is it. And if you take one thing away from this video is that everything, every one of these tricks was basically just the same idea implemented in different ways and methods. You just keep how things move and you throw away who is in the video.
21:20
Speaker A
And hopefully you guys know how Genjutsu actually works now and how to build yourself. I have attached the entire open source project in the link in the description. So just get it from there.
21:31
Speaker A
It's free and go make something with it. And in the next one, we're going to go way past the basics.
21:38
Speaker A
See you guys there and do not forget, create without limits. This is Sirio.
Topics:AI face swapHiggsfield GenjutsuSeedance 2.5reverse engineeringdepth maskvoice isolationopen source AIface trackingvideo editing AIIP bypass

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries and speaker detection. 30 minutes free.

Or transcribe another YouTube video here →