bleckhert Alissonerdx commited on
Commit
edd3965
·
0 Parent(s):

Duplicate from Alissonerdx/BFS-Best-Face-Swap-Video

Browse files

Co-authored-by: Alisson Pereira Anjos <Alissonerdx@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <<<<<<< HEAD
2
+ *.7z filter=lfs diff=lfs merge=lfs -text
3
+ *.arrow filter=lfs diff=lfs merge=lfs -text
4
+ *.bin filter=lfs diff=lfs merge=lfs -text
5
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
6
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
7
+ *.ftz filter=lfs diff=lfs merge=lfs -text
8
+ *.gz filter=lfs diff=lfs merge=lfs -text
9
+ *.h5 filter=lfs diff=lfs merge=lfs -text
10
+ *.joblib filter=lfs diff=lfs merge=lfs -text
11
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
12
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
13
+ *.model filter=lfs diff=lfs merge=lfs -text
14
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
15
+ *.npy filter=lfs diff=lfs merge=lfs -text
16
+ *.npz filter=lfs diff=lfs merge=lfs -text
17
+ *.onnx filter=lfs diff=lfs merge=lfs -text
18
+ *.ot filter=lfs diff=lfs merge=lfs -text
19
+ *.parquet filter=lfs diff=lfs merge=lfs -text
20
+ *.pb filter=lfs diff=lfs merge=lfs -text
21
+ *.pickle filter=lfs diff=lfs merge=lfs -text
22
+ *.pkl filter=lfs diff=lfs merge=lfs -text
23
+ *.pt filter=lfs diff=lfs merge=lfs -text
24
+ *.pth filter=lfs diff=lfs merge=lfs -text
25
+ *.rar filter=lfs diff=lfs merge=lfs -text
26
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
27
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
29
+ *.tar filter=lfs diff=lfs merge=lfs -text
30
+ *.tflite filter=lfs diff=lfs merge=lfs -text
31
+ *.tgz filter=lfs diff=lfs merge=lfs -text
32
+ *.wasm filter=lfs diff=lfs merge=lfs -text
33
+ *.xz filter=lfs diff=lfs merge=lfs -text
34
+ *.zip filter=lfs diff=lfs merge=lfs -text
35
+ *.zst filter=lfs diff=lfs merge=lfs -text
36
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
37
+ ltx-2/ filter=lfs diff=lfs merge=lfs -text
38
+ =======
39
+ *.7z filter=lfs diff=lfs merge=lfs -text
40
+ *.arrow filter=lfs diff=lfs merge=lfs -text
41
+ *.bin filter=lfs diff=lfs merge=lfs -text
42
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
43
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
44
+ *.ftz filter=lfs diff=lfs merge=lfs -text
45
+ *.gz filter=lfs diff=lfs merge=lfs -text
46
+ *.h5 filter=lfs diff=lfs merge=lfs -text
47
+ *.joblib filter=lfs diff=lfs merge=lfs -text
48
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
49
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
50
+ *.model filter=lfs diff=lfs merge=lfs -text
51
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
52
+ *.npy filter=lfs diff=lfs merge=lfs -text
53
+ *.npz filter=lfs diff=lfs merge=lfs -text
54
+ *.onnx filter=lfs diff=lfs merge=lfs -text
55
+ *.ot filter=lfs diff=lfs merge=lfs -text
56
+ *.parquet filter=lfs diff=lfs merge=lfs -text
57
+ *.pb filter=lfs diff=lfs merge=lfs -text
58
+ *.pickle filter=lfs diff=lfs merge=lfs -text
59
+ *.pkl filter=lfs diff=lfs merge=lfs -text
60
+ *.pt filter=lfs diff=lfs merge=lfs -text
61
+ *.pth filter=lfs diff=lfs merge=lfs -text
62
+ *.rar filter=lfs diff=lfs merge=lfs -text
63
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
64
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
65
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
66
+ *.tar filter=lfs diff=lfs merge=lfs -text
67
+ *.tflite filter=lfs diff=lfs merge=lfs -text
68
+ *.tgz filter=lfs diff=lfs merge=lfs -text
69
+ *.wasm filter=lfs diff=lfs merge=lfs -text
70
+ *.xz filter=lfs diff=lfs merge=lfs -text
71
+ *.zip filter=lfs diff=lfs merge=lfs -text
72
+ *.zst filter=lfs diff=lfs merge=lfs -text
73
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
74
+ ltx-2/ filter=lfs diff=lfs merge=lfs -text
75
+ >>>>>>> 47487e3e7ace468206bad3d0247ed81b792bf222
76
+ examples/ filter=lfs diff=lfs merge=lfs -text
77
+ workflows/workflow_head_swap_drag_and_drop.png filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,501 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: ltx-2-community-license-agreement
4
+ tags:
5
+ - ltx-2
6
+ - ic-lora
7
+ - head-swap
8
+ - video-to-video
9
+ - image-to-video
10
+ - bfs
11
+ - lora
12
+ base_model:
13
+ - Lightricks/LTX-2
14
+ - Lightricks/LTX-2.3
15
+ library_name: diffusers
16
+ pipeline_tag: image-to-video
17
+ ---
18
+
19
+ ## ⚠️ Ethical Use & Disclaimer
20
+
21
+ This model is a technical tool designed for **Digital Identity Research, Professional VFX Workflows, and Cinematic Prototyping**.
22
+
23
+ By downloading or using this LoRA, you acknowledge and agree to the following:
24
+
25
+ * **Intended Use:** Designed for filmmakers, VFX artists, and researchers exploring high-fidelity video identity transformation.
26
+ * **Consent & Rights:** You must possess explicit legal consent and all necessary rights from any individual whose likeness is being processed.
27
+ * **Legal Compliance:** You are fully responsible for complying with all local and international laws regarding synthetic media.
28
+ * **Liability Waiver:** This model is provided *"as is."* **As the creator (Alissonerdx), I assume no responsibility for misuse.** Any legal, ethical, or social consequences are solely the responsibility of the end user.
29
+
30
+ ---
31
+
32
+ # 📺 Video Examples
33
+
34
+ ## V1 Examples
35
+
36
+ Generated using the **Frame 0 Anchoring Technique**.
37
+ All examples follow the guide video motion while preserving the identity provided in the first frame.
38
+
39
+ | Example 1 | Example 2 |
40
+ | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
41
+ | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/1.mp4" controls autoplay loop muted></video> | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/2.mp4" controls autoplay loop muted></video> |
42
+
43
+ | Example 3 | Example 4 |
44
+ | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
45
+ | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/3.mp4" controls autoplay loop muted></video> | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/4.mp4" controls autoplay loop muted></video> |
46
+
47
+ | Example 5 |
48
+ | ------------------------------------------------------------------------------------------------------------------------------------------ |
49
+ | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/5.mp4" controls autoplay loop muted></video> |
50
+
51
+ ## V3 Examples
52
+
53
+ If you want to see the full setup in practice, watch here:
54
+
55
+ [https://www.youtube.com/watch?v=HBp03iu7wLA](https://www.youtube.com/watch?v=HBp03iu7wLA)
56
+
57
+ The following examples demonstrate the new **persistent-template workflow** used in V3:
58
+
59
+ | Example 6 | Example 7 |
60
+ | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
61
+ | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/6.mp4" controls autoplay loop muted></video> | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/7.mp4" controls autoplay loop muted></video> |
62
+
63
+ | Example 8 |
64
+ | ------------------------------------------------------------------------------------------------------------------------------------------ |
65
+ | <video src="https://huggingface.co/Alissonerdx/BFS-Best-Face-Swap-Video/resolve/main/examples/8.mp4" controls autoplay loop muted></video> |
66
+
67
+ The image references for the versions are stored under:
68
+
69
+ ```txt
70
+ ltx-2.3/...
71
+ ```
72
+
73
+ ---
74
+
75
+ # 🛠 Technical Background (V1)
76
+
77
+ To achieve this level of identity transfer, I **heavily modified the official LTX-2 training scripts**.
78
+
79
+ ### Key Improvements
80
+
81
+ * **Novel Conditioning Injection:** Custom latent injection methods for reference identity stabilization.
82
+ * **Noise Distribution Overhaul:** Implemented a **custom High-Noise Power Law timestep distribution**, forcing the model to prioritize target identity reconstruction over guide-video context.
83
+ * **Training Compute:** 60+ hours of training on **NVIDIA RTX PRO 6000 Blackwell GPUs**, iterating through 300GB+ of experimental checkpoints.
84
+
85
+ ---
86
+
87
+ # 📊 Dataset Specifications
88
+
89
+ ## V1 Dataset
90
+
91
+ * **300 high-quality head swap video pairs**
92
+ * Trained on **512x512 buckets**
93
+ * Primarily **landscape format**
94
+ * Optimized for **close-up framing**
95
+
96
+ Wide shots may reduce identity fidelity.
97
+
98
+ ---
99
+
100
+ # 💡 Inference Guide (V1)
101
+
102
+ ## 🔴 CRITICAL — Frame 0 Requirement
103
+
104
+ This version was trained to use **Frame 0 as the identity anchor**.
105
+
106
+ You must prepare the first frame correctly.
107
+
108
+ ### Recommended Workflow
109
+
110
+ 1. Perform a high-quality head swap on Frame 0.
111
+ 2. Use that processed frame as conditioning input.
112
+ 3. Run the full video generation.
113
+
114
+ For best results, prepare Frame 0 using my previous **BFS Image Models**.
115
+
116
+ ---
117
+
118
+ ## Optimization
119
+
120
+ ### LoRA Strength
121
+
122
+ * **1.0** → Best motion fidelity
123
+ * **>1.0** → Stronger identity and hair capture, but may distort original motion
124
+
125
+ ### Multi-Pass Workflows
126
+
127
+ You can experiment with multiple passes using different strengths.
128
+
129
+ ### Prompting
130
+
131
+ Detailed prompts currently have **no effect**.
132
+
133
+ Trigger remains:
134
+
135
+ ```txt
136
+ head swap
137
+ ```
138
+
139
+ ---
140
+
141
+ # ⚠️ Known Issues (V1 – Alpha)
142
+
143
+ * **Identity Leakage:** Hair from the guide video may reappear.
144
+ * **Hard Cuts:** Jump cuts can reset identity.
145
+ * **Portrait Format:** Performance is significantly better in landscape.
146
+
147
+ ---
148
+
149
+ # 🚀 Version 2 – Major Update
150
+
151
+ V2 introduces a **complete redesign of conditioning strategy and masking logic**, significantly improving identity robustness and reducing leakage.
152
+
153
+ ---
154
+
155
+ ## 🔹 Multiple Conditioning Modes (Using First Frame)
156
+
157
+ V2 supports multiple identity injection approaches:
158
+
159
+ ### 1️⃣ Direct Photo Conditioning
160
+
161
+ Use a clean photo of the new face as reference input.
162
+
163
+ This method works and can produce strong results. However, because the model must internally reconcile lighting, perspective, depth, and occlusion differences, it may need to fight to correctly integrate the new identity into the guide video. In some cases, this can reduce stability or identity consistency.
164
+
165
+ ### 2️⃣ First-Frame Head Swap (Recommended)
166
+
167
+ Applying a proper head swap on Frame 0 still produces **extremely strong and reliable results**.
168
+
169
+ Because the first frame is already structurally correct — pose, lighting, depth, and occlusions — the model has significantly less work to do. Instead of forcing alignment from a static photo, it simply propagates and stabilizes the identity through time.
170
+
171
+ This approach generally:
172
+
173
+ * Produces higher identity fidelity
174
+ * Reduces deformation
175
+ * Minimizes integration artifacts
176
+ * Improves overall temporal stability
177
+
178
+ ### 3️⃣ Automatic Magazine-Style Overlay
179
+
180
+ The new face is automatically cut and positioned over the guide face using mask alignment.
181
+ This simulates a magazine-cutout-style overlay, but performed automatically based on mask positioning.
182
+
183
+ ### 4️⃣ Manual Overlay
184
+
185
+ Advanced users may manually composite the new face over Frame 0 before running inference.
186
+
187
+ ---
188
+
189
+ ## 🔹 Facial Motion Behavior (Important Change)
190
+
191
+ Unlike V1:
192
+
193
+ **V2 does not follow the original guide face’s facial micro-movements.**
194
+
195
+ The guide face is fully masked to prevent identity leakage.
196
+
197
+ This makes masking quality critical.
198
+
199
+ ### Mask Requirements
200
+
201
+ * The guide face must be completely covered.
202
+ * Mask color must be a **magenta tone**.
203
+ * Any visible guide identity may leak into the final output.
204
+
205
+ ---
206
+
207
+ ## 🔹 Mask Types
208
+
209
+ Users may alternate between:
210
+
211
+ ### ▪ Square Masks
212
+
213
+ * More stable identity
214
+ * Better consistency
215
+ * Often produce stronger overall results
216
+ * May generate slightly oversized heads due to spatial padding
217
+
218
+ In most scenarios, square masks tend to perform better because they provide additional spatial context for the model to reconstruct structure and hair.
219
+
220
+ ### ▪ Tight / Adjusted Masks
221
+
222
+ * More natural head proportions
223
+ * May deform if guide head shape differs significantly
224
+ * Sensitive to long-hair mismatches
225
+
226
+ If the original guide has long hair and the new identity does not, deformation risk increases.
227
+
228
+ ---
229
+
230
+ ## 🔹 Dataset & Training Improvements (V2)
231
+
232
+ * **800+ video pairs**
233
+ * Trained at **768 resolution**
234
+ * **768** is the recommended inference resolution
235
+ * Improved hair stability
236
+ * Reduced identity leakage compared to V1
237
+ * More robust identity transfer under motion
238
+
239
+ ---
240
+
241
+ ## 🔹 First Pass vs Second Pass
242
+
243
+ You may:
244
+
245
+ * Run a single pass at 768 (recommended)
246
+ * Or run a downscaled first pass plus a second upscale pass
247
+
248
+ ⚠️ Important:
249
+
250
+ A second pass may alter identity from the first pass and reduce consistency in some cases.
251
+
252
+ ---
253
+
254
+ ## 🔹 Trigger
255
+
256
+ Trigger remains:
257
+
258
+ ```txt
259
+ head swap
260
+ ```
261
+
262
+ ---
263
+
264
+ # 🚀 Version 3 – Persistent Template Workflow
265
+
266
+ V3 introduces a new **persistent-template conditioning workflow**.
267
+
268
+ Unlike previous versions, which relied primarily on the identity being established from **Frame 0 only**, V3 uses a **custom guide-video construction step** that keeps the new face visible throughout the entire guide sequence.
269
+
270
+ This results in a much stronger and more persistent identity signal during inference.
271
+
272
+ # 🙏 Acknowledgements
273
+
274
+ Special thanks to **facy.ai** for sponsoring the GPU used to train this model.
275
+
276
+ If you want to check their platform, you can use my referral link:
277
+
278
+ [https://facy.ai/a/headswap](https://facy.ai/a/headswap)
279
+
280
+ ---
281
+
282
+ ## 🔹 How V3 Works
283
+
284
+ V3 uses a custom node from **ComfyUI-BFSNodes** to prepare the guide video before inference.
285
+
286
+ Repository:
287
+
288
+ ```txt
289
+ https://github.com/alisson-anjos/ComfyUI-BFSNodes
290
+ ```
291
+
292
+ Workflow file:
293
+
294
+ ```txt
295
+ workflows/workflow_ltx2_head_swap_drag_and_drop_v3.0
296
+ ```
297
+
298
+ The guide-video preparation process works like this:
299
+
300
+ 1. Start from the original guide video
301
+ 2. Add a **vertical green chroma-key strip** on the side
302
+ 3. Place the **reference face image** inside that strip
303
+ 4. Apply this composition to **every frame** of the original video
304
+ 5. Use this new composite video as the actual inference guide
305
+
306
+ This means the new identity remains **fully visible during all frames** of the guide video, instead of appearing only in Frame 0 like in previous versions.
307
+
308
+ That is the main reason V3 can achieve better consistency than earlier versions.
309
+
310
+ ---
311
+
312
+ ## 🔹 Why V3 Is Different
313
+
314
+ Because the identity reference stays visible during the full guide sequence, V3 gives the model a much more stable conditioning signal across time.
315
+
316
+ In practice, this can improve:
317
+
318
+ * Identity consistency
319
+ * Temporal stability
320
+ * Resistance to identity drift
321
+ * Facial motion continuity
322
+ * Lip sync behavior
323
+ * Expressive facial movement preservation
324
+
325
+ This version is especially useful for shots where the face remains visible for longer periods, or where dialogue, mouth movement, and facial acting matter more.
326
+
327
+ V3 is not just a refinement of the first-frame method. It changes the conditioning logic by giving the model access to a persistent identity template across the entire inference sequence.
328
+
329
+ ---
330
+
331
+ ## 🔹 Final Output Behavior
332
+
333
+ Even though the guide video used during inference contains the **vertical chroma-key side strip**, the **final generated result does not include that strip**.
334
+
335
+ The generated video is returned in the **original resolution and framing** of the source guide video.
336
+
337
+ So in practice:
338
+
339
+ * The green side strip exists only in the internal guide/template video
340
+ * It is used only to improve inference conditioning
341
+ * It does not appear in the final output
342
+
343
+ ---
344
+
345
+ ## 🔹 Prompting for V3
346
+
347
+ For V3, users can also pass the composite guide video into a vision-capable model to extract a structured prompt.
348
+
349
+ This is useful because the composite video contains two different information sources:
350
+
351
+ * the **reference identity** inside the side strip
352
+ * the **performance and scene information** in the main video area
353
+
354
+ This helps keep identity and action description separated more cleanly.
355
+
356
+ ### Recommended Prompt Template
357
+
358
+ ```txt
359
+ Analyze this composite video.
360
+
361
+ The video contains:
362
+ 1. a side chroma-key panel with a reference face image
363
+ 2. a main performance video showing the body, clothing, movement, hand actions, objects, framing, and environment
364
+
365
+ Your task is to extract:
366
+ - the target face identity from the side panel
367
+ - the performance/action from the main video
368
+
369
+ Critical rules:
370
+ - The side-panel face is the only valid source for identity traits and head-level accessories.
371
+ - Ignore the visible face and head appearance in the main video completely.
372
+ - Do not describe any face, hair, hairstyle, hair color, eye color, makeup, facial features, facial expression, attractiveness, headwear, hood, hat, or accessories from the main video.
373
+ - In the ACTION section, describe the performer only as "a person" and focus only on body movement, clothing, hand actions, objects, framing, and environment.
374
+ - Do not mention the chroma panel, green background, split layout, or editing structure.
375
+ - Be factual and non-creative.
376
+ - Do not guess uncertain details. If a detail is not clearly visible, omit it.
377
+
378
+ Return exactly in this format:
379
+ head_swap:
380
+
381
+ FACE:
382
+ A brief but detailed objective identity description from the side-panel face only. Include, when clearly visible: apparent gender, apparent ethnicity, skin tone or complexion, approximate age range, head shape, hair or baldness pattern, hair color, eye color, facial hair, visible skin details, headwear or head covering, visible facial accessories, and any especially distinctive facial trait. Prioritize the eyes when they are a strong defining feature.
383
+
384
+ ACTION:
385
+ A concise performance description from the main video. Include only: visible clothing, body position, movement, hand actions, objects being shown or handled, camera-facing behavior, framing, and environment. Do not include any face or head appearance from the main video.
386
+
387
+ Good example:
388
+ FACE:
389
+ Female, fair skin, approximately 20-30 years old, oval head shape, long wavy vivid blue-violet hair, bright golden-amber eyes with dark defined pupils, no facial hair, smooth skin, and pink flower hair accessories as a distinctive head adornment.
390
+
391
+ ACTION:
392
+ A person in a dark top faces the camera indoors, holds a package of false eyelashes close to the lens, peels one lash from the backing, brings it near the eye area, and examines it while making small hand movements.
393
+
394
+ Bad example:
395
+ ACTION:
396
+ A person with long curly blonde braids holds a pair of false eyelashes...
397
+ ```
398
+
399
+ ### How to Use
400
+
401
+ 1. Generate the V3 composite guide video using the node
402
+ 2. Pass that composite video into a vision-capable model
403
+ 3. Extract the structured **FACE** and **ACTION** prompt
404
+ 4. Use that output as the base prompt for the V3 workflow
405
+
406
+ ---
407
+
408
+ ## 🔹 Captions / Descriptions for V3
409
+
410
+ If you want automatic captions or prompt extraction from video, you can also use my Ollama nodes.
411
+
412
+ Repository:
413
+
414
+ ```txt
415
+ https://github.com/alisson-anjos/ComfyUI-Ollama-Describer
416
+ ```
417
+
418
+ A useful node for this workflow is:
419
+
420
+ **Ollama Video Describer**
421
+
422
+ This can help generate structured descriptions from the composite guide video and make it easier to build the final prompt for V3.
423
+
424
+ ---
425
+
426
+ ## 🔹 V3 Trigger
427
+
428
+ Trigger remains:
429
+
430
+ ```txt
431
+ head_swap:
432
+ FACE:
433
+ ....
434
+
435
+ ACTION:
436
+ ....
437
+ ```
438
+
439
+ ---
440
+
441
+ # 🔴 Critical Success Factor (V2 / V3)
442
+
443
+ Mask and preparation quality still matter enormously.
444
+
445
+ Even with improved conditioning, final quality depends on:
446
+
447
+ * Proper face coverage
448
+ * Clean compositing
449
+ * Strong alignment
450
+ * Good source and reference quality
451
+
452
+ If any portion of the original guide identity remains visible where it should not, the model may still reintroduce unwanted traits.
453
+
454
+ Take time to refine your inputs. Better preparation consistently produces better output than simply increasing LoRA strength.
455
+
456
+ ---
457
+
458
+ ## 🔧 Advanced Technique: Combine with LTX-2 Inpainting
459
+
460
+ Advanced users can experiment with combining this LoRA with the native **LTX-2 inpainting workflow**.
461
+
462
+ This can help:
463
+
464
+ * Refine problematic areas
465
+ * Correct small deformation zones
466
+ * Improve edge blending
467
+ * Recover detail in hair or jaw regions
468
+
469
+ When properly combined, inpainting can significantly enhance final output quality, especially in challenging frames.
470
+
471
+ ---
472
+
473
+ ## 🔹 Recommendation
474
+
475
+ I strongly recommend testing **both LoRAs** and comparing the final behavior.
476
+
477
+ Depending on the guide clip, framing, facial motion, and the kind of result you want, some users may prefer the look or motion style of one version over the other.
478
+
479
+ In general:
480
+
481
+ * **V2** may still be preferred for some first-frame-driven workflows
482
+ * **V3** is better when you want a stronger persistent identity signal, better consistency, and better facial/lip motion continuity
483
+
484
+ The best version will often depend on the shot and on personal preference.
485
+
486
+ ---
487
+
488
+ # 💙 Support
489
+
490
+ Maintaining R&D and renting Blackwell GPUs is expensive.
491
+
492
+ If this project helps you, consider supporting the development of:
493
+
494
+ * V3 improvements
495
+ * Advanced conditioning pipelines
496
+ * SAM 3 integration
497
+ * Full reference-photo-only workflows
498
+
499
+ Support here:
500
+
501
+ [https://buymeacoffee.com/nrdx](https://buymeacoffee.com/nrdx)
examples/1.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb5008a5750697b2ba6a105813455861c8f84949762a9a919c4f423b60bfc124
3
+ size 30056875
examples/2.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd30caefa2bbefe1bdc8927855bf472dce67bdb83e25eaa46255f073f4f79885
3
+ size 31230709
examples/3.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:545e6962243768f4a448f92e8ae72037d3f36a5d26eb2684595e79b045dbe500
3
+ size 66240119
examples/4.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1166856da2aee972247473feca89a3d2f31a90d8bbeeb882de3c5664f07fa1e2
3
+ size 38867419
examples/5.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7437dfd7d03ec41af05d1fa71b2d3b1fe107f4fa87b8e38ebc66b573ec16c66d
3
+ size 38173397
examples/6.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:55b9eef31500f4a392e1f203f098e9408fe7e47887f6a39ebee03b722d44ab7b
3
+ size 1688996
examples/7.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7a00b82f46b13fcf1eaa632a922c1b221316b3593387eb151ffe10e9613fe86f
3
+ size 1750633
examples/8.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:adee6e86236ce5c5c2e1d3f12811488875b2b0232658cb7e1d5cd90ffc53f270
3
+ size 3308271
ltx-2.3/head_swap_v3_rank_64.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a25d4fd9622d9d81f897744d1d7d7b99a17812eeb41905e68aacf5c375892608
3
+ size 654443424
ltx-2.3/head_swap_v3_rank_adaptive_fro_098.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:78083ede0004c500ac0ae7b8fee43d28a1c3b503bac4a21edfeba7c60da8460d
3
+ size 1358465856
ltx-2/download-models-head-swap-ltx2-windows.ps1 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:40f0dded61a782ccdc19c30d9ead9bef69a5a53a05a6fc596ee7e02efedbde97
3
+ size 4146
ltx-2/download-models-head-swap-ltx2.sh ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6427e10f9a4d9141c728ebef8e9c0908700bd94726be53effc85ecb99a78423b
3
+ size 5119
ltx-2/head_swap_v1_13500_first_frame.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:056373cf73418dac449fecf34a5b749deeb802a1a4a4a9fc1677cd46c2d48864
3
+ size 1308756368
ltx-2/head_swap_v1_8750_first_and_last_frame.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:48d16179c82629385fb5a812ed33182952e7755a60464ffe645d3417f5d48a71
3
+ size 1308756368
ltx-2/head_swap_v2_multimodes.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f459e03568447dcc5d6ea7c02466fece5eee3409bb23f07dfc2ecab24ac7a2fa
3
+ size 1316096704
workflows/workflow_ltx2_head_swap_drag_and_drop.json ADDED
The diff for this file is too large to render. See raw diff
 
workflows/workflow_ltx2_head_swap_drag_and_drop_v1.1.json ADDED
The diff for this file is too large to render. See raw diff
 
workflows/workflow_ltx2_head_swap_drag_and_drop_v2.0.json ADDED
The diff for this file is too large to render. See raw diff
 
workflows/workflow_ltx2_head_swap_drag_and_drop_v3.0.json ADDED
The diff for this file is too large to render. See raw diff