Dataset — ~50–60 mixed photos + video-clip frames per LoRA. Real-world inputs: phone-camera quality, multiple lighting setups, a mix of expressions and angles. No studio shoots — the dataset is what you'd actually have for a real-person LoRA.
Three LoRAs trained:
Training stack — AI Toolkit on RunPod. Stable, well-documented, runs cleanly on rented A100 / H100 instances. Cheaper than dedicated hardware, more reliable than any "train in the browser" service I tried.
The right LoRA per architecture matters. Cross-family transfer (using a Wan-trained LoRA on LTX 2) does technically work, but it loses fidelity that a dedicated LoRA recovers — clearest in the lip-sync test below.