Checkpoint text_both_0721 / epoch005-global_step40000 (EMA) · Wan2.1-T2AV 1.3B, FiLM text
conditioning, trained on both edit directions · 81 frames @ 16 fps, 256×256, 50 sampling steps, CFG 5.0.
One test-set sample and one train-set sample (overfit check), each run in both directions and both
with the first frame (target frame 0 prepended as anchor, stripped before decode) and
without the first frame (no anchor). All clips include generated audio — unmute to listen.
Addition — “birds”
test seta video with birds · epMUuqXcgeo_000030
with first frame
Source (input, object removed)
Generated (model output)
Target (ground truth)
without first frame
Source (input, object removed)
Generated (model output)
Target (ground truth)
Removal — “birds”
test seta video without birds · epMUuqXcgeo_000030
with first frame
Source (input, object present)
Generated (model output)
Target (ground truth)
without first frame
Source (input, object present)
Generated (model output)
Target (ground truth)
Addition — “trumpet”
train set (overfit check)a video with trumpet · GfeEN8LONh0_000253
with first frame
Source (input, object removed)
Generated (model output)
Target (ground truth)
without first frame
Source (input, object removed)
Generated (model output)
Target (ground truth)
Removal — “trumpet”
train set (overfit check)a video without trumpet · GfeEN8LONh0_000253