JavisDiT Edit — audio-visual object addition & removal

Checkpoint text_both_0721 / epoch005-global_step40000 (EMA) · Wan2.1-T2AV 1.3B, FiLM text conditioning, trained on both edit directions · 81 frames @ 16 fps, 256×256, 50 sampling steps, CFG 5.0. One test-set sample and one train-set sample (overfit check), each run in both directions and both with the first frame (target frame 0 prepended as anchor, stripped before decode) and without the first frame (no anchor). All clips include generated audio — unmute to listen.

Addition — “birds”

test seta video with birds · epMUuqXcgeo_000030

with first frame

Source (input, object removed)

Generated (model output)

Target (ground truth)

without first frame

Source (input, object removed)

Generated (model output)

Target (ground truth)

Removal — “birds”

test seta video without birds · epMUuqXcgeo_000030

with first frame

Source (input, object present)

Generated (model output)

Target (ground truth)

without first frame

Source (input, object present)

Generated (model output)

Target (ground truth)

Addition — “trumpet”

train set (overfit check)a video with trumpet · GfeEN8LONh0_000253

with first frame

Source (input, object removed)

Generated (model output)

Target (ground truth)

without first frame

Source (input, object removed)

Generated (model output)

Target (ground truth)

Removal — “trumpet”

train set (overfit check)a video without trumpet · GfeEN8LONh0_000253

with first frame

Source (input, object present)

Generated (model output)

Target (ground truth)

without first frame

Source (input, object present)

Generated (model output)

Target (ground truth)