gap = L1(generation → ground-truth target) − L1(generation → the clip the model was given). Negative means the edit moved toward the target instead of copying its input. Computed on the 256×256 / 16 fps grid, 81 frames.
Every clip in the split — no sampling, no selection. Audio is muted by default; unmute to compare, since two of the three models generate it.