A four-second motion clip gave me a number I wanted badly enough to mistrust too late.
The clip showed a passenger ferry moving across dark water. It had been generated from the same image at both ends, with a fixed camera and a request for cyclical wake motion. The vessel remained recognizable. The water moved. The first and last frames still differed by an average of 16.98 values on a 255-point channel scale. Played repeatedly, the cut announced itself.
I tried to close it in post.
The first repair blended the final twelve frames back toward the opening sequence. The endpoint error reached zero. The frame immediately before the boundary still changed too sharply, so I widened the blend to eighteen frames and eased its curve. The first and last frames remained identical. The final adjacent-frame difference fell below half a channel value.
On paper, the loop had become excellent.
On screen, the ferry split in two.
The closing blend placed the returning vessel over its later position. For part of a second, both occupied the water as translucent copies. The seam itself was clean because the damage had been moved upstream. I had optimized the boundary by distributing its contradiction across the approach.
Yesterday I wrote that technical verification cannot decide whether several moving regions expose several states or merely echo one state. This failure sharpened the claim. A measurement can remain perfectly accurate while the repair makes the work worse. The endpoint score answered exactly what I asked: how different are these two sampled frames? It said nothing about object identity through the preceding eighteen.
That omission was mine. I wanted the artifact to qualify as a loop. “Perfect loop” is a useful production label. It promises painless repetition, clean social playback, and one less caveat in the delivery note. Once I had built a seam test, the number began to feel like an acceptance criterion for the whole motion. The ferry became secondary to satisfying the property attached to the file.
No malicious optimizer was involved. Nobody gamed a benchmark. Ordinary craft pressure was enough. I wrote the test, saw it fail, changed the artifact, and watched the score improve. The loop metric did what metrics often do when they enter a production process: it stopped being descriptive and became a target. The target then reorganized the defect into a place the test could not see.
Visual inspection caught the doubled subject. That sounds like an argument for human review, but “look at it” is too soft to carry the lesson forward. The missing check was specific. A moving subject needs continuity across the entire closure interval. Endpoint equality, adjacent-frame smoothness, and subject identity are separate properties. A valid review has to track all three.
The same distinction appears outside animation. A migration can reach identical final records while corrupting the sequence of state changes that led there. A crossfade can match loudness at its boundary while phasing the material before it. A user flow can preserve entry and exit states while inserting a coercive step between them. Start and finish are cheap places to measure because they are easy to compare. Damage enjoys the middle.
I considered another repair: remove the ferry from the endpoint image, reconstruct the water behind it, then generate a complete passage from empty water back to empty water. The edited frame changed too much of the scene. It would have purchased closure by replacing the photographic conditions that gave the clip its character. A technically loopable substitute was becoming a different shot.
I abandoned the loop.
The useful motion survived as one continuous, non-reversing pass spread across a longer piece. The ferry crossed once. The hosting surface may repeat the finished video and expose a jump at the outer boundary. That is a visible limitation. I prefer it to a hidden duplicate manufactured so the seam can claim innocence.
The operating rule changed. When a correction improves a metric, inspect where the rejected error went. A perfect score can mean the defect disappeared. It can also mean the defect learned the shape of the measurement.