To generate a video that loops without a visible seam, 3 rules are enough: lock the camera, supply the same anchor image at the start and end of the generation, and conserve the cycle, so that what comes in goes out and nothing piles up in the frame. Applied to 20 micro-videos of 5 seconds produced with Seedance 2.0 on Magnific, they gave 17 closed loops without any editing, for 20,000 credits, or €16, and 1.5 hours of work in total, learning curve included.

A loading screen is the only place in a product where users watch without doing anything. On Qiplim, the app I'm building, those seconds show a spinner. I produced what it takes to replace it: one loop per feature, silent, on a white background, with the label placed by the app over the video. The credits went to Magnific.

It took 3 regeneration passes to get there, and the last one changed the design of the whole series. Every video in this article is fully AI-generated.

One loop from the series
The loop shown while a PDF turns into a quiz. 5 seconds, silent, pure white background, no text in the image. The app places the label in HTML on top. AI-generated video.

A loading screen is space you've already paid for

In 1995, Namco filed a patent on mini-games that keep players busy while a video game loads. The idea came from Ridge Racer, which ran a full game of Galaxian while the race loaded. The patent held for 20 years, until it expired in November 2015, and until then other studios had to pay for a license or work around the patent, according to the Electronic Frontier Foundation.

So the video game industry understood very early what business software still overlooks. Waiting time is attention space. It's already paid for, it's already being watched, and most products put a spinner there.

Qiplim turns documents into interactive activities. Importing a presentation, generating a quiz from a PDF, launching a session: each time, models are at work and the screen waits. That's the place for product education, the kind nobody reads in a help center.

The brief fit in 5 lines. 20 loops of 5 seconds, one per feature. Silent. A pure white #FFFFFF background, to blend into the interface. No text in the image. Autoplay and infinite loop, with no controls: these are interface backdrops, watched out of the corner of your eye while the app works.

5 decisions that make a series of loops manageable

Producing a video that loops is a different discipline from producing a video that tells a story. These 5 decisions are what made the series possible.

Text moves out of the video

Video models render accented French very poorly. A word comes back as a string of plausible glyphs, and an accent floats above the wrong letter. Instead of fighting it on every generation, I took the label out of the image: the app displays it in HTML over the video.

3 gains at once. The typography stays exactly on brand. Fixing a sentence no longer costs a generation. The series becomes translatable without touching a pixel.

The constraint shows up elsewhere: no text, number or logo may appear in the generated image anymore. The ban has to be repeated in every prompt, because models spontaneously add fake text to the screens and signs they draw.

The camera is locked

This was the most useful break from standard practice on the project. Video prompting guidelines discourage the static shot and frontal framing, which produce a dead image. For a loop, the rule flips. A camera that drifts by 3 degrees makes a seamless join impossible, because the last frame can no longer land on the first.

Focal length, height, distance: everything is locked in the prompt. Only the objects in the frame are allowed to move.

The cycle closes

A clean loop follows a conservation rule. What comes in goes out. What is consumed is replaced. Nothing piles up in the frame.

Without that discipline, the last frame no longer looks like the first and the loop jumps on every pass. The eye catches it immediately, even on a 200-pixel thumbnail in a corner of the interface.

Quantities are hard-coded

“About 10 hearts falling” produced more than 30 objects. Any vague quantity gets multiplied by 3, and I come back to it in the mistakes below.

The template is code

The 20 prompts aren't written by hand. A Python script generates them from a spec for each video, plus shared blocks: the style, the characters' identity lock, the format and the prohibitions. The script produces 34 files, reproducibly.

Every time I discovered a rule, I added it in one place, and all 20 prompts inherited it. That's what made it possible to correct the series 3 times without it drifting.

Beyond 5 visuals in the same series, prompts written one by one end up drifting apart. The unit of work becomes the template that builds them, and the prompts are simply its output.

Contact sheet of the 20 loading videos: blue, yellow and mauve plush mascots on a white background, one scene per feature
The contact sheet of the 20 loops, as used for review. One scene per feature, 3 mascots, a shared background. AI-generated images.

The 5 mistakes I paid for

All 5 are reproducible, and each one has a short fix. I list them in the order I figured them out.

1. References framed too tight

The symptom. The characters came out with stumps instead of legs.

The cause. I had supplied the expression sheets as references, and they frame the face and cut off the legs. The model couldn't guess what it couldn't see, so it made something up.

The fix. Reference the full-body views, and express proportions as ratios to the body instead of percentages of the frame. It took 2 full regenerations of the series before I traced the problem back to its cause.

2. Vague quantities get multiplied by 3

The symptom. “About 10 hearts falling” produced more than 30 objects, to the point of making the text displayed on top unreadable.

The cause. A density instruction filed under the format section carries no weight against a vague quantity written in the scene description.

The fix. 3 elements together: “exactly 10, no more”, a measurable spacing rule, and an explicit description of the failure to avoid.

3. The white background is never white

The symptom. On a white interface, each video formed a visible gray rectangle.

The cause. Despite the repeated instruction “pure white #FFFFFF”, the model outputs a studio cyclorama, between 236 and 250, with a slight gradient.

The fix. Systematic post-processing. A simple color distance threshold isn't enough: measured on these images, the darkest corner is at a distance of 31 from white, the contact shadows at 25, and a white felt panel at 42. No single threshold whitens the background without erasing the shadows that ground the characters. You have to target by connectivity to the edges of the image.

4. Describing a state the character can't have

The symptom. A character ended up with a plush ball sewn under its mouth.

The cause. I had written “the character has a slightly rounded belly” to suggest it had just swallowed a document. But this character is a blob: its body is a single mass, and it has no belly. The model took the instruction literally and built the missing part.

The fix. Describe the anatomy through what exists, and explicitly forbid any added protrusions.

An image model doesn't know an instruction is impossible. Name a body part that doesn't exist and it will build one.

5. The first frame already showed the result

This one cost the most, and it's the only one that changed the design of the series.

The symptom. On 3 test videos, 2 distinct behaviors. The sticky notes piled up on top of each other without the old ones leaving. The 2 transformation videos produced a near-frozen shot, with nothing moving.

The cause. My anchor image showed the middle of the cycle, with the produced objects already in place. The model reads that as stable scenery: either it has nothing left to tell and freezes the shot, or it stacks new objects on top of the old ones.

The fix. The starting image is point zero. Everything the action produces must be absent from it: a bare board, half the frame empty. A second image shows the result, and a 0.3-second crossfade in the edit joins them.

Before and after comparison for 3 loops: on the left the starting image without the product of the action, on the right the result the video has to create
The fix, on the 3 loops concerned. On the left, the starting image, emptied of everything the action must produce. On the right, the expected result at the end of the cycle. AI-generated images.

2 families of loops

That last fix revealed 2 ways of producing a loop, and they aren't generated the same way.

Family Principle Images supplied Looping
Pass-through what comes in goes out, nothing is created a single one, placed at the start and the end closes natively
Production the action makes something 2, before and after 0.3-second crossfade in the edit

17 of the 20 videos are pass-through loops and 3 are production loops. The 4 loops below are all pass-through: they loop without any editing.

Ranking. Cards that reorder themselves. The cleanest loop in the series.
The timer. An hourglass stretched then squashed. The best demonstration of elastic physics in the batch.
Live results. Bars rising without any character touching them.
Drag and drop. A card picked up and placed back in a queue. The only direct manipulation gesture.

What it costs on Magnific

Magnific, called Freepik until April 2026, bundles image and video models under a single credit-based bill. The 20 clips came out of Seedance 2.0, ByteDance's video model, at 1,000 credits per 5-second video, or 200 credits per second. That's the low end of this model's range, which climbs to 700 credits per second depending on the variant and resolution requested. For a plush toy moving on a white background, a higher resolution adds nothing you can see.

Item Value
Videos produced 20, 5 seconds each, in 16:9
Anchor image generations 40, for 23 images kept
Image rejection rate 43%
Credits used on video 20,000, or €16
Cost per video €0.80
Total working time 1.5 hours, learning curve included
Test clips before the fix 3, which revealed the design flaw

One detail changes how to read that figure: the €16 is 20,000 credits taken from the 45,000 in the monthly Premium+ plan, at €36 a month, and it covers only the video generations that were kept. It doesn't include the 40 anchor image generations. As for time, the whole project fits in 1.5 hours, including the 3 regeneration passes and learning the rules.

43% waste on images, 0 on videos once the method was set. All the waste was concentrated on the starting image, because I didn't yet know how to specify it.

If you want the narrative version of this pipeline, the one that starts from a script and produces a film instead of a loop, it's detailed in my case study on creating an AI video with Magnific and Dreamina. And for the logical next step, the edit itself, I documented a reel edited entirely by Claude. The Magnific links in this article are affiliate links, at no extra cost to you.

What I take away

Loading time is an in-house ad inventory. It's already paid for, it's already being watched, and almost nobody uses it. Putting product education there costs less than any acquisition channel.

The technical constraint produced the art direction. The pure white background came from the need to blend into the interface, with no aesthetic intent at the start. It became the signature of the series.

The value moved to the template. The whole project, including learning from the 5 mistakes, fits in 1.5 hours and €16. What remains is the script that builds the prompts: the rules are written there once, and the next series starts from it.

As of this writing, the 20 loops are finished and waiting to be integrated into the product. The mascots belong to Qiplim, which I cofounded.

Look at your own product. How many seconds a week do your users spend in front of a spinner, and what could you teach them in that time?

Frequently asked questions

3 conditions. The camera stays locked, with no lens movement at all: a drift of a few degrees keeps the last frame from landing on the first. The same anchor image is supplied at the start and end of the generation, so the model interpolates a cycle that returns to its starting point. And the cycle is conserved: what comes in goes out, what is consumed is replaced, and nothing piles up in the frame. A fourth rule applies to loops that make something: the starting image must be emptied of the result, otherwise the model reads the produced objects as stable scenery and freezes the shot. Out of 20 videos produced this way with Seedance 2.0, 17 loop without any editing. The 3 that make an object need a 0.3-second crossfade between the end and the start.
On this series, 20 videos of 5 seconds each used 20,000 credits, or €16 at the rate of a €36-a-month Premium+ subscription for 45,000 credits, which comes to €0.80 per video. They came out of Seedance 2.0 at 1,000 credits each, or 200 credits per second, the low end of this model's range, which goes up to 700 credits per second depending on the variant and resolution requested. The figure covers only the video generations that were kept: it doesn't include the 40 anchor image generations, 43% of which were discarded. The whole project, including the 3 regeneration passes and the learning curve, took 1.5 hours of work.
It's possible, but not advisable in French. Models render accented characters poorly and produce plausible glyphs instead of words. The better solution is to take the text out of the video and display it in HTML on top: the brand's typography is respected down to the pixel, fixing a sentence no longer costs a generation, and the series becomes translatable without retouching a single image. This approach comes with a constraint in return. You have to explicitly forbid any text, number and logo in every prompt, and repeat the ban on every generation, because models spontaneously add fake text to the screens, signs and documents they draw. On a series of 20 videos, this rule is written only once, in the template that builds the prompts.
Models interpret a white background instruction as a studio cyclorama and render values between 236 and 250, with a slight gradient, even when the hex code #FFFFFF is written in the prompt. On a white interface, the video then forms a visible gray rectangle. Post-processing is mandatory, and a global color threshold isn't enough: measured on this series, the darkest corner is at a distance of 31 from white, the contact shadows at 25 and a white felt panel at 42. No single threshold whitens the background without erasing the shadows that ground the characters. The method that works targets the pixels connected to the edges of the image, which leaves the light areas inside the subject untouched.
3 habits. Supply full-body references instead of expression sheets framed on the face: a model that can't see the legs invents them and outputs stumps. Express proportions as ratios to the body instead of percentages of the frame, for example legs twice the height of the body. And generate the prompts from a single template instead of writing them by hand: beyond 5 visuals in the same series, prompts written one by one drift apart, and a rule discovered late has to be copied by hand everywhere. On this project, a Python script produces the 20 prompts from a per-video spec and shared blocks, which made it possible to correct the series 3 times without it drifting.
Product education. The video game industry has filled these seconds since the 1990s: Namco filed a patent in 1995 on mini-games played during loading, which expired in November 2015. In a business application, a 5-second looping micro-video that shows a feature turns a forced wait into a demo, and reaches users at the exact moment they're looking at the screen with nothing else to do. The inventory is already paid for and already watched, which makes it cheaper than any acquisition channel. The design rule that makes the exercise sustainable: one silent loop per feature, with the label displayed by the application in HTML over the video instead of burned into the image.