How We Made an Animated Tamil Children's Song with AI, Shot by Shot
In August we published the first film on our own children's channel, NIM Kids: an original Tamil song, animated end to end with AI video models and finished by hand. It is 88 seconds long and took three attempts. The first two were rejected by me, the person who had asked for it, because the pictures matched the words only where the two happened to coincide. This is the account of what changed between the second cut and the third, written for anyone who wants to make a song film with these tools and would rather skip the two failures.
Key takeaways
- One sung line, one shot, and the shot shows literally what the line says — the picture cut to the beat matches only by accident.
- Isolate the vocal before transcribing and take the song's structure from the audio, checked against the written lyric sheet.
- Key captions to the measured vocal onset, not the bar line; on this song the gap ran from −1.9 s to +2.5 s.
- Lock each character with a clean, text-free reference image and one paragraph pasted verbatim into every prompt.
- Pick the model by whether it obeys the verb; generate every clip longer than its slot and trim, never stretch.
- Words on screen come from the lyric sheet and are composited on afterwards — a model never draws text or logos.
The rule that decided whether it worked
One sung line, one shot, and the shot shows literally what that line says. That sentence is the entire product. Everything else in this post is machinery for keeping it true.
It sounds obvious. It is not what the first two cuts did. They cut the picture on the beat grid — a new shot every bar or two, because that is how music videos feel — and then laid the lyrics over the top as captions. That guarantees musical cuts and guarantees nothing about meaning. When the line said "roll, roll, come", the character was sometimes rolling, sometimes hopping, sometimes standing in a field looking pleased. My own verdict on cut two, in the message that sent it back, was that very few scenes matched and the others did not match at all. A parent watching with a three-year-old feels that mismatch even if they cannot name it.
So the third cut started from the other end: the words first, then a shot per line, and a test for every shot — can you say in one sentence why this picture is that line? If not, the shot is wrong.
| The line says | The picture shows |
|---|---|
| roll, roll, come | the character rolling — not hopping, not walking |
| soft like butter | something physically squishing, next to real butter |
| into the tummy, go | someone actually eating, and the tummy responding |
| it jumped and fell into the pot | it jumps, and it lands in the pot |
| fried → crunchy | the frying pan, edges turning crisp |
| boiled → soft | the boiling pot, going wobbly |
Literal enactment: what the line says, what the picture must show
Measure the song before you shoot anything
The second cut also had a structural error that looked like a creative one. It contained a thirty-second instrumental intro and a verse that, according to our transcript, was never sung. Both were false. The transcription tool is a speech recogniser, and singing over instruments is not speech: on the full mix it found fifteen words and gave up two-thirds of the way through. Isolating the vocal track first and transcribing that gave fifty-three words and reached the end of the song.
The lesson generalises. Absence of a transcript is not absence of singing. Before a single shot is generated, the song's structure has to come from the audio itself — where the sections actually begin and repeat — and it has to be checked against the written lyric sheet. If the measured section count matches the number of stanzas on the sheet, the analysis can be trusted. If it does not, stop.
One more consequence of measuring: captions are keyed to the moment the singer actually starts each line, not to the bar the section starts on. On this song the gap between the two ran from almost two seconds early to two and a half seconds late. For a sing-along, captions that are out of step by that much are fatal.
Lock the cast, then paste the lock into every prompt
A song film has a small cast that must look the same for eighty seconds. Video models do not remember characters between shots, so consistency is something you build, not something you get. We generated one clean reference image per character, stripped every trace of label text out of it — any text left in a reference leaks into the generated footage — and wrote a one-paragraph description of each character that was pasted verbatim into every prompt. Paraphrasing the description between shots is the single most common reason a character's face changes halfway through a film.
The reference image and the prompt must also agree with each other. If the picture shows a yellow top and a green skirt, the prompt cannot ask for a yellow kurta with orange flowers. The model does not pick one; it splits the difference, and the character drifts.
Choose the model on whether it obeys verbs
The word in the rule is "literally", and that puts a hard requirement on the video model: it has to do what the verb says. We ran the same prompt — rolls toward the camera, does not hop — through three current models. One produced a character that walked and never rotated. One rolled but lost the character's face partway through the shot. One rolled and kept the face. That third model became the default for the whole film, and the choice was made on that evidence rather than on price or reputation.
Two smaller rules came out of the same shoot. Every shot is generated longer than the slot it will fill and then trimmed, never stretched: slowing a clip to fit is slow motion, and most of the rejected second cut was, on inspection, running in slow motion. And every prompt ends with an instruction to render no text, no captions, no logos and no watermark — the model is never allowed to draw words, because words are the one thing it gets subtly and visibly wrong.
Generate longer, then trim
A clip that runs 1.5 seconds longer than its slot can be cut to the exact beat. A clip that runs short can only be slowed, and a viewer reads slow motion instantly.
The captions are a separate film
The words on screen come from the written lyric sheet, never from the transcript — automatic Tamil spelling is wrong often enough to embarrass a children's channel. Tamil script also needs proper shaping to render, which the usual command-line video tools do not do, so every caption was rendered as an image in a real browser and composited onto the footage afterwards. The same goes for the title card and the end card: real artwork, placed on top, never generated.
Because the captions are their own layer keyed to the measured vocal onsets, the published video also carries proper Tamil and English subtitle tracks built from the same timings. That matters for a channel marked as made for kids, where parents read along.
What happened when it went out
The film went live on 18 August 2026 on a channel with no audience and no other videos. Seven days later it had 3,463 views and 47 likes, and the channel had 43 subscribers. Roughly five hundred views a day from a brand-new channel is the platform choosing to show it, not luck, and it is also perishable: a channel that promises a new video every week and then goes quiet loses that push quickly. Three vertical cuts of the same film followed in September. The next episode is the real test.
We do not treat the channel as an income line. A made-for-kids video carries no personalised advertising at all, and children's video in India earns very little per view. It is a portfolio piece for the AI video work we do for clients — the same pipeline, the same rules, aimed at a product or a brand instead of a potato — and it sits alongside our other products on the products page.
One sung line, one shot, and the shot shows literally what the line says. Everything else is machinery for keeping that true.
If you want one made
Bring the song and the written lyrics. The measuring, the shot table, the cast lock, the shoot and the caption layer are the work — and the work scales with the number of shots, which is why a tight 34-shot film is a better film than a lazy 80-shot one. What an AI video costs is written up separately. The channel itself is at youtube.com/@nimkidsofficial if you want to see the result before you read another word.
Frequently asked questions
Does the picture really have to match every line?
Yes, and it is the whole difference between a film that feels deliberate and one that feels generated. Cutting attractive footage to the beat and captioning it produces a video that matches the words only by accident. Children notice; so do their parents.
Why isolate the vocal before transcribing?
Transcription models recognise speech, and singing over instruments is not speech. On our song the full mix yielded fifteen words and stopped two-thirds through; the isolated vocal yielded fifty-three and reached the end. Without isolation you will conclude a song has silence where it has singing.
Can a video model draw the captions and the title?
It can, and it will get the letterforms subtly wrong — in Tamil, often badly wrong. Every word on screen in our film was rendered from the written lyrics in a real browser and composited on top of the footage. Models are never allowed to draw text or logos.
How long does one of these take to make?
It depends on the number of shots, and the number of shots is set by the number of sung lines, not by the song's length. Our 88-second film is 34 shots. Each shot is measured, prompted, generated longer than its slot, reviewed against its line and trimmed — a repeated chorus gets a different staging every time it comes round.
Have a project in mind?
We design, build, and ship software end-to-end — with a fixed, written quote after a free scoping call.
