The creator of the crimeacs channel built an autonomous AI short-video studio in a claymation stop-motion style from scratch in five days and released twelve films assembled by two coding agents, Claude Code and Codex. Direct costs came to $184 — roughly one cent per view. The main takeaway of the breakdown goes beyond the videos themselves: at API prices, the most expensive part of such production turned out to be not video generation, but the agents' work.

image
image

What happened

On September 27, the creator opened the crimeacs channel with short videos in a claymation stop-motion style and grew the library to twelve films in five days. Nine of them are stories about famous fraudsters: Mavrodi and MMM, Charles Ponzi, Elizabeth Holmes, Bernie Madoff, Harshad Mehta, and others; the other three are office comedies about an analyst named Sam. All videos were automatically published to Instagram, YouTube Shorts, TikTok, and Facebook and collected a total of 19,739 views with 25 subscribers. All media generation went through fal.ai: static frames at $0.039 each were created by Nano Banana, clips at $0.112 per second by Kling 3 Pro, and voiceover was handled by ElevenLabs TTS. Direct costs fit within $184, of which $61, or 20% of the budget, went to recreated films.

Context

This project's pipeline grew out of an accident, not a plan: coding agents once silently recreated three films, and those losses forced the creator to rebuild the entire process. The resulting eight-stage scheme is built around error control: scenes are defined by a deterministic JSON script film.json, a character reference sheet is attached to each frame, voiceover is verified by cross-checking with a Whisper transcript, each shot is auto-checked across eight frames, and publication goes to social networks without human involvement. The video model was chosen from ten tested options based on a subjective criterion: Kling 3 Pro won with the best, in the creator's view, "acting" of hands and eyes. The price of agent work as a separate cost category is also telling: a Claude Max 20x subscription costs $200 per month, meaning agent access today is comparable in budget to GPU spending, not "minor" API calls.

Why this matters for the industry

The main industry signal from this case is the inversion of the cost structure of an autonomous studio: at API prices, agent tokens would have cost about $268, while all media generation cost roughly $117, and Claude Code in a single film spent up to $47, mostly re-reading its own history. The industrial takeaway is that value is shifting from wrappers around video models to the orchestration layer: budget limits, checkpoints, and auto-QA determine how much money reaches the finished video. The pipeline assembled by the creator essentially works as a homemade eval harness against agent errors, and it is precisely such tools that could become products: "pipeline-as-code" templates for niche styles, stage state control, and tools to reduce re-reading — context compression, caching, schedulers. If independent replications confirm the numbers, choosing a video model will become a budget decision based on price per second, and competition will shift there.

Why this matters for users

The recipe is reproducible today without a studio or film crew: Claude Code, Codex, fal.ai with Nano Banana and Kling 3 Pro, and ElevenLabs are available via API or subscription, and the cost of one film is $7–11. A breakdown using the film about Holmes as an example shows that the budget is spent predictably: 34 frames through Nano Banana for $1.33, 15 Kling 3 Pro clips for $5.25, voiceover for $0.26, and music for $0.29. Practical prompting rules were also derived from the same experience: attach a character reference sheet to each frame, explicitly state the absence of living beings in empty scenes, and put all "acting" elements into the static frame, because the video model will draw what was missing from the frame.

What is still unknown / limitations

The case is a single self-report from one operator working in one style, claymation stop-motion animation, so transferring the conclusions to other genres and formats is not yet confirmed by anything. The distribution signal is weak: 19,739 views with 25 subscribers, and the discussion on Hacker News gained only 1 point, so the thesis about a shift in value toward the orchestration layer remains an interpretation, not an established pattern. The choice of Kling 3 Pro was made without a protocol, metrics, or blind evaluation, based on the creator's subjective impression, and as a benchmark this test is invalid. All bills and arithmetic were re-checked only by the creator himself: both primary sources are his blog and his own thread on Hacker News, and there is no independent replication yet. Finally, "demonstrated" is not equal to "deployed reliably": without human gates, for example approval of one frame and one clip per batch, and checkpoints, the stability of the pipeline remains unproven.

Sources

Author

Look at AI, editorial team