Ben Kamens, author of the olddogs.ai project, published a detailed breakdown of how his team built a corporate site and the article itself using AI agents Claude Code and Codex, achieving "zero slop." Quality here is presented not as luck, but as the result of four reproducible techniques: generative sketches before code, screenshot-to-mockup verification, a Hero Lab settings panel with JSON export back to the agent, and narrow "jig" applications for routine tasks. The entire generative design path is published as an archive of 2,460 images.

What happened
The article "How we built olddogs.ai with lots of agents but zero slop" was published on olddogs.ai, in which Ben Kamens detailed the assembly of a corporate project by Claude Code and Codex agents — covering not only the site pages but also the text of the breakdown itself. Design was sought through generative images rather than code: the agent produced 34 different "visual worlds," from which a lighthouse variant was chosen and then refined by verifying screenshots of the assembled page against the mockup; the "Why work with us" section took 38 variants and 7 rounds of such verification. The next technique — Hero Lab, a live dev panel with site parameters (cell density 55, fill 94%, tile shape round or hex), whose settings are exported as a single JSON block and passed to the agent as regular input. Routine tasks were moved to separate "jig" applications: reviewing 50+ pairs of dog images was done in a narrow application with a "Copy notes for Codex" button, which converts the reviewer's notes into structured input for the agent. Finally, inflated requirements: after the request "make it much faster," the agent rewrote the renderer in a GPU shader, frames grew from 14 to 117 fps at 1440×900 resolution, and the size increase was only 77 KB. At the end of the breakdown, a mosaic of all 2,460 intermediate images generated during the work is published.
Context
The discussion of agentic development usually swings between the poles of "neural network built a site in one evening" and "without a human, nothing will work anyway." The olddogs.ai breakdown is useful because it fixes the real place of the human in such assembly: quality here is not the result of a lucky prompt, but a function of a process built around the agent. Generative images before code cheaply expand the search space at the design stage, the "screenshot versus mockup" loop replaces trust in the model's self-assessment with iterative acceptance with human judgment, JSON export of settings makes the "human configured — agent applied" channel machine-readable and reproducible, and narrow "jig" applications turn routine review into cheap eval harnesses, in which human notes are packaged into structured input. In this view, each technique is a human-in-the-loop step: the model's autonomy is consciously limited, and it is precisely human selection that makes the quality. Notably, the public "benchmark" of a site built by agents demonstrates a measurable engineering result, not just a gallery of images.
Why this matters for the industry
For teams building products on Claude Code and Codex, this is a catalog of four reproducible patterns that can be transferred to their own pipelines without waiting for new tools. Image sketches before code save engineering time at the stage when layout has not yet proven anything; screenshot-to-mockup verification turns result acceptance into a manageable procedure rather than a final roulette; a parameters panel with JSON export sets an explicit "human — agent" contract, in which settings are not lost in the dialogue history; "jig" applications provide a cheap way to formalize routine operations, turning human notes into structured input. The likely consequence — instrumentalization of patterns: automatic comparison of screenshots with mockups, "jig" application templates, and normalized JSON context transfer to the agent; if this happens, the expectation bar for agentic web development will shift from "generated a site" to "built a reproducible process." We emphasize: this is a forecast, not a completed fact.
Why this matters for users
If you work with Claude Code or Codex, Kamens's techniques can be tried in your own project today. The article contains ready-made prompt formulations that can be taken as templates, and the Hero Lab JSON configuration, by which it is not difficult to set up your own parameters panel and pass settings to the agent in one block. The "screenshot versus mockup" verification can be built into any existing pipeline: assemble the page, take a screenshot, compare with the mockup, format the edits as text, and give it back to the agent. For your own routine — reviewing generative images, comparing variants, checking texts — you can build a narrow application in the spirit of "jig" with a button that immediately converts your notes into a prompt for the agent. The mosaic of 2,460 intermediate images shows how many iterations actually stand behind the final result: for small teams and startups without a studio budget, the breakdown calibrates expectations and confirms that a quality site is achievable by a small team working with agents.
What is not yet known / limitations
The case is field-based, not a methodology with metrics, and some effects need verification. The performance figures — growth from 14 to 117 fps at 1440×900 and an increase of 77 KB — are provided by the author himself without a description of the measurement methodology and stand, so they should be considered a point estimate. One successful translation of the renderer into a GPU shader by one bold request is n=1: from it does not follow a general conclusion that inflated requirements replace senior engineering expertise; for such a conclusion, reproducibility and independent confirmation are needed. Moreover, this is about a static corporate site with low failure risk, not a service with SLA, so the transfer of techniques to loaded or responsible systems is not demonstrated in the case.
Sources
Author
Look at AI, editorial team
