Intro
How to read this
I gave the same four-line prompt to nine AI coding setups and asked each to rebuild this homepage as a macOS desktop. Same prompt, one run per combination, no follow-up coaching. Some are frontier models, some are free; some ran inside a polished agent harness, others on a bare CLI. Then I collected the results the way you'd collect anything you actually want to learn from.
The report has three parts. First, a scorecard of what's measurable: did it build, did the site's content survive, how it scored on Lighthouse, and what it spent in tokens. Second, the nine outputs as screenshots and live links, side by side with no winner declared, so you can judge the look yourself. Third, for each run, a short read of how it got there, drawn from its transcript and checked against the raw log.
One thing to hold onto throughout: this compares combinations, not bare models. The same model behaves very differently depending on the harness it runs in and the machine it runs on, so every result describes the pairing, not the model alone. It's one person's single-shot experiment, not a benchmark. Read it for the patterns, not a leaderboard.