DigestAI news desk
Generative AI & Models updated 4 min read

AI Coding Tools Show Distinct Differences in Website Rebuilds

I tested three AI coding tools—Claude Code, Codex, and Google Antigravity—to rebuild a complex food-delivery website. Despite starting with the same brief, each tool produced vastly different results. Claude Code maintained consistent design elements throughout, while Codex excelled in its hero section but lacked overall polish. Antigravity's Gemini version stood out for its polished visuals and…

1 source

Key points

  • Claude Code maintained consistency across all design elements
  • Codex excelled in the hero section but lacked overall polish
  • Antigravity's Gemini version stood out for its polished visuals

The story so far

2 episodes →
  1. AI Coding Tools Show Distinct Differences in Website Rebuilds this story
Full story from xda-developers.com · by Parth Shah · via Search: Claude Open source ↗

I asked Claude Code, Codex, and Google Antigravity to rebuild the same website; one is in a different league

xda-developers.com · 14 September 2026

AI coding tools are getting quite good at building websites, but I wanted to know which current model can nail all the little details that I ask for in a prompt. So I gave Claude Code, Codex, and Google Antigravity the exact same challenge: rebuild a website from the same brief, with no extra hand-holding along the way.

On paper, all three should have produced something polished. In practice, the results were nowhere near as close as I expected.

A word about the prompt

And my preferred coding models

To keep this comparison as fair as possible, I am sticking with the default flagship coding agents available in each tool: Opus 5 in Claude Code, GPT-5.6 in Codex, and Gemini 3.8 in Google Antigravity. I am also giving all three the exact same prompt, with no extra hints or follow-up instructions during the first run.

This isn't a simple one-page portfolio test either. I picked a fairly complex food-delivery website with a dedicated hero section, restaurant filters, responsive layouts, animations, testimonials, and multiple interactive elements.

The prompt also asks each agent to make its own design decisions from scratch, so this test should reveal which tool turns a demanding brief into a polished website.

Google Antigravity

Gemini finally feels seriously competitive

Google clearly wasn’t kidding when it talked up the coding improvements in Gemini 3.8. I have used Gemini 3.5 inside Antigravity before, and the difference here is impressive.

The first thing that impressed me was the overall presentation. The generated food imagery looked good, the animations were smooth, without being distracting, and Gemini did a great job of giving different sections their own personality instead of turning the entire page into the usual collection of white cards.

I especially liked the offer and newsletter sections, where it used a dark gradient treatment that instantly made them stand out from the rest of the page. Those sections looked like something a designer had spent time refining rather than a generic AI-generated layout.

It wasn’t flawless, though. The FAQ section felt cramped. The mobile mockups were even worse. They looked laughably artificial compared to the rest of the site.

Still, these are minor misses. The overall output is excellent, and Antigravity generated everything at blazing-fast speed. It doesn’t quite take the crown, but it easily earns second place.

ChatGPT Codex

It looks polished until closer inspection

Codex made a strong first impression. The website looked clean, the layout was sensible, and GPT-5.6 followed my prompt closely without leaving out any of the major sections. The animations were decent, the food imagery worked well, and everything felt functional right from the first run.

The hero section was easily the highlight for me. In fact, I would say it was the best hero design out of all three tools. It used a large food image effectively, and layered in small floating details such as ratings, delivery information, and short taglines around it.

The problems started showing up once I looked more closely. Typography continues to be one of my biggest complaints with GPT-5.6 when I use it for website design. The fonts were small in several areas, like secondary text and supporting copy.

This is something I have noticed with GPT-5.6-generated interfaces, and I usually end up asking it to increase font sizes. The newsletter section was another weak point. It felt rushed, plain, and unfinished.

That sums up my experience with Codex here. It did almost everything I asked, and there were no major disasters. But it lacked the final layer of refinement that separates a good AI-generated website from one that looks professionally designed.

Claude Code

It wins on tiny details

Claude Code was the one that felt the most complete from the start. It wasn’t just that the website worked or that it followed my prompt properly. What stood out was how well it handled all the small design decisions that usually need another round of cleanup.

The typography felt right, the spacing was consistent, the padding around cards and sections was properly judged, and the image placement looked intentional.

The mobile mockups were impressive. This is one area where both GPT-5.6 and Gemini fall apart, but Claude Code got them almost spot on. They actually looked like believable app interfaces instead of generic phone frames with random content stuffed inside.

The live tracking detail made the whole experience feel much closer to a real food-delivery product. It was one of those small touches that leaves a strong impression in front of your client. Claude didn’t make any of the mistakes that I found with Gemini and GPT. The consistency carried across the entire page, which is what impressed me most.

If I had to nitpick, I still preferred Codex’s hero section. Claude Code’s version was good, but it didn’t have quite the same visual impact or those clever floating details around the main image. However, once I looked at the website as a whole, Claude Code was the obvious winner.

Three agents entered, one dominated

After putting all three through the exact same challenge, Claude Code came out on top for me. Antigravity was the biggest surprise and easily earned second place, while Codex impressed me with its hero section but fell short on overall polish.

What separated Claude Code was consistency. The typography, spacing, mockups, tracking section, FAQ, and newsletter all felt finished.

This text was published by xda-developers.com and written by Parth Shah. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories