Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?
How I AI
2025/12/03
Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?
Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?

How I AI
2025/12/03
Shownote
Shownote
I put three cutting-edge AI models to the test in a head-to-head design
competition. Using the exact same prompt, I challenged Google’s Gemini 3,
Anthropic’s Opus 4.5, and OpenAI’s Codex 5.1 to redesign my blog page,
evaluating them on visual design quality, user experience improvements, and SEO
optimization capabilities. One model produced a beautiful, polished,
production-ready redesign. One was fine. And one completely whiffed. If you’re
trying to figure out where each model fits in your workflow—design, planning,
back-end, or something else—this episode will save you a lot of trial and error.
What you’ll learn:
1. How each AI model approaches the same design challenge differently
2. Why planning capabilities dramatically impact design quality
3. The specific visual and functional improvements each model made
4. Which model excels at front-end design versus back-end functionality
5. How to strategically choose the right AI model for different parts of your
workflow
6. The importance of model-switching based on specific use cases
—
Blog design: https://www.chatprd.ai/blog [https://www.chatprd.ai/blog]
—
Brought to you by:
Lovable [https://lovable.dev/]—Build apps by simply chatting with AI
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/ [https://www.chatprd.ai/]
Website: https://clairevo.com/ [https://clairevo.com/]
LinkedIn: https://www.linkedin.com/in/clairevo/
[https://www.linkedin.com/in/clairevo/]
X: https://x.com/clairevo [https://x.com/clairevo]
—
In this episode, we cover:
(00:00) Introduction to the AI design challenge
(01:25) The question: Which model is the better designer?
(03:08) The prompt used for all three models
(04:10) Gemini 3 Pro’s approach and results
(06:00) Opus 4.5’s approach and results
(10:54) Codex 5.1’s approach and disappointing results
(14:51) Comparing the three designs side by side
(16:03) Analyzing the change logs and SEO improvements from each model
(22:43) Final verdict
(23:00) Conclusion and next steps
—
Tools referenced:
• Gemini 3 Pro: https://deepmind.google/models/gemini/pro/
[https://deepmind.google/models/gemini/pro/]
• Anthropic Opus 4.5: https://www.anthropic.com/news/claude-opus-4-5
[https://www.anthropic.com/news/claude-opus-4-5]
• OpenAI Codex 5.1: https://platform.openai.com/docs/models/gpt-5.1-codex
[https://platform.openai.com/docs/models/gpt-5.1-codex]
• Cursor: https://cursor.com/ [https://cursor.com/]
—
Production and marketing by https://penname.co/ [https://penname.co/]. For
inquiries about sponsoring the podcast, email jordan@penname.co.
Highlights
Highlights
In this episode, a direct comparison is made between three advanced AI models—Gemini 3 Pro, Opus 4.5, and Codex 5.1—tasked with redesigning a blog page from the same prompt. The focus is on evaluating their real-world performance in design quality, user experience, and technical execution.
Chapters
Chapters
Introduction to the AI design challenge
00:00The question: Which model is the better designer?
01:25The prompt used for all three models
03:08Gemini 3 Pro’s approach and results
04:10Opus 4.5’s approach and results
06:00Codex 5.1’s approach and disappointing results
10:54Comparing the three designs side by side
14:51Analyzing the change logs and SEO improvements from each model
16:03Final verdict
22:43Conclusion and next steps
23:00Transcript
Transcript
Claire Vo: Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive, here on a mission to help you build better with these new tools. Today, I have a really fun mini episode where I'm going to answer the question on everyone's mind. Which o...
