YouTube

GPT-5.5 vs Opus 4.8: The Real Winner Was /goal Mode

Brief

Matt Maher (presentation/demo) compares GPT-5.5 and Opus 4.8 by having each build a native Swift macOS radial dial launcher across two modes (straight and /goal). He documents a design workflow using Claude Design and an image model, tries Open Design, and concludes Claude (Opus 4.8) won while /goal unexpectedly flipped token-cost and performance tradeoffs.

Why it matters

On 2026-07-05 Matt Maher ran four builds (GPT-5.5 Codex Extra High and Opus 4.8 Claude Code Extra, each in straight and /goal modes) to implement a native Swift macOS radial “dial” launcher; Opus 4.8 (Claude) produced the overall winning build while one of the four builds was unusable.

Key details

  • Maher's design workflow: ideation with an image model, initial UI/design work in Claude Design, iteration to 'Design 2.0', then refinement; he also tested Open Design (an open alternative) and found it genuinely impressive.
  • /goal mode materially changed outcomes and token-cost behavior in an unexpected way—contrary to common assumptions—altering which model/mode was most effective; GPT-5.5 remained fast and high-quality but lost the head-to-head to Claude.
Source evidence

I gave GPT-5.5 and Opus 4.8 the exact same brutal job — build a real native macOS app in Swift — then ran each of them two ways: a straight build, and in /goal mode. Claude took the win. But the thing that actually surprised me wasn't which model won.

The app is a radial "dial" launcher: rings of apps and modes you fly around with the arrow keys, that can open files, apps, and websites. It's a genuinely hard spec — especially in Swift — which is exactly what makes it a good test of where these models really are.

Two models, two modes, four builds: • GPT-5.5 (Codex, Extra High) — straight + /goal • Opus 4.8 (Claude Code, Extra) — straight + /goal

One build came out unusable. One I wanted to keep using the second I opened it. And when I checked the token cost, /goal mode did the opposite of what almost everyone assumes it does.

I also walk through how I designed the app — starting in Claude Design, going wide with an image model for ideation, then back into Claude Design to lock the look — plus a newer, open alternative I just stumbled onto called Open Design that genuinely impressed me.

GPT-5.5 was excellent, for the record. This was close. But /goal changed the math.

🔗 Links
• The CARE benchmark — my open model benchmark, with live scores you can dig into right now:
https://carebenchmark.metal-sole.com/
• Claude Design: https://claude.ai/design
• Open Design: https://open-design.ai/

Going forward I'm going to let a few other models take a swing at building this same app — so we get a feel for how they actually build, not just how they score. Subscribe if you want to see who's next.
00:00 - Intro
00:45 - What we will build... sort of
02:15 - Kicking off the build... plus GOAL
03:26 - How did I design this?
03:57 - Claude Design
05:24 - Design 2.0 move
06:20 - Leveling up Claude Design
09:05 - Open Design?!
09:59 - App Builds
10:46 - GPT, as always, is faster
12:05 - Let's talk about "/goal"
14:31 - RESULTS : Claude straight
15:11 - RESULTS : Codex straight
16:39 - RESULTS : Codex GOAL
18:19 - RESULTS : Claude GOAL
20:06 - Cost? Token burn?
22:35 - Conclusion

Channel: Matt Maher
Published: 2026-07-05
Video URL: https://www.youtube.com/watch?v=uWauLhx6pWk