Twitter/X

Author ran a benchmark (one-shot) that has AIs generate a single file containing…

Brief

The post presents a one-shot benchmark that asks models to produce a single file of procedurally-generated harbor towns spanning historical periods. The author compares outputs from GPT-5.6 Pro, Fable, Kimi K3, and Inkling, links an interactive gallery (ai-harbor-town-gallery.netli…), and calls the differences "surprisingly indicative."

Why it matters

Author ran a benchmark (one-shot) that has AIs generate a single file containing procedurally-generated harbor towns 'through history'; models included GPT-5.6 Pro, Fable, Kimi K3, and Inkling.

Key details

  • Benchmark and interactive gallery are publicly accessible at ai-harbor-town-gallery.netli…; author characterizes the results as "surprisingly indicative."
Source evidence

My benchmark where I have AIs create one file procedurally-generated harbor towns through history in one shot now has GPT-5.6 Pro, Fable, Kimi K3, and Inkling. You can play with all the simulations: ai-harbor-town-gallery.netli…

I think they are surprisingly indicative.