Piotr Migdał makes interactive visualizations. He gave two coding agents the same one-shot task: build an explorable 3D rendering of the 55 imaginary cities in Italo Calvino’s Invisible Cities, with up to six hours and no follow-up questions. This is his subjective design experiment, not a controlled model test.

  • GPT-6 Astra, running in Codex, finished in 53 minutes for about $10 in API usage. It worked end to end, but Migdał found unnecessary visual noise and text that stated the obvious.
  • Claude Opus 5.5, running in Claude Code with six parallel helper agents, finished in 1 hour 25 minutes for about $74. Migdał was impressed by the result, though the agent did not actually spend the six hours it was allotted.
  • His earlier AI-assisted visualizations had needed substantial human cleanup. This one makes him ask where his own creative work fits when an agent can produce such a polished draft.

The result is a demonstration of a new starting point, not evidence that a model understands Calvino or that a designer’s job is done. Migdał’s question about his role is more interesting than the model comparison.

The 127-comment thread on Hacker News tests the gap between first impression and a closer look.

What the thread adds

  • SiempreViernes — reports inspecting a city whose text celebrates bridges, only to find several rendered bridges that did not cross anything. This is one reader’s concrete counterexample to judging the work from a thumbnail.
  • toddmorey — argues that the interactive examples can look impressive without teaching anything: “They are pretty but not helpful.” For an explorable explanation, what the visitor learns matters as much as visual polish.
  • sailingparrot — contrasts wanting to inspect the details of human-made art with finding “bridges that make no sense” when zooming into AI output. That is a critique of craft, not a measurement of either model.
  • ouchouchhaha — pushes back on judging it as a finished artwork: “they wanted to show off what Claude/GPT could do in 6hrs.” They suggest human iteration on each city would be a different project.
  • twoodfin — calls the one-shot result “the new floor,” arguing that human creativity and further work could make something beyond the first draft. That is a prediction, not a result of this test.

Several commenters, including droidjj and uludag, worry that literal renderings narrow what readers might imagine from Calvino’s book. It’s a different standard from “can the agent generate 55 working scenes?” — and the essay does not test it.

HN handles are pseudonymous; HN publishes no per-comment scores. The ordering is HN’s own ranking, so these remarks are a slice of the discussion, not a consensus.