Thibaud Colas tried to spend September coding with one efficient open-weight model: GLM 5.3 Flash. The challenge technically failed—only 1 billion of the month’s 2 billion tokens went through it—but the failure says more about operating coding agents than about model quality.

What happened

  • During the first half of the month, GLM 5.3 Flash handled Wagtail core work, sites, interface tasks, documentation, visual checks, and AI experiments for $68.
  • The provider reported about 4 kilowatt-hours of GPU energy use and 365 grams of carbon emissions for that work.
  • A prototype accidentally ran on the more expensive non-Flash GLM 5.3. It burned through 450 million tokens, $150, and about 5 kilowatt-hours nearly overnight.
  • GLM 5.3 Flash later slowed under apparent provider congestion, so Colas switched some work to DeepSeek V4.1 Flash and Qwen 3.8 Flash.
  • Benchmarking and research also required other models. The full month used about 35 kilowatt-hours instead of the 10 targeted.

The practical conclusion is not “pick one model.” It is to make cheap Flash-tier models the default for routine work, then budget separately for prototypes and model comparisons. Measure cost and energy against completed tasks, not raw token counts, because different models may use very different amounts of text to reach the same result.

The 135-comment thread on Hacker News clarifies two measurements and catches one misleading chart.

What the thread adds

  • ThibWeb — “Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget.”
  • ThibWeb — “Yes, all from Neuralwatt, GPU energy use only. Makes models’ ‘efficiency’ much more visible than tokens.”
  • andai — “That pareto graph isn’t very helpful because it shows cost per token, but some models are way more token hungry than others. There are also massive differences in speed to complete a task.” ThibWeb replied: “I’ve switched the visual to A index by cost per task.”
  • stkdump — “I have a computer with a 5090 on a smart plug at home running Qwen3.8 27B for agentic coding. On a busy day it can use around 5kWh, though on most days it is around 2kWh. […] But still I also believe that the energy use numbers you get from the inference providers are a bit ‘beautified’.”
  • surgical_fire — “GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan. It’s an excellent workhorse.”

HN handles are pseudonymous, and HN publishes no per-comment scores. The ordering is HN’s own ranking, so this is a slice of the thread rather than a consensus.