Nish Tahir takes apart the idea behind a Jev-style decision model: instead of asking a language model to write an answer, give it a fixed list and inspect which option it would generate next. His walkthrough uses a small Qwen model to show both the shortcut and the catch. Choosing a valid answer is not the same as choosing a correct one—or knowing how likely it is to be correct.

  • Less generation: restrict the model to single-token labels such as A through E. It can choose a label after processing the prompt, without generating a multi-token JSON response.
  • Accuracy still needs testing: on a held-out sample of CommonsenseQA, a collection of everyday multiple-choice questions, Tahir reports accuracy of 0.5938 before fine-tuning and 0.6241 afterward.
  • Confidence is a separate problem: in his initial calibration table, predictions in the highest confidence band average 0.9855 confidence but only 0.7009 accuracy. The probability of generating a label is not automatically the probability that the answer is right.
  • A repair, not a smarter model: temperature scaling spreads out the probabilities so they better match observed accuracy. Tahir fits a temperature parameter and shows a substantially closer match in the resulting table.

The useful distinction is between restricting outputs and earning trust in them. A model that can only answer A, B, C, D or E cannot wander off-format; it can still confidently select the wrong label. The walkthrough makes that difference visible rather than hiding it behind a clean API.

The 79-comment thread on Hacker News adds an application example, a dispute about the novelty, and questions about the calibration procedure.

What the thread adds

  • xavortm — describes using Jev to screen text before sending it to a larger model: “an LLM only starts work once I’ve chunked text and marked it as ’to be reviewed’ instead of parsing the full content.” Their reported saving is a commenter’s account, not a result from Tahir’s experiment.
  • howunfortunate — pushes back on the rebranding: “plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now”. Zero-shot means choosing categories without first training on examples of that particular task.
  • jofzar — replies with a counterargument about economics rather than novelty: “It’s the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past”. That is their assessment, not an independently verified comparison.

The calibration question left open

Two commenters probe different parts of the same issue: nl asks, “And how do you curve fit for this single example?”; demibabs asks, “Is simply changing the temperature so that the model appears calibrated over a particular benchmark after the fact “allowed”?” The article reports the fitted parameter and tables, but does not spell out a separate calibration-fitting sample and a fresh test sample. Those questions are worth keeping alongside the appealing before-and-after result.

HN handles are pseudonymous, and HN publishes no per-comment scores. Ordering is HN’s own ranking; this is a slice of the thread, not a consensus.