Cole Medin presents four practical ways to pair Jev with coding agents. His central pattern is not to replace the LLM: let the LLM build the harness and act on results, while Jev makes frequent, bounded decisions in the middle.

What Jev is for

  • Jev takes a description of the current state plus a fixed set of questions or choices, then returns answers with confidence scores.
  • Medin says this decision-only design is 20–200× faster and 40–1,000× cheaper than using a generative LLM for the same kind of classification; those are product claims, not independent measurements in this video.
  • It cannot generate free-form text or invent the action space. The workflow must define the state, legal choices, thresholds, and fallback behavior.

1. A security hook before tool use

  • Coding agents may read secrets, delete data, exfiltrate information, or drift away from the assigned task—even when instructions say not to.
  • A pre-tool-use hook can inspect the tool name, arguments, working directory, and intended effect before allowing an action.
  • Regex rules are fast but brittle: the same risky operation can be expressed as a shell command, Python script, direct file read, or many other forms.
  • An LLM judge is more flexible, but adding a model call before thousands of tool calls creates latency and cost.
  • Medin’s early comparison reports Jev at roughly a quarter-second per check, with fewer false positives than his regex hook and nearly all of his hand-labeled risky calls blocked. The video does not provide the dataset or raw results.

2. Playtesting as a user would

  • Unit tests cannot show whether a game actually feels playable or whether a feature breaks during live interaction.
  • Jev receives game state and chooses among legal actions quickly enough to attack, dodge, and navigate in real time.
  • The run can expose bugs that deterministic tests miss because it exercises the game through the same interaction loop as a player.
  • The LLM still builds the state adapter and action harness, then diagnoses and fixes whatever the playtest finds.

3. Faster browser testing

  • Browser automation is another bounded loop: inspect the current page, choose a button or field, execute the action, and observe the next state.
  • Jev can handle repeated choices such as what to focus or click; an LLM is still needed when a page requires original free-form text or broader visual judgment.
  • Medin demonstrates this split on Dino Chat: text is generated in advance, Jev chooses interface actions, and an LLM can review the resulting trace for bugs.

4. Route work before spending tokens

  • A GitHub issue can first be classified as a bug or feature, then routed to the matching skills and workflow.
  • A second decision selects a model tier based on task difficulty, avoiding an expensive frontier model for routine work.
  • Medin reports agreeing with Jev on all 12 issue-routing trials and on 15 of 16 pull-request review classifications. These are small, self-judged samples, not a production benchmark.

The reusable architecture

  1. An LLM defines the harness, state representation, and available actions.
  2. Jev makes frequent, narrow decisions inside that structure.
  3. Deterministic code enforces thresholds and executes approved actions.
  4. The LLM interprets failures, repairs code, or handles open-ended work.

“Jev is never really the end of our workflow. It’s always just making parts of our workflows faster and more efficient.”