Cole Medin presents four practical ways to pair Jev with coding agents. His central pattern is not to replace the LLM: let the LLM build the harness and act on results, while Jev makes frequent, bounded decisions in the middle.
What Jev is for
- Jev takes a description of the current state plus a fixed set of questions or choices, then returns answers with confidence scores.
- Medin says this decision-only design is 20–200× faster and 40–1,000× cheaper than using a generative LLM for the same kind of classification; those are product claims, not independent measurements in this video.
- It cannot generate free-form text or invent the action space. The workflow must define the state, legal choices, thresholds, and fallback behavior.
1. A security hook before tool use
- Coding agents may read secrets, delete data, exfiltrate information, or drift away from the assigned task—even when instructions say not to.
- A pre-tool-use hook can inspect the tool name, arguments, working directory, and intended effect before allowing an action.
- Regex rules are fast but brittle: the same risky operation can be expressed as a shell command, Python script, direct file read, or many other forms.
- An LLM judge is more flexible, but adding a model call before thousands of tool calls creates latency and cost.
- Medin’s early comparison reports Jev at roughly a quarter-second per check, with fewer false positives than his regex hook and nearly all of his hand-labeled risky calls blocked. The video does not provide the dataset or raw results.
2. Playtesting as a user would
- Unit tests cannot show whether a game actually feels playable or whether a feature breaks during live interaction.
- Jev receives game state and chooses among legal actions quickly enough to attack, dodge, and navigate in real time.
- The run can expose bugs that deterministic tests miss because it exercises the game through the same interaction loop as a player.
- The LLM still builds the state adapter and action harness, then diagnoses and fixes whatever the playtest finds.
3. Faster browser testing
- Browser automation is another bounded loop: inspect the current page, choose a button or field, execute the action, and observe the next state.
- Jev can handle repeated choices such as what to focus or click; an LLM is still needed when a page requires original free-form text or broader visual judgment.
- Medin demonstrates this split on Dino Chat: text is generated in advance, Jev chooses interface actions, and an LLM can review the resulting trace for bugs.
4. Route work before spending tokens
- A GitHub issue can first be classified as a bug or feature, then routed to the matching skills and workflow.
- A second decision selects a model tier based on task difficulty, avoiding an expensive frontier model for routine work.
- Medin reports agreeing with Jev on all 12 issue-routing trials and on 15 of 16 pull-request review classifications. These are small, self-judged samples, not a production benchmark.
The reusable architecture
- An LLM defines the harness, state representation, and available actions.
- Jev makes frequent, narrow decisions inside that structure.
- Deterministic code enforces thresholds and executes approved actions.
- The LLM interprets failures, repairs code, or handles open-ended work.
“Jev is never really the end of our workflow. It’s always just making parts of our workflows faster and more efficient.”