← Alle projekter

open source

jev-tools

Tre TypeScript-værktøjer, der sætter en billig, typet model foran den dyre

Et monorepo med værktøjer bygget på TypeSafe AI's Jev - en guardrail for agenters værktøjskald, en GitHub Action til issue-triage og et CLI til masseklassificering - der kører Jev på alt og kun sender de usikre tilfælde videre til et menneske eller en rigtig LLM.

Den fulde case er skrevet på engelsk.

Why Jev

Existing LLM bots are too slow and too expensive to run on every issue, row or tool call, and they hand back prose you still have to parse. Jev is a System One model: text in, typed probabilistic decisions out - choice, score, noul - each with a calibrated confidence, in roughly 70-500 ms at near-zero cost.

All three tools use it the same way: ask narrow questions, branch on the typed answer in code, and send anything Jev is unsure about to a human or a larger model. They share @cmaintz/jev-core (TypeSafe and Cloudflare providers, retries, timeouts, response validation), which replaced three identical copies of the provider code when the tools were merged into one repo with their history.

jev-guard

A fail-safe guardrail for an agent’s tool calls. Because Jev is effectively free per call, you can afford to guard every call, and because it reports confidence, uncertainty can fail safe. Each proposed call is scored on policy dimensions you declare (blast radius, destructiveness, exfiltration risk), and a small decision function maps the typed result to allow, block or hold.

  • One Jev request per guarded call asks all of the policy’s risk questions at once
  • Low confidence, missing answers and provider errors never resolve to allow
  • Adapters for LangChain JS (middleware) and the Vercel AI SDK, with structural types so neither framework is a dependency
  • Presets like shellPolicy(), an audit hook on every verdict, and an onHold handler for asking a human
  • Not a security boundary: think second opinion, not sandbox

jev-triage

A GitHub Action that labels every new issue in milliseconds for fractions of a cent, and escalates only the ones it’s unsure about. It asks a small set of bounded questions (type, priority, area, is-it-security) and applies labels when confidence clears the threshold.

  • Typed, confidence-gated labels, priority and a team mention. Routing is a plain map in code; Jev never picks the team
  • Low-confidence answers get triage:needs-human, or are re-asked to any OpenAI-compatible LLM, which leaves a one-line rationale
  • Duplicate detection: GitHub search finds candidates, Jev picks one or none, and nothing is auto-closed
  • Backlog sweep on a schedule, a security flag that alerts rather than labels, and a job summary with the estimated cost of the run
  • The jev-triage repo is a thin wrapper so it can sit on the Marketplace

jev-sort

jq for judgment: a CLI that streams JSONL/CSV rows through Jev and appends typed columns plus a confidence field. It earns its place at scale - 10k to 1M rows - where running an LLM per row is the problem. Below a few thousand rows a one-off LLM script is simpler, and if the rule is crisp, plain code wins.

  • Inline questions: -q 'team:choice(billing,tech,sales)', -q 'urgent:noul', -q 'size:score(low,mid,high)'
  • --escalate 'conf<0.6' writes uncertain rows to a review file, and --reject-out collects rows Jev couldn’t answer, so nothing disappears silently
  • Streams stdin to stdout, so it composes with jq, csvkit and friends. Also usable as a library

jev-eval measures where each tool’s confidence cut-point should sit, jev-rerank applies the same idea to RAG in Python, and Leash uses Jev to keep a coding agent to your project’s rules. Every package is gated by Foundry.

README.md - jev-tools▼

jev-tools

TypeSafe AI's Jev is a fast, very cheap model that takes text state and returns typed answers (choice / score / noul) with a probability or confidence instead of prose. A model that can't ramble seemed worth building on, so I wrapped it in a few TypeScript tools.

Unofficial: not affiliated with TypeSafe AI.

All three tools use Jev the same way: ask narrow questions, branch on the typed answer in code, and send anything Jev is unsure about to a human or a larger model.

Package What it does Distribution
@cmaintz/jev-guard Vets an LLM agent's tool calls before they run: allow / block / hold. Fails safe on low confidence, missing answers and provider errors. LangChain JS and Vercel AI SDK adapters. npm library
@cmaintz/jev-triage GitHub Action that labels, prioritises, routes and dedupes issues; low-confidence answers go to a human label or an LLM. GitHub Action
@cmaintz/jev-sort CLI + library that streams JSONL/CSV rows through Jev and appends typed columns and a _confidence field. npm CLI / library
@cmaintz/jev-core Shared Jev client used by the three above: question/answer types, TypeSafe and Cloudflare providers, retries, timeouts, response validation. npm library

Install

The npm packages are not published yet. They are set up for publishing (publishConfig with public access and npm provenance), but until the first release, install from source:

git clone https://github.com/CMaintz/jev-tools.git
cd jev-tools
npm ci
npm run build

Once published: npm install @cmaintz/jev-guard or npm install -g @cmaintz/jev-sort.

The Action needs no install. Reference it by path:

- uses: CMaintz/jev-tools/packages/triage@main # pin a jev-triage-v* tag or a commit SHA
  with:
    jev-api-key: ${{ secrets.JEV_API_KEY }}

Every tool needs a Jev key: a first-party key from TypeSafe, or a Cloudflare API token for Workers AI (typesafe/jev).

Usage

Guard an agent's shell tool:

import { shellPolicy, TypeSafeProvider, wrapTool } from '@cmaintz/jev-guard';

const provider = new TypeSafeProvider(process.env.JEV_API_KEY!);
const safeBash = wrapTool('bash', bash.execute, shellPolicy(), provider, {
  audit: (entry) => console.log(entry.verdict, entry.reasons),
});
await safeBash({ cmd: 'rm -rf /var/lib/postgresql/data' }); // throws GuardBlockedError unless allowed

Classify a JSONL file from the command line:

export JEV_API_KEY=...
jev-sort -q 'team:choice(billing,tech,sales)' -q 'urgent:noul' \
  --escalate 'conf<0.6' --review-out review.jsonl --reject-out rejects.jsonl \
  < tickets.jsonl > labeled.jsonl

Triage issues as they open: see packages/triage/examples/.

What Jev is and isn't

Figures below are TypeSafe's own published numbers; this repo does not benchmark them.

  • Speed and price: TypeSafe quotes 70-500 ms end to end, $0.042 per million input tokens, and free output tokens (announcement).
  • Accuracy: On TypeSafe's workflow evals, Jev scores 67.8%, measured as agreement with reference labels averaged from two frontier models across four example workflows. That is agreement with other models, not verified ground truth. Either way, treat answers as first-pass judgments; that is why every tool here gates on confidence.
  • Limits: Text only. 64k tokens per request, 32k for state plus the longest question; 40 requests/s (models page). No rationale, no generation, no arithmetic: counting, dates and routing tables stay in code.

Status

  • jev-guard 1.0, jev-triage 1.0 and jev-sort 1.0 were built as separate repos and merged here with their history (git log -- packages/<name>). jev-core is new: it replaces three identical copies of the provider code.
  • Unit tests run against a stubbed provider. The triage package has a live smoke test (packages/triage/test/live.test.ts) and the guard package a live smoke script (packages/guard/examples/smoke.mjs); both need JEV_API_KEY and are skipped in CI.
  • Not a security boundary: see the guard README. Think second opinion, not sandbox.

Development

npm ci
npm run lint        # eslint
npm run typecheck   # tsc per workspace
npm run test:coverage   # vitest across all packages, with the coverage floor
npm run build       # core → guard/sort (tsc) → triage (ncc bundle into packages/triage/dist)

mise run gate runs lint, typecheck, test and audit in the same order CI does. If you change packages/triage or packages/core, rebuild and commit packages/triage/dist: the Action runs the bundled file, and CI fails if it's out of date.

Releasing

Each package is versioned on its own and released when I decide it's ready, not on every merge. From a clean main:

bash scripts/cut-release.sh guard minor   # <core|guard|sort|triage> <major|minor|patch|X.Y.Z>, --dry-run to preview

It bumps that package's version, prepends a section to its CHANGELOG.md from the conventional commits that touched packages/<pkg> since its last tag, commits, tags <pkg>-vX.Y.Z (e.g. guard-v1.1.0), pushes and creates the GitHub release. A breaking commit without a major bump is refused. guard, sort and triage shipped 1.0.0 from their old repos, so before their first tag here the changelog starts at the commit that set 1.0.0. core has never been released; cut-release.sh core 0.1.0 tags it as is.

The tag push runs publish.yml, which publishes core, guard or sort to npm with trusted publishing (OIDC, no npm token in the repo) and provenance. triage is only tagged; Action users pin CMaintz/jev-tools/packages/triage@triage-vX.Y.Z or a SHA. Release core before anything that needs a new core, and widen the @cmaintz/jev-core range in guard, sort and triage when core leaves 0.1.x.

One-time setup per npm package (trusted publishing is configured on the package page, so the package has to exist first):

  1. Publish the current version by hand: npm login, npm run build, then npm publish --workspace packages/<pkg> --access public --provenance=false (provenance needs CI).
  2. On npmjs.com, package Settings > Trusted publishing > GitHub Actions: owner CMaintz, repository jev-tools, workflow publish.yml.

After that every tag publishes from CI. If a tag's version is already on npm (say core-v0.1.0 after the manual publish), the workflow skips it.

Other languages

Jev clients for other ecosystems live in their own repos: jev-java (Java SDK), jev-dotnet (.NET SDK) and jev-rerank (Python RAG reranker).

License

MIT © Christoffer Maintz

command.exeesc

↑↓ vælg · Tab udfyld · Enter kør · Esc luk

doom.exe

WASD move · ←→ turn · Space fire · E use · Shift run · Esc menu · click to capture mouse · licences