← All projects

open source

jev-rerank

Near-free relevance filtering and reranking for RAG

A zero-dependency Python library that scores each retrieved passage's relevance to the query with Jev in one batched call, then sorts and filters - at roughly a hundredth of the cost and latency of a hosted reranker.

Why

Rerankers like Cohere Rerank or a cross-encoder are a RAG staple, but they’re slow and cost real money per query. Jev’s Score question type gives each passage a relevance level plus a confidence, and adding questions to a call is near-free, so a whole candidate set can be scored at once.

Highlights

  • One batched Jev call per chunk: every passage becomes a Score question over shared {query, passages} state
  • min_score filters, top_n truncates, and batch_size keeps each call under Jev’s context cap
  • A few lines turn it into a LlamaIndex postprocessor or a LangChain reranker; first-class adapters are on the roadmap
  • Passes the Foundry gate (ruff, mypy, pytest, pip-audit)

Honest limits

It’s a cheap first pass, not a cross-encoder. For the last few points of ranking quality a dedicated cross-encoder still wins, so jev-rerank shines as a prefilter in front of one, or where hosted-reranker cost and latency are the actual problem. The TypeScript tools built on the same idea are jev-tools.

README.md - jev-rerank▼

jev-rerank

Fast, near-free RAG relevance filtering and reranking, powered by TypeSafe AI's Jev.

Rerankers (Cohere Rerank, cross-encoders) are a RAG staple, but they are slow and cost real money per query. jev-rerank scores each candidate passage's relevance to the query with Jev's Score primitive, in one batched call at roughly a hundredth of the cost and latency, then sorts and filters.

import os
from jev_rerank import rerank, TypeSafeProvider

provider = TypeSafeProvider(api_key=os.environ["JEV_API_KEY"])
ranked = rerank(query, passages, provider, top_n=5, min_score=2.0)
for p in ranked:
    print(p.score, p.text)

How it works

  • One batched Jev call per chunk: every passage becomes a Score question over shared {query, passages} state (adding questions is near-free, evaluated in parallel).
  • Each passage gets a relevance score on an ordered rubric (0..N) plus a confidence.
  • Results sort by score descending; min_score filters, top_n truncates. batch_size chunks large candidate sets so each call stays under Jev's context cap.

Drop-in for LlamaIndex / LangChain

The framework reranker interfaces are a few lines over rerank():

# LlamaIndex BaseNodePostprocessor
class JevPostprocessor(BaseNodePostprocessor):
    def _postprocess_nodes(self, nodes, query_bundle):
        texts = [n.get_content() for n in nodes]
        ranked = rerank(query_bundle.query_str, texts, provider, top_n=self.top_n)
        return [nodes[r.index] for r in ranked]

First-class jev-rerank[llamaindex] and [langchain] adapters are on the v0.2 roadmap.

Honest limitations

  • A cheap first pass, not a cross-encoder. Jev is ~68% accurate; for the last few points of ranking quality a dedicated cross-encoder still wins. jev-rerank shines as a cheap prefilter before an expensive reranker, or when Cohere-scale cost and latency are the actual problem.
  • Text only; each chunk must fit Jev's context (~32k tokens).
  • Needs a Jev key (JEV_API_KEY; sign up at TypeSafe).

Status

v0.1 - the core rerank() + JevReranker + zero-dependency provider (stdlib urllib), batched and chunked. Passes the Foundry gate (ruff, mypy, pytest, pip-audit). Roadmap: first-class LlamaIndex/LangChain adapters, async batching, and score caching.

MIT (c) Christoffer Maintz

command.exeesc

↑↓ select · Tab complete · Enter run · Esc close

doom.exe

WASD move · ←→ turn · Space fire · E use · Shift run · Esc menu · click to capture mouse · licences