← writing

Jev has company

Three more decision models on the same reranking test, and how much the shape of the request changes the result.

In Hev meets Jev I used Jev, TypeSafe’s general-purpose decision model, as a reranker, and it landed next to models built for search. Since then I’ve run three more decision models through the same test: Cloudflare’s Clef and Clef-flash, and Perplexity’s pplx-decider-v1-27b.

The short version:

  • Perplexity edges Jev. Asked to choose among the 30 documents, pplx-decider beats Jev batch on precision at 1 on all three corpora, on nDCG@10 on two of them and on the three-corpus mean (0.5015 to 0.5009), for $0.40 per 1,000 queries against Jev’s $0.54. Jev’s one clear quality win is NFCorpus.
  • Jev still holds up. Jev batch scores 0.501 nDCG@10 in one call over 30 documents, gives each document its own score, and was the fastest batch call: a 192 ms median side by side, against 530 ms for Clef-flash and 2,008 ms for Perplexity batch. Mixedbread v3.1 listwise still leads overall at 0.511.
  • Perplexity works best as a chooser. Its choice probabilities sum to one across the list, so they rank the documents but aren’t independent relevance scores. Asked about all 30 in one batch instead, it matches Jev’s quality but costs $11.59.
  • Clef-flash needs one call per document. Scored one document at a time it reaches 0.498. Given all 30 at once it falls to 0.283, below the BM25 order it was meant to improve.

The test

Every system reranks the same BM25 top 30 for SciFact (300 queries), NFCorpus (323) and a fixed 300-query sample of FiQA. The score is nDCG@10, averaged with equal weight per corpus. The new models can be asked in three ways:

ShapeWhat the model seesWhat I sort byCalls per query
PairQuery + one documentRelevance probability30
BatchQuery + 30 documents, 30 yes-or-no questions30 relevance probabilities1
ChoiceQuery + 30 documents, one questionProbability per option1

The yes-or-no questions use the Noul type. In choice, the model spreads one unit of probability across the documents, so several documents can’t each score 0.9.

Clef and Perplexity share a prompt that counts both supporting and contradicting evidence as relevant. I picked it on 20 SciFact training queries before the test runs. A simpler prompt scored slightly better on that slice, but I chose the one that spelled out what relevant means. Jev keeps the original post’s generic prompt and results. The goal here is to keep the prompt consistent across models and approaches. I’d expect you could improve your specific use case quite a bit by improving the prompt, perhaps in a future experiment.

Quality against price

Quality against price

mean nDCG@10 · $ per 1,000 queries

0.47 0.48 0.49 0.50 0.51 0.52 $0.25 $0.50 $1 $2 $5 $10 nDCG@10, mean of SciFact, NFCorpus, FiQA cost per 1,000 queries at list price, log scale Jev · batch Jev · pair Mixedbread v3.1 listwise Voyage rerank-3 Cohere v3.5 Mixedbread large-v2 Jina v3.5* Clef-flash · pair Clef · pair Perplexity · batch Perplexity · choice Perplexity · pair

Scroll sideways for the whole chart.

Every reranker and request shape. Pair prices include all 30 calls. Clef-flash · batch (0.283) and Clef-flash · choice (0.329) score below the BM25 order (0.404) and sit off the bottom of the chart. Jina's quality is from its open weights. Hover a point, or pick a family in the legend.

Perplexity choice is the strongest of the new decision models. It has the best three-corpus mean of any one-call decision shape, and at $0.40 per 1,000 queries it is the cheapest system on the chart that scores above 0.49. Jev batch ($0.54) and Clef-flash pair ($1.39 across its 30 calls) land close behind. The gaps are small. Perplexity choice’s leads over Jev batch on SciFact and FiQA have 95% intervals that include zero, and Jev batch’s 0.011 lead on NFCorpus is the only difference between the two that clears its interval (−0.021 to −0.002). Clef-flash pair’s intervals against Jev batch include zero on all three corpora.

Perplexity batch scores 0.500 but costs $11.59 per 1,000 queries, because its usage receipts count the shared context once for each of the 30 questions. Sending one HTTP request doesn’t mean paying for the context once. The price follows those receipts, whatever the model does internally.

Reranker / shapeMean nDCG@10$ / 1,000 queriesp50 / p95 per query
Mixedbread v3.1 listwise0.511$1.42217 / 262 ms (laptop)
Voyage rerank-30.504$0.50185 / 287 ms (laptop)
Jev · pair0.502$0.8730 calls
Perplexity · choice0.501$0.40521 / 737 ms (mini)
Jev · batch0.501$0.54223 / 1,419 ms (laptop)
Perplexity · batch0.500$11.591,783 / 2,770 ms (mini)
Clef-flash · pair0.498$1.3930 calls
Clef · pair0.494$3.6930 calls
Jina v3.5*0.492$0.47303 / 607 ms (laptop)
Perplexity · pair0.490$0.5730 calls
Cohere v3.50.486$2.00192 / 456 ms (laptop)
Mixedbread large-v20.476$1.42361 / 429 ms (laptop)
Clef-flash · choice0.329$0.25410 / 621 ms (mini)
Clef-flash · batch0.283$0.55631 / 956 ms (mini)
BM25 order0.404no reranker

Prices are each run’s recorded usage at list rates, normalized to 1,000 queries, before credits or discounts. Specialist rates and model versions, Mixedbread’s reconstructed token counts and Jina’s open-weight quality all carry over from the first post, which explains them.

Clef-flash needs the documents apart

Clef-flash by request shape

nDCG@10 · BM25 order marked

0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 nDCG@10 per corpus SciFact pair choice batch BM25 NFCorpus FiQA

Scroll sideways for the whole chart.

Clef-flash on each corpus. Pair scores one document per call; batch asks 30 yes-or-no questions in one call; choice asks it to pick among the 30. The bar is the BM25 order the reranker starts from.

Pair is competitive on every corpus, and batch and choice fall below the BM25 order on every corpus. Document order matters too. Reversing the candidate list changed Clef-flash’s top document on all 20 diagnostic queries per corpus, in both batch and choice, while repeating an unchanged request returned identical scores. The larger Clef model, run as pairs, averages 0.494 for $3.69, so size didn’t help.

Perplexity choice is also sensitive to order, though less so. Reversal changed its top result on 3 of 20 SciFact queries and 7 of 20 on each of the other corpora, and mean per-document rank correlations were 0.24 to 0.38. These are small diagnostic slices, and I haven’t ruled out API nondeterminism for Perplexity choice.

Latency

Quality against latency

mean nDCG@10 · p50 per query, whisker to p95

○ September, laptop · ◆ October, mini
0.47 0.48 0.49 0.50 0.51 0.52 100 ms 200 ms 500 ms 1 s 2 s 5 s nDCG@10, mean of SciFact, NFCorpus, FiQA latency per query, p50 with a whisker to p95, log scale Jev · batch Mixedbread v3.1 listwise Voyage rerank-3 Cohere v3.5 Mixedbread large-v2 Jina v3.5* Perplexity · batch Perplexity · choice

Scroll sideways for the whole chart.

One-call shapes only, each with the timing from its own run. Hollow circles are September runs from a laptop and diamonds are October runs from the mini, at different concurrency and pacing, so this is context rather than a speed ranking. Clef-flash · batch (0.283) and Clef-flash · choice (0.329) are off the bottom. Jina's timing is SciFact only.

This chart combines two setups. The September runs came from a laptop and the October runs from the mini, at different concurrency and pacing. Pair shapes are left out because a per-document latency doesn’t tell you how long 30 calls take, which depends on scheduling, concurrency and rate limits.

To compare batch calls fairly, the factory ran Jev, Clef-flash and Perplexity back to back on October 4, on the same 50 SciFact queries from the mini at concurrency 4, with request starts capped at 0.4 per second.

Batch latency, measured side by side

SciFact nDCG@10 · p50 per query, whisker to p95

0.4 0.5 0.6 0.7 0.8 0 500 ms 1 s 1.5 s 2 s 2.5 s 3 s 3.5 s nDCG@10, SciFact latency per query, p50 with a whisker to p95 Perplexity · batch 2.0 s / 3.2 s Jev · batch 192 ms / 233 ms Clef-flash · batch 530 ms / 885 ms

Scroll sideways for the whole chart.

Batch shapes only, on October 4: the same 50 SciFact queries from the mini, concurrency 4, request starts capped at 0.4 per second, pacing delay excluded. Quality is from the full 300-query SciFact runs. Whiskers run to p95; they are not confidence intervals.

Jev’s p95 in that window was 233 ms, where the first post measured about 1.4 s. The machine, date and request conditions all changed, so I’m not calling the tail fixed yet.

Where this leaves Jev

Perplexity’s decider deserves the credit in this round. As a chooser it ranks at least as well as Jev for less money. Jev batch keeps an independent score for each document, a lead on NFCorpus and the fastest call in the side-by-side batch test. Pick Perplexity choice when you only need an order, and Jev when you need scores you can threshold or a fast single call. Clef-flash pair ranks well if you can afford 30 calls per query.

Two general-purpose decision models now land next to purpose-built rerankers. That’s a better showing than I expected from a reranking test, and reason enough to try both on other classifier tasks. It doesn’t predict how they will do there.

The broader factory comparison is still running, so this post covers only the finished Clef and Perplexity runs against the frozen baselines. The public benchmarks may be in some models’ training data, the shortlist is only 30 deep, and the bootstrap intervals cover query sampling with no correction for multiple comparisons.

Cloudflare documents Clef-flash, and Perplexity publishes a model card. Every number here is in the measurement snapshot.

Start typing to search.