Jev has company
Three more decision models on the same reranking test, and how much the shape of the request changes the result.
In Hev meets Jev I used Jev, TypeSafe’s general-purpose decision model, as a reranker, and it landed next to models built for search. Since then I’ve run three more decision models through the same test: Cloudflare’s Clef and Clef-flash, and Perplexity’s pplx-decider-v1-27b.
The short version:
- Perplexity edges Jev. Asked to choose among the 30 documents,
pplx-deciderbeats Jev batch on precision at 1 on all three corpora, on nDCG@10 on two of them and on the three-corpus mean (0.5015 to 0.5009), for $0.40 per 1,000 queries against Jev’s $0.54. Jev’s one clear quality win is NFCorpus. - Jev still holds up. Jev batch scores 0.501 nDCG@10 in one call over 30 documents, gives each document its own score, and was the fastest batch call: a 192 ms median side by side, against 530 ms for Clef-flash and 2,008 ms for Perplexity batch. Mixedbread v3.1 listwise still leads overall at 0.511.
- Perplexity works best as a chooser. Its choice probabilities sum to one across the list, so they rank the documents but aren’t independent relevance scores. Asked about all 30 in one batch instead, it matches Jev’s quality but costs $11.59.
- Clef-flash needs one call per document. Scored one document at a time it reaches 0.498. Given all 30 at once it falls to 0.283, below the BM25 order it was meant to improve.
The test
Every system reranks the same BM25 top 30 for SciFact (300 queries), NFCorpus (323) and a fixed 300-query sample of FiQA. The score is nDCG@10, averaged with equal weight per corpus. The new models can be asked in three ways:
| Shape | What the model sees | What I sort by | Calls per query |
|---|---|---|---|
| Pair | Query + one document | Relevance probability | 30 |
| Batch | Query + 30 documents, 30 yes-or-no questions | 30 relevance probabilities | 1 |
| Choice | Query + 30 documents, one question | Probability per option | 1 |
The yes-or-no questions use the Noul type. In choice, the model spreads one unit of probability across the documents, so several documents can’t each score 0.9.
Clef and Perplexity share a prompt that counts both supporting and contradicting evidence as relevant. I picked it on 20 SciFact training queries before the test runs. A simpler prompt scored slightly better on that slice, but I chose the one that spelled out what relevant means. Jev keeps the original post’s generic prompt and results. The goal here is to keep the prompt consistent across models and approaches. I’d expect you could improve your specific use case quite a bit by improving the prompt, perhaps in a future experiment.
Quality against price
Quality against price
mean nDCG@10 · $ per 1,000 queries
Scroll sideways for the whole chart.
Perplexity choice is the strongest of the new decision models. It has the best three-corpus mean of any one-call decision shape, and at $0.40 per 1,000 queries it is the cheapest system on the chart that scores above 0.49. Jev batch ($0.54) and Clef-flash pair ($1.39 across its 30 calls) land close behind. The gaps are small. Perplexity choice’s leads over Jev batch on SciFact and FiQA have 95% intervals that include zero, and Jev batch’s 0.011 lead on NFCorpus is the only difference between the two that clears its interval (−0.021 to −0.002). Clef-flash pair’s intervals against Jev batch include zero on all three corpora.
Perplexity batch scores 0.500 but costs $11.59 per 1,000 queries, because its usage receipts count the shared context once for each of the 30 questions. Sending one HTTP request doesn’t mean paying for the context once. The price follows those receipts, whatever the model does internally.
| Reranker / shape | Mean nDCG@10 | $ / 1,000 queries | p50 / p95 per query |
|---|---|---|---|
| Mixedbread v3.1 listwise | 0.511 | $1.42 | 217 / 262 ms (laptop) |
| Voyage rerank-3 | 0.504 | $0.50 | 185 / 287 ms (laptop) |
| Jev · pair | 0.502 | $0.87 | 30 calls |
| Perplexity · choice | 0.501 | $0.40 | 521 / 737 ms (mini) |
| Jev · batch | 0.501 | $0.54 | 223 / 1,419 ms (laptop) |
| Perplexity · batch | 0.500 | $11.59 | 1,783 / 2,770 ms (mini) |
| Clef-flash · pair | 0.498 | $1.39 | 30 calls |
| Clef · pair | 0.494 | $3.69 | 30 calls |
| Jina v3.5* | 0.492 | $0.47 | 303 / 607 ms (laptop) |
| Perplexity · pair | 0.490 | $0.57 | 30 calls |
| Cohere v3.5 | 0.486 | $2.00 | 192 / 456 ms (laptop) |
| Mixedbread large-v2 | 0.476 | $1.42 | 361 / 429 ms (laptop) |
| Clef-flash · choice | 0.329 | $0.25 | 410 / 621 ms (mini) |
| Clef-flash · batch | 0.283 | $0.55 | 631 / 956 ms (mini) |
| BM25 order | 0.404 | no reranker |
Prices are each run’s recorded usage at list rates, normalized to 1,000 queries, before credits or discounts. Specialist rates and model versions, Mixedbread’s reconstructed token counts and Jina’s open-weight quality all carry over from the first post, which explains them.
Clef-flash needs the documents apart
Clef-flash by request shape
nDCG@10 · BM25 order marked
Scroll sideways for the whole chart.
Pair is competitive on every corpus, and batch and choice fall below the BM25 order on every corpus. Document order matters too. Reversing the candidate list changed Clef-flash’s top document on all 20 diagnostic queries per corpus, in both batch and choice, while repeating an unchanged request returned identical scores. The larger Clef model, run as pairs, averages 0.494 for $3.69, so size didn’t help.
Perplexity choice is also sensitive to order, though less so. Reversal changed its top result on 3 of 20 SciFact queries and 7 of 20 on each of the other corpora, and mean per-document rank correlations were 0.24 to 0.38. These are small diagnostic slices, and I haven’t ruled out API nondeterminism for Perplexity choice.
Latency
Quality against latency
mean nDCG@10 · p50 per query, whisker to p95
Scroll sideways for the whole chart.
This chart combines two setups. The September runs came from a laptop and the October runs from the mini, at different concurrency and pacing. Pair shapes are left out because a per-document latency doesn’t tell you how long 30 calls take, which depends on scheduling, concurrency and rate limits.
To compare batch calls fairly, the factory ran Jev, Clef-flash and Perplexity back to back on October 4, on the same 50 SciFact queries from the mini at concurrency 4, with request starts capped at 0.4 per second.
Batch latency, measured side by side
SciFact nDCG@10 · p50 per query, whisker to p95
Scroll sideways for the whole chart.
Jev’s p95 in that window was 233 ms, where the first post measured about 1.4 s. The machine, date and request conditions all changed, so I’m not calling the tail fixed yet.
Where this leaves Jev
Perplexity’s decider deserves the credit in this round. As a chooser it ranks at least as well as Jev for less money. Jev batch keeps an independent score for each document, a lead on NFCorpus and the fastest call in the side-by-side batch test. Pick Perplexity choice when you only need an order, and Jev when you need scores you can threshold or a fast single call. Clef-flash pair ranks well if you can afford 30 calls per query.
Two general-purpose decision models now land next to purpose-built rerankers. That’s a better showing than I expected from a reranking test, and reason enough to try both on other classifier tasks. It doesn’t predict how they will do there.
The broader factory comparison is still running, so this post covers only the finished Clef and Perplexity runs against the frozen baselines. The public benchmarks may be in some models’ training data, the shortlist is only 30 deep, and the bootstrap intervals cover query sampling with no correction for multiple comparisons.
Cloudflare documents Clef-flash, and Perplexity publishes a model card. Every number here is in the measurement snapshot.