OpenAI makes four
OpenAI's new Decisions API on the same reranking test: a credible fourth decision model that trails the other three.
In Hev meets Jev I used Jev, TypeSafe’s decision model, as a reranker. In Jev has company Cloudflare’s Clef and Perplexity’s pplx-decider took the same test. Now OpenAI has a decision model too: the Decisions API, in public beta, serving gpt-6-luna. It went through the same test, with the same shortlists and the same prompt, and no tuning.
The short version:
- It trails the other three. Its best shape, asking it to choose among the 30 documents, averages 0.472 nDCG@10 over SciFact, NFCorpus and FiQA. Jev batch and Perplexity choice both score 0.501. The paired 95% interval against Jev excludes zero on every corpus.
- It still improves on BM25 on two corpora. Choice lifts SciFact from 0.667 to 0.751 and FiQA from 0.234 to 0.355. On NFCorpus no shape clearly beats the BM25 order.
- Choice is the shape to use. It is both the best and the cheapest shape, at about $0.92 per 1,000 queries. Batch and pair cost about $1.55, and score lower.
- The probabilities are too sharp to prune with. At a 0.1 threshold, batch keeps 6% of SciFact candidates but drops a quarter of the relevant ones. The threshold that worked for Jev doesn’t carry over.
- It’s a beta. The API refused some choice questions outright, and OpenAI may change the model before general availability.
The test
Nothing changed from the last post. Every system reranks the same BM25 top 30 for SciFact (300 queries), NFCorpus (323) and a fixed 300-query sample of FiQA, scored by nDCG@10 with equal weight per corpus. OpenAI got the same three request shapes as Clef and Perplexity:
| Shape | What the model sees | What I sort by | Calls per query |
|---|---|---|---|
| Pair | Query + one document | Relevance probability | 30 |
| Batch | Query + 30 documents, 30 yes-or-no questions | 30 relevance probabilities | 1 |
| Choice | Query + 30 documents, one question | Probability per option | 1 |
The Decisions API calls its yes-or-no question a predicate, and it has no field for true and false criteria, so the criteria from the shared Clef and Perplexity prompt go into the predicate’s instructions. One 20-query probe on SciFact’s training split confirmed the wording worked. I didn’t run a prompt search and didn’t retune.
Results
Quality against price
mean nDCG@10 · $ per 1,000 queries
Scroll sideways for the whole chart.
| Reranker / shape | Mean nDCG@10 | $ / 1,000 queries | p50 / p95 per query |
|---|---|---|---|
| Mixedbread v3.1 listwise | 0.511 | $1.42 | 217 / 262 ms (laptop) |
| Voyage rerank-3 | 0.504 | $0.50 | 185 / 287 ms (laptop) |
| Jev · pair | 0.502 | $0.87 | 30 calls |
| Perplexity · choice | 0.501 | $0.40 | 521 / 737 ms (mini) |
| Jev · batch | 0.501 | $0.54 | 223 / 1,419 ms (laptop) |
| Perplexity · batch | 0.500 | $11.59 | 1,783 / 2,770 ms (mini) |
| Clef-flash · pair | 0.498 | $1.39 | 30 calls |
| Clef · pair | 0.494 | $3.69 | 30 calls |
| Jina v3.5* | 0.492 | $0.47 | 303 / 607 ms (laptop) |
| Perplexity · pair | 0.490 | $0.57 | 30 calls |
| Cohere v3.5 | 0.486 | $2.00 | 192 / 456 ms (laptop) |
| Mixedbread large-v2 | 0.476 | $1.42 | 361 / 429 ms (laptop) |
| OpenAI · choice | 0.472 | $0.92 | 255 / 498 ms (mini) |
| OpenAI · batch | 0.450 | $1.55 | 346 / 855 ms (mini) |
| OpenAI · pair | 0.441 | $1.58 | 30 calls |
| Clef-flash · choice | 0.329 | $0.25 | 410 / 621 ms (mini) |
| Clef-flash · batch | 0.283 | $0.55 | 631 / 956 ms (mini) |
| BM25 order | 0.404 | no reranker |
OpenAI prices are estimates: reported input tokens at the $0.10 per million list price, with no output charge. The rest carry over from the earlier posts.
Per corpus, choice is the best shape on SciFact (0.751) and FiQA (0.355), and batch is best on NFCorpus (0.323). Against Jev batch those best shapes trail by 0.018 on SciFact, 0.035 on NFCorpus and 0.021 on FiQA, and every interval excludes zero. They also trail Perplexity’s best shape by 0.025 to 0.042 and Clef-flash pair by 0.017 to 0.044. These gaps are small, but unlike the Jev-versus-Perplexity race in the last post they all point one way.
One comparison is closer. The first post ran gpt-5.6-luna, an older model, as a plain listwise LLM reranker on SciFact. OpenAI choice ties it there (+0.004, interval spanning zero). On the other two corpora I only kept that run’s averages, which are a bit higher than Decisions, so there is no paired comparison.
Choice beats batch
OpenAI by request shape
nDCG@10 · BM25 order and Jev batch marked
Scroll sideways for the whole chart.
With Clef-flash, batch and choice both collapsed below the BM25 order. OpenAI doesn’t collapse, but the shapes still split. On SciFact, choice scores 0.751 while batch and pair score 0.681 and 0.670, and neither of those clearly beats BM25’s 0.667. Asking the model to spread one unit of probability across the list works better than asking 30 separate yes-or-no questions about it.
Choice is also the cheapest shape, because one question over 30 options takes fewer input tokens than 30 predicates. Pair and batch cost about the same as each other. Unlike Perplexity batch, which billed the shared context once per question ($11.59 per 1,000 queries), OpenAI batch is billed close to its actual input.
Sharp probabilities, broken prune
The first post’s argument for Jev was that each document gets a calibrated probability, so one call both reranks and prunes: drop everything under 0.1 and lose almost nothing. OpenAI’s pair and batch probabilities are calibrated well on average (SciFact batch has a Brier score of 0.026), but they are sharply peaked:
- SciFact batch keeps 5.7% of candidates at p ≥ 0.1, and loses 26% of the relevant documents in the shortlist.
- NFCorpus batch keeps 19.5% and loses 29%.
- FiQA batch keeps 23.6% and loses only 6%, but at 12% precision.
Pruning at p ≥ 0.1
share of relevant documents kept · share of candidates kept
Scroll sideways for the whole chart.
Jev batch, at the same threshold, kept 94% of SciFact’s relevant documents, 91% of NFCorpus’s and nearly all of FiQA’s, while dropping 46% to 85% of the candidates. OpenAI prunes harder and loses more. So the threshold mostly fails. If you want to prune with these scores, you’d need to tune a lower cutoff per corpus, which is the work a calibrated probability was supposed to save.
Order and refusals
Same-order reissues returned identical scores, so the API is deterministic for a fixed request. Candidate order still matters. Reversing the list kept batch rankings fairly stable (mean rank correlation 0.76 to 0.79) but moved choice more (0.48 to 0.69), and changed the top document on 3 to 11 of 20 diagnostic queries per corpus. That’s less sensitive than Clef-flash and in the same range as Perplexity choice.
The API also declined some choice questions outright: 4 SciFact queries, 18 NFCorpus and 1 FiQA. I score a refusal as a uniform distribution, which leaves those queries in BM25 order. NFCorpus is medical, so that’s where a safety filter would be most likely to fire, and it is also where choice did worst.
Latency
Quality against latency
mean nDCG@10 · p50 per query, whisker to p95
Scroll sideways for the whole chart.
The full runs on the mini, at concurrency 6, put OpenAI choice at 255 ms median and 498 ms at p95, and batch at 346 ms and 855 ms, averaged over the three corpora. Choice is about twice as fast as Perplexity choice (521 ms) and close to Jev batch (223 ms median in September, 192 ms in the October side-by-side test).
I also tried to time Jev and OpenAI batch side by side, as the last post did for Jev, Clef-flash and Perplexity. OpenAI batch took 406 ms median and 698 ms at p95 over 50 SciFact calls, but the Jev account had run out of credits and returned an error, so there is no matched Jev number. The mini was also under heavy load during that window, so none of these numbers is a controlled speed ranking.
Where this leaves things
Four general-purpose decision models have now taken a reranking test built for specialist models. Jev and Perplexity sit with the specialists. Clef-flash gets there if you can afford 30 calls per query. OpenAI’s Decisions API is a credible fourth: choice beats BM25 on two of three corpora for under a dollar per 1,000 queries. But it trails the other three everywhere, and its probabilities don’t support the prune that made Jev interesting in the first place.
The usual limits apply. The public benchmarks may be in some models’ training data, the shortlist is only 30 deep, the intervals cover query sampling with no correction for multiple comparisons, and the prompt was chosen for consistency across models, not tuned for this one. The API is a beta, so I’ll rerun it when it reaches general availability.
Every run, with per-query rankings and the receipts behind the cost figures, is in the reranker comparison.