← writing

OpenAI makes four

OpenAI's new Decisions API on the same reranking test: a credible fourth decision model that trails the other three.

In Hev meets Jev I used Jev, TypeSafe’s decision model, as a reranker. In Jev has company Cloudflare’s Clef and Perplexity’s pplx-decider took the same test. Now OpenAI has a decision model too: the Decisions API, in public beta, serving gpt-6-luna. It went through the same test, with the same shortlists and the same prompt, and no tuning.

The short version:

  • It trails the other three. Its best shape, asking it to choose among the 30 documents, averages 0.472 nDCG@10 over SciFact, NFCorpus and FiQA. Jev batch and Perplexity choice both score 0.501. The paired 95% interval against Jev excludes zero on every corpus.
  • It still improves on BM25 on two corpora. Choice lifts SciFact from 0.667 to 0.751 and FiQA from 0.234 to 0.355. On NFCorpus no shape clearly beats the BM25 order.
  • Choice is the shape to use. It is both the best and the cheapest shape, at about $0.92 per 1,000 queries. Batch and pair cost about $1.55, and score lower.
  • The probabilities are too sharp to prune with. At a 0.1 threshold, batch keeps 6% of SciFact candidates but drops a quarter of the relevant ones. The threshold that worked for Jev doesn’t carry over.
  • It’s a beta. The API refused some choice questions outright, and OpenAI may change the model before general availability.

The test

Nothing changed from the last post. Every system reranks the same BM25 top 30 for SciFact (300 queries), NFCorpus (323) and a fixed 300-query sample of FiQA, scored by nDCG@10 with equal weight per corpus. OpenAI got the same three request shapes as Clef and Perplexity:

ShapeWhat the model seesWhat I sort byCalls per query
PairQuery + one documentRelevance probability30
BatchQuery + 30 documents, 30 yes-or-no questions30 relevance probabilities1
ChoiceQuery + 30 documents, one questionProbability per option1

The Decisions API calls its yes-or-no question a predicate, and it has no field for true and false criteria, so the criteria from the shared Clef and Perplexity prompt go into the predicate’s instructions. One 20-query probe on SciFact’s training split confirmed the wording worked. I didn’t run a prompt search and didn’t retune.

Results

Quality against price

mean nDCG@10 · $ per 1,000 queries

0.43 0.45 0.47 0.49 0.51 $0.25 $0.50 $1 $2 $5 $10 nDCG@10, mean of SciFact, NFCorpus, FiQA cost per 1,000 queries at list price, log scale Jev · batch Jev · pair Mixedbread v3.1 listwise Voyage rerank-3 Cohere v3.5 Mixedbread large-v2 Jina v3.5* Clef-flash · pair Clef · pair Perplexity · batch Perplexity · choice Perplexity · pair OpenAI · choice OpenAI · batch OpenAI · pair

Scroll sideways for the whole chart.

Every reranker and request shape. Pair prices include all 30 calls. OpenAI prices are estimated from reported input tokens at list price. Clef-flash · batch (0.283) and Clef-flash · choice (0.329) score below the BM25 order (0.404) and sit off the bottom of the chart. Hover a point, or pick a family in the legend.
Reranker / shapeMean nDCG@10$ / 1,000 queriesp50 / p95 per query
Mixedbread v3.1 listwise0.511$1.42217 / 262 ms (laptop)
Voyage rerank-30.504$0.50185 / 287 ms (laptop)
Jev · pair0.502$0.8730 calls
Perplexity · choice0.501$0.40521 / 737 ms (mini)
Jev · batch0.501$0.54223 / 1,419 ms (laptop)
Perplexity · batch0.500$11.591,783 / 2,770 ms (mini)
Clef-flash · pair0.498$1.3930 calls
Clef · pair0.494$3.6930 calls
Jina v3.5*0.492$0.47303 / 607 ms (laptop)
Perplexity · pair0.490$0.5730 calls
Cohere v3.50.486$2.00192 / 456 ms (laptop)
Mixedbread large-v20.476$1.42361 / 429 ms (laptop)
OpenAI · choice0.472$0.92255 / 498 ms (mini)
OpenAI · batch0.450$1.55346 / 855 ms (mini)
OpenAI · pair0.441$1.5830 calls
Clef-flash · choice0.329$0.25410 / 621 ms (mini)
Clef-flash · batch0.283$0.55631 / 956 ms (mini)
BM25 order0.404no reranker

OpenAI prices are estimates: reported input tokens at the $0.10 per million list price, with no output charge. The rest carry over from the earlier posts.

Per corpus, choice is the best shape on SciFact (0.751) and FiQA (0.355), and batch is best on NFCorpus (0.323). Against Jev batch those best shapes trail by 0.018 on SciFact, 0.035 on NFCorpus and 0.021 on FiQA, and every interval excludes zero. They also trail Perplexity’s best shape by 0.025 to 0.042 and Clef-flash pair by 0.017 to 0.044. These gaps are small, but unlike the Jev-versus-Perplexity race in the last post they all point one way.

One comparison is closer. The first post ran gpt-5.6-luna, an older model, as a plain listwise LLM reranker on SciFact. OpenAI choice ties it there (+0.004, interval spanning zero). On the other two corpora I only kept that run’s averages, which are a bit higher than Decisions, so there is no paired comparison.

Choice beats batch

OpenAI by request shape

nDCG@10 · BM25 order and Jev batch marked

0.2 0.3 0.4 0.5 0.6 0.7 0.8 nDCG@10 per corpus SciFact choice batch pair BM25 Jev NFCorpus FiQA

Scroll sideways for the whole chart.

OpenAI Decisions on each corpus, in its three shapes, against Jev batch (orange) and the BM25 order the reranker starts from (the bar). Choice asks it to pick among the 30; batch asks 30 yes-or-no questions in one call; pair scores one document per call.

With Clef-flash, batch and choice both collapsed below the BM25 order. OpenAI doesn’t collapse, but the shapes still split. On SciFact, choice scores 0.751 while batch and pair score 0.681 and 0.670, and neither of those clearly beats BM25’s 0.667. Asking the model to spread one unit of probability across the list works better than asking 30 separate yes-or-no questions about it.

Choice is also the cheapest shape, because one question over 30 options takes fewer input tokens than 30 predicates. Pair and batch cost about the same as each other. Unlike Perplexity batch, which billed the shared context once per question ($11.59 per 1,000 queries), OpenAI batch is billed close to its actual input.

Sharp probabilities, broken prune

The first post’s argument for Jev was that each document gets a calibrated probability, so one call both reranks and prunes: drop everything under 0.1 and lose almost nothing. OpenAI’s pair and batch probabilities are calibrated well on average (SciFact batch has a Brier score of 0.026), but they are sharply peaked:

  • SciFact batch keeps 5.7% of candidates at p ≥ 0.1, and loses 26% of the relevant documents in the shortlist.
  • NFCorpus batch keeps 19.5% and loses 29%.
  • FiQA batch keeps 23.6% and loses only 6%, but at 12% precision.

Pruning at p ≥ 0.1

share of relevant documents kept · share of candidates kept

● batch · ○ pair
70% 80% 90% 100% 0% 10% 20% 30% 40% 50% 60% relevant documents kept candidates kept at p ≥ 0.1 SciFact SciFact NFCorpus NFCorpus FiQA FiQA

Scroll sideways for the whole chart.

What a 0.1 threshold keeps, per corpus. Up and to the left is a better prune. Jev keeps 91% to 100% of the relevant documents; OpenAI keeps 71% to 76% on SciFact and NFCorpus. Recall counts only relevant documents already in the BM25 top 30.

Jev batch, at the same threshold, kept 94% of SciFact’s relevant documents, 91% of NFCorpus’s and nearly all of FiQA’s, while dropping 46% to 85% of the candidates. OpenAI prunes harder and loses more. So the threshold mostly fails. If you want to prune with these scores, you’d need to tune a lower cutoff per corpus, which is the work a calibrated probability was supposed to save.

Order and refusals

Same-order reissues returned identical scores, so the API is deterministic for a fixed request. Candidate order still matters. Reversing the list kept batch rankings fairly stable (mean rank correlation 0.76 to 0.79) but moved choice more (0.48 to 0.69), and changed the top document on 3 to 11 of 20 diagnostic queries per corpus. That’s less sensitive than Clef-flash and in the same range as Perplexity choice.

The API also declined some choice questions outright: 4 SciFact queries, 18 NFCorpus and 1 FiQA. I score a refusal as a uniform distribution, which leaves those queries in BM25 order. NFCorpus is medical, so that’s where a safety filter would be most likely to fire, and it is also where choice did worst.

Latency

Quality against latency

mean nDCG@10 · p50 per query, whisker to p95

○ September, laptop · ◆ October, mini
0.43 0.45 0.47 0.49 0.51 100 ms 200 ms 500 ms 1 s 2 s 5 s nDCG@10, mean of SciFact, NFCorpus, FiQA latency per query, p50 with a whisker to p95, log scale Jev · batch Mixedbread v3.1 listwise Voyage rerank-3 Cohere v3.5 Mixedbread large-v2 Jina v3.5* Perplexity · batch Perplexity · choice OpenAI · choice OpenAI · batch

Scroll sideways for the whole chart.

One-call shapes only, each with the timing from its own run. Hollow circles are September runs from a laptop and diamonds are October runs from the mini, at different concurrency and pacing, so this is context rather than a speed ranking. OpenAI's runs used concurrency 6.

The full runs on the mini, at concurrency 6, put OpenAI choice at 255 ms median and 498 ms at p95, and batch at 346 ms and 855 ms, averaged over the three corpora. Choice is about twice as fast as Perplexity choice (521 ms) and close to Jev batch (223 ms median in September, 192 ms in the October side-by-side test).

I also tried to time Jev and OpenAI batch side by side, as the last post did for Jev, Clef-flash and Perplexity. OpenAI batch took 406 ms median and 698 ms at p95 over 50 SciFact calls, but the Jev account had run out of credits and returned an error, so there is no matched Jev number. The mini was also under heavy load during that window, so none of these numbers is a controlled speed ranking.

Where this leaves things

Four general-purpose decision models have now taken a reranking test built for specialist models. Jev and Perplexity sit with the specialists. Clef-flash gets there if you can afford 30 calls per query. OpenAI’s Decisions API is a credible fourth: choice beats BM25 on two of three corpora for under a dollar per 1,000 queries. But it trails the other three everywhere, and its probabilities don’t support the prune that made Jev interesting in the first place.

The usual limits apply. The public benchmarks may be in some models’ training data, the shortlist is only 30 deep, the intervals cover query sampling with no correction for multiple comparisons, and the prompt was chosen for consistency across models, not tuned for this one. The API is a beta, so I’ll rerun it when it reaches general availability.

Every run, with per-query rankings and the receipts behind the cost figures, is in the reranker comparison.

Start typing to search.