OpenAI 决策 API 需要值得信赖的置信度。
The OpenAI Decisions API needs a confidence you can trust

原始链接: https://anth.us/blog/openai-decisions-api-preview/

一项针对 GPT-6 Luna 的测试发现,在多步骤逻辑任务中,它表现出的确定感并不可靠。Luna 是 OpenAI 限量预览版 Decisions API 背后的低成本模型。在 3,600 道 ProofWriter 题目中,Luna 在简单、可以直接读取的题目上表现良好,但随着证明链深度增加,成绩大幅下降。在包含五步连续推理的题目中,它的两项得分分别为 46% 和 45%,而 Jev 的得分分别为 81% 和 89%。总体来看,Luna 的得分为 64.1% 和 65.0%,低于 Jev 的 83.8% 和 89.3%。 Luna 的置信度也严重失准。它对 2,672 个答案给出了至少 99% 的确定度,但其中只有 68% 是正确的。在最困难的题目中,它的置信度几乎无法区分答案正确与否。相比之下,Jev 给出 99% 置信度的答案中,有 98.9% 是正确的。 这些结果涉及的是通用 API 模型,而不是为 Decisions API 提供支持的未公开专业版本。此外,ProofWriter 是一套合成数据集。即便如此,任何基于置信度的路由机制,都应在带有标注且难度符合实际的问题上进行测试,并分别衡量准确率和置信度校准情况。

讨论质疑 OpenAI 的 Decisions API 是否得到了公正的评估。评论者认为,该 API 适用于快速的单步决策,而复杂的推理任务更适合通过启用推理强度的 Responses API 来处理。因此,在将推理设置为“无”的情况下运行“GPT-6 Luna”,可能无法代表模型的真实能力,并可能导致性能大幅下降。 该基准测试中的“Alan”示例也受到批评,存在歧义。“通常”和“有时”等含糊表述、不清晰的改写以及缺失的原始提示词,使人难以判断“Alan 不是蓝色的”究竟应被判定为真还是假。有人认为,面对概率性不确定性,模型应当给出较低的置信度,而不是直接给出确定答案;也有人指出,这个示例与原文中更清晰的自逻辑谜题存在很大差异。 多位评论者还质疑模型与 API 之间的比较是否合理,并指出 Decisions API 当时尚未公开宣布。总体而言,讨论者认为这篇文章进行的是一场“苹果与橘子”的比较,而且其核心示例的解释并不充分。
相关文章

原文

Say you send an agent's next step to a decision model, and you only let it act on its own when it's at least 99% sure. Everything else goes to a person. That rule is only as good as the 99%. We gave GPT-6 Luna, the model behind OpenAI's new Decisions API, 3,600 reasoning problems and asked for exactly that kind of answer. When Luna said it was 99% sure or more, it was right 68% of the time.

OpenAI announced the Decisions API at DevDay on September 29, in limited preview. It picks one answer from a list you define, so software can branch on it: classify a message, route a request, choose what an agent does next. The Decoder and Axios report it runs on "a version of GPT-6 Luna," OpenAI's low-cost model, and OpenAI's own chart puts it at 150 milliseconds a decision. It's OpenAI's answer to Jev.

We'd already run Luna that way. In Hard-Decisions, our benchmark of decision models on multi-step logic, Luna answered the same problems as Jev, Kev and Laya, one request each, from a fixed list of options. That makes it a preview of a Luna-based decision model. It isn't the Decisions API itself: OpenAI uses a specialized version of Luna, and there's no documentation yet to test against.

Why we care

We've been routing on a model's confidence since 2023. For Call Criteria, LLMs score calls against hundreds of client scorecards, and people review the calls the system is unsure about. Everything depends on knowing which calls those are. When we read that confidence from the token log-probabilities of OpenAI's models, the numbers were so concentrated that they looked certain about nearly every answer. OpenAI's models have had confidence problems for as long as we've used them, and not the shy kind. Newer reasoning models often don't expose log-probabilities at all. So we built the confidence ourselves: extract the probability behind each label, check it against labeled answers, calibrate it, and only then route on it. Classification with Confidence shows that pipeline on GPT-4o-mini.

That's the pain the Decisions API could end. What we're hoping for most is that OpenAI now treats calibration as a product: a confidence you can set a threshold on, documented and tested, instead of something withheld from the API, which we suspect, without proof, is partly about making the models harder to distill. A decision endpoint is where that would have to happen. So we wanted to know what the model underneath does today.

Accuracy by the number of inference steps the answer needs, both ProofWriter tasks pooled.

The test

The problems come from ProofWriter, a dataset from the Allen Institute for AI. Each one is a short list of facts and if-then rules in plain English, and a statement to judge. Some answers can be read straight off the page. Others need five rules chained together. The dataset labels every problem with that number, its proof depth, which makes difficulty something you can measure instead of argue about.

There are two versions of the task. In the open-world one the answer is true, false or unknown: unknown when the rules settle neither the statement nor its opposite. In the closed-world one, anything you can't prove is false. We sampled 1,800 problems for each, about 300 at every depth, and recomputed every gold answer with an independent solver.

Luna got the same treatment as everyone else. One request per problem. The problem, the question and the meaning of each option, word for word what Jev receives. A strict JSON schema that only allows the three options. Reasoning off, which is the setting a fast decision API implies. We wrote our predictions down before any model answered.

Accuracy: fine on the page, coin flips five steps in

On problems you can answer by reading, Luna is nearly as good as Jev: 93% against 98% on the open-world task, and 98% against 99% on the closed one. Then it loses ground with every step. By five chained inferences Luna gets 46% and 45% right. On the closed-world task, that's a coin flip. Jev gets 81% and 89% at the same depth, on the same problems.

Across all depths, Jev scored 83.8% and 89.3%; Luna, 64.1% and 65.0%. The gap at depth 5 is 35 and 45 points, with 95% intervals from 27 to 43 and from 38 to 51. That isn't a close call.

You might be thinking that reasoning off is what sank Luna. Turning it on might help, but it would also make every decision slower and costlier than the one-shot job a decision API exists to do, and every model here got the same single shot. The open models we tested are stuck at depth 5 too: every model except Jev was at coin-flip accuracy by depth 5.

Wrong when the answer is on the page

Most of Luna's misses come from long chains, but not all of them. At depth 0, where the statement or its opposite is written in the text, Luna got 20 of 300 open-world problems wrong and 6 of 300 closed-world ones. At depth 1, one rule away, it missed 67 of 302 and 70 of 300. Here are the shortest of those misses at each depth, in each version of the task, with the confidence Luna gave when we asked again for log-probabilities and it gave the same answer:

  • Depth 0, open world. The text says "The lion is not big." Is "The lion is big" true, false or unknown? It's false; the text says so. Luna said unknown, 100% sure.
  • Depth 0, closed world. The text says "Charlie is nice." Is "Charlie is not nice" true or false? False. Luna said true, 100% sure.
  • Depth 1, open world. "The tiger is big. All big people are cold." Is "The tiger is cold" true? Yes, one rule away. Luna said unknown, 100% sure.
  • Depth 1, closed world. "The rabbit is big. If someone is big then they chase the rabbit." Is "The rabbit does not chase the rabbit" true or false? False: the rabbit is big, so it chases the rabbit. Luna said true, 100% sure.

They're the clearest of Luna's misses; at depth 0 it's right 93% and 98% of the time. But a decision model that can be 100% sure Charlie is not nice, when the text says Charlie is nice, needs more than a threshold on its confidence.

Can Luna tell you how sure it is?

A decision model earns its keep when you can trust its confidence. Jev returns a probability for every option. Luna, through the regular API, returns text. But with reasoning off, OpenAI's API will also return log-probabilities: for each token Luna writes, the log of the probability it gave that token, plus the most likely alternatives. With reasoning on, it refuses: "'logprobs' is not supported with this model," which matches OpenAI's guide.

The strict schema makes the reading clean. It only lets Luna reply {"answer":"<option>"}, so the token that spells the option carries the whole decision, with no "No" versus "no" to reconcile. Turn its log-probability into a probability and you have Luna's confidence in the answer it gave.

Here's one real exchange. The request, minus the long prompt:

{
  "model": "gpt-6-luna",
  "reasoning_effort": "none",
  "max_completion_tokens": 1024,
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "decision",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": { "answer": { "type": "string", "enum": ["true", "false", "unknown"] } },
        "required": ["answer"],
        "additionalProperties": false
      }
    }
  },
  "logprobs": true,
  "top_logprobs": 5
}

The problem, in its paraphrased form:

Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. … Young round people who are green are usually blue. … Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. …

Statement: Alan is not blue.

And Luna's answer, with its log-probabilities:

"content": "{\"answer\":\"unknown\"}",
"logprobs": { "content": [
  { "token": "{\"",     "logprob": 0.0,      "top_logprobs": [ { "token": "{\"",     "logprob": 0.0 } ] },
  { "token": "answer",  "logprob": 0.0,      "top_logprobs": [ { "token": "answer",  "logprob": 0.0 } ] },
  { "token": "\":\"",   "logprob": 0.0,      "top_logprobs": [ { "token": "\":\"",   "logprob": 0.0 } ] },
  { "token": "unknown", "logprob": -0.00182, "top_logprobs": [ { "token": "unknown", "logprob": -0.0018 } ] },
  { "token": "\"}",     "logprob": 0.0,      "top_logprobs": [ { "token": "\"}",     "logprob": 0.0 } ] }
]}

A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.

When Luna says 99%

One example proves nothing, so we asked for log-probabilities on all 3,600 problems with exactly the benchmark's request. It cost about 14 cents, and every request and response is kept verbatim. We preregistered four predictions first. All four held.

Each point is a band of stated probability; the dashed diagonal is where a perfectly calibrated model would sit. Both tasks pooled.

Luna stated 99% or more on 2,672 of the 3,600 answers, and 68% of those were right. Below that it's flat: whether Luna says 70%, 90% or 97%, it's right about half the time. Jev stated 99% or more on 1,667 answers and got 98.9% of them right. Its line runs close to the diagonal the whole way.

The other numbers say the same thing:

  • Wrong answers stated at 95% or more: 78% of Luna's, against 7% and 17% of Jev's on the two tasks.
  • Expected calibration error, the average gap between stated and actual: 0.32 for Luna on both tasks; 0.04 and 0.03 for Jev.
  • Separating right from wrong (AUROC, where 0.5 means the probability tells you nothing): 0.66 and 0.68 for Luna; 0.85 and 0.86 for Jev.

Past a few steps, the confidence carries no information

Overconfidence on its own is fixable. If a model's 99% always means 68%, you can relabel it, the same way Jev's confidence calibrates. What calibration can't fix is a probability that doesn't move with being right.

How well each model's stated probability separates its right answers from its wrong ones, by proof depth. 0.5 is no information at all.

On problems you can read off the page, Luna's probability is informative: right answers get higher numbers than wrong ones. Each inference step erodes that, and by five steps the AUROC is 0.51: a high stated probability is no more likely to be right than a low one. Remapping the numbers can't add information they don't carry, so on hard problems a threshold on Luna's confidence routes at random. Jev's still sorts right from wrong at 0.84.

Hidden, or just overconfident?

We'd assumed OpenAI hides Luna's log-probabilities. With reasoning off, it mostly doesn't. The alternatives it returns cover 99.8% of the probability on average; the ones it leaves out carry almost nothing. The problem in these results is overconfidence, not secrecy.

That fits what OpenAI itself reported for GPT-4: the pre-trained model's probabilities were well calibrated, and post-training made them markedly worse (GPT-4 Technical Report, figure 8). It also fits our experience with OpenAI models generally. Where they expose log-probabilities, they've been too sure of themselves to calibrate into a confidence you'd route on.

Why does the API refuse log-probabilities once reasoning is on? We don't know. Researchers have recovered part of a production OpenAI model from its API's log-probabilities and logit bias, and OpenAI changed its API in response (Carlini et al., 2024; see also Finlayson et al., 2024). OpenAI has also kept o1's raw reasoning private, partly for competitive advantage (OpenAI). Protecting the model from distillation is a plausible reason. We have no evidence that it's the reason.

What this means for the Decisions API

We'll admit these results worry us. We want the Decisions API to be good: we've waited years for a hosted model whose confidence we could route on. But the model it's built on lost to Jev by 35 to 45 points at five inference steps, and its 99% meant 68%. A specialized version has a lot of ground to make up, on accuracy as much as on confidence.

Some coverage says the Decisions API returns a confidence score. OpenAI hasn't published documentation, so we can't say. If it does, these results tell you what to check before you route on it:

  • Test it at the difficulty you'll run it at. Easy problems flatter everyone. Luna and Jev are within a few points at depth 0.
  • Check calibration against labeled answers before you set a threshold. A 99% that means 68% will push the wrong decisions past your reviewers.
  • Check that the confidence separates right from wrong on your hardest cases. If it doesn't, no calibration will rescue it.
  • Ask twice. Asked every question a second time, Luna gave a different answer on 10% of them; Jev changed on 2 to 3%, and Kev-0.8B, Kev-4B and Laya on none. Part of that is sampling: 137 of Luna's replies weren't even its own most probable answer, because OpenAI samples at its default temperature.

A specialized Luna may do better than the general one. When the Decisions API opens up, we'll put it through the same 3,600 problems and the same probe.

How we measured

Every model answered the same 3,600 problems once, with the same wording and no examples, and the predictions were written down before any answers came in. We preregistered the Luna probe separately, and it sits outside the scored benchmark. Luna ran with reasoning off at OpenAI's default sampling. ProofWriter is synthetic, templated logic, so these numbers describe multi-step reasoning of that kind; other tasks deserve their own test. The full results, the preregistration and every record are on hard-decisions.anth.us.

If you're choosing between the Decisions API, Jev and an open model, run the same kind of test on your own decisions before you trust any of them: labeled cases from your own work, hard ones included, every model asked the same way, accuracy by difficulty, and a check that each model's confidence separates its right answers from its wrong ones. ProofWriter shows how these models handle chained logic; your own cases show how they'll handle your decisions. That's work we do with teams.

联系我们 contact @ memedata.com