<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>OpenAI on KbWen Blog</title>
    <link>https://www.kbwen.com/tags/openai/</link>
    <description>KbWen is a practical technology blog about AI systems, machine learning, Python, data engineering, and software development.</description>
    <generator>Hugo</generator>
    <language>zh-tw</language>
    <image>
      <url>https://www.kbwen.com/images/og-default.png</url>
      <title>KbWen Blog</title>
      <link>https://www.kbwen.com/</link>
    </image>
    
    <lastBuildDate>Wed, 23 Sep 2026 09:50:00 +0800</lastBuildDate><atom:link href="https://www.kbwen.com/tags/openai/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>GPT-6 Astra 評測與價格：跟 GPT-5.6 Sol、Fable 5.1 的比較</title>
      <link>https://www.kbwen.com/gpt-6-astra-benchmarks-and-cost-per-task/</link>
      <pubDate>Wed, 23 Sep 2026 09:50:00 +0800</pubDate><dc:creator>KbWen</dc:creator>
      <guid>https://www.kbwen.com/gpt-6-astra-benchmarks-and-cost-per-task/</guid>
      <description>我們來看看新模型 GPT-6 Astra 跟上一代 Sol、Anthropic 的 Fable 5.1 在分數和花費上差在哪裡。</description>
      <content:encoded><![CDATA[<p>OpenAI 在美國時間 9 月 3 日發表了 GPT-6 Astra，換算成台灣時間已經是 4 日。Astra 的上一代 <a href="/gpt-5-6-sol-terra-luna-codex-zh/">GPT-5.6 Sol</a>，對手那邊則有 Anthropic 的 Fable 5.1。最近模型更新速度很快，所以新模型一出來，大家想知道的通常就是兩件事：比上一代進步多少，跟對手比起來又有甚麼差別。</p>
<p>OpenAI 在提供的對照表中，把 Astra 和 Sol、Fable 5.1 等幾個現在排名較高的模型排在一起比較。可以看到跟 Sol 比的部分沒有懸念，領先的都是 Astra。像是 Terminal-Bench 4.0 這一項（在終端機裡把交代的任務做完），Sol 完成不到四成，Astra 接近六成，這個差距是可見的（但聽說有些 AI 會偷看，因此參考就好）。</p>
<p>不過同一列換成 Fable 5.1，差距就縮到兩個百分點左右。整體看下來，Astra 贏 Fable 5.1 的項目比較多，但也有反過來的時候，像是 Humanity&rsquo;s Last Exam 這份收集各領域難題的考卷，而在可以用工具的版本裡，Fable 5.1 反而高了將近八個百分點。</p>
<p>所以光看 OpenAI 的表，Astra 對 Sol 是全面領先，跟 Fable 5.1 的差距則小得多。而獨立評測機構 Artificial Analysis 把多項測驗合起來，算成單一的綜合指數，目前的版本裡，兩邊都開到最高設定（max）時，Astra 和 Fable 5.1 都是 53 分。這種綜合指數看不出贏的項目跟輸的項目是哪些，只有總分，所以如果想知道哪個會比較適用，還是要看做的是哪一類工作。</p>
<p>那既然能力差不多，價錢就成了比較實際的問題。目前 Astra 每個 token 的價錢是 GPT-5.6 Sol 促銷價的 2.5 倍，跟 Fable 5.1 的輸入、輸出單價則是一樣的（快取讀取反而是 Fable 5.1 比較便宜），所以只看單價，從 Sol 換過來是變貴，從 Fable 5.1 換過來也沒有比較便宜。不過 Artificial Analysis 在剛剛提到的評測裡，也算了每個任務平均要花多少錢，Astra 是 3.26 美元，Fable 5.1 是 7.63 美元，Astra 大約只要四成。但是價錢這部分就只能期待各個公司的持續軍備了。</p>
<p>單價差不多，總花費卻差了許多，我們可以看得出來是完成任務到底要花多少 token。有點像搭計程車，每個 token 的價錢是跳表的費率，費率差不多的兩台車，跑的路長短不同，有些會走高速公路，有些繞小路，有時候又會遇到塞車，付的錢也就不同。</p>
<p>最後說明這篇是照 OpenAI 的公告與 API 文件、Anthropic 的 Fable 5.1 公告，以及 Artificial Analysis 目前公布的數字寫成的，讓大家參考，以後新模型可能也會有新價錢或新算法。</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>What GPT-6 Astra Costs Per Token and Per Task</title>
      <link>https://www.kbwen.com/gpt-6-astra-cost-per-token-and-per-task/</link>
      <pubDate>Wed, 23 Sep 2026 09:30:00 +0800</pubDate><dc:creator>KbWen</dc:creator>
      <guid>https://www.kbwen.com/gpt-6-astra-cost-per-token-and-per-task/</guid>
      <description>GPT-6 Astra lists at $10/$50 per million tokens against GPT-5.6 Sol&amp;#39;s $4/$20. A walk through Artificial Analysis&amp;#39;s per-task cost figures for both, and the one benchmark where Astra&amp;#39;s lower turn count sits beside a lower score.</description>
      <content:encoded><![CDATA[<p>OpenAI released GPT-6 Astra on September 3, 2026. In the API it costs $10 per million input tokens and $50 per million output tokens for prompts up to 272K input tokens, above which the whole request bills at $20 and $75. Its predecessor <a href="/gpt-5-6-sol-terra-luna-codex/">GPT-5.6 Sol</a> has been $4 and $20 since a promotional cut on August 21, held at least through November 21. On September 22 OpenAI added two cheaper models to the GPT-6 line, <code>gpt-6-sol</code> at $2 and $10 and <code>gpt-6-luna</code> at $0.10 and $0.50.</p>
<h2 id="cost-per-task">Cost per task</h2>
<p>On Artificial Analysis&rsquo;s Intelligence Index v4.3, in a run published on September 9, Astra at max effort scores 53, level with Claude Fable 5.1 at max with fallback and six points above GPT-5.6 Sol. Fallback is Anthropic&rsquo;s setting that routes safety-flagged requests to another Claude model, and it served about 4% of output tokens across the index. Astra&rsquo;s tasks come to $3.26 each and Fable 5.1&rsquo;s to $7.63, which Artificial Analysis describes as matching the score &ldquo;at ~40% of the cost per task&rdquo;. Against GPT-5.6 Sol the comparison runs the other way: &ldquo;At max effort, Astra is ~60% more expensive than GPT-5.6 Sol&rdquo;, which Artificial Analysis has since priced at $1.99 per Intelligence Index task.</p>
<p>Most of that difference comes from output tokens. At max effort Astra uses 27k output tokens per task where Fable 5.1 at max with fallback uses 78k — about a third, for the same score. The two models list at the same rates, $10 and $50 per million, which leaves the token counts to explain the gap, and Astra gets there even though its cached input costs $1.00 per million against Fable 5.1&rsquo;s $0.25. Cost per task is not purely an output-token bill, though. Artificial Analysis calculates each evaluation&rsquo;s cost &ldquo;from input, cache hit, cache write, reasoning, and answer token prices&rdquo;, so the input and cache sides of the run are in the figure too.</p>
<p>OpenAI makes a claim of its own in its GPT-6 guide, that Astra delivers &ldquo;a lower estimated API cost per task than earlier models despite its higher per-token pricing&rdquo;. In the September 9 run Artificial Analysis put every Astra effort level on its Intelligence Index cost frontier, and Astra still defines the frontier for output tokens per task. Sitting on that frontier meant nothing else in the comparison set reached a given score for less.</p>
<p>In the Coding Agent Index, where Astra runs inside Codex, it scores 62, level with Fable 5.1 in Claude Code and seven points ahead of GPT-5.6 Sol. Per task at max effort it costs $7.09, which Artificial Analysis puts at about 15% more than GPT-5.6 Sol for those seven points, and about 40% less than Fable 5.1 for the same score.</p>
<h2 id="fewer-turns">Fewer turns</h2>
<p>Not every result in the same report moves in Astra&rsquo;s favor. On GDPval-AA v2, an adaptation of OpenAI&rsquo;s dataset covering economically valuable tasks across 44 occupations, Astra comes in about 45 Elo points below GPT-5.6 Sol. Alongside that, Artificial Analysis records how many turns each model took: 24 per task for Astra at max effort, 45 for GPT-5.6 Sol, and 60 each for Fable 5.1 and Claude Opus 5. A turn is roughly one exchange between the model and the harness running it.</p>
<p>Astra took the fewest turns of the four models here, in the same direction as the output-token figures on the Intelligence Index. But on this benchmark the lower count sits next to a lower score. Artificial Analysis reports the turn counts and the Elo drop as two observations from the same runs, and whether one produced the other is not something the published figures settle. The drop is an Elo figure, so it places the models relative to each other rather than against a fixed score, and the evaluation has since moved to v2.1.</p>
<h2 id="effort-levels-and-two-weeks-later">Effort levels, and two weeks later</h2>
<p>Astra&rsquo;s costs above are max-effort numbers. Artificial Analysis prices its low effort at $0.82 per task on the same index. Two weeks after that run the picture had moved again: on September 22 Anthropic cut Claude Opus 5.5 to $4 and $20 per million, where it scores 58 on the Intelligence Index, and Artificial Analysis measured OpenAI&rsquo;s new GPT-6 Sol at $1.06 per Intelligence Index task against GPT-5.6 Sol&rsquo;s $1.99.</p>
<p>This note is put together from OpenAI&rsquo;s pricing page and GPT-6 guide and from Artificial Analysis&rsquo;s published runs, and every figure in it carries a date because at this rate none of them will hold for long. Between the run it leans on and the day this was written, both OpenAI and Anthropic shipped something cheaper than what that run measured. That is not long enough for a price comparison to go stale, or it should not be, but at the moment it seems to be.</p>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
