TL;DR
Grok 4.7 與 Claude Fable 5.1 處於同一前沿模型層級,但它們優化方向不同,側重於不同的部署經濟性。Grok 4.7 主打程式設計、代理式工作與性價比,提供 500K token 上下文視窗,官方定價為輸入 $2/M、輸出 $6/M tokens。Claude Fable 5.1 將上下文容量加倍至 1M tokens,面向高要求推理與長時程代理式工作,定價為輸入 $10/M、輸出 $50/M tokens。
基準測試表現是分化而非一面倒。xAI 的官方發佈評測 顯示 Fable 5.1 在 CursorBench 4.0 與 Terminal-Bench 4.0 領先,而 Grok 4.7 在 DeepSWE v1.1、EEBench 與 Harvey Legal Agent Benchmark 領先。實務選擇更多取決於每個被接受結果的成本、任務時長、上下文需求、工具使用與重試行為,而非單一「通吃」的勝者。
Key Takeaways
- Grok 4.7 在較低上下文定價門檻下,輸入 $2/M、輸出 $6/M tokens;已快取輸入為 $0.50/M。
- Claude Fable 5.1 輸入 $10/M、輸出 $50/M tokens,快取讀取為 $0.25/M。
- Claude Fable 5.1 提供 1M-token 上下文視窗;Grok 4.7 提供 500K tokens。
- 基準領先依任務而異:Fable 5.1 在多個長時程程式設計測試領先;Grok 4.7 在 xAI 的對比中於若干工程與領域代理測評領先。
- 最有用的生產指標是每個成功任務成本,而非單看每百萬 token 價格。
What Are Grok 4.7 and Claude Fable 5.1?
Grok 4.7 Overview
Grok 4.7 是 xAI 面向程式設計、代理式任務與知識工作的前沿模型,於 2026 年 9 月 21 日發佈。xAI 稱其採用比 Grok 4.6 更大的新基座模型,並進行了更長時程的強化學習,偏重可持續數小時完成的任務。此次發佈亦強調更佳的自我驗證與長上下文管理能力。
官方開發者文件指出,模型 ID 為 grok-4.7,提供 500,000-token 上下文視窗、文字與影像輸入、文字輸出,四種推理等級——low、medium、high、xhigh——以及包含函式呼叫、網頁搜尋、X 搜尋與程式碼執行在內的工具。
Grok 4.7 官方發佈視覺 - SpaceXAI
Claude Fable 5.1 Overview
Claude Fable 5.1 於 2026 年 9 月 1 日發佈。Anthropic 將其定位於高要求推理、長時間程式設計、多步研究、文件密集型專業工作,以及可在延長工作階段中持續運作的代理。
Anthropic 的模型文件指出其擁有 1M-token 上下文視窗、最高 128K 輸出 tokens、文字與影像輸入、自適應思考與努力程度控制,並具備 2026 年 6 月的可靠知識截止。
Claude Fable 5.1 官方發佈視覺 - Anthropic
Grok 4.7 vs Claude Fable 5.1 at a Glance
| Specification | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Release date | September 21, 2026 | September 1, 2026 |
| Provider | xAI / SpaceXAI | Anthropic |
| Model ID | grok-4.7 | claude-fable-5-1 |
| Primary positioning | Coding, agents, knowledge work | Demanding reasoning and long-horizon agents |
| Context window | 500K tokens | 1M tokens |
| Max output | No fixed text output limit documented | Up to 128K tokens |
| Input modalities | Text + image | Text + image |
| Output | Text | Text |
| Knowledge cutoff | May 2026 | June 2026 |
| Reasoning | Low / Medium / High / xhigh | Adaptive thinking + effort controls |
| Web/search tools | Web search + X search | Environment/tool dependent |
| Function/tool calling | Yes | Yes |
| Model weights | Closed | Closed |
| CursorBench 4.0 coding performance | 46.3% | 51.8% |
| DeepSWE v1.1 agent performance | 71.0% (high effort) | 70.0% |
| Terminal-Bench 4.0 agent performance | 38.0% | 57.9% |
| Typical use case | Cost-sensitive coding and web/X-connected agents | Long-horizon coding and failure-sensitive professional work |
最大的結構差異在於上下文容量與 token 經濟性。Grok 4.7 的官方 API 文件確認 500K 上下文與 $2/$6 基礎定價,Anthropic 的 Fable 5.1 文件確認 1M 上下文與 $10/$50 基礎定價。
Head-to-Head Benchmark Results
目前最乾淨的直接對比來自 xAI 的 Grok 4.7 發佈基準表,兩款模型均列於同一份已發佈評測。
以下數據為廠商公布而非獨立復現。在 xAI 的表格中,Grok 4.7 以 xhigh 努力等級評測,Claude Fable 5.1 以 max 努力等級評測;xAI 另註 Grok 4.7 的 DeepSWE 成績為 high 努力等級。可用以辨識工作負載模式,之後應在你自己的提示、工具、程式碼庫與驗收標準上進行驗證。
| Benchmark | Grok 4.7 | Claude Fable 5.1 | Observed gap |
|---|---|---|---|
| CursorBench 4.0 | 46.3% | 51.8% | Fable +5.5 pts |
| DeepSWE v1.1 | 71.0%* | 70.0% | Grok +1.0 pt |
| AA Briefcase v1.1 | 1,657 | 1,678 | Fable +21 |
| Terminal-Bench 4.0 | 38.0% | 57.9% | Fable +19.9 pts |
| Harvey Legal Agent Benchmark | 19.6% | 6.7% | Grok +12.9 pts |
| HealthBench Professional | 56.7% | 62.1% | Fable +5.4 pts |
| EEBench | 64.0% | 56.4% | Grok +7.6 pts |
* xAI 標註 Grok 4.7 的 DeepSWE 成績為 high 努力等級。更廣義的結論是任務專長差異,而非單一等級:Fable 在多個長時程程式設計與專業工作流程更強;Grok 在同一供應商表格中的若干工程與領域代理測試更強。
Professional Knowledge Work
在 xAI 的表格中,Fable 5.1 在 AA Briefcase v1.1 略微領先,Grok 4.7 則在 EEBench 與 Harvey Legal Agent Benchmark 領先。Fable 5.1 在 HealthBench Professional 領先。這一分化強化了簡單但重要的觀點:知識型工作並非單一能力。工程、醫療、法律、金融、研究與辦公自動化各自要求不同的工具與推理能力。
Coding Performance
程式設計是最微妙的對比面向。在 CursorBench 4.0 上,Fable 5.1 達到 51.8%,對比 Grok 4.7 的 46.3%。在 DeepSWE v1.1 上,Grok 4.7 達到 71.0%,Fable 5.1 為 70.0%。最大差距出現在 Terminal-Bench 4.0,Fable 5.1 為 57.9%,Grok 4.7 為 38.0%。
| Coding dimension | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Repository engineering | Very strong | Very strong |
| DeepSWE | Slight lead in xAI table | Very close |
| CursorBench 4.0 | 46.3% | 51.8% |
| Terminal autonomy | Strong | Major strength |
| Long-context codebase work | 500K context | 1M context |
| Repeated-call token cost | Much lower | Premium |
| Native xAI code execution | Supported | Depends on Claude environment/tools |
| Long unattended execution | Explicit training focus | Core product positioning |
部署含義可能是:對於終端密集、長時間運行的程式設計代理,Fable 5.1 值得重點關注;當軟體工程品質達標且 token 成本對單位經濟性影響顯著時,Grok 4.7 具有吸引力。
Agentic Workflows
兩款模型都針對代理式工作流程,而非孤立的提示-回覆會話。xAI 表示 Grok 4.7 在更艱難的長時程任務上進行了訓練並改進自我驗證。Anthropic 將Fable 5.1描述為可運行數小時並跨多個應用的工作模型。
| Agent requirement | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Long-running reasoning | Strong | Very strong |
| Context capacity | 500K | 1M |
| Self-verification | Explicit training focus | Explicit long-horizon focus |
| Tool use | Strong | Strong |
| Search-native workflow | Web + X search | Environment dependent |
| Long repeated context | Prompt cache + compaction | 1M context + low-cost cache reads |
| Raw token cost | Lower | Higher |
Pricing and Deployment Economics
| Official base pricing | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Input / 1M tokens | $2 | $10 |
| Cached input / cache read | $0.50 below 200K prompt | $0.25 |
| Output / 1M tokens | $6 | $50 |
| 10M input tokens | $20 | $100 |
| 10M output tokens | $60 | $500 |
| US regional endpoint premium | +10% | Platform dependent |
在標準未快取 token 費率下,Grok 4.7 的輸入價格比 Fable 5.1 低 80%,輸出價格低 88%。然而,Anthropic 的快取讀取定價可在重複、代理式工作負載中顯著降低 Fable 5.1 的實際成本。
定價注意:Grok 4.7 的 $2/M 輸入與 $6/M 輸出為基礎費率。SpaceXAI 文件對超過 200K 上下文的請求另有較高上下文定價。以下試算假設請求均在基礎定價層內,且不含工具收費、區域溢價與快取寫入成本。
示例工作負載:10M 輸入 tokens + 2M 輸出 tokens。
- Grok 4.7: $20 input + $12 output = $32。
- Claude Fable 5.1: $100 input + $100 output = $200。
- 名義差距(未考慮快取效果前):$168。
「更好的生產指標是每個成功完成任務的成本。」若能降低重試、失敗或人工修正,較昂貴的模型仍可能更經濟;若兩者均能通過相同驗收標準,較便宜者可能佔優。
Context and Reasoning Controls
Fable 5.1 提供 1M-token 上下文視窗,相較之下 Grok 4.7 為 500K tokens。對於超大型程式碼庫、文件集、法律取證、科學研究或需維持大量歷史的代理,這一差異可能重要。
Grok 4.7 提供 low、medium、high、xhigh 推理設定;Fable 5.1 採用自適應思考與努力等級控制。這些標籤不可直接對比,公平測試應保持任務與驗收標準一致,而非對齊設定名稱。
Multimodal and Tool Capabilities
兩者均支援文字與影像輸入。Grok 4.7 的 API 包含函式呼叫、網頁搜尋、X 搜尋與程式碼執行。Claude Fable 5.1 針對文件、試算表、簡報、研究與長時程代理工作優化,工具行為取決於 Claude 環境或應用堆疊。
| Capability | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Text input | Yes | Yes |
| Image input | Yes | Yes |
| Function/tool calling | Yes | Yes |
| Web search | Native xAI tool | Environment dependent |
| X search | Native | No equivalent proprietary X integration |
| Code execution | Native xAI tool | Environment dependent |
| Documents / spreadsheets / slides | General knowledge-work focus | Explicit product focus |
Decision Summary
| Dimension | Grok 4.7 | Claude Fable 5.1 |
|---|---|---|
| Coding quality | Frontier | Frontier |
| CursorBench 4.0 | 46.3% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 70.0% |
| Terminal-Bench 4.0 | 38.0% | 57.9% |
| EEBench | 64.0% | 56.4% |
| Harvey Legal Agent | 19.6% | 6.7% |
| HealthBench Professional | 56.7% | 62.1% |
| Context | 500K | 1M |
| Input price | $2/M | $10/M |
| Output price | $6/M | $50/M |
| Search | Web + X | Tool/environment dependent |
| Long-running agents | Strong | Major strength |
| Price-performance | Major advantage | Premium capability tier |
| Typical fit | Scale-sensitive frontier workloads | High-value long-horizon tasks |
Can You Use Grok 4.7 and Claude Fable 5.1 Through CometAPI?
可以。CometAPI 中的 Grok 4.7 API 與 CometAPI 中的 Claude Fable 5.1 API 均可作為多模型路由策略的一部分使用。這一點很重要,因為最佳的生產架構未必需要只選一款前沿模型。
一種實用的路由模式:
- 預設路由:對成本敏感的前沿任務使用 Grok 4.7。
- 升級路由:當評估未通過、任務超出上下文需求,或長時程代理需要更強的終端執行時,切換至 Claude Fable 5.1。
這讓團隊比較被接受結果率、總 tokens、延遲、重試次數與每個被接受結果的成本,而不是僅依賴單一公開排行榜。
Grok 4.7 vs Claude Fable 5.1: How to choose
Choose Grok 4.7 for Cost-Sensitive and Search-Native Workloads
- 高量程式設計助理,且 token 成本對單位經濟性影響顯著。
- 工程或技術型代理,且 Grok 的領域結果與工作負載貼合。
- 受益於原生網頁與 X 搜尋的研究工作。
- 批次知識處理與重複的背景自動化。
- 當兩者都能通過相同驗收門檻、成本成為決定因素的工作負載。
對使用多模型閘道的開發者,可透過單一整合層評估 CometAPI 中的 Grok 4.7 API 與其他前沿模型。
Choose Claude Fable 5.1 for Long-Horizon and Failure-Sensitive Workloads
- 多小時的自主程式設計與終端密集工作流程。
- 極大程式碼庫或文件集,受益於 1M-token 上下文視窗。
- 需在多步驟中維持狀態的複雜專業代理。
- 高商業價值場景,重試或失誤成本高於推理成本。
- 當 Anthropic 的長時程代理基準與生產需求緊密對齊的研究與自動化任務。
CometAPI 中的 Claude Fable 5.1 API 使用模型 ID claude-fable-5-1;由於閘道費率可能獨立於 Anthropic 標價變動,部署時應驗證定價與路由。
Conclusion
當 token 成本、與網頁/X 連結的工具,以及對規模敏感的代理工作流程是決策主導因素時,Grok 4.7 是更強的起點。當 1M-token 上下文與高要求的長時程程式設計或專業任務比基礎 token 價格更重要時,Claude Fable 5.1 是更強候選。廠商公布的基準結果呈現分化,因此不存在「通吃」的單一勝者。
在選擇前,請在相同的生產任務、工具、時間預算與驗收標準下測試兩者。同時比較每個被接受結果的成本、延遲、重試、工具失敗與人工修正時間,而非僅看公開分數。
FAQ
How Should Teams Evaluate Grok 4.7 vs Claude Fable 5.1 Fairly?
從具代表性的生產任務出發,而非泛用排行榜。對兩者使用相同的程式碼庫狀態、工具許可、時間預算與驗收標準。記錄被接受結果率、端到端延遲、輸入與輸出 tokens、快取行為、工具失敗、重試與人工修正時間。重複足夠次數以區分模型行為與任務波動。
When Can Grok 4.7 Cost More Than Claude Fable 5.1 per Accepted Task?
當較低的 token 費率導致更多重試、更長提示、更多失敗工具呼叫或更多人工審核時,總成本可能更高。反之,高階模型也非必然經濟:其較高完成率必須能抵消更高的推理成本。比較每個被接受任務的成本,而非單看每個 token 的成本。
When Should a Router Escalate from Grok 4.7 to Claude Fable 5.1?
使用可量化的觸發條件。對成本敏感的預設路由可處理能快速通過驗證的任務;當任務超出偏好的上下文範圍、自動評估失敗、需要長時間終端執行或錯誤代價高時,啟用升級。記錄觸發條件與結果,以便基於生產證據調整路由策略。
Which Grok 4.7 vs Claude Fable 5.1 Benchmark Caveats Matter Most?
最重要的注意點包括基準版本、努力等級、代理框架、工具存取、安全機制,以及結果是廠商公布還是獨立復現。例如 CursorBench 3.2.0 與 4.0 不應混為一談。公開分數有助於形成假設,但部署決策應依賴受控的內部評估。
