GPT-6 vs Fable 5.1 Coding Benchmark
Research: gpt6 vs fable 5.1 vs 3.8-flash vs... ??? — coding
TL;DR
As of September 2026, Claude Fable 5.1 is the strongest coding model on BenchLM's overall ranking (79.5), narrowly ahead of GPT-5.6 Sol (79.0), with Fable 5.1 posting the highest SWE-bench Pro score (81.2 vs. Sol's 64.6) and topping agentic benchmarks like CursorBench (73.4%) [1][3]. GPT-5.6 Sol counters with the best Terminal-Bench 2.1 score (85.77%) and better DeepSWE performance (72.7 vs. Fable 5's 70.0) [7][9]. Gemini 3.8 Flash is the standout value option: it delivers Fable-class capability on several coding benchmarks at $0.75/$3.75 per million tokens — roughly 13x cheaper on input than Fable 5.1 — though its Terminal-Bench 4.0 agentic score (19.1%) trails the frontier significantly [4][6]. Claude Opus 5 remains a serious contender, leading DeepSWE v1.1 at 74.0% and outscoring Sol on Terminal-Bench 4.0 (51.8% vs. 37.3%) [6].
The frontier picture: two leaders, a budget disruptor, and the rest
The September 2026 coding landscape is a three-way contest between Anthropic's Fable and Opus lines, OpenAI's GPT-5.6 family, and Google's Gemini 3.8 Flash. BenchLM's composite coding ranking places Claude Fable 5 at the very top (81.3), followed by Claude Fable 5.1 (79.5) and GPT-5.6 Sol (79.0), with Claude Opus 5 holding rank 5 [1]. But benchmarks diverge sharply depending on what dimension of coding they stress — software engineering, terminal agent tasks, long-horizon repository work, or cost-scaled value.
Anthropic's own evaluation materials emphasize Fable 5.1's step-change over its predecessor on agentic coding. The company reports Terminal-Bench 4.0 scores of 55.8% for Fable 5.1 versus 42.0% for Fable 5 and 37.3% for GPT-5.6 Sol, while Mythos 5.1 reaches 60.9% [3]. On CursorBench 3.2.0, Fable 5.1 lands at 73.4%, ahead of Fable 5 (70.5%), Opus 5 (70.0%), and GPT-5.6 Sol (67.2%) [3][5]. The claim from a Jane Street research head — "Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5" and "remains readable over long, multi-step tasks" — is a vendor-supplied quote, not independent verification, but it is concrete about the qualitative dimension: readability over long horizons [5].
GPT-5.6 Sol's strengths lie elsewhere. It leads Terminal-Bench 2.1 at 85.77% accuracy, ahead of Fable 5.1 (85.02%) and Opus 5 (84.64%) [7][10]. On the hard split, Sol ties Opus 5 at 84.44%, with Fable 5.1 at 82.22% [7]. Sol also beats Fable 5 on DeepSWE: 72.70 versus 70.00, ranking 3rd of 34 models [9]. Notably, Sol's reported SWE-bench Pro score of 64.60 (rank 8 of 60) is dramatically below Fable 1's 81.2 — a gap large enough that either the measuring conditions differ meaningfully or the benchmark captures something Sol handles relatively poorly [1][9]. BenchLM explicitly rates Fable 1 at 81.2 on SWE-bench Pro versus Sol's 64.6, and OpenAI's own preview page is unhelpful here — its content is a string of "5" and "6" characters with no legible benchmark data [1][8].
Gemini 3.8 Flash, launched September 2, 2026, is the price-performance story of this generation. Its terminal benchmark trajectory — 90.8% on Terminal-Bench 2.1 per DataCamp versus 89.4% per emergent.sh — sits within striking distance of the frontier, while its DeepSWE v1.1 score of 73.7% trails Opus 5's 74.0% by just 0.3 points and actually beats Sol's 72.7% [4][6]. The pre-release characterization — "Fable 5 level performance at Flash pricing" — captures the positioning, though it is exactly the kind of unverifiable leak DataCamp reports without attribution rigor [4]. Google's internal Jetski testing, with engineers preferring 3.8 Flash to Anthropic's Opus for coding, is again self-reported [4].
The benchmark divergence matters more than the headline numbers
A striking pattern across this evidence: the same model pair reorders itself depending on benchmark family. Terminal-Bench 2.1 and Terminal-Bench 4.0 tell different stories. On 2.1, the order is Sol (85.77%) > Fable 5.1 (85.02%) > Opus 5 (84.64%) > Gemini 3.8 Flash (81.27%) [10]. On 4.0, the only direct comparisons available come from different sources: Anthropic claims Fable 5.1 at 55.8% and Sol at 37.3% [3], while emergent.sh reports Opus 5 at 51.8% and Gemini 3.8 Flash at a distant 19.1% [6]. These gaps are large enough to suggest Terminal-Bench 4.0's longer-horizon tasks expose a weakness in Flash that 2.1's 89-task suite does not [7][6]. My inference: the Flash family excels at bounded terminal tasks but degrades on extended agentic autonomy, a pattern consistent with Google prioritizing mid-range models for throughput rather than endurance.
SWE-bench Pro is another fault line. BenchLM puts Fable 5.1 at 81.2 and Sol at 64.6 — a 16.6-point advantage for Anthropic — while DataCamp has Gemini 3.8 Flash at 61.6% and 3.7 Flash at 60.4% [1][4]. The source mismatch is real: BenchLM and DataCamp may be evaluating with different harnesses, different subset selections, or different pass criteria. DeepSWE flips the Fable/Sol relationship partially: Sol at 72.7 beats Fable 5 at 70.0 [9], though Fable 5.1-specific DeepSWE numbers under impractically comparable conditions aren't clearly reported.
Terminal-Bench-Science 0.1 shows Fable 5.1 at 52.6%, more than double Fable 5's 24.7% and ahead of Sol's 22.4%, suggesting Anthropic concentrated real capability gains in scientific coding — a fact consistent with the Mythos sibling model's 60.9% on the science-heavy Terminal-Bench 4.0 [3]. In contrast, AutomationBench shows Fable 5.1 (31.4%) ahead of Opus 5 (26.9%) and Sol (19.6%) but below what the scale of improvement elsewhere might predict [3][5].
The price dimension: Flash redefines cost while Fable defends premium
Gemini 3.8 Flash's economics are not just cheap — they are a different order of magnitude. At $0.75 per million input tokens and $3.75 per million output [4][6], Flash undercuts Fable 5.1's $10/$50 by approximately 13x on both axes [4], and it does so while staying within a few points of Fable on DeepSWE v1.1 and Terminal-Bench 2.1. Google's CWE-Bench cybersecurity patching numbers tell the same story from the other direction: Flash Cyber scores 47.2% pass@1, versus 47.8% for an unnamed "leading frontier model" — a 0.6-point gap for what DataCamp doesn't disclose as a model name but near-certainly is Anthropic or OpenAI [4][6]. If that's the comparison, then Flash provides near-frontier security patching at a thirteenth of the price.
Anthropic's response on economics is parametrics, not price cuts. It dropped cache reads from $1 to $0.25 per million tokens — meaningful for agentic loops that re-read large contexts — while holding $10 input and $50 output [3]. But Artificial Analysis disputes Anthropic's "up to 45 percent less" claim: at max effort, Fable 5.1 produces 1.7x more output tokens per task, driving per-task cost to $3.76 on the Intelligence Index versus Opus 5's $2.34 — i.e., Fable 5.1 actually costs 20% more per task than Fable 5 [3]. This is a measured finding, and it directly contradicts Anthropic's marketing. My inference: Fable 5.1 "thinks longer" to achieve its gains, which is neither a defect nor an accident but a design point with real cost consequences. Cognition's reported move from Opus 5 to Fable 5.1 in Devin for code review cites "the new cache read pricing" as the enabling factor — again a vendor-supplied quote [5].
GPT-5.6 Luna enters as OpenAI's value play at $6 output per 1M tokens — still 60% above Flash's output price, and its Terminal-Bench 2.1 score (79.03%) is about two points below Flash's 81.27% [1][10]. For value-focused buyers, Flash is measurably ahead on both price and terminal performance; Luna's niche appears to be OpenAI-ecosystem integration rather than raw cost-per-token.
The build-up of claims and their sourcing
The data quality across these sources varies considerably. BenchLM.ai and evals.report are aggregators that publish leaderboards with varying verification standards — evals.report explicitly labels Claude Fable 5's 84.6% Terminal-Bench 2.1 as "unverified" [12]. Google's Gemini 3.8 Flash benchmark report is self-computed, with competitor scores self-reported; emergent.sh flags this openly [6]. Anthropic's Fable 5.1 blog discloses a serious confound: "On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench," with fallback models completing interventions [5]. This likely understates Fable's nominal capability on safety-gated benchmarks but reflects production behavior — a choice Anthropic made deliberately and disclosed, which at least frames the comparison honestly. OpenAI's GPT-5.6 Sol preview page [8] is effectively inert as a source: it contains no parseable benchmark data, only repeated numeric glyphs, making it useless for direct verification of OpenAI's own claims. Grok 4.6 appears at 78.28% on Terminal-Bench 2.1 — respectable but behind Flash and all three frontier leaders [10] — and a Hacker News thread on Grok 4.5 is about structured output restrictions and moderation philosophy, contributing nothing measurable about coding performance [11]. DeepSeek V4 Flash 0731 scores 67.04% on Terminal-Bench 2.1, a meaningful tier below the pack [10]; its other numbers are not in this source set.
What is contested or uncertain
Several headline numbers cannot be reconciled. SWE-bench Pro: BenchLM credits Fable 5.1 with 81.2 while DataLearnerAI credits Fable 5 with 80.30 "with Deep Thinking Mode and Tools" — different models, different settings, and the 0.9-point gap between them may reflect either the model upgrade or evaluation variance [1][9]. Meanwhile Sol's 64.6 from BenchLM is the only Sol SWE-bench Pro datapoint available; no OpenAI-confirmed number exists in this corpus [8]. Terminal-Bench 2.1 Flash: DataCamp says 90.8%, emergent.sh says 89.4%, BenchLM says 81.27% — a 9.5-point spread that no single source resolves [4][6][10]. The likely explanation is different subsets or harnesses, but "likely" is an inference, not a measurement. Gemini 3.8 Flash's Terminal-Bench 4.0 score of 19.1% versus Fable 5.1's 55.8% is the largest single-model gap on agentic coding and is reported from different sources measuring the same benchmark name — a serious comparability risk [3][6]. Artificial Analysis's Intelligence Index puts Fable 5.1 at 66, Opus 5 at 63, Sol at 61, and Flash at 59 (high effort), yet per-task cost ordering at max effort puts Fable 5.1 far above Opus 5 [3][6]. The cost claims from Anthropic and the rebuttal from Artificial Analysis cannot both be correct as framed; they measure different things (token-level savings vs. per-successful-task cost). Finally, Fable 5's overall BenchLM rank (81.3) sitting above its successor Fable 5.1 (79.5) is unexplained in the sources — an anomaly worth flagging unless it reflects a composite heavily weighted toward aspects where Fable 5 was competitive or a model-family variance artifact.
Open questions
-
What does GPT-5.6 Sol actually score on SWE-bench Pro under OpenAI's own harness with code-execution scaffolding enabled? The single third-party figure (64.6) is far below Fable 5.1's 81.2 and looks inconsistent with Sol leading Terminal-Bench 2.1 [1][7].
-
Why does Gemini 3.8 Flash score 89–90% on Terminal-Bench 2.1 but only 19.1% on Terminal-Bench 4.0, while Opus 5 moves from 84.64% to 51.8% and Fable 5.1 from 85.02% to 55.8%? Is Terminal-Bench 4.0 discriminating between bounded terminal command execution and the very different skill of sustained agentic self-direction, and does Flash's deficit reflect architecture or deliberate scope-limiting? [3][6][10]
-
Fable 5 vs. Fable 5.1: why does Fable 5 lead the BenchLM composite at 81.3 while Fable 5.1 scores 79.5, despite Fable 5.1 dominating on agentic benchmarks like CursorBench and Terminal-Bench 4.0 [1][3]? What composite weighting produces that inversion?
-
The Qwen blog source [2] promised FrontierSWE, MLS-Bench-Lite, PaperBench, AndroidBench, and proprietary Qwen benchmarks — but no actual scores are present in the source text. Where does Qwen3.8-Max actually land against Fable 5.1 and Sol on externally verifiable coding benchmarks?
-
Is SwE-bench Pro score of 61.6% for Gemini 3.8 Flash [4] compatible with the pre-release "Fable 5 level" characterization [4], given Fable 5.1's 81.2? If the gap is this wide, the "Fable-level" framing would be misleading in any context except price-per-dev-unit.
Figures
Terminal-Bench 2.1 accuracy (%) — GPT-5.6 Sol, Claude Fable 5.1, Opus 5, Gemini 3.8 Flash, GPT-5.6 Luna, Grok 4.6, DeepSeek V4 Flash
DeepSWE v1.1 score (%) — Claude Opus 5, Gemini 3.8 Flash, GPT-5.6 Sol, Claude Fable 5
Sources
- Best LLM for Coding (September 2026): SWE-bench & LiveCodeBench Ranked | BenchLM.ai
- Qwen3.8-Max: A New Bar for Coding and Cowork
- Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent less
- Gemini 3.8 Flash: Features, Benchmarks, and Pricing | DataCamp
- Anthropic Launches Claude Fable and Mythos 5.1 Models
- Gemini 3.8 Flash Benchmarks: Scores, Cost, and Meaning
- Terminal-Bench 2.1
- Previewing GPT-5.6 Sol: a next-generation model
- GPT-5.6 Sol Benchmark Results Analysis & Model Comparisons | DataLearnerAI
- Terminal-Bench 2.1 Leaderboard & Scores — September 2026 | BenchLM.ai
- Grok 4.5 | Hacker News
- Terminal-Bench 2.1 leaderboard & model scores | evals.report
Generated by tink · sources are web pages; verify anything important.