Grok 4.7 Explained: Benchmarks, Pricing, and How It Stacks Up Against Claude Fable 5.1 and GPT-6 Astra
Written by
Aerin Kim

SpaceXAI's Grok 4.7 keeps Grok 4.6's pricing on a larger base model, but new benchmarks show a real gap to Claude Fable 5.1 and GPT-6 Astra. Here is what the numbers actually show.
SpaceXAI shipped Grok 4.7 on September 21, 2026, calling it the company's most capable model yet for coding and knowledge work [1]. The launch post leads with a bigger base model and a longer reinforcement learning run, but leaves out two things a technical reader actually wants: a parameter count and a straight answer on how the model performs against independently run benchmarks rather than xAI's own numbers [2][6]. Once you pull in Artificial Analysis's own measurements, DataCamp's and MarkTechPost's breakdowns of the launch materials, and The Decoder's read on the competitive picture, a more complicated and more useful story emerges: a genuinely stronger coding model at an unchanged price, that still trails the two current frontier leaders by a wide margin on the composite score that most buyers will see first [4].

Why This Specific Release Is Worth Reading Closely
Frontier model launches happen often enough now that most do not deserve a full breakdown. Grok 4.7 earns one for a specific reason: it is a case study in the gap between a lab's own launch benchmarks and what an independent evaluator measures on the same tests, and that gap is large enough here to change the buying decision. SpaceXAI's own harness reports Terminal-Bench 4.0 at 38.0 percent. Artificial Analysis, running the identical benchmark independently, measures 26 percent [2][9]. That is not rounding error, and it is exactly the kind of discrepancy worth understanding in detail before routing production traffic to any model based on a launch-day chart. The rest of this post works through where Grok 4.7 actually gains ground, where the self-reported and independent numbers diverge, and what that means if you are choosing between this model, Claude Fable 5.1, and GPT-6 Astra for a real workload.
Step 1: The Architecture Question, What "Larger Base Model" Actually Means
xAI's own release notes describe Grok 4.7 as built on a new, larger base model than Grok 4.6, trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete, and trained specifically to natively understand the Grok Bot conversational harness [1][6]. That is a meaningfully different framing from Grok 4.6's own release, which reused the same V9 pretraining foundation as Grok 4.5 and spent its entire compute budget on a targeted agentic RL stage. Grok 4.7 instead grew the base model itself, then layered a similar RL stage on top, which is a more expensive way to gain ground and the more plausible explanation for why the model's own launch materials emphasize multi-hour task performance rather than a single headline architecture claim.
The Parameter Count Nobody Will Confirm
Here is the specific gap in the official materials: xAI's Grok 4.7 launch post does not publish a parameter count [3]. Several outlets have filled that gap with a figure of 2.1 trillion parameters, roughly 40 percent larger than Grok 4.6's reported 1.5 trillion, but that number traces back to an earlier claim Elon Musk made on July 28, 2026, months before this release, and it remains unconfirmed by xAI in Grok 4.7's own launch notes [6][10]. That distinction matters for anyone citing this figure themselves. A specific number with a specific provenance, an unconfirmed comment from the company's own founder months before the model shipped, is a different kind of fact than a company publishing its own architecture spec, and treating the two the same is exactly the kind of small error that erodes trust in a technical writeup once a reader checks the source.
Context Window, Modality, and Knowledge Cutoff
The practical specs are easier to pin down, though notably these too come from independent measurement rather than xAI's own launch copy. Artificial Analysis reports a 500,000 token context window, unchanged from Grok 4.6, and a knowledge cutoff of May 2026 [6][2]. The model remains text-based on output, with no multimodal generation improvements called out in the launch materials, and no image, video, or audio output claims worth repeating here [6]. If you were hoping the 500,000 token window would grow alongside the larger base model, it did not. Whatever compute went into this release, it went into depth of reasoning and training-task difficulty rather than context length, which lines up with the release notes' own emphasis on tasks that take many hours to work through rather than tasks that need to read an unusually large amount of source material at once.
Step 2: The Headline Number, Artificial Analysis Intelligence Index v4.3.2
A single composite score is the first thing most buyers will see, and it is worth understanding exactly what it does and does not tell you before anything else. The Artificial Analysis Intelligence Index version 4.3.2 combines ten separate benchmarks into one number, and on that index Grok 4.7 scores 46, a two-point gain over Grok 4.6's 44 on the same version of the index, but landing squarely mid-pack against the field [4][7]. Claude Fable 5.1 and GPT-6 Astra both score 53 on the same index, an eight-point gap over Grok 4.7 that Artificial Analysis's own coverage frames plainly as "a wide gap to Claude and GPT-6" despite Grok 4.7's bargain pricing [4].

That said, the mid-pack framing needs one honest qualifier. Artificial Analysis's own model page notes that Grok 4.7's score of 46 is well above the median score of 24 among reasoning models in a similar price tier [7]. In other words, Grok 4.7 is not competing well against the two most expensive frontier models on the market, and it is not really trying to. It is competing against a field of similarly priced models, and against that field, a score nearly double the median is a genuinely strong result. Both framings are true at once, and which one matters to you depends entirely on whether your actual alternative is Claude Fable 5.1 and GPT-6 Astra at roughly five to eight times the price, or a cheaper open-weight or budget-tier model at a comparable price point.
| Benchmark | Grok 4.7 | Claude Fable 5.1 | GPT-6 Astra | GPT-5.6 Sol Max |
|---|---|---|---|---|
| AA Intelligence Index v4.3.2 | 46 | 53 | 53 | — |
| CursorBench 4.0 | 46.3% | 51.8% | — | 41.7% |
| DeepSWE v1.1 (high effort) | 71.0% | 70.0% | — | 72.7% |
| EEBench | 64.0% | 56.4% | — | 39.4% |
| Harvey Legal Agent Benchmark | 19.6% | 6.7% | — | 2.5% |
| AA-Briefcase v1.1 | 1,657 | 1,678 | — | 1,487 |
Read the table by column rather than by row and the pattern gets clearer. Against GPT-5.6 Sol Max specifically, an older and cheaper OpenAI tier rather than the current GPT-6 Astra flagship, Grok 4.7 wins outright on CursorBench 4.0, EEBench, Harvey Legal Agent Benchmark, and AA-Briefcase v1.1, and loses only on DeepSWE v1.1 by a narrow 1.7 points [3][5]. Against Claude Fable 5.1, the picture flips: Fable 5.1 leads on CursorBench 4.0, DeepSWE v1.1, and AA-Briefcase v1.1, while Grok 4.7 leads on EEBench and Harvey Legal Agent Benchmark by wide margins [2]. None of this shows up in the single Intelligence Index number, which is exactly why a composite score is a starting point for evaluation, not a substitute for it.
Step 3: The Coding-Agent Focus, Benchmark by Benchmark
Coding is the domain SpaceXAI explicitly built this release around, and it is the domain with the most benchmark coverage across independent outlets, so it deserves a closer look than a single table row.
CursorBench 4.0: Real IDE-Style Coding Tasks
CursorBench 4.0 scores a model on realistic pull-request-sized coding tasks that mirror how a developer actually works inside an editor like Cursor: read existing code, make a targeted change, pass the test suite. Grok 4.7 scores 46.3 percent here, up from Grok 4.6's 40.4 percent, and ahead of GPT-5.6 Sol Max's 41.7 percent, but behind Claude Fable 5.1's 51.8 percent [2][3]. A nearly six-point generational gain on a benchmark this specific is a real signal that the extended RL run SpaceXAI describes actually targeted this exact task shape, and it is consistent with the same pattern Grok 4.6 showed on its own coding-focused benchmarks a full point release earlier, covered in our Grok 4.6 breakdown.

DeepSWE v1.1: The Clearest Generational Win
DeepSWE v1.1, run at high reasoning effort, is where the generational jump from Grok 4.6 to Grok 4.7 shows up most cleanly: 71.0 percent versus Grok 4.6's 65.2 percent, a gain of nearly six points on a benchmark that specifically measures end-to-end software engineering task completion rather than a single edit [2]. That score edges out Claude Fable 5.1's 70.0 percent, though it still trails GPT-5.6 Sol Max's 72.7 percent by a narrow margin [5]. Beating Fable 5.1 on any coding benchmark is worth noting given how consistently Fable 5.1 leads on agentic coding elsewhere, and it is the strongest single data point in this entire comparison for anyone specifically evaluating end-to-end software engineering agents.
Terminal-Bench 4.0: Two Numbers, One Benchmark
This is the benchmark where the self-reported and independently measured numbers genuinely diverge, and it is worth walking through both rather than picking whichever number fits a narrative. SpaceXAI's own internal harness reports Terminal-Bench 4.0 at 38.0 percent, nearly double Grok 4.6's 20.3 percent on the same internal measurement, and ahead of GPT-5.6 Sol Max's 37.3 percent [2][3]. Artificial Analysis, running the same benchmark independently rather than relying on a vendor-reported number, measures Grok 4.7 at 26 percent, well behind Claude Fable 5.1's 55 percent and GPT-6 Astra's 60 percent [9][4].
| Terminal-Bench 4.0 source | Grok 4.7 | Grok 4.6 | Claude Fable 5.1 | GPT-6 Astra / Sol Max |
|---|---|---|---|---|
| SpaceXAI's own harness (self-reported) | 38.0% | 20.3% | — | Sol Max: 37.3% |
| Artificial Analysis (independent) | 26% | — | 55% | Astra: 60% |

Both numbers can be accurate at once without contradicting each other, and the likely explanation is methodology rather than dishonesty: a self-run harness can differ from an independent evaluator's harness in task selection, timeout limits, tool-call scaffolding, and scoring rubric, all of which move a benchmark's absolute number without anyone involved doing anything improper. The practical lesson is not to distrust either number specifically, it is to never treat a self-reported benchmark score on a hard, agentic, tool-using task as directly comparable to an independently measured one, and to specifically seek out the independent number before making a purchasing decision around autonomous shell execution. Terminal-Bench, more than any other benchmark in this comparison, measures exactly the kind of unsupervised, multi-step agent behavior where scaffolding differences compound the most.
The Vals Index Regression Nobody at xAI Is Talking About
One more data point worth including precisely because it complicates the story rather than supporting it cleanly: on the independent Vals Index ranking, Grok 4.7 scored 54.15 percent and placed 24th out of 59 models evaluated, a worse showing than Grok 4.6's own 59.17 percent and 14th place finish on the same ranking [9]. That same independent report also places Grok 4.7's Terminal-Bench 4.0 performance below DeepSeek V4.1 Flash, an open-weight model, on the same benchmark [9]. Losing ground against your own predecessor on one respected independent leaderboard, and losing outright to an open-weight competitor covered in our DeepSeek V4.1 Flash breakdown, is not the story SpaceXAI's own launch post tells, and it is exactly the kind of result that only surfaces once you go looking for independent corroboration rather than reading the launch materials on their own.
Step 4: Where Grok 4.7 Actually Wins
The picture is not uniformly discouraging. Grok 4.7 posts some of its strongest results, relative to every named competitor in this comparison, on benchmarks that get far less attention than the Intelligence Index or Terminal-Bench.
EEBench: The Surprise Domain
EEBench measures performance on electrical engineering problems, and Grok 4.7 scores 64.0 percent here, up sharply from Grok 4.6's 53.0 percent, and ahead of both Claude Fable 5.1's 56.4 percent and GPT-5.6 Sol Max's 39.4 percent [6][3]. An 11-point generational jump on a specialized engineering benchmark, and a lead of nearly eight points over the current agentic-coding leader, is a genuinely notable result for anyone working in a hardware-adjacent or electrical engineering domain specifically, even if it will not move most software teams' decisions on its own.

Harvey Legal Agent Benchmark
The Harvey Legal Agent Benchmark tests a model's ability to handle realistic legal-reasoning and document-review agent tasks, and Grok 4.7 jumped to 19.6 percent here, up from Grok 4.6's 15.8 percent, and well ahead of GPT-5.6 Sol Max's 2.5 percent [3][6]. Claude Fable 5.1 scores 6.7 percent on the same benchmark, meaning Grok 4.7 nearly triples Fable 5.1's result on this specific legal-reasoning task despite trailing Fable 5.1 on general coding [2]. All three of these absolute numbers remain low in a way that should temper any enthusiasm, no model here is close to reliable on autonomous legal-agent work yet, but the relative gap between Grok 4.7 and both named competitors is real and large enough to matter for anyone specifically evaluating models for legal-adjacent tooling.
AA-Briefcase and GDPval: Knowledge Work
On AA-Briefcase v1.1, Artificial Analysis's private benchmark for long-horizon agentic knowledge work, Grok 4.7 scores 1,657 Elo, a gain of 111 Elo points over Grok 4.6 at the same high-effort setting, placing it just behind Claude Fable 5.1's 1,678 and ahead of GPT-5.6 Sol Max's 1,487 [1][2]. On GDPval, a benchmark spanning real professional tasks across law, nursing, and finance, Grok 4.7 scores 1,695, ahead of GPT-6 Astra's 1,542 but behind Claude Fable 5.1's 1,735 [5][9]. Put together with the EEBench and Harvey results above, a real shape emerges: Grok 4.7 is at its strongest on long-horizon knowledge work and specialized professional domains, and closer to mid-pack on the general coding and shell-automation benchmarks that got most of the launch-day attention.
Step 5: Pricing That Didn't Move, and the Real Cost That Did
Pricing is the one number every source agrees on completely: $2 per million input tokens and $6 per million output tokens, identical to Grok 4.6, with a faster variant available at double the price for roughly double the throughput [1][2][3].
| Model | Input ($ per 1M tokens) | Output ($ per 1M tokens) |
|---|---|---|
| Grok 4.7 (standard) | $2.00 | $6.00 |
| Grok 4.7 (fast variant) | $4.00 | $12.00 |
| Grok 4.6 (for comparison) | $2.00 | $6.00 |
| GPT-5.6 Sol Max | $4.00 | $20.00 |
| Claude Fable 5.1 | $10.00 | $50.00 |
| GPT-6 Astra | $10.00 | $50.00 |
Against the two models it trails on the Intelligence Index, the pricing gap is enormous rather than incremental. Claude Fable 5.1 and GPT-6 Astra both list at $10 per million input tokens and $50 per million output tokens, meaning Grok 4.7's output pricing sits at roughly one-eighth of either model's list rate [5][2]. That is the trade the entire release is built around: an eight-point gap on the composite Intelligence Index in exchange for an eight-fold reduction in output token cost.
Why Hold Pricing Flat on a Bigger Model
Holding the price line on a materially larger base model is a deliberate positioning choice, not an accident of unit economics. A bigger model with a longer training run almost always costs more to serve at inference time, and xAI absorbing that cost rather than raising the sticker price signals that the company is optimizing for adoption and price-performance perception ahead of near-term margin on this specific SKU [6]. It also keeps Grok 4.7 positioned against the field the same way Grok 4.6 was: not as a head-to-head competitor to Claude Fable 5.1 or GPT-6 Astra on raw capability, but as the far cheaper option for teams whose workload does not need the extra eight points on the Intelligence Index badly enough to pay five to eight times more per token.
The Token Consumption Problem
Here is the catch, and it is a real one rather than a theoretical caveat. Artificial Analysis's own measurement found Grok 4.7 consumes roughly 2.25 times more output tokens per task than Grok 4.6 did to reach a comparable answer, which means actual per-task spend rose meaningfully even though the per-token rate did not move at all [6]. A separate independent comparison put a concrete number on the same effect against a competitor: Grok 4.7 consumed approximately 81,000 output tokens on a representative task where GPT-6 Astra used only about 27,000, roughly a three-fold difference in token consumption for what should be a comparable unit of work [9].

That token-verbosity gap is the real reason "same price per token" does not mean "same price per task," and it is the single most important number in this whole pricing section for anyone actually budgeting a migration. A model that generates three times the tokens to answer the same question is not three times cheaper just because its list price looks that way, and any cost comparison that stops at the per-million-token rate without measuring actual token consumption on your own workload will be wrong, potentially by a wide margin.
python/code def real_task_cost(input_tokens: int, output_tokens: int, output_token_multiplier: float = 1.0) -> float: """Grok 4.7 lists the same $2/M input and $6/M output pricing as Grok 4.6, but Artificial Analysis measured roughly 2.25x more output tokens generated per task on Grok 4.7 than on Grok 4.6 for comparable work. A flat per-token price does not mean a flat per-task price once that multiplier is factored in.""" input_price_per_million = 2.0 output_price_per_million = 6.0 adjusted_output_tokens = output_tokens * output_token_multiplier return (input_tokens / 1_000_000 * input_price_per_million) + ( adjusted_output_tokens / 1_000_000 * output_price_per_million ) # A representative agentic coding task: ~12k input tokens, ~1,800 output tokens # on Grok 4.6's own consumption profile grok_4_6_cost = real_task_cost(input_tokens=12_000, output_tokens=1_800, output_token_multiplier=1.0) grok_4_7_cost = real_task_cost(input_tokens=12_000, output_tokens=1_800, output_token_multiplier=2.25) print(f"Grok 4.6 estimated task cost: ${grok_4_6_cost:.5f}") print(f"Grok 4.7 estimated task cost at measured token multiplier: ${grok_4_7_cost:.5f}") print(f"Real cost increase despite identical per-token pricing: {(grok_4_7_cost / grok_4_6_cost - 1) * 100:.1f}%")
Step 6: Calling Grok 4.7 From the API
Grok 4.7 is available through the SpaceXAI API as grok-4.7, and the API surface remains OpenAI-compatible, so an existing integration built for Grok 4.6 or an OpenAI-format client generally needs only a base URL and model name change [1]. Reasoning effort remains a configurable parameter, and given the token-consumption finding above, it is worth setting deliberately rather than accepting whatever default your client library ships with.
python/code from openai import OpenAI client = OpenAI( api_key="YOUR_SPACEXAI_API_KEY", base_url="https://api.x.ai/v1", ) response = client.chat.completions.create( model="grok-4.7", messages=[ {"role": "system", "content": "You are a senior engineer reviewing a pull request for correctness and risk."}, {"role": "user", "content": "Review this diff and flag any behavior change before it merges."}, ], extra_body={"reasoning_effort": "high"}, # low | medium | high | xhigh ) print(response.choices[0].message.content)
Inference Speed and Latency
Artificial Analysis's own measurement puts Grok 4.7 at xhigh effort generating output at 39.3 tokens per second, on the slower side relative to a median of 72.5 tokens per second among comparably priced reasoning models, but with a time-to-first-token of 0.88 seconds, well ahead of the 3.70-second median for that same peer group [7]. In practical terms, Grok 4.7 starts responding quickly but then generates its full answer more slowly than most peers in its price tier, a profile that fits a model spending real effort on multi-step reasoning rather than racing to a short answer. Combined with the higher token consumption noted above, that slower generation speed compounds directly into wall-clock latency on long agentic loops, worth benchmarking against your own timeout budgets before committing a latency-sensitive path to this model.
Step 7: Where It Falls Short, and Why That Matters for Buyers
The Intelligence Index Gap
The single most important number for anyone comparing Grok 4.7 against the two most capable models on the market is the eight-point gap on the Artificial Analysis Intelligence Index v4.3.2, 46 against 53 for both Claude Fable 5.1 and GPT-6 Astra [4]. That gap is real and it is not explained away by any single benchmark in this post, it shows up as a genuine composite difference across ten separate evaluations. Anyone choosing Grok 4.7 specifically for raw general capability, rather than for its price or its specific strengths in EEBench, Harvey, or AA-Briefcase, is choosing a model that independent measurement places meaningfully behind the current frontier.
HealthBench Professional
On HealthBench Professional, a benchmark for professional medical reasoning tasks, Grok 4.7 scores 56.7 percent, a real improvement over Grok 4.6's 48.5 percent, but still trailing both GPT-6 Astra and Claude Fable 5.1, which land in the low-to-mid 60s on the same benchmark [3][2]. For any application touching medical or health-adjacent reasoning specifically, that gap is large enough to warrant testing the more expensive alternatives before defaulting to the cheaper model on cost alone.
Losing to an Open-Weight Model on Terminal-Bench
The result already mentioned above deserves repeating on its own, because it is arguably the most damaging single data point in this entire comparison for SpaceXAI's own coding-agent positioning: on Terminal-Bench 4.0, independent measurement placed Grok 4.7 below DeepSeek V4.1 Flash, an open-weight model, not just behind the two closed-weight frontier leaders [9]. A model marketed specifically around agentic coding and terminal automation losing to a freely available open-weight competitor on exactly that benchmark is a genuinely notable weak point, not a minor footnote, and it should carry real weight for any team evaluating Grok 4.7 specifically for unsupervised shell automation rather than IDE-assisted coding with a human reviewing along the way.
Step 8: Safety Claims Worth a Skeptical Read
SpaceXAI's own launch materials make several safety-related claims worth noting, with the caveat that these figures come from the company's own launch post rather than independent verification, and should be read with the same caution applied to any self-reported benchmark elsewhere in this post [1][3].
HackerBench and LatchBio Biosafety
xAI describes Grok 4.7 as the strongest model it has tested on refusals and jailbreak resistance, citing a HackerBench v0.3 score of 3.3 percent risky prompts allowed through, and a LatchBio Biosafety score of 62.4 percent that the company frames as a leading result [1][3]. The stated design goal is distinguishing benign from genuinely dangerous requests within dual-use domains rather than broadly refusing an entire topic area, which is a reasonable goal on its face, but one that is difficult to verify from a company's own launch post alone. Treat these two figures as a claim worth following up on through independent red-teaming coverage as it emerges, the same way the composite Intelligence Index score required Artificial Analysis's own measurement to properly contextualize, rather than as a settled fact about the model's real-world safety behavior.
When to Actually Pick Grok 4.7 vs the Alternatives
Rather than searching for a single "best" answer, the numbers above point toward a genuine decision framework based on task shape and budget:
- Long-horizon knowledge work, legal-adjacent tooling, or electrical engineering domains: Grok 4.7's results on AA-Briefcase, Harvey Legal Agent Benchmark, and EEBench make it worth a direct trial, especially given the price gap against Claude Fable 5.1 and GPT-6 Astra.
- IDE-assisted coding with a human reviewing every change: Grok 4.7's CursorBench and DeepSWE gains are real, though Claude Fable 5.1 still leads CursorBench by roughly five points. Worth an A/B test against your own repository rather than assuming either model wins on your specific codebase.
- Unsupervised shell automation or long autonomous terminal sessions: this is Grok 4.7's clearest weak point. The Terminal-Bench gap against both Claude Fable 5.1 and GPT-6 Astra is wide, and independent measurement even places it behind an open-weight competitor on this exact benchmark. Lean toward GPT-6 Astra or Claude Fable 5.1 here, or add substantially stronger automated verification if cost forces you to stay on Grok 4.7.
- Maximum raw capability regardless of price: Claude Fable 5.1 and GPT-6 Astra's shared eight-point lead on the Intelligence Index makes either one worth the premium for workloads where a capability gap matters more than a five-to-eight-fold price difference.
- Budget-constrained deployments where price-performance in-tier matters more than beating the frontier: Grok 4.7's score of 46 against a same-price-tier median of 24 is a genuinely strong result here, and this is the segment the pricing strategy is actually built for.
python/code def choose_model_for_task(task_type: str, budget_sensitive: bool = True) -> str: """A simple task-shape router based on the benchmark gaps documented in this post. Not a substitute for testing against your own task distribution, but a reasonable starting default.""" unsupervised_shell_tasks = {"autonomous_shell", "long_terminal_session", "unsupervised_ops"} knowledge_work_tasks = {"legal_agent", "electrical_engineering", "long_horizon_research"} if task_type in unsupervised_shell_tasks: # Terminal-Bench 4.0: Grok 4.7 26% (AA) vs Claude Fable 5.1 55%, GPT-6 Astra 60% return "claude-fable-5.1" if not budget_sensitive else "gpt-6-astra" if task_type in knowledge_work_tasks: # EEBench 64.0%, Harvey Legal Agent Benchmark 19.6%, AA-Briefcase 1657 return "grok-4.7" if task_type == "ide_assisted_coding": # CursorBench 4.0: Grok 4.7 46.3% vs Claude Fable 5.1 51.8% return "grok-4.7" if budget_sensitive else "claude-fable-5.1" return "grok-4.7" # default: budget-tier general work
Common Mistakes When Evaluating This Release
- Citing the parameter count as confirmed fact. The widely repeated 2.1 trillion figure traces back to an unconfirmed Elon Musk comment from months before launch, not to xAI's own Grok 4.7 release notes. Cite it as reported and unconfirmed, not as a published spec.
- Comparing self-reported and independently measured benchmark numbers as if they were the same figure. Terminal-Bench 4.0 alone has two publicly cited numbers, 38.0 percent from SpaceXAI's own harness and 26 percent from Artificial Analysis, and conflating them in a single chart misrepresents the result either way.
- Reading the flat $2/$6 pricing as flat total cost. Artificial Analysis's own measurement of roughly 2.25 times higher token consumption than Grok 4.6 means actual per-task spend can rise meaningfully even when the per-token rate does not move.
- Treating the Intelligence Index score in isolation. A 46 looks unremarkable next to 53, but the same score is well above the median for Grok 4.7's own price tier. Which framing matters depends entirely on what you are actually comparing it against.
- Assuming a win on one specialized benchmark like EEBench or Harvey Legal Agent Benchmark generalizes to coding-agent reliability. These are genuinely different task shapes, and the Terminal-Bench and Vals Index results show real weakness in exactly the autonomous-agent domain SpaceXAI marketed this release around.
Production Notes: Fitting a Coding Model Like This Into a Real Pipeline
Any team building an automated pipeline on top of a model like Grok 4.7 runs into the same lesson that applies across this entire comparison: a launch-day benchmark chart is a starting point for evaluation, not a substitute for testing on your own task distribution. That is true whether the pipeline is a coding agent routing tasks between Grok 4.7, Claude Fable 5.1, and GPT-6 Astra based on task type, or a content pipeline like Text2Shorts in Miraflow AI, which turns a topic into a script, then scene visuals, then a finished vertical video in one flow. A model that is cheap per token but generates three times as many tokens to finish a task is a different real-world cost profile than the sticker price suggests, in exactly the same way a fast, cheap model that occasionally produces a confidently wrong answer is a different engineering problem than a slower model that reliably flags its own uncertainty. Teams evaluating Grok 4.7 for an internal coding agent or knowledge-work pipeline should measure token consumption and task success rate on their own workload rather than trusting either xAI's launch numbers or Artificial Analysis's aggregate figures alone, the same discipline that mattered when comparing Claude Fable 5.1's own pricing and benchmarks or GPT-6 Astra's cybersecurity benchmarks against their own predecessors. If you are exploring how a fast, budget-tier coding model like Grok 4.7 fits into a broader content or automation workflow, Miraflow AI covers the visual side of that pipeline directly in the browser: AI video, AI images, YouTube thumbnails, and AI music, with no local setup required.
Here is a Wan-style video prompt built around the core tension in this release, a larger model shipping at an unchanged price, useful if you want to illustrate the concept rather than just describe it:
A tall pastel-toned toy building-block tower stands beside a shorter matching tower on a softly lit tabletop, both towers carrying an identical small round price tag clipped near their base. Over a few smooth seconds, the shorter tower slowly grows additional blocks stacking upward until it nearly matches the taller tower's height, while both price tags remain perfectly still and unchanged throughout. Clean pastel editorial illustration style rendered as smooth motion, gentle even studio lighting, no readable text, no logos, no people, slow steady camera push in toward the two towers as the growth completes.
Frequently Asked Questions
Is Grok 4.7 better than Claude Fable 5.1 or GPT-6 Astra? Not on the composite Artificial Analysis Intelligence Index, where it scores 46 against 53 for both. It does lead both on specific benchmarks, including EEBench and the Harvey Legal Agent Benchmark, and it beats Claude Fable 5.1 specifically on DeepSWE v1.1. Which model is actually better depends heavily on the specific task shape you are evaluating.
How many parameters does Grok 4.7 have? xAI has not officially confirmed a parameter count in the Grok 4.7 launch notes. A widely repeated figure of 2.1 trillion parameters traces back to an unconfirmed comment from Elon Musk made months before this release and should be treated as reported, not official.
Is Grok 4.7 cheaper than Claude Fable 5.1 and GPT-6 Astra? Yes, substantially, on a per-token basis. Grok 4.7 lists at $2 per million input tokens and $6 per million output tokens, versus $10 and $50 for both Claude Fable 5.1 and GPT-6 Astra. Real per-task cost is closer than the sticker price suggests, however, since Grok 4.7 was independently measured generating roughly 2.25 times more output tokens per task than Grok 4.6.
What is the context window and knowledge cutoff? 500,000 tokens of context, unchanged from Grok 4.6, with a knowledge cutoff of May 2026, according to independent measurement rather than xAI's own launch materials.
Why did Artificial Analysis measure a different Terminal-Bench score than xAI's own launch post? SpaceXAI's own internal harness reports Terminal-Bench 4.0 at 38.0 percent. Artificial Analysis, running the benchmark independently with its own task selection and scoring rubric, measured 26 percent. Both numbers can be accurate under their own methodology, which is why an independently run benchmark is worth seeking out before making a decision based on a self-reported score alone.
Should I switch a coding agent from Grok 4.6 or another model to Grok 4.7? It depends on the task. Grok 4.7 shows real generational gains on CursorBench 4.0 and DeepSWE v1.1 for IDE-assisted, human-reviewed coding work. For unsupervised shell automation or long autonomous terminal sessions, the Terminal-Bench and Vals Index results suggest testing GPT-6 Astra or Claude Fable 5.1 alongside Grok 4.7 before committing.
Conclusion
Grok 4.7 is a genuinely more capable coding and knowledge-work model than Grok 4.6, and it earns that upgrade at exactly the same price per token, a real achievement given that the base model itself grew. It is not, on the evidence gathered here, a model that closes the gap to Claude Fable 5.1 or GPT-6 Astra on general capability, and the eight-point difference on the Artificial Analysis Intelligence Index is not explained away by any single benchmark in this comparison. Where Grok 4.7 genuinely shines is narrower and more specific than the launch materials suggest: long-horizon knowledge work, electrical engineering problems, and legal-agent tasks, exactly the domains reflected in its EEBench, Harvey, and AA-Briefcase scores. Where it struggles is just as specific: independently measured Terminal-Bench performance, a same-tier composite score, and a token-consumption profile that means the flat sticker price does not translate into a flat real-world cost. Teams that test their own workload against all three models, rather than trusting a single launch-day chart from any of the three labs involved, will end up with a genuinely useful answer. Teams that pick based on the Intelligence Index number alone will miss both Grok 4.7's real strengths and its real weaknesses.
References and Sources
[1] SpaceXAI. "Introducing Grok 4.7."
[2] DataCamp. "Grok 4.7: Features, Benchmarks, and Pricing."
[3] MarkTechPost. "SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6."
[4] The Decoder. "xAI Launches Grok 4.7 at Bargain Prices, But Benchmarks Reveal a Wide Gap to Claude and GPT-6."
[5] officechai. "SpaceXAI Launches Grok 4.7, Beats GPT-5.6 Sol And Fable 5.1 On Some Benchmarks."
[6] CellCog. "Grok 4.7 Released: Specs, Price, Benchmarks, Rumors Graded."
[7] Artificial Analysis. "Grok 4.7: Intelligence, Performance & Price Analysis."
[8] Artificial Analysis. "Benchmarking Grok 4.7."
[9] VGTimes. "Grok 4.7 Beats GPT-6 Astra in Office Tasks, But Comes Out Behind the Previous Version in an Independent Ranking."
[10] Kingy AI. "Grok 4.7 Benchmarks vs GPT, Claude & Gemini."


