SpaceXAI released Grok 4.7 on 21 September 2026, and the headline everyone repeated was that it costs exactly what Grok 4.6 cost: $2 per million input tokens and $6 per million output tokens. That part is true. The token price did not move by a cent.

What moved is the bill. Independent measurements from Artificial Analysis put the cost of running one benchmark task at $2.73, against Grok 4.6's $1.86. That is a rise of roughly 47 percent, and it happens because the new model writes far more to get to the same place. This article works through what actually changed, what it costs in practice rather than per token, where it quietly got worse, and who can use it today.

The Key Takeaways

  • Same token price, bigger bill: $2/$6 per million is unchanged from Grok 4.6, but cost per benchmark task rises from $1.86 to $2.73 because output roughly doubles.
  • A two point gain: Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index v4.3.2 against Grok 4.6's 44, measured at the same reasoning setting.
  • The headline chart is uneven but fair: SpaceXAI runs Grok 4.7 at xhigh against Grok 4.6 at high, yet Grok 4.6 scores 44 either way, so the gap holds up.
  • xhigh buys nothing on either model: both score the same at high and at xhigh, yet on Grok 4.7 the setting costs $3.74 per task and drops output to 39.5 tokens per second.
  • The 200,000 token cliff survives: cross that prompt length and the entire request re-bills at $4/$1/$12, double the advertised rate.

What Grok 4.7 Actually Is

From the publisher

Every AI model in one app

Fello AI puts GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 and more in one native Mac and iPhone app.

Download now!

It is the frontier model from SpaceXAI, the company formed when SpaceX absorbed xAI and which now trades under the combined name. The chatbot, the apps and the developer API all keep the Grok brand. The company set the model out in its official Grok 4.7 announcement on 21 September, describing it as its "most capable model for coding and knowledge work".

The vendor's framing is deliberately modest about the upgrade itself. Its own announcement says Grok 4.7 is "served at the same price and speed as Grok 4.6" and calls it "highly competitive in its class" rather than claiming a lead.

The specification sheet

Per the SpaceXAI developer documentation, the model takes text and images in and returns text only. It carries a 500,000 token context window, a knowledge cutoff of May 2026, and four reasoning effort levels: low, medium, high and xhigh, with high as the default. Rate limits are 150 requests per second and 50 million tokens per minute, and it is served from three US regions. There is no Batch API support.

SpecificationGrok 4.7Grok 4.6
Context window500,000 tokens500,000 tokens
Reasoning effortslow, medium, high, xhighlow, medium, high, xhigh
Default efforthighhigh
Modalitiestext and image in, text outtext and image in, text out
Batch APINot supportedNot supported
Knowledge cutoffMay 2026Not stated

That table is worth pausing on. The two models' documentation pages are identical apart from the model name. Same context, same effort levels, same default, same rate limits, same regions, same prices. Whatever changed happened inside the weights, not in the envelope around them. SpaceXAI describes it as "a new, larger base model" trained with "a longer reinforcement learning run on a harder mix of tasks", and gives no parameter count. Several outlets have printed a figure of 2.1 trillion parameters; it appears on no SpaceXAI page, so treat it as unconfirmed. For the previous release, see our breakdown of what changed in Grok 4.6.

Grok 4.7 Pricing and the 200,000 Token Cliff

The advertised rate is the easy part. The part that catches people is what happens to a long prompt.

Prompt sizeInput / 1MCached input / 1MOutput / 1M
Under 200k tokens$2.00$0.50$6.00
200k tokens and above$4.00$1.00$12.00

Why the cliff hurts more than it looks

The SpaceXAI pricing documentation is explicit that requests crossing the threshold "are billed at the higher rate for all tokens in the request". It is not a surcharge on the tokens past 200,000. A prompt of 199,000 tokens and a prompt of 201,000 tokens are priced an entire tier apart, and the second one re-prices everything, including the cached portion. Anyone feeding a large repository or a long document set into a 500,000 token window is operating above that line by default and paying double the rate they read on the announcement. This is the same structure Grok 4.6 used, and we covered the mechanics in our guide to Grok pricing across plans and the API.

Two smaller billing details

First, the US regional endpoint applies a 1.1x multiplier to every rate, taking the standard tier to $2.20, $0.55 and $6.60, and the long context tier to $4.40, $1.10 and $13.20. Second, there is a Fast variant, the same model on faster infrastructure at twice the rates, so $4.00, $1.00 and $12.00 below 200k tokens and $6.00, $1.50 and $18.00 above. The pricing documentation is blunt about who gets it: it is "not available on the public xAI API, and Grok Build's free tier does not include it". The announcement mentions the fast variant without that restriction, so the docs are the operative source. Third, list price is not the only price. As of 22 September 2026 the model router OpenRouter was serving the model from SpaceXAI endpoints at $1.60 input and $4.80 output per million, a fifth below list, with the long context tier discounted in step at $3.20 and $9.60. Grok 4.6 sat at full list price on the same router, so the discount is specific to the new model and may not last. For how these rates sit against other providers, see our comparison of AI model pricing.

Grok 4.7 Against Grok 4.6, Measured Fairly

This is where most coverage of the launch goes wrong, so it is worth being precise.

SpaceXAI's headline benchmark table labels its two leading columns Grok 4.7 xHigh and Grok 4.6 High. Artificial Analysis headlines the same pairing, stating in its launch post that "We evaluated the new model at xhigh reasoning effort". Both therefore front a comparison that runs the new model at a harder setting than the old one, and as the specification table above shows, Grok 4.6 supports xhigh too.

That looks like a stacked deck, and it is worth saying plainly that it is not one. Artificial Analysis quietly publishes all four combinations, and the fourth settles it: Grok 4.6 at xhigh also scores 44, exactly what it scores at high. The extra effort setting earns the older model nothing, so the two point gap is a real improvement at every setting rather than an artifact of the chart. The asymmetry is a presentation choice, not a distortion.

The like-for-like numbers

Artificial Analysis publishes a separate page for the model at the default high setting, which makes an equal comparison possible. All figures below are from Intelligence Index v4.3.2, read on 22 September 2026.

MeasureGrok 4.6 (high)Grok 4.6 (xhigh)Grok 4.7 (high)Grok 4.7 (xhigh)
Intelligence Index44444646
Rank of 202 models23rd24th17th16th
Output speed60.4 tokens/sec58.2 tokens/sec55.3 tokens/sec39.5 tokens/sec
Cost per Index task$1.86$2.32$2.73$3.74
Output tokens used94M97M200M240M

Read across that table and two things fall out. At the setting both models share, the newer one gains two index points and pays for them with 47 percent more cost per task, 8 percent less speed and more than double the output tokens.

The second is the one nobody reported. Raising the effort to xhigh adds nothing to either model's score. Grok 4.6 sits at 44 either way and the new model at 46 either way, yet the setting costs another dollar per task and a third of the speed. On this evidence xhigh is a pure expense, which makes the default the right place to leave it.

What that means for a bill

Verbosity is the mechanism, and the table isolates it. Holding the effort setting constant, Grok 4.6 emits 94 million output tokens across the benchmark and the new model emits 200 million, so the extra writing is built in rather than bought with a dial. Grok 4.6 pushed to xhigh only reaches 97 million. Artificial Analysis measured the new model at xhigh using "approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high)". Output tokens are the expensive ones at $6 per million. A model that reasons at greater length to reach a similar answer costs more even when its price list is frozen. "Same price" is true per token and misleading per job, and per job is how anyone actually pays. If benchmark terminology is unfamiliar, our guide to how AI benchmarks work unpacks the scoring.

Elon Musk framed the release around cost when it landed.

Where It Really Did Improve

None of the above means the model is a dud. The gains are real; they are just narrower than the launch coverage suggested, and they cluster in one area.

Artificial Analysis found it "gains +111 Elo over Grok 4.6 (high) on AA-Briefcase", its private benchmark for long horizon agentic knowledge work, scoring 1657 Elo and landing alongside the leading Claude models. On GDPval-AA, which asks models to produce real work products such as documents, spreadsheets and slides, it scores 1695 Elo against Grok 4.6's 1605.

SpaceXAI's own numbers point the same way. Its table reports 46.3 percent on CursorBench 4.0 against Grok 4.6's 40.4, 38.0 percent on Terminal-Bench 4.0 against 20.3, and 64.0 percent on EEBench against 53.0. These are vendor run and carry the effort asymmetry described above, but the direction is consistent with the independent tests. The pattern is a model tuned for tasks that take hours rather than seconds.

Where It Got Worse

This is the section the vendor page does not have, and it matters if your work is not agentic.

Artificial Analysis reported that "Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks". It also logged two outright regressions: AA-LCR down 3.7 percentage points and AutomationBench-AA down 1.1 percentage points. Alongside those, output speed fell at every setting and token use more than doubled.

So for a reader running short prompts, quick lookups or single file edits, this is a model that scores the same as its predecessor on most tasks, answers more slowly and writes more. It therefore costs more. The upgrade pays off if your work runs long, and not if it does not.

Musk himself set expectations a week before launch, replying on 14 September that "Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance." He put the real jump further out, adding that "Grok 4.8 will be a noticeable improvement" and that "Grok 4.9 is probably Astra/Fable class". The independent data landed close to his own forecast, at 46 on the index against a higher mark for Claude Opus 5.

Safety and Cybersecurity

SpaceXAI put unusual weight on safety in this release, saying the model "was built with an entirely new safeguard stack" and calling it "the strongest model we've tested on refusals and jailbreak resistance". It reports 62.4 percent on LatchBio's biosafety benchmark and says the model allows "only 3.3% of risky dual-use prompts through" on HackerBench v0.3, its own internal cyber benchmark.

Two caveats belong with those figures. Both benchmarks are run or owned by the vendor, and neither has independent replication yet. More importantly, a refusal rate is a measure of behaviour under test, not a capability rating, so a strong score here says nothing about how well the model does your work. The company has also begun giving "select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research".

How to Use Grok 4.7 Today

The access story is narrower than most coverage implied, and the vendor is specific about it. The announcement says the model is "available today in Cursor and Grok Build" and adds that it also reaches the Grok API alongside "third-party coding harnesses, and model routers and cloud platforms".

Every one of those is a developer surface. SpaceXAI names no consumer app and no subscription tier anywhere in the announcement, the model documentation or the release notes. Reports that it went live in the consumer Grok app on day one are not supported by any vendor page. If you subscribe to the Grok app, treat your access as unconfirmed until the company says otherwise. Our guide to what Grok costs and what the free tier includes covers the consumer plans as they stand.

Running it alongside other models on a Mac

Because the confirmed routes are all API based, the practical way to put a model like this next to the ones you already use is a client that speaks to several providers at once. Fello AI does exactly that on Mac, iPhone and iPad. It puts Grok, GPT, Claude, Gemini, Perplexity, DeepSeek, Kimi, GLM and Qwen behind one subscription at $9.99 a month or $79.99 a year, with a free tier to try it. It holds 4.7 stars from more than 27,000 reviews, and it means a release like this one becomes a model you switch to for a single question rather than another account to manage.

For the wider field rather than one vendor, our regularly updated rundown of the best AI models available right now tracks where each one currently lands.

The Verdict

Grok 4.7 is a narrow, honest upgrade wearing a wide marketing headline. If your work is long horizon agentic knowledge work, the kind measured by AA-Briefcase and GDPval, it is a genuine step up and the price list is unchanged. If your work is anything else, you are buying two index points for roughly half again as much money per task, a slower answer and twice the tokens, and you should stay on Grok 4.6 until that changes.

The more useful lesson is about how to read a launch. When a vendor holds the price list still, check the cost per task rather than the cost per token, and check which reasoning setting produced the chart. On this release both of those questions changed the answer. For what comes next, we are tracking everything known about Grok 5.

Frequently Asked Questions

Is Grok 4.7 free?

SpaceXAI has not announced a free consumer tier for Grok 4.7. The announcement lists Cursor, Grok Build, the Grok API, third-party coding harnesses and model routers, all of which are paid developer surfaces. No consumer app or subscription tier is named on any vendor page.

How much does Grok 4.7 cost?

$2.00 per million input tokens, $0.50 per million cached input tokens and $6.00 per million output tokens, for prompts under 200,000 tokens. At or above 200,000 tokens the whole request is billed at $4.00, $1.00 and $12.00. The US regional endpoint adds a 1.1x multiplier.

Is Grok 4.7 better than Grok 4.6?

Measured at the same reasoning setting, Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index v4.3.2 against Grok 4.6's 44. The gains concentrate in long running agentic work. Artificial Analysis found it broadly matches Grok 4.6 elsewhere and regressed on two tasks.

What is the Grok 4.7 context window?

500,000 tokens, unchanged from Grok 4.6. Note that this is separate from the billing threshold: once a prompt reaches 200,000 tokens, the entire request is charged at the higher long context rate even though the window runs to 500,000.

What does xhigh reasoning effort do?

It lets the model work longer before answering. On Grok 4.7 it does not raise the Artificial Analysis Intelligence Index score, which stays at 46, but it raises cost per task from $2.73 to $3.74 and cuts output speed from 55.3 to 39.5 tokens per second.