P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Grok 4.7: 38% on Terminal-Bench in xAI's harness, 26% in an independent one, and on by default in Copilot

xAI reports 38.0% for Grok 4.7 on Terminal-Bench 4.0 and Artificial Analysis measures 25.8% on the same 66 tasks. The model went generally available in GitHub Copilot the same day, where unconfigured models are enabled by default.

By Parminder Kumar Sharma · · 18 min read

Editorial illustration for the briefing: Grok 4.7: 38% on Terminal-Bench in xAI's harness, 26% in an independent one, and on by default in Copilot

The same benchmark, twelve points apart

xAI's model card for Grok 4.7 reports 38.0% on Terminal-Bench 4.0 at its highest reasoning effort. Artificial Analysis, running the same 66 task benchmark against the same effort setting, records 25.8%. Applied to 66 tasks that is roughly 25 tasks solved against 17, a difference of about eight tasks. Both figures were published on launch day, 21 September 2026, and on that same day GitHub made the model generally available in Copilot, where a model nobody has configured is switched on by a policy rather than by a person.

What the gap does not establish: that either number is wrong, or that anyone inflated anything. xAI's card names the harness behind its figure and warns in the same paragraph that absolute scores "remain sensitive to the agent harness". Artificial Analysis publishes its harness too. These are two measurements of the same model taken with different equipment, and the difference belongs to the equipment at least as much as to the model. The question for a UK security or platform lead is narrower: which of those two setups resembles the way your developers will actually run the model, and did anyone in your organisation decide they could?

What xAI claims and what the independent record shows

The launch post is short. It carries a benchmark table, two price lines, a subtitle that promises "Twice as fast, at half the price of comparable models", and a body sentence that the model is "Served at the same price and speed as Grok 4.6". The 30 page model card is longer and more careful. Artificial Analysis publishes a third view, built from its own runs against the first party API. The table below sets the claims against that independent record, fetched on 22 September 2026, the day after launch.

xAI's launch claims against the Artificial Analysis record for Grok 4.7 (xhigh), fetched 22 September 2026. Sources: x.ai/news/grok-4-7, the Grok 4.7 model card, artificialanalysis.ai/models/grok-4-7.

ClaimIndependent recordWhat the record does not establish
Terminal-Bench 4.0: 38.0% at xhigh25.8% at xhigh, 24.7% at highWhich of the two harnesses predicts your own results
"Twice as fast, at half the price of comparable models"39.2 output tokens per second, ranked 152 of 206 models. GPT-6 Astra (max) 60.7, Claude Fable 5.1 (max) 65.7, GPT-5.6 Sol (max) 82.0What the speed claim is measured against: the page does not say
"Served at the same price and speed as Grok 4.6"Price identical at $2 and $6 per million tokens. Speed: Grok 4.6 (high) 58.6 tokens per second, against 39.2 for Grok 4.7 (xhigh) and 46.5 at highWhether the measured speed gap survives launch week capacity
Card: reaches results with "fewer steps and fewer output tokens than other frontier models"240M output tokens across the Intelligence Index against a median of 92M, described as "very verbose"Token use on CursorBench 4.0, where xAI's own chart measures it
GDPval Elo 1,695 at xhighGDPval-AA 1,695.21, the same numberAnything beyond this one agreed figure

That last row matters for fairness. Where xAI quotes GDPval it is quoting Artificial Analysis runs, and the numbers match to the decimal: 1,695 against 1,695.21 for Grok 4.7, 1,735 against 1,734.67 for Claude Fable 5.1, 1,542 against 1,541.89 for GPT-6 Astra. This is not a vendor that ignores the independent scoreboard. On one benchmark, the one it ran itself in its own agent, its number and the independent number come apart.

On the headline aggregate, Artificial Analysis puts Grok 4.7 (xhigh) at 46.4 on its Intelligence Index v4.3.2, ranked 20th of 206 models in its class, against 52.7 for GPT-6 Astra (max) and 53.4 for Claude Fable 5.1 (max). Grok 4.6 (high) scored 44.3, so the independent gain generation on generation is about two points. Terminal-Bench 4.0 is one of the ten evaluations inside that index, so the lower terminal score is part of the 46.

The price claim holds per token and gets more complicated per task. Grok 4.7 lists at $2 input and $6 output per million tokens, against $4 and $20 for GPT-5.6 Sol and $10 and $50 for both GPT-6 Astra and Claude Fable 5.1. On the Artificial Analysis cost per Intelligence Index task, which prices the tokens each model actually consumed, Grok 4.7 (xhigh) cost $3.74 per task, against $7.63 for Claude Fable 5.1 (max), $3.26 for GPT-6 Astra (max) and $1.99 for GPT-5.6 Sol (max). Against Anthropic's model the "half the price" line holds on that measure, at 49% of the cost. Against OpenAI's two it does not: Grok 4.7 cost 15% more per task than GPT-6 Astra and 88% more than GPT-5.6 Sol, because verbosity eats the cheaper rate. The same arithmetic runs through time: Artificial Analysis estimates 1,126 seconds per index task for Grok 4.7 against 449 for GPT-6 Astra, 751 for Claude Fable 5.1 and 353 for GPT-5.6 Sol.

Why one benchmark gives two answers

For each of four models, two dots joined by a line show the Terminal-Bench 4.0 task success rate in xAI's Grok 4.7 model card and the rate measured by Artificial Analysis. Grok 4.7 at xhigh is 38.0 per cent in the card and 25.8 per cent independently. Grok 4.6 at high is 20.3 and 21.2 per cent. GPT-5.6 Sol at max is 37.3 and 39.9 per cent. Claude Fable 5.1 at max is 57.9 and 52.0 per cent. Notes give each harness and say the card does not state the harness used for the other models.
Drawn from the Grok 4.7 model card section 2.3 and the Artificial Analysis model page and benchmarking methodology, both fetched 22 September 2026.

Three things routinely move a benchmark score without anyone behaving badly, and only one of them is in play here.

  • The harness. xAI evaluated Grok 4.7 on Terminal-Bench 4.0 "with the Grok Build harness", with the runs "conducted by Harbor". Artificial Analysis uses the mini-swe-agent harness, over the full 66 task set, with pass at one scoring averaged over three repeats, a cap of 500 agent steps, and, in its own words, "no context compaction or summarization: the agent always sees its full transcript". Grok Build is xAI's own coding agent and the card says Grok 4.7 was trained to understand its own harness natively. A model tuned to its house agent, measured in that agent, is measured at its best. That is a reasonable thing for a vendor to publish and a poor thing for a buyer to assume.
  • The effort setting. Not the explanation here. Both figures are at xhigh. Artificial Analysis also publishes the high variant at 24.7%, so dialling effort down costs about one point, not twelve.
  • The version and the date. Also not the explanation. Both are Terminal-Bench 4.0, both use the 66 task set described in the card, and both were published within a day of each other.

Terminal-Bench 4.0, as reported in the Grok 4.7 model card and as measured by Artificial Analysis. Differences are our arithmetic.

Model and effortxAI model cardArtificial AnalysisDifference
Grok 4.7 (xhigh)38.0%25.8%12.2 points lower
Grok 4.6 (high)20.3%21.2%0.9 points higher
GPT-5.6 Sol (max)37.3%39.9%2.6 points higher
Claude Fable 5.1 (max)57.9%52.0%5.9 points lower

Read the column of differences rather than any single row. For three of the four models the two scorekeepers land within six points of each other, and they disagree in both directions. The figure xAI publishes for Anthropic's Claude Fable 5.1 is nearly six points higher than the independent one, which is the opposite of the pattern a reader looking for bad faith would expect. Only Grok 4.7 moves by twelve. The card does not state which harness produced the competitor scores, so the honest reading is that a house harness score and a generic harness score are different quantities, and the difference is largest for the model whose house harness it is.

One plausible mechanism, and this is inference rather than something either source establishes: a generic harness that never compacts context and stops at 500 steps will punish a verbose model on long tasks, and Grok 4.7 is measurably verbose. Neither xAI nor Artificial Analysis says this is the cause, and nobody has published a run of Grok 4.7 in a third harness that would settle it. If your developers work in Copilot, Cursor or another third party agent, they are not using Grok Build, which is the harness behind the 38.0%.

The model card exists, and it is better than the launch post

This briefing began from the premise that xAI had shipped a model to general availability without publishing a model card. That premise was wrong, and correcting it is worth more than the story it was meant to support. The card exists: a 30 page PDF headed "Model Card: Grok 4.7, September 21, 2026, Revision: 2026-09-21", built on launch day, served from media.x.ai with a last modified timestamp of 15:32 GMT on 21 September. The launch page, in the HTML served to us, does not link to it. Neither did the Grok 4.6 announcement link the Grok 4.6 card. Cards for Grok 4, Grok 4 Fast, Grok 4.1 and Grok 4.20 sit on a different host again, data.x.ai, under dated filenames.

What the Grok 4.7 model card states, and what it leaves out, on the points a buyer is likely to ask about.

TopicStated in the cardNot stated in the card
Model sizeNothing. The Grok 4.6 card described a "1.5T-scale model family"; the announcement says only that 4.7 uses "a new, larger base model"Any parameter count for Grok 4.7, including the 2.1 trillion figure carried by press reports of Elon Musk's posts
Training dataPublic data, internally generated data, licensed data, and "supplemental training on anonymized Cursor workflow data". Pretraining cutoff June 2026Any SpaceX engineering, Starlink telemetry or failure log corpus, which appears only in secondary coverage
Terminal-Bench methodGrok Build harness, runs conducted by Harbor, 66 tasks, agents given up to eight hours, scores "sensitive to the agent harness"The harness behind the competitor scores, and the number of trials per task
Where it runsxAI API, Grok Build, Cursor on every plan tier, Microsoft Office add-ins, and gateways including OpenRouter, Vercel and CloudflareGitHub Copilot, although GitHub announced general availability the same day
Cyber postureCyberGym 80.3% and CVE-Bench 37.7% at high, both unsafeguarded; HackerBench harmful compliance 3.31% at high and 4.02% at xhighAny independent replication of HackerBench, which is an internal xAI suite
Safety movementSelf-harm compliance 1.05%, up from 0.84% for Grok 4.6; general refusal compliance 1.10%, up from 0.93%; jailbreak compliance downConfidence intervals, so whether those small rises are meaningful cannot be judged

The card earns credit on method. It names nine third party evaluation partners, states the harness for each benchmark it reports, discloses that Grok 4.6 was used as the grader model for the Grok 4.7 HealthBench results, and says plainly that its cyber capability tests were run without the production safeguards, so that the scores measure ability rather than refusal. It also flags an unrestricted configuration supplied to third party evaluators. Those are the disclosures a reviewer needs in order to discount a number, and most of them are absent from the launch page.

Two small inconsistencies are worth knowing before you quote a figure. The announcement table gives EEBench 64.0% for Grok 4.7; the card gives 66.0% at xhigh. The developer documentation states a knowledge cutoff of May 2026 while the card states a pretraining data cutoff of June 2026, which may be two different measurements rather than a contradiction. Neither changes a purchasing decision. Both are reasons to cite the card rather than the post.

In Copilot, the default is not a decision

GitHub's changelog entry of 21 September puts Grok 4.7 in the model picker across Visual Studio Code, Visual Studio, the Copilot CLI, the cloud agent, the Copilot app, JetBrains, Xcode and Eclipse, on Pro, Pro+, Max, Business and Enterprise plans, billed at provider list pricing under usage based billing. GitHub's own pricing reference lists it as generally available at $2.00 input, $0.50 cached input and $6.00 output per million tokens up to 200,000 input tokens, and $4.00, $1.00 and $12.00 above that threshold. The sentence that should reach your change advisory board is the one about enablement: "new models are automatically enabled unless an administrator has turned off the global default or explicitly disables this model".

A flow diagram of how a new generally available model reaches users in GitHub Copilot. If an administrator has configured the model, that setting applies. If not, it is labelled Delegate to Default Policy: with the default availability policy on it is enabled for every Business and Enterprise seat in scope, and with the policy off it stays disabled. A panel lists what stays off regardless: pre-release and open weight models, Claude Fable 5 and 5.1, and models failing residency rules.
Drawn from GitHub's Copilot documentation on default model availability and on managing model availability in an enterprise, fetched 22 September 2026.

Look at the label in the middle of that flow. A model that nobody has configured shows in enterprise settings as Delegate to Default Policy. The words describe an act of delegation, which sounds like a choice somebody made. It is the name for the absence of a choice. The same trap sits in the release status: GA is a statement about GitHub's rollout, not an assurance that the model suits your risk appetite, your regulated code or your data classification. Neither label is dishonest. Both read, to a busy administrator scanning a settings page, like evidence that somebody already thought about it.

GitHub Copilot model controls for Business and Enterprise plans, from GitHub's documentation on default model availability and enterprise model management.

ControlWhat it doesWhat it does not do
Default availability for released models, left onEnables every unconfigured generally available model for seats in scope as soon as it shipsEvaluate the model, notify anyone, or record who accepted the risk
Setting a model to DisabledBlocks that model for every user and agent app in scope, and keeps it out of automatic enablementStop the same model reaching your code through Cursor, a gateway or a personal API key
Delegate to organisationsHands the decision to each organisation ownerGuarantee a decision: an unconfigured model then inherits that organisation's default policy instead
Data resident or FedRAMP restrictionsExcludes models that do not respect those policies from default enablementApply to models you have explicitly enabled yourself

On one axis Grok 4.7 arrives in a stronger position than the strongest Anthropic model. GitHub's hosting documentation says xAI operates its models in Copilot "under a zero data retention API policy", with prompts and completions held only in memory for the life of the request. Claude Fable 5 and Claude Fable 5.1 are the exception to GitHub's retention agreement: Anthropic retains prompts and outputs by default to run safety classifiers, eligible enterprises can request zero retention only through a time limited exemption running to the end of 2026, and those models are therefore excluded from default enablement altogether. If your reason for scrutinising a new model is data handling, that is an argument in xAI's favour and against Anthropic's flagship, and it should be weighed as one. It is also a commitment reported by GitHub about a third party, not something you can inspect. Ask for it in your contract paperwork rather than reading it off a documentation page.

Cursor works the other way round. Its enterprise documentation states that "When new models become available, Cursor doesn't immediately enable them for all enterprise teams", and that enterprise teams opt in. Model access control there is described as an Enterprise feature, configured under Team Settings and Organization Groups, and the two layers merge on a "most-permissive (union)" rule, so a group can widen access but never tighten it below what the team allows. Two consequences follow. Below the Enterprise tier the documentation describes no administrative model restriction at all, while the Grok 4.7 card lists Cursor availability as "all users, on every plan tier". And if you run Cursor Enterprise alongside Copilot, your strictest setting must live on the team baseline, because the union rule means a permissive group wins.

One pricing detail that gets misread. The "fast variant with twice the output speed at twice the price" from the announcement is, per xAI's documentation, "the same model served on faster infrastructure", available only in Cursor and Grok Build, excluded from Grok Build's free tier, and not offered on the public xAI API. Copilot bills Grok 4.7 at the standard $2 and $6 rates, which implies the standard service rather than the fast one; GitHub does not say so directly, so treat that as inference and confirm it if latency matters to your business case.

Procurement questions for a UK buyer

Who you are contracting with, and where the processing happens, are answerable from primary sources only in part. The model card records that "SpaceXAI is a doing-business-as (DBA) name of XAI LLC", so the brand on the launch page is not the contracting entity. For direct API use, xAI's regions documentation says the global endpoint may route "between regions for capacity or reliability, so the processing location is not guaranteed", and the only regional alternative it documents is a United States endpoint, serving grok-4.7 and grok-4.6 at a 10% token premium, whose guarantee covers request handling, inference, moderation and retained request data but not files, collections or server side tools. There is no documented endpoint that keeps processing in the United Kingdom or the European Economic Area. The data processing addendum is linked from x.ai's footer, but repeated requests for it returned an HTTP 403 bot challenge, so this briefing does not characterise its terms; that is a gap you should close before signing, not one to fill with assumption.

Take this with you

Questions to put in writing before a UK deployment

  • Which legal entity is the processor, and does the data processing addendum incorporate the UK International Data Transfer Addendum or the standard contractual clauses for restricted transfers?
  • Which endpoint will our traffic use, and if it is the global endpoint, what is the transfer mechanism for processing outside the United Kingdom?
  • If we route through GitHub Copilot or a gateway, whose agreement governs the data, and does the zero data retention commitment appear in a document we hold?
  • Does the supplemental training on anonymised Cursor workflow data include any tenant like ours, and what settings exclude us from any such corpus?
  • Who maintains the subprocessor list, and how will we be notified of changes to it?
  • What is the retention and deletion position for prompts, outputs and request metadata on the endpoint we will actually call?

What to do, in order

Take this with you

Actions in the order worth doing them

  • Open your Copilot enterprise AI controls and look at the model list. If Grok 4.7 shows as Delegate to Default Policy and the Default availability policy is on, it is already live for your developers.
  • Make the status explicit either way: set Grok 4.7 to Disabled pending review, or Enabled for a named pilot organisation with a budget ceiling.
  • Decide the standing question behind the specific one: whether Default availability for released models should stay on for organisations handling regulated or sensitive code.
  • Check Cursor separately. Confirm on Enterprise that your team has not opted in, and establish which teams are below Enterprise and therefore have no model restriction control.
  • Read the model card before the launch post, in particular the Terminal-Bench harness note, the cyber and jailbreak sections, and the self-harm and refusal numbers.
  • Run your own evaluation in the harness your developers use, on your repositories, and compare it with the harness the vendor used.
  • Model the cost on completed tasks rather than list prices, and watch the long context tier that doubles the Copilot rate above 200,000 input tokens.
  • Send the procurement questions above to whoever holds the contract, whether that is xAI, GitHub or a gateway.
  • Record the decision, the evidence and the date in your change record, so the answer to who approved it is a name rather than a policy toggle.

The question that exposes the gap

Benchmark scores move. Artificial Analysis will refresh its numbers, other harnesses will publish, and the twelve point gap on Terminal-Bench 4.0 will narrow, widen, or be explained by someone who runs the model three ways. That argument will resolve itself without you. The governance point will not. A frontier model became generally available inside your development tooling on the day it was announced, priced per token against a bill you have not modelled, scored by its maker in a harness your developers do not use, and switched on by a policy setting whose name describes a decision nobody made.

So: if Grok 4.7 is live for your developers today, can you name the person who approved it, and the evidence they looked at?

Sources

  1. PrimaryIntroducing Grok 4.7, the launch announcement: used for pricing, the benchmark table, the speed and price claims and the availability listxAI (SpaceXAI)accessed 2026-09-22
  2. PrimaryModel Card: Grok 4.7, revision 2026-09-21: used for the Terminal-Bench harness statement, training data, deployment channels, cyber and safety evaluationsxAI (SpaceXAI)accessed 2026-09-22
  3. PrimaryModel Card: Grok 4.6, revision 2026-08-17: used for the 1.5T-scale description and the earlier card publishing patternxAI (SpaceXAI)accessed 2026-09-22
  4. PrimaryGrok 4.7 developer documentation: context window, default reasoning effort, token prices and the fast variant termsxAI (SpaceXAI)accessed 2026-09-22
  5. PrimaryRegional endpoints documentation: global endpoint routing and the United States endpoint guarantee and premiumxAI (SpaceXAI)accessed 2026-09-22
  6. PrimaryGrok 4.7 (xhigh) model page: Intelligence Index, Terminal-Bench 4.0, output speed, verbosity, cost per task and context windowArtificial Analysisaccessed 2026-09-22
  7. PrimaryIntelligence benchmarking methodology: the Terminal-Bench 4.0 harness, task count, repeats and step limitArtificial Analysisaccessed 2026-09-22
  8. PrimaryGrok 4.7 is now available in GitHub Copilot: availability, clients, billing and the default model enablement sentenceGitHubaccessed 2026-09-22
  9. PrimaryAbout default availability of Copilot models: which models inherit the default policy and which are excludedGitHubaccessed 2026-09-22
  10. PrimaryManaging availability of models in your enterprise: model statuses, delegation and targeted model rulesGitHubaccessed 2026-09-22
  11. PrimaryHosting of models for GitHub Copilot: the xAI zero data retention statement and the Claude Fable retention exceptionGitHubaccessed 2026-09-22
  12. PrimaryModels and pricing for GitHub Copilot: Grok 4.7 release status and the default and long context token ratesGitHubaccessed 2026-09-22
  13. PrimaryModel and integration management: enterprise model access control, new model rollout and the union merge ruleCursoraccessed 2026-09-22
  14. PrimaryGrok 4.1 model card, 17 November 2025: used to establish the earlier model card publishing patternxAIaccessed 2026-09-22
  15. PrimaryGrok 4.20 model card, 7 April 2026: used to establish the earlier model card publishing patternxAIaccessed 2026-09-22
  16. PrimaryData processing addendum linked from the xAI footer: requested but returned an HTTP 403 bot challenge, so its terms are not characterised herexAI (SpaceXAI)accessed 2026-09-22
  17. Reported byCoverage of the launch, used as a pointer to the Artificial Analysis figuresThe Decoderaccessed 2026-09-22
  18. Reported byLaunch coverage carrying the 2.1 trillion parameter and SpaceX training data claims that we could not verify in xAI primary sourcesDecrypt via Yahoo Techaccessed 2026-09-22
  19. Reported byReport of Elon Musk's pre-launch claims about Grok 4.7 parameters, timing and SpaceX dataBeInCrypto via Yahoo Techaccessed 2026-09-22

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.