Meta released Muse Spark 1.3 on September 2, 2026, the fourth Muse Spark model in five months. It scores 61 on the Artificial Analysis Intelligence Index v4.1.1, up four points from Muse Spark 1.2, and on Meta’s own launch chart it wins every coding row and both long-context rows outright.

It also loses every single agent row. That half of the chart went almost unreported, and it is not the only thing worth a second look. The column Meta benchmarked is labelled Muse Spark 1.3 (max), a variant that is in limited preview for Meta’s partners and that you cannot use today. Meta’s own evaluation methodology says the comparison runs 1.3 at max reasoning and its predecessor at xhigh. This article covers what actually changed, the full benchmark table read from Meta’s chart rather than a summary of it, what it costs, and whether you can run it at all right now.

The Key Takeaways

  • Released September 2, 2026 in Muse Code and the Meta Model API. Meta’s announcement names no consumer surface at all, so this is a developer release.
  • It wins coding, it loses agents. On Meta’s own scorecard it takes DeepSWE v1.1, SWEAtlas CodeBase QnA and both MRCR long-context bands, ties Terminal-Bench 2.1, and loses all six agent rows to Claude Opus 5 or GPT 5.6 Sol.
  • The benchmarked model is not the shipping model. The chart runs Muse Spark 1.3 (max), still in limited preview. Meta says max reasoning arrives only after further safety testing.
  • Two index scores, one you can use. 1.3 (max) scores 62 and 1.3 (xhigh) scores 61 on Artificial Analysis Intelligence Index v4.1.1. Only xhigh is generally available.
  • Pricing has not moved since July: $1.25 input and $4.25 output per million tokens, with a 1M token context window.

What Muse Spark 1.3 Actually Changes

De l'éditeur

Tous les modèles d'IA dans une seule app

Fello AI réunit GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 et plus dans une seule app native pour Mac et iPhone.

Téléchargez maintenant !

Meta frames 1.3 as a usability release rather than a raw capability release. The pitch is that the model sustains longer work without falling over, which is a different problem from scoring higher on a single-turn test.

It asks before it acts

The headline behavioural change is collaboration. Meta says the model asks clarifying questions when prompts are ambiguous, invokes help from the user when it gets stuck, and confirms before taking consequential actions. On long tasks it adapts to preference, either reporting frequently or working silently in the background.

This matters more than it sounds. The failure mode that makes agentic coding tools frustrating is not usually a wrong answer, it is a confident wrong answer executed twelve steps deep. Meta also says it trained the model to have a better sense of what it does not know, so it flags hurdles instead of hallucinating an outcome.

It juggles more than one thread

The second change is multitasking inside a single long conversation. Meta claims 1.3 maps an incoming prompt to the correct task more accurately in messy single-threaded contexts, whether the user is steering an earlier request or interrupting it outright. Anyone who has watched a coding agent apply a correction to the wrong file will recognise the problem being solved.

Safety work gated the best mode

Meta reports stronger adversarial robustness and improved resistance to prompt injection, plus better calibration on what counts as an irreversible action. That work is also why the launch is incomplete. In Meta’s words, previously available reasoning modes shipped on day one, with max reasoning coming shortly after it finishes additional safety testing.

The Full Muse Spark 1.3 Benchmark Table

Below is Meta’s complete launch scorecard, read directly from the chart Meta published rather than from a secondary write-up. Most coverage reproduced only the rows Meta won and dropped the Muse Spark 1.2 column entirely, which is where the size of the jump actually lives.

BenchmarkMuse Spark 1.3 (max)Muse Spark 1.2 (xhigh)GPT 5.6 Sol (max)Opus 5 (max)
GDPVal-AA v2 (knowledge work)1754161517101824
JobBench (professional tool use)64.961.645.465.7
OSWorld 2.0 (computer use)66.947.662.768.3
DeepSearchQA (agentic browsing)89.485.993.090.4
Agentic IF Index (internal)57.846.260.559.1
AutomationBench (business workflows)49.438.246.750.3
MRCR 256K to 512K (long context)98.566.391.5Not reported
MRCR 512K to 1M (long context)98.155.573.8Not reported
DeepSWE v1.1 (agentic coding)75.455.073.074.0
SWEAtlas CodeBase QnA59.446.253.552.7
Terminal-Bench 2.1 (terminal coding)88.882.988.886.7

Where Muse Spark 1.3 wins outright

The long-context result is the most striking number on the page and nobody is arguing about it. Muse Spark 1.2 scored 55.5 on the 512K to 1M band of MRCR. Muse Spark 1.3 scores 98.1. That is not a tuning gain, that is a different capability, and it beats GPT 5.6 Sol’s 73.8 by a margin no amount of configuration fiddling explains away. Opus 5 has no published score in either band.

Coding is the other clean win. DeepSWE v1.1 climbs from 55.0 to 75.4, edging past Opus 5 at 74.0, and SWEAtlas CodeBase QnA at 59.4 leads both rivals by roughly six points. Terminal-Bench 2.1 lands at 88.8, a dead tie with GPT 5.6 Sol and ahead of the 86.7 Meta records for Claude Opus 5. Third-party trackers put Opus 5 higher on that test, which is the usual caveat about a vendor running a rival’s model. Against its own predecessor the jump is large and consistent, which is the fairest reading of this chart.

Where it loses, and nobody mentioned it

Meta groups the first six rows under the heading Agent.

Muse Spark 1.3 loses all six. Opus 5 takes GDPVal-AA v2, JobBench, OSWorld 2.0 and AutomationBench. GPT 5.6 Sol takes DeepSearchQA and the Agentic IF Index. Several are close, and JobBench at 64.9 against 65.7 is close enough to be noise, but a sweep is still a sweep.

This is the gap between the headline and the chart. Zuckerberg called 1.3 the biggest jump Meta has made on coding and agentic work. The coding half of that sentence is supported by Meta’s own numbers. The agentic half is supported only against Muse Spark 1.2, not against the competition Meta chose to put in the table.

The configuration problem

There is one more thing in the table that Meta discloses and most coverage skipped. The evaluation methodology document states it plainly: “We use max reasoning effort for Muse Spark 1.3, Claude Opus 5 and GPT-5.6 Sol, and xhigh for Muse Spark 1.2.”

So the 1.2 column is not the same setting as the 1.3 column. Part of every improvement in that comparison is a reasoning-effort change rather than a model change, and Meta does not break out how much. The comparison against Opus 5 and GPT 5.6 Sol is like-for-like at max. The comparison against Meta’s own previous release, which is the one generating the excited numbers, is not.

The Muse Spark 1.3 You Can Run Scores Lower

Artificial Analysis lists two entries for this release, and the difference is the whole story. Muse Spark 1.3 (max) scores 62 on Intelligence Index v4.1.1 and is in limited preview for Meta’s partners. Muse Spark 1.3 (xhigh) scores 61 and is the one in Muse Code and the Meta Model API today.

One point is not much.

What matters is which model the marketing describes. Every figure in Meta’s launch table is the max variant, and max reasoning is the mode Meta says is still in safety testing. If you sign up this morning, you are running the 61.

ModelIntelligence Index v4.1.1Availability
Claude Fable 5.1 (max)66Generally available
Claude Opus 5 (max)63Generally available
Muse Spark 1.3 (max)62Limited preview, Meta partners
Claude Fable 5 (max)62Generally available
Muse Spark 1.3 (xhigh)61Generally available
GPT 5.6 Sol (max)61Generally available
Grok 4.6 (high)61Generally available
Gemini 3.8 Flash (high)59Generally available
Muse Spark 1.2 (xhigh)57Superseded

Read as a trend line, Meta’s progress is the real headline. Muse Spark 1.1 scored 53 in July, 1.2 scored 57 in August, and 1.3 scores 61 in September. Eight points in two months is a pace nobody else on that table is matching, and it puts Meta level with GPT-5.6 Sol and within striking distance of Claude Fable 5.1. The company that was written off as a frontier player eighteen months ago is now two points off second place.

Meta Says Fewer Tokens, the Measurements Say More

Meta makes a specific efficiency claim: “In comparisons by Meta engineers, it proved to be significantly faster and more efficient, using ~20% fewer tool calls and ~25% fewer tokens.” Zuckerberg’s framing was blunter, calling the model less yappy than its predecessor.

Artificial Analysis measured something different. Running its index, Muse Spark 1.3 (xhigh) produced 100 million output tokens against a 71 million median for comparable models. Its model page calls the result notably fast but somewhat verbose. The full index cost $810.19 to run.

Both statements hold, and the reconciliation is the useful part. Meta’s figure is an internal comparison by its own engineers on coding workflows against Muse Spark 1.2, not a published evaluation, and it is scoped to that. Artificial Analysis measured the publicly available variant across a broad benchmark suite that is mostly not coding. If you are budgeting for a coding agent, Meta’s number may hold. If you are budgeting for general agentic work, plan for a verbose model and do not assume a 25% saving.

Muse Spark 1.3 Pricing and the Contributor Trade

Pricing has not changed since Muse Spark 1.1 launched the paid API in July, which is notable given three capability bumps since. The 1M token context window is also unchanged, and Artificial Analysis clocks output at roughly 209 tokens per second.

TierInput per 1MCached inputOutput per 1MCondition
Standard$1.25$0.15$4.25None
Contributor$0.10$0.002$0.20Permission to train Meta’s future models on your prompts and completions

The Contributor tier is roughly 92% cheaper and it is not a promotion. You buy the discount with your data, and on a terminal coding agent your prompts and completions are your codebase. That was our reading when the tier appeared alongside Muse Spark 1.2 and Muse Code, and nothing about 1.3 changes it. For a personal side project it is close to free frontier coding. For work under an employment contract or a client NDA, read the terms before you export an API key. Our AI pricing comparison sets the standard tier against the rest of the field.

Developers are taking the trade. Meta AI chief Alexandr Wang told Axios that a “meaningful double digit” percentage of coders are choosing the Contributor tier, and characterised the unchanged standard pricing as “aggressive”. That is the clearest signal yet that the discount is working as a data acquisition strategy rather than as a loss leader.

How to Use Muse Spark 1.3 Today

This is where expectations need managing. Meta’s announcement lists exactly two places the model runs: Muse Code and the Meta Model API. It names no consumer surface. Not the Meta AI app, not WhatsApp, not Instagram, not Facebook, not Messenger.

That is a real change of shape from the April launch of the original model, which arrived in the consumer apps first and was free to everyone. Some coverage has carried that April rollout language over to this release, and it is worth checking any claim that 1.3 is landing in WhatsApp or Instagram against what Meta actually published. If you are not a developer, there is currently nothing to open, and Meta has given no date for when that changes.

If you are a developer, the path is the same one Muse Code has used since August, so the install walkthrough in our Muse Code guide still applies, with the model updated underneath. Note that the max reasoning mode behind every number in Meta’s chart is not part of that path yet. For background on how the line got here, our explainer on Meta Muse Spark covers the original release and Muse Spark 1.1 covers the move to paid API access.

The Open Weights Promise Has Now Slipped

Muse Spark 1.3 is a closed model. Meta’s roadmap paragraph promises “bigger models, the Muse Spark open weights release, and more,” with no date attached, and Zuckerberg’s launch post repeated it.

It is worth putting his two statements side by side. On August 10, announcing the open-weight Muse Glimmer, he wrote: “Soon we’ll also release the weights for Muse Spark 1.2, our latest foundation model.” On September 2, twenty-three days later, the commitment had become “Muse Spark open weights releases coming soon,” with no version named. Muse Spark 1.2 weights never shipped, and 1.2 is no longer the latest foundation model.

The distinction that will confuse people is Muse Glimmer, which is open. It is a 30 billion parameter dense model under Apache 2.0, distilled from Muse Spark, released on August 10 and sized to run locally. It is not the frontier model and Meta has never said it was. If you want Meta weights you can download today, Glimmer is what exists. Our roundup of the best open source AI models puts it against the Chinese open-weight releases that currently set the pace.

The Verdict on Muse Spark 1.3

Muse Spark 1.3 is the best coding model Meta has shipped and the long-context result is the most impressive number of the release. At $1.25 and $4.25 per million tokens, set against Opus 5 and Fable 5.1 pricing, it is also the cheapest way to buy something close to frontier coding performance. The four-point index gain in a month is real.

What it is not is the across-the-board win the launch reads as. Meta lost every agent row on its own chart, the numbers come from a variant still in limited preview, and the comparison against its own predecessor changes the reasoning setting mid-table. If you write code and you are comfortable with the standard tier, it is worth an evaluation this week. If you want it in the Meta AI app, or you want the weights, you are waiting on a date Meta has not given for either.

Sources: Meta’s launch post and evaluation methodology at Meta AI Research, independent measurements from Artificial Analysis, and Alexandr Wang’s launch-day interview with Axios, linked above. Both Zuckerberg posts quoted above were read in full rather than via a summary.

Frequently Asked Questions

What is Muse Spark 1.3?

Muse Spark 1.3 is Meta’s frontier AI model, released on September 2, 2026 by Meta Superintelligence Labs. It scores 61 on the Artificial Analysis Intelligence Index v4.1.1, has a 1 million token context window, and ships in Muse Code and the Meta Model API.

Is Muse Spark 1.3 better than Claude Opus 5?

It depends on the task. On Meta’s own chart, Muse Spark 1.3 beats Opus 5 on all three coding benchmarks and both long-context bands. Opus 5 beats it on four of the six agent benchmarks and scores higher on the Artificial Analysis Intelligence Index, 63 to 62.

How much does Muse Spark 1.3 cost?

Standard pricing is $1.25 per million input tokens, $0.15 per million cached input tokens and $4.25 per million output tokens, unchanged since Muse Spark 1.1. The Contributor tier costs $0.10 and $0.20 in exchange for permission to train Meta’s future models on your prompts and completions.

Can I use Muse Spark 1.3 in the Meta AI app?

Not yet. Meta’s announcement names only Muse Code and the Meta Model API, and mentions no consumer product. Coverage claiming a rollout to WhatsApp, Instagram and Facebook is repeating language from the original April 2026 Muse Spark launch, not this one.

Is Muse Spark 1.3 open source?

No. Muse Spark 1.3 is closed and available only through Meta’s API and Muse Code. Mark Zuckerberg has said an open-weights Muse Spark release is coming soon, most recently on September 2, 2026, without a date. Meta’s open-weight model today is Muse Glimmer, a 30 billion parameter Apache 2.0 model released on August 10, 2026.