Muse Spark 1.3, Meta’s newest model, scores 48.1 on the Artificial Analysis Intelligence Index, which puts it just ahead of GPT-6 Sol at 47.5 and well short of the top. Claude Opus 5.5 leads that index at 57.6, Claude Fable 5.1 follows at 53.4 and GPT-6 Astra at 52.7, while Google’s best, Gemini 3.8 Flash, sits at 40.9. The catch is where you can use each one. Meta launched 1.3 in its paid API and the Muse Code terminal agent, then put it behind Muse, its free-to-start personal agent, but it has never said which version the free Meta AI app runs. That split frames the whole Muse Spark vs ChatGPT vs Claude vs Gemini question. If your question is which free assistant to open rather than which model scores highest, that is the app-level comparison in Meta AI vs ChatGPT.

Raw benchmark scores still do not tell you which model to open when you need to write an email, debug code, or read a medical label. We compared Muse Spark vs ChatGPT, Claude, and Gemini across the tasks that actually matter, from coding and writing to reasoning and visual analysis. We also ran one identical prompt through all four, so you can see how they behave on the same job.

Update, September 3, 2026. This comparison was first published when Muse Spark launched, and the field has moved several times since. Meta shipped Muse Spark 1.1 on July 9, 2026 alongside its first paid API, then Muse Spark 1.2 and the Muse Code terminal agent on August 5, 2026. Meta then shipped Muse Spark 1.3 on September 2, 2026. Version 1.2 scores 57 on the Intelligence Index against 1.1’s 53, and 1.3 scores 61, though only the paid API and Muse Code run it. OpenAI put GPT-5.6 into broad release in three tiers, and since August 6, 2026 free ChatGPT defaults to GPT-5.6 Luna rather than GPT-5.5 Instant. Anthropic shipped Claude Opus 5 on July 24, 2026 and Claude Fable 5.1 on September 1, 2026, which now tops the Intelligence Index. Google’s newest Flash model is Gemini 3.8 Flash, shipped September 2, 2026 at the same $0.75 / $3.75 API rate as Gemini 3.7 Flash, and its newest Pro model is still Gemini 3.1 Pro. The Claude and GPT figures below were re-read on September 2, 2026; the Muse Spark and Gemini figures date from August 7. The hands-on test further down is kept as the dated record it is, with the models we actually ran.

Update, September 25, 2026. Every figure on this page was re-read today, because the boards it rested on changed shape. Artificial Analysis moved to Intelligence Index v4.3.2, which rescales every score (Muse Spark 1.1’s 53 is now 33.7), and it no longer publishes the separate Coding Index and Agentic Index this comparison used to quote, so those rows are gone and coding now leans on Terminal-Bench 4.0, which sits inside the index. Anthropic shipped Claude Opus 5.5 on September 22 and moved Opus 5 to its legacy models. OpenAI shipped GPT-6 Astra on September 3 and GPT-6 Sol and Luna on September 22, which halved the Sol rate. Meta made 1.3 with max reasoning generally available on September 4 and launched Muse, a personal agent that runs on Muse Spark, on September 8. The nutrition-label test further down is still the April record.

The Key Takeaways

  • Muse Spark 1.3 scores 48.1 on the Artificial Analysis Intelligence Index, just ahead of GPT-6 Sol at 47.5, and costs $1.25 / $4.25 per million tokens through the Meta Model API. Meta’s Muse agent, free to start, runs on it, but Meta has not said which version the free Meta AI app runs.
  • Claude Opus 5.5 leads the Intelligence Index at 57.6 at $4 / $20 per million tokens, ahead of Claude Fable 5.1 on 53.4 and GPT-6 Astra on 52.7 at $10 / $50. Claude Opus 5 is now a legacy model.
  • Coding is contested, not settled. On Terminal-Bench 4.0, as Artificial Analysis runs it, Opus 5.5 scores 59.6% and GPT-6 Astra 59.1%, and Muse Spark 1.3 sits at 33.3%.
  • There is no Gemini 3.5 Pro, 3.6 Pro or 3.8 Pro. Google’s newest Pro model is still Gemini 3.1 Pro, labelled preview, and its newest Flash model, Gemini 3.8 Flash, outscores it 40.9 to 29.7.
  • No single model wins everything. Matching the model to the task beats picking one and forcing every job through it.

Muse Spark vs ChatGPT vs Claude vs Gemini at a Glance

De l'éditeur

Tous les modèles d'IA dans une seule app

Fello AI réunit GPT-6, Claude 5, Gemini 3.8, Grok 4.7 et plus dans une seule app native pour Mac et iPhone.

Téléchargez maintenant !

Before breaking down individual categories, here is how the four sit against each other today. Each column is the highest-scoring model the vendor sells now. Every score comes from the same source, the Artificial Analysis Intelligence Index v4.3.2, so they are directly comparable; Terminal-Bench 4.0 is the agentic coding test inside that index, and prices are list API rates per million tokens. The cheaper models most people meet sit a step lower: GPT-6 Sol scores 47.5 at $2 / $10, and Gemini 3.1 Pro, Google’s Pro model, scores 29.7 at $2 / $12.

Muse Spark 1.3GPT-6 AstraClaude Opus 5.5Gemini 3.8 Flash
AA Intelligence Index v4.3.248.152.757.640.9
Terminal-Bench 4.0 (AA run)33.3%59.1%59.6%19.7%
AA cost per index task$1.60$3.26$5.98$1.24
Best forCheap API work, health, chartsCoding, Codex and ChatGPT WorkAgentic work, long documentsGoogle ecosystem, multimodal
Context window1M1.05M1M1M
API price (in / out per 1M)$1.25 / $4.25$10 / $50$4 / $20$0.75 / $3.75
Consumer priceFree, with usage limits$20/mo (Plus, in Work and Codex)$20/mo (Pro)$19.99/mo (Google AI Pro)
Mac appYes (Meta AI, Muse)YesYesYes
iOS appYes (Meta AI, Muse)YesYesYes

Sources: Artificial Analysis leaderboard, Intelligence Index v4.3.2, Meta’s Muse Spark 1.3 announcement, Anthropic pricing, OpenAI API pricing. Every score and price re-read on September 25, 2026. Scores are each model’s highest reasoning setting, which for Gemini 3.8 Flash is high. Its $0.75 / $3.75 rate runs through December 31, 2026 and rises to $1.50 / $7.50 after that.

Muse Spark vs ChatGPT: Coding and Software Development

Coding is where the gap between Muse Spark and its paid rivals is most visible, and it is also the category where the leaderboard has shifted most since April. On Terminal-Bench 4.0, the agentic coding test inside the Artificial Analysis index, Muse Spark 1.3 scores 33.3%. Claude Opus 5.5 leads at 59.6%, with GPT-6 Astra half a point behind at 59.1%.

Who Actually Leads Coding Right Now

Nobody leads it outright, which is the honest answer and a useful one. Half a point on one test is not a crown, and the two models get there differently. Artificial Analysis spends $3.26 per index task running Astra against $5.98 for Opus 5.5, so Astra reaches the same coding score for a little over half the cost, while Opus 5.5 is the stronger model overall, 57.6 against 52.7. If cost per job matters most, Astra edges it. If you hand a model a repository and walk away, Opus 5.5 is the better bet. One step down, GPT-6 Sol scores 43.9% for $1.06 a task.

Anthropic has not published a SWE-bench Verified score for Opus 5 or Opus 5.5, so treat any figure you see quoted for either with suspicion. Anthropic does publish a Terminal-Bench 4.0 figure for Opus 5.5, 66.4% on its own harness; the 59.6% above is Artificial Analysis’s independent run, and both are right on their own terms. For the detail behind the model itself, see our Claude Opus 5.5 breakdown.

Where That Leaves Muse Spark

Muse Spark has closed real ground since April, and 1.3 is the version that did it. Version 1.1 was Meta’s stated answer to the coding criticism, and on today’s index it scores 33.7; Muse Spark 1.3, released September 2, 2026, jumps to 48.1. On Meta’s own launch chart 1.3 posts 75.4 on DeepSWE v1.1 and 88.8 on Terminal-Bench 2.1, ahead of Claude Opus 5 on both, though those are vendor-reported numbers against a model Anthropic has since replaced. They come from the max reasoning setting, which Meta made generally available on September 4. Independent measurement is less generous: on Artificial Analysis’s Terminal-Bench 4.0 run, 1.3 scores 33.3%, a little over half what Opus 5.5 manages.

If coding is your primary use case, Muse Spark is still not a replacement for ChatGPT or Claude. It is now a credible, cheap second opinion rather than a distant one. You can reach both GPT and Claude through Fello AI on your Mac, which is useful if you switch between coding and non-coding tasks through the day.

Writing and Creative Tasks

Writing quality is harder to benchmark than coding because it depends on tone, style, and what you are trying to produce. In blind preference tests, Claude has consistently ranked as the most human-sounding AI writer, and that has not changed with the move to Sonnet 5 and Opus 5.5.

ChatGPT is the best all-rounder for writing. It handles emails, blog posts, social content and scripts reliably. It does not have Claude’s distinctive voice, but it rarely produces awkward output either. Know which model is doing the writing: ordinary chats on Free and Plus still answer on GPT-5.6, and the GPT-6 models reach Plus subscribers inside ChatGPT Work and Codex.

Muse Spark writes competently with a noticeable lean toward conversational, social-media-friendly tone. TechRadar described the original release as “ChatGPT built for the social internet,” and that is still a fair summary. For Instagram captions or casual posts that tone fits. For business reports or long-form content, Claude and GPT produce more polished results.

Gemini 3.1 Pro is solid for factual, research-heavy writing where accuracy matters more than voice. Its 1 million token context window lets you feed entire documents as reference material, though every model in this comparison now matches that window.

Reasoning and Problem-Solving

This is where Muse Spark’s Contemplating mode makes its strongest case. Instead of one model thinking for longer, the way GPT Pro or Gemini Deep Think do, Contemplating mode spins up multiple reasoning agents that work in parallel and synthesises their outputs. Meta’s argument is that thinking wider produces comparable or better answers at lower latency than thinking deeper.

At launch that approach lifted Muse Spark on Humanity’s Last Exam from behind the field to 50.2%, ahead of both GPT-5.4 Pro at 43.9% and Gemini Deep Think at 48.4%. The idea held up well enough that Meta kept it in 1.1.

The catch is abstract reasoning. On ARC-AGI-2, which tests novel pattern recognition rather than recall, Muse Spark scored 42.5 against scores above 76 for the paid flagships of the time. For structured, well-defined problems Contemplating mode competes with the best. For open-ended abstract challenges it still falls behind.

Health, Medical, and Vision Tasks

This is Muse Spark’s strongest category by a wide margin. It scored 42.8 on HealthBench Hard, beating GPT-5.4’s 40.1 and more than doubling Gemini 3.1 Pro’s 20.6. Meta has kept health and science a stated priority through the 1.1 release.

For visual understanding, the original Muse Spark scored 80.5% on MMMU-Pro and 86.4 on CharXiv Reasoning for chart and figure analysis, which put it at the top of the chart-understanding field when it launched. If your work involves reading scientific charts or interpreting visual information, it remains the best free option we have tested.

At launch, Gemini 3.1 Pro was the only rival here that came close on vision, scoring 82.4% on MMMU-Pro. Its medical performance was far weaker, which is why Muse Spark stays our pick for health-related work.

We Tested All Four Models on a Real Nutrition Label

Benchmarks do not tell you which model will read a label correctly and give you a useful answer. So we ran an identical prompt across all four models, using a photo of an instant ramen cup. It is a Vegan Society registered product at 436 kcal, 14g fat, 6.8g saturated fat, 69g carbs, 8.4g protein, 3.6g salt per 100g.

This test was run in April 2026, so the models named in it are the ones that were current then. We have left the result exactly as we recorded it rather than restaging it against newer versions, because the behaviours it exposed are the point.

The Prompt We Used

I’m sharing the nutrition label from a pack of instant ramen noodles. Read it carefully and answer:

  1. What are the three nutrition facts a health-conscious buyer should notice, and why do they matter?
  2. Ramen is often marketed as a cheap, filling meal. Based on this label, is it a reasonable everyday food or an occasional treat? Take a clear position.
  3. Who is this product actually a good fit for, and who should avoid it? Be specific.

Reference actual numbers from the label. No generic nutrition advice. No disclaimers about consulting a doctor. I want short output in bullets and table.

The Result

CriterionMuse SparkGPT-5.4Claude Opus 4.6Gemini 3.1 Pro
Stuck to 3 key factsListed all 7 firstYesYesSkipped protein
Specific cup-size math2.5-2.9g salt per 70-80g cupGeneric2.3-2.7g salt per typical cupGeneric
Caught “deep-fried” inferenceYesNoNoYes
Caught Vegan Society logoYesNoYesYes
Instruction adherencePartialGoodBestGood
Memorable framingNoNoYesNo

Winner: Claude. It kept to exactly three nutrients as asked, gave the sharpest math for a real cup size, and delivered the only memorable bottom line, “It’s a legitimate pantry item, not a legitimate staple. Treat it like frozen pizza, not like rice.” That is the kind of answer you remember the next time you are in a grocery aisle.

The surprise was that Muse Spark and Gemini both caught visual details that Claude and ChatGPT missed. Both noticed the noodles are deep-fried, an inference from 14g total fat with 6.8g saturated, and both spotted the Vegan Society logo on the packaging. That is visual chain-of-thought in action, and it matches Muse Spark’s chart-understanding scores.

The bigger surprise was how ChatGPT was the weakest performer on this specific test. It followed the format and took a clear position, but it missed the visual inferences and skipped the cup-size math that made Claude’s answer sharper.

The takeaway. For visual analysis and health reasoning, Muse Spark punches above its benchmark score. For sharp judgment and clean instruction-following, Claude wins. No single model reads a label perfectly, which is exactly why access to more than one matters.

Muse Spark vs ChatGPT: Pricing and Platform Access

The pricing picture changed materially in July 2026 and again in September. Muse Spark is still free to use, but Meta now also sells it, and every vendor in this comparison, Meta included, now ships a native Mac app.

Muse Spark 1.3GPT-6Claude Opus 5.5GeminiFello AI
Consumer priceFree, with usage limits$20/mo (Plus)$20/mo (Pro)$19.99/mo (Google AI Pro)$9.99/mo
API price (in / out per 1M)$1.25 / $4.25$2 / $10 (Sol), $10 / $50 (Astra)$4 / $20$0.75 / $3.75 (3.8 Flash), $2 / $12 (3.1 Pro)N/A
Free tierMeta AI app, version not namedGPT-5.6 Luna, plus GPT-6 Luna in the desktop appSonnet 5 (limited)Gemini 3 Flash-Lite, Flash and Pro, 32K contextFree model included
Mac desktop appYes (Meta AI, Muse)YesYesYesYes
iOS appYes (Meta AI, Muse)YesYesYesYes
Web accessmeta.ai, muse.aichatgpt.comclaude.aigemini.google.comN/A
APIMeta Model APIYesYesYesN/A

What Free Actually Gets You

Muse Spark is free in the Meta AI app and at meta.ai, up to a daily limit. Past that limit, Meta’s help centre points to paid Meta One Core and Premium plans, which are still in limited testing. When Meta launched Muse Spark 1.1 in July it made it free in Thinking mode on the same surfaces, but Meta has not said which version the app runs since. Muse, the personal agent Meta launched in the US on September 8, is free up to a usage limit too, and Meta describes the Muse Spark model behind it as its most capable to date.

What changed is the other side of it. Since July 9, 2026 Meta also sells the model through the Meta Model API, which now serves 1.3 at $1.25 input / $4.25 output per million tokens, the same rate 1.1 launched at with $20 in free credits. That is Meta’s first paid model, and it prices well under Claude Opus 5.5 and GPT-6 Astra, though GPT-6 Sol at $2 / $10 is no longer far off.

The other free tiers have moved too. Free ChatGPT answers ordinary chats on GPT-5.6 Luna, and since September 22 free users also get GPT-6 Luna in the desktop app. Free Claude runs Sonnet 5 with tight limits. Free Gemini gives you all three Gemini 3 models, Flash-Lite, Flash and Pro, but with a 32K-token context window rather than the 1M a paid plan unlocks. Because Meta does not name the model inside its free app, the free-tier comparison comes down to limits and features, not index scores.

Mac and Desktop Access

This matters if you work on a Mac. ChatGPT and Claude both ship native Mac apps with companion windows, keyboard shortcuts and system-wide access. Google shipped its Gemini Mac app on April 15, 2026, free, for macOS 15 and up, summoned with Option and Space. Meta has caught up on paper. It released a free Meta AI Mac app on August 19, 2026, still labelled beta, and a Mac version of its Muse agent in September, downloaded from Meta rather than the Mac App Store, which since September 23 can operate other Mac apps with your permission. The Meta AI app does not say which Muse Spark version it runs, so if you want to pick the model yourself, a third-party Muse Spark desktop client for Mac is the other route.

There is another way round that if you want several models from one place on your Mac. Fello AI is a native Mac, iPhone and iPad app that puts ChatGPT, Claude, Gemini, Grok and DeepSeek behind one subscription, along with Perplexity, Kimi, GLM, Qwen and Muse Spark. It starts at $9.99/month against $20 for ChatGPT Plus, $20 for Claude Pro and $19.99 for Google AI Pro, with a free tier to try first and 4.7 stars across 27,000+ reviews. You switch models inside a single conversation instead of paying for three of them separately, and if it is ChatGPT alone you want, our guide to the ChatGPT desktop client for Mac covers the native app.

Which AI Model Should You Use for What?

No single model wins everything. Here is the practical split based on the boards above and how these models actually behave. For the paid flagships on their own, without Muse Spark in the mix, see the three-model comparison.

Pick Muse Spark When

Open Muse Spark when you need a capable model and do not want to pay anything, and especially when the job is health-related or visual. It is the strongest free option we have tested for reading medical information, interpreting charts and figures, and pulling detail out of a photographed label. Its conversational tone suits social copy better than the paid flagships do.

Contemplating mode is the other reason to reach for it. On structured problems with several valid approaches, running reasoning agents in parallel gets closer to the paid models than the headline index score suggests. On open-ended abstract problems it does not, so keep the expectation matched to the task.

Pick ChatGPT (GPT-6) When

ChatGPT is the reliable all-rounder, and GPT-6 Astra is the coding pick when cost per job matters, within half a point of Opus 5.5 on Terminal-Bench 4.0 at a little over half the cost per task. It is also the most polished general-purpose experience of the four, with the deepest set of integrations and third-party tooling around it. Know which model you are getting, though. Ordinary chats on Plus still answer on GPT-5.6 Sol, Astra reaches Plus inside ChatGPT Work and Codex, and GPT-6 Pro, powered by Astra, is in Chat only on the Pro, Business and Enterprise plans.

The tiering got a lot cheaper on September 22, 2026, when OpenAI shipped GPT-6 Sol and Luna. On the API, GPT-6 Astra is the flagship at $10 / $50 per million tokens, GPT-6 Sol costs $2 / $10, half the $4 / $20 that GPT-5.6 Sol still sells at, and GPT-6 Luna runs $0.10 / $0.50 for high-volume work where speed matters more than depth.

Pick Claude (Opus 5.5 or Sonnet 5) When

Claude is the pick when you hand the model a long, multi-step job rather than a single question. Opus 5.5 tops the whole Artificial Analysis Intelligence Index at 57.6, agent tasks make up 30% of that index, and it also leads Terminal-Bench 4.0. It holds a 1M token context window for the long documents that come with that kind of work, and at $4 / $20 it costs less per token than Opus 5 did, which Anthropic now lists as a legacy model.

It is also still the best writer of the four if you care about voice rather than correctness alone. For desktop work, Claude Cowork and Computer Use on Mac both give Claude a way to act on your machine rather than just answer in a chat window.

Pick Gemini (3.8 Flash or 3.1 Pro) When

Gemini earns its place when your work is image, video or document heavy, or when you already live inside Google Workspace, Search and Drive. It is a strong pick for factual, research-heavy writing where accuracy matters more than voice, and its free tier gives you all three Gemini 3 models rather than a cut-down one, though with a 32K-token context window.

Be clear-eyed about where it sits on the boards, though. Google’s strongest model is not its Pro. Gemini 3.8 Flash scores 40.9 on the Intelligence Index and Gemini 3.1 Pro only 29.7, both well behind the other three here. Gemini wins on ecosystem, multimodal range and price, not on raw measured capability.

One correction worth stating plainly, because it circulates constantly. There is no Gemini 3.5 Pro, 3.6 Pro or 3.8 Pro. Google’s published model catalog lists Gemini 3.1 Pro as its newest Pro model, still labelled preview. Its newest stable Flash model is Gemini 3.8 Flash, shipped September 2, 2026. Our review of the Gemini 3.5 generation covers what Google actually shipped.

If you find yourself switching between two or three of these depending on the day, that is normal. Our best AI models ranking tracks which model leads in each category as things change.

The Bottom Line

Muse Spark has become a serious model. Version 1.3 scores 48.1 on the Intelligence Index, just ahead of GPT-6 Sol, for $1.25 / $4.25 through the API, and the Meta AI app still costs nothing up to its daily limit. Its health, medical and chart-reading work led the free field at launch, and Contemplating mode remains a real idea rather than a marketing one.

But a strong model is not the same as the best model, and the distance has held. Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7 are 9.5 and 4.6 points clear, and both come with mature APIs and ecosystems Muse Spark does not have. If you code, write professionally, or need deep Mac integration, the subscriptions still justify themselves. Our guides to the Claude desktop client for macOS and the Gemini desktop client for macOS show what running each one natively actually looks like.

The smartest approach is not choosing one model. It is having access to the right model for each task. Whether that means switching between free tiers or using one Mac app to reach all of them, the winners in 2026 are the people who match the tool to the job.

For a deeper breakdown of Muse Spark’s benchmarks and features, check our full explainer, and our guide to Muse Spark 1.2 and Muse Code covers what Meta shipped next. And if you want to see how Claude stacks up against ChatGPT or how ChatGPT compares to Gemini in more detail, we have dedicated comparisons for those matchups too.

FAQ

Is Muse Spark really free?

Yes, up to a daily limit, in the Meta AI app and at meta.ai. The original release made all three reasoning modes (Instant, Thinking, Contemplating), voice input and image analysis free with a Meta account, and Muse Spark 1.1 launched free in Thinking mode on the same surfaces; Meta has not said which version the app runs now. Past the free limit, Meta points to paid Meta One plans, still in limited testing. Since July 9, 2026 Meta also sells the model through the paid Meta Model API at $1.25 input and $4.25 output per million tokens, but that is for developers, not app users.

Can I use Muse Spark on Mac?

Yes. Meta released a free Meta AI Mac app on August 19, 2026, still labelled beta, and a Mac version of its Muse agent in September, and meta.ai works in any browser. ChatGPT, Claude and Gemini all ship Mac apps too, Google’s since April 15, 2026.

Is Muse Spark better than ChatGPT for coding?

No. On Terminal-Bench 4.0, as Artificial Analysis runs it, GPT-6 Astra scores 59.1% and Muse Spark 1.3 scores 33.3%. Version 1.3 closed much of the gap on the overall index, where it now edges GPT-6 Sol, but ChatGPT and Claude are both still well ahead on coding.

Which AI is best for coding in 2026?

It depends on the shape of the work, and no model sweeps it. On Terminal-Bench 4.0, Claude Opus 5.5 scores 59.6% and GPT-6 Astra 59.1%, a tie in practice. Opus 5.5 leads the full Artificial Analysis index, 57.6 to 52.7, while Astra gets its coding score for a little over half the cost per task. Long multi-step jobs favour Opus 5.5; cost-sensitive coding favours Astra.

What is Contemplating mode?

Contemplating mode runs multiple reasoning agents in parallel instead of one agent thinking for longer. At launch it scored 50.2% on Humanity’s Last Exam, ahead of both GPT-5.4 Pro and Gemini Deep Think. It is best for complex problems with several valid approaches.

Should I switch from ChatGPT to Muse Spark?

For coding or professional writing, no; ChatGPT and Claude still win. For health questions, chart analysis or casual chat without paying, yes. Meta now has Mac apps too, but if you rely on deep Mac desktop integration, ChatGPT and Claude are still the more mature choice.