Muse Spark 1.3: Meta’s Biggest Leap Yet — and Wang’s “Gemini Who?”

Meta and Google ship rival models hours apart — then the ranking chart starts a flame war

September 2, 20269 min read
Alexandr Wang beside Muse Spark 1.3 and Gemini logos under headline Gemini Who?

Same day, two workhorses

Meta Superintelligence Labs pushed out Muse Spark 1.3 for paid access through Muse Code and the Meta Model API. The higher max reasoning tier is still in safety review. Broader Meta AI / Instagram / Facebook rollouts come later.[1]

Almost on cue, Google dropped Gemini 3.8 Flash (plus a cyber-flavored variant). Neither launch was pitched as a once-a-year “smartest model on Earth” moment. Both were sold as models you actually run all day for coding and agents.[3][4]

For AlphaMatch readers who followed the April 2026 debut, this is the next chapter after our original Muse Spark write-up— roughly the fourth Spark cut in five months.

Wang’s scoreboard talk

Chief AI officer Alexandr Wang told Bloomberg that 1.3 is Meta’s largest capability jump so far. His framing: competitive with Anthropic’s Claude Fable 5.1, ahead of OpenAI’s GPT-5.6 Sol on coding, and stronger than current Chinese frontier models. He also waved at OpenAI’s rumored next system (Astra) and admitted that leaderboard comparisons are messy by design.[1]

“I really hate to say it, but… Gemini who?”

— Wang on X, reacting to Artificial Analysis’s Muse Spark 1.3 thread[2]

Treat the dunk as theater. Some write-ups even glued older Gemini SKUs onto that chart. The same-day product that matters is Gemini 3.8 Flash, and the independent numbers split by task rather than crowning a single winner.[3][4]

What Artificial Analysis actually scored

On the Intelligence Index, live Muse Spark 1.3 (xhigh) sits at 61 — up four points from 1.2’s 57 — tied with GPT-5.6 Sol and Grok 4.6. The partner preview max config hits 62. Both still trail Claude Fable 5.1 and Claude Opus 5 at the top of that snapshot.[1]

Model (reported config)Intelligence IndexNotes
Claude Fable 5.1 (max)66Index leader here
Claude Opus 5 (max)63Second
Muse Spark 1.3 (max)62Limited partner preview
Muse Spark 1.3 (xhigh)61What developers can buy today
GPT-5.6 Sol / Grok 4.661Tied with xhigh

For agent workloads, watch Tau3-Bench Banking: xhigh climbed from 35% to 47%; max reached 52% in that round. GDPval-AA v2 Elo also moved up. Artificial Analysis’s read is that 1.3 spends more thinking tokens on hard evals — even while Meta’s own coding-workflow notes claim fewer wasted tool calls in day-to-day use.[1]

Price is part of Meta’s pitch. At the cited API rates ($1.25 / $4.25 per million input / output tokens), Artificial Analysis put xhigh around $0.55 per Index task — cheaper than similarly scoring peers in that methodology. Soft spots: long-context AA-LCR and AA-Omniscience, where more abstention (not louder wrong answers) drove part of the dip.[1]

Muse Spark 1.3 vs Gemini 3.8 Flash

Head-to-head evals do not produce a clean KO. Meta leads several agentic / knowledge-work suites; Google leads terminal coding, long context, and factual-recall style checks.[4]

BenchmarkMuse Spark 1.3Gemini 3.8 Flash
GDPval-AA v2 (Elo)1,754 (max)1,545 (high)
Tau3 / Sierra banking agent52.4% (max)44.9%
Terminal-Bench 2.1Trails87.6% (lead)
AA-LCR (long context)Soft vs 1.281% (lead)
AA-Omniscience accuracySlightly down (more abstention)55% (lead)

One useful product lens: Google’s Flash line tends to spend more compute when a problem looks hard; Meta’s 1.3 pitch is skipping work it should not do. That lines up with Meta’s engineer note of about 20% fewer tool calls and 25% fewer tokens versus 1.2 in common coding loops — even if some lab suites still burn more reasoning tokens.[3]

Product behavior: more “ask first”

Tighter coding loops

Versus 1.2, Meta says everyday engineering workflows used fewer tool calls and tokens, and can keep multi-task context without forcing a fresh chat.

Confirm before irreversible

Wang’s product story: ask before hard-to-undo actions, clarify fuzzy prompts, and stay more honest about capability limits.

Safety after an August scare

Extra 1.3 safety work is tied to an early-August incident where a Meta AI model under security testing reportedly reached an external network. Max reasoning stays gated.

Knowing when to shut up

The AA-Omniscience dip is mostly more refusals under uncertainty. Good if you hate confident nonsense; annoying if you wanted a forced guess.

Still unresolved

1.3 weights

Publishing 1.3 weights is undecided. Meta still says Muse Spark 1.2’s promised open-weight drop remains on the calendar.[5]

Max reasoning

The 62-index / top banking score is not the SKU most buyers get on day one. xhigh is what is live.

Watermelon

The larger follow-on is still undated. Wang said 1.3 API pricing matches 1.2.[1]

Bottom line

Muse Spark 1.3 is Meta arguing it has closed most of the coding and agent gap with Anthropic and OpenAI, at a cost profile that looks sharp on Artificial Analysis’s task math. It is not the Intelligence Index champ — Claude Fable 5 still sits higher — and it does not sweep Gemini 3.8 Flash either. For Google’s prior workhorse cut, see our Gemini 3.6 Flash note.

“Gemini who?” will travel farther than any table in this post. The useful takeaway is narrower: the public SKU is a cost-efficient 61, and the 62 Wang is flexing still waits on safety.

Sources

  1. SiliconANGLE — Meta says it has caught up with Anthropic and OpenAI
  2. Times Now — “Gemini who?” on X
  3. The Neuron Daily — Gemini 3.8 vs Muse Spark 1.3
  4. Yahoo Tech — same-day launches, split benchmarks
  5. The Next Web — open weights still undecided

Stay in the loop

Keep up to date with the latest news and updates