Meta's Muse Spark 1.3 surpasses Google's Gemini in hours, signaling intensifying AI competition
Meta released Muse Spark 1.3 this week, claiming its strongest performance gains yet in coding and agent tasks, while also formally launching its Muse Code agent with new subscription tiers. The new model briefly dethroned Google's freshly released Gemini 3.8 Flash from the cost-efficiency frontier.

Meta has had a significant week advancing its artificial intelligence offerings, moving its Muse Code coding agent from beta status and introducing three subscription tiers. The company then revealed Muse Spark 1.3 on Wednesday evening, its latest reasoning model iteration, with executives asserting it represents the company's most substantial improvements in coding capabilities and agent-based operations.
Accessible via Muse Code and the Meta Model API, Muse Spark 1.3 represents the newest generation of the reasoning model Meta debuted in April. The release comes within weeks of Muse Spark 1.2's introduction. CEO Mark Zuckerberg promoted the model on social media, characterizing it as delivering "frontier performance almost too cheap to meter."
This is the biggest jump we've made so far on coding and agentic work.
Mark Zuckerberg
Open-weight release 'coming soon'
Zuckerberg announced that open-weight versions of Muse Spark would arrive "coming soon," potentially enabling developers to obtain and execute the model on their own systems, circumventing Meta's API infrastructure.
The extent of this openness remains uncertain, as Meta has not yet disclosed the licensing framework accompanying the release. Previous open-weight offerings from Meta have included different usage limitations, meaning these specifications will ultimately dictate the degree to which Spark can be altered, shared or deployed.
Zuckerberg additionally hinted at Meta's anticipated next model, internally referred to as Watermelon, using a watermelon emoji. This system is expected to be substantially larger than Spark and was reportedly still undergoing training as of July, per a Business Insider report from that period. No official timeline has been announced for Watermelon's availability, though Zuckerberg's messaging suggests it may not be distant.
https://x.com/finkd/status/2095232032896946311?ref_src=twsrc%5Etfw
Meta is making substantial performance assertions regarding Spark 1.3. Zuckerberg shared a benchmark comparison table contrasting the model against OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus 5 across coding, agentic, computer-use and extended-context assessments.
Notable benchmark outcomes include Muse Spark 1.3 achieving 75.4% on the DeepSWE coding benchmark and 98.1% on the 512K-1M variant of the MRCR long-context test.
These figures originate from Meta's own testing methodology. Meta states that Spark 1.3 scores came from the Meta Model API, while competing model data derives from a combination of Meta's internal assessments, public leaderboards and results shared by competing providers. Meta characterizes its evaluation of third-party models as "best-effort," indicating the table should not be interpreted as an independent, controlled comparison under uniform conditions.

An important qualification applies to these outcomes: Meta evaluated Spark 1.3 using its new "max" reasoning configuration for primary comparisons, whereas the maximum reasoning setting currently accessible to most developers is "xhigh." The max setting remains in restricted preview while Meta completes additional safety evaluations.
"Gloves are off": How Spark 1.3 stacks up
Independent analysis from Artificial Analysis offers an external assessment. The San Francisco-based benchmarking firm rates the publicly available Muse Spark 1.3 xhigh at 61 on its Intelligence Index, surpassing Spark 1.2 by four points and matching GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. Testing the restricted-preview max version yielded a score of 62, positioning it behind only Claude Fable 5.1 and Claude Opus 5 among compared models at launch.
Alex Volkov, AI evangelist at infrastructure provider CoreWeave, references the benchmark as evidence of Meta's rapid convergence with leading competitors.

Damn, gloves are off!
Alex Volkov
Volkov characterizes the performance as "quite the statement" from Meta while forecasting "busy weeks ahead of us!"
Pricing represents another dimension of Meta's competitive positioning. Artificial Analysis identifies Spark 1.3 xhigh as the "most cost-efficient model" at its measured intelligence tier, with its 61 Intelligence Index score combined with relatively modest per-task expenses placing it on the company's Pareto frontier.

According to Artificial Analysis calculations, Spark 1.3 xhigh costs approximately $0.55 per Intelligence Index task, the lowest among models scoring 59 or above on its index. GPT-5.6 Sol max and Grok 4.6 high, both tied with Spark at 61, cost $0.95 and $0.94 respectively.
Spark 1.3 carries higher per-task expenses than Spark 1.2's $0.40, with Artificial Analysis attributing this rise primarily to the newer iteration consuming roughly 57% additional input tokens during agentic testing.

Muse meets Gemini: A 'playground slapfight'
Shortly before Meta's Muse Spark 1.3 announcement, Google introduced Gemini 3.8 Flash, its third Flash iteration in six weeks. Google characterized its model as its most advanced reasoning and coding Flash variant to date, while maintaining introductory pricing at $0.75 per million input tokens and $3.75 per million output tokens.
Artificial Analysis initially positioned Gemini 3.8 Flash (high) on its Intelligence-versus-Cost Pareto frontier—the collection of models for which no alternative exists that is simultaneously more capable and less expensive. Gemini achieved 59 on the Intelligence Index at $0.58 per task.
Within hours, however, Muse Spark 1.3 xhigh arrived with 61 and $0.55 per task, surpassing Gemini 3.8 Flash on both dimensions and displacing it from that frontier.
This development did not escape notice within the AI community. AI researcher Benjamin Marie observed on X that "Google was ahead only a few hours."
Google was ahead only a few hours.
Benjamin Marie
https://x.com/bnjmn_marie/status/2095251751737971190?ref_src=twsrc%5Etfw
Florian Brand, research engineer at Prime Intellect, captured the reversal concisely on X: "Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours."
Meta's chief AI officer Alexandr Wang himself participated in the commentary, leveraging the Artificial Analysis report to criticize Google.
Corey Quinn, co-founder and chief cloud economist at Duckbill, quickly dismissed the exchange, contending that neither Meta nor Google are substantially driving the broader AI competition. "Meta casting shade at Google in AI is a playground slapfight outside a MMA championship," Quinn posted on X.
https://x.com/alexandr_wang/status/2095249704888197175?ref_src=twsrc%5Etfw
Muse Code enters the scene
Muse Spark 1.3 represents the fourth iteration Meta has delivered since April, following Muse Spark 1.1 in July and 1.2 in August. The more strategically significant element of Meta's initiative, however, involves the agent framework surrounding the model.
This framework materializes through Muse Code, Meta's terminal-based coding agent, which debuted in beta alongside Spark 1.2 in early August. The agent officially launched Tuesday with subscription plans beginning at $5 monthly and a new SDK entering developer preview.
While Meta is clearly pursuing competition with leading research organizations on model performance, as evidenced by Muse Spark's progression, the company is also competing at a higher tier, where Anthropic's Claude Code and OpenAI's Codex have established benchmarks for how developers interact with agents in practice.
This context explains why the 1.3 release messaging emphasizes coding and agentic capabilities specifically. The model constitutes a central component of a broader developer platform Meta is promoting with emphasis on cost-effectiveness alongside performance.