AI & Machine Learning
·By Seedwire Editorial·

Muse Spark 1.3 Ships Without the Mode Behind Its Best Scores

Based on reporting by venturebeat.com. Analysis and framing are Seedwire's own.

Meta released Muse Spark 1.3 this week and pitched it as a frontier model, but the configuration that produced its most impressive benchmark numbers is not the one developers can use. According to a report first reported by venturebeat.com, the model's strongest results come from a "max" reasoning setting that Meta says is still undergoing additional safety testing and will arrive "shortly." What is actually rolling out through the Muse Code harness and the Meta Model API uses Meta's existing reasoning settings, including the one it calls xhigh.

Mark Zuckerberg announced the release on X with the line that Muse Spark 1.3 offers "frontier performance almost too cheap to meter," and called it Meta's "biggest jump" yet in coding and agentic work. VentureBeat's reporting supports both halves of that claim in part, while adding the caveats that matter for anyone planning to deploy the model.

Two configurations, one set of launch slides

Meta does publish scores for both versions in its evaluation report, so the gap is disclosed rather than concealed. But the launch materials lean on the max numbers. VentureBeat lists the differences: on GDPval-AA v2, max scores 1,754 Elo against 1,709 for xhigh. On OSWorld 2.0, the gap is 66.9 versus 57.2. On JobBench, it is 64.9 versus 61.2. Not every test favors max. DeepSearchQA is tied at 89.4, and on Terminal-Bench 2.1 the shipping xhigh version actually edges ahead, 89.2 to 88.8.

The independent picture from Artificial Analysis is similar. The firm evaluated max in a limited partner preview and currently lists no API provider for it. It scores max at 62 on its Intelligence Index and xhigh at 61. That 61 puts the shipping version level with GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. Anthropic still holds the top spots, with Claude Fable 5.1 at 66 in max mode and 65 in xhigh, and Claude Opus 5 at 63 in both modes. Muse Spark 1.3 as shipped is in the frontier cluster, in other words, but it is not the model defining that frontier.

That is still a real move from where Meta stood a month ago. VentureBeat's coverage of Muse Spark 1.2 described a credible coding challenger that generally trailed Anthropic's best. Version 1.2 posted 82.9 percent on Terminal-Bench 2.1 against 86.7 percent for Opus 5, and lost the other main coding comparisons Meta chose to present. With 1.3, the report says Meta is trading wins with OpenAI and Anthropic on several coding and agentic evaluations rather than simply appearing in the comparison.

Meta also claims the model is easier to run in practice. It says Muse Spark 1.3 is trained to hold multiple workflows in a single long thread, pull context with tools, notice gaps in its own plans, ask for clarification when it needs to, and confirm before taking consequential actions. In internal comparisons by Meta engineers, it used roughly 20 percent fewer tool calls and 25 percent fewer tokens than 1.2 on coding work. For a company running agent loops at scale, that kind of behavioral change can matter more than a single leaderboard point. Related: Meta's AI Mode: A New Frontier in Social Media Intelligence.

What "too cheap to meter" does and does not mean

Muse Spark 1.3 did not get a price cut. Standard pricing is unchanged from 1.2 at $1.25 per million input tokens, $4.25 per million output tokens and $0.15 per million cached input tokens. That combined $5.50 per million sits in the middle of the pricing table VentureBeat compiled. Meta's own Contributor tier is far cheaper at $0.10 input and $0.20 output, and models such as GPT-5.6 Luna at $1.40 combined and MiMo-V2.5 Flash at $0.40 undercut the Standard rate. Above it, Claude Opus 5 lists at $5.00 input and $25.00 output, and Claude Fable 5.1 at $10.00 and $50.00. Related: Zuckerberg's AI Concerns.

So Zuckerberg's phrase is really a claim about work per dollar rather than dollars per token. Artificial Analysis gives that argument some backing. It measures Muse Spark 1.3 xhigh at 235.2 output tokens per second and estimates a cost of $0.55 per Intelligence Index task, which it says is the lowest cost per task of any currently measured model at a score of 61.

There is a wrinkle, though. Muse Spark 1.2 cost only $0.40 per Artificial Analysis task while scoring 57. With per-token pricing flat, the cost of finishing an average task went up between generations. Artificial Analysis attributes the increase mainly to heavier input-token consumption on agentic evaluations. VentureBeat notes this does not directly contradict Meta's 25 percent token-reduction figure, because Meta is describing its own coding workflows while Artificial Analysis is measuring a broader suite. Both things can be true at once: the model may be leaner on the tasks Meta optimized for and hungrier on the tasks an outside firm chose. Related: Gemini 3.8 Flash: Google's Third Budget AI Model in 6 Weeks.

The question buyers should actually ask

The more interesting question for enterprises is not whether Muse Spark 1.3 can reach the frontier, but how close the deployable version gets and at what all-in cost. On that framing, the shipping model looks like a strong value: level with several rivals' top modes on the Intelligence Index, with a per-task cost the benchmarking firm ranks best in its class. The premium max configuration exists, but for now it is a promise attached to a safety review rather than a product with a price.

That distinction suggests a few things to watch. One is whether max arrives at the same $1.25 and $4.25 rates or carries its own pricing, since a heavier reasoning mode that consumes more tokens could shift the cost-per-task math again. Another is whether the cost increase Artificial Analysis observed between 1.2 and 1.3 holds up as more independent evaluators measure the model. And the third is the softer claim about operating behavior. Fewer tool calls and clarifying questions before consequential actions are exactly what teams running unattended agents want, but Meta's numbers here are internal, and outside confirmation would go a long way.

For now, the honest read is that Meta has closed most of the gap with the leaders on the model people can actually call today, kept its prices where they were, and put its most eye-catching numbers on a version that is still in the lab. That is a stronger position than Muse Spark 1.2 occupied a month ago. It is also a reminder that "frontier performance" in a launch post and "frontier performance" in a production API are not always the same configuration.

Meta
Muse Spark 1.3
Artificial Analysis
AI model pricing
agentic coding
Mark Zuckerberg
frontier models
Seedwire Newsletter

Stay ahead of the curve

Get the most important tech stories delivered to your inbox. No spam, unsubscribe anytime.