On this page
The short answer
Muse Spark 1.3 is Meta’s current hosted candidate for agentic and coding workflows. Evaluate the exact tier, reasoning setting, and input modality. Its mathematics results provide useful evidence, but the Standard versus Contributor data-use distinction and the audio limitation may matter more to your deployment decision.
What to take away
- Standard and Contributor have different training-use terms.
- Do not mistake a math benchmark for a coding success rate.
- Audio understanding has a documented limitation in version 1.3.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
A dated model snapshot. Published metrics describe the named model, settings, and test. They are not Rubrex experiments or a universal ranking. Prices and availability can change after the research date.
Compare models and check the data date →Which Meta model this article evaluates
Meta’s current API documentation recommends Muse Spark 1.3 for new agentic and coding work. It lists text, image, video, audio, and PDF inputs, but explicitly warns that audio understanding in version 1.3 is not fully supported and quality may be degraded. This is a hosted model briefing, not a claim that every current Meta model has downloadable weights. Choose the exact product and deployment surface before drawing conclusions from the Meta brand name.
Evidence: Meta Model API: models [1]
Choose the tier before comparing the price
Meta documents different data-use terms for Standard and Contributor variants: Standard does not use data for training, while Contributor grants training permission for prompts and completions in exchange for a lower price. That distinction matters to an AI SaaS handling customer information. Verify the actual endpoint and terms before submitting data. The catalog row below identifies its own price source and should not be treated as a quote for every direct Meta tier or an assurance about application logging.
Evidence: Meta Model API: models [1]OpenRouter public model catalog [2]
| Model or property | Published value | How to interpret it |
|---|---|---|
| meta/muse-spark-1.3 | $1.25 input / $4.25 output | OpenRouter catalog entry, not a Contributor-tier quote |
| Context window | 1,048,576 tokens | Meta documented capacity |
| Audio input | Not fully supported in 1.3 | Provider warns of degraded quality |
Interpret the mathematics evidence
The saved Epoch AI FrontierMath Tiers 1-3 v2 records include both xhigh and max configurations. The benchmark uses expert mathematics problems with Python access; it does not directly measure repository maintenance or customer support. The two point estimates are close relative to their standard errors. Do not infer that xhigh is inherently better than max from this table. Preserve both rows and their dates rather than selecting the higher number as if it were a universal model score.
Evidence: Epoch AI: FrontierMath Tiers 1–3 v2 [3]
| Benchmark | Exact configuration | Score | Standard error | Run date |
|---|---|---|---|---|
| FrontierMath Tiers 1–3 v2 | Muse Spark 1.3 (xhigh) | 74.39% | 2.59 percentage points | 2026-09-16 |
| FrontierMath Tiers 1–3 v2 | Muse Spark 1.3 (max) | 74.04% | 2.60 percentage points | 2026-09-18 |
Test the workflow you plan to ship
For an image-assisted engineering tool, test whether the model can identify the relevant visual detail, select the right tool, and verify the resulting change. Include misleading screenshots, incomplete context, and tools that return partial results. Our recommendation is to measure accepted completion, unnecessary actions, unauthorized-action attempts, and complete task cost. Keep audio out of a claimed capability until your own test supports the use case and current provider guidance permits it. A successful text-only test says little about a multimodal workflow.
Limits of the evidence
These are provider specifications and independent Epoch AI results, not Rubrex test outcomes. OpenRouter prices can differ from direct-provider tiers. A no-training term does not establish zero retention or secure behavior across the rest of your SaaS application.
Common questions
Is Muse Spark 1.3 the same as an open-weight Llama model?
No. This briefing covers the hosted Muse Spark API model. Verify the license and deployment model separately for any downloadable Meta model.
Which tier should handle customer data?
Review the required data terms and exact account configuration. Do not select Contributor merely for price when its training permission conflicts with your commitments.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Meta Model API: models Meta · Living reference · Maintained documentationReviewed: Muse Spark version, modalities, context, and Standard versus Contributor tiers. Accessed September 25, 2026.
- OpenRouter public model catalog OpenRouter · Living reference · Maintained documentationReviewed: catalog fields; numerical prices taken from the saved September 25 API snapshot. Accessed September 25, 2026.
- Epoch AI: FrontierMath Tiers 1–3 v2 Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.