Back to glossary

Term

Muse Spark

Muse Spark is Metas closed-weight flagship from April 2026 — the first model of the Muse line out of Meta Superintelligence Labs, focused on efficient reasoning via thought compression.

Muse Spark — explained in more detail

Muse Spark is the first model of Metas Muse line and the very first model to come out of Meta Superintelligence Labs, founded the year before under AI chief Alexandr Wang. It was unveiled on April 8, 2026. It follows the disappointingly received Llama 4 and marks a double strategy shift: away from the open Llama heritage toward closed weights, and away from a pure text model toward multimodal perception, reasoning and agentic tasks. Meta ran the model internally under the code name Avocado. What is known about Muse Spark so far comes from the vendor itself and from reporting — broad independent testing did not exist at the time of writing.

Key facts

  • Release: April 8, 2026, the first model from Meta Superintelligence Labs; internal code name Avocado.
  • Access: closed-weight, at first only in the Meta AI app, on meta.ai and in a private API preview; API access for third-party developers is planned per Meta and intended as a revenue stream of its own.
  • Focus: natively multimodal (text, image, speech), reasoning and agentic tasks; per Meta small and fast by design, yet capable enough for complex questions in science, math and health.
  • Thought compression: during RL training a second signal penalizes the length of the chain of thought — the model learns to reach a solution with far fewer reasoning tokens. Meta states this needs roughly an order of magnitude less compute than Llama 4 Maverick at comparable accuracy.
  • Efficiency per third party: Artificial Analysis reports that Muse Spark completed its Intelligence Index run with around 58M output tokens — less than half the footprint of comparably placed models. A pricing model for broad use was not published at launch.

Example / Practical use

The value lies less in a single answer than in the interplay of efficiency and agency. Meta describes a “Contemplating” mode in which the model distributes sub-tasks to sub-agents running in parallel. For users, the thought-compression approach mainly means the same tasks at a lower token and therefore cost footprint — provided the vendor figures hold up in practice.

Delimitation

Muse Spark is not another Llama version but the break with it: Meta gives up the open weights that had bound a large developer community since 2023. That moves the model closer to the practice of proprietary providers such as OpenAI or Google. Comparisons with their flagships should stay cautious until independent benchmarks exist; what is solid so far is mainly the compute-efficiency claim, not a consistent quality lead.

See everything in one place:Muse Spark