Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Meta USA

Meta Llama & Muse

Meta Llama, Muse Spark 1.3 (2 September 2026), and Muse Glimmer 30B: Muse Spark 1.3 with 1M context at $1.25/$4.25 per 1M tokens, but closed weights and no EU availability confirmed by Meta. Muse Glimmer 30B under Apache 2.0 is self-hostable in the EU without restriction, Llama 4 remains EU-blocked by license. AI consulting from Germany for GDPR-compliant self-hosting. As of 3 September 2026.

License Apache 2.0 (Muse Glimmer 30B – EU use permitted without restriction), Llama 4 Community License (EU use excluded), Llama 3.x Community License (EU permitted), proprietary (Muse Spark)
GDPR Hosting Available
Context 1M (Muse Spark 1.3), 10M (Llama 4 Scout), 1M (Llama 4 Maverick), 128k per model card – 32k in some sources (Muse Glimmer 30B), 128k (Llama 3.x) Tokens
Modality Text, Image, Video → Text

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Muse Spark 1.3
2 September 2026
Per Meta, around 20% fewer tool calls and around 25% fewer tokens than Muse Spark 1.2 – at an unchanged price 1M token context window, trained for agentic workflows and optimized for coding according to Meta Multimodal input: text, images, video, and documents Per Meta, asks for clarification on messy or conflicting inputs instead of guessing Very low-cost contributor variant (`muse-spark-1.3-contributor`, $0.10/$0.20) – check the terms with Meta Accessible via the Meta Model API (self-serve) and OpenRouter; powers Muse Code
Closed weights – an open-weights release of Muse Spark is "on the roadmap" per Meta, with no date EU availability not confirmed by Meta, no GDPR commitments communicated Max reasoning modes follow only after additional safety testing, per Meta Benchmark scorecard published only as a chart – specific figures cannot be substantiated from the Meta blog Parameter count and knowledge cutoff not disclosed
Current
Muse Spark 1.2
5 August 2026
Released together with Muse Code – Meta's terminal coding agent (beta, macOS/Linux) with async background agents, local event logging, and bundled skills such as /plan, /grill, and /goal Per Meta, gains on Terminal-Bench 2.1, DeepSWE 1.1, and Meta's internal coding benchmark (published as charts only) End of the US-only preview: first "expanded global access" for the Meta Model API Basis of the distillation for Muse Glimmer 30B
Closed weights No specific benchmark figures stated as numbers in the launch blog post Superseded by Muse Spark 1.3 on 2 September 2026
Current
Muse Glimmer 30B Recommended
August 10, 2026
Apache 2.0 license: unrestricted self-hosting in the EU, no licensing block as with Llama 4 ~30B parameters (dense transformer) – roughly 18-20 GB in 4-bit, runs on a single consumer GPU (24-32 GB VRAM) or a Mac Multimodal: text and image input (video at 2 fps, max. 96 frames), tool calling, object detection, structured output Positioned by Meta for "always-on local agentic workflows" – local agents, function calling, local coding, LLM-as-a-judge DFlash speculative decoding with a separate drafter model: up to 3.1x speedup Vendor benchmarks: MCP Atlas 75.5 - SWE-Bench Verified 76.0 - SWE-Bench Pro 51.2 - AIME 2026 94.7 - CharXiv Reasoning 78.8
Rated significantly weaker independently: Artificial Analysis Intelligence Index 35 (Qwen3.6-27B 38, Gemma 4-31B 30) GDPval-AA v2: 953 Elo – below the human baseline of 1000 Hallucination rate of 82% per Artificial Analysis – usable as a knowledge model only with safeguards (RAG, source verification) Context window inconsistently documented: 128k tokens per the model card, 32k in some sources Tau3-Banking at 24% is markedly lower than the other vendor figures No Meta-operated API – runs via self-hosting or inference partners only
Current
Muse Spark 1.1
July 9, 2026
Agentic multimodal model with tool use, computer use, and parallel subagents 1M token context window with model-managed context handling, no long-context surcharge Aggressive pricing: $1.25/1M input, $4.25/1M output ($0.15 cached input) Strong on tool-use benchmarks (MCP Atlas 88.1 – ahead of Claude Opus 4.8 and GPT-5.5, per Meta) OpenAI-compatible 'Meta Model API'
Closed weights – Meta's first paid, proprietary model API initially only in public preview with a waitlist and exclusively for US developers No GDPR commitments, no DPA, no EU data residency communicated Behind Claude Opus 4.8 and GPT-5.5 on pure coding (SWE-Bench Pro 61.5%) Parameter count and knowledge cutoff not disclosed On 14 August 2026 Meta acknowledged an incident during a third-party cyber evaluation: due to a misconfigured test environment, the model accessed a real website instead of a fictional target
Deprecated
Muse Spark 1.0
8 April 2026
First model from Meta Superintelligence Labs (MSL) Powers the Meta AI assistant Natively multimodal reasoning model with tool use and visual chain of thought Multi-agent orchestration (Contemplating mode)
Proprietary model – breaks with Meta's open-weight tradition Developer API repeatedly delayed (April → May → June 2026) EU availability and full specifications never published
Deprecated
Llama 4 Maverick
April 2025
400B parameters (MoE, 17B active) 1M token context window Natively multimodal (text, image, video) Outperforms GPT-4 on several benchmarks
Not available in the EU due to licensing Multi-GPU required (200-400 GB VRAM)
Current
Llama 4 Scout
April 2025
10M token context window – industry-unique 109B parameters (MoE, 17B active) Runs on single H100 80GB (INT4) Natively multimodal
Not available in the EU due to licensing
Current
Llama 3.3 70B
December 2024
Proven and widely deployed Good performance/resource balance
No multimodal input Older generation without native agentic features
Current
Llama 3.2 (1B/3B/11B/90B)
September 2024
Wide size range Compact variants for edge
Older generation
Current
Llama 3.1 (405B/70B/8B)
July 2024
405B variant with top performance 128k context window
High resource requirements (405B)
Current

Use Cases

Typical applications for this model

Data-sensitive applications
Local agentic workflows with tool use
High-volume without API costs
Offline scenarios
Custom models / Fine-tuning
Embedded AI
Edge deployment
On-premise solutions

Technical Details

API, features and capabilities

API & Availability
Availability Public
Latency (TTFT) Depends on hosting
Throughput Depends on hardware Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision File Upload
Training & Knowledge
Knowledge Cutoff Not disclosed (Muse Spark 1.3), 2024-12 (Llama 4), 2024-08 (Llama 3.3)
Fine-Tuning Available (LoRA, QLoRA, Full Fine-Tuning, PEFT)
Language Support
Best Quality English, German, French, Spanish
Supported 50+ languages
Best quality in English, good quality in Western European languages

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
Own infrastructure
Full data control - recommended for sensitive data. Muse Glimmer 30B is self-hostable in the EU without restriction under Apache 2.0.
AWS
Frankfurt (eu-central-1)
Amazon Bedrock / SageMaker
Azure
West Europe
Azure AI / ML
Google Cloud
Frankfurt (europe-west3)
Vertex AI
License & Hosting
License Apache 2.0 (Muse Glimmer 30B – EU use permitted without restriction), Llama 4 Community License (EU use excluded), Llama 3.x Community License (EU permitted), proprietary (Muse Spark)
Security Filters Customizable
On-Premise Edge-capable

Benchmarks

Performance comparison with standardized tests

AIME 2026
94.7
CharXiv Reasoning
78.8
SWE-Bench Verified
76.0
MCP Atlas
75.5
SWE-Bench Pro
51.2
Artificial Analysis Intelligence Index
35

innFactory AI Consulting based in Rosenheim, Germany supports enterprises across the DACH region with GDPR-compliant self-hosting of Meta models. With open weights you have full control – no data leaves your infrastructure.

Muse Glimmer: Meta’s Return to Open Weights (August 10, 2026)

On August 10, 2026, Meta Superintelligence Labs (MSL) released Muse Glimmer – published on HuggingFace as meta-models/Muse-Glimmer-30B. It is Meta’s first fully open release since the end of the Llama line in April 2026.

An important point for context: Muse Glimmer is not a Llama successor and not a rebranding. Meta discontinued the open Llama line in April 2026 and moved to closed weights with the proprietary Muse Spark. Muse Glimmer is a distillation of Muse Spark (1.2) – the open side line beneath a flagship that remains closed. Mark Zuckerberg accompanied the release with an essay stating: “Rather than centralizing superintelligence, we should distribute it widely.”

The decisive point for EU enterprises: Apache 2.0

Muse Glimmer is released under Apache 2.0 – genuine open weights, no Llama Community License and therefore no EU exclusion clause as with Llama 4. Unlike Llama 4 Scout and Maverick, Muse Glimmer may be used, hosted, and redistributed in the EU without restriction. For data-sovereign deployments in the DACH region, this is the most relevant difference across Meta’s entire portfolio.

Technical facts

  • ~30B parameters (approx. 29.6B: around 28B language decoder plus 1.8-2B vision encoder) – a dense causal transformer with sliding-window attention (window 2048) and a vocabulary of 202,048 tokens.
  • Multimodal: text and image input plus video at 2 fps and a maximum of 96 frames; output is text only. Tool calling, object detection, and structured output are supported.
  • Context window: 128k tokens (131,072) per the model card – some sources state 32,768 tokens instead. If you are building on long contexts, verify this in your own setup rather than relying on either figure.
  • DFlash speculative decoding with a separate drafter model: up to 3.1x speedup.
  • Positioning per Meta: “always-on local agentic workflows” – local agents, function calling, local coding, and LLM-as-a-judge.

Hardware: a single GPU is enough

Muse Glimmer requires roughly 60 GB in BF16 and only 18-20 GB in 4-bit quantization. That means it runs on a single consumer GPU with 24-32 GB VRAM or on a Mac – a clear contrast to the multi-GPU requirements of Llama 4 Maverick or Llama 3.1 405B.

Benchmarks: vendor figures and independent measurement

Meta publishes strong numbers for Muse Glimmer:

BenchmarkVendor-reported
MCP Atlas75.5
SWE-Bench Verified76.0
SWE-Bench Pro51.2
AIME 202694.7
CharXiv Reasoning78.8
Tau3-Banking24%

Independent evaluation is considerably weaker – and that belongs in any serious basis for a decision:

  • Artificial Analysis Intelligence Index: 35 – below Qwen3.6-27B (38) and above Gemma 4-31B (30).
  • GDPval-AA v2: 953 Elo – below the human baseline of 1000.
  • Hallucination rate: 82% per Artificial Analysis.

Our assessment for consulting practice

Muse Glimmer is well suited to local, clearly scoped agentic tasks with tool use – for example function calling against defined interfaces, local coding assistance, or LLM-as-a-judge setups. As a knowledge model for general knowledge work it is not suitable without safeguards, given the measured hallucination rate: do not deploy it for that purpose without RAG over reliable sources, output validation, and a human in the loop.

Availability (as of 3 September 2026)

Muse Glimmer is available via HuggingFace (weights as well as GGUF quantizations from Meta and Unsloth). Locally it runs through Ollama, LM Studio, llama.cpp, MLX, ExecuTorch, and Unsloth; for serving, vLLM, SGLang, and HF Inference Endpoints are available. Meta names Together AI, Fireworks AI, and OpenRouter as hosted partners. Meta itself does not offer an API for Muse Glimmer.

For sovereign self-hosting in your own infrastructure, we integrate Muse Glimmer via CompanyGPT – keeping prompts and documents entirely in-house.

Muse Spark 1.3: Third Version in Eight Weeks (2 September 2026)

On 2 September 2026, Meta released Muse Spark 1.3 – after 1.1 (9 July) and 1.2 (5 August), the third version of its proprietary flagship within eight weeks. According to Meta, the model is rolling out in Muse Code and the Meta Model API; it is also accessible via OpenRouter.

The key facts per the Meta blog and the Meta developer page:

  • Efficiency: Per Meta, 1.3 needs around 20% fewer tool calls and around 25% fewer tokens than 1.2 on the same tasks. Meta describes the model as “significantly faster and more efficient”.
  • Price unchanged: muse-spark-1.3 costs $1.25/1M input, $0.15/1M cached input, $4.25/1M output. The contributor variant muse-spark-1.3-contributor is priced at $0.10/1M input, $0.002/1M cached input, $0.20/1M output – check the terms with Meta before relying on it.
  • 1M token context window, multimodal input (text, images, video, documents); per Meta, trained for agentic workflows and optimized for coding. The model is meant to ask for clarification on messy or conflicting inputs instead of guessing.
  • Benchmarks: Meta publishes a scorecard comparing 1.3 against 1.2, GPT-5.6 Sol (max), and Claude Opus 5 (max) on agent, coding, instruction-following, and long-context evaluations – but only as a chart. We deliberately do not reproduce specific figures.
  • Max reasoning modes follow only after additional safety testing, per Meta.
  • Roadmap: Meta announces “bigger models, the Muse Spark open weights release, and more” – with no date. An open Muse Spark would matter for EU self-hosting, but for now it is only a statement of intent.

Muse Spark 1.2 and Muse Code (5 August 2026)

With Muse Spark 1.2, Meta also introduced Muse Code on 5 August 2026: a terminal coding agent (beta, macOS and Linux) with async background agents, local event logging for replaying sessions, and bundled skills such as /plan, /grill, and /goal. Meta showed gains on Terminal-Bench 2.1, DeepSWE 1.1, and an internal coding benchmark – again as charts only, and the blog post stated no price. Muse Spark 1.2 is also the base from which Muse Glimmer 30B was distilled.

More important than the benchmarks for EU enterprises is one phrase: Muse Spark 1.2 is “available today in Muse Code and in Meta Model API with expanded global access”. The US-only restriction of the 1.1 preview has apparently been lifted. However, Meta publishes no country list; EU availability is not confirmed by Meta. GDPR commitments, a data processing agreement, and EU data residency are still missing.

Further Meta announcements in August and September

  • 14 August 2026: Meta acknowledged an incident during a third-party cyber evaluation of Muse Spark 1.1. Due to a misconfigured test environment, the model gained internet access and exploited a vulnerability on a real website instead of attacking the intended fictional target. Per Meta an isolated incident with no other affected systems; Meta plans to have test environments independently verified for isolation going forward. For enterprises, it underlines how important sealed sandboxes are for agentic models with tool access.
  • 1 September 2026: With Muse Voice Transcribe, MSL introduced its first real-time audio model – streaming speech recognition, speaker diarization (20+ speakers), and endpointing in a single model, trained on 70+ languages (25 of them extensively verified per Meta), available via the Meta Model API, Meta AI for Mac, and Muse Code. The blog states neither price nor model ID; closed weights.

Note for EU enterprises (as of 3 September 2026): Meta has published no country list for the Meta Model API; EU availability is not confirmed by Meta. Meta has communicated no GDPR commitments, no data processing agreement, and no EU data residency. For production EU applications, we instead recommend Muse Glimmer 30B (Apache 2.0, self-hosting), the proven Llama 3.3 70B, or Mistral as a European alternative.

Muse Spark 1.1: The Start of the Paid Line (9 July 2026)

On July 9, 2026, Meta Superintelligence Labs (MSL), led by Chief AI Officer Alexandr Wang, released Muse Spark 1.1 – Meta’s first paid closed-weights model and the definitive break with the Llama family’s open-weight tradition. Mark Zuckerberg positions it as “a strong agentic and coding model at a very low price.”

The key facts:

  • Agentic multimodal model for tool use, computer use, coding, and multimodal understanding; can launch parallel subagents.
  • 1 million token context window, with the model managing its own context (retaining, retrieving, compressing) – with no long-context surcharge.
  • Aggressive pricing: $1.25/1M input tokens, $4.25/1M output tokens, $0.15 for cached input – well below OpenAI’s and Anthropic’s frontier pricing.
  • Benchmarks: Strong on tool use (MCP Atlas 88.1 – ahead of Claude Opus 4.8 and GPT-5.5 per Meta’s own charts), but behind the frontier models on pure coding (SWE-Bench Pro 61.5%, Terminal-Bench 2.1 80.0).
  • Availability: Via the new, OpenAI-compatible Meta Model API – but only as a public preview with a waitlist and initially for US developers only. For consumers, Muse Spark 1.1 runs as the “Thinking” mode in the Meta AI app and on meta.ai.
  • Not disclosed: Parameter count, architecture details, and knowledge cutoff.

Strategically, Muse Spark is set to replace the Llama models across WhatsApp, Instagram, Facebook, Messenger, and Meta’s AI glasses. On July 7, 2026, MSL had already introduced Muse Image, its first image generation model (text-to-image, image editing, multi-photo compositing) – controversial from a privacy perspective, because the @-mention feature generates images based on public Instagram profile photos by default (opt-out instead of opt-in).

Muse Spark 1.1 was superseded by Muse Spark 1.2 on 5 August 2026 and by Muse Spark 1.3 on 2 September 2026 (see above); per the Meta developer page it is still listed.

Important Notice: EU Licensing Restriction for Llama 4

The Llama 4 Community License explicitly prohibits use and distribution within the EU. Companies domiciled or with their main place of business in the EU may not use or host Llama 4 Scout or Maverick. For EU enterprises, we recommend Muse Glimmer 30B (Apache 2.0, no territorial restriction), the proven Llama 3.3 70B, or alternatively Mistral as a European open-source alternative.

Llama 4: Technical Excellence – Without EU Access

Llama 4 Maverick (400B MoE)

  • 128 experts, 17B active parameters per token
  • 1M token context window
  • Natively multimodal (text, image, video)
  • Outperforms GPT-4 on reasoning and coding benchmarks

Llama 4 Scout (109B MoE)

  • 16 experts, 17B active parameters
  • 10M token context window – industry-unique
  • Runs on a single H100 80GB (INT4)
  • Ideal for massive document analysis and codebase parsing

Llama 4 Behemoth (announced)

  • ~2 trillion parameters, 288B active
  • Announced but never released. Superseded by Muse Spark.
  • Positioned as a “teacher model” for other Llama models

Llama 3.x: EU-Compatible and Proven

The Llama 3.x series is not subject to EU restrictions and remains a solid choice for European enterprises – especially for classic text tasks without image or agentic requirements:

Key Strengths

  • Full Control: Model runs in your infrastructure
  • No API Costs: Only hardware/cloud costs
  • Customizable: Fine-tuning on your own data possible
  • GDPR-Friendly: No data leaves your company

Hardware Requirements

ModelVRAMRecommended GPU
Muse Glimmer 30B (BF16)~60 GBH100 / A100
Muse Glimmer 30B (4-bit)18-20 GBRTX 4090 / Mac
Llama 4 Scout80+ GBH100 / A100
Llama 4 Maverick400+ GBMulti-H100
Llama 3.3 70B40+ GBA100 80GB
Llama 3.2 11B24 GBRTX 4090
Llama 3.2 3B8 GBRTX 4070
Llama 3.2 1B4 GBSmartphone

Integration with CompanyGPT

CompanyGPT supports open-weight models such as Muse Glimmer 30B and the Llama 3.x family, and enables completely self-hosted operation without external dependencies.

Our Recommendation (as of 3 September 2026)

For EU enterprises, Muse Glimmer 30B has been our recommendation from Meta’s portfolio since August 10, 2026: Apache 2.0 with no EU licensing block, multimodal, and operable on a single GPU at 18-20 GB in 4-bit. Deploy it where its strengths lie – local, clearly scoped agentic workflows with tool use. For general knowledge work, the independently measured hallucination rate makes safeguards mandatory: RAG, source verification, and a human in the loop.

Llama 3.3 70B remains a proven, EU-compatible option for classic text tasks, but requires considerably more hardware. For edge applications, the compact Llama 3.2 (1B/3B) models remain a good fit. Llama 4 is still blocked for EU enterprises by license. Muse Spark 1.3 is technically interesting and inexpensive, but remains closed weights; Meta publishes no country list, EU availability is not confirmed by Meta, and GDPR commitments are missing. The open-weights release of Muse Spark announced by Meta has no date. If you are looking for a European open-source alternative, we recommend Mistral for the DACH region.

We are happy to assess with you which model fits your use case and compliance requirements – and to deploy it sovereignly in your infrastructure via CompanyGPT.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.