Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Meta USA

Meta Llama

Meta Llama and Muse Spark – from open-source LLMs to the proprietary Muse Spark 1.1 (July 9, 2026, Meta's first paid closed-weights model). Note: Llama 4 is not available in the EU due to licensing, Muse Spark is US-developers only. AI consulting from Germany for Llama 3.x deployment.

License Llama 4 Community License (EU use excluded), Llama 3.x Community License (EU permitted)
GDPR Hosting Available
Context 10M (Llama 4 Scout), 1M (Llama 4 Maverick), 128k (Llama 3.x) Tokens
Modality Text, Image, Video → Text

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Muse Spark 1.1
July 9, 2026
Agentic multimodal model with tool use, computer use, and parallel subagents 1M token context window with model-managed context handling, no long-context surcharge Aggressive pricing: $1.25/1M input, $4.25/1M output ($0.15 cached input) Strong on tool-use benchmarks (MCP Atlas 88.1 – ahead of Claude Opus 4.8 and GPT-5.5, per Meta) OpenAI-compatible 'Meta Model API'
Closed weights – Meta's first paid, proprietary model API only in public preview with waitlist, initially US developers only – no EU availability No GDPR commitments, no DPA, no EU data residency communicated Behind Claude Opus 4.8 and GPT-5.5 on pure coding (SWE-Bench Pro 61.5%) Parameter count and knowledge cutoff not disclosed
Current
Muse Spark 1.0
8 April 2026
First model from Meta Superintelligence Labs (MSL) Powers the Meta AI assistant Natively multimodal reasoning model with tool use and visual chain of thought Multi-agent orchestration (Contemplating mode)
Proprietary model – breaks with Meta's open-weight tradition Developer API repeatedly delayed (April → May → June 2026) EU availability and full specifications never published
Deprecated
Llama 4 Maverick
April 2025
400B parameters (MoE, 17B active) 1M token context window Natively multimodal (text, image, video) Outperforms GPT-4 on several benchmarks
Not available in the EU due to licensing Multi-GPU required (200-400 GB VRAM)
Current
Llama 4 Scout
April 2025
10M token context window – industry-unique 109B parameters (MoE, 17B active) Runs on single H100 80GB (INT4) Natively multimodal
Not available in the EU due to licensing
Current
Llama 3.3 70B Recommended
December 2024
Proven and widely deployed Good performance/resource balance
No multimodal input
Current
Llama 3.2 (1B/3B/11B/90B)
September 2024
Wide size range Compact variants for edge
Older generation
Current
Llama 3.1 (405B/70B/8B)
July 2024
405B variant with top performance 128k context window
High resource requirements (405B)
Current

Use Cases

Typical applications for this model

Data-sensitive applications
High-volume without API costs
Offline scenarios
Custom models / Fine-tuning
Embedded AI
Edge deployment
On-premise solutions

Technical Details

API, features and capabilities

API & Availability
Availability Public
Latency (TTFT) Depends on hosting
Throughput Depends on hardware Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision File Upload
Training & Knowledge
Knowledge Cutoff Not disclosed (Muse Spark 1.1), 2024-12 (Llama 4), 2024-08 (Llama 3.3)
Fine-Tuning Available (LoRA, QLoRA, Full Fine-Tuning, PEFT)
Language Support
Best Quality English, German, French, Spanish
Supported 50+ languages
Best quality in English, good quality in Western European languages

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
Own infrastructure
Full data control - recommended for sensitive data
AWS
Frankfurt (eu-central-1)
Amazon Bedrock / SageMaker
Azure
West Europe
Azure AI / ML
Google Cloud
Frankfurt (europe-west3)
Vertex AI
License & Hosting
License Llama 4 Community License (EU use excluded), Llama 3.x Community License (EU permitted)
Security Filters Customizable
On-Premise Edge-capable

innFactory AI Consulting based in Rosenheim, Germany supports enterprises across the DACH region with GDPR-compliant self-hosting of Meta Llama. With open weights you have full control – no data leaves your infrastructure.

Muse Spark 1.1: Meta’s First Paid Model (July 9, 2026)

On July 9, 2026, Meta Superintelligence Labs (MSL), led by Chief AI Officer Alexandr Wang, released Muse Spark 1.1 – Meta’s first paid closed-weights model and the definitive break with the Llama family’s open-weight tradition. Mark Zuckerberg positions it as “a strong agentic and coding model at a very low price.”

The key facts:

  • Agentic multimodal model for tool use, computer use, coding, and multimodal understanding; can launch parallel subagents.
  • 1 million token context window, with the model managing its own context (retaining, retrieving, compressing) – with no long-context surcharge.
  • Aggressive pricing: $1.25/1M input tokens, $4.25/1M output tokens, $0.15 for cached input – well below OpenAI’s and Anthropic’s frontier pricing.
  • Benchmarks: Strong on tool use (MCP Atlas 88.1 – ahead of Claude Opus 4.8 and GPT-5.5 per Meta’s own charts), but behind the frontier models on pure coding (SWE-Bench Pro 61.5%, Terminal-Bench 2.1 80.0).
  • Availability: Via the new, OpenAI-compatible Meta Model API – but only as a public preview with a waitlist and initially for US developers only. For consumers, Muse Spark 1.1 runs as the “Thinking” mode in the Meta AI app and on meta.ai.
  • Not disclosed: Parameter count, architecture details, and knowledge cutoff.

Strategically, Muse Spark is set to replace the Llama models across WhatsApp, Instagram, Facebook, Messenger, and Meta’s AI glasses. On July 7, 2026, MSL had already introduced Muse Image, its first image generation model (text-to-image, image editing, multi-photo compositing) – controversial from a privacy perspective, because the @-mention feature generates images based on public Instagram profile photos by default (opt-out instead of opt-in).

Note for EU enterprises: The Meta Model API is not available in the EU (preview restricted to US developers). Meta has communicated no GDPR commitments, no data processing agreement, and no EU data residency; there is no EU rollout timeline. For production EU applications, we continue to recommend Llama 3.3 70B (self-hosting) or Mistral as a European alternative.

Important Notice: EU Licensing Restriction for Llama 4

The Llama 4 Community License explicitly prohibits use and distribution within the EU. Companies domiciled or with their main place of business in the EU may not use or host Llama 4 Scout or Maverick. For EU enterprises, we recommend Llama 3.3 70B or alternatively Mistral as a European open-source alternative.

Llama 4: Technical Excellence – Without EU Access

Llama 4 Maverick (400B MoE)

  • 128 experts, 17B active parameters per token
  • 1M token context window
  • Natively multimodal (text, image, video)
  • Outperforms GPT-4 on reasoning and coding benchmarks

Llama 4 Scout (109B MoE)

  • 16 experts, 17B active parameters
  • 10M token context window – industry-unique
  • Runs on a single H100 80GB (INT4)
  • Ideal for massive document analysis and codebase parsing

Llama 4 Behemoth (announced)

  • ~2 trillion parameters, 288B active
  • Announced but never released. Superseded by Muse Spark.
  • Positioned as a “teacher model” for other Llama models

Llama 3.x: Recommended for EU Enterprises

The Llama 3.x series is not subject to EU restrictions and remains the recommended choice for European enterprises:

Key Strengths

  • Full Control: Model runs in your infrastructure
  • No API Costs: Only hardware/cloud costs
  • Customizable: Fine-tuning on your own data possible
  • GDPR-Friendly: No data leaves your company

Hardware Requirements

ModelVRAMRecommended GPU
Llama 4 Scout80+ GBH100 / A100
Llama 4 Maverick400+ GBMulti-H100
Llama 3.3 70B40+ GBA100 80GB
Llama 3.2 11B24 GBRTX 4090
Llama 3.2 3B8 GBRTX 4070
Llama 3.2 1B4 GBSmartphone

Integration with CompanyGPT

CompanyGPT supports Llama 3.x models and enables completely self-hosted operation without external dependencies.

Our Recommendation

For EU enterprises, Llama 3.3 70B is the best choice from the Llama family – proven, EU-compatible, and with a strong performance profile. For edge applications, the compact Llama 3.2 (1B/3B) models are ideal. If you’re looking for a powerful European open-source alternative, we recommend Mistral as a Llama 4 replacement for the DACH region.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Consultation for this model?

We help you select and integrate the right AI model for your use case.