Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM Microsoft USA

Microsoft Phi & MAI

Microsoft MAI family and Phi: MAI-Transcribe-2 (new on 3 September 2026, 60 languages, USD 0.10 per hour, EU region North Europe), MAI-Thinking-1, MAI-Code-1-Flash, MAI-Voice-2 with German voices, MAI-Image-2.5 plus Phi-4-reasoning-vision and Phi-4 for edge and local deployments. GDPR assessment per model (Foundry regions, Global Standard vs. regional). As of 3 September 2026. AI consulting from Germany.

License MIT
GDPR Hosting Available
Context 256k (MAI-Thinking-1), 16k (Phi-4), up to 128k (Phi-3 variants) Tokens
Modality Text, Image, Audio → Text, Image, Audio

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
MAI-Transcribe-2 Recommended
3 September 2026 (public preview)
60 languages including German (MAI-Transcribe-1.5: 43), multilingual mode with automatic language identification by default FLEURS benchmark: per Microsoft rank 1 across 60 languages at an average 5.2% word error rate; rank 2 on the Artificial Analysis WER leaderboard Per Microsoft up to 10x faster than OpenAI GPT-Transcribe, 7x faster than ElevenLabs Scribe v2 and 5x faster than Gemini 3.5 Transcribe Speaker diarization, word-level timestamps, keyword biasing (phraseList), code switching (e.g. Hinglish, Spanglish), noise robustness Selectable transcription style: verbatim (with fillers, for compliance and QA) or clean (for minutes and captions) Also usable as input transcription in the Voice Live API
Public preview without SLA – per Microsoft not recommended for production workloads In the EU only North Europe; no Germany West Central, West Europe or Sweden Central for this model File transcription up to 300 MB (WAV, MP3, FLAC) via the REST API only; parameters currently documented only for the REST path Introductory price limited until the end of 2026
Preview
MAI-Transcribe-1.5
2 June 2026 (Build 2026)
43 languages including German, per Microsoft rank 1 on FLEURS at launch Content biasing and improved accuracy over MAI-Transcribe-1
Superseded by MAI-Transcribe-2 (3 September 2026): fewer languages, higher price MAI-Transcribe-1 has been deprecated since 20 August 2026
Current
MAI-Voice-2 / MAI-Voice-2-Flash
2 June 2026 (Build 2026, public preview)
15 languages and 18 locales, including German with the voices de-DE-Klaus and de-DE-Mia Fine-grained emotion and style control via SSML (mstts:express-as), e.g. happy, empathy, whispering MAI-Voice-2-Flash for latency-critical voice agents, IVR and call centers; MAI-Voice-2 for long-form content with speaker consistency Instant voice cloning from 5 to 60 seconds of reference audio – only after limited-access approval and with consent
Public preview without SLA No Germany West Central – EU operation via West Europe, Sweden Central or France Central Voice cloning gated (Custom Neural Voice limited access)
Preview
MAI-Image-2.5 / Flash / Pro
2 June 2026 (Pro: 19 June 2026), preview
Text-to-image and image-to-image with precise, local edits (per Microsoft 'control with preservation') Per Microsoft rank 2 on Arena.ai for image editing (Build 2026) Flash variant for speed, Pro variant for photo-realistic output; integrated in PowerPoint and OneDrive Up to 1,048,576 total pixels (e.g. 1024x1024 or 768x1365), PNG output
Preview; prompts currently English only (languages: en) Global Standard only – no EU data zone commitment
Preview
MAI-Thinking-1 Recommended
2 June 2026 (public preview)
First fully in-house Microsoft reasoning model (no OpenAI distillation) Sparse Mixture-of-Experts with 35B active parameters 256k context window 97.0% AIME 2025, 94.5% AIME 2026 Matches Claude Opus 4.6 on SWE-Bench Pro Trained on commercially licensed data OpenAI-compatible chat completions API (endpoint /mai/v1), function calling, encrypted chain-of-thought to carry across turns
Public preview without SLA – limited enterprise track record Global Standard only (no PTU, no EU data zone); output capped at 64k tokens Not open source
Preview
MAI-Code-1-Flash Recommended
2 June 2026
5B parameter coding model – fast and cost-efficient Beats Claude Haiku 4.5 on all four tested coding benchmarks 16-point lead on SWE-Bench Pro (51.2% vs. 35.2%) Integrated directly in GitHub Copilot (Free, Pro, Pro+, Max)
Specialized for coding – not a general-purpose model Not open source
Current
Phi-4-reasoning-vision-15B Recommended
4 March 2026
15B parameters with vision and reasoning Decides autonomously when deep reasoning is needed 84.8 AI2D, 83.3 ChartQA, 75.2 MathVista, 88.2 ScreenSpot v2 MIT license, on Microsoft Foundry, HuggingFace and GitHub
Higher resource requirements than Phi-4 16k context window
Current
Phi-4-reasoning (14B)
April 2025
Pure reasoning model (no vision) MIT license
No vision support
Current
Phi-4-mini
February 2025
3.8B parameters – edge-capable MIT license
Lower capacity
Current
Phi-4-multimodal
February 2025
Multimodal inputs (text, image, audio) MIT license
Current
Phi-4
December 2024
14B parameters - very efficient Strong reasoning capabilities
Smaller than frontier models
Current
Phi-3.5-MoE
2024
Mixture-of-Experts 42B parameters (6.6B active)
Current
Phi-3.5-mini
2024
3.8B parameters Runs on smartphones
Limited capacity
Current
Phi-3.5-vision
2024
Multimodal Image understanding
Current
Phi-3-medium
2024
14B parameters Balance of size and performance
Current

Use Cases

Typical applications for this model

Edge AI & IoT
Mobile Applications
Offline Scenarios
Embedded Systems
Resource-Constrained Environments
Coding Assistants
Local LLM Deployments

Technical Details

API, features and capabilities

API & Availability
Availability Public
Latency (TTFT) ~100ms (local)
Throughput Hardware-dependent Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision Reasoning Mode File Upload
Training & Knowledge
Knowledge Cutoff February 2026 (Phi-4-reasoning-vision), 2024-10 (Phi-4)
Fine-Tuning Available (LoRA, QLoRA, Full Fine-Tuning)
Language Support
Best Quality English, German
Supported 20+ languages
Optimized for English, usable quality in German

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-Hosted
Own Infrastructure
Recommended - full data control
Azure
West Europe / Germany
Azure AI Model Catalog
Ollama
Local
Easy local deployment
License & Hosting
License MIT
Security Filters Customizable
Enterprise Support Yes
SLA Available Yes
On-Premise Edge-capable

Benchmarks

Performance comparison with standardized tests

MMLU
84.8%
HumanEval
82.6%
GSM8K
91.1%
MATH
55.5%

As AI consultants based in Rosenheim, Germany, we recommend Microsoft Phi for enterprises in the DACH region (Germany, Austria, Switzerland) that want to run powerful AI on resource-constrained hardware. Since the release of Phi-4 in December 2024, the Phi models have established themselves as reliable solutions for edge and on-premise deployments, offering impressive performance with minimal requirements. With the MAI family (Build 2026), Microsoft now extends its in-house model portfolio with fully proprietary reasoning and coding models.

New: MAI-Transcribe-2 – Speech-to-Text in 60 Languages (3 September 2026)

On 3 September 2026 Microsoft AI introduced MAI-Transcribe-2 and made it available as a public preview in Microsoft Foundry, the MAI Playground and via OpenRouter. The model transcribes 60 languages – including German – and per Microsoft ranks first on the FLEURS benchmark at an average 5.2% word error rate; on the independent Artificial Analysis WER leaderboard it ranks second according to Microsoft. On speed, Microsoft cites up to 10x faster than OpenAI GPT-Transcribe, 7x faster than ElevenLabs Scribe v2 and 5x faster than Gemini 3.5 Transcribe. The price is USD 0.10 per hour of audio as a limited-time offer until the end of 2026 (MAI-Transcribe-1: USD 0.36 per hour) – diarization, word timestamps and keyword biasing are included.

Features: speaker diarization, word-level timestamps, keyword biasing via a phrase list, automatic language identification, code switching within an utterance, noise robustness and a selectable transcription style (verbatim with fillers for compliance and QA, clean for minutes and captions). The model is called via the fast transcription API of Azure Speech (enhancedMode.model = "MAI-Transcribe-2", files up to 300 MB in WAV, MP3 or FLAC) or as input transcription in the Voice Live API. MAI-Transcribe-1.5 (43 languages, Build 2026) remains available, MAI-Transcribe-1 has been deprecated since 20 August 2026.

EU assessment: Per the Microsoft Learn region list the MAI-Transcribe model is available in eastus, northeurope, southeastasia and westus – in the EU therefore only via North Europe (Ireland). Azure Speech processes and stores data exclusively in the region of the Speech resource, so a Speech resource in North Europe confines processing to the EU. A deployment in Germany West Central is currently not possible for this model. Because it is a preview without SLA, we recommend MAI-Transcribe-2 for evaluations and non-critical workloads for now; production use should wait for the GA announcement.

The MAI Family at a Glance (as of 3 September 2026)

ModelTypeStatusEU regions per Microsoft LearnNote
MAI-Transcribe-2Speech-to-textPreview, 3 Sep 2026North Europe60 languages, USD 0.10/h until end of 2026
MAI-Transcribe-1.5Speech-to-textsince 2 June 2026North Europe43 languages
MAI-Voice-2 / -2-FlashText-to-speechPreview, 2 June 2026France Central, Sweden Central, West Europe15 languages, German voices Klaus and Mia
MAI-Image-2.5 / -Flash / -ProImage generationPreview, 2/19 June 2026Sweden Central, West Europe (Global Standard)English prompts only
MAI-Thinking-1Reasoning LLMPreview, 2 June 2026all EU regions (Global Standard)256k context, 64k output
MAI-Code-1-FlashCoding modelsince 2 June 2026GitHub Copilot, Foundry, OpenRouter5B parameters

For the GDPR assessment the deployment type is decisive: the speech models (Transcribe, Voice) process regionally in the chosen Speech region, whereas the Foundry models MAI-Thinking-1 and MAI-Image-2.5 run only as Global Standard – inference may therefore take place in any Azure region, and Microsoft does not yet offer an EU data zone for the MAI models.

MAI-Thinking-1 and MAI-Code-1-Flash (June 2026)

At Build 2026 (June 2, 2026), Microsoft unveiled its first fully in-house foundation models – trained without OpenAI distillation and exclusively on commercially licensed data. From a DACH perspective, two are especially relevant:

  • MAI-Thinking-1: Sparse Mixture-of-Experts with 35B active parameters and a 256k context window (64k output). Per Microsoft it scores 97.0% on AIME 2025, 94.5% on AIME 2026 and 52.8% on SWE-Bench Pro. In Microsoft Foundry as a public preview via an OpenAI-compatible chat completions API (/mai/v1) with function calling and encrypted chain-of-thought – exclusively as Global Standard (no PTU), per the region list deployable from all European Foundry regions including Germany West Central, but without an EU data zone.
  • MAI-Code-1-Flash: 5B-parameter coding model that per Microsoft outperforms Claude Haiku 4.5 on all four tested coding benchmarks (51.2% vs. 35.2% on SWE-Bench Pro). Rolling out in GitHub Copilot (Free, Pro, Pro+, Max) for VS Code from June 2, 2026, additionally via OpenRouter, Fireworks and Baseten.

Both models are part of the MAI family (see table above) and significantly reduce Microsoft’s dependence on OpenAI. They are, however, not open source – for full data control, Phi remains the better choice.

New: Phi-4-reasoning-vision (March 2026)

Microsoft has released Phi-4-reasoning-vision-15B, a new model that combines vision and reasoning in a 15B-parameter model. The model can autonomously decide when deeper reasoning is required — a key capability for efficient local deployments.

  • 15B Parameters: More capacity than Phi-4 (14B)
  • Vision + Reasoning: Understands images and can draw complex conclusions
  • MIT License: Full commercial use permitted
  • Self-Hosting: Runs on RTX 4090 or M2 Mac

Small Language Models (SLMs)

Microsoft Phi is Microsoft’s family of “Small Language Models” - compact but powerful models specifically optimized for efficiency.

Why Phi for Enterprises?

  • Compact: 3.8B to 14B parameters
  • Efficient: Runs on consumer hardware
  • MIT License: Full commercial use permitted
  • Self-Hosting: Full data control
  • Edge-Ready: Smartphones, IoT, embedded

Key Strengths

Efficiency

Phi models achieve impressive performance at minimal size:

ModelParametersComparable Performance to
Phi-414BGPT-4 (partially)
Phi-3.5-MoE42B (16 active)Llama 3 70B
Phi-3.5-mini3.8BLlama 3 8B

Local Deployment

Phi models can be operated completely locally:

  • Ollama: ollama run phi4
  • LM Studio: Simple GUI
  • vLLM: Production deployment
  • ONNX: Optimized inference

Hardware Requirements

ModelRAM/VRAMRecommended Hardware
Phi-416 GBRTX 4070 / M2 Mac
Phi-3.5-MoE24 GBRTX 4090
Phi-3.5-mini4 GBLaptop / Smartphone
Phi-3.5-vision8 GBRTX 3060

Reasoning Strength

Phi-4 shows particularly strong reasoning capabilities:

  • Mathematics and logic
  • Coding tasks
  • Structured analysis
  • Chain-of-thought

Comparison to Other SLMs

FeaturePhi-4Llama 3.2 3BGemma 2 2B
Parameters14B3B2B
ReasoningStrongMediumMedium
CodingStrongGoodGood
VisionYes (3.5)NoNo
LicenseMITCommunityApache 2.0

Integration with CompanyGPT

Microsoft Phi can be integrated in CompanyGPT as a self-hosted option - ideal for enterprises that want to operate AI without cloud dependency.

Our Recommendation

For transcription in the Microsoft stack, MAI-Transcribe-2 has been the first choice for evaluations since 3 September 2026: 60 languages including German, diarization and timestamps included in the base price of USD 0.10 per hour, EU processing via a Speech resource in North Europe. Because the model is still a public preview without SLA, we plan it for production workloads only once it reaches GA; until then GPT-Realtime-Whisper and Whisper or ElevenLabs Scribe remain the production alternatives. For German-language speech output, MAI-Voice-2 (voices Klaus and Mia) from West Europe or Sweden Central is worth testing.

For reasoning-heavy cloud workloads on Azure / Foundry, MAI-Thinking-1 has been our pick in the Microsoft stack since June 2026 – with the caveat that it runs only as Global Standard and therefore offers no EU processing commitment. For coding assistants in GitHub Copilot, MAI-Code-1-Flash offers a very good price-performance ratio.

For local and edge deployments with vision and reasoning requirements, we still recommend Microsoft Phi-4-reasoning-vision-15B. For pure text tasks, Phi-4 remains an excellent choice. For smartphones and IoT, Phi-3.5-mini is ideal, for simpler multimodal applications Phi-3.5-vision.

For applications requiring maximum quality where cloud hosting is acceptable, we recommend OpenAI GPT or Anthropic Claude instead.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is MAI-Transcribe-2 and what does it cost?

MAI-Transcribe-2 is the speech-to-text model from Microsoft AI introduced on 3 September 2026, available as a public preview via Azure Speech in Microsoft Foundry, in the MAI Playground and via OpenRouter. It transcribes 60 languages including German with speaker diarization, word timestamps, keyword biasing and code switching. The price is USD 0.10 per hour of audio as a limited-time offer until the end of 2026 (MAI-Transcribe-1: USD 0.36 per hour).

Can MAI-Transcribe-2 be used GDPR-compliant in the EU?

Per the Microsoft Learn region list the MAI-Transcribe model is available in eastus, northeurope, southeastasia and westus, in the EU therefore via North Europe (Ireland). Azure Speech processes and stores data exclusively in the region of the Speech resource, so processing stays in the EU with a resource in North Europe. Germany West Central is currently not offered for this model, and the preview runs without SLA.

Which MAI models does Microsoft offer?

As of 3 September 2026: MAI-Transcribe-2 and MAI-Transcribe-1.5 (speech-to-text), MAI-Voice-2 and MAI-Voice-2-Flash (text-to-speech with 15 languages, including German), MAI-Image-2.5 with the Flash and Pro variants (image generation), MAI-Thinking-1 (reasoning LLM with 256k context) and MAI-Code-1-Flash (5B coding model in GitHub Copilot). All are proprietary and run via Microsoft Foundry; MAI-Transcribe-1 has been deprecated since 20 August 2026.

Does MAI-Voice-2 offer German voices?

Yes. Per Microsoft Learn, MAI-Voice-2 and MAI-Voice-2-Flash offer the German voices de-DE-Klaus and de-DE-Mia with controllable emotions via SSML. In the EU the MAI voices are available in France Central, Sweden Central and West Europe; instant voice cloning is possible only after limited-access approval. Both models are public preview.

Do MAI-Thinking-1 and MAI-Image-2.5 run with EU data residency?

No. Both models are deployable in Microsoft Foundry exclusively as Global Standard, MAI-Thinking-1 from all European Foundry regions including Germany West Central, MAI-Image-2.5 from Sweden Central and West Europe. Global Standard means inference may run in any Azure region; Microsoft does not yet offer an EU data zone for the MAI models. Regional processing is available only for the speech models MAI-Transcribe and MAI-Voice.

When does Phi make more sense than MAI?

Phi models (Phi-4, Phi-4-reasoning-vision-15B, Phi-3.5) are open source and run locally on your own hardware or at the edge, i.e. with full data control and no cloud dependency. The MAI models are proprietary, more capable and usable only via Microsoft Foundry or Azure Speech. If you need maximum data sovereignty, choose Phi; if you want frontier-level reasoning, transcription or speech output in the Microsoft stack, evaluate MAI.

Consultation for this model?

We help you select and integrate the right AI model for your use case.