Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
LLM OpenAI USA

OpenAI GPT

New on 3 September 2026: GPT-6 Astra (gpt-6-astra, $10/$50 per 1M tokens, 1.05M context, cutoff 30 April 2026) – initially only for the Trusted Access Program, API and ChatGPT plans to follow; in Microsoft Foundry only Global and US Data Zone, no EU Data Zone. GPT-5.6 (Sol, Terra, Luna) remains the highest EU-data-zone-capable OpenAI model (all nine European Foundry regions); Sol at the promotional price of $4/$20, recommendation for most workloads GPT-5.6 Luna ($0.20/$1.20). AI consulting from Rosenheim, Germany. As of 3 September 2026.

License Proprietary
GDPR Hosting Available
Context 1.05M (GPT-6 Astra, GPT-5.6/5.5/5.4), 400k (GPT-5.2) Tokens
Modality Text, Image, Audio, PDF → Text, Image, Audio, Video

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
GPT-6 Astra
3 September 2026 (limited rollout)
Per OpenAI its most capable model to date ('generational leap'), first model of the sixth generation Benchmarks per OpenAI: ARC-AGI-3 98.6%, FrontierMath Tier 4 v2 97.6%, GPQA Diamond 96%, DeepSWE v1.1 74.1%, OSWorld 2.0 72.6% (GPT-5.6 Sol: 65.7%) Computer use across browsers, spreadsheets and desktop applications (incl. CRM, notebooks, Power BI, KiCad, FreeCAD) rather than APIs only 1.05M token context window (922k input, 128k output), knowledge cutoff 30 April 2026 Five reasoning levels (low to max); new in the Responses API: async tool calling, mid-turn steering via WebSocket, changing reasoning effort mid-conversation without cache loss Tools: web search, file search, code interpreter, hosted shell, apply patch, skills, computer use, MCP, tool search; structured outputs, function calling, prompt caching
First OpenAI model rated 'Critical' for cybersecurity under the Preparedness Framework – the generally available variant refuses advanced cybersecurity tasks per OpenAI, and monitoring may slow or pause legitimate work No general access yet: Trusted Access Program first, API and ChatGPT plans 'in the coming days' (as of 3 September 2026) No EU Data Zone in Microsoft Foundry (Global and US only), not yet listed on Bedrock – not suitable for EU residency requirements No reasoning effort none, no custom temperature/top_p values, no logprobs; tool calling only via the Responses API, not via Chat Completions Long-context surcharge above 272k tokens ($20.00 / $75.00) and double price in fast mode No Realtime, no Assistants API, no fine-tuning
Preview
GPT-5.6 Sol Recommended
July 9, 2026 (GA, preview since June 26)
Flagship ('Sol' tier in the Sol/Terra/Luna naming scheme) – flagship of the 5.6 generation for coding, knowledge work, cybersecurity and science; the alias 'gpt-5.6' routes to Sol; superseded as the top model by GPT-6 Astra since 3 September 2026, but remains the highest EU-data-zone-capable model Terminal-Bench 2.1: 88.8%, 91.9% with the 'Ultra' configuration (subagents) – new state of the art 80 points on the Artificial Analysis Coding Agent Index v1.1 (max reasoning) 1.05M token context window, 128k output, knowledge cutoff February 16, 2026 Six reasoning effort levels (none through max) plus reasoning.mode 'pro' as the replacement path for gpt-5-pro / o3-pro Promotional pricing since 21 August 2026: $4.00/1M input, $0.40/1M cached input, $20.00/1M output – per OpenAI at least through 21 November 2026 (previously $5.00/$0.50/$30.00) Deployable in Microsoft Foundry as EU Data Zone Standard – data stays in the EU Since 21 August 2026, regional processing per request in the OpenAI API (prefixed domain with an API key from a project with 'Global' geography) Ultrafast tier up to 14x faster than Standard per OpenAI (limited preview since 13 August 2026)
SWE-Bench Pro only 64.6% vs. 80% for Claude Fable 5 – though OpenAI considers the benchmark partially broken Long-context surcharge: prompts above 272k tokens are billed at $8.00/1M input and $30.00/1M output Promotional pricing is time-limited – OpenAI guarantees it only 'at least through November 21, 2026' 'High' classification in cybersecurity and bio/chem in the OpenAI Preparedness Framework Foundry deployment requires a quota request on lower quota tiers (tier 5/6 have default quota) On Bedrock: US regions only, reduced context window (272k), and only the reasoning levels 'medium' and 'maximum' Most expensive tier in the family – usually oversized for routine workloads
Current
GPT-5.6 Terra
July 9, 2026 (GA, preview since June 26)
Balanced tier between Sol and Luna Pricing $2.00/1M input, $0.20/1M cached input, $12.00/1M output (above 272k tokens: $4.00/$18.00) 1.05M token context window, knowledge cutoff February 16, 2026 Full reasoning scale (none through max) and full tool support Deployable in Microsoft Foundry as EU Data Zone Standard and Data Zone Provisioned Managed
On Bedrock: US regions only and reduced context window (272k) Considerably more expensive than Luna for many standard workloads, without the reasoning depth of Sol
Current
GPT-5.6 Luna Recommended
July 9, 2026 (GA, preview since June 26)
Price-performance recommendation: $0.20/1M input, $0.02/1M cached input, $1.20/1M output Same context window as Sol and Terra: 1.05M tokens, 128k output, knowledge cutoff February 16, 2026 Full tool support (web search, file search, code interpreter, MCP, computer use) and the full reasoning scale Default model for ChatGPT Free and Go since August 6, 2026 according to media reports Deployable in Microsoft Foundry as EU Data Zone Standard across all nine European regions Ideal for sub-agents, high volumes, and cost-sensitive workloads
Less reasoning depth than Sol and Terra – not the first choice for the hardest coding and analysis tasks Cost trap on very long context: prompts above 272k tokens are billed at $0.40/1M input and $1.80/1M output (2x input, 1.5x output – applies to all three tiers per the OpenAI price list) Not listed in the Data Zone Provisioned Managed table – currently not an option for PTU customers On Bedrock: US regions only and reduced context window (272k)
Current
GPT-5.5 Instant
May 2026
New default chat model for ChatGPT (replaces GPT-5.3 Instant) 52.5% fewer hallucinations than GPT-5.3 Instant Stronger factual accuracy and tool calling Available in Microsoft Foundry as 'gpt-chat-latest'
Not a dedicated reasoning model
Current
GPT-Realtime-2.1
2026 (successor to GPT-Realtime-2 from May 2026)
Current realtime voice generation (`gpt-realtime-2.1`), also available as `gpt-realtime-2.1-mini` Target model for migrating the legacy audio and realtime models More natural speech synthesis
Audio tokens are considerably more expensive than text tokens Legacy audio/realtime models retire on January 20, 2027 – migration required
Current
GPT-5.6 Cyber / Daybreak Red & Blue
August 2026
Specialized cybersecurity models: GPT-5.6 Cyber, Daybreak Red (vulnerability research), Daybreak Blue (defensive) In the OpenAI API since 7 August 2026 as Daybreak Blue and Daybreak Red tiers (`gpt-daybreak-red-latest`) for approved defenders, on Amazon Bedrock for eligible customers since 11 August 2026
Accessible only to eligible customers after approval On Bedrock: US regions only – no EU region
Current
GPT Image 2
April 21, 2026
Current image generation model (`gpt-image-2`) Successor to `gpt-image-1`, which retires on October 23, 2026
Image-only model – no text or reasoning capabilities
Current
Sora 2 / Sora 2 Pro
2025/2026
Video generation via the Videos API
Retirement announced: `sora-2`, `sora-2-pro` and the Videos API shut down on September 24, 2026 – with no successor model
Deprecated
GPT-Realtime-Translate
May 2026
Live real-time translation 70+ input languages, 13 output languages Billed per minute ($0.034/min)
Limited output languages
Current
GPT-Realtime-Whisper
May 2026
Live speech-to-text in the Realtime API Billed per minute ($0.017/min)
Specialized for transcription
Current
GPT Transcribe / GPT Live Transcribe
2026
Target models named by OpenAI for migrating from whisper-1 and the gpt-4o-transcribe family Batch transcription (gpt-transcribe) and streaming transcription (gpt-live-transcribe)
Specialized for transcription
Current
Whisper-1 / GPT-4o Transcribe
2023–2025
Proven transcription models, usable until shutdown
Deprecated on 26 August 2026 – shutdown on 26 February 2027; successors gpt-transcribe and gpt-live-transcribe
Deprecated
GPT-5.5
April 2026
Flagship (codename 'Spud') More efficient than GPT-5.4 Improved coding capabilities Variants: GPT-5.5 Thinking, GPT-5.5 Pro EU-available via Microsoft Foundry (EU Data Zone) and AWS Bedrock
High cost at large context
Current
GPT-5.4
March 2026
Flagship – 1M token context window Native computer use (desktop & browser) 33% fewer hallucinations than GPT-5.2 GDPval 83%, OSWorld-Verified 75% EU-available via Microsoft Foundry (EU Data Zone)
Premium pricing ($2.50/1M input, $15/1M output) Superseded as recommendation by GPT-5.6
Current
GPT-5.4 Pro
March 2026
Deepest reasoning of all OpenAI models Maximum precision for complex tasks
Significantly higher cost ($30/1M input, $180/1M output) Slowest variant
Current
GPT-5.4 mini
March 2026
2x faster than predecessor Ideal for quick code edits and classification
Lower capacity than GPT-5.4
Current
GPT-5.4 nano
March 2026
Lowest latency Ideal for sub-agents and repetitive tasks
Limited functionality
Current
GPT-5.3 Codex
February 2026
Agentic coding model 25% faster than GPT-5.2 Self-optimizing
Specialized for development
Current
o3
2025
Reasoning focused
Slower Retires on December 11, 2026 – successor: GPT-5.6 Sol (o3-pro via reasoning.mode 'pro')
Deprecated
o4-mini
2025
Reasoning focused Compact reasoning model
Specialized for reasoning Retires on October 23, 2026 – successor: GPT-5.6 Terra
Deprecated
GPT-5.2
December 2025
Proven model 400k token context window
Being superseded by GPT-5.4
Deprecated
GPT-5.2 pro
January 2026
Higher precision
Replaced by GPT-5.4 Pro
Deprecated
GPT-4.1
2025
Strong general model
Deprecated
GPT-4o
May 2024
Multimodal
Deprecated

Use Cases

Typical applications for this model

Coding & Software Development
Customer Service & Chatbots
Content Creation
Data Analysis
Translation
Agentic Workflows
Native Computer Use & Desktop Automation
Image Generation
Video Generation
Voice Assistance

Technical Details

API, features and capabilities

API & Availability
Availability Public
Requests/Min 10000
Tokens/Min 2000000
Latency (TTFT) ~300ms
Throughput ~200 Tokens/Sec
Features & Capabilities
Tool Use Function Calling Structured Output Vision Reasoning Mode Code Execution Web Browsing File Upload Realtime API
Training & Knowledge
Knowledge Cutoff April 30, 2026 (GPT-6 Astra), February 16, 2026 (GPT-5.6 Sol/Terra/Luna), varies by model
Fine-Tuning Available (Fine-tuning API, Custom Models)
Language Support
Best Quality English, German, French, Spanish, Chinese
Supported 100+ languages
Best quality in English, very good quality in European languages

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Microsoft Foundry
Data Zone Standard (EU) – all nine European regions: France Central, Germany West Central, Italy North, Norway East, Poland Central, Spain Central, Sweden Central, Switzerland North, West Europe
GPT-5.6 Sol, Terra and Luna are listed in all nine regions as Data Zone Standard; prompts and responses are processed within the EU data zone. Data Zone Batch is available as well. GPT-5.5 and GPT-5.4 also via Data Zone Standard, GPT-5.5 Instant as 'gpt-chat-latest'.
Microsoft Foundry
Data Zone Provisioned Managed (EU)
Only GPT-5.6 Sol and Terra are listed – GPT-5.6 Luna is not included in the PTU table.
Microsoft Foundry
Global Standard in EU regions (incl. Germany West Central, Sweden Central, Poland Central)
Deployable, but inference can happen worldwide. Choose Data Zone Standard (EU) explicitly for EU data processing.
OpenAI (first party)
Data residency Europe (EEA + Switzerland)
For models released on or after March 5, 2026 – including GPT-5.6 – at a 10 percent price premium. Zero data retention only after prior approval by OpenAI. Since 21 August 2026, regional processing can also be selected per request (prefixed domain with an API key from a project with 'Global' geography).
Azure OpenAI
West Europe (Netherlands), Germany West Central (Frankfurt), Sweden Central
Azure OpenAI Service – EU Data Boundary, classic GPT-5 family
AWS
US regions only (us-east-1, us-east-2, partly us-west-2)
Amazon Bedrock – GPT-5.5/5.4/Codex since June 1, GPT-5.6 GA since July 2026. Per AWS documentation no EU region available (as of 3 September 2026) – not suitable for GDPR workloads with EU residency requirements.
License & Hosting
License Proprietary
Security Filters Customizable
Enterprise Support Yes
SLA Available Yes
Cloud Only

Benchmarks

Performance comparison with standardized tests

ARC-AGI-3 (GPT-6 Astra)
98.6
FrontierMath Tier 4 v2 (GPT-6 Astra)
97.6
GPQA Diamond (GPT-6 Astra)
96
DeepSWE v1.1 (GPT-6 Astra)
74.1
OSWorld 2.0 (GPT-6 Astra)
72.6
GDPval (GPT-5.4)
83
SWE-bench Pro (GPT-5.4)
57.7
OSWorld-Verified (GPT-5.4)
75
Investment Banking Modeling (GPT-5.4)
87.3

As an AI consulting firm based in Rosenheim, Germany, we help enterprises across the DACH region (Germany, Austria, Switzerland) integrate OpenAI models in a GDPR-compliant way. With our CompanyGPT you can run GPT models securely in your own infrastructure.

What is GPT?

GPT (Generative Pre-trained Transformer) is OpenAI’s model family. New on 3 September 2026: OpenAI introduced the sixth generation with GPT-6 Astra (gpt-6-astra) – initially only for enterprises in the Trusted Access Program, with API and ChatGPT plans to follow ‘in the coming days’ per OpenAI; there is no EU Data Zone for it yet (details in the GPT-6 Astra section). Since July 9, 2026, GPT-5.6 (tiers Sol, Terra, Luna) is generally available – in ChatGPT, Codex, and the API – and remains the highest EU-data-zone-capable OpenAI model. The launch was staggered: after a government-cleared limited preview starting June 26, the US Department of Commerce concluded its review on July 8 and cleared the model for public launch. The most important news for EU customers: Microsoft Foundry lists all three tiers – gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna – as Data Zone Standard (EU) across all nine European Foundry regions, including Germany West Central (Frankfurt), West Europe, and Sweden Central. Prompts and responses are processed within the EU data zone. This makes GPT-5.6 deployable in EU data zones in a GDPR-compliant way, replacing GPT-5.5 as the highest EU-available OpenAI model. Amazon Bedrock offers GPT-5.6 GA since July 2026 – but exclusively in US regions; the AWS documentation lists no EU region for either GPT-5.6 or GPT-5.5 (as of 3 September 2026). Since 21 August 2026, GPT-5.6 Sol is also available at a promotional price of $4/1M input and $20/1M output (previously $5/$30), which OpenAI guarantees at least through 21 November 2026. Beyond that, GPT-5.5 (April 2026) and GPT-5.4 (March 2026, 1 million token context window, native computer control) remain proven EU options via Microsoft Foundry, and GPT-5.5 Instant is available as gpt-chat-latest. For specialized coding tasks, GPT-5.3 Codex remains available; the o-series with o3 and o4-mini, by contrast, has been deprecated and migrates to GPT-5.6 by the end of 2026.

GPT-6 Astra: Sixth Generation in a Phased Rollout (3 September 2026)

On 3 September 2026 OpenAI introduced GPT-6 Astra (gpt-6-astra). OpenAI president Greg Brockman speaks of a “generational leap” and welcomes the “AGI era”; OpenAI describes the model as its most intelligent and best-aligned to date. The codename backstory: on 1 August 2026 OpenAI had already named “Astra” as its next major model, and on 18 August it paused training for two weeks because of the model’s cyber capabilities.

Access – deliberately phased: At launch, enterprises in the Trusted Access Program (the Daybreak program for cybersecurity defenders) get access. Access via the API and the ChatGPT Plus, Pro, Business and Enterprise plans follows “in the coming days” per OpenAI; media reports name 9 September 2026 as the expected broad launch – not confirmed by OpenAI. Microsoft Foundry starts in parallel via the Foundry Limited Access Program, Amazon Bedrock is announced but not yet listed in the AWS model documentation as of 3 September.

Performance (figures from OpenAI): ARC-AGI-3 98.6%, FrontierMath Tier 4 v2 97.6% (Claude Fable 5.1 per OpenAI’s comparison 87.8%), GPQA Diamond 96%, DeepSWE v1.1 74.1%, BenchCAD 95.9%, OSWorld 2.0 72.6% at around 40 minutes per task (GPT-5.6 Sol: 65.7% at 75 minutes). The focus is on computer use: Astra operates browsers, spreadsheets, websites and desktop applications – forms, CRMs, calendars, Python notebooks, Power BI, KiCad, FreeCAD – through the human interface rather than via APIs. On the DeepSWE comparison OpenAI cites roughly 57% lower estimated API cost per task than the best GPT-5.6 Sol configuration. Notably, the GDPval benchmark for real-world occupational tasks is absent from the launch materials.

Technical data: context window 1,050,000 tokens (922k input, 128k output), knowledge cutoff 30 April 2026, text and image input, text output. Five reasoning levels (low, medium, high, xhigh, max) – the none level is gone. Endpoints: Responses, Chat Completions and Batch; no Realtime, no Assistants API, no fine-tuning. Custom temperature or top_p values and logprobs are not supported, and tool calling runs exclusively via the Responses API. New in the Responses API are async tool calling, mid-turn steering via WebSocket and changing reasoning effort mid-conversation while preserving the prompt cache.

Pricing (OpenAI first party, per 1M tokens): $10.00 input, $1.00 cached input, $12.50 cache write, $50.00 output up to 272k tokens; above that $20.00 / $2.00 / $75.00. Batch and flex 50% each, fast mode $20.00 / $100.00 (long context $40.00 / $150.00). That puts Astra at 2.5 times Sol’s promotional price. OpenAI argues the comparison should shift to price per completed task rather than per token – for budget planning we recommend exactly that measurement before migrating a workload.

Safety – the point compliance leads should know: Per OpenAI, Astra is the first model to reach the “Critical” level for cybersecurity under the Preparedness Framework – it can find previously unknown vulnerabilities and develop exploits. The consequence is a defense-in-depth approach of model behavior, classifiers, security controls, monitoring and post-deployment response: the generally available variant refuses advanced cybersecurity tasks, and OpenAI acknowledges that monitoring may slow or pause legitimate work. Advanced cyber capabilities are available only via Daybreak Blue for vetted defenders. Anyone planning to use Astra in security or engineering teams should account for these interventions in the operating concept.

EU availability (as of 3 September 2026): For GPT-6 Astra in Microsoft Foundry the Azure blog lists only Standard Global ($10.00 / $50.00) and Standard Data Zone (US) ($11.00 / $55.00; long context $22.00 / $82.50) – no EU Data Zone, no provisioned, no batch. On Bedrock the model is still missing. For workloads with EU residency requirements, GPT-5.6 via Foundry’s EU data zones therefore remains the highest available OpenAI model; we will update this page as soon as Microsoft documents an EU Data Zone for Astra.

GPT-5.6 (Sol, Terra, Luna): Generally Available Since July 9, 2026

On July 9, 2026, OpenAI publicly launched the GPT-5.6 family – in ChatGPT, Codex, and the API. The new naming scheme separates the generation number (5.6) from durable capability tiers; there are no more mini/nano variants:

  • GPT-5.6 Sol (gpt-5.6-sol) – flagship, according to OpenAI its strongest model to date (focus on coding, knowledge work, cybersecurity, science). Terminal-Bench 2.1: 88.8%, 91.9% with the Ultra configuration (subagents) – new state of the art. The alias gpt-5.6 routes to Sol.
  • GPT-5.6 Terra (gpt-5.6-terra) – balanced tier between Sol and Luna.
  • GPT-5.6 Luna (gpt-5.6-luna) – fast, very affordable tier with the same context window and feature set as Sol and Terra.

Pricing per 1 million tokens (OpenAI first party, USD)

ModelInputCached inputCache writeOutputInput / output above 272k tokens
GPT-5.6 Sol (promotional)$4.00$0.40$5.00$20.00$8.00 / $30.00
GPT-5.6 Terra$2.00$0.20$2.50$12.00$4.00 / $18.00
GPT-5.6 Luna$0.20$0.02$0.25$1.20$0.40 / $1.80

Price cut for Sol: Since 21 August 2026, GPT-5.6 Sol costs $4.00/1M input (down 20 percent) and $20.00/1M output (down 33 percent), previously $5.00/$30.00. OpenAI explicitly calls this promotional pricing, available at least through 21 November 2026 – for budgets that extend beyond that date, plan with the list price.

For all three tiers, per the OpenAI price list: prompts above 272,000 tokens are billed at 2x the input price and 1.5x the output price (as of 3 September 2026). If you regularly process very long contexts, factor that into your cost model. Since 5 August 2026, Fast mode also supports prompts above 272k tokens for Sol, Terra, and Luna (up to 2.5x faster than Standard per OpenAI); on 13 August an Ultrafast tier for Sol followed as a limited preview (up to 14x faster than Standard per OpenAI).

All three models offer a context window of 1,050,000 tokens (922k input, 128k output) with a knowledge cutoff of February 16, 2026. In terms of modalities, GPT-5.6 supports text and image input and text output – audio and video output are covered by other models in the family (Realtime, GPT Image). Reasoning effort can be controlled across six levels – none, low, medium (default), high, xhigh, and max – plus reasoning.mode: pro as the replacement path for the former Pro models. Supported tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search, along with streaming, structured outputs, function calling, and prompt caching. Also new in the API are programmatic tool calling, expanded multi-agent capabilities, and prompt cache breakpoints. Alongside the launch, OpenAI introduced ChatGPT Work – a work agent powered by Sol with Codex integration, initially for Pro, Enterprise, and Edu customers.

GPT-5.6 in ChatGPT: Luna as the new default for Free and Go

According to media reports, OpenAI changed the model assignment in ChatGPT on August 6, 2026: GPT-5.6 Luna is now the default model for Free and Go users, replacing GPT-5.5 Instant there. In the API changelog of the same day, OpenAI confirms an update of the chat-latest snapshot to the current ChatGPT model and recommends GPT-5.6 Sol for production workloads. Free and Go accounts receive unlimited text chats; limits remain in place for uploads, image generation, and tool usage. GPT-5.6 Sol is the model for Plus, Pro, Business, and Enterprise. In addition, a “Think” button lets users request more reasoning for a specific query.

An honest look at the benchmarks is part of the picture: on SWE-Bench Pro, Sol reaches only 64.6% according to independent analysis, versus around 80% for Claude Fable 5 – although OpenAI considers roughly 30% of the SWE-Bench Pro tasks broken. On agentic benchmarks like Terminal-Bench, Sol leads; on the Artificial Analysis Coding Agent Index v1.1, Sol scores 80 points at maximum reasoning. OpenAI also states that Sol is around 54 percent more token-efficient on coding tasks than its predecessor – that figure comes from secondary sources and we have not verified it independently.

From Government Clearance Process to Public Launch

Update – as of 3 September 2026: GPT-5.6 is deployable in EU data zones. Microsoft Foundry lists all three tiers (Sol, Terra, Luna) as Data Zone Standard (EU) across all nine European regions. Amazon Bedrock also offers GPT-5.6 GA since July 2026, but US regions only there. For GDPR-compliant EU deployments, GPT-5.6 via Microsoft Foundry is the reference – with GPT-5.6 Luna as the new price-performance recommendation.

The path to launch was unusual: GPT-5.6 initially started on June 26, 2026 only as a government-cleared limited preview for a small group of vetted organizations. The background was a US cybersecurity order under which the US Department of Commerce (Center for AI Standards and Innovation) could review the model before public release – triggered by the “High” classification in cybersecurity in the OpenAI Preparedness Framework. On July 8, 2026, the review concluded and the restriction was lifted; the public launch followed one day later. This repeated the Anthropic pattern (Claude Fable 5 / Mythos 5, Anthropic Claude) just days later – including the lifting of restrictions. Read more on the backstory in our blog post on GPT-5.6 and Claude Fable 5.

GPT-5.5 – The Previous-Generation Flagship

GPT-5.5 (April 23, 2026, codename “Spud”) was OpenAI’s top model until the GPT-5.6 launch. It is more efficient than GPT-5.4 and offers improved coding capabilities. In addition to the base model, the GPT-5.5 Thinking and GPT-5.5 Pro variants are available. Since April 24, 2026, GPT-5.5 is also available in the API ($5/1M input, $30/1M output, 1M context window).

GPT-5.5 Instant (May 2026)

On May 5, 2026, OpenAI introduced GPT-5.5 Instant as the new default chat model in ChatGPT, replacing GPT-5.3 Instant. In internal evaluations, GPT-5.5 Instant produces 52.5 percent fewer hallucinations than GPT-5.3 Instant on high-stakes prompts (medicine, law, finance). In the API it is available as chat-latest and in Microsoft Foundry as gpt-chat-latest – making it accessible for GDPR-compliant enterprise deployments in EU regions (depending on Foundry region configuration).

Realtime Voice and Transcription Models

The realtime family has moved on a generation since May 2026. Current status (3 September 2026):

  • GPT-Realtime-2.1 (gpt-realtime-2.1) and GPT-Realtime-2.1-mini – the current realtime voice generation. They supersede GPT-Realtime-2, introduced in May 2026 (128k token context window, $32/1M audio in, $64/1M audio out).
  • GPT-Realtime-Translate – Real-time live translation, 70+ input languages, 13 output languages. Pricing: $0.034/min.
  • Transcriptiongpt-transcribe ($0.0045/min), gpt-live-transcribe ($0.017/min), and gpt-realtime-whisper (live speech-to-text in the Realtime API, $0.017/min).

These models are well-suited for voice agents, conference translation, and real-time meeting notes. Important for ongoing projects: the legacy audio and realtime models retire on January 20, 2027; the migration targets are gpt-realtime-2.1 and gpt-realtime-2.1-mini.

New since 26 August 2026 – Whisper deprecation: OpenAI has deprecated whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize. The models will be removed from the API on 26 February 2027; OpenAI names gpt-live-transcribe (streaming) and gpt-transcribe (batch) as successors. If you use Whisper via the OpenAI API for meeting minutes, call-center transcripts, or diarization, plan the migration now – especially since output formats and diarization logic can change with the model switch.

Embeddings

For retrieval and RAG scenarios, text-embedding-3-large ($0.13/1M tokens) and text-embedding-3-small ($0.02/1M tokens) remain the current models – no successor has been announced so far.

Cybersecurity models

On 7 August 2026, per the API changelog, OpenAI released the Daybreak Blue and Daybreak Red tiers with the model gpt-daybreak-red-latest for approved defenders in the API; since 11 August 2026, eligible customers can also access GPT-5.6 Cyber, Daybreak Red (vulnerability research), and Daybreak Blue (defensive) on Amazon Bedrock. Access requires approval in each case, and no EU region is available on Bedrock. Since 3 September 2026 Daybreak program participants are the first to receive access to GPT-6 Astra; per OpenAI its advanced cyber capabilities are available exclusively via Daybreak Blue (see the GPT-6 Astra section).

GPT-5.4 – The Proven Flagship

GPT-5.4 (March 5, 2026) combines all key capabilities in a single model and sets new benchmarks across multiple domains:

Native Computer Use

GPT-5.4 can control desktop applications and browsers natively – a breakthrough for automating real-world workflows. Scoring 75 percent on the OSWorld-Verified benchmark, it surpasses the human baseline (72.4 percent) for GUI automation.

1 Million Token Context Window

With up to 1,050,000 tokens (922K input + 128K output), GPT-5.4 processes documents spanning thousands of pages – ideal for extensive contract analysis, code reviews, or research documents.

Tool Search

Instead of loading all tool definitions upfront, GPT-5.4 can dynamically search and use tools as needed. This reduces token costs in tool-heavy workflows by approximately 47 percent.

Model Variants

VariantStrengthPrice (Input/1M tokens)
GPT-5.4All-round flagship$2.50
GPT-5.4 ProDeepest reasoning$30.00
GPT-5.4 miniFast tasksAffordable
GPT-5.4 nanoSub-agents & repetitive tasksVery affordable

GPT-5.3 Codex: Agentic Coding

GPT-5.3 Codex (February 2026) remains the specialized model for agentic coding. It was the first OpenAI model that helped build itself and delivers over 1,000 tokens per second in the Codex-Spark variant.

Additional APIs

Realtime API

Real-time conversations with low latency:

  • Speech-to-Speech: Natural conversations
  • Text, Audio, Image: Multimodal inputs in real-time

Videos API: Sora 2 is being discontinued

OpenAI has announced the retirement of its video generation: sora-2, sora-2-pro, and the Videos API shut down on September 24, 2026 – with no successor model. If you use video generation in production, now is the time to evaluate an alternative. We support the selection and migration.

GPT Image 2 (Image Generation)

OpenAI’s current image generation model (gpt-image-2, available since April 21, 2026):

  • High-Fidelity: High-quality image output
  • Image Editing: Modification of existing images
  • Migration: The predecessor gpt-image-1 retires on October 23, 2026

Deprecations and migration roadmap

OpenAI has published a dense deprecation calendar for the second half of 2026. The first date has already passed: the Assistants API was shut down on 26 August 2026 – anyone who has not migrated yet must now move to the Responses API and the Conversations API. For existing projects it pays to look at the remaining dates early:

DateAffectedSuccessor
August 26, 2026 (done)Assistants APIResponses API + Conversations API
September 24, 2026sora-2, sora-2-pro, Videos APIno successor
October 23, 2026gpt-image-1gpt-image-2
October 23, 2026o1, o3-miniGPT-5.6 Sol
October 23, 2026o4-miniGPT-5.6 Terra
October 23, 2026various GPT-3.5 and GPT-4 modelsGPT-5.6 family
November 30, 2026Reusable Prompts, Evals platform, Agent Builder
December 11, 2026gpt-5, o3GPT-5.6 Sol
December 11, 2026gpt-5-miniGPT-5.6 Terra
December 11, 2026gpt-5-nanoGPT-5.6 Luna
December 11, 2026gpt-5-pro, o3-proGPT-5.6 Sol with reasoning.mode: pro
January 20, 2027legacy audio/realtimegpt-realtime-2.1, gpt-realtime-2.1-mini
February 26, 2027whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarizegpt-transcribe, gpt-live-transcribe

Data retention, zero data retention, and data residency

For GDPR projects, what happens to the data matters just as much as the region:

OpenAI directly (Platform API): by default, OpenAI stores abuse monitoring logs for up to 30 days. Zero data retention (ZDR) is possible but requires prior approval by OpenAI plus additional contractual terms; it can then be enabled at organization or project level, with the store parameter forcibly treated as false. ZDR-eligible endpoints include /v1/chat/completions, /v1/responses, /v1/images/*, /v1/embeddings, /v1/audio/*, /v1/realtime, and /v1/moderations – Assistants, Conversations, and Vector Stores are not ZDR-eligible. For Europe (EEA and Switzerland), OpenAI additionally offers data residency; for models released on or after March 5, 2026 – which includes GPT-5.6 – this carries a 10 percent price premium.

New since 21 August 2026 – regional processing per request: per the API changelog, API customers can now select regional processing for an individual request by using a prefixed domain together with an API key from a project with “Global” geography. This matters for mixed architectures: a single project can route EU-relevant requests to regional processing without running a separate residency project for everything. Which regions and models are supported in detail is something we verify per project against the OpenAI documentation. In addition, mutual TLS (mTLS) and X.509 workload identity federation have been generally available for the OpenAI API since 29 August 2026 – useful for enterprise authentication requirements without long-lived API keys.

Microsoft Foundry / Azure OpenAI: here, too, a 30-day abuse monitoring retention applies by default and cannot be switched off by the customer. Modified abuse monitoring or ZDR is available only after approval by Microsoft through the Azure OpenAI Limited Access Program and generally requires an Enterprise Agreement or MCA-E – there is no self-service toggle.

We clarify these points before go-live and document them in the record of processing activities.

GDPR-Compliant Deployment in the EU

As of 3 September 2026: GPT-5.6 (Sol/Terra/Luna) is available in Microsoft Foundry as Data Zone Standard (EU) across all nine European regions – including Germany West Central (Frankfurt). This makes GPT-5.6 the highest OpenAI model deployable in EU data zones in a GDPR-compliant way – GPT-6 Astra, introduced today, is so far available in Foundry only as Global and US Data Zone. Amazon Bedrock, by contrast, offers the OpenAI models (GPT-5.6 since July 2026, GPT-5.5/5.4/Codex since June 1) in US regions only – currently not an option for EU workloads with residency requirements.

Available now (EU): Microsoft Foundry

GPT-5.6 (Sol, Terra, Luna), GPT-5.5, and GPT-5.4 are generally available in Microsoft Foundry. For GDPR-compliant deployments, the deployment type is decisive:

  • Data Zone Standard (EU) – per the Microsoft Learn documentation, all three GPT-5.6 tiers are listed in all nine European Foundry regions: France Central, Germany West Central (Frankfurt), Italy North, Norway East, Poland Central, Spain Central, Sweden Central, Switzerland North, and West Europe. Prompts and responses are processed within the stated data zone (“European Union: data processed within any EU member nation”), and data at rest stays within the Azure geography. GPT-5.5 and GPT-5.4 are available here as well.
  • Data Zone Provisioned Managed (EU) – for predictable throughput with reserved capacity. Only Sol and Terra are listed – Luna is not. Anyone who wants to run Luna in production is therefore limited to Data Zone Standard or Batch.
  • Data Zone Batch (EU) – for asynchronous bulk processing, also within the EU data zone.
  • Global Standard – deployable, but inference can happen worldwide. For EU data processing, choose Data Zone Standard (EU) explicitly.

For GPT-5.6, tier 5 and tier 6 subscriptions have default quota; lower quota tiers must submit a quota request. Alongside this, the classic GPT-5 family remains available via the Azure OpenAI Service in West Europe (Netherlands, EU Data Boundary), Germany West Central (Frankfurt), and Sweden Central. We verify the specific region and deployment availability per project.

Amazon Bedrock: GPT-5.6 GA – but US regions only

Since July 2026, GPT-5.6 Sol, Terra, and Luna are generally available on Amazon Bedrock (model IDs openai.gpt-5.6-sol, openai.gpt-5.6-terra, openai.gpt-5.6-luna via the bedrock-mantle endpoint with the OpenAI Responses API). GPT-5.5, GPT-5.4, and Codex have been available on Bedrock since June 1, 2026. The background is the strategic partnership between Amazon and OpenAI ($50B investment), making AWS the exclusive third-party cloud distribution partner for OpenAI Frontier.

Important for EU customers – and a correction of our earlier assessment: the AWS documentation lists US regions only for the OpenAI models on Bedrock – GPT-5.6 Sol in us-east-1 (N. Virginia) and us-east-2 (Ohio), Terra and Luna additionally in us-west-2 (Oregon); cross-region inference profiles (Geo/Global) are not supported. An EU region (such as eu-central-1 Frankfurt) is not available so far. In addition, the context window on Bedrock is 272k tokens – considerably smaller than the 1.05M tokens via the OpenAI API and Foundry – and Sol can only be run there with the reasoning levels medium and maximum, not the full six-level scale. On the plus side: pricing matches OpenAI first-party rates, prompt caching is supported with a 90 percent discount on cached input, and AWS describes “zero-operator access” at the chip level for the OpenAI models; traffic flagged by classifiers is stored for up to 30 days. For GDPR workloads with EU residency requirements, Bedrock remains off the table for OpenAI models for now – we continuously verify EU region availability.

Integration with CompanyGPT

With CompanyGPT you can use GPT models GDPR-compliant in your company – without your data being used for training.

Our Recommendation

On GPT-6 Astra: watch, do not migrate yet. As of 3 September 2026 the model is unlocked only for the Trusted Access Program, has no EU Data Zone in Foundry, and comes with cyber monitoring that may slow legitimate work. What makes sense is an evaluation slot once the API is open – measured by price per completed task, not per token, because Astra sits at 2.5 times Sol’s promotional price.

With the current pricing and the broad data zone availability, our recommendation has shifted – we now recommend two models with clearly separated roles:

  • GPT-5.6 Luna via Microsoft Foundry (Data Zone Standard, EU) – the new default for the majority of enterprise workloads. Luna offers the same context window as the flagship (1.05M tokens), full tool support, and the EU data zone across all nine European Foundry regions – at a fraction of the cost ($0.20/1M input, $1.20/1M output). For chat assistants, classification, extraction, summarization, RAG, and sub-agents this is usually the economically right choice.
  • GPT-5.6 Sol via Microsoft Foundry (Data Zone Standard, EU) – the flagship for the most demanding tasks. For deep reasoning chains, complex coding, and autonomous agentic workflows, Sol remains the reference – since 21 August 2026 at a promotional price of $4.00/1M input and $20.00/1M output (at least through 21 November 2026; list price previously $5.00/$30.00). Whether Foundry pricing follows the OpenAI promotional price is something we verify per project. Quota note: tier 5/6 have default quota, below that a quota request is required.
  • GPT-5.6 Terra – the middle ground ($2.00/1M input, $12.00/1M output). Useful when Luna’s quality is not enough but Sol is oversized – or when Data Zone Provisioned Managed is required, which is not listed for Luna.

Honest about Luna’s limits: its reasoning depth is below that of Sol and Terra – for demanding analysis and coding tasks it is worth benchmarking against the larger tier. Prompts above 272,000 tokens are billed at 2x input and 1.5x output pricing – as with Sol and Terra – which erodes the cost advantage on very long contexts. And for PTU customers Luna is currently not an option, because it is not listed in the Data Zone Provisioned Managed table. Our recommendation is therefore: build workloads on Luna by default and escalate selectively to Terra or Sol where quality demands it.

Further options:

  • GPT-5.5 / GPT-5.4 via Microsoft Foundry (Data Zone Standard) – proven options for existing projects; migrating to GPT-5.6 is usually worthwhile short-term given the identical context window (1.05M tokens) and better benchmarks.
  • GPT-5.5 Instant in Microsoft Foundry (gpt-chat-latest) – still a good choice for chat workloads in EU regions.
  • Amazon Bedrock – GPT-5.6/5.5/5.4 are GA there, but US regions only, with a reduced 272k context window and only two reasoning levels for Sol. Currently not an option for EU workloads with residency requirements; interesting as a multi-cloud path for US workloads though.

The platform landscape is moving fast right now: Microsoft Foundry and AWS Bedrock keep expanding model and region coverage. We verify the current EU availability per project. For specialized coding tasks, choose GPT-5.3 Codex (also GA on Bedrock). If you are still running gpt-5, o3, or o4-mini, plan the migration now – the deprecation calendar runs through February 2027 (Whisper transcription models). The Assistants API has already been shut down since 26 August 2026.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is GPT-6 Astra and who can use it?

GPT-6 Astra (gpt-6-astra) is the sixth-generation OpenAI model introduced on 3 September 2026. At launch only enterprises in the Trusted Access Program (Daybreak) get access; access via the API and the ChatGPT Plus, Pro, Business and Enterprise plans follows in the coming days per OpenAI. Microsoft Foundry starts via the Limited Access Program, Amazon Bedrock is announced.

What does GPT-6 Astra cost?

In the OpenAI API USD 10 per 1 million input tokens, USD 1 for cached input and USD 50 per 1 million output tokens up to 272k tokens of context; above that 20 / 2 / 75 USD. Batch and flex cost half, fast mode double (20 / 100 USD). In Microsoft Foundry the US Data Zone costs 11 / 55 USD. That is 2.5 times the promotional price of GPT-5.6 Sol (4 / 20 USD).

Is GPT-6 Astra available in the EU with data residency?

No, as of 3 September 2026. Per the Azure blog Microsoft Foundry offers GPT-6 Astra only as Standard Global and Standard Data Zone (US), no EU Data Zone is listed; on Amazon Bedrock the model is not yet available. For EU residency requirements GPT-5.6 (Sol, Terra, Luna) via Data Zone Standard (EU) across all nine European Foundry regions remains the highest OpenAI model.

What are the limitations of GPT-6 Astra?

Astra is the first OpenAI model rated Critical for cybersecurity under the Preparedness Framework. The generally available variant refuses advanced cybersecurity tasks, and monitoring may slow or pause legitimate work. Technically the reasoning effort none, custom temperature and top_p values and logprobs are gone; tool calling runs only via the Responses API, and Realtime, Assistants API and fine-tuning are not supported.

Which OpenAI model do we recommend for GDPR-bound workloads?

GPT-5.6 Luna via Microsoft Foundry as Data Zone Standard (EU) for the majority of enterprise workloads (USD 0.20 / 1.20 per 1M tokens, 1.05M context) and GPT-5.6 Sol for the most demanding tasks (promotional price USD 4 / 20 at least through 21 November 2026). Both run in all nine European Foundry regions, including Germany West Central. We recommend evaluating GPT-6 Astra only once the API is open and an EU Data Zone is documented.

How does GPT-6 Astra perform on benchmarks?

Per OpenAI, GPT-6 Astra reaches 98.6 percent on ARC-AGI-3, 97.6 percent on FrontierMath Tier 4 v2, 96 percent on GPQA Diamond, 74.1 percent on DeepSWE v1.1 and 72.6 percent on OSWorld 2.0 (GPT-5.6 Sol: 65.7 percent). The knowledge cutoff is 30 April 2026, the context window 1,050,000 tokens. Independent measurements were not yet available at launch.

Consultation for this model?

We help you select and integrate the right AI model for your use case.