As an AI consulting firm based in Rosenheim, Germany, we help enterprises across the DACH region (Germany, Austria, Switzerland) integrate OpenAI models in a GDPR-compliant way. With our CompanyGPT you can run GPT models securely in your own infrastructure.
What is GPT?
GPT (Generative Pre-trained Transformer) is OpenAI’s model family. New on 3 September 2026: OpenAI introduced the sixth generation with GPT-6 Astra (gpt-6-astra) – initially only for enterprises in the Trusted Access Program, with API and ChatGPT plans to follow ‘in the coming days’ per OpenAI; there is no EU Data Zone for it yet (details in the GPT-6 Astra section). Since July 9, 2026, GPT-5.6 (tiers Sol, Terra, Luna) is generally available – in ChatGPT, Codex, and the API – and remains the highest EU-data-zone-capable OpenAI model. The launch was staggered: after a government-cleared limited preview starting June 26, the US Department of Commerce concluded its review on July 8 and cleared the model for public launch. The most important news for EU customers: Microsoft Foundry lists all three tiers – gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna – as Data Zone Standard (EU) across all nine European Foundry regions, including Germany West Central (Frankfurt), West Europe, and Sweden Central. Prompts and responses are processed within the EU data zone. This makes GPT-5.6 deployable in EU data zones in a GDPR-compliant way, replacing GPT-5.5 as the highest EU-available OpenAI model. Amazon Bedrock offers GPT-5.6 GA since July 2026 – but exclusively in US regions; the AWS documentation lists no EU region for either GPT-5.6 or GPT-5.5 (as of 3 September 2026). Since 21 August 2026, GPT-5.6 Sol is also available at a promotional price of $4/1M input and $20/1M output (previously $5/$30), which OpenAI guarantees at least through 21 November 2026. Beyond that, GPT-5.5 (April 2026) and GPT-5.4 (March 2026, 1 million token context window, native computer control) remain proven EU options via Microsoft Foundry, and GPT-5.5 Instant is available as gpt-chat-latest. For specialized coding tasks, GPT-5.3 Codex remains available; the o-series with o3 and o4-mini, by contrast, has been deprecated and migrates to GPT-5.6 by the end of 2026.
GPT-6 Astra: Sixth Generation in a Phased Rollout (3 September 2026)
On 3 September 2026 OpenAI introduced GPT-6 Astra (gpt-6-astra). OpenAI president Greg Brockman speaks of a “generational leap” and welcomes the “AGI era”; OpenAI describes the model as its most intelligent and best-aligned to date. The codename backstory: on 1 August 2026 OpenAI had already named “Astra” as its next major model, and on 18 August it paused training for two weeks because of the model’s cyber capabilities.
Access – deliberately phased: At launch, enterprises in the Trusted Access Program (the Daybreak program for cybersecurity defenders) get access. Access via the API and the ChatGPT Plus, Pro, Business and Enterprise plans follows “in the coming days” per OpenAI; media reports name 9 September 2026 as the expected broad launch – not confirmed by OpenAI. Microsoft Foundry starts in parallel via the Foundry Limited Access Program, Amazon Bedrock is announced but not yet listed in the AWS model documentation as of 3 September.
Performance (figures from OpenAI): ARC-AGI-3 98.6%, FrontierMath Tier 4 v2 97.6% (Claude Fable 5.1 per OpenAI’s comparison 87.8%), GPQA Diamond 96%, DeepSWE v1.1 74.1%, BenchCAD 95.9%, OSWorld 2.0 72.6% at around 40 minutes per task (GPT-5.6 Sol: 65.7% at 75 minutes). The focus is on computer use: Astra operates browsers, spreadsheets, websites and desktop applications – forms, CRMs, calendars, Python notebooks, Power BI, KiCad, FreeCAD – through the human interface rather than via APIs. On the DeepSWE comparison OpenAI cites roughly 57% lower estimated API cost per task than the best GPT-5.6 Sol configuration. Notably, the GDPval benchmark for real-world occupational tasks is absent from the launch materials.
Technical data: context window 1,050,000 tokens (922k input, 128k output), knowledge cutoff 30 April 2026, text and image input, text output. Five reasoning levels (low, medium, high, xhigh, max) – the none level is gone. Endpoints: Responses, Chat Completions and Batch; no Realtime, no Assistants API, no fine-tuning. Custom temperature or top_p values and logprobs are not supported, and tool calling runs exclusively via the Responses API. New in the Responses API are async tool calling, mid-turn steering via WebSocket and changing reasoning effort mid-conversation while preserving the prompt cache.
Pricing (OpenAI first party, per 1M tokens): $10.00 input, $1.00 cached input, $12.50 cache write, $50.00 output up to 272k tokens; above that $20.00 / $2.00 / $75.00. Batch and flex 50% each, fast mode $20.00 / $100.00 (long context $40.00 / $150.00). That puts Astra at 2.5 times Sol’s promotional price. OpenAI argues the comparison should shift to price per completed task rather than per token – for budget planning we recommend exactly that measurement before migrating a workload.
Safety – the point compliance leads should know: Per OpenAI, Astra is the first model to reach the “Critical” level for cybersecurity under the Preparedness Framework – it can find previously unknown vulnerabilities and develop exploits. The consequence is a defense-in-depth approach of model behavior, classifiers, security controls, monitoring and post-deployment response: the generally available variant refuses advanced cybersecurity tasks, and OpenAI acknowledges that monitoring may slow or pause legitimate work. Advanced cyber capabilities are available only via Daybreak Blue for vetted defenders. Anyone planning to use Astra in security or engineering teams should account for these interventions in the operating concept.
EU availability (as of 3 September 2026): For GPT-6 Astra in Microsoft Foundry the Azure blog lists only Standard Global ($10.00 / $50.00) and Standard Data Zone (US) ($11.00 / $55.00; long context $22.00 / $82.50) – no EU Data Zone, no provisioned, no batch. On Bedrock the model is still missing. For workloads with EU residency requirements, GPT-5.6 via Foundry’s EU data zones therefore remains the highest available OpenAI model; we will update this page as soon as Microsoft documents an EU Data Zone for Astra.
GPT-5.6 (Sol, Terra, Luna): Generally Available Since July 9, 2026
On July 9, 2026, OpenAI publicly launched the GPT-5.6 family – in ChatGPT, Codex, and the API. The new naming scheme separates the generation number (5.6) from durable capability tiers; there are no more mini/nano variants:
- GPT-5.6 Sol (
gpt-5.6-sol) – flagship, according to OpenAI its strongest model to date (focus on coding, knowledge work, cybersecurity, science). Terminal-Bench 2.1: 88.8%, 91.9% with the Ultra configuration (subagents) – new state of the art. The aliasgpt-5.6routes to Sol. - GPT-5.6 Terra (
gpt-5.6-terra) – balanced tier between Sol and Luna. - GPT-5.6 Luna (
gpt-5.6-luna) – fast, very affordable tier with the same context window and feature set as Sol and Terra.
Pricing per 1 million tokens (OpenAI first party, USD)
| Model | Input | Cached input | Cache write | Output | Input / output above 272k tokens |
|---|---|---|---|---|---|
| GPT-5.6 Sol (promotional) | $4.00 | $0.40 | $5.00 | $20.00 | $8.00 / $30.00 |
| GPT-5.6 Terra | $2.00 | $0.20 | $2.50 | $12.00 | $4.00 / $18.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 / $1.80 |
Price cut for Sol: Since 21 August 2026, GPT-5.6 Sol costs $4.00/1M input (down 20 percent) and $20.00/1M output (down 33 percent), previously $5.00/$30.00. OpenAI explicitly calls this promotional pricing, available at least through 21 November 2026 – for budgets that extend beyond that date, plan with the list price.
For all three tiers, per the OpenAI price list: prompts above 272,000 tokens are billed at 2x the input price and 1.5x the output price (as of 3 September 2026). If you regularly process very long contexts, factor that into your cost model. Since 5 August 2026, Fast mode also supports prompts above 272k tokens for Sol, Terra, and Luna (up to 2.5x faster than Standard per OpenAI); on 13 August an Ultrafast tier for Sol followed as a limited preview (up to 14x faster than Standard per OpenAI).
All three models offer a context window of 1,050,000 tokens (922k input, 128k output) with a knowledge cutoff of February 16, 2026. In terms of modalities, GPT-5.6 supports text and image input and text output – audio and video output are covered by other models in the family (Realtime, GPT Image). Reasoning effort can be controlled across six levels – none, low, medium (default), high, xhigh, and max – plus reasoning.mode: pro as the replacement path for the former Pro models. Supported tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search, along with streaming, structured outputs, function calling, and prompt caching. Also new in the API are programmatic tool calling, expanded multi-agent capabilities, and prompt cache breakpoints. Alongside the launch, OpenAI introduced ChatGPT Work – a work agent powered by Sol with Codex integration, initially for Pro, Enterprise, and Edu customers.
GPT-5.6 in ChatGPT: Luna as the new default for Free and Go
According to media reports, OpenAI changed the model assignment in ChatGPT on August 6, 2026: GPT-5.6 Luna is now the default model for Free and Go users, replacing GPT-5.5 Instant there. In the API changelog of the same day, OpenAI confirms an update of the chat-latest snapshot to the current ChatGPT model and recommends GPT-5.6 Sol for production workloads. Free and Go accounts receive unlimited text chats; limits remain in place for uploads, image generation, and tool usage. GPT-5.6 Sol is the model for Plus, Pro, Business, and Enterprise. In addition, a “Think” button lets users request more reasoning for a specific query.
An honest look at the benchmarks is part of the picture: on SWE-Bench Pro, Sol reaches only 64.6% according to independent analysis, versus around 80% for Claude Fable 5 – although OpenAI considers roughly 30% of the SWE-Bench Pro tasks broken. On agentic benchmarks like Terminal-Bench, Sol leads; on the Artificial Analysis Coding Agent Index v1.1, Sol scores 80 points at maximum reasoning. OpenAI also states that Sol is around 54 percent more token-efficient on coding tasks than its predecessor – that figure comes from secondary sources and we have not verified it independently.
From Government Clearance Process to Public Launch
Update – as of 3 September 2026: GPT-5.6 is deployable in EU data zones. Microsoft Foundry lists all three tiers (Sol, Terra, Luna) as Data Zone Standard (EU) across all nine European regions. Amazon Bedrock also offers GPT-5.6 GA since July 2026, but US regions only there. For GDPR-compliant EU deployments, GPT-5.6 via Microsoft Foundry is the reference – with GPT-5.6 Luna as the new price-performance recommendation.
The path to launch was unusual: GPT-5.6 initially started on June 26, 2026 only as a government-cleared limited preview for a small group of vetted organizations. The background was a US cybersecurity order under which the US Department of Commerce (Center for AI Standards and Innovation) could review the model before public release – triggered by the “High” classification in cybersecurity in the OpenAI Preparedness Framework. On July 8, 2026, the review concluded and the restriction was lifted; the public launch followed one day later. This repeated the Anthropic pattern (Claude Fable 5 / Mythos 5, Anthropic Claude) just days later – including the lifting of restrictions. Read more on the backstory in our blog post on GPT-5.6 and Claude Fable 5.
GPT-5.5 – The Previous-Generation Flagship
GPT-5.5 (April 23, 2026, codename “Spud”) was OpenAI’s top model until the GPT-5.6 launch. It is more efficient than GPT-5.4 and offers improved coding capabilities. In addition to the base model, the GPT-5.5 Thinking and GPT-5.5 Pro variants are available. Since April 24, 2026, GPT-5.5 is also available in the API ($5/1M input, $30/1M output, 1M context window).
GPT-5.5 Instant (May 2026)
On May 5, 2026, OpenAI introduced GPT-5.5 Instant as the new default chat model in ChatGPT, replacing GPT-5.3 Instant. In internal evaluations, GPT-5.5 Instant produces 52.5 percent fewer hallucinations than GPT-5.3 Instant on high-stakes prompts (medicine, law, finance). In the API it is available as chat-latest and in Microsoft Foundry as gpt-chat-latest – making it accessible for GDPR-compliant enterprise deployments in EU regions (depending on Foundry region configuration).
Realtime Voice and Transcription Models
The realtime family has moved on a generation since May 2026. Current status (3 September 2026):
- GPT-Realtime-2.1 (
gpt-realtime-2.1) and GPT-Realtime-2.1-mini – the current realtime voice generation. They supersede GPT-Realtime-2, introduced in May 2026 (128k token context window, $32/1M audio in, $64/1M audio out). - GPT-Realtime-Translate – Real-time live translation, 70+ input languages, 13 output languages. Pricing: $0.034/min.
- Transcription –
gpt-transcribe($0.0045/min),gpt-live-transcribe($0.017/min), andgpt-realtime-whisper(live speech-to-text in the Realtime API, $0.017/min).
These models are well-suited for voice agents, conference translation, and real-time meeting notes. Important for ongoing projects: the legacy audio and realtime models retire on January 20, 2027; the migration targets are gpt-realtime-2.1 and gpt-realtime-2.1-mini.
New since 26 August 2026 – Whisper deprecation: OpenAI has deprecated whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize. The models will be removed from the API on 26 February 2027; OpenAI names gpt-live-transcribe (streaming) and gpt-transcribe (batch) as successors. If you use Whisper via the OpenAI API for meeting minutes, call-center transcripts, or diarization, plan the migration now – especially since output formats and diarization logic can change with the model switch.
Embeddings
For retrieval and RAG scenarios, text-embedding-3-large ($0.13/1M tokens) and text-embedding-3-small ($0.02/1M tokens) remain the current models – no successor has been announced so far.
Cybersecurity models
On 7 August 2026, per the API changelog, OpenAI released the Daybreak Blue and Daybreak Red tiers with the model gpt-daybreak-red-latest for approved defenders in the API; since 11 August 2026, eligible customers can also access GPT-5.6 Cyber, Daybreak Red (vulnerability research), and Daybreak Blue (defensive) on Amazon Bedrock. Access requires approval in each case, and no EU region is available on Bedrock. Since 3 September 2026 Daybreak program participants are the first to receive access to GPT-6 Astra; per OpenAI its advanced cyber capabilities are available exclusively via Daybreak Blue (see the GPT-6 Astra section).
GPT-5.4 – The Proven Flagship
GPT-5.4 (March 5, 2026) combines all key capabilities in a single model and sets new benchmarks across multiple domains:
Native Computer Use
GPT-5.4 can control desktop applications and browsers natively – a breakthrough for automating real-world workflows. Scoring 75 percent on the OSWorld-Verified benchmark, it surpasses the human baseline (72.4 percent) for GUI automation.
1 Million Token Context Window
With up to 1,050,000 tokens (922K input + 128K output), GPT-5.4 processes documents spanning thousands of pages – ideal for extensive contract analysis, code reviews, or research documents.
Tool Search
Instead of loading all tool definitions upfront, GPT-5.4 can dynamically search and use tools as needed. This reduces token costs in tool-heavy workflows by approximately 47 percent.
Model Variants
| Variant | Strength | Price (Input/1M tokens) |
|---|---|---|
| GPT-5.4 | All-round flagship | $2.50 |
| GPT-5.4 Pro | Deepest reasoning | $30.00 |
| GPT-5.4 mini | Fast tasks | Affordable |
| GPT-5.4 nano | Sub-agents & repetitive tasks | Very affordable |
GPT-5.3 Codex: Agentic Coding
GPT-5.3 Codex (February 2026) remains the specialized model for agentic coding. It was the first OpenAI model that helped build itself and delivers over 1,000 tokens per second in the Codex-Spark variant.
Additional APIs
Realtime API
Real-time conversations with low latency:
- Speech-to-Speech: Natural conversations
- Text, Audio, Image: Multimodal inputs in real-time
Videos API: Sora 2 is being discontinued
OpenAI has announced the retirement of its video generation: sora-2, sora-2-pro, and the Videos API shut down on September 24, 2026 – with no successor model. If you use video generation in production, now is the time to evaluate an alternative. We support the selection and migration.
GPT Image 2 (Image Generation)
OpenAI’s current image generation model (gpt-image-2, available since April 21, 2026):
- High-Fidelity: High-quality image output
- Image Editing: Modification of existing images
- Migration: The predecessor
gpt-image-1retires on October 23, 2026
Deprecations and migration roadmap
OpenAI has published a dense deprecation calendar for the second half of 2026. The first date has already passed: the Assistants API was shut down on 26 August 2026 – anyone who has not migrated yet must now move to the Responses API and the Conversations API. For existing projects it pays to look at the remaining dates early:
| Date | Affected | Successor |
|---|---|---|
| August 26, 2026 (done) | Assistants API | Responses API + Conversations API |
| September 24, 2026 | sora-2, sora-2-pro, Videos API | no successor |
| October 23, 2026 | gpt-image-1 | gpt-image-2 |
| October 23, 2026 | o1, o3-mini | GPT-5.6 Sol |
| October 23, 2026 | o4-mini | GPT-5.6 Terra |
| October 23, 2026 | various GPT-3.5 and GPT-4 models | GPT-5.6 family |
| November 30, 2026 | Reusable Prompts, Evals platform, Agent Builder | — |
| December 11, 2026 | gpt-5, o3 | GPT-5.6 Sol |
| December 11, 2026 | gpt-5-mini | GPT-5.6 Terra |
| December 11, 2026 | gpt-5-nano | GPT-5.6 Luna |
| December 11, 2026 | gpt-5-pro, o3-pro | GPT-5.6 Sol with reasoning.mode: pro |
| January 20, 2027 | legacy audio/realtime | gpt-realtime-2.1, gpt-realtime-2.1-mini |
| February 26, 2027 | whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize | gpt-transcribe, gpt-live-transcribe |
Data retention, zero data retention, and data residency
For GDPR projects, what happens to the data matters just as much as the region:
OpenAI directly (Platform API): by default, OpenAI stores abuse monitoring logs for up to 30 days. Zero data retention (ZDR) is possible but requires prior approval by OpenAI plus additional contractual terms; it can then be enabled at organization or project level, with the store parameter forcibly treated as false. ZDR-eligible endpoints include /v1/chat/completions, /v1/responses, /v1/images/*, /v1/embeddings, /v1/audio/*, /v1/realtime, and /v1/moderations – Assistants, Conversations, and Vector Stores are not ZDR-eligible. For Europe (EEA and Switzerland), OpenAI additionally offers data residency; for models released on or after March 5, 2026 – which includes GPT-5.6 – this carries a 10 percent price premium.
New since 21 August 2026 – regional processing per request: per the API changelog, API customers can now select regional processing for an individual request by using a prefixed domain together with an API key from a project with “Global” geography. This matters for mixed architectures: a single project can route EU-relevant requests to regional processing without running a separate residency project for everything. Which regions and models are supported in detail is something we verify per project against the OpenAI documentation. In addition, mutual TLS (mTLS) and X.509 workload identity federation have been generally available for the OpenAI API since 29 August 2026 – useful for enterprise authentication requirements without long-lived API keys.
Microsoft Foundry / Azure OpenAI: here, too, a 30-day abuse monitoring retention applies by default and cannot be switched off by the customer. Modified abuse monitoring or ZDR is available only after approval by Microsoft through the Azure OpenAI Limited Access Program and generally requires an Enterprise Agreement or MCA-E – there is no self-service toggle.
We clarify these points before go-live and document them in the record of processing activities.
GDPR-Compliant Deployment in the EU
As of 3 September 2026: GPT-5.6 (Sol/Terra/Luna) is available in Microsoft Foundry as Data Zone Standard (EU) across all nine European regions – including Germany West Central (Frankfurt). This makes GPT-5.6 the highest OpenAI model deployable in EU data zones in a GDPR-compliant way – GPT-6 Astra, introduced today, is so far available in Foundry only as Global and US Data Zone. Amazon Bedrock, by contrast, offers the OpenAI models (GPT-5.6 since July 2026, GPT-5.5/5.4/Codex since June 1) in US regions only – currently not an option for EU workloads with residency requirements.
Available now (EU): Microsoft Foundry
GPT-5.6 (Sol, Terra, Luna), GPT-5.5, and GPT-5.4 are generally available in Microsoft Foundry. For GDPR-compliant deployments, the deployment type is decisive:
- Data Zone Standard (EU) – per the Microsoft Learn documentation, all three GPT-5.6 tiers are listed in all nine European Foundry regions: France Central, Germany West Central (Frankfurt), Italy North, Norway East, Poland Central, Spain Central, Sweden Central, Switzerland North, and West Europe. Prompts and responses are processed within the stated data zone (“European Union: data processed within any EU member nation”), and data at rest stays within the Azure geography. GPT-5.5 and GPT-5.4 are available here as well.
- Data Zone Provisioned Managed (EU) – for predictable throughput with reserved capacity. Only Sol and Terra are listed – Luna is not. Anyone who wants to run Luna in production is therefore limited to Data Zone Standard or Batch.
- Data Zone Batch (EU) – for asynchronous bulk processing, also within the EU data zone.
- Global Standard – deployable, but inference can happen worldwide. For EU data processing, choose Data Zone Standard (EU) explicitly.
For GPT-5.6, tier 5 and tier 6 subscriptions have default quota; lower quota tiers must submit a quota request. Alongside this, the classic GPT-5 family remains available via the Azure OpenAI Service in West Europe (Netherlands, EU Data Boundary), Germany West Central (Frankfurt), and Sweden Central. We verify the specific region and deployment availability per project.
Amazon Bedrock: GPT-5.6 GA – but US regions only
Since July 2026, GPT-5.6 Sol, Terra, and Luna are generally available on Amazon Bedrock (model IDs openai.gpt-5.6-sol, openai.gpt-5.6-terra, openai.gpt-5.6-luna via the bedrock-mantle endpoint with the OpenAI Responses API). GPT-5.5, GPT-5.4, and Codex have been available on Bedrock since June 1, 2026. The background is the strategic partnership between Amazon and OpenAI ($50B investment), making AWS the exclusive third-party cloud distribution partner for OpenAI Frontier.
Important for EU customers – and a correction of our earlier assessment: the AWS documentation lists US regions only for the OpenAI models on Bedrock – GPT-5.6 Sol in us-east-1 (N. Virginia) and us-east-2 (Ohio), Terra and Luna additionally in us-west-2 (Oregon); cross-region inference profiles (Geo/Global) are not supported. An EU region (such as eu-central-1 Frankfurt) is not available so far. In addition, the context window on Bedrock is 272k tokens – considerably smaller than the 1.05M tokens via the OpenAI API and Foundry – and Sol can only be run there with the reasoning levels medium and maximum, not the full six-level scale. On the plus side: pricing matches OpenAI first-party rates, prompt caching is supported with a 90 percent discount on cached input, and AWS describes “zero-operator access” at the chip level for the OpenAI models; traffic flagged by classifiers is stored for up to 30 days. For GDPR workloads with EU residency requirements, Bedrock remains off the table for OpenAI models for now – we continuously verify EU region availability.
Integration with CompanyGPT
With CompanyGPT you can use GPT models GDPR-compliant in your company – without your data being used for training.
Our Recommendation
On GPT-6 Astra: watch, do not migrate yet. As of 3 September 2026 the model is unlocked only for the Trusted Access Program, has no EU Data Zone in Foundry, and comes with cyber monitoring that may slow legitimate work. What makes sense is an evaluation slot once the API is open – measured by price per completed task, not per token, because Astra sits at 2.5 times Sol’s promotional price.
With the current pricing and the broad data zone availability, our recommendation has shifted – we now recommend two models with clearly separated roles:
- GPT-5.6 Luna via Microsoft Foundry (Data Zone Standard, EU) – the new default for the majority of enterprise workloads. Luna offers the same context window as the flagship (1.05M tokens), full tool support, and the EU data zone across all nine European Foundry regions – at a fraction of the cost ($0.20/1M input, $1.20/1M output). For chat assistants, classification, extraction, summarization, RAG, and sub-agents this is usually the economically right choice.
- GPT-5.6 Sol via Microsoft Foundry (Data Zone Standard, EU) – the flagship for the most demanding tasks. For deep reasoning chains, complex coding, and autonomous agentic workflows, Sol remains the reference – since 21 August 2026 at a promotional price of $4.00/1M input and $20.00/1M output (at least through 21 November 2026; list price previously $5.00/$30.00). Whether Foundry pricing follows the OpenAI promotional price is something we verify per project. Quota note: tier 5/6 have default quota, below that a quota request is required.
- GPT-5.6 Terra – the middle ground ($2.00/1M input, $12.00/1M output). Useful when Luna’s quality is not enough but Sol is oversized – or when Data Zone Provisioned Managed is required, which is not listed for Luna.
Honest about Luna’s limits: its reasoning depth is below that of Sol and Terra – for demanding analysis and coding tasks it is worth benchmarking against the larger tier. Prompts above 272,000 tokens are billed at 2x input and 1.5x output pricing – as with Sol and Terra – which erodes the cost advantage on very long contexts. And for PTU customers Luna is currently not an option, because it is not listed in the Data Zone Provisioned Managed table. Our recommendation is therefore: build workloads on Luna by default and escalate selectively to Terra or Sol where quality demands it.
Further options:
- GPT-5.5 / GPT-5.4 via Microsoft Foundry (Data Zone Standard) – proven options for existing projects; migrating to GPT-5.6 is usually worthwhile short-term given the identical context window (1.05M tokens) and better benchmarks.
- GPT-5.5 Instant in Microsoft Foundry (
gpt-chat-latest) – still a good choice for chat workloads in EU regions. - Amazon Bedrock – GPT-5.6/5.5/5.4 are GA there, but US regions only, with a reduced 272k context window and only two reasoning levels for Sol. Currently not an option for EU workloads with residency requirements; interesting as a multi-cloud path for US workloads though.
The platform landscape is moving fast right now: Microsoft Foundry and AWS Bedrock keep expanding model and region coverage. We verify the current EU availability per project. For specialized coding tasks, choose GPT-5.3 Codex (also GA on Bedrock). If you are still running gpt-5, o3, or o4-mini, plan the migration now – the deprecation calendar runs through February 2027 (Whisper transcription models). The Assistants API has already been shut down since 26 August 2026.
