Sol, Terra, Luna, Fable, Opus, Sonnet, Flash, Scout, Maverick. The names of AI models mean little to a regular user, and yet it matters which one you use: the model determines how good the answer is, how much text you can feed it and what you pay for it.
This page is meant as a reference. Below you'll find every model that matters as of August 2026, grouped by provider, with the release date, the context window and the price per million tokens. You don't need to memorize them. But when you come across a model name somewhere, you can look up here who makes it, what it's meant for and whether it fits your work.
Model, app and subscription are three different things
Most of the confusion comes from three things getting mixed up. The model is the engine: GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.1 Pro. The app is what you open: ChatGPT, Claude, Gemini. The subscription determines which models you can use in that app and how often. So one app contains several models, and the same models often show up in other apps too. Which subscription unlocks which models is laid out in our comparison of AI subscriptions.
You'll run into four terms in every overview.
Tokens. The unit a language model counts in. A token is a piece of a word; as a rule of thumb, a thousand tokens is about 750 words. Providers charge per million tokens, separately for what you put in (input) and what comes out (output).
Context window. How much text the model can keep in view at once: your question, the whole conversation so far, your attachments and the answer combined. A window of a million tokens comes down to several thousand pages. If a conversation goes over it, the beginning drops out and the model starts forgetting things.
Open weights. The model can be downloaded as a file, so you can run it yourself on your own hardware or with a hosting provider of your choice. With closed models that isn't possible: they only run at the provider itself. Open weights say something about freedom and privacy, not automatically about quality.
Reasoning. A mode in which the model first works out step by step how to reach an answer before it responds. That takes longer and costs more, but gives noticeably better results in logic, math and code. For simple questions it's a waste of time.
| Model | From | Released | Context | Dutch | Best for |
|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 9 July 2026 | 1,05M tokens | Good | Heavy analysis, agents and professional work |
| GPT-5.6 Terra | OpenAI | 9 July 2026 | 1,05M tokens | Good | Daily work at a lower price |
| GPT-5.6 Luna | OpenAI | 9 July 2026 | 1,05M tokens | Fair | Quick questions and high volumes |
| Claude Fable 5 | Anthropic | 9 June 2026 | 1M tokens | Excellent | Hard, long-running work where quality beats cost |
| Claude Opus 5 | Anthropic | June 2026 | 1M tokens | Excellent | Complex reasoning tasks and long documents |
| Claude Sonnet 5 | Anthropic | 30 June 2026 | 1M tokens | Excellent | Daily writing, coding and agents |
| Gemini 3.5 Flash | 19 May 2026 | 1,05M tokens | Good | Fast multimodal work and large volumes | |
| Gemini 3.1 Pro | preview | 1M tokens | Good | Research and document analysis | |
| Grok 4.5 | xAI | 8 July 2026 | 500K tokens | Fair | Programming, engineering and current information |
| Mistral Medium 3.5open | Mistral AI | April 2026 | 256K tokens | Good | European deployments and agent work |
| Mistral Small 4open | Mistral AI | March 2026 | 256K tokens | Fair | Cheap instructing, reasoning and coding |
| DeepSeek V4 Proopen | DeepSeek | April 2026 | 1M tokens | Poor | Cheap reasoning and programming |
| Qwen 3.8 Max | Alibaba | 3 August 2026 | 1M tokens | Poor | Multimodal agents and code |
| Llama 4 Scout / Maverickopen | Meta | April 2025 | 10M tokens | Fair | Self-hosting with extremely long context |
| Muse Glimmeropen | Meta | August 2026 | 256K tokens | Poor | Local agents on your own laptop |
OpenAI: GPT-5.6 in three flavors
On 9 July 2026, OpenAI released one generation in three sizes. They share the same context window of 1.05 million tokens and differ mainly in how heavy and how expensive they are.
| Model | Released | Context window | API input–output per million tokens |
|---|---|---|---|
| GPT-5.6 Sol | 9 July 2026 | 1.05 million | $5 / $30 |
| GPT-5.6 Terra | 9 July 2026 | 1.05 million | $1.25 / $10 |
| GPT-5.6 Luna | 9 July 2026 | 1.05 million | $0.10 / $0.80 |
Sol is the flagship. Since the update of 6 August 2026 it handles facts better, and ChatGPT has an effort slider that lets you decide how long the model may think about an answer. Terra is the mid-range model and roughly as strong as last year's flagship, GPT-5.5, at a fraction of the price: on 30 July 2026 another 20% came off. Luna is the cheapest and, since 6 August 2026, the default model for free users and for the Go plan. Luna's price dropped by 80% on 30 July 2026.
That free tier is now serious: unlimited text chat on Luna plus a Think button for harder questions. The limits are now on uploads, images and tools, no longer on the number of messages. Paying mainly makes sense if you use the top model a lot — Plus gives you 160 messages per three hours on the heaviest model, compared with historically ten messages per five hours on the free plan.
For images and video, OpenAI has GPT Image 2 and Sora 2, both included in ChatGPT Plus. That's the clearest argument for sticking with OpenAI: no other provider puts video generation in a subscription at this price point.
Anthropic: Claude Fable, Opus and Sonnet
Anthropic uses the same three-way split, but with its own emphasis: all three models have a context window of a million tokens, and the focus is on text, code and reliability.
| Model | Released | Context window | API input–output per million tokens |
|---|---|---|---|
| Claude Fable 5 | 9 June 2026 | 1 million | $10 / $50 |
| Claude Opus 5 | June 2026 | 1 million | $5 / $25 |
| Claude Sonnet 5 | 30 June 2026 | 1 million | $2 / $10 |
Fable 5 is the heaviest and by far the most expensive model on the market. Note: in the Claude apps, Fable 5 is not included in the Pro plan; you pay for it with prepaid credits. Opus 5 is the reasoning model for hard problems. Sonnet 5 is the workhorse for daily use and also the model you get in the free version of Claude — which makes free Claude one of the strongest free options around.
What else sets Claude apart: Claude Pro includes Claude Code and Claude Cowork, and training on your conversations is off by default on the paid consumer plans. If you do turn training on, your data is kept longer. More on how Anthropic works and why the model writes such good Dutch in Claude AI explained.
Google: Gemini 3.5 Flash and 3.1 Pro
Google works on two tracks: a fast, general-purpose model and a heavier model for thinking work.
| Model | Released | Context window | API input–output per million tokens |
|---|---|---|---|
| Gemini 3.5 Flash | 19 May 2026 | 1.05 million | $1.50 / $9 |
| Gemini 3.1 Pro | preview | 1 million | $2 / $12 |
Gemini 3.5 Flash is the stable workhorse and multimodal: it handles text, images and documents in the same conversation. Gemini 3.1 Pro is still in preview and is the heavy model in Google AI Plus and Google AI Pro.
The reason many Dutch users end up with Google isn't the model but what comes with it. Google AI Pro (€21.99 per month) gives you a context window of a million tokens, Deep Research, NotebookLM and 2 TB of storage. Google AI Plus costs €7.99 and comes with 200 GB of storage. For images and video, Google has Nano Banana Pro and Veo 3.1. How that subscription is put together — and why you may already have it without knowing — is covered in Gemini AI explained.
xAI and Mistral: Grok and the European option
Grok 4.5 from xAI came out on 8 July 2026, has a context window of 500,000 tokens and costs $2 input and $6 output per million tokens. It's strong at code and the only model with direct access to live data from X. Its context window is smaller than the competition's, and at €31.53 per month its subscription is the most expensive standard plan among the major providers.
Mistral is the French provider and relevant to Dutch users for a different reason: processing happens in the EU, in France and Sweden. All three models have a context window of 256,000 tokens.
| Model | Released | License | API input–output per million tokens |
|---|---|---|---|
| Mistral Medium 3.5 | April 2026 | open weights, modified MIT | $1.50 / $7.50 |
| Mistral Large 3 | December 2025 | Apache 2.0 | $0.50 / $1.50 |
| Mistral Small 4 | March 2026 | Apache 2.0 | $0.15 / $0.60 |
Medium 3.5 is the model you use in Le Chat Pro (€15.76 per month), including a mode without telemetry. Small 4 is multimodal and so cheap that it's often the most sensible starting point for your own applications. To be fair: on the heaviest tasks, the Mistral models don't reach the level of Sol or Opus 5. You choose them for where your data lives, not for top performance.
The open models: DeepSeek, Qwen, Llama and the rest
You rarely use this group through a consumer subscription. You come across them when you build something yourself or want to run something on your own hardware. Prices then depend on your hosting provider, not on the maker.
| Model | Maker | Released | Context window | What it's for |
|---|---|---|---|---|
| DeepSeek V4 Pro / Flash | DeepSeek | April 2026 | 1 million | Very cheap: $0.435 / $0.87 and $0.14 / $0.28 |
| Qwen 3.8 Max | Alibaba | 3 August 2026 | 1 million | Newest Alibaba model, tiered pricing |
| Llama 4 Scout / Maverick | Meta | April 2025 | up to 10 million | The standard for self-hosting |
| Muse Glimmer | Meta | August 2026 | 256,000 | 30B, Apache 2.0, runs on a single consumer GPU |
| GLM-5.2 | Z.ai | 16 June 2026 | 1 million | Open engineering model |
| Kimi K2.7 Code | Moonshot | June 2026 | 256,000 | Long runs of coding agents |
| MiniMax M3 | MiniMax | 1 June 2026 | up to 1 million | Efficient multimodal, $0.30 / $1.20 |
Two notes on this. With up to ten million tokens, Llama 4 Scout has by far the largest context window of them all, which makes it attractive for work with huge document collections. And Muse Glimmer stands out because, with 30 billion parameters under Apache 2.0, it runs on a single consumer GPU: your AI conversations then never leave your own computer.
If you want to build something with these models yourself, the prices per million tokens above are your starting point. In building an AI chatbot in Dutch we work out what that costs in practice.
Which model should you use for what?
The short summary of everything above, by task.
| Task | Strongest choice | Cheap alternative |
|---|---|---|
| Writing Dutch texts | Claude Opus 5 or Sonnet 5 | Free Claude (Sonnet 5) |
| Programming | Claude Opus 5, Grok 4.5 | Kimi K2.7 Code |
| Long documents and PDFs | Gemini 3.1 Pro | Gemini 3.5 Flash |
| Everyday questions | GPT-5.6 Terra | GPT-5.6 Luna (free) |
| Image generation | GPT Image 2, Nano Banana Pro | — |
| Video generation | Sora 2, Veo 3.1 | — |
| EU processing | Mistral Medium 3.5 | Mistral Small 4 |
| Running it yourself | Llama 4 | Muse Glimmer |
Two tasks fall outside this list. If you want answers with sources cited, a model isn't enough and you need an AI search engine like the one in Perplexity AI: the AI that actually cites its sources. And if you're specifically looking for the model that writes the best Dutch, we tested that separately in which AI speaks the best Dutch.
| Provider | Plan | List price | Per month in € | Per year |
|---|---|---|---|---|
| GoedkopeAIAll models below | Plus | €12.95 incl. VAT | €12.95 | €155 |
| ChatGPTOpenAI | Plus | $20 excl. VAT | €21.02 | €252 |
| ClaudeAnthropic | Pro | $20 excl. VAT | €21.02 | €252 |
| GeminiGoogle | Google AI Pro | €21.99 incl. VAT | €21.99 | €264 |
If you want to pick the best model for each task, you quickly end up tied to two or three subscriptions. How that choice plays out per task is worked out in GPT-5.6 vs Claude 5 vs Gemini 3.5.
Conclusion
Frequently asked questions
How many AI models are there, really?
There are hundreds of models, but in practice about twenty matter. They come from ten providers: OpenAI, Anthropic, Google, xAI, Mistral, DeepSeek, Alibaba, Meta, Z.ai and Moonshot. The rest are variants, older versions or models built on top of one of these.
What is the difference between ChatGPT and GPT-5.6?
ChatGPT is the app you type into, GPT-5.6 is the model that comes up with the answer. The same goes for Claude (app) and Claude Sonnet 5 (model), and for Gemini (app) and Gemini 3.1 Pro (model). One app usually contains several models, and you can often choose which one you use.
What does a context window of 1 million tokens mean?
The context window is the amount of text a model can keep in view at once: your question, the earlier conversation, your attachments and the answer combined. A token is roughly three quarters of a word, so a million tokens comes down to several thousand pages of text. Go over it and the model forgets the start of the conversation.
Which AI model is the best in 2026?
No model wins at everything. Claude Opus 5 and Sonnet 5 write the best Dutch, GPT-5.6 Sol has the broadest package around it, Gemini 3.1 Pro is strong at documents and research and Grok 4.5 is good at code. For simple questions, cheap models like GPT-5.6 Luna are now more than good enough.