Deep Dive10 March 202613 min read

The LLM Landscape: A Practical Guide to Choosing the Right Model

There are now dozens of large language models from OpenAI, Anthropic, Google, Meta, Mistral, and others. This guide cuts through the noise and explains what each does well, what it costs, and how to choose.

The LLM Landscape: A Practical Guide to Choosing the Right Model

A note on currency. This guide was updated in September 2026 and covers releases up to and including 22 September 2026, including Claude Opus 5.5, which Anthropic released the day before this update. Models get renamed, repriced, and replaced on a regular basis. The selection framework below will stay useful longer than any individual price or model name in it. If you are making a purchasing decision, verify current pricing and capabilities directly with the provider.

What an LLM actually does

A large language model is the engine behind most of the AI tools you hear about: chatbots, copilots, content generators, document analysers, code assistants. At its core, an LLM takes text in and produces text out. You give it a question, instruction, or document, and it generates a response based on patterns learned from vast amounts of training data.

For enterprise use, LLMs power three broad categories of capability:

  • Content and communication: Drafting emails, generating reports, writing product descriptions, summarising documents, translating languages
  • Analysis and reasoning: Interpreting data, answering questions about documents, extracting structured information from unstructured text, evaluating options against criteria
  • Automation and workflow: Powering customer-facing chatbots, routing support tickets, classifying inputs, generating code, and acting as the reasoning layer in automated workflows

You do not interact with most LLMs directly. They sit behind products and APIs. When you use ChatGPT, Claude, Gemini, or Copilot, you are using an interface built on top of an LLM. When your developers build AI features into your own products, they call an LLM via an API.

The choice of which LLM to use affects how well those tools perform, what they cost to run, where your data goes, and how locked in you become to a specific vendor. That is what this guide covers.

The current market

If you are running an AI transformation, you will need to choose which models to use. The problem is not a lack of options. There are too many, changing too quickly, with marketing claims that make everything sound equivalent. This guide is designed to help you make practical decisions, not to declare a winner.

Server room housing the GPU infrastructure that powers large language models
Server room housing the GPU infrastructure that powers large language models

The major model families

OpenAI (GPT-6 Astra, Sol, Luna)

OpenAI's current lineup is GPT-6 Astra as the flagship, released 3 September 2026, with GPT-6 Sol and GPT-6 Luna as the mid and budget tiers, both released 22 September 2026 at roughly half the price of the generation they replaced. GPT-5.4, the o3 reasoning models, GPT-5 Mini and Nano, and GPT-4.1 are all still listed and still usable, but none of them are the current generation anymore.

ModelInput (per 1M tokens)Output (per 1M tokens)Context Window
GPT-6 Astra$10 (up to 272K), $20 beyond$50 (up to 272K), $75 beyond1.05M (922K max input)
GPT-6 Sol$2 ($4 beyond 272K)$10 ($18 beyond 272K)1.05M
GPT-6 Luna$0.10 ($0.20 beyond 272K)$0.50 ($0.75 beyond 272K)1.05M
o3 (legacy reasoning)$2.00$8.00-
GPT-4.1 (legacy)$2.00$8.001M

See OpenAI's full pricing table and the GPT-6 Astra model card for the current detail. Astra is the first OpenAI model to hit "Critical" on the company's own cybersecurity capability scale, and OpenAI reports it saturating benchmarks such as FrontierMath Tier 4 (98%) and ARC-AGI-3 (99.9%), treat those as vendor-reported figures rather than independent results.

Best for: Astra for the hardest reasoning and coding work, Sol for general-purpose business tasks, Luna for high-volume cost-sensitive applications, GPT-4.1 where you specifically need its 1M context window at legacy pricing.

Watch out for: the model naming problem has got worse, not better. GPT-6 (Astra, Sol, Luna), GPT-5.6, GPT-5.5, GPT-5.4, GPT-5, GPT-4.1, and the o-series all sit on the pricing page at once. Check which generation you are actually buying before you commit a workload to it.

Available via: OpenAI API, ChatGPT, and Microsoft Foundry, the renamed and now multi-vendor successor to Azure AI Foundry (see below).

Anthropic Claude (Fable 5.1, Opus 5.5, Sonnet 5, Haiku 4.5)

Anthropic split its top tier this year. Claude Fable 5.1, and a restricted sibling called Claude Mythos 5.1, now sit above Opus for the hardest reasoning and long-horizon agentic work, both released 1 September 2026. Claude Opus 5.5 followed on 22 September 2026, the most recent frontier release from any provider as of this update, and Anthropic now positions it as the default choice for most workloads rather than a specialist tool reserved for hard problems. Anthropic's own model comparison guidance puts it plainly: start with Opus 5.5 for most workloads, and move to Fable 5.1 for demanding reasoning and agentic work, or when your evals on Opus 5.5 at higher effort still fall short.

ModelInput (per 1M tokens)Output (per 1M tokens)Context Window
Claude Fable 5.1$10$501M
Claude Opus 5.5$4$201M
Claude Sonnet 5$2$101M
Claude Haiku 4.5$1$5200K

Anthropic's own Opus 5.5 launch page says it performs at the level of Fable 5.1 on most work at 40% lower cost than its predecessor, and it is itself cheaper per token than the Opus model it replaced. Sonnet 5, released 30 June 2026, is described as the most agentic Sonnet model yet, with improved browser and terminal tool use. Haiku 4.5, unchanged since October 2025, delivers around 90% of Sonnet 4.5's performance at a third of the cost and more than twice the speed.

Best for: Fable 5.1 for the hardest reasoning and agentic work, Opus 5.5 as the general recommendation for most business tasks, Sonnet 5 for agentic and tool-heavy workflows at a lower price point, Haiku 4.5 for high-volume chat, customer service, and orchestration work.

Watch out for: this inverts the old advice to treat Opus as the expensive specialist and Sonnet as the generalist. Opus is no longer Anthropic's most capable or most expensive tier. 1M context is now standard on Fable 5.1, Opus 5.5, and Sonnet 5, not a paid beta; only Haiku 4.5 stays at 200K.

Available via: Anthropic API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

Google Gemini (3.1 Pro, 3.8 Flash)

Google has moved on to the Gemini 3 generation. Gemini 3.1 Pro, still labelled a preview, is the current flagship reasoning tier, and Gemini 3.8 Flash, with Gemini 3.5 Flash-Lite as the cut-price option, is the current fast tier. Gemini 2.5 Pro and Flash remain available to existing users, but Google is steering new projects to Gemini 3.

ModelInput (per 1M tokens)Output (per 1M tokens)
Gemini 3.1 Pro (Preview)$2.00 (up to 200K), $4.00 beyond$12.00 (up to 200K), $18.00 beyond
Gemini 3.8 Flash$0.75, rising to $1.50 from 1 Jan 2027$3.75, rising to $7.50 from 1 Jan 2027
Gemini 3.5 Flash-Lite$0.30$2.50

Full detail on Google's current pricing page.

Best for: multimodal tasks involving video, image, or audio, cost-sensitive high-volume applications (Flash-Lite), Google Workspace-integrated businesses. Google's own documentation did not state a clear context window figure for the two current models at the time of writing, so check that directly before you plan around it.

Watch out for: Gemini 3.1 Pro is still a preview release, and it is more expensive per token than the Gemini 2.5 Pro figures this article previously quoted. Long context still increases latency and cost. The enterprise ecosystem remains less established than OpenAI's or Anthropic's.

Available via: Google AI Studio, Google Cloud Vertex AI.

Meta Llama 4 (Scout, Maverick)

Meta's open-weight models remain live and unchanged since March. Llama 4 Scout still carries a 10M token context window, the largest in the industry, and Llama 4 Maverick sits at 1M. Licensing is unchanged too: the Llama 4 Community Licence still restricts companies with 700M+ monthly active users.

ModelPricingContext Window
Llama 4 Scout (17B active)Free (open-weight)10M
Llama 4 Maverick (17B active)Free (open-weight)1M

Meta's own developer site now foregrounds a new proprietary model family, Muse, built by its newly formed Superintelligence Labs, available only through the Meta AI app and website and a private API preview. Meta's public statement on Llama's future was noncommittal: existing Llama models stay open source, with no commitment given on future versions. If you are selecting Llama for a multi-year roadmap, that uncertainty is a genuine input to the decision, separate from anything about the current Scout and Maverick models themselves.

Best for: Cost-conscious enterprises with ML capability, fine-tuning for domain-specific tasks, maximum data privacy through on-premises deployment.

Watch out for: "Open-weight" is not the same as open source. You need infrastructure and ML expertise to self-host, or you can access these models via cloud providers at cost. There is no official enterprise support or SLA from Meta, and Meta's own roadmap commitment beyond the current models is now unclear.

Available via: Self-hosted, Amazon Bedrock, Microsoft Foundry, Together AI, and others.

Mistral AI (Large 3, Medium 3.5, Small 4)

The leading European AI company, based in Paris, has moved its recommended lineup on from Large/Medium 3/Small 3.2 to Large 3, Medium 3.5, and Small 4. Mistral's own pricing page confirms Large 3 at $0.50 input and $1.50 output per 1M tokens. Medium 3.5 and Small 4 pricing did not render on that page directly at the time of writing, so check current pricing with the provider before you budget against either.

EU AI Act enforcement powers for general-purpose AI model providers, covering documentation requests, evaluations, and fines of up to €15M or 3% of global turnover, took effect on 2 August 2026, alongside Article 50 transparency obligations. That is live ground now rather than a future consideration, and it strengthens Mistral's EU-jurisdiction pitch for anyone weighing data sovereignty. See the EU AI Act implementation timeline for the detail, and our own guide to the EU AI Act and GDPR for what the obligations mean for your business.

Best for: European businesses with GDPR or data sovereignty requirements, self-hosted enterprise use, organisations needing EU-jurisdictional guarantees.

Watch out for: a smaller ecosystem than the US providers, and a narrowing but still real performance gap versus frontier models on the hardest tasks.

Available via: La Plateforme (Mistral API), Microsoft Foundry, Amazon Bedrock, self-hosted.

Amazon Nova

AWS announced a Nova 2 family (Lite, Pro, Sonic, Omni) alongside Nova Forge and Nova Act. The original Nova Pro, Lite, and Micro pricing this article quoted in March has been superseded, but AWS's own pricing pages did not return consistent per-token figures for Nova 2 at the time of writing, so check current pricing directly with AWS before you budget against it.

Best for: AWS-native enterprises, multimodal applications, companies wanting to build custom models via Nova Forge.

Watch out for: an AWS-locked ecosystem, less proven track record on frontier benchmarks than the major labs, a smaller developer community.

Available via: Amazon Bedrock only.

DeepSeek

DeepSeek has moved on from R1 and V3.2 to DeepSeek-V4.1-Flash and DeepSeek-V4-Pro. Pricing now runs on a peak and off-peak schedule: costs roughly double during Asia business hours (01:00-04:00 and 06:00-10:00 UTC, weekdays, excluding Chinese public holidays) compared with the rest of the day. That is new since March, and it changes how you should schedule batch workloads: run them off-peak and the bill roughly halves.

ModelInput, off-peak / peakOutput, off-peak / peakContext
DeepSeek-V4.1-Flash$0.15 / $0.30$0.60 / $1.201M
DeepSeek-V4-Pro$0.66 / $1.32$1.98 / $3.96Not stated by DeepSeek

Full pricing, including the cache-hit discounts, on DeepSeek's pricing page.

Best for: coding tasks, cost-sensitive applications, research and analysis, particularly where batch work can be scheduled into off-peak hours.

Watch out for: DeepSeek is a Chinese company, which remains a material consideration for some enterprises on data sovereignty and compliance grounds, more so now that EU AI Act GPAI enforcement is live. It is available on Microsoft Foundry for those wanting Western infrastructure.

Available via: DeepSeek API, Microsoft Foundry.

xAI Grok

Built by Elon Musk's xAI with deep integration into the X platform. Grok 4.7 is the current flagship. Grok 5 has not shipped: Musk's own roadmap, stated 14 September 2026, puts point releases 4.8 and 4.9 ahead of it, after missed windows in late 2025 and the first half of 2026.

ModelInput, under 200K / at or aboveOutput, under 200K / at or aboveContext Window
Grok 4.7$2.00 / $4.00$6.00 / $12.00500K
Grok 4.3 / 4.20$1.25 / $2.50$2.50 / $5.001M

Full model and pricing detail on xAI's model documentation.

Best for: social media monitoring and sentiment analysis, market intelligence, real-time news tracking, brand monitoring.

Watch out for: tied to the X platform with limited interoperability. Claims about coding performance and enterprise certifications relative to competitors were not independently re-checked for this update, so treat those as directional rather than confirmed. Less suitable for core business operations.

Available via: xAI API, X platform.

How to choose: the decision framework

The right model depends on your use case, not on benchmark scores. Here is how to think about it.

Match the model to the task

Most businesses do not need one model. They need different models for different jobs:

  • Customer-facing chatbot: Prioritise speed and cost. Gemini 3.5 Flash-Lite, GPT-6 Luna, Claude Haiku 4.5, or Amazon Nova.
  • Document analysis and summarisation: Prioritise context window and recall. Gemini 3.1 Pro or Claude Sonnet 5/Opus 5.5 with extended context.
  • Complex reasoning and strategy: Prioritise accuracy over speed. Claude Fable 5.1, GPT-6 Astra, or Gemini 3.1 Pro.
  • Content generation at scale: Prioritise quality and cost balance. Claude Opus 5.5, GPT-6 Sol, or Gemini 3.1 Pro.
  • Coding and software engineering: Claude Opus 5.5 or Fable 5.1, DeepSeek V4 Pro, or GPT-6 Astra.
  • Data privacy-sensitive tasks: Self-host Llama 4 or Mistral, or use EU-jurisdictional Mistral.

Consider your existing infrastructure

Cloud platforms have become multi-vendor by default since March. If you are already on AWS, Nova and Bedrock, which now also host Anthropic, Meta, and Mistral models, are the path of least resistance. Microsoft Foundry, renamed from Azure AI Foundry on 1 January 2026, now serves OpenAI, Anthropic, Google, Meta, and xAI models from the same surface rather than OpenAI alone. On Google Cloud, Gemini integrates natively. Switching cloud providers purely to access a specific model is less important than it was in March, now that the big three have opened up to each other's models.

Do not over-index on benchmarks

Traditional benchmarks like MMLU and HumanEval are saturated, and the newest frontier evals are heading the same way. OpenAI cites GPT-6 Astra scoring 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3, its own figures, not independently verified. The differences that matter for business use are in instruction-following, consistency, latency, and how well a model handles your specific domain. A benchmark score does not tell you whether the model will write good emails for your customer service team.

Start with the cheapest model that works

A common mistake is defaulting to the most capable, and most expensive, model for every task. GPT-6 Luna at $0.10/1M input tokens can handle classification, routing, and simple summarisation perfectly well, a generation ahead of the budget models this article previously recommended, at a similar price point. Reserve frontier models like Fable 5.1 or GPT-6 Astra for tasks that genuinely need them.

Plan for model switching

Do not build your entire stack around a single provider. Use abstraction layers (LangChain, LiteLLM, or your own routing logic) that let you swap models without rewriting your application. Models improve, pricing changes, and new entrants appear regularly, as the past six months have shown across every provider in this guide. The organisation that can switch models quickly has a structural advantage.

Open source vs proprietary: the 2026 reality

The gap between open-weight and proprietary models has closed dramatically over the past two years. The decision is not purely about capability:

  • Choose open-weight when data privacy is paramount, you need fine-tuning for a specific domain, you want to avoid vendor lock-in, or you have the ML engineering capability to self-host.
  • Choose proprietary when you need enterprise support and SLAs, you want managed infrastructure, you need the absolute frontier capability, or you lack the team to self-host.

Weigh Meta's uncertain commitment to Llama's future versions into this decision if you are choosing open-weight specifically to protect a long-term roadmap: Mistral, the other major open-weight anchor named in this guide, has not signalled any comparable shift.

Most enterprises will use both. A proprietary model for customer-facing applications where reliability is critical, and open-weight models for internal tools where cost and privacy take priority.

The cost reality

LLM pricing now ranges from around $0.10/1M tokens (GPT-6 Luna input) to $75/1M tokens (GPT-6 Astra output on longer prompts), roughly a 750x range within OpenAI's own lineup alone, before you compare across providers. DeepSeek's new peak and off-peak pricing adds a time dimension to that spread that did not exist in March.

The smart approach is tiered: route simple, high-volume tasks to cheap, fast models, and reserve frontier models for the queries that genuinely need them. Most user interactions do not need the most capable model available.

What this means for your programme

If you are early in your AI transformation, do not spend weeks agonising over model selection. Pick a reputable provider that fits your cloud infrastructure, start with a mid-tier model such as Claude Opus 5.5, GPT-6 Sol, or Gemini 3.1 Pro, and optimise later based on real usage data.

The model you choose today will not be the model you use in twelve months. Every named flagship in this guide's March version has been superseded at least once, in some cases twice, in six months. The more valuable investment is building your applications in a way that makes switching easy, developing prompt engineering and evaluation frameworks, and understanding your actual usage patterns so you can optimise cost and performance over time.

Section 5: Technology Landscape covers how to evaluate and select AI tools for your programme. Section 7: Experimentation provides the framework for testing different approaches before committing at scale.

AI Transformation Playbook

Ready to put this into practice?

The playbook gives you 100+ practical tools, checklists, templates, and facilitation guides for every stage of an AI transformation programme.