State of Models — Aug 2026
The frontier model landscape this month, and which one to actually use for each job.
- Buy once, yours for good
- 17 lessons across 5 sections
- Progress tracking across devices
What you'll learn
- Name the models actually at the frontier this month and what separates them
- Read a benchmark table without being misled by it
- Pick the right model for coding, everyday work, high-volume jobs and long documents
- Work out the real cost of a model, not just its sticker price
- Know which trends this month are load-bearing and which are noise
- Re-run this decision yourself next month when the board changes again
Course content
5 sections · 17 lessons
Where the board standsThe frontier this month, who is on it, and how to read the tables everyone is arguing about.4 lessons
- How to read this courseFree preview2m
Read this lesson
Three rules that make the rest of it useful.
A benchmark number is evidence, not a verdict. Every table below is one lab's measurement of one thing on one day. Two reputable leaderboards can rank the same models differently and both be honest — they are weighting different tasks. Treat a two-point gap as noise and a ten-point gap as signal.
The frontier is a tier, not a rank. The top four or five models are close enough that your prompt, your effort setting, and your harness matter more than which one you picked. The interesting decisions are one tier down, where price differences are 10× rather than 2×.
Nothing here is stable. This course is dated on purpose. The recommendations are good for roughly a month; the method for making them is good indefinitely, which is why the last section teaches the method rather than the answer.
- The frontier tierFree preview4m
Read this lesson
Five models are credibly at the top as of early August 2026.
Model Lab Released Context Notes Claude Opus 5 Anthropic 24 Jul 2026 1M Current top of Artificial Analysis' Intelligence Index Claude Fable 5 Anthropic 2026 1M Highest capability tier; leads SWE-bench Verified GPT-5.6 Sol OpenAI 9 Jul 2026 1.05M Leads Artificial Analysis' coding index Claude Mythos 5 Anthropic 2026 1M Same capability as Fable 5, restricted access programme Kimi K3 Moonshot AI 2026 1M The open-weight model that made the top tier On Artificial Analysis' Intelligence Index at maximum effort, Claude Opus 5 scores 61, Claude Fable 5 60, and GPT-5.6 Sol 59. Those three numbers are within measurement noise of each other. Anyone telling you one of them is decisively smarter than the others is selling something.
Two absences are worth naming. Google has no Pro-tier frontier model on the board this month — Gemini 3.5 Flash shipped at I/O on 19 May 2026, but Gemini 3.5 Pro slipped from June to July and, as of early August, remains unreleased; Google's own position on 21 July was that it is "testing with partners." xAI's Grok 4.5 is present but a tier down, sitting around 11th on llm-stats' composite while being one of the cheapest models in the top group.
- The comparison table5m
- Where the leaderboards disagree3m
What it costsSticker prices, then the three costs that make sticker prices misleading.2 lessons
- The price table4m
- The three costs people forget4m
What changed this month2 lessons
- Releases since July4m
- Trends worth tracking4m
Which model for which jobThe section to actually act on. Each recommendation names a first pick, a runner-up, and the condition under which the runner-up wins.6 lessons
- Best for coding5m
- Best for everyday knowledge work4m
- Best bang for the money4m
- Best for high-volume and structured jobs3m
- Best for long documents and large codebases3m
- Best open-weight models4m
Deciding, and re-deciding3 lessons
- The one-page decision table3m
- What to re-check next month3m
- Sources3m
Requirements
- You have used an AI assistant before — that is the whole prerequisite
Description
Model rankings rot. A comparison written in February is actively misleading by August — prices halve, a lab ships a model that reorders the top five, and the advice everyone repeats is advice about a model that has been superseded twice.
This is the August 2026 edition: a snapshot of the board as it stands, what moved since last month, and — the part that actually matters — which model to reach for on which job. It is deliberately short. You should be able to read the whole thing in under an hour and come out with a decision you can defend.
Everything here is dated 4 August 2026. Numbers are quoted with their source so you can re-check them rather than trust them. Where the field genuinely disagrees, this course says so instead of picking a winner.