AI in Review — August 2026: A Price U-Turn, the Open-Weight Wave and an Enforceable EU AI Act
August 2026 wasn’t a month of one big headline — it was a month of many small shifts, and together they add up to a clear picture. Anthropic walks back an announced price increase, a whole wave of open models pushes the feature-set expectation upward, Google cleans up how AI visibility is measured, and in the EU AI Act announced law becomes enforceable law. Here’s what sticks — and what it means for day-to-day work with AI.
August 2026 in five points
- Price U-turn: Anthropic cancels the planned Sonnet 5 increase — 2/10 US dollars stay permanent.
- Open-weight wave: GLM-5.3, DeepSeek V4-Pro and Qwen3.8 bring near-frontier performance under MIT/Apache licenses.
- New frontier models: Qwen3.8-Max (2.4-trillion MoE) and Meta Muse Spark 1.2 (agentic coding) keep the fast cadence going.
- AI search: Google AI Mode runs on Gemini 3.7 Flash from 14 August — and only visible links count as an impression.
- Law: Since 2 August the EU AI Act’s GPAI and transparency obligations are enforceable; high-risk is pushed to December 2027.
Pricing: Anthropic pulls the increase
The most practically relevant news of the month was a reversal. At launch on 1 July, Anthropic had explicitly declared the price of Claude Sonnet 5 an introductory price — 2 US dollars per million input tokens, 10 US dollars per million output tokens, valid through the end of August. From 1 September the regular price of 3/15 US dollars was meant to kick in, a 50-percent surcharge.
On 10/11 August, that increase was cancelled: the introductory price becomes permanent. For anyone running Sonnet 5 at volume — agent runs, coding pipelines, bulk content — a cost bump that was already budgeted for simply drops away. The real lesson lies less in the single price than in the volatility of the announcement: within a few weeks, “introductory price with an end date” turned into “permanent price”. If you plan model costs, treat announced price changes as provisional — in both directions. Only the billed price is reliable.
It’s reasonable to assume competitive pressure in the mid-tier played the deciding role. Exactly where Sonnet 5 sits, cheap, capable models from Google, DeepSeek and the open camp are pushing in — and those held their prices steady.
Models: the open-weight wave becomes a pattern
If one theme defines August, it’s this one. Between July and the end of August, an unusually dense run of open models appeared — models whose weights are freely downloadable and self-hostable: Z.AI GLM-5.3, DeepSeek V4-Pro (GA), Alibaba’s Qwen3.8 line, plus open models from Meta and NVIDIA and an announced new open-weight family from Mistral.
What stands out isn’t the volume alone but the combination: permissive licenses (MIT, Apache 2.0) meeting numbers that, per their makers, come close to the proprietary frontier. That shifts the old story — “the strongest models are closed” — noticeably. For agencies this is strategically relevant: a self-hosted MIT model keeps client data on your own network and is immune to the big providers’ price and plan changes. The cost is operational overhead — GPU capacity, maintenance, ops — which has to be counted honestly against the savings.
Two names from the same movement deserve their own look:
- Qwen3.8-Max (2 August) — Alibaba’s proprietary flagship: a Mixture-of-Experts model with a claimed 2.4 trillion parameters, a 1-million-token context and native image and video input, but only via a closed API. The parameter count is a vendor figure; what drives cost and latency is the number of active parameters per token, not the total size.
- Meta Muse Spark 1.2 (5 August) — a coding model built explicitly for agentic work at the repository level: 1M context, co-trained with a code agent, parallel tool calls. The core is a training philosophy — the line between “model” and “agent” blurs because the interplay is baked in during training.
The same caution applies to all three: the figures come from release trackers and vendor communication. Performance only becomes reliable with independent benchmarks and a test on your own task.
AI search: Google clarifies the measurement
For anyone watching AI-search visibility, 14 August was the more important date. Google communicated two things at once: the AI Mode starts answering a share of queries with Gemini 3.7 Flash — and, more consequential in practice, in AI Overviews and AI Mode only visible links count as an impression. Plain brand mentions without a linked click path do not.
This is where GEO reporting has to get honest. The Search Console impression is a narrower metric than what many GEO tools report as “visibility” or “share of voice”. The two don’t measure the same thing: a tool can count a brand mention that never shows up as an impression in the Search Console. Put both numbers side by side without naming the difference and you produce reports that don’t add up internally. The consequence for agency work: keep linked impressions (GSC) and brand mentions (GEO tools) cleanly separated and label both.
Law: the EU AI Act becomes enforceable
On 2 August, another stage of the EU AI Act went live: the GPAI and transparency obligations from Articles 5 and 50 are now applicable and enforceable. An announced roadmap becomes applicable law with a sanctions framework. At the same time, the EU pushed the toughest obligations for high-risk systems to December 2027.
The shift isn’t in the wording — that was settled long ago — but in the status change from “announced” to “enforceable”. For content production that means: the pressure to act today is on labeling synthetic media (AI images, audio, video) and on disclosing chatbots, not on elaborate risk assessments. The exemption for classic marketing and SEO copy remains, but it’s narrow — content on matters of public interest should be checked separately. (Not legal advice; the concrete obligations depend on the individual case.)
And at boostN?
Two of August’s movements land right on what boostN is built around.
The price volatility on Sonnet 5 and the open-weight wave are both arguments against committing to a single model. That’s exactly what our model choice instead of model lock-in is for: one agent, one execution pipeline — and underneath it the model that fits the task and the price right now. Tie yourself to a single provider and you’re tied to its pricing and quality curve. A pipeline that knows several models lets you switch without rebuilding the workflow. Google Gemini is already live as an agent, and the Mistral integration is close to shipping.
And the open-weight debate hits the same nerve as our work on token and cost discipline: it isn’t the biggest model that wins, but the setup that makes the right choice per task — and counts the real costs (active parameters, hosting, ops) honestly.
What to take away from August
- Redo the math: the Sonnet 5 September surcharge is gone — remove the line from your budget and treat future price announcements as provisional until they’re billed.
- Take open models seriously: for privacy-critical tasks, evaluate a 27B/30B dense model under an Apache 2.0 license as a realistic self-hosting entry point — license first, size second.
- Untangle GEO reporting: separate linked GSC impressions from GEO-tool brand mentions and label both so no one reads them as the same figure.
- Introduce labeling: if you ship AI media or run chatbots, disclosure is now compliance, not a nice-to-have.
- Distrust benchmarks: treat vendor figures as a marketing signal and test candidates on real tasks before you switch.
The bottom line: August confirmed the direction of travel — more model choice, shakier price commitments, more honest visibility measurement, and a rulebook that now has teeth. Whoever stays flexible rather than locked in comes through best.