Skip to content
BlackOak AgencyBlackOak.agency

// AI RADAR

AI Radar

What is shipping in AI across models, agents, video and capital — with what it is, why it matters and what to do with it on Monday. Below the brief sits the landscape: prices, tools and numbers that hold up longer.

The week in artificial intelligence

The rules landed before the habit did.

Since late July the frontier moved twice — Fable 5.1 on 1 September, GPT-6 Astra two days later, both at $10 in and $50 out — but the news that reaches your Monday is older than either: since 2 August, AI-made content has to mark itself as such. Here is what changed in six weeks, and what to do about it.

August–September 2026 — the release calendar

← drag to scroll

  1. EU / CaliforniëTransparency rules take effect02 AUG
  2. AlibabaQwen3.8-Max generally available03 AUG
  3. OpenAIGPT-5.6-Cyber10 AUG
  4. xAIGrok 4.612 AUG
  5. DeepSeekV4 Pro out of preview12 AUG
  6. GoogleGemini 3.7 Flash13 AUG
  7. AnthropicWatermark in all output14 AUG
  8. SalesforceClaudeforce26 AUG
  9. AnthropicModel Hardware Standard27 AUG
  10. AnthropicFable 5.1 and Mythos 5.101 SEP
  11. OpenAIGPT-6 Astra03 SEP
  12. OpenAISora API shuts down24 SEP

The brief

Six developments that change your stack

By date, with what happened, why it matters and what to do next.

02.08.2026Rules

From now on your content has to say a machine made it

On 2 August, Article 50 of the EU AI Act and California’s AI Transparency Act (SB 942, as amended by AB 853) became enforceable. Both require machine-readable marking on AI-generated image, audio, video and — for matters of public interest — text. The European Commission is working on a standardised label distinguishing "fully AI-generated" from "AI-assisted". Anthropic turned it into policy immediately: every Claude product has marked its output since 2 August, worldwide rather than only in the EU. Penalties run to €15 million or 3% of global annual turnover.

Why it matters

This is the first time using AI in your work becomes an administrative obligation rather than a choice. For an agency it means the question "was this made with AI?" is no longer a conversation with the client but a field in the deliverable. And because the marking is machine-readable, a client can check it afterwards without asking you.

What next

Do an inventory this month: which deliverables are fully generated, which AI-assisted, which made by hand. Put the marking in the production process, not in the small print of a proposal — a label that first appears on the invoice arrives too late.

13.08.2026Models

Gemini 3.7 Flash doubles on office work — and doubles in price on 1 January

Google shipped Gemini 3.7 Flash on 13 August, twenty-three days after 3.6 Flash. On AutomationBench, which measures real business workflows, it moves from 17.0% to 30.4%; on DeepSWE v1.1 from 49.0% to 65.3%. Introductory pricing is $0.75 per million input tokens and $3.75 per million output, with a context window over a million tokens. That rate expires on 31 December 2026: from 1 January 2027 it is $1.50 and $7.50. The promised 3.5 Pro still has not appeared; Google is now training Gemini 4.

Why it matters

The workhorse tier is where an agency’s margin lives — that is where the volume runs. A doubling on a benchmark that mimics office processes, at under a dollar per million input tokens, shifts which tasks you still do by hand at all. The catch is in the small print: budgeting on $0.75 today is budgeting on a price that disappears in four months.

What next

Move bulk work here — extraction, classification, first drafts — but run your business case on the 2027 rate. Anything that only works at the introductory price does not work.

26.08.2026Agents

Claudeforce: the CRM moves into the agent, not the other way round

Salesforce and Anthropic announced Claudeforce on 26 August. Two directions at once: Claude becomes the reasoning model inside Salesforce, and Salesforce becomes a plugin inside Claude, with 37 prebuilt sales skills — from meeting prep and deal health checks to pipeline reviews. The plugin is with pilot customers, open beta in September; more skills follow in late 2026. It is the first time Salesforce has attached its "force" suffix to another company’s product.

Why it matters

Until now the agent sat beside the stack: you gave it access to your systems. Here the customer context sits inside the conversation, with the CRM’s governance wrapped around it. That moves the question from "which agent do we use?" to "where does the data that makes it useful live?" — and the answer decides who owns the customer relationship. Benioff did not brush off the SaaSpocalypse question in the same week by accident.

What next

For your largest clients, check where the customer context actually sits and who is allowed to open it up. An agent that lives in your client’s CRM is not a tool you choose but a channel you plug into.

01.09.2026Models

Fable 5.1 leaves the token price alone and makes memory four times cheaper

On 1 September Anthropic shipped Claude Fable 5.1, alongside its restricted twin Mythos 5.1 for vetted cybersecurity and life-sciences organisations. List price stays $10 in and $50 out per million tokens; the cache-read rate drops 75%, from $1 to $0.25. Anthropic puts the saving at roughly 25% on ordinary token work and up to about 45% on heavily agentic workloads. On Terminal-Bench-Science 0.1 it goes from 24.7% to 52.6%, with 55.8% on Terminal-Bench 4.0 and 73.4% on CursorBench 3.2. One million tokens of context, 128K output, available through the Claude API, AWS, Google Cloud and Azure as claude-fable-5-1.

Why it matters

The interesting price cut is not on tokens but on memory. An agent that re-reads the same brief, brand guide and client file thirty times pays its bill exactly there — and that line item got four times cheaper without the list price moving. If you are not using caching, none of that 45% reaches you.

What next

Check whether your long-running jobs have prompt caching switched on, and whether your fixed context — brand guidelines, tone, client facts — sits at the front of the prompt where it can be cached. That is the cheapest optimisation available this month.

03.09.2026Models

GPT-6 Astra can do more on a computer than anything before it — and shows less of how

OpenAI announced GPT-6 Astra on 3 September, first to Daybreak programme participants and then in phases to Plus, Pro, Business, Enterprise and the API. List price $10 in and $50 out per million tokens, jumping to $20/$75 once a request passes 272K input tokens; context 1.05 million tokens, 128K output. On OSWorld 2.0 it reaches 72.6% at roughly 47% less time per task than Sol. It is the first model to hit OpenAI’s own "Critical" threshold on cyber capability. The technique behind it — opaque recurrence, looping tokens through the same layers instead of a readable chain of thought — makes the reasoning less auditable; chief scientist Jakub Pachocki calls that a consequence of rising capability.

Why it matters

There are now two frontier models at exactly the same list price: Astra and Fable 5.1, both $10/$50. So the choice is no longer about budget but about the task — and about what you have to be able to explain. Astra is faster and stronger at computer work; it is also the model whose reasoning you can least retell. For work going to a client or a regulator, that has become a real criterion.

What next

Re-test your two heaviest workflows on both Astra and Fable 5.1 this month, and compare on outcome, not on benchmark. Make human review mandatory on anything outbound: with a model whose reasoning you cannot read, the review is your only audit trail.

H2 2026Market

A quarter of Google now shows an AI answer, and nine in ten sessions end without a click

Across 21.9 million queries, Conductor measured AI Overviews appearing in 25.11% of Google results, up from 13.14% in March 2025. AI referral traffic is still small — 1.08% of all website traffic — but grows about a percentage point a month, and ChatGPT accounts for 87.4% of it. Semrush counted roughly 93% of AI search sessions ending without a website visit; Conductor simultaneously finds that visitors who do arrive from an LLM convert at twice the rate in a third of sessions.

Why it matters

Those two numbers are each other’s mirror image, and together they are the whole story: traffic falls, value per visitor rises. Anyone still judging AI visibility on sessions sees a falling line and draws the wrong conclusion. Anyone judging it on citations and conversion sees a channel that is smaller and better at once.

What next

Measure citations, not just visits. Record each quarter which questions you are cited for in ChatGPT, Perplexity, Gemini and Google’s AI overviews, from which source, and put the conversion of that traffic beside it. That is one report, and it answers the only question your client is asking.

Signals

Short, but not small

10 AUG · OPENAI

A model that writes exploits on purpose

GPT-5.6-Cyber, built on Sol, completes 95.0% of advanced security requests where Sol stalls at 1.5% — exploit chains, privilege escalation, authentication bypass. OpenAI used it to find two unknown flaws in Chrome’s V8 (CVE-2026-15903). Access only through the Daybreak programme, in two tiers: Blue for defence, Red for the rest. Individual Daybreak accounts have needed a hardware key since 1 September.

12 AUG · PRIJZEN

Two releases, one day, a fifteenfold price gap

Grok 4.6 arrived at $2 in and $6 out per million tokens, doubling above 200K tokens. The same day DeepSeek V4 Pro dropped its preview label at roughly $0.42 in and $0.87 out. On the Artificial Analysis Intelligence Index that is a 7.9-point gap (60.9 against 53.0) at about fifteen times the cost per completed task. For most production work, those 7.9 points are not the difference you invoice.

27 AUG · ANTHROPIC

MCP, but for hardware

The Model Hardware Standard entered research preview with research labs and manufacturers: one driver that translates "read" and "write" to microscopes, liquid handlers and robotic arms. Carnegie Mellon ran dose-response dilutions three times faster with it; QuEra had an agent recover a laser’s operating frequency without a human 99.3% of the time. Wiring that took weeks now takes hours.

24 SEP · OPENAI

The Sora API goes dark in two weeks

Web and app closed back on 26 April; the API follows on 24 September. Anyone still building on it is genuinely out of time. Veo 3.1 remains the strongest all-rounder with native audio, Kling 3.0 and Seedance hold several top-ten slots, and Runway bundles Kling 3.0 Pro, Veo 3.1, Seedance and its own Gen-4.5 under one subscription and one credit system — precisely the answer to the vendor risk Sora exposed.

AUG 2026 · KAPITAAL

The money is going to atoms and power

Crunchbase counted $42B into just over 1,500 startups in August: 25% below July’s $56B, but 122% above last August. The billion-dollar rounds went to defence tech (Hadrian), nuclear power (Valar Atomics), satellites (Yuanxin), home batteries (Base Power) and coding models (poolside). Five of the seven had last raised less than twelve months earlier.

SEP 2026 · AUTONOMIE

Eleven days of work without a human

Through the Prove2Me platform, Anthropic had Claude work largely autonomously for eleven days on the first end-to-end machine-checked proof of Fermat’s Last Theorem in Lean: 13 million lines of code and 30,300 theorems proved. It is not a marketing task, but it is the clearest measure so far of how long an agent holds its course without intervention.

Radar

The landscape, not the news

What is standing right now, with prices and characteristics. This part ages slower than the brief above — but always verify rates with the vendor.

Frontier models

price per million tokens · checked September 7, 2026
ModelLabDateContextIn $Out $Where it stands out
GPT-6 AstraOpenAI3 sep1,05M1050Above 272K input tokens the whole request bills at $20/$75. OSWorld 2.0 72.6%, ~47% faster per task than Sol.
Claude Fable 5.1Anthropic1 sep1M1050Cache reads $0.25 (−75%). Terminal-Bench-Science 52.6% against 24.7% for Fable 5. Mythos 5.1 is the restricted twin.
Claude Opus 5Anthropic24 jul1M525Half the price of Fable 5.1 and still strong on agent work. Fast mode $10/$50, ~2.5× faster.
GPT-5.6 SolOpenAI9 jul1M530The previous flagship tier; the base for GPT-5.6-Cyber.
GPT-5.6 TerraOpenAI9 jul1M2,5015The workhorse: GPT-5.5-level at about half the cost.
GPT-5.6 LunaOpenAI9 jul1M16Volume work: classification, extraction, bulk summarisation.
Gemini 3.7 FlashGoogle13 aug1M0,753,75Introductory rate through 31 Dec 2026; then $1.50/$7.50. AutomationBench 30.4% against 17.0% for 3.6 Flash.
Qwen3.8-MaxAlibaba3 aug1M262.4T param MoE, ~95B active. One flat rate across the whole context; cached input $0.25.
Grok 4.6xAI12 aug—26Double rate ($4/$12) once a request passes 200K tokens. 60.9 on the Artificial Analysis index.
DeepSeek V4 ProDeepSeek12 aug—≈ 0,42≈ 0,87Out of preview, priced in yuan (3 and 6 CNY). 53.0 on the index at about a fifteenth of the cost per task.
Kimi K3Moonshot16 jul1M——2.8T param MoE, ~50B active. Open weights since 27 Jul (modified MIT), but ~1.4TB of fast memory to run.

Moving image

in the final weeks of the Sora API · checked September 7, 2026
ModelStrengthAudioPrice per secondStatus
Veo 3.1The strongest all-rounder: 4K, prompt adherence, reference controlNative 48 kHz synchronised dialogueLite $0,05 · Fast $0,15 · Std $0,40Available
Kling 3.0 TurboPhoneme-level lip-sync, including multiple characters—≈ $0,11 – $0,14Available
Seedance 2.5Strongest at narrative: multiple shots in one pass, up to 50 referencesGenerated by default—Available
Runway Gen-4.5The most control over shot and direction; also bundles Veo, Kling and Seedance——Available
FLUX 3 Video20 seconds of picture + sound in one generation, multilingual dialogueNative, in sync—Still gated early access
SoraOpened the category——App ended 26 Apr · API ends 24 Sep

The agent stack

what sits where · checked September 7, 2026

Protocol

  • MCP — de facto standard, with the Agentic AI Foundation (Linux Foundation) since Dec 2025
  • MHS — in research preview since 27 Aug: the same idea for microscopes, robotic arms and production lines
  • Natively supported by Anthropic, OpenAI, Google and Microsoft

Workspace

  • Claude Cowork — desktop, web and mobile; remote sessions and scheduled tasks
  • Codex — OpenAI, with GPT-6 Astra since 3 Sep
  • Agentic Computer — Alibaba’s sandbox for secure execution

Orchestration

  • LangGraph · CrewAI · Dify — code-first multi-agent
  • AgentTeams — Alibaba’s multi-agent layer in Agent Native Cloud
  • n8n — visual, for teams without engineering

Inside the existing stack

  • Claudeforce — Salesforce inside Claude with 37 sales skills, open beta Sep 2026
  • HubSpot Agent Hub — public beta, agents share customer context
  • SAP Joule · ServiceNow · UiPath · Copilot Studio — RPA that went agentic

Where attention is going

numbers for the pitch · checked September 7, 2026
25,11%

of Google searches show an AI Overview, against 13.14% in March 2025

Conductor, AEO/GEO Benchmarks 2026 (21,9 mln zoekopdrachten)
93%

of AI search sessions end without a single website visit

Semrush
1,08%

of all website traffic comes from AI — growing ~1 point a month

Conductor, AEO/GEO Benchmarks 2026
87,4%

of that AI referral traffic comes from ChatGPT alone

Conductor, AEO/GEO Benchmarks 2026
$42B

venture investment in August 2026: −25% on July, +122% year on year

Crunchbase
€15M

or 3% of global turnover: the maximum penalty under EU AI Act Article 50

EU AI Act, art. 50

Sources

Where this comes from

Introducing Claude Fable 5.1 and Claude Mythos 5.1anthropic.com · Anthropic, 1 sep 2026Anthropic’s Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for cache readsVentureBeat · 1 sep 2026OpenAI launches Astra, its powerful (and controversial) new modelTechCrunch · 3 sep 2026OpenAI announces rollout of GPT-6 Astra modelCNBC · 3 sep 2026Gemini 3.7 Flash: our most intelligent workhorse modelblog.google · Google, 13 aug 2026Google’s Gemini 3.7 Flash arrives before Gemini 3.5 ProAxios · 13 aug 2026Salesforce and Anthropic announce Claudeforcesalesforce.com · Salesforce, 26 aug 2026Salesforce, Anthropic expand partnership as Benioff responds to ‘SaaSpocalypse’ concernsCNBC · 26 aug 2026Previewing the Model Hardware Standardanthropic.com · Anthropic, 27 aug 2026How Claude’s text watermark worksanthropic.com · Anthropic, 14 aug 2026EU compliance, delivered globally: Anthropic to watermark Claude’s output worldwideEuronews · 11 aug 2026The EU AI Act’s transparency rules: a practical guide to Article 50artificialintelligenceact.euEU finalises transparency rules for AI-generated contentPaul, WeissOpenAI unveils new cybersecurity model GPT-5.6-CyberSecurityWeek · 10 aug 2026OpenAI launches two-tier access programme alongside GPT-5.6-CyberInfosecurity Magazine · 10 aug 2026Grok 4.6 vs DeepSeek V4 Pro: 14.9× the cost, 7.9 more index pointsSpeedway Media · 17 aug 2026Qwen3.8-Max: Alibaba’s 2.4T model, generally availableDataNorth · 3 aug 2026Global venture funding jumps 122% in August as streak of billion-dollar deals continuesCrunchbase News · sep 2026The 2026 AEO / GEO benchmarks reportConductorAI search statistics 2026: 60+ data points on visibility, citations and trafficSuperlinesSora 2 shuts down September 24: move your prompts nowPrompt ArchitectsAI news for September 1, 2026 — daily editionAI Weekly · 1 sep 2026

About this page. The brief above carries the date of its edition; each block in the radar below states when it was last checked. Everything comes from public sources, with the primary source where one is available — the list sits above. Prices, benchmark scores and release dates partly come from trade press and trackers; always verify rates with the vendor before budgeting on them. Benchmarks are snapshots and say nothing about your specific work: re-test on your own tasks.