In this issue
The AI price war moved down the stack this fortnight, from the frontier models you show off to the cheap ones you actually run all day. Google shipped Gemini 3.6 Flash with a 17% cut in output tokens and said it has already started training Gemini 4. OpenAI put a managed voice-and-chat agent named Presence in front of enterprises, and Microsoft and Mistral made controllable, even air-gapped, frontier AI something you can buy inside the stack you already run.
The story underneath the launches is the one I care about: everyone keeps cutting the price per token, but agents burn tokens in bulk through planning, retries, and tool calls, so your bill can climb while the rate card falls. The real operator skill this fortnight is measuring cost per finished job, not per token. And Anthropic quietly shipped the friendliest way yet to teach an agent your work, you record your screen doing it once. Let us get into it.
Topics of the day:
What's happening: Google released Gemini 3.6 Flash this week, its high-volume workhorse model, at $1.50 per million input tokens and $7.50 per million output, and says it now uses 17% fewer output tokens than 3.5 Flash. The same post confirmed Google has "started our most ambitious pre-training run yet, for Gemini 4."
In practice:
Bottom line: Flash is the model most teams route their volume through, so a 17% output cut is a line item, not a leaderboard point. Turn it on where you already pay per token.
What's happening: OpenAI launched Presence, a managed system for running production voice and chat agents with policies, guardrails, escalation, and evals built in. It is not self-serve, OpenAI deploys it with its own engineers and partners, and it already runs OpenAI's English phone support, where it resolves 75% of inbound issues without a human.
In practice:
Bottom line: OpenAI stopped selling only the API and started selling the finished support agent. If you run a help desk, this turns into a buy-or-build decision this quarter, not someday.
What's happening: Microsoft and Mistral expanded their partnership, putting Mistral Medium 3.5 and OCR 4 into Microsoft Foundry and Copilot Studio and letting teams deploy across cloud, cloud-connected, and fully disconnected environments. Microsoft's Brad Smith framed it as Europe getting "the world's most capable AI without compromising control over their data, operations or digital future."
In practice:
Bottom line: For anyone stuck between needing frontier AI and not being allowed to lose control of the data, this puts capable models inside the stack you already run. Control just became a deployment setting instead of a blocker.
Cheaper AI tokens do not guarantee cheaper enterprise agents - The cost-per-outcome argument, expanded. It is also today's closer, so read it if the last story lands for you.
The State of AI Agents in 2026 - Jon Radoff's nine findings on why most organizations have not captured AI's value yet, even as cost and capability keep improving.
Databricks State of AI Agents - Telemetry from thousands of orgs, with one number worth pinning up: teams with AI governance ship over 12x more agents to production.
Google's Nano Banana 2 Lite and Gemini Omni Flash - If you missed it, Google's near-instant image model and its $0.10-per-second video model are live inside Firefly, Figma Weave, and more.
What's happening: Anthropic added Record a Skill to Claude Cowork. You screen-record yourself doing a task once while narrating it, and Claude turns the demo into a reusable skill it can run again on its own. It is on Pro, Max, and Team plans, in the desktop app.
In practice:
Bottom line: This drops the bar for automating a workflow from writing a spec to just doing it once while you talk. It is the most approachable way to teach an agent your job that I have seen.
What's happening: A Forbes analysis makes the case that this month's price cuts hide a trap: agents multiply token use through planning, retrieval, tool calls, and retries, so the price per token can fall while the cost per completed task rises. It notes an OpenAI audit where roughly 30% of one benchmark's tasks were broken, and that 73% of enterprises overshot their AI cost projections last year.
In practice:
Bottom line: Every price cut in this issue is real, and none of them lowers your bill by itself. The gap at the top of the leaderboard is a rounding error. The gap in how you route and cap your agents is the whole budget.
GitHub added Gemini 3.6 Flash to Copilot on day one, so any team on a paid plan can route dev work through this week's cheaper model once an admin flips the policy.
Mistral shipped OCR 4 into Microsoft Foundry, a document model with 170 languages, bounding boxes, and confidence scores that self-hosts in a single container, handy for RAG pipelines that cannot leave your walls.
Google is piloting Gemini 3.5 Flash Cyber, a cybersecurity-tuned model limited to government and trusted partners, an early sign of labs shipping vertical, security-specific tiers.
HITL engraving style, blue ink on cream. The Gemini image auto-gen step needs the Mac toolchain, so on this box the Midjourney prompts below are the portable deliverable (same fallback as the 2026-07-09 issue):
Option 1 (lead, cheaper workhorse): A single sturdy draft horse standing calm, rendered in fine engraving lines, crisp outline with subtle interior detail, blue ink on white. Soft gradient field in #F7F8FA / #4B7BEC with #0B1221 accents. Minimal, centered composition, no text --ar 16:9
Option 2 (cost per outcome): A balance scale weighing a small stack of coins against a dense tangle of gears, crisp outline with subtle interior detail, blue ink on white. Soft gradient field in #F7F8FA / #4B7BEC with #0B1221 accents. Minimal, centered composition, no text --ar 16:9
Option 3 (record a skill): A hand-crank film camera on a tripod pointed at a small seated robot, crisp outline with subtle interior detail, blue ink on white. Soft gradient field in #F7F8FA / #4B7BEC with #0B1221 accents. Minimal, centered composition, no text --ar 16:9
The AI price war moved to the models you actually run all day. Here is what changed for your stack this fortnight.
I broke it down in today's Human in the Loop:
The catch, and today's closer: cheaper tokens do not mean cheaper agents. Agents burn tokens in bulk, so your bill can climb while the rate card falls. The skill now is measuring cost per finished job, not per token.
Full breakdown in today's issue ↓
https://hitl.plyolab.com
Audience filter applied: operators using AI at work, "what changes for my Monday workflow", not model research or lab gossip.
Window: July 10 to 24, 2026 (last issue shipped 2026-07-09). De-duped against the July 8-9 launch cluster already covered last issue (GPT-5.6 GA, Grok 4.5, Muse Spark 1.1 initial), plus GPT-Live, Microsoft 365 Copilot GPT-5.6, Claude reflection, and the local-models thread.
Story picks:
1. Gemini 3.6 Flash (LEAD): the model most operators route volume through got a real price cut, plus the Gemini 4 tease. Highest "changes my Monday" density. Rejected: "no Gemini 3.5 Pro yet" angle (roadmap gossip, folded into the Shortlist Cyber item instead).
2. OpenAI Presence: moves OpenAI from API to a finished, managed support agent, a concrete buy-vs-build decision. Rejected: industrial-espionage and "model escaped sandbox" Reddit threads (unverified, drama not workflow).
3. Microsoft + Mistral: controllable and air-gapped frontier AI is a genuine procurement unlock for regulated and EU teams. Primary Microsoft source.
4. Claude Record a Skill: requested-shaped operator win, teach an agent by doing the task once. Secondary sourcing only (see note below).
5. Cheaper tokens do not mean cheaper agents (CLOSER): synthesizes every price cut in the issue into one budgeting rule, the real takeaway.
Subject picks:
PICKED: "Google cuts Gemini Flash prices and teases Gemini 4" (8 words, named entity + verb, lead-only).
ALT 1: "Gemini 3.6 Flash makes your cheap tier cheaper" (8 words, benefit-first, vaguer entity).
ALT 2: "The price war reached the models you run all day" (9 words, theme-first, no named entity).
Pre-header pick:
PICKED: "PLUS: OpenAI Presence, Claude learns your workflow from one recording, and why cheaper tokens can still raise your bill" (teases stories 2, 4, 5, no trailing period).
Read Later picks: Forbes (also the closer, cross-listed on purpose), Radoff and Databricks agent-economics reads (both loaded and confirmed), and the Nano Banana / Omni Flash catch-up (Google Cloud primary; note the launch itself was July 1, just before the window, included as an evergreen creative-models pointer).
Sourcing notes (verification, per house rule):
Hero images: NOT generated this run, no GEMINI_API_KEY on this box (Linux/Orca). Midjourney prompts above are the deliverable, same fallback as the 2026-06-25 and 2026-07-09 issues.