The Agency AI Stack Just Split. Here’s What We’re Shipping With.
- Aseem Singh
- 2 days ago
- 2 min read
This week the agency AI stack actually split. Not in a keynote. In contracts, model cards, and client Slack threads.
On one side: Tencent open-sourced Hy4-preview — a 770B-parameter mixture-of-experts model with a 1-million-token context window. Alibaba pushed Qwen3.8-Flash-Next in the same cheap-fast lane. Long briefs, brand books, and campaign archives can now sit in one pass instead of getting chopped into forgetful chunks.
On the other side: OpenAI told SpaceX it will cut model supply to Cursor on November 12 after the $60B buyout. That is not gossip. That is a change-of-control clause hitting a tool half the industry built workflows on.
I’m using this at Dzine Prodigy as a reminder, not a panic button: if your production line lives inside one closed door, you do not have a stack. You have a lease.
What actually changed this week
Three moves matter if you ship work for clients, not slide decks.
Open weights got long-context. 1M tokens is not a demo number. It is a full brand system + last quarter’s ads + the brief, in one run.
Closed tools showed the kill switch. Cursor is the loud example. Any agency still single-homed on one model vendor is taking unnecessary production risk.
Agents left the chat box. Anthropic previewed a Model Hardware Standard so Claude can operate lab tools and physical devices. Production is moving from “write copy” to “run the workflow.”
BEFORE: one hero model, one coding IDE, one image generator, hope the invoice stays friendly.
AFTER: a two-layer stack. Layer 1 is portable — open or multi-vendor models you can swap in a week. Layer 2 is the expensive specialist you rent only when the output actually needs it.
The marketing piece nobody should skip
While labs fight over chips and contracts, buyers already changed how they find you. Forrester has 94% of B2B buyers using AI in purchasing. Semrush has been showing AI-referred visitors converting around 4.4× organic. A 2026 B2B study put GEO experimentation at 92%, with most teams still missing a named owner.
Clients are no longer asking “what’s our keyword rank?” First question in the room is: do we show up when someone asks ChatGPT, Gemini, or Perplexity who to hire?
That is generative engine optimization. Not a rebrand of SEO. A second scoreboard.
What we actually run for brands right now:
Entity-first pages. One clear who / what / for whom. Models cite clean entities, not clever headlines.
Proof that can be quoted. Original numbers, named case studies, author pages. If a model cannot attribute you, it will attribute your competitor.
Citation tracking across ChatGPT, Gemini, Perplexity, and AI Overviews. If you cannot see the mention, you cannot fix the brief.
SEO gets you the blue link. GEO gets you the answer. In 2026, the answer is where the budget starts.
Do this before Friday
Pick one live client workflow. Write down every model and tool it depends on. Mark anything you cannot replace in seven days.
Then open ChatGPT, Gemini, and Perplexity and ask the exact buying question your client cares about. Screenshot who gets cited. That screenshot is the brief for next month’s content.
Finish something this week. The stack already moved. Waiting for a “settled” vendor map is how agencies lose the quarter.


Comments