top of page

AT&T Cut AI Costs 80%. Marketing Is Still Paying Frontier Prices.

  • Writer: Aseem Singh
    Aseem Singh
  • 11 minutes ago
  • 2 min read

AT&T burns through 45 billion AI tokens a day.

This year they moved a huge slice of that work off OpenAI and Anthropic. Open models now handle about 40% of their AI use. The target is 60%. The saving: up to 80% versus earlier this year. Same week, Mistral raised $3.5 billion and four labs dumped another point release.

If you run an agency, the headline is not GPT-6 Astra or Claude Fable 5.1. The headline is the bill.

What actually shipped this week

I am not collecting model names for a newsletter. I am watching where the money moved.

  • Open models went from 10% of usage on OpenRouter a year ago to 58% last month. Airbnb and Deloitte are in the same shift.

  • AT&T routes work through an AI gateway. Coding tasks dropped about 56% in cost with quality off by roughly 2%. Some telecom-tuned models hit 90% inference savings.

  • Mistral took $3.5 billion so Europe still has a seat. Nvidia agreed to buy Hugging Face for $12.9 billion. The cheap path just got infrastructure.

  • Anthropic, Meta, Google, and OpenAI still shipped this week. Claude Fable 5.1. Muse Spark 1.3. Gemini 3.8 Flash. GPT-6 Astra. Noise. Not the operating system.

The marketing and AI section nobody should skip

Here is the part most AI recaps skip, and it is the one that hits agency P&L first.

Most marketing work is not frontier work. First drafts. Twelve ad variations. Call-transcript summaries. Competitor scans. Brief expansions. That is router work. You do not need the most expensive model in the room to ship a Tuesday deck.

What we actually use at Dzine Prodigy is boring on purpose.

LAYER 1 — Route the volume. Open or cheap models for research dumps, first-pass copy, variation volume, and internal summaries.

LAYER 2 — Keep one frontier model for the brief that has to be right. Strategy, brand voice lock, client-facing reasoning.

LAYER 3 — Do not rebuild the stack every Tuesday. Lock the router for 30 days. Measure cost per shipped asset, not model FOMO.

AI for marketing agencies in 2026 is not “which lab won this week.” It is production speed with a bill you can defend.

Do this before Friday

  1. List the 10 jobs that eat your tokens this month.

  2. Mark which ones a cheaper model already ships at an acceptable quality bar.

  3. Move those off the expensive path this week. Leave one frontier model for the work that still needs it.

The labs will keep declaring eras. The agencies that ship on a routed stack will own the next one.

Recent Posts

See All

header.all-comments


bottom of page