New State of AI 2026: Mid-Year Reality Check is live. Read the report
Skip to content

More Agents Than Employees: How Zapier Disrupted Itself Before AI Could

The best AI model in the world just scored 18.1%. On Zapier's own benchmark for real business work — the cross-app tasks any white-collar worker does every day — even the top frontier model completes them barely one time in five. That's the number Wade Foster keeps pointing at, and he runs an automation company that stands to gain from the hype. Instead, he makes the case for what actually works right now: not turning a model loose, but blending deterministic workflows with agents where each is strong.

In this episode of Talking AI, Matt Paige sits down with Wade Foster, co-founder and CEO of Zapier, who built a scrappy Y Combinator startup into the $5 billion plumbing of the SaaS era on barely a million dollars raised. Foster called a company-wide “code red” the week GPT-4 launched, and he's spent the years since rewiring how Zapier — and its customers — actually use AI.

The conversation covers why he shut the company down for a week in 2023, how AI habits actually stick, what Zapier's AutomationBench reveals about the gap between benchmark scores and real-world reliability, why coding models improve faster than knowledge-work models, how to tell a workflow from an agent, and the difference between individual AI and the institutional AI almost no company has cracked.

In this episode, you'll hear about:

  • The three things about GPT-4 that triggered Zapier's first-ever code red
  • How daily AI use jumped from 11% to over 50% in a single hackathon week
  • The moves that make AI habits stick: show-and-tell, repeat hackathons, and “not yet”
  • Why the best model on AutomationBench still scores only 18.1%
  • Why coding is easy to verify — and subjective knowledge work isn't
  • The power of hybrid setups that blend deterministic workflows with agents
  • Wade's prediction: most tokens on open-source models, most spend on the frontier
  • What actually makes a good eval — hard for models, easy for humans, private data
  • A plain-English definition of an “agent” versus a deterministic workflow
  • The daily recap workflow Wade thinks everyone is sleeping on
  • Floor raisers vs. ceiling raisers — and why individual AI isn't enough
  • Why the six-month product roadmap is dead

Key Moments

  • 00:00:30 — Why Wade called a company-wide code red when GPT-4 dropped
  • 00:04:40 — Making AI habits stick: show-and-tell and repeat hackathons
  • 00:06:38 — Differentiation when AI is best at the thing you sell
  • 00:09:34 — AutomationBench: the best model scores just 18.1%
  • 00:11:31 — Why the top model stalls: verifiable code vs. subjective work
  • 00:14:19 — Getting squeezed on both sides: AI in the company and the product
  • 00:15:20 — Model efficiency, Coinbase, and the token-maxing debate
  • 00:17:18 — What makes a good eval
  • 00:19:30 — What actually counts as an “agent”
  • 00:23:12 — Iterating on workflows with your own mini-evals
  • 00:26:15 — The kind of worker thriving right now
  • 00:27:36 — Wade's favorite workflow: the daily recap
  • 00:30:44 — Floor raisers vs. ceiling raisers for AI adoption
  • 00:34:55 — From individual AI to institutional AI
  • 00:37:58 — Why the six-month roadmap is dead
  • 00:40:15 — Where to find Wade and the best way to try Zapier

Get the best of our content
straight to your inbox!

Don’t worry, we don’t spam!
Related Posts
Categories
Explore by topic
Trending now
More topics

No topics match that.

View the full blog