This episode of Unsupervised Learning features Jacob Efron (Redpoint) in conversation with Benedict Evans, a prominent tech analyst, discussing the current AI hype cycle, foundation model economics, enterprise adoption realities, job automation fears, consumer AI challenges, and historical parallels to previous platform shifts.
Historical perspective grounds the AI hype cycle
Evans argues against ranking AI against the internet, mobile, or industrial revolution; instead, he examines patterns from past platform shifts to understand competitive dynamics, value capture, and adoption curves.
Every major technology shift (electricity, semiconductors, mobile, cloud, operating systems) looked different but shared structural dynamics: escalating costs, marginal costs, value moving up-stack, and commoditization at the infrastructure layer.
Mobile data traffic grew ~2,000x over 15 years into a trillion-dollar industry with $200B annual capex, yet carriers captured little value — most went to application-layer companies (Uber, YouTube, banking).
Semiconductors followed “Rock’s Law” (fab costs doubling every 4 years), consolidating from dozens of cutting-edge players to essentially one (TSMC); foundation models may follow a similar consolidation if scaling continues.
A key difference: previous shifts had known physical limits (PC prices, fiber deployment timelines), whereas we lack scientific understanding of why LLMs work so well or where they hit limits.
Evans distinguishes between the “binary” AGI/ASI scenario (where employment concerns become irrelevant) and the more analyzable near-term enterprise software transformation.
He references 1990s internet utopianism (Barlow’s “Declaration of Independence of Cyberspace”) as a reminder that millenarian narratives accompany every platform shift; the practical work is building software and discovering use cases.
Foundation models face commoditization pressure and uncertain value capture
Currently 3–6 companies produce frontier models with similar capabilities; leadership rotates every few weeks, signaling low barriers to entry beyond capital.
SpaceX (xAI) jumping back to the top of leaderboards after failing is a negative signal for moats — it suggests anyone with billions and talent can compete.
If scaling continues, compute becomes the primary moat, and the number of frontier players should shrink (semiconductor analogy), but a price/performance collapse (e.g., 10x model size for same cost) could disrupt this.
Value capture question: how many use cases need the absolute frontier vs. “good enough” commodity models? Dictation runs free on-device; Amazon’s recommendation ROI sits somewhere in the middle.
Evans compares foundation models to TSMC, Windows, or AWS — great businesses, but none own the full stack up to end-user applications.
OpenAI and Anthropic lack distribution, infrastructure, and legacy revenue; they must build the entire stack (chips, data centers, products) while conducting cutting-edge research — akin to Bill Gates building PCs, enterprise software, and broadband simultaneously in 1980.
Incumbents (Google, Microsoft, Apple) try to make AI a feature of existing products; model labs must create new distribution and product layers.
OpenAI’s chaotic product launches (browser, social, shopping, ads, multiple app stores) reflect experimentation in radical uncertainty; Anthropic’s narrower focus on coding emerged from constraints.
Execution matters: network effects and commoditization are not deterministic — MySpace preceded Facebook, Google had all the data but lost social; someone must actually build the winning product.
Enterprise adoption is early, fragmented, and requires structural reimagination
Current state: ~10–15% daily active users (often once/twice daily), 20–40% weekly/monthly — most people use LLMs occasionally, not as a continuous computing substrate.
Analogy: accountants saw spreadsheets as life-changing in the 1970s; lawyers saw them as timesheet tools. Software developers have clear product-market fit; knowledge workers with autonomy use LLMs heavily; most employees don’t yet.
Three barriers to broad enterprise adoption: (1) most people aren’t tool builders, (2) most don’t recognize the problems tools could solve, (3) most lack authority/data access to deploy tools (regulated data, systems of record, 1,500 stakeholders).
AI shuffles the existing stack: improvised tools (Excel, CSV, Tableau) ↔ horizontal systems (SAP, Workday) ↔ 400–500 SaaS apps. New AI-native SaaS will compete with AI-enhanced incumbents and AI-augmented spreadsheets.
Deployment so far: Step 1 — give everyone Copilot (low adoption). Step 2 — pilots (work but are one-off). Step 3 — structural reimagination of workflows (the real opportunity, but requires consulting/integration skills most companies lack).
Vertical AI companies and consulting firms (Bain, BCG, Accenture) are converging: both reimagine processes, one via software + forward-deployed engineers, the other via strategy + growing engineering teams.
Enterprise sales cycles are long (cloud took 20 years to reach ~30% of workflows); non-tech companies have other priorities (regulation, infrastructure, droughts, leadership conflicts) — AI isn’t their only concern.
Some industries (aggregates, heavy machinery) may see minimal AI impact, just as the internet barely changed Caterpillar’s core business; physical AI/robotics adoption will follow its own curve.
Job impact is jagged, unpredictable, and historically overestimated in the short term
Hinton’s 2014 “stop training radiologists” claim failed because ML couldn’t do the task and he misunderstood the role (radiologists do far more than image classification).
Automation historically targets tasks, not jobs; accountants increased in number throughout the 20th century despite punch cards, mainframes, and spreadsheets — the work changed, not the headcount (Jevons paradox / price elasticity).
“Lump of labor” fallacy: we see jobs disappearing but not the new ones created (e.g., internet demolished local newspaper advertising, not journalism itself; Uber created ride-hailing demand no one forecasted).
Current models have jagged capabilities: they can do some 17-hour tasks but fail at 5-minute ones; intelligence isn’t linear.
The “expert-in-the-system” fallacy: you cannot measure what a law associate does via a radar chart, nor measure if a model replicates it — the task definition is the error.
Consumer surplus / enterprise equivalent: legal research that took a day now takes minutes, but clients pay the same for the same output — value accrues to clients, not necessarily law firms.
Industry variability: Uber destroyed ~75% of NYC taxi medallion value; Airbnb is ~10–15% of hotel market. Hotels have business/conference travel; taxis didn’t. AI impact will be equally uneven.
Physical-world jobs and implicit/tacit knowledge (hard to explain, hard to validate, hard to get training data) resist automation; simulation may eventually bridge this, but it’s a slog.
Consumer AI lacks a breakout product; usage is shallow and gimmick-prone
Most consumers don’t have a computer as primary device (smartphone-first); “computer use” agents feel geeky and insecure (analogy: 1990s PCs freezing, crawling under desk to reboot).
Early internet required portals/handholding (Yahoo, AOL) before habits formed; today’s chatbot “tiles” may be the equivalent portal phase.
Use cases must be invented by entrepreneurs (Flickr, Instagram, TikTok), not emerge spontaneously from broadband/LLM access.
Sora and Midjourney follow the consumer social/gimmick cycle: brief excitement (drones, 3D printing), then fade unless integrated into a real product (e.g., virtual try-on for fashion).
Image/video generation remains a hobbyist/tool category, not a mass consumer habit; the product layer hasn’t been invented yet.
Voice/chat interfaces suffer from articulation burden: users must isolate, describe, and validate tasks — a skill most lack.
Coding is the breakout use case; model labs are competing at the application layer
Coding works because: (1) verifiable (tests pass/fail), (2) scalable validation (run millions of times), (3) tool builders are the users (developers building for developers), (4) code is uniquely suited to LLMs (structured, logical, massive training data).
Model providers (Claude Code, GitHub Copilot) are winning the coding application layer — analogous to spell-check, charts, and printing moving from standalone apps into the OS/editor.
The feedback loop: real-world coding usage improves the underlying model (unlike AWS, where pharma vs. grocery workloads don’t change the compute).
Cursor, Windsurf, etc., concluded they must own the model layer to survive; Anthropic’s coding focus gave it a temporary lead, but everyone is now crowding coding.
Coding is a relatively small industry (~millions of developers); a billion-user consumer use case would break current infrastructure (marginal cost per query vs. zero-marginal-cost software).
Uncertainty demands experimentation, not deterministic forecasts
Evans emphasizes radical uncertainty: MCP, current architectures, and leaderboards will likely be replaced; presume most things being built today will fail (slide of 1995/2010 failed acronyms/companies).
Execution is non-deterministic: history isn’t a force; people must build the winning products (Mark Andreessen: “history isn’t a person”).
OpenAI’s recent chaotic app launches show writing code isn’t the hard part — product sense, integration, and UX are.
Cargo cult risk: hiring Meta execs from 2010 (post-product-market-fit) to run 2024 pre-PMF playbooks (app stores, ads, commerce) without the early-stage instincts.
Advice to young AI researchers: (1) assume radical turnover — today’s acronyms/concepts will fail; (2) enterprise adoption is slow and messy (cloud at 30% after 20 years); (3) other industries have hard problems you don’t see (water company CEO: “innovation maybe top 5 priority next year” amid lead pipes, drought, regulation); (4) AI won’t change every industry (internet didn’t change aggregates or Caterpillar much).
Final heuristic: “It’s a low-margin reseller, one-time sales” — sometimes the core business model is simpler than the hype suggests.