Mati Staniszewski, co-founder and CEO of ElevenLabs, describes building an AI-native audio company from 2022 — before ChatGPT launched — by combining frontier research with product deployment, staying relentlessly focused on voice as the core modality, and using forward-deployed engineers to turn customer problems into platform capabilities.
Origin: Poland’s Single-Voice Dubbing Problem Sparked the Idea
The founding insight came from growing up in Poland, where every film — regardless of character gender or emotion — is narrated by a single voice actor, stripping all emotional intonation.
In 2021 this was still the norm; the founders saw a future where content could be experienced in any language with the original voice, emotion, and intonation preserved.
The initial plan was dubbing, but early creator feedback revealed a more immediate need: high-quality speech generation for narration, post-production fixes, and pre-video script preview — so they paused the language-shift layer and focused on making generated audio sound human.
Research + Product Deployment: The Core Organizational Model
ElevenLabs is structured as a combination of a research lab (frontier audio models: text-to-speech, speech-to-text, orchestration) and a product platform that helps companies transform how they communicate with the world.
Research side is 100% audio; product side integrates audio with knowledge, LLMs, and creative workflows for marketing (e.g., Ramp Super Bowl ad), support (Deutsche Telekom voice agents), sales qualification, and government operations (Polish healthcare appointment agent).
The company runs on small, autonomous teams (usually <10 people) with a flat hierarchy (max 5 layers, aiming to reduce over time) and a “best idea wins” culture inherited from Palantir.
Focus as Competitive Advantage: Saying No to Video Avatars
They evaluate every product opportunity by asking whether their audio models provide a unique advantage in that experience; if the bottleneck is video quality, not audio, they don’t build it.
Two years ago they explored lip-dubbing/avatars but stopped because video quality was the limiting factor — even perfect audio couldn’t save the experience.
Today they’re revisiting that space via partnerships (e.g., London Ads Engine) using existing open-source video models, staying at the intersection where audio adds unique value.
Building an Ecosystem: Voice Marketplace with Creator Compensation
They built a marketplace where users create voices, ElevenLabs authenticates them, and creators earn recurring compensation every time their voice is used.
The marketplace now has ~20,000 voices covering diverse accents, styles, languages, ages, and genders; ElevenLabs also fills gaps by commissioning voices for underrepresented pockets.
Example: “George” (a professional voice actor) earns each time his voice is used in ElevenReader, the app that turns documents into podcast-like audio.
Communication Platform Vision: Deutsche Telekom Case Study
Deutsche Telekom started with marketing (podcasts, ads via ElevenCreative), then expanded to call-center voice agents (ElevenAgents) integrated with CRM and telephony (SIP/Twilio), and recently deployed an in-network agent for T-Mobile subscribers that can book appointments and do real-time translation.
Expansion path: prove impact in one use case → build integrations for the next → deploy with forward-deployed engineers (FDEs) working side-by-side with customer teams → trial, simulate, scale gradually.
Four FDEs were embedded in Germany for months to integrate with Deutsche Telekom’s systems and ensure the agent followed correct logic and had the right knowledge.
FDEs sit in Product (not Go-to-Market) so they feed customer learnings back into the roadmap — e.g., turning a custom healthcare integration into a reusable product feature for all customers.
From Palantir, Mati adopted: small deployment teams (5 people considered large), “best idea wins” culture, and obsession with being on the frontline with the customer to understand the real problem.
Also from Palantir: the belief that art and taste matter as much as science in AI — voice quality is subjective, and great product experience depends on design language, class, and emotional resonance.
Flat, Transparent, AI-Native Organization
Near-zero titles, max 5 management layers, wide spans of control (~10 direct reports) — enabled by AI summarizing granular updates from every team so leaders can be proactive, not reactive.
Full internal transparency: almost everyone has access to all documents and data, creating a system where information flows directly from the person closest to the problem.
Non-technical teams (talent, ops, legal) embed engineers to automate workflows and teach AI adoption bottom-up; centralized functions (legal, RevOps) are engineering-heavy to scale knowledge across the org.
Where Revenue Is Growing Fastest
Largest clients are on the agent side (conversational AI) and creative side (content localization/production).
Fastest-growing vertical: fintech (Revolut, Klarna, Pug Bank, Customers Bank) — moving at “superfast” speed despite regulatory complexity.
Next: healthcare, telecom, retail/e-commerce (retail wave started this year).
Using AI to Amplify Human Potential: Voice Restoration Stories
Restored voices for 10,000+ people who lost speech to ALS, throat cancer, or other conditions — e.g., a musician touring again with his AI voice, a Congresswoman addressing Parliament for the first time after voice loss, author Tim Green (ALS) launching a podcast and winning an Emmy.
Partnered with 800+ organizations in education and culture to expand access.
Safety/authenticity maintained: imperfections (“um,” pauses) deliberately kept in voice agents because users trust natural, slightly flawed voices more than perfect ones.
Becoming an Entrepreneur & Building With Piotr
Mati didn’t know entrepreneurship was a path growing up in Poland; exposure at BlackRock and Palantir (where he worked on-site with customers in Aberdeen) showed him he could solve problems directly.
He and co-founder Piotr Dąbkowski (ex-Google Knowledge Graph, university vision research) spent years doing weekend hack projects before committing to a problem they were “obsessed with” — the 2021 dubbing insight was that moment.
Piotr stays out of public view but is deeply aligned on the long-term vision and AGI trajectory; both see this as a once-in-history opportunity.
Why Mati Won’t Sell: No Price Can Replace the Work
Received 3-4 serious acquisition offers (last in June 2024), all declined — not because numbers weren’t life-changing, but because selling was never an option.
Influenced by founders like Evan Spiegel, Scott Wu (Cognition), and Travis Kalanick who regretted or resisted exits; the work itself is the reward, and the current moment (right people, right skills, right time) is unrepeatable.
Goal: build an independent, enduring company that defines the voice/AI communication layer for the next decade.
Voice as the Primary AI Interface
Within 12 months, voice will be a core way people interact with AI — combining IQ and EQ, understanding emotion, pausing, thinking, and responding with natural conversational timing.
The shift: humans used to learn technology’s language (keyboards, code); now technology learns human language (voice), returning to the most natural communication mode.
Vision: screens and phones recede; ambient devices respond to voice and gesture — the “Babel fish” from Hitchhiker’s Guide realized not as a single device but as a capability embedded everywhere.
Conferences as a Business Tool, Not a Distraction
Attends only 2-3 events per quarter, almost exclusively partner-hosted (Ramp, Dell) where customers and prospects are already gathered.
Preparation is key: meetings scheduled in advance, clear business intent, no “chilling” or session-hopping — learned the hard way after early conferences yielded fun but no productive output.
The wide applicability of conversational agents (every business with customers) makes targeted events an efficient way to deepen relationships and expand deals.