How I would define product health and run experiments across ElevenCreative, ElevenAgents and ElevenAPI, grounded in live public adoption data, plus a working desk that reads an A/B test, sizes one, runs certified metric SQL in the browser and measures SDK and Voice Library adoption.
The analyst behind discovery, adoption and retention for three platforms, working on top of a certified dbt KPI layer and with Product and Growth on experiments.
One company, three usage shapes: creators spending a shared credit pool, businesses running agents by the minute, developers calling an API through SDKs. One "active user" definition cannot cover all three.
A six-tab desk: experiment readout with SRM and Holm, a power planner with CUPED, a DuckDB metrics lab over a labelled synthetic warehouse, and live SDK and Voice Library adoption.
| Takeaway | What it means for product analytics |
|---|---|
| Activation differs per platform | Creative activation is a habit (repeat generation with a chosen voice), Agents activation is a deployment, API activation is sustained request volume. Each needs its own aha search and a test before it is certified. |
| Credit rollover changes what retention means | Unused credits roll over for two months. A paying creator can skip a month and stay healthy, so a calendar-month "active" flag overstates churn risk. |
| The SDK long tail is an adoption ceiling | 59% of last week's Python SDK installs and 43% of JS SDK installs were on versions more than 180 days old. Any feature that needs a new SDK reaches developers slowly. |
| Read the test before the result | Sample ratio checks by segment, a pre-registered primary metric, guardrails and a multiple-testing correction come before the ship call. The demo shows a planted logging bug caught this way. |
A voice model company that now sells three platforms, with enterprise the larger half of revenue and agents the fastest-growing line.
| Platform | How it is billed | Who uses it | What product health has to capture |
|---|---|---|---|
| ElevenCreative | Plans from Free (10k credits) to Business (6M credits); one shared credit pool across speech, music, image, video and dubbing; two-month rollover capped at 3x quota | Creators, marketers, studios; web and mobile apps; Voice Library on both sides | Habit and breadth: repeat generation days, voices chosen, products touched, credit burn against quota |
| ElevenAgents | $0.08 per call minute on every plan, $0.16 in burst; LLM usage passed through; telephony at cost | Businesses building voice and chat agents; partners; enterprise deployments | Build to production: agent created, tested, deployed, then live minutes and resolution quality |
| ElevenAPI | Per character or per minute by model (v4 $0.08 per 1K chars at list; 72% off until 12 Oct) | Developers through the JS and Python SDKs and raw REST | Integration depth: key created, first successful call, sustained volume, model mix, SDK version |
Sources: ElevenLabs blog (Series D, $500M ARR, Sept 2026 tender, creator payouts), elevenlabs.io/pricing, /pricing/agents, /pricing/api, read 6 Oct 2026.
Public package downloads are the one adoption signal anyone can read for ElevenAPI and ElevenAgents. Downloads include CI runs, so shares and trends matter more than levels. Week to 4 October 2026; the demo refreshes them live.
| Version-lag measure, last 7 days | JS (npm) | Python (PyPI) | Reading |
|---|---|---|---|
| Installs on versions older than 180 days | 42.8% | 59.1% | Most Python installs run code from before March. New model ids and fixes reach this group late or never. |
| Median installed version age | 147 days | 246 days | Python lags JS by about three months. |
| Most-installed single version | 2.68.0 (7.7%) | 2.1.0 (16.3%) | Python 2.1.0 is from May 2025, 79 releases behind. One version holding a sixth of installs points to pinned CI or a popular template. |
| Versions above 0.1% of installs | 66 | 67 | A wide support surface for any deprecation. |
Feature adoption for an API feature has a ceiling set by the share of developers on an SDK that supports it. Version mix belongs in the denominator of every ElevenAPI adoption metric.
Join SDK version (from the user-agent header) to API keys and accounts. If the 2.1.0 tail is a few high-volume keys it is a CI artefact; if it is many small accounts it is an onboarding template still in circulation.
Sources: api.npmjs.org downloads and per-version endpoints, registry.npmjs.org, ClickPy (ClickHouse public PyPI dataset), pypi.org release dates. Live in the SDK adoption tab.
Each platform faces a different set of rivals, so each has a different switching risk to watch in retention data.
| Rival | Platform | Their angle | Signal to watch in our data |
|---|---|---|---|
| Vapi | Agents | Model-agnostic developer stack, $0.05 per minute plus providers at cost | Agent builders who test but never deploy; price-sensitive cohorts in burst pricing |
| Retell AI | Agents | Orchestration that resells third-party voices, ElevenLabs among them | Agents traffic arriving through partners instead of direct |
| Cartesia | Agents, API | Low-latency positioning, $0.06 per agent minute | Latency-sensitive API accounts reducing volume after incidents |
| Sierra, Decagon | Agents | Outcome-priced CX for enterprise support leaders | Enterprise deals priced on resolution, which minutes do not measure |
| OpenAI Realtime, Gemini Live | Agents, API | Native speech-to-speech from the model provider | Accounts that already use those LLMs as the agent brain |
| Deepgram, AssemblyAI | API (STT) | Speech-to-text first, priced below Scribe for batch | Scribe adoption among accounts that already use our TTS |
| OpenAI TTS, Google Chirp and Gemini TTS | API, Creative | Cheap, steerable TTS bundled with a cloud | Model mix shifting to Flash tiers; churn on price-led plans |
| Speechify | Creative (reader) | Larger consumer reader base | ElevenReader retention against reader-first habits |
| Suno, Udio | Creative (music) | Consumer music leaders; licensing still contested for Suno | Music as a second product for existing creators vs a new-user door |
| HeyGen, Synthesia | Creative (video) | Avatar video with voice as a feature | Dubbing users who also need lip-sync |
ElevenLabs v4 and v4 Turbo hold the top two places on the Artificial Analysis Speech Arena as read on 6 Oct 2026; the top slot changed hands several times in 2026. Rival prices from their public pricing pages.
Small inconsistencies across public pages, noted while building. Each is a metric-definition question in miniature: the same quantity counted three ways.
| Item | What the pages say | Likely explanation |
|---|---|---|
| Supported languages | Job postings: 70+. About page: 90+. Pricing table: 74. Voice Library post: 32. Docs: v4 90+, v3 70+, Multilingual v2 29, Flash v2.5 32. | The count depends on the model. "90+" follows v4; posting boilerplate predates it. A single canonical figure per model would remove the drift. |
| v4 launch credits | Pricing page banner: "3x credits included on Creator+ until October 12". Same page lower down: "Try v4 with 2x credits for two weeks". | Two promo texts live at once. Worth one owner, since promo terms feed straight into credit-burn metrics. |
| Legacy npm package | elevenlabs on npm is deprecated ("moved to @elevenlabs/elevenlabs-js") and still draws 274k downloads a week. | Expected migration tail. It has eased slowly, from about 25% of JS SDK downloads in May to 19% now, so a nudge may be worth testing. |
| JD duty | How I would do it | Detailed in |
|---|---|---|
| Define and own product health metrics across the three platforms | One metric tree with a separate activation and retention unit per platform, each with a written definition, grain, owner and the SQL that computes it. Candidates come from an aha search and are confirmed by a test. | §05, demo Metrics lab |
| Design, analyse and communicate experiments | Pre-register the primary metric, guardrails and duration from a power calculation; read SRM by segment before any result; Holm across secondaries; a one-page memo with the ship call. | §06, demo Experiment readout and Power planner |
| Partner with Analytics Engineering on the certified dbt KPI layer | Bring metric definitions as dbt-ready specs (grain, filters, tests such as uniqueness and accepted values) and let Analytics Engineering own the models. Reconcile against PostHog event counts before certifying. | §05, §07 |
| Turn ambiguous product questions into trusted analysis | Restate the question as a decision with a metric and a threshold before writing SQL. Push back when the ask has no decision attached. | §08 |
| Work with Growth data scientists on marketing- and revenue-adjacent questions | Share the activation definitions as the conversion event for marketing measurement, and paid conversion and credit burn with the finance side, so all three teams count the same user the same way. | §05 |
Starting hypotheses from public product mechanics. Each activation definition is a candidate until the aha search and an experiment back it.
| Platform | Activation candidate (7 days) | Retention unit | Engagement depth | Guardrail |
|---|---|---|---|---|
| ElevenCreative | Generated on 3+ distinct days and used a chosen voice (library, cloned or designed) | Weekly generation days; for paid plans, credit burn over a 60-day window to respect rollover | Products touched (speech, music, dubbing, video); share of quota used | Refund and chargeback rate; moderation flags |
| ElevenAgents | First agent deployed (live number, widget or API) after a test call | Weeks with live call minutes per workspace | Minutes per deployed agent; resolution or transfer rate | Burst-minute share; failed call starts |
| ElevenAPI | Key created plus sustained requests (for example 50+ successful calls) inside 7 days | Weeks with successful API volume per account | Model mix (v4, Flash, Scribe); SDK version currency | Error rate and p95 latency per account |
| Voice Library (two-sided) | Creator side: a shared voice cloned by at least one other user within 30 days | Creators with payouts in a month | Clones per voice; usage concentration (top 1% share) | Reported or removed voices |
Agents are built by teams. Retention at the user grain undercounts a workspace where one builder hands off to an operator.
A retention cell stays empty until every user in the cohort has been observed for the full week. The demo triangle does this.
Users who start in Creative and try Agents are the expansion path. Tag each user's first platform and track second-platform adoption as its own metric.
Five tests I would propose in the first quarter, each with a unit, a primary metric and the trap specific to it.
| Test | Unit | Primary metric | Guardrail | The trap |
|---|---|---|---|---|
| Template-first agent onboarding | Workspace | Agent deployed within 7 days | Creative activation for users who arrive with a Creative intent | Mobile assignment logging can leak; check SRM per platform before reading (the demo's planted bug). |
| SDK upgrade prompt on outdated versions | API key | Share of requests from SDK versions under 90 days old, 28 days after | Error rate after upgrade | A few CI keys dominate volume; cap or stratify by request volume. |
| Showing rolled-over credits in the app | Paying user | Paid retention at day 60 | Top-up revenue | Effects land after a full rollover cycle; the readout date has to be set before launch. |
| Featured shelf holdout in the Voice Library | Voice (cluster) | New-voice activation (first clone in 30 days) | Generation volume per user | Two-sided interference: randomise voices, read both sides, never split users only. |
| Annual plan default on the pricing page | Visitor | Paid conversion within 14 days | Weekly retention of new payers (the guardrail ElevenLabs has named for pricing tests) | Annual payers churn later by construction; compare like-for-like cash and retention windows. |
Assignment health, then the pre-registered primary, then guardrails, then secondaries with Holm. Inconclusive is a valid call; the memo states the smallest effect the test could have seen.
The v4 launch discount, which ends 12 October, is a natural experiment: compare v4 adoption and retention for accounts active before and after the price step against v3 users over the same weeks.
AI makes the first draft of a query cheap. The certified layer keeps the answer consistent.
| Rule | How it works in the demo and how I would run it |
|---|---|
| Show the SQL first | The demo's Ask box turns a question into read-only SQL with Workers AI and shows the query above the number, labelled as uncertified. |
| Point the model at certified views | The model sees the schema and the certified views (activation, weekly activity). Questions that match a certified metric should return that metric's SQL unchanged. |
| Read-only and bounded | Only SELECT or WITH statements run; anything else is rejected before execution. |
| Promote by review | A generated query that answers a recurring question becomes a dbt model through Analytics Engineering review. It does not become a dashboard on its own. |
Public sources only, read on 6 October 2026 (public job posting, 2026). Live adoption figures come from public package registries. The demo is my own prototype built for this application; its warehouse is synthetic and labelled as such.
Company: Series D · $500M ARR · Sept 2026 tender · creator payouts
Docs: models · changelog · about
Analytics stack: PostHog customer story (2024)
Adoption data: npm · PyPI · ClickPy
Benchmarks: Artificial Analysis Speech Arena
Competitors: Vapi · Retell · Cartesia · Deepgram · AssemblyAI
Independent homework for the ElevenLabs Data Scientist, Product Analytics role · 2026 · edwardtay.com