PRODUCT ANALYTICS HOMEWORK ·Data Scientist, Product Analytics, ElevenLabs
ⓘ Independent job-application page. Not affiliated with, endorsed by, or operated by ElevenLabs. Public information as of 6 October 2026.

Homework for my application to ElevenLabs Data Scientist, Product Analytics.

How I would define product health and run experiments across ElevenCreative, ElevenAgents and ElevenAPI, grounded in live public adoption data, plus a working desk that reads an A/B test, sizes one, runs certified metric SQL in the browser and measures SDK and Voice Library adoption.

2.24M
weekly PyPI downloads of the Python SDK (week to 4 Oct)
59%
of those installs on SDK versions older than 180 days
2
billing units to reconcile: credits and agent minutes
8
certified metrics running as SQL in the demo
00

Summary

The seat

The analyst behind discovery, adoption and retention for three platforms, working on top of a certified dbt KPI layer and with Product and Growth on experiments.

The hard part

One company, three usage shapes: creators spending a shared credit pool, businesses running agents by the minute, developers calling an API through SDKs. One "active user" definition cannot cover all three.

The work sample

A six-tab desk: experiment readout with SRM and Holm, a power planner with CUPED, a DuckDB metrics lab over a labelled synthetic warehouse, and live SDK and Voice Library adoption.

TakeawayWhat it means for product analytics
Activation differs per platformCreative activation is a habit (repeat generation with a chosen voice), Agents activation is a deployment, API activation is sustained request volume. Each needs its own aha search and a test before it is certified.
Credit rollover changes what retention meansUnused credits roll over for two months. A paying creator can skip a month and stay healthy, so a calendar-month "active" flag overstates churn risk.
The SDK long tail is an adoption ceiling59% of last week's Python SDK installs and 43% of JS SDK installs were on versions more than 180 days old. Any feature that needs a new SDK reaches developers slowly.
Read the test before the resultSample ratio checks by segment, a pre-registered primary metric, guardrails and a multiple-testing correction come before the ship call. The demo shows a planted logging bug caught this way.
01

ElevenLabs in 2026

A voice model company that now sells three platforms, with enterprise the larger half of revenue and agents the fastest-growing line.

$500M+
ARR, May 2026
55%
of revenue from enterprise, Sept 2026
15M
agent conversations a week, up 3x since February
$22M+
paid to 10,400+ Voice Library creators
800+
people, up from about 120 in Jan 2025
PlatformHow it is billedWho uses itWhat product health has to capture
ElevenCreativePlans from Free (10k credits) to Business (6M credits); one shared credit pool across speech, music, image, video and dubbing; two-month rollover capped at 3x quotaCreators, marketers, studios; web and mobile apps; Voice Library on both sidesHabit and breadth: repeat generation days, voices chosen, products touched, credit burn against quota
ElevenAgents$0.08 per call minute on every plan, $0.16 in burst; LLM usage passed through; telephony at costBusinesses building voice and chat agents; partners; enterprise deploymentsBuild to production: agent created, tested, deployed, then live minutes and resolution quality
ElevenAPIPer character or per minute by model (v4 $0.08 per 1K chars at list; 72% off until 12 Oct)Developers through the JS and Python SDKs and raw RESTIntegration depth: key created, first successful call, sustained volume, model mix, SDK version

Sources: ElevenLabs blog (Series D, $500M ARR, Sept 2026 tender, creator payouts), elevenlabs.io/pricing, /pricing/agents, /pricing/api, read 6 Oct 2026.

02

Developer adoption, from live data

Public package downloads are the one adoption signal anyone can read for ElevenAPI and ElevenAgents. Downloads include CI runs, so shares and trends matter more than levels. Week to 4 October 2026; the demo refreshes them live.

1.30M
@elevenlabs/client weekly (Agents), +129% in 12 weeks
1.15M
@elevenlabs/elevenlabs-js weekly, +82% in 12 weeks
2.24M
PyPI elevenlabs weekly, +6% in 12 weeks
80.7%
of JS SDK downloads on the scoped package; 274k a week still on the deprecated one
Version-lag measure, last 7 daysJS (npm)Python (PyPI)Reading
Installs on versions older than 180 days42.8%59.1%Most Python installs run code from before March. New model ids and fixes reach this group late or never.
Median installed version age147 days246 daysPython lags JS by about three months.
Most-installed single version2.68.0 (7.7%)2.1.0 (16.3%)Python 2.1.0 is from May 2025, 79 releases behind. One version holding a sixth of installs points to pinned CI or a popular template.
Versions above 0.1% of installs6667A wide support surface for any deprecation.
Why a product analyst cares

Feature adoption for an API feature has a ceiling set by the share of developers on an SDK that supports it. Version mix belongs in the denominator of every ElevenAPI adoption metric.

What I would check first

Join SDK version (from the user-agent header) to API keys and accounts. If the 2.1.0 tail is a few high-volume keys it is a CI artefact; if it is many small accounts it is an onboarding template still in circulation.

Sources: api.npmjs.org downloads and per-version endpoints, registry.npmjs.org, ClickPy (ClickHouse public PyPI dataset), pypi.org release dates. Live in the SDK adoption tab.

03

Competitive map

Each platform faces a different set of rivals, so each has a different switching risk to watch in retention data.

RivalPlatformTheir angleSignal to watch in our data
VapiAgentsModel-agnostic developer stack, $0.05 per minute plus providers at costAgent builders who test but never deploy; price-sensitive cohorts in burst pricing
Retell AIAgentsOrchestration that resells third-party voices, ElevenLabs among themAgents traffic arriving through partners instead of direct
CartesiaAgents, APILow-latency positioning, $0.06 per agent minuteLatency-sensitive API accounts reducing volume after incidents
Sierra, DecagonAgentsOutcome-priced CX for enterprise support leadersEnterprise deals priced on resolution, which minutes do not measure
OpenAI Realtime, Gemini LiveAgents, APINative speech-to-speech from the model providerAccounts that already use those LLMs as the agent brain
Deepgram, AssemblyAIAPI (STT)Speech-to-text first, priced below Scribe for batchScribe adoption among accounts that already use our TTS
OpenAI TTS, Google Chirp and Gemini TTSAPI, CreativeCheap, steerable TTS bundled with a cloudModel mix shifting to Flash tiers; churn on price-led plans
SpeechifyCreative (reader)Larger consumer reader baseElevenReader retention against reader-first habits
Suno, UdioCreative (music)Consumer music leaders; licensing still contested for SunoMusic as a second product for existing creators vs a new-user door
HeyGen, SynthesiaCreative (video)Avatar video with voice as a featureDubbing users who also need lip-sync

ElevenLabs v4 and v4 Turbo hold the top two places on the Artificial Analysis Speech Arena as read on 6 Oct 2026; the top slot changed hands several times in 2026. Rival prices from their public pricing pages.

★

Public-facts findings

Small inconsistencies across public pages, noted while building. Each is a metric-definition question in miniature: the same quantity counted three ways.

ItemWhat the pages sayLikely explanation
Supported languagesJob postings: 70+. About page: 90+. Pricing table: 74. Voice Library post: 32. Docs: v4 90+, v3 70+, Multilingual v2 29, Flash v2.5 32.The count depends on the model. "90+" follows v4; posting boilerplate predates it. A single canonical figure per model would remove the drift.
v4 launch creditsPricing page banner: "3x credits included on Creator+ until October 12". Same page lower down: "Try v4 with 2x credits for two weeks".Two promo texts live at once. Worth one owner, since promo terms feed straight into credit-burn metrics.
Legacy npm packageelevenlabs on npm is deprecated ("moved to @elevenlabs/elevenlabs-js") and still draws 274k downloads a week.Expected migration tail. It has eased slowly, from about 25% of JS SDK downloads in May to 19% now, so a nudge may be worth testing.
04

JD duties, my plan

JD dutyHow I would do itDetailed in
Define and own product health metrics across the three platformsOne metric tree with a separate activation and retention unit per platform, each with a written definition, grain, owner and the SQL that computes it. Candidates come from an aha search and are confirmed by a test.§05, demo Metrics lab
Design, analyse and communicate experimentsPre-register the primary metric, guardrails and duration from a power calculation; read SRM by segment before any result; Holm across secondaries; a one-page memo with the ship call.§06, demo Experiment readout and Power planner
Partner with Analytics Engineering on the certified dbt KPI layerBring metric definitions as dbt-ready specs (grain, filters, tests such as uniqueness and accepted values) and let Analytics Engineering own the models. Reconcile against PostHog event counts before certifying.§05, §07
Turn ambiguous product questions into trusted analysisRestate the question as a decision with a metric and a threshold before writing SQL. Push back when the ask has no decision attached.§08
Work with Growth data scientists on marketing- and revenue-adjacent questionsShare the activation definitions as the conversion event for marketing measurement, and paid conversion and credit burn with the finance side, so all three teams count the same user the same way.§05
05

Metric tree, per platform

Starting hypotheses from public product mechanics. Each activation definition is a candidate until the aha search and an experiment back it.

PlatformActivation candidate (7 days)Retention unitEngagement depthGuardrail
ElevenCreativeGenerated on 3+ distinct days and used a chosen voice (library, cloned or designed)Weekly generation days; for paid plans, credit burn over a 60-day window to respect rolloverProducts touched (speech, music, dubbing, video); share of quota usedRefund and chargeback rate; moderation flags
ElevenAgentsFirst agent deployed (live number, widget or API) after a test callWeeks with live call minutes per workspaceMinutes per deployed agent; resolution or transfer rateBurst-minute share; failed call starts
ElevenAPIKey created plus sustained requests (for example 50+ successful calls) inside 7 daysWeeks with successful API volume per accountModel mix (v4, Flash, Scribe); SDK version currencyError rate and p95 latency per account
Voice Library (two-sided)Creator side: a shared voice cloned by at least one other user within 30 daysCreators with payouts in a monthClones per voice; usage concentration (top 1% share)Reported or removed voices
Workspace grain for Agents

Agents are built by teams. Retention at the user grain undercounts a workspace where one builder hands off to an operator.

Censor young cohorts

A retention cell stays empty until every user in the cohort has been observed for the full week. The demo triangle does this.

One cross-platform view

Users who start in Creative and try Agents are the expansion path. Tag each user's first platform and track second-platform adoption as its own metric.

06

Experiment program

Five tests I would propose in the first quarter, each with a unit, a primary metric and the trap specific to it.

TestUnitPrimary metricGuardrailThe trap
Template-first agent onboardingWorkspaceAgent deployed within 7 daysCreative activation for users who arrive with a Creative intentMobile assignment logging can leak; check SRM per platform before reading (the demo's planted bug).
SDK upgrade prompt on outdated versionsAPI keyShare of requests from SDK versions under 90 days old, 28 days afterError rate after upgradeA few CI keys dominate volume; cap or stratify by request volume.
Showing rolled-over credits in the appPaying userPaid retention at day 60Top-up revenueEffects land after a full rollover cycle; the readout date has to be set before launch.
Featured shelf holdout in the Voice LibraryVoice (cluster)New-voice activation (first clone in 30 days)Generation volume per userTwo-sided interference: randomise voices, read both sides, never split users only.
Annual plan default on the pricing pageVisitorPaid conversion within 14 daysWeekly retention of new payers (the guardrail ElevenLabs has named for pricing tests)Annual payers churn later by construction; compare like-for-like cash and retention windows.
Readout order

Assignment health, then the pre-registered primary, then guardrails, then secondaries with Holm. Inconclusive is a valid call; the memo states the smallest effect the test could have seen.

When a test is not possible

The v4 launch discount, which ends 12 October, is a natural experiment: compare v4 adoption and retention for accounts active before and after the price step against v3 users over the same weeks.

07

Self-serve analysis with AI

AI makes the first draft of a query cheap. The certified layer keeps the answer consistent.

RuleHow it works in the demo and how I would run it
Show the SQL firstThe demo's Ask box turns a question into read-only SQL with Workers AI and shows the query above the number, labelled as uncertified.
Point the model at certified viewsThe model sees the schema and the certified views (activation, weekly activity). Questions that match a certified metric should return that metric's SQL unchanged.
Read-only and boundedOnly SELECT or WITH statements run; anything else is rejected before execution.
Promote by reviewA generated query that answers a recurring question becomes a dbt model through Analytics Engineering review. It does not become a dashboard on its own.
08

First 90 days

Days 1 to 30: learn the numbers
Read the certified dbt layer and the event taxonomy; list every current activation and retention definition in use.
Reconcile one metric end to end (signups) across PostHog, the warehouse and billing.
Audit the last ten experiments for SRM, power and guardrails.
Days 31 to 60: define
Run the aha search per platform and propose activation definitions with lift and coverage.
Ship the specs to Analytics Engineering as dbt models with tests.
Publish a pre-registration template and readout memo format.
Days 61 to 90: test
Run two of the §06 tests to a decision.
Weekly product health review: activation, retention and adoption per platform with confidence intervals.
Hand Growth and Finance one shared definition of an activated user.
09

Method & sources

Public sources only, read on 6 October 2026 (public job posting, 2026). Live adoption figures come from public package registries. The demo is my own prototype built for this application; its warehouse is synthetic and labelled as such.

Company: Series D · $500M ARR · Sept 2026 tender · creator payouts

Pricing: plans · agents · API

Docs: models · changelog · about

Analytics stack: PostHog customer story (2024)

Adoption data: npm · PyPI · ClickPy

Benchmarks: Artificial Analysis Speech Arena

Competitors: Vapi · Retell · Cartesia · Deepgram · AssemblyAI

Demo: elevenlabs-product-ds.leverlabs.workers.dev

Independent homework for the ElevenLabs Data Scientist, Product Analytics role · 2026 · edwardtay.com