Deep Dive: North Star Metrics
Picking a north star metric sounds simple until you try to defend it. Most candidates can generate a plausible candidate. Fewer can explain why that candidate beats the alternatives, what it gets wrong, and what they'd watch alongside it. This lesson is about building that second layer.
Asked at Meta •
Asked at Stripe •
Asked at Meta • What interviewers are looking for
- Criteria-driven selection: Do you have a principled reason for choosing this metric, or did you just pick something familiar?
- Awareness of what the metric misses: Can you articulate where it might mislead, without being asked?
- Company-type pattern recognition: Do you understand why a media company and a marketplace would choose different north stars for similar goals?
- A specific, defined signal: Not a category ("engagement"), not a vague outcome ("user success"), but a measurement you could actually instrument.
What makes a strong north star metric
A north star metric is the single metric that best represents whether your product is achieving its core goal. It's the thing a PM team rallies around, the number that shows up in the weekly review, and the signal that tells you whether a big bet is working.
Three criteria determine whether a metric is strong enough to fill that role.
Criteria 1: Customer value
This is where most candidates start, and for good reason. The metric has to represent users actually achieving something meaningful, not just arriving.
A user opening the app is not customer value. A user completing a task, consuming content they sought out, or connecting with someone they care about is.
This matters because metrics that measure presence without purpose are easy to game. If you track weekly active users, you can move that number with push notifications, re-engagement emails, or friction in the offboarding flow. If you track users who complete at least one meaningful action per week, you have made it much harder to inflate the number without actually helping users.
Spotify's north star is time spent listening, not app opens. The distinction matters: a user who opens Spotify and leaves after 30 seconds counts in one metric but not the other. Time spent listening is a signal that users found something worth staying for.
Criteria 2: Business value
The metric has to connect to the company's mission and underlying business model.
This is why "time spent" means different things for different businesses. For an ad-supported media product, time spent directly predicts ad inventory and revenue. For a subscription productivity tool, what matters is not raw time but deep usage of the workflows that make the product hard to leave. A metric that captures customer value but floats free of the business model will not survive the executive review.
Slack tracks messages sent within a team, not logins. Message volume drives team stickiness, and team stickiness justifies the per-seat price. A team that logs in once a day to share a link is not a team that renews. A team that routes all coordination through Slack is.
If your north star metric increased by 20% while the business was still struggling, that is a sign the metric lacks business value. Run this test before committing to a metric in your answer.
Criteria 3: Team value
The metric has to be one that the product team can actually move.
A metric that responds to macroeconomic shifts, seasonality, or platform changes beyond the team's control is hard to learn from. You cannot run a proper experiment if the outcome variable is swayed by forces outside your control.
DoorDash tracks orders completed per week per active user, not overall GMV. GMV swings with restaurant partnerships and local market conditions: things the product team cannot touch in a standard sprint. Order frequency responds to improvements in experience: better search, better recommendations, and a smoother checkout. The team can run experiments against it and learn from the results.
The "so what" test
Before you commit to a metric in your answer, run this test. If the metric went up 10%, ask: Does that unambiguously mean the product is working?
If the answer requires caveats, it is probably not the right north star. "Daily active users went up 10%" could mean growth, a re-engagement campaign, a viral moment, or a competitor's outage. "Orders completed per active user went up 10%" means users are finding more value in the product. The latter tells a clear story.
This test also protects you against vanity metrics: numbers that look good but do not connect to anything real.
Top-line metrics by category
Before landing on a north star, use your diagnosis of the product's lifecycle stage and company type to narrow the candidate set. The table below maps common metric categories to the specific signals worth considering.
We’re aware that the common framework you may have heard of for use here is AARRR. In our research, it is less effective than what we have scoped below. This is one of the more opinionated tactics in our course, and one we’re quite bullish on.

A few notes on using this table well. First, "engagement" and "retention" are categories, not metrics. Any cell in the table that reads as a category needs one more level of specificity before it becomes a north star candidate. Second, the right category is determined by the product's lifecycle stage. A new product brainstorms from Adoption. A growth product brainstorms from Growth and Engagement. A mature product brainstorms from Retention and Monetization. Skipping this narrowing step is how candidates end up proposing an acquisition metric for a product that's been live for a decade.
North star patterns by company type
Different product categories have different north star conventions and, more importantly, different failure modes. The table below covers the most common types you'll encounter in interviews, with the key tradeoffs for each.
Knowing these patterns doesn't mean memorizing the "right answer" for every category. It means walking into a question about a media company and immediately knowing that time spent is the obvious candidate, that its addiction and passivity risks are well-documented, and that any strong answer needs to name those tradeoffs unprompted.

More in-depth view of this framework:

Applying the framework: Spotify
The question: "How would you define a north star metric for Spotify?"
Step 1: Narrow the candidate set by stage and category.
Spotify is a mature consumer media product. Its growth story is largely about engagement depth and retention, not new user acquisition. That points to the Engagement and Retention rows of the top-line table and to the Media row of the company type table.
Step 2: Generate candidates.
From the media category, the obvious candidates are time spent listening per month, number of streams per user per month, and number of songs saved or playlists created per user per month.
Step 3: Apply the customer, business, and team value
Time spent per month is the conventional media north star, and for good reason. It reflects customer value (users are choosing to listen), business value (more time spent drives ad impressions for free-tier users and justifies subscription price for paid users), and team value (the discovery, recommendations, and social features team can move this in a release cycle).
But time spent has a known weakness for Spotify specifically: it doesn't distinguish between active listening and background playing. A user who leaves Spotify running while they work may show high time spent without ever actively engaging with a recommendation, exploring new music, or taking an action that reflects genuine satisfaction.
Number of songs saved per user per week scores better on the customer value criterion. Saving a song is an intentional act that signals the user found something they wanted to keep. It maps reasonably well to business value, because users who are building a library are more likely to retain. And the recommendations and discovery teams can directly influence it.
Step 4: Land on a choice and name the tradeoff
Once you have run the three criteria, commit to a metric. Name it specifically and say why it wins over the alternatives you considered. Then, before the interviewer asks, surface the tradeoff yourself.
"I'd go with number of songs saved per user per week as my north star. It represents a meaningful user action: the user heard something new, decided it was worth keeping, and took a deliberate step. That is a stronger signal of customer value than passive play time. It maps to Spotify's business model because library depth is one of the strongest predictors of subscription retention. And the discovery and recommendations teams can directly influence it through Discover Weekly, Radio, and Daily Mixes. The tradeoff I'd name: it doesn't capture users who are satisfied but not actively building a library, for instance commuters who play the same playlists repeatedly. So I'd pair it with a counter metric: monthly listening hours per retained user, segmented by cohort. That combination tells me whether users are both discovering new music and staying engaged over time."
Notice that an earlier section of this lesson cites time spent listening as Spotify's north star, and that is also a defensible answer. Both can be correct. The right choice depends on what Spotify is optimizing for right now. A Spotify focused on growing new subscriber engagement will weight discovery signals more heavily. A Spotify focused on retention among long-tenured subscribers may care more about listening hours. Two candidates can give different north star metrics and both be excellent, as long as each answer is grounded in a clear reading of the company's current goals.
When your interviewer gives you a product and asks for a north star, they are not expecting a predetermined right answer. They are watching whether you can reason from a specific strategic context to a principled choice. Naming your assumptions about the company's current priority before you name the metric is a signal that you understand this.
How to defend your choice
The most important thing to understand about north star questions is that interviewers are not looking for a predetermined correct answer. They are watching how you reason when pushed. The skill being tested is whether you can hold a position with logic, update it with new information, and navigate apparent contradictions.
When the interviewer pushes back on your north star, the right move is not to immediately offer alternatives. It is to acknowledge the tension directly, explain what your metric does and does not capture, and describe the counter metric you would use to address the gap. That is the answer of a PM who has actually worked with metrics under pressure, not one who memorized a framework.
The thing most candidates miss: naming a tradeoff in your north star before the interviewer surfaces it is one of the clearest signals of senior-level thinking in an analytical round. It is not hedging. It is demonstrating that you understand the tool well enough to know where it breaks.
A complete answer picks a north star from the right category for the product stage and explains why it connects to the product goal using one or two of the three criteria. It acknowledges that tradeoffs exist when probed.
A senior+ answer applies all three criteria explicitly and names the tradeoffs of two or three candidates before landing on one. It also gets ahead of the conflicting-signal follow-up rather than waiting to be asked: "The risk I'd watch for is saves going up among new users but declining among users past the six-month mark, which would tell me we're winning at discovery but losing at long-term retention."
A senior+ answer treats the north star not as a final answer but as a decision that will need to be monitored and potentially revised.
Common pitfalls
Picking a metric because it sounds impressive, not because it fits. "Revenue" sounds like a serious answer for any mature product, but revenue is usually a lagging indicator that the PM team can't move directly. If you're asked for a north star and you pick revenue, you'd better be able to explain how the team influences it within a normal experiment window.
Not defining the metric precisely. "Number of active users" is not a metric. "Number of users who complete at least one listening session per day" is. If the interviewer has to ask you what counts as "active," you've already lost ground.
Treating the category as the north star. "Engagement" is not a north star. "Time spent" is not specific enough to be a north star unless you define the interval, the user population, and what counts as time spent. Get specific.
Ignoring one side of the marketplace. In any two-sided product, a north star that only captures consumer behavior misses supply health. If bookings are up but host satisfaction is declining, you're burning through supply. The north star doesn't have to capture both sides directly, but the counter metrics need to.
Never naming the tradeoff. Every metric has one. Candidates who present a north star with no caveats either haven't thought it through or are afraid that acknowledging weakness signals uncertainty. The opposite is true: naming the tradeoff clearly is what makes the choice feel earned.