Skip to main content
Back
Netherlands4 months ago
N

Senior Technical Product Manager Interview Experience

Nebius·Senior / L5
Result
Rejected
Interview date
4 months ago
Timespan
2 weeks
Difficulty
Very difficult

Interview process

I had a good recruiter round. In the second product case round with HM, they asked me deep technical questions on AI

  • Recruiter screen
  • Final round

Interview tips

- Deep-dive on AI tools and concepts a lot - Brush up GPU optimisation and cloud platforms - Know difference between training and inference patterns

Company culture

Its a company where everyone is an ML engineer, irrespective of title (PM / Management all are ML engineers)

Questions asked

Specific questions asked

Design an inference batching system for a single GPU that can handle up to 100 inputs per batch while users wait synchronously, maximizing utilization under compute constraints.

Clarifying questions: Digital-native Companies Foundation model providers are not the target Geo: US, EU, UK Time : 1 Q Resourcing : couple of engineers Vision / Mission of Nebius: To provide AI cloud services in a scalable and vertically integrated manner to various types of customers ( foundational providers, enterprise, digital (startups incl), AI startups) Stage of product (Token Factory): Hypergrowth stage Goal of product (Token Factory): Goal of Batch Inference in Token Factory is to increase adoption of token factory amongst digital startups and digital companies. More details: To create a scalable, reliable, easy to use (low code / no code) service that can run process data and inference in batches Why does Nebius pursue Batch inference? Many digital companies have use cases where Cost is a limitation (can onboard customers with limited token spent) GPU Capacity is a limitation Use cases can tolerate latency Why can we give better cost for Batch inference? How to calculate the discount? GPU Capacity (idle GPUs are used) Segments Digital companies and digital startups Pain points Cost spent per use case (has limits) as companies have limited budget H H UX ( product should be usable by personas such as PM / PM/ BA, not just only MLE ) needs to low-code /no code and user friendly VH H Quality of responses M M Solution Product & User Experience Graphical user interface ( drag and drop style ) to enable everyone in digital companies to create Select open source LLM model with guidance on cost and TTFT Implement prefix caching / semantic caching Call batch inference as a module inside the UI (Call GenAI gateways) - MVP Vector store from embeddings - MVP Prompt-based training of the batch inference Test output anytime using a chatbot-style UI Technical & Operational Strategy Observability using another UI Allow them to measure E2E time for every request Allow user to see the response and the input Allow user to see tokens consumed Go to Market & Execution Adoption metric Number of users using batch inference UI product for min x days Number of live use cases running built using batch inference UI product

Unlock more real interview experiences

Get full access with a membership, or share your experience to try it free.

formerly Exponent

Get updates in your inbox with the latest tips, job listings, and more.

Follow Us

Products
Courses
Interview Questions
Interview Experiences
Popular articles
Guides
Coaching
For Partners
Company
Exponent Labs, LLC © 2026
Terms of Service | Privacy