Skip to main content
All Questions

Design an inference batching system for a single GPU that can handle up to 100 inputs per batch while users wait synchronously, maximizing utilization under compute constraints.

How would you batch and process inputs? How would you return responses? Trade-offs? Points of failure? Scaling?

Interview experiences

5 shared
N
Nebius · 4 months ago
"I had a good recruiter round. In the second product case round with HM, they asked me deep technical questions on AI"
Product Manager · Senior / L5
Anthropic · 5 months ago
"Some of the interviews are very formulaic. When they ask about management experience they first start by asking what is the largest org you've managed, and then all following questions are about that org. The largest org I had managed was at a bank which doesn't have the greatest parallels to draw from. If I had better understood the format I would have answered with the bank, but explicitly stated that it's not a good analogy lets talk about this slightly smaller org because it's closer to you guys. Instead, this conversation happened at the end."
Engineering Manager
Anthropic · a year ago
"I applied on the careers site with no referral, and the recruiter was in my inbox quickly, so the whole process started fast. The loop was a 15-minute recruiter call, then a phone design round, then a five-round onsite with system design, coding, project deep dive, behavioral, and a separate culture round. Everybody was responsive and genuinely nice, and the process felt way more efficient than most big tech loops, but the bar also felt extremely high. The AI-looking technical questions were mostly normal infra questions if I abstracted them correctly, but the coding round blindsided me because it was a practical project exercise, not the multithreaded coding problem I had prepared for. I came out feeling good about everything except coding, and I ended up rejected with zero feedback."
Software Engineer · Staff / L6
Anthropic · a year ago
"I applied, and did a normal recruiter screen followed by a coding phone screen. After I passed that, hiring paused for the team, and then time passed until a different recruiter picked me back up for a virtual onsite. My onsite was a coding round, a system design round, a hiring manager round, a culture round, and a one-person technical presentation. At the time the question bank was small enough that I basically knew the kinds of questions coming, but Anthropic was one of my first loops, so I wasn't as sharpened as I should have been. I got pretty close but didn't get the offer, and my read is that the system design round hurt me the most."
Software Engineer · Staff / L6
Anthropic · a year ago
"I interviewed for an infrastructure software engineer role that was basically staff-scoped even though the title stayed generic. The process was a recruiter screen, a live coding phone screen, then two onsite loops: system design, coding, culture, and later experiences/goals plus a technical project deep dive. The phone screen was the only question I had seen before online. Everything else felt genuinely novel and very tied to GPU infrastructure and Anthropic's safety culture. The culture interview in particular was unusually interrogative, and they even warned me in advance that some questions might make me uncomfortable. I felt strongest in the project deep dive and weakest in the later coding round."
Software Engineer · Staff / L6

Community answers

No answers contributed by the community yet.

Related courses

Course

AI Engineering Interview Prep

Ace the AI engineering interview at top companies like Google, Anthropic, Meta, and OpenAI. Learn from detailed frameworks, example questions, and mock interviews with senior engineers. Practice system design, AI/ML concepts, AI-assisted coding, and behavioral rounds.

Course

Generative AI Interviews

Ace your interviews at AI-first and AI-powered companies. Learn to explain how Gen AI really works, where it fails, and how to use it responsibly. Understand how top tech companies build and deploy AI products and how generative AI works in production, from LLMs and RAG systems to real-world deployment.