Senior Machine Learning Engineer, Personalization, Magenta
You never pay to apply. Your application happens on jobs.lever.co, the employer's official hiring system.
The Personalization team makes deciding what to play next easier and more enjoyable for every listener. From Blend to Discover Weekly, we’re behind some of Spotify’s most-loved features. We built them by understanding the world of music and podcasts better than anyone else. Join us and you’ll keep millions of users listening by making great recommendations to each and every one of them.
The Sessions Department within Personalization is building a portfolio of agentic and conversational products that define how hundreds of millions of people discover and experience audio, such as prompted Playlists or DJ, all powered by a single layer that understands music, culture, and the user’s taste
You'll join a team of four engineers actively building the agent and the strategy behind it. We work closely with the broader Sessions organization on one of the most highly-leveraged bets at Spotify right now: making it possible to have natural language conversation with Spotify across the entire app!
The team moves fast by staying hyper-focused: we pick a focused set of problems, ship new features to users weekly, and learn in the wild. We constantly dogfood our product and learn from users' data and feedback to find the most important next thing to build or improve, together.
What You'll Do
- You'll build and improve the core agentic capabilities that power the agent behind Talk to Spotify (memory, context management, multi-step tool use)
- You'll design and calibrate evaluation frameworks (including LLM-as-judge) that accelerate our confidence in the agent's behavior, and increase our offline-to-online success
- You'll work in a very dynamic space: the team prototypes, dogfoods, ships, learns, and refines in tight loops with real users, as our understanding of the problem and users' expectations of agentic products and Spotify evolve
Who You Are
- You're excited by agentic experiences — building agents, evaluating agents, and the hard problems in between (context handling, multi-step reasoning, ambiguity at scale)
- You like getting your hands dirty: shipping quickly, testing ideas against real usage, and learning from the wild rather than over-indexing on offline evaluation
- You have 5+ years of production ML experience deploying highly impactful products, or equivalent experience in other roles with a deep ML background
- You know how to evaluate ML systems rigorously — designing metrics, building eval pipelines, judge alignment, and can develop intuition through dogfooding and looking at user behavior
- You're comfortable debugging the messy interactions between models, tools, and system constraints like latency
Where You'll Be
- This role is based in New York
- We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home
How this role compares
- Restricted to named countries, like 1,542 of the 4,122 live roles we track (37%).
- 74 of the 467 AI and Data roles we track publish a pay range. This one does not.
- Posted 83 days ago, with 365 newer AI and Data roles arriving here since.
- Spotify has 34 live roles here across 7 fields.
Counted from the 4,122 remote roles live on CitizenHire right now, so these figures move as the market does. All AI and Data roles · How we verify jobs
Get every job like this, the day it opens
We email you every remote job we index, every day, with where you can apply from marked on each one. Free, and one click to leave.
Free · unsubscribe in one click