Own and operate Yelp's real-time streaming platform built on Kafka and Flink at massive scale. Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure across multi-cloud environments. Work with distributed systems handling 300M+ reviews and 100K+ photo uploads daily.
Solid SRE or infrastructure engineering foundation with IaC and cloud platforms
Production experience with Kafka at scale including upgrades and capacity planning
Programming proficiency in Python, Java, or similar for tooling and automation
Strong debugging and systems-thinking skills across distributed systems
Own reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments
Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery
Partner with engineering teams to enable streaming use cases and ensure data pipeline reliability
Troubleshoot complex issues affecting data flow, performance, or stability and lead root cause analyses
Execute Kafka version upgrades and platform migrations with minimal disruption
Participate in follow-the-sun on-call rotations with geographically distributed SRE teams
Fully remote opportunity across all locations in Canada
Role posted to fill an existing position
Yelp engineering culture values individual authenticity and creative problem solving
New engineers deploy working code in their first week
Systems handle over 300 million business reviews and 100,000 photo uploads daily
186,368 β 223,642 CAD
/ year
184,000 β 356,500 USD
/ year
168,000 β 304,750 USD
/ year
124,000 β 195,500 USD
/ year
135,482 β 227,700 USD
/ year