Lead Affirm's Resilience Engineering team to ensure production system safety through chaos engineering and load testing. Build platforms for controlled production experimentation with strong safeguards and observability.
Proven experience leading engineering teams in reliability, infrastructure, or distributed systems
Hands-on experience with production load testing, chaos engineering, or large-scale system validation
Experience with chaos engineering vendors such as Gremlin, Harness, or similar
Strong understanding of failure modes in distributed systems
Experience building systems with strong safety guarantees
Familiarity with cloud-native environments and observability tooling
Strong programming background in Python, Kotlin, Java, or similar
Define and drive the vision for resilience engineering with focus on load testing and chaos engineering
Lead and mentor engineers building platforms for safe production experimentation
Partner with infrastructure, product, and security leadership to embed resilience validation
Own design and evolution of platforms for controlled production load testing and fault injection
Ensure safeguards including isolation boundaries, approval workflows, and automated rollback
Build systems for end-to-end observability, traceability, and auditability of experiments
Drive reliability improvements by identifying weaknesses through testing and experiments
Work with engineering teams to safely design and execute production experiments
Enable teams to adopt resilience practices through reusable tooling and frameworks
This posting is for an existing vacancy
Pay Grade - P, Equity Grade - 7
New employees typically start at the beginning of the pay range
Affirm is a remote-first company with majority remote roles
Total compensation includes monthly stipends for health, wellness and tech spending
153,000 – 213,000 CAD
/ year
40,000 – 60,000 USD
/ year
100,000 – 150,000 CAD
/ year
108,000 – 135,000 CAD
/ year
64,000 – 84,000 EUR
/ year