JobsBackendCloud Engineer (Azure Platform Engineer)- Remote Ontario
Smile Digital Health

Smile Digital Health

·

Backend

Cloud Engineer (Azure Platform Engineer)- Remote Ontario

RemotemidPosted Aug 17, 2026
AzureKubernetesAKSTerraformHelmDockerGrafanaPrometheus+25 more

About this role

Design, deploy, and operate production-grade Azure Kubernetes Service (AKS) clusters with focus on networking, security hardening, and reliability. Act as SME for Kubernetes deployments and own platform end-to-end from architecture through incident response.

Must have

Expertise in Azure Kubernetes Service cluster design and operations

Strong experience with Terraform infrastructure as code

Proficiency with Helm charts and containerized deployments

Deep knowledge of AKS networking including Azure CNI and ingress controllers

Experience with observability tools including Grafana Prometheus and Loki

Familiarity with GitOps workflows using Flux or ArgoCD

Understanding of security hardening network policies and vulnerability management

Technologies

AzureGitDockerRESTLinuxShellTerraformKafkaNginxGrafanaPrometheusHelmKubernetesCI/CDAzure DevOpsFHIRArgoCDAzure StorageGitOpsRBACAKSFluxAzure MonitorLog AnalyticsApplication InsightsLokiTempoAzure SQLAzure Key VaultAzure Container RegistryAzure PolicyAzure CNITraefik

Responsibilities

Act as SME for Kubernetes deployments and production issues across all environments

Design and deploy AKS clusters with private configurations managed identities and RBAC

Own health scaling and lifecycle management of production AKS clusters

Configure AKS networking including Azure CNI load balancers and ingress controllers

Design integrations between AKS and Azure services like ACR Key Vault and Monitor

Manage containerized deployments using Docker and Helm with reusable standards

Harden AKS environments through policy enforcement and network policies

Own container and cluster vulnerability management and remediation coordination

Contribute to Terraform-based infrastructure as code for Azure resources

Support CI/CD pipelines including GitOps workflows to reduce deployment risk

Design and manage observability across Grafana stack and Azure-native tooling

Collaborate with teams to diagnose and resolve performance bottlenecks

Contribute to disaster recovery and business continuity planning procedures

Provide escalation support for production incidents and participate in on-call rotation

Document runbooks post-incident reviews and operational knowledge base

Benefits

Flexible time offTraining and career developmentEmployee Assistance ProgramHealth insuranceRetirement benefits with matchRemote work

Amenities

Remote work environmentFlexible time away including PTO personal and sick daysLife and disability coverage

Recruitment process

1

Company ranked 19 on Deloitte's Technology Fast 50 for 2024

2

FHIR-based health data platform used in over 20 countries

3

New role created to support continued growth and operational excellence

4

AI may be used in portions of recruitment process with human oversight

5

Core values include respect inclusion and celebrating differences

6

FHIR Study Program and Skillsoft Learning available

7

Super HAPI Fun Club employee group