Lead Platform Reliability Engineer, Global AI Platform & Solutions
Manulife — Toronto, Ontario
Posted 2026-08-25
The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of the shared platform that supports internal AI solution development. It combines software engineering, SRE practices, and operations to keep the platform reliable and developer-friendly. Position Responsibilities: Reliability and performance: Define SLOs/SLIs, track operations budgets, reduce MTTR, capacity plan, and tune autoscaling. Observability: Build and maintain logging, metrics, tracing, and alerting; instrument platform components; create runbooks and dashboards. Incident response: On-call for platform incidents; triage, mitigate, root-cause, and drive postmortems and corrective actions. Automation and tooling: Develop self-service capabilities, AIOps/MLOps/GitOps/CICD pipelines, and operational automations (provisioning, upgrades, backups). Infrastructure as code: Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift. Security and compliance: Enforce identity/RBAC, secrets management, supply chain security, and regulatory controls; collaborate with risk and audit. Scalability and cost: Optimize resource usage, plan capacity, contr