Senior Software Engineer, Site Reliability Engineering
AWSLinuxDNSTLSHTTP/STCP/IPMicroservicescloud-based systemsCapacity Planningobservability
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer, Site Reliability Engineering based in Canada. This is a senior-level opportunity to help build and operate the reliable, secure, and highly scalable infrastructure behind a high-impact digital platform. You’ll work across the technology stack, from Linux systems and cloud infrastructure to distributed applications and platform services. The role combines hands-on engineering with architectural leadership, enabling product and infrastructure teams to deliver resilient systems at scale. You’ll collaborate closely with engineers across product development, developer experience, and backend infrastructure. Your work will directly influence system availability, performance, observability, and the overall engineering experience. You’ll also help shape platform capabilities, improve operational practices, and anticipate future capacity and reliability needs. This is an ideal environment for an experienced SRE who enjoys solving complex technical problems and creating high-leverage engineering solutions. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer, Site Reliability Engineering based in Canada. This is a senior-level opportunity to help build and operate the reliable, secure, and highly scalable infrastructure behind a high-impact digital platform. You’ll work across the technology stack, from Linux systems and cloud infrastructure to distributed applications and platform services. The role combines hands-on engineering with architectural leadership, enabling product and infrastructure teams to deliver resilient systems at scale. You’ll collaborate closely with engineers across product development, developer experience, and backend infrastructure. Your work will directly influence system availability, performance, observability, and the overall engineering experience. You’ll also help shape platform capabilities, improve operational practices, and anticipate future capacity and reliability needs. This is an ideal environment for an experienced SRE who enjoys solving complex technical problems and creating high-leverage engineering solutions. Accountabilities Design, develop, and maintain software and infrastructure that improve service availability, scalability, performance, and operational efficiency. Establish architectural direction for infrastructure and platform services while providing technical guidance and support to engineering teams. Build and improve tools, processes, and systems for deployment, infrastructure, service, and change management. Troubleshoot and resolve complex production issues across the software development lifecycle, with a focus on minimizing downtime and service disruption. Develop and evolve platform capabilities that enable engineering teams to build, deploy, operate, and observe services more effectively. Conduct capacity planning and demand forecasting to anticipate system growth, identify performance bottlenecks, and proactively address scalability challenges. Instrument, operate, and monitor distributed microservices and cloud-based systems to maintain strong reliability and observability. Participate in a rotating on-call schedule and contribute to incident response, service recovery, and continuous reliability improvements. Partner with cross-functional engineering teams and stakeholders to identify opportunities, balance technical trade-offs, and deliver high-impact platform solutions. Requirements 5+ years of experience managing infrastructure and systems, ideally within large-scale or distributed production environments. Extensive hands-on expertise with AWS and Linux-based systems. Strong ability to read, write, debug, and maintain production-facing software and systems. Deep understanding of large-scale distributed systems and web technologies, including DNS, TLS, HTTP/S, TCP/IP, and related networking concepts. Demonstrated experience operating, instrumenting, and observing distributed microservices in production cloud environments. Strong problem-solving skills, with the ability to break down complex technical challenges and make thoughtful trade-offs based on business and engineering impact. Experience designing resilient, scalable infrastructure and platform services with a focus on availability, performance, and reliability. Strong communication and collaboration skills, with the ability to work effectively with technical and non-technical stakeholders across different levels of an organization. Ability to operate independently while contributing effectively to a collaborative, cross-functional engineering environment. A proactive mindset and strong ownership of production systems, operational excellence, and continuous improvement. Benefits Remote work opportunity from eligible locations in Ontario and British Columbia, Canada. Expected total cash compensation of CAD $180,200–$233,200 , depending on location, qualifications, skills, competencies, and experience. Opportunity to work on large-scale, distributed systems with significant impact across the engineering organization. Collaborative environment with exposure to product engineering, developer experience, backend infrastructure, and platform teams. Opportunity to influence architectural direction and shape the evolution of engineering infrastructure and platform capabilities. Participation in meaningful reliability, scalability, observability, and infrastructure initiatives. Commitment to diversity, inclusion, and equal employment opportunity. Reasonable accommodations available throughout the recruitment process for candidates who require them. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
C$180K–C$233K
Location
Canada
Job type
Full-time
Category
Site Reliability Engineering
Experience
5+ years
Posted
Today
Job Highlights
- C$180K–C$233K salary
- 5+ years level role
- 100% Remote — open to candidates in Canada
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.