Staff Software Engineer, Infrastructure
GoTerraformKubernetesGitOpsArgo CDGrafanaEKSOpenTelemetryPrometheusCI/CD
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Software Engineer, Infrastructure based in United States . This is a Staff-level infrastructure engineering role focused on building the platform foundations that enable engineering teams to move faster and operate more reliably. You’ll combine hands-on software engineering with technical leadership, shaping self-service infrastructure, cloud platforms, deployment systems, and operational tooling. A major focus will be replacing expert-driven provisioning and operational workflows with paved roads that provide clear ownership, safe defaults, strong guardrails, and measurable adoption. You’ll help evolve multi-region, cross-account infrastructure and Kubernetes foundations while improving reliability, security, scalability, and cost efficiency. The role also includes developing AI-assisted operational workflows that reduce toil while keeping production changes safe, auditable, and human-reviewed. You’ll influence architecture across teams through RFCs, design reviews, technical standards, and pragmatic engineering decisions. Working in a small, growing team, you’ll have substantial ownership and the opportunity to drive infrastructure investments through real production adoption. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Software Engineer, Infrastructure based in United States . This is a Staff-level infrastructure engineering role focused on building the platform foundations that enable engineering teams to move faster and operate more reliably. You’ll combine hands-on software engineering with technical leadership, shaping self-service infrastructure, cloud platforms, deployment systems, and operational tooling. A major focus will be replacing expert-driven provisioning and operational workflows with paved roads that provide clear ownership, safe defaults, strong guardrails, and measurable adoption. You’ll help evolve multi-region, cross-account infrastructure and Kubernetes foundations while improving reliability, security, scalability, and cost efficiency. The role also includes developing AI-assisted operational workflows that reduce toil while keeping production changes safe, auditable, and human-reviewed. You’ll influence architecture across teams through RFCs, design reviews, technical standards, and pragmatic engineering decisions. Working in a small, growing team, you’ll have substantial ownership and the opportunity to drive infrastructure investments through real production adoption. Accountabilities: Turn ambiguous infrastructure challenges into clear technical proposals and drive them through RFCs, architecture reviews, and cross-team alignment. Design and build self-service platform capabilities and APIs, primarily in Go, covering onboarding, provisioning, deployment, observability defaults, and day-2 operations. Establish reliable delivery standards using Terraform, GitOps with Argo CD, progressive delivery, automated testing, and continuous deployment practices. Evolve multi-tenant EKS infrastructure to improve reliability, security, scalability, and cost efficiency. Develop and improve ingress and traffic-routing capabilities, including Envoy Gateway and multi-region, cross-account connectivity. Strengthen SLOs, alerting, incident response, and operational follow-up through Grafana Cloud and improved observability practices. Measure success through outcomes for consuming engineering teams, including faster provisioning and deployment, greater self-service, and improved operational reliability. Develop AI-assisted operational workflows such as alert enrichment, incident context gathering, runbook-assisted diagnosis, remediation recommendations, and onboarding assistants. Maintain appropriate human oversight for AI-assisted operational actions, with an emphasis on safety, auditability, and responsible automation. Participate in the on-call rotation after onboarding and shadowing, while helping improve the overall health of on-call through better alerts, runbooks, automation, and blameless postmortems. Lead strategic platform initiatives from initial design through production adoption and establish durable technical patterns across engineering teams. Requirements: 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering. Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Strong software engineering expertise in Go or a comparable programming language, including system design, testing, debugging, code review, and long-term maintainability. Proven experience designing, delivering, and operating cloud services or infrastructure platforms in production. Deep expertise in at least one area such as Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms. Strong Linux, networking, and production operations fundamentals. Experience setting technical direction and leading initiatives that require alignment across multiple engineering teams. Strong written and verbal communication skills, particularly in remote environments and through RFCs, design documents, and incident writeups. Experience with EKS, ingress, CNI, or service-mesh technologies is valuable. Familiarity with OpenTelemetry, Prometheus, Grafana, CI/CD, progressive delivery, GitHub Actions, Argo CD, and canary deployments is a plus. Experience leading large-scale migrations, platform adoption programs, or cross-team infrastructure initiatives is beneficial. Strong systems judgment, curiosity, pragmatic decision-making, and the ability to develop deep expertise while navigating adjacent technical domains. Willingness to participate in an operational on-call rotation and contribute to improving its effectiveness. Visa sponsorship may be considered on a case-by-case basis depending on business needs. Benefits: CA$238,250–CA$382,250 + equity for Canada-based candidates. Remote-first work arrangement. Flexible scheduling and autonomy in managing your working hours. Generous paid time off, quarterly wellness days, and an end-of-year wellness break. Home-office support to help create an effective remote workspace. Technology stipend equivalent to US$100 net per month . Annual learning and development stipend covering conferences, courses, certifications, and continued professional learning. 16 weeks of paid parental leave after six months of employment. Equity participation for full-time employees. Medical, retirement, and paid-holiday benefits, with details varying by country. Opportunities to lead major infrastructure initiatives involving self-service provisioning, multi-region networking, continuous deployment, Kubernetes, and platform engineering. Significant technical ownership within a small, growing infrastructure team. Exposure to AI-assisted and agentic operational workflows, with a focus on safe and auditable automation. Fully remote collaboration with distributed engineering teams. Offices available in Seattle and Paris for connection and collaboration. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
C$238K–C$382K
Location
United States
Job type
Full-time
Category
Infrastructure Engineering
Experience
8+ years
Posted
Today
Job Highlights
- C$238K–C$382K salary
- 8+ years level role
- 100% Remote — open to candidates in Canada, United States
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.