Lead Platform & Infrastructure Engineer
DockerKubernetesK3sHelmLinuxTerraformAnsiblePackerGitHub ActionsGitLab CI
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Platform & Infrastructure Engineer based in the United States. This role offers the opportunity to lead the infrastructure foundation behind scalable, secure, and highly available software and AI-driven solutions. You’ll design and operate modern platforms spanning cloud, single-tenant, and customer-managed environments. The position combines platform engineering, infrastructure automation, container orchestration, security, and operational excellence. You’ll establish Infrastructure-as-Code standards and deployment practices that make environments more consistent, reliable, and scalable. Working closely with engineering, security, data, and AI/ML teams, you’ll help deliver resilient platforms for demanding production workloads. You’ll also play a key role in observability, incident response, capacity planning, disaster recovery, and production readiness. This is a high-impact opportunity for an experienced infrastructure leader who enjoys solving complex platform challenges and improving how technology is delivered and operated. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Platform & Infrastructure Engineer based in the United States. This role offers the opportunity to lead the infrastructure foundation behind scalable, secure, and highly available software and AI-driven solutions. You’ll design and operate modern platforms spanning cloud, single-tenant, and customer-managed environments. The position combines platform engineering, infrastructure automation, container orchestration, security, and operational excellence. You’ll establish Infrastructure-as-Code standards and deployment practices that make environments more consistent, reliable, and scalable. Working closely with engineering, security, data, and AI/ML teams, you’ll help deliver resilient platforms for demanding production workloads. You’ll also play a key role in observability, incident response, capacity planning, disaster recovery, and production readiness. This is a high-impact opportunity for an experienced infrastructure leader who enjoys solving complex platform challenges and improving how technology is delivered and operated. Accountabilities Design, implement, and maintain scalable platform infrastructure supporting cloud, single-tenant, and customer-managed deployment environments. Lead infrastructure architecture and operational practices for secure, highly available production platforms. Build and manage Infrastructure-as-Code solutions, automated provisioning, configuration management, deployment automation, and environment lifecycle processes. Develop and operate containerized platforms using Docker, Kubernetes, K3s, Helm, and related technologies. Establish and maintain CI/CD pipelines, release automation, environment promotion strategies, rollback procedures, and modern DevOps practices. Partner with security teams to implement encryption, network segmentation, secrets management, certificate management, access controls, and infrastructure-hardening standards. Drive platform observability through monitoring, logging, alerting, capacity planning, reliability practices, and operational readiness. Support production deployments, incident response, troubleshooting, and on-call activities while maintaining high standards of service reliability. Contribute to disaster recovery, resilience, and infrastructure continuity strategies. Collaborate with engineering, security, data, and AI/ML teams to ensure infrastructure effectively supports evolving application and platform requirements. Identify opportunities to automate repetitive operational processes and continuously improve platform reliability, scalability, and efficiency. Requirements 8+ years of experience building, deploying, and operating enterprise infrastructure and platform solutions, including leadership experience in Platform Engineering, Infrastructure Engineering, DevOps, or Site Reliability Engineering. Deep hands-on expertise with Docker, Kubernetes, K3s, Helm, Linux, networking , and production-grade container orchestration. Strong experience with Infrastructure-as-Code and automation technologies, including Terraform, Ansible, Packer , configuration management, automated provisioning, and deployment pipelines. Proven experience designing and operating secure, highly available infrastructure, including secrets management, certificate management, observability, disaster recovery, and operational resilience. Experience with CI/CD platforms and modern DevOps practices, including GitHub Actions, GitLab CI, Gitea Actions , release automation, monitoring, and infrastructure lifecycle management. Strong understanding of infrastructure security, access controls, network architecture, encryption, and platform hardening. Demonstrated ability to troubleshoot complex production environments and respond effectively to incidents and operational challenges. Strong technical leadership and communication skills, with the ability to collaborate across engineering, security, data, AI/ML, and other technical functions. Experience supporting regulated industries, customer-managed or on-premises environments, object storage, GPU infrastructure, or AI/ML platform operations is preferred. Strong analytical, problem-solving, prioritization, and continuous-improvement mindset. Benefits Full-time, fully remote position within the United States. Opportunity to lead infrastructure and platform engineering initiatives supporting modern software and AI-driven solutions. High-impact technical leadership role spanning cloud, on-premises, customer-managed, and containerized environments. Exposure to Kubernetes, Infrastructure-as-Code, CI/CD, automation, security, observability, and AI/ML infrastructure. Opportunity to establish engineering standards and influence platform architecture and operational practices. Collaboration with multidisciplinary teams across engineering, security, data, and AI/ML. Opportunity to work on complex enterprise infrastructure challenges with a strong focus on reliability, scalability, and security. Compensation and additional benefits are not specified in the provided job description. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
Not disclosed
Location
United States
Job type
Full-time
Category
Platform Engineering
Experience
8+ years
Posted
Today
Job Highlights
- 8+ years level role
- 100% Remote — open to candidates in United States
- Full-time position
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.