Platform Operations Engineer
LinuxKubernetesCI/CDTerraformAWSAzureGoogle CloudDatadogBashHelm
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Platform Operations Engineer based in United States. This role provides first-line operational support for a production SaaS platform within a DevOps/SRE organization. You will help ensure reliable day-to-day operations across production and software delivery environments. Your responsibilities will include monitoring systems, responding to alerts, managing deployments and builds, and troubleshooting operational issues. You will work with Linux, Kubernetes, cloud platforms, CI/CD pipelines, infrastructure-as-code, and observability tools. The position is designed for an engineer who can work independently while knowing when and how to escalate complex incidents. You will collaborate with globally distributed engineering teams and participate in scheduled operational coverage and on-call responsibilities. As you grow in the role, you will have opportunities to take on broader infrastructure, automation, reliability, and SRE responsibilities. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Platform Operations Engineer based in United States. This role provides first-line operational support for a production SaaS platform within a DevOps/SRE organization. You will help ensure reliable day-to-day operations across production and software delivery environments. Your responsibilities will include monitoring systems, responding to alerts, managing deployments and builds, and troubleshooting operational issues. You will work with Linux, Kubernetes, cloud platforms, CI/CD pipelines, infrastructure-as-code, and observability tools. The position is designed for an engineer who can work independently while knowing when and how to escalate complex incidents. You will collaborate with globally distributed engineering teams and participate in scheduled operational coverage and on-call responsibilities. As you grow in the role, you will have opportunities to take on broader infrastructure, automation, reliability, and SRE responsibilities. Accountabilities: Monitor production and SaaS environments and proactively identify potential operational issues. Respond to monitoring alerts, conduct initial investigations, and perform appropriate remediation. Execute routine application and platform deployments using established processes and tooling. Create, trigger, and monitor software builds and release workflows. Handle incoming operational requests from engineering and other internal teams. Take ownership of operational requests through resolution or ensure effective handoff to the appropriate team. Follow documented procedures and runbooks for common operational tasks and incidents. Perform first-line troubleshooting using logs, metrics, Kubernetes tooling, and other diagnostic information. Escalate complex or higher-risk incidents to senior DevOps/SRE or development engineers when appropriate. Initiate incident or war-room coordination when required and ensure the appropriate technical teams are engaged. Perform known and approved production remediation activities, such as restarting or scaling workloads, when appropriate. Participate in follow-the-sun operational coverage and provide clear handoffs for active incidents, deployments, and unresolved requests. Participate in an on-call rotation and fulfill assigned operational coverage responsibilities. Create, maintain, and improve operational documentation and runbooks based on recurring issues and evolving procedures. Identify opportunities to improve operational processes, documentation, and recurring troubleshooting workflows. Progressively take on greater ownership across infrastructure, Kubernetes, CI/CD, automation, observability, and reliability engineering as experience develops. Requirements Approximately 1–3 years of experience in DevOps, SRE, cloud operations, infrastructure operations, production support, or a related technical role. Strong entry-level candidates with relevant hands-on experience and solid technical fundamentals may also be considered. Hands-on experience with Linux and Kubernetes. Working familiarity with most of the following: Helm, Git, CI/CD pipelines, deployment workflows, Terraform, infrastructure-as-code concepts, public cloud platforms, monitoring and logging systems, networking fundamentals, and Bash scripting. Experience with at least one major cloud platform such as AWS, Azure, or GCP; transferable cloud and infrastructure fundamentals are valued over experience with a specific provider. Familiarity with monitoring, logging, and alerting tools such as Datadog or similar platforms. Basic understanding of networking and technical troubleshooting concepts. Familiarity with databases, storage, IAM, DNS, and cloud networking is helpful but not required. Comfortable working with production systems and following controlled operational procedures. Ability to investigate technical issues, collect useful diagnostic information, and recognize when escalation is appropriate. Strong written and verbal communication skills, particularly for incident documentation, technical handoffs, and escalations. Ability to work independently during assigned shifts while collaborating effectively with a globally distributed engineering organization. Willingness to participate in on-call responsibilities and provide operational coverage as required. Strong learning mindset and willingness to develop deeper DevOps/SRE expertise over time. No specific degree or professional certification is required. Must already be authorized to work in the United States, as visa sponsorship is not specified for this position. Benefits 100% remote work opportunity. Structured operational coverage designed to support teams across regions. Participation in an on-call rotation, with on-call arrangements compensated separately or supported through time off in lieu according to the source role terms. Opportunity to develop hands-on experience across Linux, Kubernetes, cloud infrastructure, CI/CD, Terraform, monitoring, and production operations. Clear career growth path toward broader DevOps and Site Reliability Engineering responsibilities. Increasing opportunities to work on infrastructure, automation, observability, reliability engineering, and production architecture. Exposure to a globally distributed engineering environment. Opportunity to contribute to improved operational processes, runbooks, and reliability practices. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
Not disclosed
Location
United States
Job type
Full-time
Category
DevOps / SRE
Experience
1–3 years
Posted
Today
Job Highlights
- 1–3 years level role
- 100% Remote — open to candidates in United States
- Full-time position
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.