Jobgether

Lead SRE

Jobgether

AWSAzureTerraformAnsibleJenkinsGitHub ActionsLinuxDockerKubernetesAppDynamics

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead SRE based in India. This role offers the opportunity to lead and scale a high-performing Site Reliability Engineering function supporting critical production systems. You will be responsible for strengthening reliability, scalability, performance, and operational resilience across cloud-native platforms. The position combines technical leadership with strategic planning, team development, automation, observability, and incident management. You will guide SRE teams across multiple time zones while partnering closely with engineering, product, and infrastructure stakeholders. A major focus will be on embedding SRE best practices, reducing operational toil, and driving continuous improvement. You will help establish a culture of operational excellence, learning, and reliable service delivery across a global technology environment. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead SRE based in India. This role offers the opportunity to lead and scale a high-performing Site Reliability Engineering function supporting critical production systems. You will be responsible for strengthening reliability, scalability, performance, and operational resilience across cloud-native platforms. The position combines technical leadership with strategic planning, team development, automation, observability, and incident management. You will guide SRE teams across multiple time zones while partnering closely with engineering, product, and infrastructure stakeholders. A major focus will be on embedding SRE best practices, reducing operational toil, and driving continuous improvement. You will help establish a culture of operational excellence, learning, and reliable service delivery across a global technology environment. Accountabilities: Lead, mentor, and develop a high-performing team of SREs working across multiple time zones, fostering strong collaboration, technical excellence, and continuous learning. Execute the SRE roadmap in alignment with broader business and engineering objectives, ensuring reliability priorities are clearly defined and delivered. Partner with engineering, product, and infrastructure teams to improve system reliability, scalability, performance, and operational resilience. Drive adoption of SRE principles, including Service Level Indicators, Service Level Objectives, error budgets, reliability metrics, and data-driven operational practices. Lead capacity planning, performance optimization, disaster recovery, incident response, root cause analysis, and blameless postmortems for critical systems. Champion automation and infrastructure improvements to reduce operational toil and increase engineering efficiency. Build and maintain CI/CD pipelines, observability capabilities, infrastructure-as-code solutions, and operational tooling while ensuring alignment with security, privacy, and regulatory requirements. Establish and monitor key metrics covering system health, service reliability, operational performance, and team effectiveness. Requirements: Hold a bachelor’s degree in Computer Science, Engineering, or a related field, or demonstrate equivalent practical experience. Bring 10+ years of experience across software engineering, DevOps, or Site Reliability Engineering, including at least 3 years in a technical leadership or people-management capacity. Demonstrate proven experience managing large-scale, distributed systems within cloud-native environments such as AWS or Azure. Possess strong expertise in monitoring and observability technologies such as AppDynamics and Splunk, as well as automation tools including Terraform and Ansible. Have strong experience designing and managing CI/CD pipelines using technologies such as Jenkins and GitHub Actions. Demonstrate deep knowledge of Linux systems, networking, containers, Docker, Kubernetes, and modern infrastructure engineering practices. Possess excellent communication, collaboration, leadership, and stakeholder-management skills, with the ability to influence teams across a complex organization. Experience working in a global, matrixed organization and contributions to open-source SRE or DevOps tools are considered valuable. Benefits: Opportunities to lead and shape a growing SRE function within a global, technology-driven environment. Professional development opportunities, including structured leadership and manager-development programs. Access to mentorship opportunities supporting technical growth, leadership development, and career progression. Generous paid time off and Volunteer Time Off opportunities, subject to applicable eligibility requirements. Competitive healthcare and employee benefits, including medical and dental coverage. Access to a Global Employee Assistance Program providing additional support and resources. Opportunities to participate in Employee Impact Groups and employee experience initiatives that encourage connection, inclusion, collaboration, and community involvement. The opportunity to work with distributed teams across multiple time zones while contributing to critical reliability and infrastructure initiatives. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

Not disclosed

Location

India

Job type

Full-time

Category

Site Reliability Engineering

Experience

10+ years

Posted

Today

Job Highlights

  • 10+ years level role
  • 100% Remote — open to candidates in India
  • Full-time position

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.