Engineering Manager, SRE
KubernetesAWSPostgreSQLCI/CDInfrastructure as Codereliability engineeringDockerTerraformGitLab CIGitHub Actions
About the Role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Spain. This is a hands-on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform. You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges. The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering. You will shape how the team balances operational excellence, incident response, reliability improvements, and longer-term engineering initiatives. A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization. You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership. The environment is fully remote and asynchronous, offering significant autonomy in a fast-growing, globally distributed organization. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Spain. This is a hands-on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform. You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges. The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering. You will shape how the team balances operational excellence, incident response, reliability improvements, and longer-term engineering initiatives. A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization. You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership. The environment is fully remote and asynchronous, offering significant autonomy in a fast-growing, globally distributed organization. Accountabilities Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring. Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus. Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly. Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed. Design and maintain effective support rotations and on-call processes while strengthening incident response practices. Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform. Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post-incident improvements. Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations. Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership. Remain hands-on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalate. Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early. Continuously improve team health, collaboration, conflict resolution, and retrospective practices. Requirements: Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team, with direct responsibility for performance and career development. Strong hands-on background in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review designs, challenge implementation decisions, and contribute during production incidents. Production experience with Kubernetes and AWS at meaningful scale, including the operational realities of running cloud infrastructure. Hands-on experience building, enabling, or scaling AI infrastructure and working with AI-related engineering workloads. Strong understanding of observability principles and practices, infrastructure as code with Terraform, and CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins. Experience with Docker, shell scripting, and production infrastructure operations. Proven ownership of reliability practices including incident response, on-call operations, SLOs, error budgets, and turning incidents into lasting engineering improvements. Experience working in regulated environments, with an understanding of infrastructure controls, compliance, and security requirements. Exceptional prioritization skills, particularly when operational workloads compete with project commitments. Excellent written communication and documentation skills, with the ability to lead effectively in a highly distributed and asynchronous environment. Strong relationship-building, collaboration, conflict-resolution, and stakeholder-management capabilities. A coaching-oriented leadership style, with evidence of developing engineers both technically and professionally. Strong judgment, accountability, adaptability, curiosity, and commitment to high-quality execution. Nice-to-have experience with Elixir, Java, Clojure, Node.js, Python, or another backend programming language. Additional desirable experience includes OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora performance optimization, connection pool management, query tuning, Linux systems administration, security, FinOps, and cloud cost management. Experience growing an engineering team from a small base and establishing a strong hiring bar is advantageous. Ability to work effectively across global teams and time zones. Benefits: Annual salary range of USD $75,450–$169,700 , with actual compensation determined by location, experience, skills, training, business needs, and market conditions. Fully remote, work-from-anywhere environment. Flexible working hours within an asynchronous work culture. Flexible paid time off. 16 weeks of paid parental leave . Budget for coworking spaces, learning, and wellness activities, including gym memberships. Mental health support services. Stock options. Home office budget and IT equipment. Global exposure through collaboration with colleagues across multiple continents. Opportunities to travel internationally and meet colleagues at company events. A high-autonomy environment where employees are encouraged to organize their schedules around their lives and personal commitments. Opportunity to influence the maturity of reliability engineering practices while working on complex infrastructure and platform challenges. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
You'll be redirected to Jobgether's application page
Job Details
Salary
$75K–$170K
Location
Spain
Job type
Full-time
Category
Site Reliability Engineering
Experience
Senior
Posted
Today
Job Highlights
- $75K–$170K salary
- Senior level role
- 100% Remote — open to candidates in France, Germany, Ireland, Netherlands, Spain, Switzerland, United Kingdom
About Jobgether
This job is hosted by Jobgether. Clicking Apply opens their site.
Remote Work Style
Mixed
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.