Site Reliability Engineer (SRE)
About the Role
We are looking for an experienced Site Reliability Engineer (SRE) to strengthen observability and operational resilience across a Microsoft Azure environment. This long-term Contract role will work closely with DevOps and engineering teams to establish monitoring standards, expand telemetry coverage, and improve service reliability across cloud-based platforms. The ideal candidate brings deep expertise in Azure operations, modern observability tooling, and production support, with the ability to turn data into actionable insight for faster troubleshooting and stronger system performance.
Responsibilities:
• Create and advance an observability framework for Azure-hosted systems and integrated third-party platforms, ensuring scalable monitoring coverage.
• Develop meaningful dashboards, alerting rules, log analysis views, and distributed tracing to provide actionable insight into application and infrastructure behavior.
• Utilize Azure services such as Azure Monitor, Log Analytics, Application Insights, Managed Prometheus, and Azure Managed Grafana to expand end-to-end visibility.
• Work alongside DevOps and software engineering teams to strengthen platform stability, incident response readiness, and service performance.
• Assess existing monitoring practices to uncover blind spots, reduce unnecessary alert volume, and support quicker root-cause identification.
• Improve insight into the health of applications, infrastructure components, and dependent services across production environments.
• Support reliability-focused engineering efforts by applying SRE principles such as service measurement, alert strategy refinement, and operational readiness improvements.
Requirements
• At least 5 years of hands-on experience working with Microsoft Azure in production environments.• Strong practical knowledge of observability platforms such as Datadog, Dynatrace, or comparable monitoring solutions.
• Demonstrated experience with Azure Monitor, Log Analytics, Application Insights, Prometheus, and Grafana.
• Background using GitHub Actions, Terraform, Kubernetes, and Infrastructure as Code practices.
• Solid understanding of metrics, logging, tracing, alerting, and reliability engineering for cloud-based systems.
• Ability to independently lead observability initiatives from technical design through implementation and optimization.
• Experience supporting live production systems, troubleshooting operational issues, and improving incident response processes.
• Bachelor’s degree in Computer Science, Information Technology, or a related discipline.
You'll be redirected to Robert Half's application page
Job Details
Salary
Not disclosed
Location
United States - Maumee OH
Job type
Temporary
Category
Site Reliability Engineering
Experience
5+ years
Posted
2w ago
Job Highlights
- 5+ years level role
- 100% Remote — open to candidates in United States
- Temporary position
About Robert Half
This job is hosted by Robert Half. Clicking Apply opens their site.
Remote Work Style
Mix of flexible and scheduled meetings
Your Match
See how well your skills line up with this role, and what you're missing.
AI Cover Letter
Generate a cover letter tailored to this job from your profile.