Senior Site Reliability Engineer II

relx groupPlano, TX

today

Occupations

Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems Administrators

Industries

Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesCustom Computer Programming Services
APPLY NOW

About the role

Overview In this role you design, build, and operate highly available systems in AWS, owning observability and automation to improve reliability. You will work closely with application teams to boost performance, deployability, and incident resilience. You’ll lead incident response and post-incident reviews, driving prevention and automation. The role offers hybrid or fully remote options and emphasizes ownership of production systems within a culture of automation and continuous improvement. Compensation / Benefits Competitive salary and comprehensive benefits Flexible work location with hybrid or fully remote options Real ownership of production systems and reliability outcomesA culture that values automation, learning, and continuous improvement Responsibilities Design, build, and operate highly available, scalable systems in AWSWrite, maintain, and review Terraform to provision and manage infrastructure Own and improve monitoring, alerting, and observability using Grafana, Pingdom, and Uptrends Participate in a rotating on-call schedule, responding to production incidents and driving issues to resolution Lead incident response, root cause analysis, and post-incident reviews with a focus on prevention and automation Define and manage SLOs, SLIs, and error budgets Build and improve CI/CD pipelines and operational workflows using Azure Dev Ops and Git Hub Work directly with application teams to improve reliability, performance, and deployability Automate manual operational tasks to reduce toil Maintain clear, actionable runbooks and documentation in Confluence Track work, incidents, and operational improvements using Jira and Service Now Mentor other engineers and help set SRE standards and best practices Key requirements 5+ years of hands-on experience in SRE, Dev Ops, or Infrastructure Engineering roles Strong production experience in AWSSignificant hands-on experience with Terraform in real-world environments Experience operating monitoring and uptime platforms such as Grafana, Pingdom, and Uptrends Strong Linux systems, networking, and troubleshooting skills Experience supporting production systems through incident response and on-call rotations Proficiency with Git Hub and modern Git workflows Experience building or maintaining CI/CD pipelines with Azure Dev Ops Familiarity with ITSM and incident workflows using Service Now Strong written communication skills with experience documenting systems and processes in Confluence Ability to work independently in a remote or hybrid environmentstrong written communicationindependence in remote or hybrid environmentsmentoring others TerraformAWSGrafana

Matching similar jobs

JOB OVERVIEW

Experience level

Lead

Location

Plano, TX

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

today

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE
Senior Site Reliability Engineer II at relx group | Johnson Jobs