Do you wish to view this page in English? Change language

Senior Site Reliability Engineer – 12-month contract

  • 12-month contract
  • Hybrid 2 days on site in South Dublin
  • Competitive Day rate

Oliver James’ client is seeking an experienced Senior Site Reliability Engineer (SRE) to ensure the reliability, scalability, availability, and performance of critical applications and services. The role will act as a production-readiness steward, working closely with development teams to build resilient, fault-tolerant, and scalable products. Key areas include operational design, automation, capacity planning, monitoring, observability, incident management, and continuous improvement.

You will support day-to-day operations through incident triage, root-cause analysis, business-impact assessment, and blameless post-mortems, while proactively identifying and addressing reliability risks throughout the development lifecycle.

Key Responsibilities

  • Drive production readiness, reliability, scalability, and operational excellence across applications and platforms.
  • Troubleshoot and resolve routine and complex production issues, including incident triage and root-cause analysis.
  • Develop automation and scripting to improve operational workflows and incident response.
  • Implement and improve monitoring, observability, alerting, and performance management.
  • Support capacity planning, performance optimisation, and disaster recovery.
  • Collaborate with development teams and stakeholders to embed SRE and DevOps best practices.
  • Contribute to operational standards, documentation, knowledge sharing, and continuous improvement.
  • Participate in change management, quality reviews, and blameless post-incident reviews.
  • Mentor colleagues and contribute technical expertise to projects and initiatives.

Key requirements

Top 3 Must-Have Skills

  • Cloud: AWS preferred/Azure/GCP: Hands-on experience deploying, managing, and troubleshooting cloud infrastructure and applications, with a strong focus on availability, scalability, security, and performance.
  • ITSM: Practical experience with incident, problem, and change management processes, preferably using BMC Helix.
  • Monitoring & Observability – Splunk / Dynatrace: Hands-on experience using Splunk and Dynatrace for monitoring, observability, troubleshooting, and application performance management.

Technical Skills

  • Cloud: AWS, Azure, or GCP; AWS preferred.
  • Observability: Splunk, Dynatrace, metrics, logs, traces, alerting, and APM.
  • Programming/Scripting: Python, Go, Bash, or similar.
  • Systems: Linux/Unix and networking.
  • DevOps: CI/CD, containerisation, orchestration, and automation.
  • Reliability: High availability, fault tolerance, disaster recovery, and scalability.
  • ITSM: Incident, problem, and change management.
  • Performance: Capacity planning, resource utilisation, and performance optimisation.
  • Troubleshooting: Systematic diagnosis and resolution of application, infrastructure, and network issues.

Experience Required

  • 7-8 years of relevant experience in SRE, DevOps, Production Engineering, Cloud Engineering, Infrastructure Engineering, or a related field.
  • Hands-on experience supporting production environments and business-critical applications.
  • Strong experience with cloud technologies, preferably AWS.
  • Experience with ITSM processes, preferably BMC Helix.
  • Practical experience with Splunk, Dynatrace, or similar monitoring and observability platforms.

Desirable Skills

  • Kubernetes and container platforms.
  • Infrastructure as Code and configuration management.
  • Strong automation and scripting experience.
  • Experience implementing SRE practices and production-readiness standards.
  • Experience with blameless post-mortems and continuous improvement.
  • Experience working in Agile environments.

Key Performance Measures

  • Application availability and reliability.
  • MTTD and MTTR.
  • Reduction in recurring incidents.
  • Monitoring and observability coverage.
  • Automation and operational efficiency.
  • Production readiness and successful change management.
  • Performance, scalability, and capacity optimisation.

Education & Certifications

  • Bachelor’s degree in Computer Science, IT, Engineering, or a related field preferred.
  • AWS, SRE, DevOps, ITSM, or other relevant certifications are advantageous.