Do you wish to view this page in English? Change language

Site Reliability Engineer – 12 month contract

  • 12-month contract (view to extend)
  • Hybrid 2 days on site (usually Mon & Tue)
  • Day rate
  • Dublin

We are looking for a Site Reliability Engineer II (SRE) to help ensure the reliability, availability, scalability, and performance of critical applications and infrastructure. You will collaborate with development and operations teams to improve system resilience, automate operational processes, monitor production environments, and support incident response.

Preferred Tools

  • AWS
  • Splunk / Dynatrace
  • ITIL Framework

Key Responsibilities

  • Ensure production systems are reliable, scalable, and highly available.
  • Partner with development teams to build resilient, fault-tolerant applications.
  • Automate operational tasks using scripting and infrastructure tools.
  • Monitor applications and infrastructure using observability solutions.
  • Investigate, troubleshoot, and resolve production issues, conducting root cause analysis and post-incident reviews.
  • Support capacity planning, performance optimisation, and disaster recovery initiatives.
  • Maintain operational documentation and promote best practices.
  • Contribute to change, incident, problem, and risk management processes.
  • Continuously improve system reliability through proactive monitoring and automation.

Required Skills

  • Experience with observability tools for monitoring metrics, logs, and traces.
  • Proficiency in scripting/programming (e.g., Python, Go, Bash).
  • Strong Linux/Unix system administration and networking knowledge.
  • Hands-on experience with cloud platforms, particularly AWS.
  • Understanding of high availability, scalability, and disaster recovery.
  • Familiarity with DevOps practices, CI/CD pipelines, containers, and orchestration.
  • Strong troubleshooting and root cause analysis skills.
  • Experience with capacity planning and performance tuning.
  • Knowledge of IT service management concepts, including incident, problem, and change management.