Site Reliability Engineer (Remote)

Copenhagen, Capital Region
Posted 11 hours ago
Engineering

About the role

Job summary

The role involves setting the technical direction for reliability across a cloud-based infrastructure, focusing on the platform that provisions and manages AI agents for payment processing globally. The position requires a senior-level individual contributor who will drive architectural decisions and establish reliability standards across engineering teams.

Qualifications

  • Minimum of 7 years of experience in site reliability engineering or related fields.
  • Proven experience with event-driven architecture and messaging systems, including message queues like Kafka or RabbitMQ.
  • Deep knowledge of AWS services such as EC2, VPC, IAM, S3, and RDS.
  • Proficiency in Infrastructure as Code tools like Terraform or Pulumi.
  • Experience with Kubernetes and Docker in production environments.
  • Strong skills in observability tools and defining SLOs and error budgets.
  • Hands-on experience with chaos engineering and resilience testing.
  • Solid debugging skills for distributed systems and experience with SQL and NoSQL databases.
  • Advanced proficiency in English, both written and spoken.

Responsibilities

  • Define and implement the reliability strategy and standards across engineering teams.
  • Drive architectural evolution and make key technology decisions for the platform.
  • Design and manage the messaging layer for inter-service communication.
  • Automate cloud infrastructure provisioning and ensure scalability.
  • Build monitoring and alerting systems to maintain platform health.
  • Lead incident response and conduct postmortems to improve reliability.
  • Foster a culture of chaos engineering to identify and mitigate potential failures.

Skills

  • Technical leadership and mentorship capabilities.
  • Strong understanding of distributed systems and debugging techniques.
  • Familiarity with AI and MLOps infrastructure is a plus.

Education

  • Relevant degree in Computer Science, Engineering, or a related field is preferred but not explicitly required.

Tools

  • Experience with Datadog or similar observability tools.
  • Familiarity with chaos engineering tools like Gremlin or AWS FIS.
Full Access

Ready to apply for this role?

Full Access gives you the company name, full job description, and a direct link to apply. On the 1- and 3-month plans, CV Tailor rewrites your CV for this exact role.

Share this job