Site Reliability Engineer
About the Role
WOXA GROUP is looking for a deeply technical and resilient Site Reliability Engineer to architect, build, and protect the reliability of the company’s complex hybrid and multi-cloud infrastructure.
In this role, you will act as both a builder and an uptime guardian, taking ownership of the infrastructure lifecycle across on-premise data centers and cloud platforms such as AWS, GCP, and DigitalOcean. You will work extensively with Kubernetes and Infrastructure as Code while defining reliability standards and improving production operations.
You will also play an important role in capacity management, business continuity, disaster recovery, incident management, and the company’s ISO 27001 compliance journey. The position requires someone who can reduce repetitive operational work through automation and promote continuous improvement through a blameless incident culture.
Responsibilities
Architect, build, and maintain reliable hybrid and multi-cloud infrastructure.
Take ownership of the complete infrastructure lifecycle and maintain production-system availability.
Manage infrastructure through Infrastructure as Code tools such as Terraform or OpenTofu.
Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
Manage workloads across on-premises data centers, AWS, GCP, and DigitalOcean.
Build and operate production workloads running extensively on Kubernetes.
Support the organization’s ISO 27001 compliance initiatives.
Own capacity management, business continuity, and disaster recovery planning.
Manage production incidents and support effective incident-response processes.
Replace repetitive operational tasks with automation.
Promote blameless post-incident reviews and continuous improvement after production failures.
Qualifications
At least 5 years of experience in Site Reliability Engineering, Cloud Infrastructure, DevOps, or a related field.
Experience leading or scaling technical teams responsible for 24/7 production environments.
Expert-level proficiency in building modular, scalable, and secure infrastructure using Terraform or OpenTofu.
Strong experience building and managing production Kubernetes clusters, including kubeadm, Amazon EKS, Google Kubernetes Engine, or similar platforms.
Strong understanding of service meshes and ingress controllers.
Extensive experience designing VPCs, IAM roles, databases, and network infrastructure across bare-metal, on-premise, and cloud environments.
Expert knowledge of modern monitoring, logging, tracing, and observability tools.
Strong programming or scripting skills in Go, Python, or Bash.
Ability to approach infrastructure and operational challenges as software-engineering problems.
Strong analytical and leadership skills.
Ability to communicate and negotiate with stakeholders when system stability must take priority over feature-delivery speed.
Work Information
Location: Khon Kaen, Thailand
Work Arrangement: On-site
Employment Type: Full-time
Experience: 5–8 years
Education: Bachelor’s degree or higher
Salary: Negotiable
Interested in this role?
Send your CV and tell us why you're a fit — we review every application.
Apply now