Cloud Support Engineer (AWS Infrastructure & Operations)
Maintain AWS infrastructure for high availability and performance, diagnose and resolve complex AWS issues, and automate operational tasks across EC2, S3, VPC, IAM, RDS and CloudWatch. May include on-call cover for critical production incidents outside standard hours.
Engineering · Malaysia / Indonesia · Full-time
About the role
We are seeking a high-potential Cloud Support Engineer to support, operate, and optimise cloud infrastructure environments, primarily on AWS. This role is critical in ensuring high availability, performance, and security of production workloads while delivering excellent technical support to customers.
The position combines hands-on engineering, operational excellence, and customer-facing support, making it ideal for candidates who are passionate about cloud technologies, problem-solving, and continuous learning. We welcome both strong fresh graduates with exceptional logical thinking and early-career engineers with relevant cloud experience.
Note: this role may require on-call support and handling of critical production incidents outside standard working hours.
Cloud infrastructure management
- Manage and maintain AWS cloud infrastructure to ensure high availability, reliability, and performance.
- Implement and maintain Infrastructure as Code using tools such as CloudFormation, CDK, or Terraform.
- Monitor and optimise cloud resources for scalability, performance, and cost efficiency.
- Apply best practices for resource provisioning, tagging, and lifecycle management.
Cloud support and operations
- Provide end-to-end technical support by diagnosing and resolving complex AWS infrastructure issues across multiple services.
- Own and manage the incident response lifecycle, including troubleshooting, escalation, and resolution during critical outages.
- Proactively monitor cloud environments to detect, prevent, and resolve issues before impact.
- Collaborate with internal teams and customers to troubleshoot and resolve infrastructure-related challenges.
- Utilise tools such as AWS Systems Manager Patch Manager to automate OS-level patching for Windows and Linux environments.
- Participate in on-call rotations or off-hours support for critical incidents when required.
- Develop and maintain runbooks, documentation, and knowledge base articles to improve operational efficiency.
Performance optimisation and reliability
- Analyse system metrics, logs, and performance data to identify bottlenecks and inefficiencies.
- Implement tuning strategies across compute, storage, and networking layers.
- Drive improvements in system reliability, observability, and operational excellence.
Security and compliance
- Enforce cloud security best practices: IAM policies with least-privilege access, network segmentation (VPC, subnets, security groups), and encryption at rest and in transit.
- Monitor and respond to security alerts and vulnerabilities.
- Support compliance with relevant standards such as PDPA, GDPR and HIPAA where applicable.
- Ensure responsible and secure usage of cloud resources.
Continuous improvement and automation
- Identify opportunities to automate repetitive operational tasks using scripting and cloud-native tools.
- Contribute to improving support processes, patching strategies, and operational workflows.
- Support adoption of emerging AWS capabilities, including exposure to Generative AI tools and services.
Qualifications
- One to three years of experience in cloud engineering, system administration, or infrastructure support. Strong fresh graduates with relevant skills are encouraged to apply.
- Solid understanding of AWS core services, including EC2, S3, VPC, IAM, RDS and CloudWatch.
- Basic understanding of networking concepts: VPC, subnets, routing, and DNS (Route 53).
- Familiarity with Infrastructure as Code: CloudFormation, CDK, or Terraform.
- Proficiency in at least one scripting language: Python, Bash, or similar.
- Understanding of cloud security fundamentals: access control, encryption, and network isolation.
- Strong analytical thinking and problem-solving skills, with the ability to troubleshoot systematically.
- Good communication skills and the ability to work collaboratively in a team environment.
- Willingness to learn, adapt, and operate in a fast-paced environment.
Preferred
- AWS certifications: AWS Certified Solutions Architect (Associate) or AWS Certified SysOps Administrator.
- Hands-on experience with monitoring tools (CloudWatch, logging systems) and AWS Systems Manager, including Patch Manager.
- Exposure to multi-cloud environments such as Azure or GCP.
- Familiarity with automation and DevOps practices.
- Exposure to AWS Generative AI services or frameworks is an added advantage.
What we look for
- Strong ownership mindset in resolving issues end to end.
- Ability to operate under pressure during production incidents.
- Passion for cloud technologies, automation, and continuous improvement.
- Attention to detail in documentation, troubleshooting, and execution.
- Growth mindset with a commitment to learning and skill development.
Why Axrail?
AWS Premier Projects
Work on cutting-edge cloud/AI implementations.
Career Accelerator
Grow into Program Management or Technical PM roles.
Innovation Culture
Experiment with GenAI tools for project delivery.
Hybrid Flexibility
Balance office collaboration with remote work.
Ready to apply?
Submit your application and tell us why this role fits you. Attach whatever shows your work best: resume, GitHub profile, code samples, links to personal projects.