Role Summary:
We are seeking a talented and motivated Platform Engineer to join our core infrastructure team. You will play a critical role in designing, building, automating, and maintaining the foundational platform that powers our applications and services. Your primary focus will be on ensuring the overall health, reliability, scalability, and security of our cloud infrastructure and container orchestration systems, enabling our customers to ship features quickly and safely.
Key Responsibilities:
- Design, implement, and manage scalable and resilient infrastructure on both Azure and AWS cloud platforms.
- Develop, operate, and maintain Kubernetes clusters, ensuring high availability and performance.
- Utilize Docker for containerization best practices across the development lifecycle.
- Implement and manage Infrastructure as Code (IaC) using Terraform and CUE for provisioning and managing cloud resources consistently and reliably.
- Develop automation scripts and tooling to streamline operations and improve platform efficiency.
- Monitor platform health, performance, and cost using appropriate observability tools (logging, metrics, tracing).
- Implement and maintain CI/CD pipelines to support automated application deployments onto the platform.
- Troubleshoot and resolve issues across the platform stack, from cloud resources to containerized applications.
- Collaborate closely with software development teams and customers to understand their needs and provide a robust, self-service platform experience.
- Ensure security best practices are implemented and maintained across the platform.
- Document infrastructure design, operational procedures, and troubleshooting guides.
- Participate in on-call rotation to support platform availability.
Required Skills & Qualifications:
- Proven experience working as a Platform Engineer, DevOps Engineer, Site Reliability Engineer (SRE), or similar role.
- Strong hands-on experience with at least one major cloud provider (Azure or AWS), with experience in both being a significant plus.
- Deep understanding and practical experience managing Kubernetes environments (e.g., cluster setup, management, networking, security, deployments).
- Proficiency with Docker and containerization concepts.
- Strong experience writing, managing, and deploying Infrastructure as Code using Terraform.
- Experience with or strong interest in learning CUE for configuration management.
- Programming/scripting skills, particularly with Go, for building automation and platform tooling.
- Solid understanding of CI/CD principles and tooling (e.g., GitHub Actions, Azure DevOps).
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana).
- Good understanding of networking concepts (TCP/IP, DNS, HTTP, Load Balancing).
- Familiarity with Linux operating systems and shell scripting.
- Excellent problem-solving and troubleshooting abilities.
- Strong communication and collaboration skills.
Desired Skills & Qualifications:
- Experience building internal developer platforms or portals.
- Knowledge of service mesh technologies (e.g., Istio, Linkerd).
- Experience with GitOps principles and tools (e.g., Argo CD, Flux).
- Understanding of cloud security best practices and ISO27001 compliance.
- Experience with building custom Kubernetes controllers using Go.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
What We Offer:
- Competitive salary and benefits package
- Opportunity to work with modern, cloud-native technologies
- A collaborative and supportive team environment
- Opportunities for professional growth and development
- Flexible working arrangements
Darwin Recruitment is acting as an Employment Agency in relation to this vacancy.
Kiara Cumming