Ensure high availability and uptime of Commodities Technology services and applications
Automate and streamline manual processes
Contribute to root cause analysis and post-mortem reports for production incidents
Liaise with the business-facing application support, development, and infrastructure teams to establish proper tracking of requirements and priorities
Monitor mission- critical systems to ensure service level objectives are met
What’s required
Proven work experience as a site reliability engineer or similar role.
Experience with Windows and Linux based operating systems, Cloud-based services and infrastructure (AWS), Infrastructure as Code (Terraform), and/or Configuration Management (Ansible)
Experience with container technologies such as Docker, Kubernetes, AWS, EKS, and ECS
Experience with observability tools such as DataDog
Proficient in coding and scripting with Python, PowerShell with an ability to comprehend C#
Self-motivated individual with great communication and interpersonal skills