Site Reliability Engineer III
We are seeking a highly skilled and experienced professional to join our team as a Site Reliability Engineer III. This role involves collaborating with cross-functional teams to ensure the reliability and performance of critical systems. The ideal candidate will have a strong technical background and a passion for improving system efficiency and reliability.
Required Skills & Qualifications
- Experience with APM tools such as DynaTrace.
- Familiarity with cloud platforms like AWS, Azure, or Google Cloud.
- Knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools.
- Knowledge in monitoring and logging tools such as Prometheus, Grafana, ELK stack, or Splunk.
- Prior experience designing and supporting Enterprise applications.
Preferred Skills & Qualifications
- Expertise in APM tools, i.e., DynaTrace.
- Strong problem-solving and troubleshooting skills, with the ability to analyze and resolve complex technical issues.
- Excellent communication and collaboration skills to work effectively with cross-functional teams.
- Understanding of networking principles and protocols (TCP/IP, HTTP, DNS, etc.).
- Strong attention to detail and ability to work in a fast-paced, dynamic environment.
- Prior work experience at client or in client's industry.
- Applicants must be able to work directly for Artech on W2.
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).
- Strong knowledge of Linux/Unix systems and command line tools.
- Proficiency in scripting languages such as Python, Shell, or Perl.
Day-to-Day Responsibilities
- Collaborate with cross-functional teams to define and establish service level objectives (SLOs) and service level agreements (SLAs) for critical systems.
- Monitor systems and applications, proactively identifying and resolving any performance bottlenecks or availability issues.
- Develop and maintain monitoring tools, alerts, and dashboards to provide visibility into system health and performance.
- Conduct post-incident analyses to identify root causes and implement preventive measures to avoid future incidents.
- Automate repetitive tasks and processes to improve efficiency and reduce manual intervention.
- Create and maintain documentation for system architecture, configuration, and troubleshooting procedures.
- Collaborate with development teams to implement and deploy new features and enhancements, ensuring they meet reliability and performance standards.
Company Benefits & Culture
- Comprehensive health, dental, and vision insurance.
- Flexible work schedule and remote work options.
- Opportunities for professional growth and development.