Performance & Reliability Engineer
Location: Richfield, MN Hybrid Duration: 8 Months of Contract
Key Responsibilities:
- Handling incidents and pagers for business-critical applications
- Identifying and resolving issues related to Dotcom applications, services and resources
- Monitoring system health and performance
- Perform system monitoring and log analysis
- Collaborate with developers and QA teams to resolve deployment defects
- Ensure compliance with security and change management standards
- Provide on-call and production support (24X7)
- Act as a technical subject-matter-expert in Dotcom technology and applications
Must Have Skills:
- 3–5 years of hands on experience in Site Reliability Engineering, DevOps, or Production Support
- Experience supporting production systems with high availability requirements
- Strong understanding of incident response, monitoring, and alerting
- Familiarity with SLOs, SLIs, and error budgets (practical application preferred)
- Basic automation/scripting skills (Python, Bash, or similar)
Frontend & API Awareness:
- Practical understanding of React/Next.js performance concepts
- Familiarity with GraphQL APIs and common performance issues
- Ability to troubleshoot frontend to backend request flows (not expected to build features)
Cloud & Platform:
- Experience working in AWS based environments
- Understanding of CDN behavior and caching strategies
- Exposure to microservices and distributed systems