Site Reliability Engineer
V2Soft is a global leader in IT services and business solutions, delivering innovative and cost-effective technology solutions worldwide since 1998. We have headquartered in Bloomfield Hills, MI and have 16 offices spread across six countries. We partner with Fortune 500 companies to address complex business challenges. Our services span AI, IT staffing, cloud computing, engineering, mobility, testing, and more. Certified with CMMI Level 3 and ISO standards, V2Soft is committed to quality and security. Beyond our work, we actively support local communities and non-profits, reflecting our core values. Join us to be part of a dynamic and impactful global company!
Position Description: Employees in this job function are responsible for ensuring availability, reliability and performance of cloud and network systems and services by automating routine manual tasks
Key Responsibilities:
- Collaborate with Infrastructure teams in implementing critical solutions by automating routine tasks
- Monitor and manage production environments, proactively identifying and resolving issues.
- Participate in building advanced tooling for system access monitoring, log session recording, administration of reliability across multiple geographically distributed data centers.
- Engage with engineering teams to improve on-call efficiencies, drive incident management and post-mortem analysis.
- Perform capacity planning and optimization to support growing demands and traffic patterns.
- Maintaining, monitoring and alerting systems for proactive system health checks.
- Continuously improve system performance, stability, and security through data-driven analysis and optimization.
- Facilitate knowledge sharing by creating and maintaining comprehensive documentation & diagrams
Skills Required: Big Query, Dynatrace, GCP
Skills Preferred: GCP Cloud Run, Python, Troubleshooting (Problem Solving)
Experience Required: Engineer 2 Exp.: Practitioner: 1 coding language or framework. 4+ years in IT; 3+ years in development; Hands-on experience with Google Cloud Platform (GCP). Proficiency with monitoring/observability tools, ideally Dynatrace (or comparable, e.g., Datadog, New Relic). Familiarity with ITSM tools such as ServiceNow (incident, problem, change management)
Experience Preferred: Familiarity with the use of AI tools – agents, skills, LLMs, copilot. Experience defining and tracking SLAs/SLOs/SLIs
Education Required: Bachelor's Degree
Additional Information : Hybrid – Monday -Thursday in the office
Site Reliability Engineer in dearborn at Unknown Company
This position is listed as full time and hybrid.