Experience Required
n10 - 20 years
nMinimum Education Required
nBachelor Degree
nCompensation
n$150,000.00 - $170,000.00 / Yearly
nHours Per Week
n40
nNumber Of Positions
n1
nShift
nFirst Shift (Day)
nJob Description
nJob Description
nAs a Principal Site Reliability Engineering, you will be responsible for building a SRE practice, monitoring and performance engineering best practices which will be aligned to our agile teams to help drive availability, resiliency and stability of Quest products, platform and services.
nYou are an engineering technical leader who has a passion for reliability and have a wide breath of experience. Ideally, you will have had experience as a Site Reliability and Observability Engineer where you made significant improvements to the products/services/platforms and customer experience. You will also partner with architecture, engineers, security, and operations to design and build reusable patterns to deploy reliable and resilient solutions.
nYou will also have responsibility to attract, retain and grow top SRE engineering talent, providing guidance and mentorship to team members.
nYou will bring empathy, humility, and a continuous learning mindset to every interaction. You are motivated to innovate and create, to always do the right thing, and to improve both what we build and how we build it.
nPay Range: $150,000-170,000, plus yearly bonus (New Jersey)
nSalary offers are based on a wide range of factors including relevant skills, training, experience, education, and, where applicable, certifications obtained. Market and organizational factors are also considered. Successful candidates may be eligible to receive annual performance bonus compensation.
nRemote: This position supporting Epic can be 100% remote if not located near a hub location within certain criteria.
nBenefits Information: We are proud to offer best-in-class benefits and programs to support employees and their families in living healthy, happy lives. Our pay and benefit plans have been designed to promote employee health in all respects physical, financial, and developmental. Depending on whether it is a part-time or full-time position, some of the benefits offered may include:
nDay 1 Medical, supplemental health, dental & vision for FT employees who work 30+ hours
nBest-in-class well-being programs
nAnnual, no-cost health assessment program
nBlueprint for Wellness
nhealthyMINDS mental health program
nVacation and Health/Flex Time
n6 Holidays plus 1 MyDay off
nFinFit financial coaching and services
n401(k) pre-tax and/or Roth IRA with company match up to 5% after 12 months of service
nEmployee stock purchase plan
nLife and disability insurance, plus buy-up option
nFlexible Spending Accounts Annual incentive plans
nMatching gifts program
nEducation assistance through MyQuest for Education Career advancement opportunities and so much more!
nResponsibilities:
nExperience in transforming an organization by designing and implementing SRE capabilities, including monitoring, performance and chaos engineering. You will set the strategy for overall Site Reliability Engineering (SRE)/Development alignment
nLead initiatives to implement service levels (SLIs, SLOs, SLAs) and error budgets. You will initiate, influence and drive SRE within the organization and work with product and service teams to enable this model.
nProvides guidelines/patterns and establishes proper metrics for building highly scalable, reliable, high performing systems
nStrategizes best in class monitoring frameworks to accomplish end to end flow monitoring and meaningful alerting.
nCoaches and mentors' teams of monitoring, performance and SRE engineers.
nProven ability to implement processes, solutions and engineering capabilities at scale.
nPrior experience in large scale digital technologies, where uptime and continuous availability was core to the business.
nStrong acumen of public cloud and / or private cloud implementation and application adoption
nStrong understanding of Cloud, API, Event Driven, and Microservices technologies for large scale environments.
nInfluences other leaders, principals, and engineers opening the discussion and adoption for implementing SRE best practices.
nBuilds relationships with other leaders and groups across the company, providing understanding of SRE concepts and value.
nWork with other team leads to identify improvements outside of SRE, i.e. DevOps, Quality, etc.
nPartners with the Director of SRE to build platform roadmaps, frameworks, and identify team/process improvements.
nTechnical owner of SRE tools with expertise and understanding of current and other widely used industry tools.
nEvaluates other tools/solutions for SRE to ensure IT is being cost aware and tool egnostic.
nQualifications:
nRequired WorkExperience:
n10+ years of experience in developing enterprise software and proficiency in multiple languages e.g., Java and web technologies (Python, Go, Perl, Ruby or shell scripting)
n5+ years in implementing SRE solutions/practices.
n5+ years in mentoring and coaching.
nExpert knowledge of Dynatrace as product owner, user, and
nExpert with a proven track record in delivering technology solutions and leading a high performing SRE team in automating manual work.
nExpert knowledge of reliability and production management domains
nExperience in public cloud environments (AWS/Azure/Google Cloud).
nExperience in leading operations, leveraging key event streaming, messaging and DB services e.g., Casandra, MQ/JMS/Kafka, Aurora, RDS, Cloud SQL, BigTable, DynamoDB, Cloud Spanner, Kinesis, Cloud Pub/Sub, etc.
nExperience in either SAFe agile, Scrum or Kanban model
nExpertise in DevSecOps practices and tools e.g. CI/CD, Gitlab, and any security scanning tools.
nExperience with cloud-based technologies and tools especially in deployment, monitoring and operations
nStrong experience and technical skills in developing/managing APIs and Microservices
nExpert practitioner in multiple technology domains, may be a cross-domain expert able to solve complex and mission critical problems within a business or across the firm
nPreferred Work Experience:
nExperience with containerization (Docker, Kubernetes)
nExperience with Terraform and Ansible
nExperience with SEIM
nExperience with other APM tools
nHealthcare industry experience
nPhysical and Mental Requirements:
nAbility to sit for long periods of time
nKnowledge:
nCompliance requirements e.g. NIST, CFR21, ISO, GDPR, HIPAA, SOX
nHL7 specifications
nIntegration Platform technologies (Mulesoft, Informatica, SnapLogic, Jitterbit, etc.)
nSkills:
nSelf-driven
nProblem solving
nAdaptable
nNegotiation
nPrioritization
nEducation
nBachelor's Degree Bachelor's in computer engineering or something similar or equivalent work experience (Required)
nMaster's Degree Master's in computer engineering (Preferred)
nLanguages
nEnglish (Preferred)
nLicenses and Certifications
nAWS (Preferred)
nAzure (Preferred)
nGPC (Preferred)
nWork Requirements
nTravel Required up to 30%
n56391
nQuest Diagnostics honors our service members and encourages veterans to apply.
nWhile we appreciate and value our staffing partners, we do not accept unsolicited resumes from agencies. Quest will not be responsible for paying agency fees for any individual as to whom an agency has sent an unsolicited resume.
nor any other legally protected status.
nEqual Opportunity Employer: Race/Color/Sex/Sexual Orientation/Gender Identity/Religion/National Origin/Disability/Vets
nPlace of Work
nOn-site
nRequisition ID
n56391
nJob Type
nFull Time
Epic Principal Site Reliability Engineer in secaucus at Unknown Company
This position is listed as part time and able to be worked remotely.