Operations Tech Lead/EngineerLocation: Whitehouse Station, NJ or Jersey City, NJ - Hybrid Duration: 6 months CTHThe Advanced Engineering department has a need for a Platforms Operations Tech Lead who will work closely with our strategic partners and will provide support for all the department's platforms including Client.IO and some legacy platforms like App Works, IBM BPM, and others. As these platforms are being consolidated across IT areas, dedicated focus is needed to ensure the highest levels of service to all application areas that utilize them.ResponsibilitiesAct as Operations Lead with our strategic partners that will be providing support for our platformsEngage with Engineering and Architecture with regards to platform usage to determine appropriate support-level activities per platformEngage with Delivery teams on specific monitoring and alerts needed for pre-production and post-production support and work closely with the APM team to ensure proper monitoring is in place and current for each new initiativeEstablish dashboards for platform operational health and specific dashboards for critical Tier 1 application/platform usersProvide ongoing monitoring of all platforms and establish alerts to the L1/L2 support team where neededExperience working with infrastructure on platform configuration, optimization, and troubleshootingMonitor system performance, and research solutions to any potential bottlenecksProvide expert technical insight for Sev 1/2 infrastructure (unplanned outages) and application (job performance or stability) issuesTroubleshoot and support Sev 3 application (coding, design) issues that gets escalatedAssist the application teams with any design, code refactoring, and performance analysis requestsPartnering with agile development and application teams to influence application modernization and migration to cloud platformsCollaborate with architects, engineering, client managers, project managers, applications and infrastructure teams to plan and coordinate changesWhen incident contact procedures do not work, act as primary escalation for problems and incidentsAccountable for root cause analysis for Severity 1, Severity 2 and chronic recurring incidentsCreate and follow processes to implement planned changes to production and non-production systemsPerform incident response and break-fix triage for serviceability issuesPartner with the business to fulfill requests for serviceExecute and improve vulnerability and resiliency management programs including patch management, infrastructure testing and other proactive maintenance tasksIdentify and document opportunities and solutions that enhance service delivery efficiency and improve the customer experienceBe available when scheduled (rotational with others on team) to provide on-call support after hours and on weekendsProvide thought leadership to promote the continuous improvement of service delivery and operational practices such as:Propose, design, and implement enhancements to current environmentDevelop actionable plans to improve procedures and systems to mitigate risk, ensure compliance with established industry rules, regulations and best practicesAssist in developing engineering and operational service metricsCreate and maintain system run-books; documenting day-to-day support, maintenance, and troubleshooting knowledgebase of the infrastructureConduct peer review analysis and acceptance for new or modified processes to ensure sustainability and repeatabilityAnalyze performance metrics to ensure timely and accurate delivery of servicesPerform implementation, maintenance, troubleshooting, and remediation services for supported platformsCollaborate with application owners to install, configure and deliver third party product installationsNegotiate dependencies and priorities with stakeholders and internal customersProvide subject matter expertise in cross-functional strategic and tactical effortsPartner with key vendors to escalate and remediate business-impacting issuesProvide leadership and mentoring to emerging talent on the teamAnalyze and execute configuration and audit strategies to ensure compliance with policies, standards, and security directionRequirementsMinimum 7 to 10 years' experience in technology engineering or operations with expertise in working with technology platformsUnderstanding of security and risk as it relates to patch and configuration management, logging, alerting and monitoringExtensive troubleshooting, triage, root cause analysis and performance monitoring skills.Experience with ITIL and/or other similar IT best practice frameworksExperience with LEAN methodologies to drive problem resolution and service improvementsExperience in troubleshooting technology problems and active participation in Sev1/Sev2 callsExperience in usage of the following tools/technologies:Dynatrace, AppDynamics, OpenSearch, Azure PortalExperience working with vendors as key strategic partners seeking to maximize their experience and knowledge based on our ongoing and growing needs.Experience with Agile and/or Kanban methodologiesBackground working in organizations that provide 24x7x365 supportDemonstrated ability to achieve successful outcomes in difficult situations and work with application teams, business customers, and various levels of managementMust be able to communicate effectively with technical and non-technical audiences.Must be a self-starter with the ability to work independently and in a collaborative team environment.Java programming experience is a huge plus as Client.IO will be largely Java-based.Ability to learn and apply new technologies to solve critical business problems.Experience in technology operations with regards to monitoring and support.Experience working with infrastructure on platform configuration, optimization, and troubleshooting.Insurance Domain experience is a plusKnowledge of Client's application portfolio, legacy and modern systems and platforms is a plus.
Unknown Company