Description
Application Operational Services is seeking Site Reliability Engineer. This role requires a strong IT professional focused on establishing and improving monitoring to measure end-to-end performance and end-user availability of systems via a suite of common monitoring tools. Interface with business partners and operations teams to develop business and technical monitoring requirements. As part of this role, the person will primarily be responsible for supporting production or operations of critical applications. They will ensure the application's operational readiness by evaluating its performance, reliability, scale, resiliency & observability. They will be responsible for identifying issues in production, triaging identified issues, partnering with other engineers on the team to identify the root cause. Possess strong analytical ability in solving IT problems, working towards automation, and elimination of systems and or process bottlenecks.
Skills
dynatrace, observability, monitoring tools, devops tools
Top Skills Details
dynatrace,observability,monitoring tools,devops tools
Additional Skills & Qualifications
* As part of the SRE team, perform full stack triaging of alerts and engage other engineers to identify root cause of application performance & stability issues.
* Work with stakeholders such as product owners to define service level objectives (SLOs) for application features and services.
* Track performance against SLOs in partnership with development teams or other stakeholders, and ensure systems continue to meet SLOs over time.
* Design, develop dashboards and reports to communicate key metrics.
* Identify opportunities to improve alerting posture and create/update alerts accordingly.
* Work closely with the Engineering team to understand application architecture and perform Single point of failure analysis and create scenarios for testing resiliency of the application.
* Create/derive NFR/Workload model and ensure performance & resiliency is considered early in the SDLC.
* Execute performance/chaos tests, analyze using APM and other tools to identify performance & stability issues.
* Document any findings/analysis/results, communicate and present to stakeholders.
* Perform analytics on previous incidents to understand root causes and use automation to reduce the probability and/or impact of problem recurrence.
* Demonstrate proficiency with DevOps tools, JIRA, ServiceNow, MS Project and perform tasks using the tools.
* Be a technical expert with expertise across multiple technology areas and the ability to diagnose complex issues throughout many technologies and apply this knowledge to effective monitoring of applications.
Experience Level
Expert Level
Job Type & Location
This is a Contract position based out of Plano, TX.
Pay and Benefits
The pay range for this position is $60.00 - $75.00/hr.
Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.
* Medical, dental & vision
* Critical Illness, Accident, and Hospital
* 401(k) Retirement Plan - Pre-tax and Roth post-tax contributions available
* Life Insurance (Voluntary Life & AD&D for the employee and dependents)
* Short and long-term disability
* Health Spending Account (HSA)
* Transportation benefits
* Employee Assistance Program
* Time Off/Leave (PTO, Vacation or Sick Leave)
Workplace Type
This is a hybrid position in Plano,TX.
Application Deadline
This position is anticipated to close on Sep 3, 2026.
The company is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
The company is an equal opportunity employer and will consider all applications without regard to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
San Francisco Fair Chance Ordinance: Pursuant to the San Francisco Fair Chance Ordinance, for all positions located in the city and county of San Francisco, we will consider for employment qualified applicants with arrest and conviction records.
Massachusetts Lie Detector: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools. #J-18808-Ljbffr
Application Operational Services is seeking Site Reliability Engineer. This role requires a strong IT professional focused on establishing and improving monitoring to measure end-to-end performance and end-user availability of systems via a suite of common monitoring tools. Interface with business partners and operations teams to develop business and technical monitoring requirements. As part of this role, the person will primarily be responsible for supporting production or operations of critical applications. They will ensure the application's operational readiness by evaluating its performance, reliability, scale, resiliency & observability. They will be responsible for identifying issues in production, triaging identified issues, partnering with other engineers on the team to identify the root cause. Possess strong analytical ability in solving IT problems, working towards automation, and elimination of systems and or process bottlenecks.
Skills
dynatrace, observability, monitoring tools, devops tools
Top Skills Details
dynatrace,observability,monitoring tools,devops tools
Additional Skills & Qualifications
* As part of the SRE team, perform full stack triaging of alerts and engage other engineers to identify root cause of application performance & stability issues.
* Work with stakeholders such as product owners to define service level objectives (SLOs) for application features and services.
* Track performance against SLOs in partnership with development teams or other stakeholders, and ensure systems continue to meet SLOs over time.
* Design, develop dashboards and reports to communicate key metrics.
* Identify opportunities to improve alerting posture and create/update alerts accordingly.
* Work closely with the Engineering team to understand application architecture and perform Single point of failure analysis and create scenarios for testing resiliency of the application.
* Create/derive NFR/Workload model and ensure performance & resiliency is considered early in the SDLC.
* Execute performance/chaos tests, analyze using APM and other tools to identify performance & stability issues.
* Document any findings/analysis/results, communicate and present to stakeholders.
* Perform analytics on previous incidents to understand root causes and use automation to reduce the probability and/or impact of problem recurrence.
* Demonstrate proficiency with DevOps tools, JIRA, ServiceNow, MS Project and perform tasks using the tools.
* Be a technical expert with expertise across multiple technology areas and the ability to diagnose complex issues throughout many technologies and apply this knowledge to effective monitoring of applications.
Experience Level
Expert Level
Job Type & Location
This is a Contract position based out of Plano, TX.
Pay and Benefits
The pay range for this position is $60.00 - $75.00/hr.
Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.
* Medical, dental & vision
* Critical Illness, Accident, and Hospital
* 401(k) Retirement Plan - Pre-tax and Roth post-tax contributions available
* Life Insurance (Voluntary Life & AD&D for the employee and dependents)
* Short and long-term disability
* Health Spending Account (HSA)
* Transportation benefits
* Employee Assistance Program
* Time Off/Leave (PTO, Vacation or Sick Leave)
Workplace Type
This is a hybrid position in Plano,TX.
Application Deadline
This position is anticipated to close on Sep 3, 2026.
The company is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
The company is an equal opportunity employer and will consider all applications without regard to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
San Francisco Fair Chance Ordinance: Pursuant to the San Francisco Fair Chance Ordinance, for all positions located in the city and county of San Francisco, we will consider for employment qualified applicants with arrest and conviction records.
Massachusetts Lie Detector: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools. #J-18808-Ljbffr
Dynatrace SRE in plano at Unknown Company
This position is listed as contract and hybrid.