Site Reliability Engineer I
nSunrise, FL, United States
n(Hybrid)
nJob Description
nSite Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.
nResponsibilities
n- n
- n
Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
n n - n
Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
n n - n
Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
n n - n
Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
n n - n
Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
n n - n
Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
n n - n
Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
n n - n
Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items
n n - n
Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements
n n - n
Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services
n n - n
Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
n n - n
Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation
n n - n
Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic
n n - n
Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
n n - n
Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
n n - n
Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
n n - n
Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
n n - n
Participate and lead the Development change review and change validation processes
n n - n
Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
n n - n
Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers
n n
Qualifications
nEducation Qualifications:
n- n
- n
Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
n n - n
Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
n n - n
Strong knowledge of operating systems and application runtimes such as Java and .NET
n n - n
Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective
n n - n
Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
n n - n
Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
n n - n
Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
n n - n
Knowledge of scripting and automation using languages such as PowerShell and Python
n n - n
Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus
n n
Work Experience:
n- n
- n
Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
n n - n
Experience supporting production systems in large‑scale enterprise environments with a focus on reliability and availability
n n - n
Experience in system administration, infrastructure operations, and network troubleshooting
n n - n
Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
n n - n
Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
n n - n
Knowledge of cloud‑based Site Reliability Engineering (SRE) practices with hands‑on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
n n - n
Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices‑based architectures
n n - n
Experience using enterprise monitoring and alerting platforms such as ELF
n n - n
Exposure to AI‑assisted monitoring, automation, or AIOps tools is a plus• Proficiency in connecting to and administering servers via SSH (Secure Shell)
n n - n
Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access
n n
Licenses & Certifications
n- n
- n
Certification in at least one programming language or runtime such as Java, .NET, or Python
n n - n
Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
n n - n
Public cloud certification in AWS or GCP is a plus
n n - n
Certification or training related to AI platforms, analytics platforms, or AIOps is a plus
n n
Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.
nJob Info
n- n
- n
Job Identification
n n - n
Job CategoryTechnology
n n - n
Posting Date08/31/2026, 07:54 PM
n n - n
Apply Before09/08/2026, 07:00 AM
n n - n
Degree LevelHigh School Graduate
n n - n
Job ScheduleFull time
n n - n
Job ShiftDay
n n - n
Locations1500 NW 136th Avenue, Sunrise, FL, 33323, US(Hybrid)
n n - n
Salary Range$78000 - $ annually + bonus + benefits
n n - n
Career AreaTechnology
n n
Return to Jobs List (
Site Reliability Engineer I in fort lauderdale at Unknown Company
This position is listed as full time and able to be worked remotely.