Unknown Company

Site Reliability Engineer I

fort lauderdale, fl • Posted 2 days ago
Remote Full Time Architecture and Engineering Occupations

Site Reliability Engineer I

n

Sunrise, FL, United States

n

(Hybrid)

n

Job Description

n

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.

n

Responsibilities

n
    n
  • n

    Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms

    n
  • n
  • n

    Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives

    n
  • n
  • n

    Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics

    n
  • n
  • n

    Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)

    n
  • n
  • n

    Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure

    n
  • n
  • n

    Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures

    n
  • n
  • n

    Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability

    n
  • n
  • n

    Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items

    n
  • n
  • n

    Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements

    n
  • n
  • n

    Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services

    n
  • n
  • n

    Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements

    n
  • n
  • n

    Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation

    n
  • n
  • n

    Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic

    n
  • n
  • n

    Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates

    n
  • n
  • n

    Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies

    n
  • n
  • n

    Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability

    n
  • n
  • n

    Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions

    n
  • n
  • n

    Participate and lead the Development change review and change validation processes

    n
  • n
  • n

    Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues

    n
  • n
  • n

    Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers

    n
  • n
n

Qualifications

n

Education Qualifications:

n
    n
  • n

    Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.

    n
  • n
  • n

    Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus

    n
  • n
  • n

    Strong knowledge of operating systems and application runtimes such as Java and .NET

    n
  • n
  • n

    Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective

    n
  • n
  • n

    Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana

    n
  • n
  • n

    Knowledge of observability practices including logging, monitoring, tracing, and performance analysis

    n
  • n
  • n

    Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra

    n
  • n
  • n

    Knowledge of scripting and automation using languages such as PowerShell and Python

    n
  • n
  • n

    Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus

    n
  • n
n

Work Experience:

n
    n
  • n

    Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools

    n
  • n
  • n

    Experience supporting production systems in large‑scale enterprise environments with a focus on reliability and availability

    n
  • n
  • n

    Experience in system administration, infrastructure operations, and network troubleshooting

    n
  • n
  • n

    Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar

    n
  • n
  • n

    Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms

    n
  • n
  • n

    Knowledge of cloud‑based Site Reliability Engineering (SRE) practices with hands‑on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform

    n
  • n
  • n

    Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices‑based architectures

    n
  • n
  • n

    Experience using enterprise monitoring and alerting platforms such as ELF

    n
  • n
  • n

    Exposure to AI‑assisted monitoring, automation, or AIOps tools is a plus• Proficiency in connecting to and administering servers via SSH (Secure Shell)

    n
  • n
  • n

    Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access

    n
  • n
n

Licenses & Certifications

n
    n
  • n

    Certification in at least one programming language or runtime such as Java, .NET, or Python

    n
  • n
  • n

    Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus

    n
  • n
  • n

    Public cloud certification in AWS or GCP is a plus

    n
  • n
  • n

    Certification or training related to AI platforms, analytics platforms, or AIOps is a plus

    n
  • n
n

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.

n

Job Info

n
    n
  • n

    Job Identification

    n
  • n
  • n

    Job CategoryTechnology

    n
  • n
  • n

    Posting Date08/31/2026, 07:54 PM

    n
  • n
  • n

    Apply Before09/08/2026, 07:00 AM

    n
  • n
  • n

    Degree LevelHigh School Graduate

    n
  • n
  • n

    Job ScheduleFull time

    n
  • n
  • n

    Job ShiftDay

    n
  • n
  • n

    Locations1500 NW 136th Avenue, Sunrise, FL, 33323, US(Hybrid)

    n
  • n
  • n

    Salary Range$78000 - $ annually + bonus + benefits

    n
  • n
  • n

    Career AreaTechnology

    n
  • n
n

Return to Jobs List (

Site Reliability Engineer I in fort lauderdale at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search