Unknown Company

Production Systems Engineer – SRE, Monitoring & Root Cause Analysis (AI) | Remote

Remote • Posted 1 weeks ago
Remote Contract Engineering

Position: Software Engineer (Site Reliability Engineer)

Type: Hourly contract

Compensation: $100–$160/hour

Location: Remote

Role Responsibilities

  • Create and review realistic production incident scenarios for AI evaluation
  • Analyze system failures including root cause analysis, monitoring, and remediation
  • Evaluate AI model performance on infrastructure and reliability problem-solving
  • Develop scenarios involving observability, alerting, and capacity planning
  • Assess best practices in incident response and post-mortem processes
  • Contribute to improving AI reasoning in production system environments

Requirements

  • Strong experience in SRE, DevOps, or production engineering
  • Strong experience with on-call operations and managing production incidents
  • Strong experience with root cause analysis and post-mortem processes
  • Strong experience with observability tools such as Prometheus, Grafana, Datadog, or PagerDuty
  • Strong knowledge of Linux systems, networking, and container orchestration
  • Strong experience with infrastructure-as-code and CI/CD pipelines
  • Strong debugging skills across application and system levels
  • Ability to evaluate complex system reliability scenarios and technical reasoning

#J-18808-Ljbffr

Production Systems Engineer – SRE, Monitoring & Root Cause Analysis (AI) | Remote in Remote at Unknown Company

This position is listed as contract and able to be worked remotely.

Back to Job Search