Working remotely in a full-time capacity, the HPC Support Engineer will serve as a senior technical escalation point, troubleshooting complex infrastructure and platform issues while collaborating with engineering teams to implement permanent fixes. Key responsibilities Troubleshoot hardware, driver, and kernel-level issues, ensuring accurate resolution of customer workload misconfigurations Proactively identify and address gaps in processes, tooling, and documentation, while creating scripts and automations to enhance operations Perform root-cause analysis across distributed systems and GPU infrastructure, contributing to the documentation of solutions and support procedures Required qualifications 3+ years of hands-on HPC experience in an administration, support, or engineering role Strong understanding of Linux system administration, with experience in HPC environments and cluster orchestration tools like Kubernetes or Slurm Proficiency in coding and CI/CD practices, with experience using AI-assisted tools Experience with monitoring/logging tools such as Prometheus, Grafana, or Datadog Knowledge of distributed AI/ML or HPC workloads and high throughput networking technologies (IB/RoCE)