Onsite
Full Time Software Architecture & Engineering
We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team
This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems
The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters
The ideal new team member will be a hands-on problem solver who is comfortable working independently
The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet
Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems
Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X
Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware
In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment
Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance
Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success
Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems
Benefits
Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
Time away: Paid time off for vacations, family bonding, and unexpected needs
401(k) match: Build your financial future with our 401(k) matching program
Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges
Ability to lean in and assist team members working on critical or complex technical initiatives
Ability to work independently and within a team
Strong analytical and problem-solving skills
Ability to set the technical direction for a specific project and execute
The ability to identify a problem, rapidly develop a scalable solution and ship it
4-6 years of software engineering experience
Strength in at least one programming language - Go, Python, Java, Rust
Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
Excellent communication and collaboration skills
Experience with Temporal and Kubernetes
Experience working directly with hardware vendors
Background in large-scale GPU fleet operations or hyperscale data center environments