Unknown Company

Staff Data Center Operations Engineer

denver, co • Posted 2 days ago
Onsite Contract General

Staff Data Center Operations EngineerCrusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org — based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio.This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as SiteOps' technical presence at headquarters — visible to engineering, procurement, and leadership in a way that a field-based role cannot be.You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring-ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site.What You'll DoCross-Site Platform Operations & EscalationOwn Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolutionTravel to sites as needed for complex platform issues, new hardware bring-ups, and deployment supportIdentify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requestsRoot-cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps orgHand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packagesOEM & ODM Technical PartnershipDevelop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM's as growing platforms — at the engineering and field escalation levelServe as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleetBuild familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprintSupport vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadershipPlatform Standards & Org DevelopmentOwn OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sitesOwn the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platformsDesign and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolveContribute to the technician certification program and technical leveling standards across the orgSupport new site bring-up efforts providing platform readiness and deployment execution expertiseHQ Presence & Cross-Functional CollaborationServe as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operationsPartner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structuresProvide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond)Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadershipWhat We're Looking ForRequired7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experienceDirect hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms requiredDeep familiarity with server platform architecture and OEM escalation and RMA processesExperience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoffDemonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolutionExperience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teamsStrong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reportingWillingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)PreferredDirect experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plusFamiliarity with ASUS or Quanta server platforms and ODM engagement modelsExperience with liquid-cooled GPU platforms and CDU integrationFamiliarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X)Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operatorExperience contributing to technician certification programs or IC leveling standards within a DC ops organizationLocationThis role is based at Crusoe Cloud's headquarters in Denver, CO, with regular travel to our data center sites across the US and internationally. Domestic relocation support is available.BenefitsCompetitive compensation and equity packagesRestricted Stock UnitsPaid time off, paid holidays & leave of absence programsComprehensive health, dental & vision insuranceEmployer contributions to HSA accountPaid parental leavePaid life insurance, short-term and long-term disabilityProfessional development & tuition reimbursementMental health & wellness supportCommuter benefits (parking & transit)Cell phone stipend401(k) Retirement plan with company match up to 4% of salaryVolunteer time offGlobal travel insurance & emergency assistanceDaily meals allowanceAdditional perks & programs specific to locationCompensation Range: Compensation will be paid in the range of up to $150,000 -$170,000 + Bonus. Restricted Stock Units are included in all offers.

Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Staff Data Center Operations Engineer in denver at Unknown Company

This position is listed as contract and onsite.

Back to Job Search