Unknown Company

HPC Technical Consultant, Onsite (LANL) Los Alamos, NM

los alamos, nm • Posted 2 weeks ago
Remote Full Time General

HPC Technical Consultant, Onsite (LANL) Los Alamos, NMThis role has been designed as "Onsite" with an expectation that you will primarily work from an HPE partner/customer office.Join a dedicated on-site team supporting operations and hardware maintenance for HPE supercomputers in one of the nation's premier High-Performance Computing facilities. US Citizenship required Onsite daily work required in Los Alamos, NM. This is not a remote position Days/Hours: M-F, 8am to 5pm or 7am to 4pm Key ResponsibilitiesMonitor and maintain system health across large-scale HPC compute, network, and storage infrastructureTroubleshoot and repair hardware issues on HPC servers and supporting systemsPerform basic Linux system administration tasks as neededCreate, monitor, update, and close support ticketsPerform hardware component replacements using sparesOperate hand tools and low-power tools for server maintenanceTrack and document hardware repairs, part replacements, and returnsCreate, update, and maintain site documentation, processes, and workflowsAssist with new system installation and expansion activitiesRead system documentation and diagrams to locate componentsCollaborate with team members using email, Teams, Slack, and in-person communicationParticipate in on-call schedule to support 24x7 operationsMaintain tools and workspace in an organized mannerMinimum Qualifications Candidates must meet all of the following requirements:Ability to obtain a Q Clearance (required)US Citizenship (required)Must be able to work onsite 5 days per week in Los Alamos, NM, with additional onsite work for on-call support.

This is not a remote positionStrong mechanical aptitude and comfort using common hand tools (screwdrivers, pliers, wrenches, cable tools, etc.) for assembling, disassembling, and maintaining server hardware and related equipmentAbility to lift up to 50 lbs individually and up to 75 lbs with assistanceSolid understanding of computer hardware components (servers, drives, memory modules, power supplies, cabling, and peripherals)Proficiency with basic computer operations on Windows and macOS (MacBook), including OS navigation, file management, and standard productivity tools such as Slack, SharePoint, Microsoft Office (Word, Excel, Outlook, and Teams)Preferred Qualifications A combination of the following is preferred:Associate's degree, some college, or technical training (BS preferred)2+ years of Linux System Administration Experience, including strong command-line navigation, log analysis and monitoring (journalctl, syslog, log files), troubleshooting system and application issues, and scripting/automation using Bash or Python.Experience using Redfish (along with IPMI) for out-of-band server hardware management and monitoring. This includes utilizing the Redfish RESTful API for querying system health, power/thermal monitoring, firmware inventory, component status (processors, memory, drives, NICs), event logs, and performing actions such as system resets, power control, and BIOS configuration.2+ years of hands-on experience troubleshooting and maintaining server hardware in a datacenter environment, including diagnosing hardware faults (power, thermal, storage, networking), performing component replacements (drives, memory, CPUs, PSUs, HBAs, NICs), rack mounting/decommissioning servers, and managing cable infrastructure1+ year of experience with high-speed networking concepts and troubleshooting for Ethernet, HPE Slingshot, and InfiniBand fabrics, including link diagnostics, performance tuning, cable/fiber management, switch configuration, and fault isolation in large-scale HPC environments.Previous experience in a 24x7 production support environmentStrong troubleshooting and problem-solving skills with the ability to work independently, including systematically diagnosing complex hardware, software, and network issues through log analysis, debugging tools, and root cause analysis while minimizing downtime in high-availability environmentsExperience reading technical diagrams, schematics, and working with ticketing systemsExperience with Git for version control of code, scripts, configuration files, and documentation (including cloning, branching, committing, merging, and resolving conflicts)Experience with High-Performance Computing (HPC) systems, clusters, or large-scale AI infrastructureExperience with large-scale storage systems, including installation, configuration, monitoring, and troubleshooting of parallel file systems, enterprise SAN/NAS solutions, object storage, and high-capacity disk arraysHighly Desired Industry Certifications (any of the following):CompTIA Linux+CompTIA Security+CompTIA Server+CompTIA A+CompTIA Network+ITIL FoundationWhat We Can Offer You:Health & WellbeingWe strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.Personal & Professional DevelopmentWe also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.Unconditional InclusionWe are unconditionally inclusive in the way we work and celebrate individual uniqueness.

We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

HPC Technical Consultant, Onsite (LANL) Los Alamos, NM in los alamos at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search