TheHigh-Performance ComputingStorage Engineer is primarily responsible for the overall health and maintenanceofstoragetechnologiesin our managed servicescustomer'senvironments. OurStorageEngineers are a valued member of the Managed Services Infrastructure Practice responsible for Tier 3 incident management, service requestmanagementand change management infrastructure support for all Managed Services customers.
Key Responsibilities
- Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities
- Administer parallel and distributed filesystems such asLustre, GPFS,BeeGFS, Ceph, Weka, or Vast
- Optimizestorage performance, throughput, metadata operations, and data locality for AI training and inference
- Build andmaintainautomation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations
- Plan and perform maintenance activities
- Assess customer environments for performance and design issues and propose resolutions
- Work across technical teams to troubleshoot complex infrastructure issues
- Create andmaintaindetailed documentation
- Serve as a subject matter expert and escalation point for storage technologies
- Work with vendors to resolve storage issues
- Communicate with customers and internal team with transparency
- Support data movement workflows including ingest, replication, caching, tiering, and archiving
- Troubleshoot storage, Linux, network, and I/O bottlenecks acrossstorageclusters and fabrics
- Partner with infrastructure, platform, and research teams to support production AI/HPC workloads
- Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency
- Communicate with customers and internal team with transparency
- Participate in on-call rotation
Required Qualifications
- 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering
- Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education
- Strong experience with Linux systems administration
- Hands-on experienceconfiguring, managing, and tuningdistributed or parallel filesystems
- Experience tuning storage for performance-sensitive workloads
- Knowledge of HPC schedulers such asSlurmand/or container platforms such as Kubernetes
- Familiarity with high-speed interconnects such as InfiniBand or RDMA
- Ability to troubleshoot complex issues across storage, compute, and networking layers
- Understandingofdata protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments
- Experience with machine learning or data science workflows in HPC environments
- Managed Services or consulting experience
- Strong background with customer service
- High level problem-solving and communication skills
- Strong oral and written communications skills
- Managed Services or consulting experience
Preferred Qualifications
- Experience supportingstorage solutions forGPU clusters and AI/ML workflows
- Familiarity with object storage such as S3,MinIO, or Ceph Object Gateway
- Experience with Terraform, Ansible, Helm, orGitOpsworkflows
- Knowledge of observability platforms such as Prometheus and Grafana
- Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems
- Experience with machine learning or data science workflows in HPC environments
- Scripting or programming experience with Python and Bash
- Related Storage certifications are a bonus
Engineer, Storage and Data Protection in new york at Unknown Company
This position is listed as full time and onsite.