Owning scheduler behavior across federated LSF cells, the full-time Senior Compute Platform Engineer will diagnose performance issues, set technical designs for cell topology, and collaborate with teams on complex workloads in a remote environment. Key responsibilities Manage scheduling behavior across 15-25 federated LSF cells, including tuning and analysis of scheduling cycles Diagnose MultiCluster forwarding problems and remote queue sizing issues reported by users Design technical specifications for cell topology and federation as the compute farm expands Required qualifications BS or MS in Computer Science, Computer Engineering, or equivalent experience 8+ years in HPC or large-scale batch compute, with 5+ years specifically on IBM Spectrum LSF Proven expertise in LSF internals and debugging beyond standard documentation Hands-on experience with MultiCluster in a production, multi-site environment Strong fundamentals in Linux systems and proficiency in programming/scripting languages such as Python, Perl, and shell
Senior Compute Platform Engineer in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.