Description
nWe are seeking a Lab Operations Engineer to support the bring-up, validation, diagnostics, and operational readiness of next-generation high-performance computing (HPC) and AI infrastructure. This role is highly hardware-focused and will play a critical part in evaluating system reliability, reproducing field issues, validating hardware platforms, and enabling internal engineering teams to successfully deploy workloads on on-prem systems.
nThe ideal candidate combines hands-on lab experience with strong Linux systems administration skills and a passion for hardware diagnostics, validation, and system performance.
nResponsibilities:
nHardware Diagnostics & Failure Analysis
nInvestigate hardware failures and operational issues observed in production environments.
nReproduce large-scale system issues within a controlled lab environment.
nPerform root cause analysis on server, memory, storage, networking, and thermal-related failures.
nExecute component-level diagnostics, stress testing, and hardware validation procedures.
nAnalyze hardware fallout trends and collaborate with engineering teams to identify corrective actions.
nSystem Bring-Up & Platform Validation
nPerform system bring-up activities for new server platforms and AI/HPC infrastructure.
nConfigure and validate hardware from OEM partners including Dell, Supermicro, and other server vendors.
nInstall, configure, and validate firmware, BIOS, BMC, and hardware management components.
nExecute validation plans and test scripts for next-generation hardware platforms.
nSupport qualification efforts for new system architectures and deployment models.
nTest Automation & Workflow Enablement
nBuild and maintain recurring test pipelines for hardware validation and operational readiness.
nExecute engineering design automation (EDA) and HPC workloads to verify system stability and performance.
nDevelop repeatable processes for system qualification and regression testing.
nPartner with engineering teams to ensure hardware platforms meet operational requirements before deployment.
nSystems Administration & Infrastructure Support
nAdminister Linux-based systems used for validation and testing.
nConfigure system-level settings including kernel parameters, drivers, and hypervisor configurations.
nSupport HPC cluster environments and distributed computing infrastructure.
nTroubleshoot operating system and platform-level issues impacting performance or reliability.
nCross-Functional Collaboration
nWork closely with hardware engineering, datacenter operations, platform architecture, and vendor partners.
nCollaborate with OEMs and suppliers to resolve hardware issues and improve platform quality.
nEnable internal users and engineering teams to successfully deploy workloads on validated systems.
nSupport operational readiness efforts for emerging hardware technologies.
nSkills
nLinux, Hardware Validation, System Validation, BIOS, Python, Bash, Automation
nTop Skills Details
nLinux,Hardware Validation,System Validation,BIOS,Python,Bash
nAdditional Skills & Qualifications
nExperience with HPC, AI infrastructure, or large-scale compute clusters.
nvalidation engineering, system testing, or platform qualification.
nExposure to GPU-based systems and accelerated computing environments.
nExperience with Dell, Supermicro, or other enterprise server platforms.
nKnowledge of Linux kernel tuning, performance optimization, and hypervisor technologies.
nExperience developing automated test frameworks and validation pipelines.
nUnderstanding of datacenter infrastructure, rack-scale deployments, and operational workflows.
nFamiliarity with EDA, AI/ML, or other high-performance workloads.
nJob Type & Location
nThis is a Contract position based out of Santa Clara, CA.
nPay and Benefits
nThe pay range for this position is $70.00 - $75.00/hr.
nIndividual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.
nEligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to specific elections, plan, or program terms. If eligible, the benefits available for this temporary role may include the following: - Medical, dental & vision - Critical Illness, Accident, and Hospital - 401(k) Retirement Plan - Pre-tax and Roth post-tax contributions available - Life Insurance (Voluntary Life & AD&D for the employee and dependents) - Short and long-term disability - Health Spending Account (HSA) - Transportation benefits - Employee Assistance Program - Time Off/Leave (PTO, Vacation or Sick Leave)
nWorkplace Type
nThis is a fully onsite position in Santa Clara,CA.
nApplication Deadline
nThis position is anticipated to close on Sep 14, 2026.
nAbout TEKsystems
nWe're partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a team of 80,000 strong, working with over 6,000 clients, including 80% of the Fortune 500, across North America, Europe and Asia. As an industry leader in Full-Stack Technology Services, Talent Services, and real-world application, we work with progressive leaders to drive change. That's the power of true partnership. TEKsystems is an Allegis Group company.
nThe company is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
nAbout TEKsystems and TEKsystems Global Services
nWe're a leading provider of business and technology services. We accelerate business transformation for our customers. Our expertise in strategy, design, execution and operations unlocks business value through a range of solutions. We're a team of 80,000 strong, working with over 6,000 customers, including 80% of the Fortune 500 across North America, Europe and Asia, who partner with us for our scale, full-stack capabilities and speed. We're strategic thinkers, hands-on collaborators, helping customers capitalize on change and master the momentum of technology. We're building tomorrow by delivering business outcomes and making positive impacts in our global communities. TEKsystems and TEKsystems Global Services are Allegis Group companies. Learn more at TEKsystems.com.
nThe company is an equal opportunity employer and will consider all applications without regard to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
nSan Francisco Fair Chance Ordinance: Pursuant to the San Francisco Fair Chance Ordinance, for all positions located in the city and county of San Francisco, we will consider for employment qualified applicants with arrest and conviction records.
nMassachusetts Lie Detector: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
nUse of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools.
Lab Operations Engineer in santa clara at Unknown Company
This position is listed as contract and onsite.