Intel is seeking an experienced reliability engineer to define and own pod-level reliability specs across compute, memory, storage, network, power and cooling for AI hardware data centers.
You will lead FMEA, telemetry, and KPI tracking; collaborate with facilities on N+1 redundancy and disaster recovery, and apply Weibull/FIT reliability methods. AI cluster operations and data analytics are a plus.
#J-18808-LjbffrAI Hardware Reliability Engineer: Pod-Level RAS Telemetry in santa clara at Unknown Company
This position is listed as full time and onsite.