Unknown Company

Staff GPU Inference Engineer — Real-Time AI Systems

Remote • Posted 5 days ago
Onsite Full Time IT & Technology

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, and rack-scale AMD GPU infrastructure.

This hands-on role requires deep debugging and optimization across application, runtime, distributed systems, and hardware layers to improve time to first token, throughput, tail latency, and capacity efficiency.

#J-18808-Ljbffr
Back to Job Search