Hewlett Packard Enterprise in Spring, Texas, with hybrid work expectations, seeks a Senior Software Engineer to advance the model runtime for AI Essentials. You will work on engine integration, batching, KV cache reuse, and distributed execution, collaborating across teams to optimize latency and GPU utilization on customer-owned hardware.
The role requires deep knowledge of LLM inference, Kubernetes, and Go/Python proficiency, with opportunities to mentor teammates and contribute to code
#J-18808-LjbffrSenior LLM Inference Engineer — Hybrid (Go/Python) in spring at Unknown Company
This position is listed as full time and hybrid.