Hewlett Packard Enterprise seeks a Senior Software Engineer to build and evolve the model runtime for the AI Essentials inference platform used by enterprises to operate large language models on customer-owned hardware, including air-gapped environments. Emphasis is on sustained execution efficiency, low tail latency, and high GPU utilization.
The role involves engine integration, batching, KV cache management, and distributed execution, with collaboration across inference teams and a Kubernetes
#J-18808-LjbffrSenior Software Engineer - LLM Inference Runtime (Hybrid) in spring at Unknown Company
This position is listed as full time and onsite.