Data Direct Networks is seeking an engineer to build and optimize LLM serving and inference systems for production environments. You will work on performance across GPU and CPU pathways and tackle bottlenecks in memory, storage, and throughput.
You will design and scale systems to support RAG and retrieval-heavy AI workloads, contributing to infrastructure where storage architecture and system efficiency directly affect AI performance.
#J-18808-Ljbffr