d-Matrix inc. is hiring end-to-end inference engineers to innovate LLM solutions in Santa Clara, CA. You will work with novel hardware and software to optimize performance and deploy advanced inference systems.
Ideal candidates possess strong Python and C/C++ skills, along with a deep understanding of generative AI architectures and extensive experience in performance optimization of LLM frameworks. Join a small, senior team committed to cutting-edge AI development.
#J-18808-Ljbffr