Unknown Company

Low-Latency ML Inference Engineer

new york, new york • Posted 4 days ago
Onsite Full Time General

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

#J-18808-Ljbffr

Low-Latency ML Inference Engineer in new york at Unknown Company

This position is listed as full time and onsite.

Back to Job Search