Unknown Company

Low-Latency ML Inference Engineer

new york, ny • Posted 5 days ago
Onsite Full Time Engineering

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

#J-18808-Ljbffr
Back to Job Search