Unknown Company

Low-Latency ML Inference Engineer (GPU/FPGA)

new york, ny • Posted 6 days ago
Onsite Full Time IT & Technology

Tower Research Capital is hiring for a Core Engineering role focused on building and optimizing ML inference pipelines that push microsecond latency limits. Lead evaluation of CPUs, GPUs, and FPGAs, optimize memory hierarchies, and collaborate with cross-functional teams to deploy low-latency trading workloads.

Required are 2+ years in latency-sensitive DL inference, deep knowledge of PyTorch/JAX, Python/C++, and GPU kernel development.

#J-18808-Ljbffr
Back to Job Search