Unknown Company

Principal Machine Learning Engineer

wa • Posted 4 days ago
Remote Full Time Electrical & Energy Engineering

Job Title

Principal Machine Learning Engineer

Job Description

Protingent Staffing is offering a direct‑hire Principal Machine Learning Engineer role that is fully remote. This role is a deep technical authority responsible for designing and evolving the most critical ML systems across training, inference, evaluation, and infrastructure.

Responsibilities

  • Architect and build large‑scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Design reproducible, high‑performance training pipelines across GPU infrastructure.
  • Architect inference systems that balance latency, throughput, cost, and reliability at scale.
  • Design and maintain data systems for high‑quality synthetic and real‑world training data.
  • Implement evaluation pipelines for performance, robustness, safety, and bias in partnership with research leadership.
  • Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.
  • Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products.
  • Make pragmatic trade‑offs and ship improvements quickly, learning from real usage.
  • Work under real production constraints: latency, cost, reliability, and safety.
  • Ensure ML systems (training, inference, evaluation) are reliable, scalable, and meet defined performance targets.
  • Deploy models that achieve measurable quality improvements and meet user‑impact goals.
  • Proactively monitor, debug, and resolve production issues with clear root‑cause analysis.
  • Provide clear guidance, best practices, and scalable ML solutions for teammates and cross‑functional collaborators.
  • Enable efficient, safe research‑to‑production cycles that continuously improve the product experience.

Qualifications

  • Strong background in deep learning and transformer‑based architectures.
  • Hands‑on experience training, fine‑tuning, or deploying large‑scale ML models in production.
  • Proficiency with at least one modern ML framework (e.g., PyTorch, JAX) and ability to learn others quickly.
  • Experience with distributed training and inference frameworks (DeepSpeed, FSDP, Megatron, ZeRO, Ray).
  • Strong software engineering fundamentals – write robust, maintainable, production‑grade systems.
  • Experience with GPU optimization, including memory efficiency, quantization, and mixed precision.
  • Comfort owning ambiguous, zero‑to‑one ML systems end‑to‑end.
  • Bias toward shipping, learning fast, and improving systems through iteration.
  • Must have experience with LLM inference frameworks such as vLLM, TensorRT‑LLM, or FasterTransformer.
  • Contributions to open‑source ML or systems libraries.
  • Background in scientific computing, compilers, or GPU kernels.
  • Experience with RLHF pipelines (PPO, DPO, ORPO).
  • Experience training or deploying multimodal or diffusion models.
  • Experience with large‑scale data processing (Apache Arrow, Spark, Ray).

Job Details

  • Job Type: Direct Hire
  • Pay Range: Market Rate
  • Location: Fully Remote

#J-18808-Ljbffr
Back to Job Search