Unknown Company

Remote AI Model Inference Engineer - GPU & LoRA

Remote • Posted 4 days ago
Onsite Full Time Engineering

Tether is seeking an experienced AI Model Engineer with deep expertise in kernel development, model optimization, fine‑tuning, and GPU acceleration. The engineer will extend the inference framework to support inference and fine‑tuning for language models, with a focus on mobile and integrated GPU acceleration (Vulkan).

Hands‑on experience with quantization, LoRA architectures, Vulkan backend, and mobile GPU debugging is required.

#J-18808-Ljbffr
Back to Job Search