Unknown Company

Software Engineer – Platform Security

san francisco, ca • Posted 3 days ago
Hybrid Full Time General

Forward Deployed Engineer (FDE)FriendliAI is seeking a Forward Deployed Engineer (FDE) to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.You will work directly on our customers' projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring.

This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.Key ResponsibilitiesDesign and implement large-scale deployment architectures for LLM and multimodal inferenceDeploy and manage containerized workloads across Kubernetes clustersDiagnose production issues, such as performance bottlenecks, and implement temporary fixes as neededCollaborate with customers' DevOps teams to integrate FriendliAI's infrastructure into their CI/CD workflowsDevelop scripts, Helm charts, and Terraform modules that simplify repeated deploymentsContribute field insights to shape our platform reliability, observability, and scaling strategiesLead workshops, technical sessions, or webinars to help customers master infrastructure best practicesQualifications3+ years of experience in cloud infrastructure, DevOps, or reliability engineeringBachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalentProficiency with Kubernetes, Docker, Terraform, and HelmStrong foundation in distributed systems, networking, and performance tuningExperience with GPU-based computing and generative AI model serving workloadsStrong technical background in backend systems or AI toolingExperience operating workloads on AWS, GCP, or OCIExcellent problem-solving and debugging skills in real-world environmentsPreferred ExperienceExperience deploying large models (LLMs, diffusion models) on GPUs or clustersFamiliarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)Understanding of networking security and compliance frameworks (e.g., SOC 2)Experience supporting on-prem or hybrid-cloud deploymentsBenefitsA front-row seat to the generative AI infrastructure revolutionCompetitive compensation and benefits packageDaily lunch and dinner provided; unlimited snacks and beveragesHealth check-up and top-tier hardware supportFlexible working hours and a highly collaborative environmentFriendliAI is building the next-generation AI inference platform that accelerates the deployment of large language and multimodal models with unmatched performance and efficiency. Our infrastructure powers high-throughput, low-latency workloads for global organizations and integrates directly with Hugging Face, providing instant access to over 510,000 open-source models.

We are on a mission to deliver the world's best platform for AI inference.

Back to Job Search