AI Infrastructure EngineerArchitect and build custom Artificial Intelligence (AI) infrastructure solutions leveraging the client Kubernetes Platform and AI. You will be responsible for designing high-performance computational stacks that integrate AI, high-speed software-defined storage, and GPU-accelerated nodes. Your mission is to make AI infrastructure "invisible" by optimizing for performance, power consumption, and seamless hybrid-multicloud scalability across on-prem.As an AI Infrastructure Engineer, you will design tailored AI solutions that bridge the gap between private data centers and public cloud.
Your day-to-day will involve optimizing the client computational stack for large language models (LLMs) and generative AI workloads. You will serve as the SME for AI, ensuring that compute, storage (client Objects/Files), and networking (Flow) are perfectly tuned for AI model training and inference.ResponsibilitiesHybrid Multicloud Architecture: Design seamless AI workflows using NC2 on Prem, allowing for rapid bursting of AI workloads from on-prem AHV clusters to the public cloud.Data Services for AI: Architect high-performance storage backends using Objects (S3-compatible) to handle the massive datasets required for AI/ML.Kubernetes & Orchestration: Deploy and manage AI workloads using client Kubernetes Platform to ensure containerized AI models are scalable and resilient.Infrastructure-as-Code: Implement IaC using Calm or Terraform to automate the lifecycle of GPU-enabled nodes.Observability: Design frameworks (monitoring, logging, alerting) for proactive issue detection. Hands on experience on Prometheus, Grafana, ELK, and OpenTelemetry.Ensure high availability, disaster recovery, and fault tolerance across all systems.Networking & Security: Familiarity with Zero-Trust architectures, enterprise networking, storage, and virtualization.Invisible Infrastructure: Modernize legacy 3-tier AI silos into a unified, web-scale environment.Professional & Technical SkillsCore: Deep proficiency in AOS (Acropolis Operating System) and AHV (Native Hypervisor).AI Performance: Experience with GPU Passthrough and vGPU configurations to optimize AI training performance.Security: Applying work Flow for micro segmentation to secure sensitive AI training data.Cost Management: Using Cloud Manager (NCM) Cost Governance to monitor and optimize spend across hybrid environments.
AI Infrastructure Engineer in washington at Unknown Company
This position is listed as full time and hybrid.