Senior MLOps & AI Infrastructure EngineerAt Altera™, our independence as the world's largest pure-play FPGA solutions provider gives us the focus, speed, and agility to innovate without compromise. With more than four decades of industry-leading FPGA expertise, our singular mission is to deliver the programmable technologies that help customers differentiate, innovate, and scale across rapidly evolving markets like AI, cloud, networking, and edge. As an independent company, we move faster, invest deeper, and partner more closely—empowering our teams to drive breakthrough innovation and shape the future of the FPGA industry.About the RoleWe are looking for a Senior MLOps & AI Infrastructure Engineer to architect, build, and operationalize machine learning systems at scale.
This role sits at the intersection of data science, software engineering, and infrastructure—combining deep ML expertise with the DevOps/MLOps discipline required to ship models reliably into production.You will partner closely with software, data, and infrastructure teams to design end-to-end ML pipelines, automate model lifecycle management, and deliver AI-powered capabilities across our EDA, HPC, and cloud environments.Key Responsibilities:ML Platform & Pipeline EngineeringDesign, build, and maintain scalable ML pipelines for training, evaluation, and deployment across cloud and on-prem HPC environmentsBuild MLOps infrastructure including experiment tracking, model registry, feature stores, and automated retraining workflowsImplement CI/CD/CT (Continuous Training) pipelines for ML models using tools such as Kubeflow, MLflow, Airflow, or similarContainerize ML workloads with Docker and orchestrate at scale using Kubernetes and GPU node poolsModel Development & OptimizationDevelop, fine-tune, and deploy large-scale models including LLMs, GNNs, and reinforcement learning agents for EDA and chip design applicationsApply advanced techniques: transfer learning, quantization, pruning, distillation, and RLHF for production-grade model efficiencyImplement A/B testing frameworks and shadow deployments for safe model rolloutBenchmark and optimize model inference performance on GPU/TPU clustersData Engineering & Feature ManagementBuild and maintain data pipelines for large-scale structured and unstructured datasets (terabyte-scale)Collaborate with data teams to design feature engineering systems and maintain data quality for ML trainingImplement data versioning and lineage tracking (DVC, Delta Lake, or similar)Infrastructure & OperationsManage cloud ML infrastructure on AWS (SageMaker), Azure (AML), or GCP (Vertex AI) with cost and performance optimizationAutomate infrastructure provisioning using Terraform or CloudFormation for GPU-backed ML environmentsBuild monitoring, alerting, and observability systems for model performance drift, data quality, and system healthSupport HPC schedulers (LSF, Slurm) for large-scale distributed training jobsCollaboration & LeadershipPartner with research scientists to productionize experimental models with engineering rigorMentor junior engineers and define ML engineering best practices across the organizationDrive adoption of AI/ML solutions within semiconductor, EDA, and simulation workflowsTechnology StackML Frameworks:PyTorch • TensorFlow • JAX • Hugging Face • scikit-learn • XGBoostMLOps & Pipelines:MLflow • Kubeflow • Airflow • Weights & Biases • DVC • FeastInfrastructure & Cloud:AWS SageMaker / GCP Vertex AI / Azure ML • Terraform • Docker • Kubernetes • Slurm / LSFLanguages:Python • Bash • Go • SQLMonitoring & Observability:Prometheus • Grafana • ELK Stack • Evidently AI • ArizeKey CompetenciesStrong ownership mindset — you drive ML initiatives from prototype to production without being askedBias toward automation: if you do it twice, you automate itAbility to bridge research and engineering — translating papers into production-grade systemsThrives in fast-paced, ambiguous environments typical of deep-tech and semiconductor companiesClear communicator who can explain complex ML concepts to non-technical stakeholdersSalary Range$149,100 - $215,925 USDWe use artificial intelligence to screen, assess, or select applicants for the position. Applicants must be eligible for any required U.S. export authorizations.Qualifications:Required QualificationsBachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or related field and 10+ years of industry experience10+ years of experience across ML engineering, data science, and MLOps — including frameworks (PyTorch, TensorFlow, JAX, Hugging Face) and production model deployment at scale8+ years of experience experience with parallelism strategies (FSDP, DeepSpeed, data/model parallelism)10+ years of experience and proficiency in Python programming8+ years of experience in cloud ML platforms (AWS, GCP, Azure), Docker/Kubernetes, and CI/CD pipelines5+ years of hands-on experience with MLflow, W&B, or Neptune for tracking and reproducibilityPreferred QualificationsPhD in Computer Science, Machine Learning, Statistics, or related fieldExperience applying ML/AI to semiconductor, EDA, or chip design domains (e.g., timing prediction, place & route optimization, DRC closure)Familiarity with HPC schedulers such as LSF or Slurm and GPU cluster management for training workloadsKnowledge of LLM fine-tuning, Retrieval-Augmented Generation (RAG) architectures, and AI agent frameworks such as LangChain or AutoGenExperience with graph neural networks (GNNs) or geometric deep learning for circuit and netlist analysisBackground in reinforcement learning for optimization problemsExposure to zero-trust security, DevSecOps, and compliance automation for ML systemsExperience working with large-scale simulation pipelines and synthetic data generationExperience at organizations such as NVIDIA, AMD, Intel, Google DeepMind, or similar AI/HPC-focused companiesPublished research or open-source contributions in ML, MLOps, or AI for EDAExperience building AI-powered developer tools or copilot-style productsFamiliarity with Synopsys, Cadence, or Siemens EDA toolchains and associated data formatsJob Type:RegularShift:Shift 1 (United States of America)Primary Location:San Jose, California, United StatesAdditional Locations: