Staff Software Engineer On ML PlatformAs a Staff Software Engineer on ML Platform, you will work across teams to design, build, and operate the infrastructure foundations that power model training, serving, experimentation, and observability for our Machine Learning and AI Systems teams.You will partner with ML engineers, data engineers, and AI Systems engineers to understand production needs, build reliable infrastructure, and deliver tooling that accelerates the team.Key challenges you will tackle:Outerbounds Migration: Own end-to-end migration of CNN's ML training and orchestration workloads to Outerbounds-managed Metaflow, with zero production disruption and a clear path to self-service for ML practitioners.Observability and Cost Attribution: Establish comprehensive observability across all ML products, AI applications, and shared infrastructure — including the monitoring, alerting, and diagnostic tooling engineers need to operate production systems reliably, and cost attribution that lets us understand and govern ML/AI spend by team, product, and use case.Model and Application Registry:Design and operate a unified registry and versioning system for ML models and AI applications, providing lineage, reproducibility, and a clean handoff between development and production.Feature Resolution Framework: Evolve our existing feature resolution framework from a loosely coupled set of pipelines into a real platform — one that lets ML and AI Systems engineers configure content types and events to intercept, register feature-generation APIs, and land resolved features as durable data products in the feature store. The features this framework produces — ML- and LLM-generated alike — power everything from analytics to training to inference to user-facing rendering.What You'll DoDesign and own infrastructure, deployment tooling, and developer experience for ML and AI Systems teamsLead architectural decisions across orchestration, serving, observability, and experimentation infrastructureBuild self-service tooling that lets ML practitioners move from prototype to production without platform team dependenciesEstablish engineering standards for ML/AI infrastructure, including reliability, cost governance, and operational excellenceReview designs and code, mentor engineers, and lead cross-team initiativesPartner with Data Platform on infrastructure coordination and data access patternsCommunicate effectively across audiences — technical documentation, design reviews, and stakeholder interactionsThe Essentials8+ years building production infrastructure or platform systems, with a Bachelor's degree in Computer Science, Information Technology, or a related technical field (or 6+ years with a Master's degree)Deep expertise in distributed systems, with a track record of shipping highly available, low-latency infrastructureStrong proficiency in Python and at least one of Go, Java, or C++Expertise with cloud infrastructure and IaC, especially AWS and TerraformExperience with ML or data infrastructure — orchestration, serving, deployment tooling, observability, or experimentation frameworksProven track record of leading complex platform projects from concept to production — knowing when to own decisions, when to rally the right people for alignment, and when to escalateCollaborative mindset, understanding that great platform work depends on deep partnership with the teams you serveA passion for helping CNN's engineering organization grow through mentorship, talent acquisition, and professional developmentThe Nice to HavesExperience with Metaflow, SageMaker, or comparable ML orchestration platformsExperience with model registries, feature stores, or experimentation frameworksBackground in cost governance, FinOps, or multi-tenant infrastructurePractical experience supporting LLM-based or GenAI production systemsPrior experience working closely with machine learning engineers