Own the architecture and design of reliable, scalable, cost-effective, and performant AI inference and training infrastructure, partnering with development teams and technical leads. Lead production incident response and blameless postmortems, promote AI-first development, and collaborate with SRE leaders across Google to share scalable solutions.
Requirements
Requires a bachelor's degree in computer science, a related field, or equivalent practical experience; at least eight years of software development experience; four years leading projects and providing technical leadership; and three years designing, analyzing, and troubleshooting distributed systems. Experience with machine learning or AI in a software development environment is required, while a master's degree in computer science or engineering is preferred.
Key Skills
Software Development, Technical Leadership, Distributed Systems, Machine Learning, Artificial Intelligence, Site Reliability Engineering, AI Infrastructure, Inference Infrastructure, Training Infrastructure, System Architecture, Large-Scale System Design, Incident Response, Blameless Postmortems, Automation, Scalability, Performance Optimization
Benefits
Bonus, Equity, Benefits
#J-18808-LjbffrSenior Staff Software Engineer, Site Reliability Engineering, Workspace AI in sunnyvale at Unknown Company
This position is listed as full time and onsite.