Architect, build, and continuously evolve a truly declarative, self-healing, cloud-native, Kubernetes-native, and highly observable enterprise-grade platform. This strategic leadership role defines and drives the fundamental architecture of our technology stack, including: Go-based microservices architecture, ensuring high throughput and resilience; next-generation CI/CD pipelines utilizing GitHub Actions and Azure DevOps, engineered for speed, reliability, governance, and release integrity; infrastructure-as-Code (IaC) automation using Terraform and Ansible to deliver consistent, secure, and repeatable environments at scale; world-class Observability strategy, integrating metrics, logging, and tracing, and intelligent alerting to enable proactive operations and optimization and governance of MongoDB data persistence and distributed systems, ensuring performance, durability, and operational efficiency.
1. CI/CD Strategy & Governance (The Quality Pipeline)
Responsibility & Authority (R&A) for the end-to-end architecture, standardization, and all YAML-based delivery pipelines (build, test, security scanning, artifact management). Establish and enforce mandatory, non-negotiable quality gates (unit, integration, security, performance) at every stage of the pipeline to ensure "shift-left" quality. Secure and harden the CI/CD ecosystem against tampering and supply chain risks through strong access controls, artifact integrity, and policy enforcement.
2. Deployment Excellence & Safety (Production Integrity)
Define, standardize, and govern safe, high-integrity deployment strategies (e.g., Canary Releases, Rolling Updates, Blue/Green, utilization of Feature Flags). Implement robust, fast, and fully automated rollback mechanisms, ensuring absolute production safety and minimizing Mean Time To Recovery (MTTR). Continuously reduce deployment risk and complexity through automation, standardization, and release discipline.
3. Production Reliability Engineering (SRE & Resilience)
Establish, evangelize, and enforce enterprise-grade standards for Stability, Performance, Scalability, Observability, and Disaster Recovery (DR). Develop, monitor, and report on key Service Level Objectives (SLOs) and Service Level Agreements (SLAs). Lead and mature the incident response lifecycle, focusing on root cause analysis and preventative, long-term remediation. The focus is on achieving production resilience.
4. AI-Accelerated DevOps (Responsible Innovation)
Champion the responsible and ethical integration of generative AI tools to accelerate repeatable engineering tasks such as IaC generation, test case development, and documentation creation. Implement prompt engineering practices to maximize AI effectiveness and reliability. Establish and enforce rigorous verification and validation processes for all AI-generated artifacts. Increased velocity must never compromise engineering correctness, security, or compliance.
5. Platform Architecture & Optimization (Scalability & Performance)
Oversee the continuous evolution of the scalable, resilient architecture for the foundational Kubernetes control plane. Drive the design of resilient Go microservices and REST APIs, emphasizing security and performance. Provide expertise in optimizing MongoDB performance, indexing strategies, and distributed systems design to ensure 24/7 availability and data integrity.
Education and Experience Requirements
BS degree in Computer Science, Computer Engineering or related plus 10 years of experience as a Engineering Manager, Software Engineer, or related.
Special Skills Requirements
Requires 10 years of experience in software engineering, DevOps, or Site Reliability Engineering (SRE). Requires 5 years of programming experience in Python, Go, Java, or a comparable enterprise-grade language; in operating, scaling, and troubleshooting Kubernetes within large, complex production environments; in MongoDB schema design, query optimization, and performance tuning for high volume, mission-critical environments; designing, governing, and securing modern CI/CD pipelines (e.g., Jenkins, GitHub Actions, Azure DevOps) with strong emphasis on automation and compliance controls; and Infrastructure-as-Code (Terraform and/or Ansible) and with one or more major cloud platforms (AWS, Azure, or GCP). Requires 3 years of direct technical leadership experience, leading teams supporting high-scale distributed systems.
Please copy and paste your resume in the email body (do not send attachments, we cannot open them) and email it to candidates at placementservicesusa.com with reference #0048-0005 in the subject line.
Thank you.