Why Join GEICO?
At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities. Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide. Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the UnitedStates. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.
Position Summary
Technology Operations Center is at the core of GEICO’s application and platform resiliency. It assists GEICO’s engineering teams with maintaining high availability of our customer and internally facing services while driving down time to detect and recover from incidents. It is governing our incident management processes and builds platforms that allow GEICO to manage and recover from incidents. GEICO is seeking an experienced SRE Software Engineer with a passion for building, operating and troubleshooting high-performance, low-maintenance, zero-downtime complex distributed platforms and applications. You will help drive our transformation to a tech organization with engineering excellence and site reliability as its mission, by defining, implementing and operating our incident management processes and in-house technology platforms that automate them. This role focuses on improving Incident Management tooling and process across GEICO. It is a hands-on technical leadership role focused on better incident management, faster time to detect, troubleshoot and recover from incidents and fewer repeat incidents, by designing, developing and operating the the tools and processes that help all GEICO engineering teams manage their on-call, runbooks, troubleshooting and BCDR.
Why This Role Is Different
This role blends deep technical understanding, hands on execution and ability to build SW with real time incident leadership, platform and process improvements. You will be driving our incident management and response processes and the platforms that automate them. You are shaping how GEICO service engineering teams detect, manage, resolve and learn from the incidents. Your work directly impacts the availability of GEICO’s critical applications and platforms, experiences and satisfaction of millions of customers and tens of thousands of associates.
Position Responsibilities
- Own and Evolve Enterprise-Critical Platforms Design, develop and operate automation, self-service tools, dashboards, and data pipelines that automate and scale our incident management, on-call, paging and troubleshooting processes.
- Build shared services, APIs, data contracts, automation, and integrations that standardize incident response and reduce systemic operational risk.
- Set and uphold engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness.
- Champion safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls that support frequent and reliable delivery.
- Evaluate, select, and implement modern technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness.
- Serve as a Senior Technical Authority During Incidents Act as a technical leader and escalation point during high-severity incidents, bringing architectural judgment, system-level problem solving, and calm execution under pressure.
- Guide troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision making to restore service safely and efficiently.
- Lead or heavily influence post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
- Develop and maintain incident response strategies, operational runbooks, readiness criteria, triage models, and resilience practices across multiple integration points.
- Influence Technical Direction and Engineering Culture Lead complex design and architecture reviews spanning multiple teams, services, dependencies, and operational domains.
- Partner with SRE, platform, product, infrastructure, security, and business stakeholders to align operational tooling with practical engineering needs and enterprise reliability goals.
- Translate technical concepts, risks, and tradeoffs clearly for executives, technical leaders, and non-technical stakeholders.
- Mentor senior engineers, staff engineers, and teams through technical leadership, example, code/design reviews, documentation, and operational coaching.
- Reinforce a culture of ownership, accountability, continuous improvement, psychological safety, learning and operational excellence.
- We have adopted “You Build it You Run it strategy. All our senior technologists take an active role in leading and managing high-severity incidents requiring strong technical judgment, clear communication and calm execution under pressure.
- All our engineers have on-call responsibilities as part of a 24x7 rotation supporting incident response and production support for mission-critical platforms and processes they build and operate.
Qualifications
- Hands on proficiency in multiple languages, including Go, Java, Python, C# for building production-grade full stack applications on Kubernetes and serverless technologies (KNative) in Azure and AWS.
- Experience with SQL and NoSQL technologies and cloud-native services for storing and analyzing incident data.
- Experience with building and using data pipelines analytics, and dashboards for operational metrics, trends, and KPIs using technologies such as Spark, Trino, Grafana, Superset, PowerBI Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, Azure Monitor.
- Experience with incident management platforms such as PagerDuty.
- Proficiency with AI assisted development processes and tools such as Claude Code, Cursor and GitHub Copilot.
- Experience improving incident, post incident review, or reliability processes at scale through automation, data and cross-team influence.
- Deep incident forensics and root cause analysis skills, with the ability to raise COE quality through clear action items and follow-through across teams.
- Strong understanding of observability, reliability engineering, incident management and post-incident improvement practices.
- Experience supporting incident response and high-severity production incidents in complex environments.
- Strong software engineering fundamentals and system design skills, with experience building reliable production systems at scale.
- Ability to lead technical design and architecture decisions in complex distributed systems.
- Strong communication skills and the ability to coach engineering teams and present findings clearly to leadership.
- Experience 10+ years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems or operational tooling.
- 8+ years of experience with architecture, design, system reliability, scalability and technical leadership for production systems.
- 6+ years of experience with open-source frameworks, modern engineering practices or platform technologies.
- 4+ years of experience with Azure, AWS, GCP or another cloud service provider or equivalent experience in complex hybrid environments.
- Demonstrated ownership of mission-critical systems operating in 24x7 production environments.
- Education Bachelor's degree in Computer Science, Information Systems or equivalent education or work experience.
Additional Job Requirements
Ability to influence engineering outcomes across teams in complex organizations. Must be able to communicate in a clear concise professional oral and written manner with customers, clients, co workers and leadership and other employees of the organization. Must be able to perform effectively under pressure and in stressful situations, including during production support and high-severity incident response. Must be able to participate in a 24x7 on call rotation for incident response and production support of mission-critical platforms.
Annual Salary
#J-18808-LjbffrSenior Staff Engineer in bethesda at Unknown Company
This position is listed as contract and hybrid.