Unknown Company

Observability Engineer

irving, tx • Posted 3 days ago
Onsite Full Time General

Senior Observability EngineerWe are seeking a Senior Observability Engineer to architect and implement robust observability frameworks. The ideal candidate will translate platform telemetry into actionable insights for engineering, operations, and leadership. This role involves architecting end-to-end observability and alerting frameworks and building the dashboards and alerts that utilize them.Key ResponsibilitiesArchitect end-to-end observability frameworks spanning cloud platforms, on-premises infrastructure, networking, databases, middleware, and application layers.Evaluate, select, and integrate observability tooling across the stack, establishing reference architectures.Design and maintain standardized Grafana dashboards for infrastructure, platform, and workload health across OCP, AKS, and GKE.Define golden signals and platform health KPIs aligned to availability, performance, capacity, and reliability.Act as an advanced Splunk user, writing complex SPL for investigation, correlation, and root-cause analysis.Correlate metrics, logs, and events across Grafana and Splunk to shorten mean time to resolution (MTTR).Implement platform-specific observability patterns for OCP, AKS, and GKE.Configure and manage ThousandEyes for synthetic monitoring and network intelligence.Administer and tune BigPanda for AIOps-driven event correlation and noise reduction.Design and maintain ServiceNow integrations for automated incident creation and alert enrichment.Document observability standards, dashboard conventions, and onboarding patterns.Required QualificationsExperience:7+ years of total IT experience, including IT operations, site reliability engineering, systems engineering, or a related infrastructure discipline.5+ years of hands-on experience designing and implementing observability, monitoring, or telemetry solutions across distributed systems.5+ years of experience working with Kubernetes platforms in production environments.Technical Skills:Observability experience with logging events, tracing, and enhanced monitoring.Strong hands-on experience with Grafana for dashboarding, variables, alerts, and data sources.Advanced proficiency with Splunk SPL, dashboards, and investigations.Ability to use various querying languages (e.g., SQL, PromQL).Deep understanding of Kubernetes internals and hands-on experience with OpenShift (OCP), AKS, and/or GKE.Experience with Prometheus, OpenTelemetry, and Kubernetes exporters.Solid understanding of OS-level monitoring (Linux/Windows) and networking fundamentals.Experience with ThousandEyes, BigPanda, and ServiceNow ITSM workflows.Preferred QualificationsExperience designing SLOs/SLIs and reliability scorecards.Familiarity with service mesh metrics (Istio / mTLS).Experience with capacity planning and trend analysis using observability data.Exposure to multi-cloud observability strategy and platform standardization.Experience monitoring databases, message brokers, or middleware.Familiarity with AIOps or ML-driven anomaly detection in observability pipelines.

Back to Job Search