Job Description
nWe are seeking two Site Reliability Engineers (SREs) to join our team supporting a new Azure-based product. This role focuses on system reliability, observability, and monitoring for a data-driven application that provides KPIs and insights to end users daily. The product leverages Azure services, APIs, Databricks, and AI/ML models to process customer data and populate dashboards refreshed once per day.
nThe SREs will ensure the reliability of the entire pipeline, provide hypercare support, and collaborate with engineering teams to streamline monitoring and alerting processes.
nWe are a company committed to creating inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity employer that believes everyone matters. Qualified candidates will receive consideration for employment opportunities without regard to race, religion, sex, age, marital status, national origin, sexual orientation, citizenship status, disability, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to Human Resources Request Form. The EEOC "Know Your Rights" Poster is available here.
nWe are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
nSkills and Requirements
nExperience in SRE or similar reliability-focused roles.
nStrong knowledge of Azure services and cloud-based architectures.
nHands-on experience with observability, monitoring, and alerting tools (App Insights, Elastic).
nAbility to work with REST APIs and understand event-driven architectures (e.g., Service Bus).
nProficiency in C# for troubleshooting and minor coding tasks.
nExcellent communication and ownership mindset-able to manage issues end-to-end. Experience with Terraform and infrastructure-as-code.
nFamiliarity with Databricks, AI/ML pipelines, and data engineering concepts.
nKnowledge of React for front-end troubleshooting.
nExposure to event-driven distributed systems.
nAbility to streamline monitoring processes across multiple teams.