Unknown Company

Principal System Engineer

plano, tx • Posted 4 days ago
Hybrid Full Time General

Principal System Engineer– EngOps T2Location: TechM US Texas PlanoYears of Experience: 7–10 Years SREAs a Tier 2/Site Reliability Engineer (SRE), you will translate core business requirements into robust, scalable, and reliable technical solutions. You'll play a pivotal role in designing and implementing applications, platforms, and services that power critical business operations, with a strong emphasis on high availability, performance, and compliance in cloud, messaging, and data environments. Individual will possess the experience skills that includes a hybrid of traditional T2/SRE operations technical skills to support our Project Growth apps and new, evolving Generative AI and Workflow Automation skillsets needed to drive operational efficiency and scalability.

Provide technical expertise and best practices for Java, Python, JavaScript, and Perl-based solutions. Strong knowledge of network and telecom standards (3GPP, TM Forum, etc.). Practical understanding of AI/ML concepts and their integration in enterprise platforms.Key ResponsibilitiesThe EngOps Tier 2/SRE team ensures applications and systems are highly reliable, scalable, and performant while fostering a collaborative culture between development and operations.Work with T1 team on incident as Triage lead during outages or critical issues Pager duty issues Minimize downtime and user impact during incidents.Conduct detailed After Action Reviews involving all stakeholders and chalk out short term and long-term resiliency options.Eliminate recurrence of similar issues through systemic fixes.Define and implement monitoring and alerting strategies tailored to the launch.Collaborate with Product development teams to gain deep insight into the application architecture, flows and critical dependencies.Monitor and evaluate key performance metrics like latency, throughput, and error rates and update alerts.Propose architectural or operational changes to prevent reoccurrence.Reduce Mean Time to Resolution (MTTR) for incidents.Required QualificationsEducation: Bachelor's degree in Computer Science, Information Systems, or a related discipline.Experience: Over 10 years hands-on experience in architecting and building scalable platforms and applications in cloud/data environments.Nice-to-Have Practical understanding of AI/ML concepts and their integration in enterprise platforms.

Back to Job Search