Unknown Company

Site Reliability - Software Engineer (II)

pittsburgh, pa • Posted 4 days ago
Remote Contract Architecture and Engineering Occupations
Site Reliability - Software Engineer (II)

Duration: 6 months contract Location: Pittsburgh, PA (Hybrid: Tues-Thurs in the office. Mon & Friday remote)

As a Site Reliability Engineer (SRE-SWE), you deliver medium sized projects from start to finish with minimal supervision. You work within one or more teams to share knowledge related to building solutions and tools to configure, maintain, and scale systems and services. You produce initial project designs for your team and/or make significant contributions to larger designs and work to execute on those designs along with other SREs. You are able to recognize issues in a large service area, plan and execute the solution, and address commonly escalated issues or triage when required.

Note: In accordance with regional legislation, in certain locations the job title for this role might be "Developer" or "Programmer" rather than "Engineer".

Responsibilities include:

  • Maintain standards of the team towards production excellence. Assist in defining improved standards for production excellence within a team.
  • Promote insights, best practices, and standards that can be applied to improve operations on the immediate team. Advocate and drive adoption of best practices and SRE converged solutions for systems design and management.
  • Set up or improve test, monitoring, or other infrastructure to ensure system scalability, reliability, and efficiency over time. Contribute to existing documentation or educational content and keep content updated over time. Triage product or system issues and debug/track/resolve by analyzing the sources of issues and the impact on hardware, network, or service operations and quality. Identify and contribute to remediation and preventative work.
  • Write and improve code in tools, libraries, or systems with a focus on reliability, resilience, scalability, and toil reduction with minimal guidance and direction. Develop interfaces between system components that may span team boundaries or RPC boundaries. Simplify code by choosing appropriate algorithms and abstractions, and by utilizing a substantial subset of standard google3 libraries in the corresponding language, as appropriate. Review code developed by others and provide feedback to ensure best practices. Test non-standard scenarios to cover all cases and any potential design issues, validate resilience, and write test case descriptions to ensure coverage of critical components with minimal guidance.
  • Review technical designs within a team with an emphasis on reliability, efficiency, simplicity, and maintainability. Design solutions to problems that may have multiple approaches to support the team's needs, or in response to reliability - and technology-related incidents. Understand constraints, identify solutions, and defend designs.

Minimum role qualification requires proficiency in: - Code Comprehension (Knowledge) - Architecture reliability - Architectural patterns - Data structures and algorithms - Software development lifecycle (SDLC) - Systems thinking - Programming - Code health and tools (IC) - Reliability debugging and bug-fixing - Data analysis and synthesis - Simplification - Risk analysis and management

Back to Job Search