Unknown Company

Site Reliability Engineer, Reliability Team - USDS

san jose, ca • Posted 6 days ago
Onsite Full Time Engineering

Role Overview

The Site Reliability Engineering (SRE) team at TikTok builds and operates large‑scale, fault‑tolerant systems that power TikTok’s core services worldwide. The team focuses on enhancing observability and operability, using data insights to maintain 24/7 business stability.

Responsibilities

  • Design and optimize high‑concurrency distributed systems, collaborating with development teams to ensure scalability, reliability, and high availability.
  • Build and maintain automation tools to reduce toil, streamline deployments, and manage infrastructure as code.
  • Develop and refine monitoring, alerting, and logging systems (SLIs/SLOs) for deep visibility into service health and performance.
  • Lead global disaster‑recovery drills, simulate complex failure scenarios, and validate failover mechanisms to keep the platform operational under extreme conditions.
  • Respond to high‑priority incidents as a key responder, coordinating cross‑functional war rooms, driving technical troubleshooting, and leading service restoration.
  • Facilitate blameless post‑mortems and root‑cause analysis; transform incident insights into engineering requirements to harden systems.
  • Manage capacity planning, resource allocation, and performance bottlenecks to accommodate organic growth and traffic surges.

Minimum Qualifications

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • Proficiency in one or more programming languages (e.g., Go, Python, Java, or C++).
  • Strong understanding of Linux system internals, networking (TCP/IP, DNS, load balancing), and distributed systems.
  • Experience managing containerized environments such as Kubernetes or Docker.

Preferred Qualifications

  • Experience in a high‑traffic production environment with a focus on incident response and site stability.
  • Hands‑on knowledge of disaster‑recovery strategies, including multi‑region failover and data consistency in distributed databases.
  • Familiarity with observability and monitoring tools.
  • Experience with Infrastructure as Code.

Benefits & Compensation

Compensation for this position ranges from $122,574 to $259,200 annually, with potential bonuses, incentives, and restricted stock units. Benefits include medical, dental, and vision insurance; a 401(k) plan with company match; paid parental leave; short‑term and long‑term disability coverage; life insurance; and wellbeing benefits. Employees receive 10 paid holidays, 10 paid sick days, and 17 days of paid personal time (prorated on hire and increasing with tenure).

Work Environment

The company is shifting to a fully in‑person schedule, up to five days a week, to enhance speed, alignment, and agility across teams.

Reasonable Accommodation

USDS is committed to providing reasonable accommodations for candidates with disabilities, pregnancy, sincerely held religious beliefs, or other protected reasons. Assistance can be requested at

#J-18808-Ljbffr
Back to Job Search