Senior Site Reliability Engineer, Gaming - USDS
Responsibilities
We are the Gaming SRE team, responsible for the stability and performance of our global gaming infrastructure. Our mission is to ensure 24/7 seamless gameplay for millions of users through automation, robust monitoring, and rapid incident response. We bridge the gap between engineering and operations to build a reliable and scalable gaming ecosystem.
We are looking for a Site Reliability Engineer who is passionate about building resilient systems that power seamless multiplayer experiences. As an SRE at USDS, you will bridge the gap between game development and infrastructure. You will be responsible for the health, performance, and scalability of our global game servers and backend services, ensuring that players around the world have a lag-free experience 24/7.
Responsibilities
- Scalability & Performance: Design and implement auto-scaling solutions for game dedicated servers (DGS) to handle massive spikes during game launches and seasonal events.
- Infrastructure as Code (IaC): Manage and provision multi-region cloud infrastructure using tools like Terraform or Pulumi.
- Monitoring & Observability: Build robust monitoring dashboards and alerting systems to detect "micro-stutters" or latency issues before they impact the player base.
- Incident Response: Participate in an on-call rotation to troubleshoot and resolve high-priority production issues, followed by blameless post-mortems.
- Cost Optimization: Balance high-performance hardware requirements (like high-clock speed CPUs for game sims) with cloud cost-efficiency.
- CI/CD Pipelines: Streamline the deployment of game builds and backend microservices to ensure rapid, safe releases.
Qualifications
Minimum Qualifications
- Bachelor's degree in Computer Science or a related technical background involving software/system engineering, or equivalent working experience.
- 2+ years of SRE or DevOps experience in large scale online services
- Programming experience with at least one of the following languages: C, C++, Java, Python, C# or Go.
Preferred Qualifications
- Expert knowledge of AWS, Google Cloud (GCP), Azure or OCI (specifically focused on global networking and compute).
- Deep understanding of Linux internals, networking protocols (UDP/TCP), and performance tuning
- Extensive knowledge of networking, operation systems, database systems and container technology.
- Good understanding of every aspect of microservice architecture, and hands on experience in troubleshooting in large scale distributed systems.
- Hands on experience in common opensource systems such as Linux, MySQL, MongoDB, Redis and ELK.
- Experience in building solutions with AWS, Google, Azures and other cloud services is a plus.
- Passionate, self-motivated and good teamwork skills.
Job Information
[For Pay Transparency] Compensation Description (annually)
The base salary range for this position in the selected city is $187040 - $438000 annually.
Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.
Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
Want more jobs like this?
Get jobs in San Jose, CA delivered to your inbox every week.

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.
For Los Angeles County (unincorporated) Candidates:
Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:
1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
3. Exercising sound judgment.
Perks and Benefits
Health and Wellness
- Health Insurance
- Dental Insurance
- Vision Insurance
- HSA
- Life Insurance
- Fitness Subsidies
- Short-Term Disability
- Long-Term Disability
- On-Site Gym
- Mental Health Benefits
- Virtual Fitness Classes
Parental Benefits
- Fertility Benefits
- Adoption Assistance Program
- Family Support Resources
Work Flexibility
- Flexible Work Hours
- Hybrid Work Opportunities
Office Life and Perks
- Casual Dress
- Snacks
- Pet-friendly Office
- Happy Hours
- Some Meals Provided
- Company Outings
- On-Site Cafeteria
- Holiday Events
Vacation and Time Off
- Paid Vacation
- Paid Holidays
- Personal/Sick Days
- Leave of Absence
Financial and Retirement
- 401(K) With Company Matching
- Performance Bonus
- Company Equity
Professional Development
- Promote From Within
- Access to Online Courses
- Leadership Training Program
- Associate or Rotational Training Program
- Mentor Program
Diversity and Inclusion
- Diversity, Equity, and Inclusion Program
- Employee Resource Groups (ERG)
Company Videos
Hear directly from employees about what it is like to work at TikTok.