MNC InsiderMNC Insider

Lead Site Reliability Engineer, Platforms

Zoom

Lead Site Reliability Engineer, Platforms

full-timePosted: Sep 3, 2026San Jose (CA)

Job Description

What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc).About the TeamThe DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG).The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat.ResponsibilitiesDesign and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environmentsDefine technical roadmaps and architectural direction for infrastructure automation and securityPartner with service teams to understand platform needs and deliver solutions that improve reliability and efficiencyEstablish and advocate for SRE best practices including infrastructure as code, monitoring, and incident managementMentor team members through design, implementation, and production deployment of complex systemsWhat We're Looking ForBring 8+ years of SRE or DevOps experience building and operating production infrastructure at scaleCode proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java)Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrogOperate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and KubernetesImplement observability solutions with logging and monitoring tools such as ELK, Prometheus, or GrafanaCommunicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditorsParticipate in on-call rotations and lead incident response to maintain system reliabilityHold a degree in Computer Science or related field, or equivalent practical experienceHold US citizenship, or Greencard statusPreferredHave experience with security from an SRE perspectiveHave experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta)Have experience operating Government environments and understanding their compliance requirementsHave experience with system design and distributed computing at scaleAbility to speak Chinese/Mandarin is a plus, but not requiredSalary Range or On Target Earnings:Minimum:$124,000.00Maximum:$271,200.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learn for more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.

Locations

  • San Jose (CA)

Skills Required

  • and operating production infrastructure at scaleintermediate
  • security from an SRE perspectiveintermediate
  • Identity securityintermediate
  • system designintermediate

Required Qualifications

  • Bring 8+ years of SRE or DevOps experience building and operating production infrastructure at scale (experience, 8 years)
  • Code proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java) (experience)
  • Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrog (experience)
  • Operate cloud infrastructure on AWS, OCI, or similar platforms using Terraform and Kubernetes (experience)
  • Implement observability solutions with logging and monitoring tools such as ELK, Prometheus, or Grafana (experience)
  • Communicate complex technical concepts clearly to diverse audiences including security teams, senior leadership, and external auditors (experience)
  • Participate in on-call rotations and lead incident response to maintain system reliability (experience)
  • Hold a degree in Computer Science or related field, or equivalent practical experience (experience)
  • Hold US citizenship, or Greencard status (experience)

Preferred Qualifications

  • Have experience with security from an SRE perspective (experience)
  • Have experience with Identity security (e.g. IAM, workload identity, zero trust) and tools (e.g. Teleport, Okta) (experience)
  • Have experience operating Government environments and understanding their compliance requirements (experience)
  • Have experience with system design and distributed computing at scale (experience)
  • Ability to speak Chinese/Mandarin is a plus, but not required (experience)

Responsibilities

  • Design and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environments
  • Define technical roadmaps and architectural direction for infrastructure automation and security
  • Partner with service teams to understand platform needs and deliver solutions that improve reliability and efficiency
  • Establish and advocate for SRE best practices including infrastructure as code, monitoring, and incident management
  • Mentor team members through design, implementation, and production deployment of complex systems

Target Your Resume for "Lead Site Reliability Engineer, Platforms" , Zoom

Get personalized recommendations to optimize your resume specifically for Lead Site Reliability Engineer, Platforms. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Lead Site Reliability Engineer, Platforms" , Zoom

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Lead Site Reliability Engineer, Platforms @ Zoom.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.