MNC InsiderMNC Insider

Senior Production Engineer - SRE

Zoom

Senior Production Engineer - SRE

full-timePosted: Sep 1, 2026Updated: Sep 3, 2026San Jose (CA)

Job Description

What You Can ExpectIn this role, you will be responsible for the reliability, scalability, and operational excellence of Zoom's government and military-facing products. A core focus is deploying and evolving our observability platform, working across Kubernetes, Terraform, and cloud infrastructure to build the systems and frameworks that set the standard for how the team operates. This is a hands-on, on-call engineering role: you will own the systems you build, respond to production incidents with urgency, and deliver permanent improvements.About the TeamWe build and own the cloud infrastructure behind Zoom's government products, operating with full end-to-end accountability. Our team moves fast, solves hard problems, and makes a direct impact where reliability is non-negotiable.ResponsibilitiesArchitecting and owning the end-to-end deployment lifecycle of Zoom's observability platform across AWS GovCloud and Oracle OCI, streamlining processes, eliminating manual steps, and setting the standard for how teams deploy at scaleDesigning and building Terraform modules and providers from scratch, alongside GitLab CI/CD pipelines, to automate infrastructure provisioning and configuration management across development and production environmentsOperating and improving production Kubernetes workloads at scale, leading incident response, driving root cause analysis, and delivering permanent fixes that reduce outage risk and improve system resilienceDeveloping reusable frameworks, runbooks, and operational standards that engineering teams across the organization can adopt — raising the bar on observability onboarding, alerting coverage, and operational readinessCollaborating with engineering teams to define and track SLOs and SLIs, contribute to architecture and design reviews, and systematically eliminate toil through automation and thoughtful infrastructure designWhat We're Looking ForHolds U.S. Citizenship or Lawful Permanent Resident (Green Card) statusBrings 4–5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, or Production Engineering in a live production environmentDemonstrates production-grade Kubernetes experience, including workload operations, troubleshooting, and cluster-level architectural design at scaleAuthors Terraform modules and providers from scratch, not limited to applying or executing existing configurationsOwns incidents end-to-end - proven on-call experience including structured triage, resolution, and post-incident review with documented root cause analysisDeploys, operates, or improves observability or monitoring platforms at production scale, with scripting proficiency in Python and/or Bash for automation and operational toolingCommunicates clearly in writing - produces runbooks, incident reports, methods of procedure, and architectural documentation to a high standardPreferred:Has hands-on Datadog experience building observability frameworks in a production environment, ideally within a government, FedRAMP, or high-compliance cloud contextExperience operating within FedRAMP, AWS GovCloud, or other regulated or compliance-driven cloud environmentsHands-on experience with popular open-source observability tooling in a production contextBackground supporting 24/7 or mission-critical production operationsExperience driving cross-team adoption of operational standards or tooling frameworksOracle OCI hands-on operational experienceFamiliarity with container image security and CVE remediation workflowsSalary Range or On Target Earnings:Minimum:$98,900.00Maximum:$228,700.00In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.We also have a location based compensation structure; there may be a different range for candidates in this and other locationsAt Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!Anticipated Position Close Date:09/17/26Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click Learn for more information.About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom Contact Center, Zoom Phone, Zoom Events, Zoom Apps, Zoom Rooms, and Zoom Webinars.We’re problem-solvers, working at a fast pace to design solutions with our customers and users in mind. Find room to grow with opportunities to stretch your skills and advance your career in a collaborative, growth-focused environment.Our CommitmentAt Zoom, we believe great work happens when people feel supported and empowered. We’re committed to fair hiring practices that ensure every candidate is evaluated based on skills, experience, and potential. If you require an accommodation during the hiring process, let us know—we’re here to support you at every step.If you need assistance navigating the interview process due to a medical disability, please submit an Accommodations Request Form and someone from our team will reach out soon. This form is solely for applicants who require an accommodation due to a qualifying medical disability. Non-accommodation-related requests, such as application follow-ups or technical issues, will not be addressed.Our interviews are supported by BrightHire, a tool that helps us create a consistent and thoughtful interview experience and may include recordings. Please refer to our candidate privacy statement for more information of how we use your data.

Locations

  • San Jose (CA)

Skills Required

  • Platform Engineeringintermediate
  • Python and/or Bash for automationintermediate
  • observability frameworks in a production environmentintermediate
  • container image securityintermediate

Required Qualifications

  • Holds U.S. Citizenship or Lawful Permanent Resident (Green Card) status (experience)
  • Brings 4–5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, or Production Engineering in a live production environment (experience, 5 years)
  • Demonstrates production-grade Kubernetes experience, including workload operations, troubleshooting, and cluster-level architectural design at scale (experience)
  • Authors Terraform modules and providers from scratch, not limited to applying or executing existing configurations (experience)
  • Owns incidents end-to-end - proven on-call experience including structured triage, resolution, and post-incident review with documented root cause analysis (experience)
  • Deploys, operates, or improves observability or monitoring platforms at production scale, with scripting proficiency in Python and/or Bash for automation and operational tooling (experience)
  • Communicates clearly in writing - produces runbooks, incident reports, methods of procedure, and architectural documentation to a high standard (experience)
  • Holds U.S. Citizenship or Lawful Permanent Resident (Green Card) status (experience)
  • Brings 4–5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, or Production Engineering in a live production environment (experience, 5 years)
  • Demonstrates production-grade Kubernetes experience, including workload operations, troubleshooting, and cluster-level architectural design at scale (experience)
  • Authors Terraform modules and providers from scratch, not limited to applying or executing existing configurations (experience)
  • Owns incidents end-to-end - proven on-call experience including structured triage, resolution, and post-incident review with documented root cause analysis (experience)
  • Deploys, operates, or improves observability or monitoring platforms at production scale, with scripting proficiency in Python and/or Bash for automation and operational tooling (experience)
  • Communicates clearly in writing - produces runbooks, incident reports, methods of procedure, and architectural documentation to a high standard (experience)

Preferred Qualifications

  • Has hands-on Datadog experience building observability frameworks in a production environment, ideally within a government, FedRAMP, or high-compliance cloud context (experience)
  • Experience operating within FedRAMP, AWS GovCloud, or other regulated or compliance-driven cloud environments (experience)
  • Hands-on experience with popular open-source observability tooling in a production context (experience)
  • Background supporting 24/7 or mission-critical production operations (experience)
  • Experience driving cross-team adoption of operational standards or tooling frameworks (experience)
  • Oracle OCI hands-on operational experience (experience)
  • Familiarity with container image security and CVE remediation workflows (experience)
  • Has hands-on Datadog experience building observability frameworks in a production environment, ideally within a government, FedRAMP, or high-compliance cloud context (experience)
  • Experience operating within FedRAMP, AWS GovCloud, or other regulated or compliance-driven cloud environments (experience)
  • Hands-on experience with popular open-source observability tooling in a production context (experience)
  • Background supporting 24/7 or mission-critical production operations (experience)
  • Experience driving cross-team adoption of operational standards or tooling frameworks (experience)
  • Oracle OCI hands-on operational experience (experience)
  • Familiarity with container image security and CVE remediation workflows (experience)

Responsibilities

  • Architecting and owning the end-to-end deployment lifecycle of Zoom's observability platform across AWS GovCloud and Oracle OCI, streamlining processes, eliminating manual steps, and setting the standard for how teams deploy at scale
  • Designing and building Terraform modules and providers from scratch, alongside GitLab CI/CD pipelines, to automate infrastructure provisioning and configuration management across development and production environments
  • Operating and improving production Kubernetes workloads at scale, leading incident response, driving root cause analysis, and delivering permanent fixes that reduce outage risk and improve system resilience
  • Developing reusable frameworks, runbooks, and operational standards that engineering teams across the organization can adopt — raising the bar on observability onboarding, alerting coverage, and operational readiness
  • Collaborating with engineering teams to define and track SLOs and SLIs, contribute to architecture and design reviews, and systematically eliminate toil through automation and thoughtful infrastructure design
  • Architecting and owning the end-to-end deployment lifecycle of Zoom's observability platform across AWS GovCloud and Oracle OCI, streamlining processes, eliminating manual steps, and setting the standard for how teams deploy at scale
  • Designing and building Terraform modules and providers from scratch, alongside GitLab CI/CD pipelines, to automate infrastructure provisioning and configuration management across development and production environments
  • Operating and improving production Kubernetes workloads at scale, leading incident response, driving root cause analysis, and delivering permanent fixes that reduce outage risk and improve system resilience
  • Developing reusable frameworks, runbooks, and operational standards that engineering teams across the organization can adopt — raising the bar on observability onboarding, alerting coverage, and operational readiness
  • Collaborating with engineering teams to define and track SLOs and SLIs, contribute to architecture and design reviews, and systematically eliminate toil through automation and thoughtful infrastructure design

Target Your Resume for "Senior Production Engineer - SRE" , Zoom

Get personalized recommendations to optimize your resume specifically for Senior Production Engineer - SRE. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Senior Production Engineer - SRE" , Zoom

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Senior Production Engineer - SRE @ Zoom.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.