MNC InsiderMNC Insider

Site Reliability Engineer III

Yum! Brands

Site Reliability Engineer III

full-timePosted: Jul 8, 2026Updated: Aug 29, 2026Ho Chi Minh, Dong Nam Bo, Viet Nam

Job Description

The Site Reliability Engineer (Level 7) is an experienced mid-level individual contributor responsible for the reliability, scalability, performance, and operational excellence of one or more markets, platforms, or critical services. Incident Management and ReliabilityIndependently lead complex incidents involving multiple systems, teams, or dependencies.Coordinate incident response activities, facilitate communication, and drive timely resolution.Lead or contribute to post-incident reviews and root cause analysis activities.Ensure corrective and preventive actions are identified, prioritized, tracked, and completedMonitoring, Alerting, and ObservabilityDesign, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.Develop meaningful alerts based on service behavior, customer impact, and business priorities.Build and maintain dashboards that provide actionable insights into system performance and reliability.Platform and Market OwnershipOwn SRE responsibilities for one or more markets, platforms, or critical services end-to-end.Establish and maintain operational excellence standards for assigned domains.Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.Continuously assess platform health, identify reliability risks, and drive improvements before incidents occur.Partner with engineering teams to ensure new features and services meet reliability requirements before production release.Define and track reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.Automation, and AIDevelop and maintain tools, scripts, and automation solutions that reduce manual effort and improve operational efficiency.Identify and eliminate repetitive tasks through automation and self-service capabilities.Establish and promote best practices for the responsible use of AI within SRE workflowsTeam Contribution and MentoringMentor and support Level 5 and Level 6 engineers in incident management, monitoring, automation, AI adoption, and operational best practices.Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.Share knowledge through training sessions, documentation, and post-incident learning activities.Contribute to the continuous improvement of team processes, standards, and ways of working. Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end.Establish and maintain operational excellence standards for assigned domains.Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.Coordinate incident response activities, facilitate communication, and drive timely resolution.Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.Develop meaningful alerts based on service behaviours, customer impact, and business priorities.Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.Lead or contribute to post-incident reviews and root cause analysis activities.Influence technical decisions that improve platform stability, scalability, and operational efficiency

Locations

  • Ho Chi Minh, Dong Nam Bo, Viet Nam

Required Qualifications

  • Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end. (experience)
  • Establish and maintain operational excellence standards for assigned domains. (experience)
  • Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date. (experience)
  • Coordinate incident response activities, facilitate communication, and drive timely resolution. (experience)
  • Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions. (experience)
  • Develop meaningful alerts based on service behaviours, customer impact, and business priorities. (experience)
  • Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards. (experience)
  • Lead or contribute to post-incident reviews and root cause analysis activities. (experience)
  • Influence technical decisions that improve platform stability, scalability, and operational efficiency (experience)

Target Your Resume for "Site Reliability Engineer III" , Yum! Brands

Get personalized recommendations to optimize your resume specifically for Site Reliability Engineer III. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Site Reliability Engineer III" , Yum! Brands

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

DigitalDigital

Answer 10 quick questions to check your fit for Site Reliability Engineer III @ Yum! Brands.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.