MNC InsiderMNC Insider

Site Reliability Engineer

Procter & Gamble

Site Reliability Engineer

full-timePosted: Aug 3, 2026Updated: Aug 28, 2026Taguig City

Job Description

Job LocationTaguig CityJob DescriptionInformation Technology (IT) at Procter & Gamble is where business, innovation and technology integrate to build a competitive advantage for P&G. Our mission is clear -- you deliver IT to help P&G win with consumers. Do you love implementing continuous improvement in IT solutions to drive efficiency and agility in meeting constantly evolving business needs? Then this job might be for you! As a Site Reliability Engineer, you will be instrumental in ensuring the high availability and reliability of our digital IT products in P&G. Your primary focus will be on enhancing system performance through faster detection, response, and resolution of issues, while also implementing strategies to prevent recurrence and reduce operational toil. You will use robust Observability and Monitoring tools, automate incident response systems, and optimize IT architecture to create a resilient and reliable infrastructure. This is a Managerial position. Being a manager at P&G involves leading teams and / or end-to-end processes, managing P&G resources, and driving business results. Managers are responsible for overseeing various aspects of the business, including strategy, operations, and team performance. They play a crucial role in ensuring that P&G's brands continue to grow and succeed in the market. Managers at P&G are expected to have strong leadership skills, a growth mindset, and the ability to make data-driven decisions roles lead and initiatives, significantly impacting business results through independent judgment and minimal guidance.Responsibilities: Implement and lead comprehensive monitoring solutions and tools to provide real-time insights into system performance, enabling proactive incident detection and ensuring accurate, actionable alerts for prompt responses. Continuously refine monitoring strategies and develop automation scripts to address recurring issues, enhancing system visibility, resource optimization, and overall efficiency. Establish and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to improve service quality and reliability, Collect and share data and insights from observability tools to drive continuous improvement initiatives. Work closely with Software Engineers, Product Teams, and Infrastructure Teams to develop and implement initiatives that enhance IT reliability. Engage with customers to understand their needs and difficulties regarding Observability and Monitoring tools, providing exceptional support in all interactions, including communications, updates, and feedback. Stay updated on industry trends and effective strategies in Site Reliability Engineering while continuously enhancing technical skills in system architecture, automation, cloud technologies, and operational processes. Job QualificationsCandidates must demonstrate strong leadership in the application of technical expertise to drive business results. We are looking for candidates who possess the following core qualities: A Bachelor's degree in related field such as Engineering, Information Technology and Computer Science discipline, and up to 5 years experience at most. Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana) Knowledge and familiarity in system administration, including Linux/Unix environments, cloud platforms (Azure or GCP preferred, but AWS is acceptable) Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform) Proficiency in at least one programming language (e.g., Python, C#) and a background in scripting for automation tasks Understanding of networking protocols, network infrastructures, load balancing, and DNS management Familiarity with containerization and Orchestration Technologies (e.g., Docker, Kubernetes) Familiarity with databases and proficiency in writing SQL queries Understanding of best practices in security and experience with implementing secure systems Knowledge of incident response methodologies, root cause analysis, and implementing preventive measures (ITIL and/or SRE) Familiarity with ticketing systems and task management (preferably ServiceNow) Problem-solving skills with ability to analyze complex issues and devise effective solutions Learning agility as there will be new topics to learn and new spaces to understand Communication and collaboration skills to work effectively with multi-functional teams, partners, and customers Teamwork and interpersonal skills, with an ability to build relationships and work effectively in a collaborative environment Operational excellence / execution skills as the work requires discipline Job ScheduleFull timeJob NumberR000156530Job SegmentationEntry Level

Locations

  • Taguig City

Skills Required

  • monitoringintermediate
  • configuration management toolsintermediate
  • at least one programming languageintermediate
  • scripting for automation tasksintermediate
  • containerizationintermediate
  • writing SQL queriesintermediate
  • databasesintermediate
  • implementing secure systemsintermediate
  • incident response methodologiesintermediate
  • ticketing systemsintermediate

Required Qualifications

  • Candidates must demonstrate strong leadership in the application of technical expertise to drive business results. (experience)
  • We are looking for candidates who possess the following core qualities: (experience)
  • A Bachelor's degree in related field such as Engineering, Information Technology and Computer Science discipline, and up to 5 years experience at most. (experience, 5 years)
  • A Bachelor's degree in related field such as Engineering, Information Technology and Computer Science discipline, and up to 5 years experience at most. (experience, 5 years)
  • Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana) (experience)
  • Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana) (experience)
  • Knowledge and familiarity in system administration, including Linux/Unix environments, cloud platforms (Azure or GCP preferred, but AWS is acceptable) (experience)
  • Knowledge and familiarity in system administration, including Linux/Unix environments, cloud platforms (Azure or GCP preferred, but AWS is acceptable) (experience)
  • Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform) (experience)
  • Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform) (experience)
  • Proficiency in at least one programming language (e.g., Python, C#) and a background in scripting for automation tasks (experience)
  • Proficiency in at least one programming language (e.g., Python, C#) and a background in scripting for automation tasks (experience)
  • Understanding of networking protocols, network infrastructures, load balancing, and DNS management (experience)
  • Understanding of networking protocols, network infrastructures, load balancing, and DNS management (experience)
  • Familiarity with containerization and Orchestration Technologies (e.g., Docker, Kubernetes) (experience)
  • Familiarity with containerization and Orchestration Technologies (e.g., Docker, Kubernetes) (experience)
  • Familiarity with databases and proficiency in writing SQL queries (experience)
  • Familiarity with databases and proficiency in writing SQL queries (experience)
  • Understanding of best practices in security and experience with implementing secure systems (experience)
  • Understanding of best practices in security and experience with implementing secure systems (experience)
  • Knowledge of incident response methodologies, root cause analysis, and implementing preventive measures (ITIL and/or SRE) (experience)
  • Knowledge of incident response methodologies, root cause analysis, and implementing preventive measures (ITIL and/or SRE) (experience)
  • Familiarity with ticketing systems and task management (preferably ServiceNow) (experience)
  • Familiarity with ticketing systems and task management (preferably ServiceNow) (experience)
  • Problem-solving skills with ability to analyze complex issues and devise effective solutions (experience)
  • Problem-solving skills with ability to analyze complex issues and devise effective solutions (experience)
  • Learning agility as there will be new topics to learn and new spaces to understand (experience)
  • Learning agility as there will be new topics to learn and new spaces to understand (experience)
  • Communication and collaboration skills to work effectively with multi-functional teams, partners, and customers (experience)
  • Communication and collaboration skills to work effectively with multi-functional teams, partners, and customers (experience)
  • Teamwork and interpersonal skills, with an ability to build relationships and work effectively in a collaborative environment (experience)
  • Teamwork and interpersonal skills, with an ability to build relationships and work effectively in a collaborative environment (experience)
  • Operational excellence / execution skills as the work requires discipline (experience)
  • Operational excellence / execution skills as the work requires discipline (experience)

Responsibilities

  • Implement and lead comprehensive monitoring solutions and tools to provide real-time insights into system performance, enabling proactive incident detection and ensuring accurate, actionable alerts for prompt responses.
  • Implement and lead comprehensive monitoring solutions and tools to provide real-time insights into system performance, enabling proactive incident detection and ensuring accurate, actionable alerts for prompt responses.
  • Continuously refine monitoring strategies and develop automation scripts to address recurring issues, enhancing system visibility, resource optimization, and overall efficiency.
  • Continuously refine monitoring strategies and develop automation scripts to address recurring issues, enhancing system visibility, resource optimization, and overall efficiency.
  • Establish and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to improve service quality and reliability,
  • Establish and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to improve service quality and reliability,
  • Collect and share data and insights from observability tools to drive continuous improvement initiatives.
  • Collect and share data and insights from observability tools to drive continuous improvement initiatives.
  • Work closely with Software Engineers, Product Teams, and Infrastructure Teams to develop and implement initiatives that enhance IT reliability.
  • Work closely with Software Engineers, Product Teams, and Infrastructure Teams to develop and implement initiatives that enhance IT reliability.
  • Engage with customers to understand their needs and difficulties regarding Observability and Monitoring tools, providing exceptional support in all interactions, including communications, updates, and feedback.
  • Engage with customers to understand their needs and difficulties regarding Observability and Monitoring tools, providing exceptional support in all interactions, including communications, updates, and feedback.
  • Stay updated on industry trends and effective strategies in Site Reliability Engineering while continuously enhancing technical skills in system architecture, automation, cloud technologies, and operational processes.
  • Stay updated on industry trends and effective strategies in Site Reliability Engineering while continuously enhancing technical skills in system architecture, automation, cloud technologies, and operational processes.

Target Your Resume for "Site Reliability Engineer" , Procter & Gamble

Get personalized recommendations to optimize your resume specifically for Site Reliability Engineer. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Site Reliability Engineer" , Procter & Gamble

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Site Reliability Engineer @ Procter & Gamble.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.