MNC InsiderMNC Insider
Oracle logo

Principal Site Reliability Engineer

Oracle

Principal Site Reliability Engineer

full-timePosted: Aug 21, 2026Updated: Aug 27, 2026ZAPOPAN, JALISCO, Mexico

Job Description

As a Principal member of the Site Reliability Engineering (SRE) team, you'll take ownership of highly available systems, influence service design, and work across teams to drive resiliency, automation, and operational excellence. This is a hands-on engineering role where deep infrastructure knowledge meets software engineering expertise, ideal for experienced SREs ready to take the lead. This is not a fully remote role but a hybrid role. Does require in office at least 3 days a week in Guadalajara. What You’ll Do:Lead the design, automation, and support of OCI services with a focus on resiliency, security, scalability, and performance.Own and improve the end-to-end reliability metrics (SLOs, SLAs, KPIs) for your services.Design and implement high-availability architectures and standards for large-scale distributed systems.Serve as the ultimate escalation point for complex operational issues, using a deep understanding of service topologies and interdependencies.Architect and build automation and orchestration tools that reduce manual work and prevent problem recurrence.Collaborate with development teams to improve service designs, optimize deployments, and implement best practices for operational efficiency.Guide technical decision-making and mentor junior SREs and developers across teams.Participate in and lead postmortems, root cause analysis, and preventative design changes.Contribute to capacity planning, demand forecasting, and long-term service scalability strategies.Participate in a rotational on-call schedule to ensure the health and availability of production services. What We’re Looking For:Advanced experience with Linux systems administrationThis is not a fully remote role but a hybrid role. Does require in office at least 3 days a week in Guadalajara. Strong programming skills in Python (with automation libraries)Advanced Bash/Shell scriptingDeep understanding of distributed systems, networking, and service architectureSolid knowledge of databases and how they behave in production (SQL or NoSQL)Strong understanding of CI/CD pipelines, Agile methodologies, and DevOps best practicesExperience writing and maintaining unit tests and production-grade softwareProven ability to lead cross-functional efforts and technical problem-solving in live environments Nice to Have:Hands-on experience with monitoring and observability tools (Grafana, Prometheus, New Relic, etc.)Familiarity with Oracle Cloud Infrastructure (OCI) or other cloud platforms (AWS, Azure, GCP)Experience with Infrastructure-as-Code (Terraform, Ansible) and container orchestration (Kubernetes) Career Level - IC4

Locations

  • ZAPOPAN, JALISCO, Mexico

Skills Required

  • Linux systems administrationintermediate
  • databasesintermediate
  • and maintaining unit testsintermediate
  • monitoringintermediate
  • Oracle Cloud Infrastructureintermediate
  • Infrastructure-as-Codeintermediate

Required Qualifications

  • Advanced experience with Linux systems administration (experience)
  • This is not a fully remote role but a hybrid role. Does require in office at least 3 days a week in Guadalajara. (experience)
  • Strong programming skills in Python (with automation libraries) (experience)
  • Advanced Bash/Shell scripting (experience)
  • Deep understanding of distributed systems, networking, and service architecture (experience)
  • Solid knowledge of databases and how they behave in production (SQL or NoSQL) (experience)
  • Strong understanding of CI/CD pipelines, Agile methodologies, and DevOps best practices (experience)
  • Experience writing and maintaining unit tests and production-grade software (experience)
  • Proven ability to lead cross-functional efforts and technical problem-solving in live environments (experience)

Preferred Qualifications

  • Hands-on experience with monitoring and observability tools (Grafana, Prometheus, New Relic, etc.) (experience)
  • Familiarity with Oracle Cloud Infrastructure (OCI) or other cloud platforms (AWS, Azure, GCP) (experience)
  • Experience with Infrastructure-as-Code (Terraform, Ansible) and container orchestration (Kubernetes) (experience)

Responsibilities

  • Lead the design, automation, and support of OCI services with a focus on resiliency, security, scalability, and performance.
  • Own and improve the end-to-end reliability metrics (SLOs, SLAs, KPIs) for your services.
  • Design and implement high-availability architectures and standards for large-scale distributed systems.
  • Serve as the ultimate escalation point for complex operational issues, using a deep understanding of service topologies and interdependencies.
  • Architect and build automation and orchestration tools that reduce manual work and prevent problem recurrence.
  • Collaborate with development teams to improve service designs, optimize deployments, and implement best practices for operational efficiency.
  • Guide technical decision-making and mentor junior SREs and developers across teams.
  • Participate in and lead postmortems, root cause analysis, and preventative design changes.
  • Contribute to capacity planning, demand forecasting, and long-term service scalability strategies.
  • Participate in a rotational on-call schedule to ensure the health and availability of production services.

Target Your Resume for "Principal Site Reliability Engineer" , Oracle

Get personalized recommendations to optimize your resume specifically for Principal Site Reliability Engineer. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Principal Site Reliability Engineer" , Oracle

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

PRODEV-ENGSVCSPRODEV-ENGSVCS

Answer 10 quick questions to check your fit for Principal Site Reliability Engineer @ Oracle.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.