MNC InsiderMNC Insider

Senior Site Reliability & Observability Engineer (SRE)

TransUnion

Senior Site Reliability & Observability Engineer (SRE)

full-timePosted: Aug 4, 2026Updated: Sep 3, 2026Eagle - D.F. Del. Miguel Hidalgo

Job Description

TransUnion's Job Applicant Privacy NoticeTeam OverviewThe Reliability Engineering team ensures the stability, availability, and performance of Buró de Crédito’s mission-critical platforms and services. Through observability, automation, and incident management practices, the team drives operational excellence and continuous improvement. Working closely with Infrastructure, Development, Database, and Security teams, they help maintain resilient systems that support critical business operations. This job is assigned as On-Site Essential and requires in- person work at an assigned TU office location as a condition of employment.Role Overview And Core ResponsibilitiesEnsure the reliability, stability, and availability of mission-critical services through Site Reliability Engineering (SRE) best practices.Operate and continuously improve the organization's observability platform, including metrics, logs, traces, and alerting capabilities.Monitor critical systems proactively and respond to operational incidents to minimize service disruption and business impact.Lead major incident response activities, coordinating recovery efforts and stakeholder communication during service outages.Conduct root cause analysis (RCA) and facilitate blameless post-mortems to identify systemic improvements and prevent recurrence.Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.Automate operational processes and reduce manual effort (toil) through scripting, infrastructure automation, and platform engineering practices.Support and troubleshoot distributed environments including Apache Cassandra, Kafka, Kubernetes, relational databases, and NoSQL platforms.Collaborate with development and infrastructure teams to improve deployment reliability, operational readiness, and platform performance.Drive continuous improvement initiatives focused on monitoring, availability, scalability, resilience, and operational efficiency.Required Knowledge And Experiences​​​Bachelor's degree in Computer Science, Systems Engineering, Telecommunications Engineering, or a related technical discipline, providing the foundation required to manage complex distributed environments.Proven experience in Site Reliability Engineering (SRE), Production Operations, Platform Engineering, Infrastructure Engineering, or Reliability-focused roles.Strong knowledge of distributed systems, observability practices, incident management, and troubleshooting methodologies for mission-critical environments.Experience managing major incidents, root cause analysis processes, service restoration activities, and operational excellence initiatives.Understanding of reliability frameworks including SLOs, SLIs, error budgets, continuous improvement, and service management best practices.​#LI-SG4TransUnion Overview:At TransUnion, we encourage and are committed to creating a real, positive impact and shared sense of purpose within our Workforce for Good, which empowers our people to grow, innovate and contribute to a better future for our communities and customers. We strive to build an environment where our associates are in the driver’s seat of their professional development— while having access to help along the way. We recognize that success comes when our associates thrive both professionally and personally; that’s why we prioritize work/life flexibility and offer resources for our teams across the globe to collaborate and drive excellence. Be a part of our Workforce for Good – you’ll work with great people, pioneering products and cutting-edge technology.TransUnion Job TitleConsultant, IT Support

Locations

  • Eagle - D.F. Del. Miguel Hidalgo

Skills Required

  • Site Reliability Engineeringintermediate
  • distributed systemsintermediate

Required Qualifications

  • ​​​Bachelor's degree in Computer Science, Systems Engineering, Telecommunications Engineering, or a related technical discipline, providing the foundation required to manage complex distributed environments. (degree in computer science)
  • Proven experience in Site Reliability Engineering (SRE), Production Operations, Platform Engineering, Infrastructure Engineering, or Reliability-focused roles. (experience)
  • Strong knowledge of distributed systems, observability practices, incident management, and troubleshooting methodologies for mission-critical environments. (experience)
  • Experience managing major incidents, root cause analysis processes, service restoration activities, and operational excellence initiatives. (experience)
  • Understanding of reliability frameworks including SLOs, SLIs, error budgets, continuous improvement, and service management best practices.​ (experience)
  • ​​​Bachelor's degree in Computer Science, Systems Engineering, Telecommunications Engineering, or a related technical discipline, providing the foundation required to manage complex distributed environments. (degree in computer science)
  • Proven experience in Site Reliability Engineering (SRE), Production Operations, Platform Engineering, Infrastructure Engineering, or Reliability-focused roles. (experience)
  • Strong knowledge of distributed systems, observability practices, incident management, and troubleshooting methodologies for mission-critical environments. (experience)
  • Experience managing major incidents, root cause analysis processes, service restoration activities, and operational excellence initiatives. (experience)
  • Understanding of reliability frameworks including SLOs, SLIs, error budgets, continuous improvement, and service management best practices. (experience)

Responsibilities

  • Ensure the reliability, stability, and availability of mission-critical services through Site Reliability Engineering (SRE) best practices.
  • Operate and continuously improve the organization's observability platform, including metrics, logs, traces, and alerting capabilities.
  • Monitor critical systems proactively and respond to operational incidents to minimize service disruption and business impact.
  • Lead major incident response activities, coordinating recovery efforts and stakeholder communication during service outages.
  • Conduct root cause analysis (RCA) and facilitate blameless post-mortems to identify systemic improvements and prevent recurrence.
  • Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Automate operational processes and reduce manual effort (toil) through scripting, infrastructure automation, and platform engineering practices.
  • Support and troubleshoot distributed environments including Apache Cassandra, Kafka, Kubernetes, relational databases, and NoSQL platforms.
  • Collaborate with development and infrastructure teams to improve deployment reliability, operational readiness, and platform performance.
  • Drive continuous improvement initiatives focused on monitoring, availability, scalability, resilience, and operational efficiency.
  • Ensure the reliability, stability, and availability of mission-critical services through Site Reliability Engineering (SRE) best practices.
  • Operate and continuously improve the organization's observability platform, including metrics, logs, traces, and alerting capabilities.
  • Monitor critical systems proactively and respond to operational incidents to minimize service disruption and business impact.
  • Lead major incident response activities, coordinating recovery efforts and stakeholder communication during service outages.
  • Conduct root cause analysis (RCA) and facilitate blameless post-mortems to identify systemic improvements and prevent recurrence.
  • Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Automate operational processes and reduce manual effort (toil) through scripting, infrastructure automation, and platform engineering practices.
  • Support and troubleshoot distributed environments including Apache Cassandra, Kafka, Kubernetes, relational databases, and NoSQL platforms.
  • Collaborate with development and infrastructure teams to improve deployment reliability, operational readiness, and platform performance.
  • Drive continuous improvement initiatives focused on monitoring, availability, scalability, resilience, and operational efficiency.

Target Your Resume for "Senior Site Reliability & Observability Engineer (SRE)" , TransUnion

Get personalized recommendations to optimize your resume specifically for Senior Site Reliability & Observability Engineer (SRE). Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Senior Site Reliability & Observability Engineer (SRE)" , TransUnion

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Senior Site Reliability & Observability Engineer (SRE) @ TransUnion.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.