MNC InsiderMNC Insider
Capital One logo

Sr. Manager SRE (Individual Contributor)

Capital One

Sr. Manager SRE (Individual Contributor)

full-timePosted: Jul 28, 2026Updated: Aug 27, 2026Mexico, Mexico City

Job Description

WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de MexicoSr. Manager SRE (Individual Contributor)We're building a Site Reliability Engineering center in Mexico City, and we're hiring a Senior Manager-level SRE to serve as the technical anchor for the site - defining the reliability vision, driving cross-team execution, and pioneering automation and AI-driven approaches that transform how we operate three payment networks at scale.This is a strategic technical leadership role. You won't manage people directly, but you'll shape how multiple teams work - setting architectural direction for observability, automation, and operational excellence, alert signal reduction, and reliability platform convergence. You'll be the most senior IC engineer in Mexico City, partnering with the Director (people leader) to translate organizational goals into technical roadmaps and ensuring the engineering quality bar stays high as the site scales.You'll operate across the full landscape: batch settlement systems processing every domestic and international credit/debit transaction, real-time observability platforms that must detect failures before customers do, and AI-powered automation that eliminates the toil standing between us and a proactive reliability culture.What You'll DoDefine and maintain a 12-18 month technical vision and roadmap for GPN SRE in Mexico City - decompose destination architecture into deliverable steps, sequence investments, and align execution across teamsDrive reliability transformation across settlement, observability, and automation domains - establish SLOs, error budgets, severity frameworks, and operational standards that teams build againstPioneer AI and agentic automation approaches - design and build AI-driven solutions (using Claude Code, Copilot CLI, and LLM frameworks) for alert classification, runbook generation, automated remediation, and incident analysis; set patterns that other engineers extendOwn the technical strategy for domain-specific knowledge ramp-up: identify which domain expertise requires deep engineering investment vs. documentation, and architect systems that reduce reliance on tribal knowledgeLead cross-team technical initiatives - drive observability platform convergence, standardize on COF tooling, and eliminate arbitrary uniqueness across towersServe as the senior escalation point for complex production incidents - diagnose cascading failures across distributed systems (storage, network, application), drive resolution, and ensure durable fixes landArchitect automation for high-risk operational processes - certificate rotation, compliance artifact generation, settlement cycle validation - ensuring security and reliability are built in from designMentor and elevate engineers across teams - conduct design reviews, establish engineering standards, coach on debugging and system thinking, and create an environment where Principal Associates and Managers grow into domain expertsIntroduce and advocate for engineering practices that raise the bar - AI engineering, innersourcing, reuse over rebuild, open source contribution, blameless postmortems, and chaos engineeringInfluence beyond the CDMX site - partner with US and UK leadership on architectural decisions, represent CDMX engineering in cross-org forums, and shape GPN-wide reliability strategyWhat Success Looks LikeTechnical roadmap established and executing - teams are delivering against a clear, sequenced plan with measurable reliability OKRsAt least one domain (alert signal reduction or settlement automation) where CDMX operates autonomously without US/UK escalation, driven by systems and patterns you architectedAI-powered automation deployed in production - incident classification models, generated runbooks, or automated remediation that demonstrably reduces MTTR or toilEngineering standards and patterns documented and adopted - design review process, observability standards, incident response framework, and automation patterns that scale with the teamRecognized as the technical authority for GPN SRE reliability - sought out across towers and geographies for architectural guidance, incident escalation, and strategic inputMultiple engineers grown through your mentorship - visible skill development in system design, debugging, and operational judgment across the CDMX teamsThe EnvironmentYou'll operate across hybrid on-prem and cloud infrastructure supporting real-time and batch financial transaction systems at global scale. The stack spans Python, Java, shell scripting, AWS, Kubernetes, OpenShift, CI/CD pipelines, and API automation frameworks. Observability runs on Datadog and Observe with complex dashboard configuration across three payment networks. Secret management and certificate automation use HashiCorp Vault. You'll design and build agentic AI automation solutions using Claude Code and LLM frameworks - this is central to the role, not an add-on. The systems span multiple on-prem data centers with mainframe, Linux, and containerized workloads alongside AWS. You'll need deep troubleshooting and debugging skills across all layers of the stack and the judgment to know when to go deep vs. when to delegate.Basic QualificationsProfessional English fluencyBachelor's degreeAt least 8+ years of experience in SRE, production operations, or reliability engineeringExperience in DevOps Engineering (internship experience does not apply)8+ years of experience in at least one of the following: Java, Python, GoAt least 6 years of experience with Cloud Native technologies (Amazon Web Services, Microsoft Azure, Google Cloud Platform)5+ years of experience with container orchestration services including Docker or KubernetesExperience with Shell or Bash scriptingAt least 5 years of Unix or Linux system administration experiencePreferred QualificationsExperience developing automation solutions using agentic AI tools (Claude Code, Copilot CLI)Troubleshooting and debugging skills across distributed systemsFamiliarity with payments, financial services, or other regulated high-availability domains Knowledge or experience of Networking concepts (TCP/DNS/TLS)At Capital One, we respect individual differences in culture, religion, and ethnicity. Likewise, we promote equal opportunities and development for all personnel. In the hiring process, we seek to provide equal employment opportunities to candidates, regardless of race, color, religion, gender, sexual orientation, marital or civil status, national origin, disability, or any other situation protected by federal, state, or local laws. For technical support or questions about Capital One's recruiting process, please send an email to Careers@capitalone.comCapital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site.Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe, any position posted in the Philippines is for Capital One Service Corp (COPSSC), and any position posted in Mexico is for Capital One Technology Labs Mexico.

Locations

  • Mexico, Mexico City

Skills Required

  • SREintermediate
  • DevOps Engineeringintermediate
  • at least one of the following: Javaintermediate
  • Cloud Native technologiesintermediate
  • container orchestration services including Dockerintermediate
  • Shellintermediate
  • paymentsintermediate

Required Qualifications

  • Professional English fluency (experience)
  • Bachelor's degree (degree)
  • At least 8+ years of experience in SRE, production operations, or reliability engineering (experience, 8 years)
  • Experience in DevOps Engineering (internship experience does not apply) (experience)
  • 8+ years of experience in at least one of the following: Java, Python, Go (experience, 8 years)
  • At least 6 years of experience with Cloud Native technologies (Amazon Web Services, Microsoft Azure, Google Cloud Platform) (experience, 6 years)
  • 5+ years of experience with container orchestration services including Docker or Kubernetes (experience, 5 years)
  • Experience with Shell or Bash scripting (experience)
  • At least 5 years of Unix or Linux system administration experience (experience, 5 years)
  • Professional English fluency (experience)
  • Bachelor's degree (degree)
  • At least 8+ years of experience in SRE, production operations, or reliability engineering (experience, 8 years)
  • Experience in DevOps Engineering (internship experience does not apply) (experience)
  • 8+ years of experience in at least one of the following: Java, Python, Go (experience, 8 years)
  • At least 6 years of experience with Cloud Native technologies (Amazon Web Services, Microsoft Azure, Google Cloud Platform) (experience, 6 years)
  • 5+ years of experience with container orchestration services including Docker or Kubernetes (experience, 5 years)
  • Experience with Shell or Bash scripting (experience)
  • At least 5 years of Unix or Linux system administration experience (experience, 5 years)

Preferred Qualifications

  • Experience developing automation solutions using agentic AI tools (Claude Code, Copilot CLI) (experience)
  • Troubleshooting and debugging skills across distributed systems (experience)
  • Familiarity with payments, financial services, or other regulated high-availability domains (experience)
  • Knowledge or experience of Networking concepts (TCP/DNS/TLS) (experience)
  • Experience developing automation solutions using agentic AI tools (Claude Code, Copilot CLI) (experience)
  • Troubleshooting and debugging skills across distributed systems (experience)
  • Familiarity with payments, financial services, or other regulated high-availability domains (experience)
  • Knowledge or experience of Networking concepts (TCP/DNS/TLS) (experience)
  • For technical support or questions about Capital One's recruiting process, please send an email to Careers@capitalone.com (experience)
  • Capital One does not provide, endorse nor guarantee and is not liable for third-party products, services, educational tools or other information available through this site. (experience)
  • Capital One Financial is made up of several different entities. Please note that any position posted in Canada is for Capital One Canada, any position posted in the United Kingdom is for Capital One Europe, any position posted in the Philippines is for Capital One Service Corp (COPSSC), and any position posted in Mexico is for Capital One Technology Labs Mexico. (experience)

Responsibilities

  • Define and maintain a 12-18 month technical vision and roadmap for GPN SRE in Mexico City - decompose destination architecture into deliverable steps, sequence investments, and align execution across teams
  • Drive reliability transformation across settlement, observability, and automation domains - establish SLOs, error budgets, severity frameworks, and operational standards that teams build against
  • Pioneer AI and agentic automation approaches - design and build AI-driven solutions (using Claude Code, Copilot CLI, and LLM frameworks) for alert classification, runbook generation, automated remediation, and incident analysis; set patterns that other engineers extend
  • Own the technical strategy for domain-specific knowledge ramp-up: identify which domain expertise requires deep engineering investment vs. documentation, and architect systems that reduce reliance on tribal knowledge
  • Lead cross-team technical initiatives - drive observability platform convergence, standardize on COF tooling, and eliminate arbitrary uniqueness across towers
  • Serve as the senior escalation point for complex production incidents - diagnose cascading failures across distributed systems (storage, network, application), drive resolution, and ensure durable fixes land
  • Architect automation for high-risk operational processes - certificate rotation, compliance artifact generation, settlement cycle validation - ensuring security and reliability are built in from design
  • Mentor and elevate engineers across teams - conduct design reviews, establish engineering standards, coach on debugging and system thinking, and create an environment where Principal Associates and Managers grow into domain experts
  • Introduce and advocate for engineering practices that raise the bar - AI engineering, innersourcing, reuse over rebuild, open source contribution, blameless postmortems, and chaos engineering
  • Influence beyond the CDMX site - partner with US and UK leadership on architectural decisions, represent CDMX engineering in cross-org forums, and shape GPN-wide reliability strategy
  • Define and maintain a 12-18 month technical vision and roadmap for GPN SRE in Mexico City - decompose destination architecture into deliverable steps, sequence investments, and align execution across teams
  • Drive reliability transformation across settlement, observability, and automation domains - establish SLOs, error budgets, severity frameworks, and operational standards that teams build against
  • Pioneer AI and agentic automation approaches - design and build AI-driven solutions (using Claude Code, Copilot CLI, and LLM frameworks) for alert classification, runbook generation, automated remediation, and incident analysis; set patterns that other engineers extend
  • Own the technical strategy for domain-specific knowledge ramp-up: identify which domain expertise requires deep engineering investment vs. documentation, and architect systems that reduce reliance on tribal knowledge
  • Lead cross-team technical initiatives - drive observability platform convergence, standardize on COF tooling, and eliminate arbitrary uniqueness across towers
  • Serve as the senior escalation point for complex production incidents - diagnose cascading failures across distributed systems (storage, network, application), drive resolution, and ensure durable fixes land
  • Architect automation for high-risk operational processes - certificate rotation, compliance artifact generation, settlement cycle validation - ensuring security and reliability are built in from design
  • Mentor and elevate engineers across teams - conduct design reviews, establish engineering standards, coach on debugging and system thinking, and create an environment where Principal Associates and Managers grow into domain experts
  • Introduce and advocate for engineering practices that raise the bar - AI engineering, innersourcing, reuse over rebuild, open source contribution, blameless postmortems, and chaos engineering
  • Influence beyond the CDMX site - partner with US and UK leadership on architectural decisions, represent CDMX engineering in cross-org forums, and shape GPN-wide reliability strategy
  • What Success Looks Like
  • Technical roadmap established and executing - teams are delivering against a clear, sequenced plan with measurable reliability OKRs
  • At least one domain (alert signal reduction or settlement automation) where CDMX operates autonomously without US/UK escalation, driven by systems and patterns you architected
  • AI-powered automation deployed in production - incident classification models, generated runbooks, or automated remediation that demonstrably reduces MTTR or toil
  • Engineering standards and patterns documented and adopted - design review process, observability standards, incident response framework, and automation patterns that scale with the team
  • Recognized as the technical authority for GPN SRE reliability - sought out across towers and geographies for architectural guidance, incident escalation, and strategic input
  • Multiple engineers grown through your mentorship - visible skill development in system design, debugging, and operational judgment across the CDMX teams
  • Technical roadmap established and executing - teams are delivering against a clear, sequenced plan with measurable reliability OKRs
  • At least one domain (alert signal reduction or settlement automation) where CDMX operates autonomously without US/UK escalation, driven by systems and patterns you architected
  • AI-powered automation deployed in production - incident classification models, generated runbooks, or automated remediation that demonstrably reduces MTTR or toil
  • Engineering standards and patterns documented and adopted - design review process, observability standards, incident response framework, and automation patterns that scale with the team
  • Recognized as the technical authority for GPN SRE reliability - sought out across towers and geographies for architectural guidance, incident escalation, and strategic input
  • Multiple engineers grown through your mentorship - visible skill development in system design, debugging, and operational judgment across the CDMX teams

Target Your Resume for "Sr. Manager SRE (Individual Contributor)" , Capital One

Get personalized recommendations to optimize your resume specifically for Sr. Manager SRE (Individual Contributor). Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Sr. Manager SRE (Individual Contributor)" , Capital One

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Sr. Manager SRE (Individual Contributor) @ Capital One.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.