MNC InsiderMNC Insider

Senior Staff Platform Engineer

Zscaler

Senior Staff Platform Engineer

full-timePosted: Aug 4, 2026Updated: Sep 3, 2026Bangalore, IND

Job Description

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler.Role We are looking for a Sr. Staff AI Platform Engineer to join our team. This is an On-site role based in Bangalore/Pune, reporting to the Senior Manager in the IT Data Strategy department. In this position, you will design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads. Operating within the IT Data Strategy team, you will drive infrastructure automation, observability, and platform governance while empowering AI and data engineering teams to deliver high-impact solutions reliably and securely. What you’ll do (Role Expectations) Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS) Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews Who You Are (Success Profile) You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution. You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust. You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges. You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication. You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability. What We’re Looking for (Minimum Qualifications) Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3) Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer What Will Make You Stand Out (Preferred Qualifications) Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments #LI-RR1 #LI-HybridAt Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In-office perks, and more! Learn more about Zscaler's hybrid working model and benefits here. By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Pay Transparency Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy-related support.

Locations

  • Bangalore, IND

Skills Required

  • AWSintermediate
  • Infrastructure as Code using Terraformintermediate
  • centralized observability using Prometheusintermediate
  • LiteLLMintermediate
  • vector database platformsintermediate

Required Qualifications

  • You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution. (experience)
  • You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust. (experience)
  • You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges. (experience)
  • You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication. (experience)
  • You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability. (experience)
  • You act like an owner with a passion for the mission, operating with integrity and navigating seamlessly between high-level platform strategy and hands-on execution. (experience)
  • You are a high-trust collaborator who is ambitious for the overall team, fostering an open feedback culture delivered with clarity and respect to build lasting trust. (experience)
  • You are driven by innovation and deep technical curiosity, continuously seeking secure, scalable, and modern solutions to complex platform engineering challenges. (experience)
  • You champion simplicity by distilling complex technical architecture, user needs, and operational concepts into clear, actionable plans and focused communication. (experience)
  • You are data-driven, leveraging analytics and measurable metrics to guide informed engineering decisions, evaluate truth, and optimize reliability. (experience)
  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving (experience)
  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3) (experience, 8 years)
  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation (experience)
  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services (experience)
  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement (experience)
  • Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer (experience)
  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving (experience)
  • 8+ years of experience as a Platform Engineer / Site Reliability Engineer / DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure, with deep hands-on expertise in AWS (EKS, Lambda, ECS, VPC, IAM, S3) (experience, 8 years)
  • Strong hands-on experience with Infrastructure as Code using Terraform (modules, workspaces, remote state) and designing/maintaining CI/CD pipelines with GitLab CI/CD (or equivalent), including runners and deployment automation (experience)
  • Proven expertise in centralized observability using Prometheus and Grafana (custom exporters, dashboard design, alerting rules), along with experience designing and tuning alerting/on-call systems (Alertmanager, PagerDuty, OpsGenie) and owning incident response for production services (experience)
  • Solid, hands-on understanding of DevOps/DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR) and how to use them to drive engineering improvement (experience)
  • Experience deploying and operating containerized applications on Kubernetes (EKS), including Helm charts and autoscaling (HPA, Cluster Autoscaler, Karpenter), combined with strong scripting skills in Python (and/or Go/Bash) and the ability to work independently and lead complex infrastructure initiatives with excellent communication skills befitting a Staff-level engineer (experience)

Preferred Qualifications

  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex (experience)
  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix (experience)
  • Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments (experience)
  • Experience with LiteLLM (including multi-geo/multi-region deployment patterns), distributed compute frameworks like Ray.io/Anyscale, workflow engines like Temporal, and agentic frameworks such as LangGraph, LangChain, or LlamaIndex (experience)
  • Working knowledge of vector database platforms (Qdrant, Pinecone, Weaviate), memory/context layers like ZEP, LLMOps/MLOps frameworks and practices (model registry, prompt versioning, feedback loops, MLflow/Kubeflow), and LLM observability/evaluation tooling such as Arize Phoenix (experience)
  • Exposure to GitOps workflows (ArgoCD/Flux), log aggregation and distributed tracing (Loki, ELK/OpenSearch, Jaeger/Tempo), cloud cost optimization (FinOps), and security/compliance frameworks (SOC 2, ISO 27001) in cloud-native environments (experience)
  • At Zscaler, we are committed to building a team that reflects the communities we serve and the customers we work with. We foster an inclusive environment that values all backgrounds and perspectives, emphasizing collaboration and belonging. Join us in our mission to make doing business seamless and secure. (experience)
  • Our Benefits program is one of the most important ways we support our employees. Zscaler proudly offers comprehensive and inclusive benefits to meet the diverse needs of our employees and their families throughout their life stages, including: (experience)
  • Various health plans (experience)
  • Time off plans for vacation and sick time (experience)
  • Parental leave options (experience)
  • Retirement options (experience)
  • Education reimbursement (experience)
  • In-office perks, and more! (experience)
  • Learn more about Zscaler's hybrid working model and benefits here. (experience)
  • By applying for this role, you adhere to applicable laws, regulations, and Zscaler policies, including those related to security and privacy standards and guidelines. (experience)
  • Pay Transparency (experience)
  • Zscaler complies with all applicable federal, state, and local pay transparency rules. (experience)

Responsibilities

  • Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
  • Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
  • Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
  • Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
  • Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews
  • Design, build, and maintain scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) for AI/ML workloads using Terraform, following IaC best practices including reusable modules, remote state management, and environment-based blueprints
  • Own and continuously evolve GitLab CI/CD pipelines for AI platform services, automating build, test, security scanning, and multi-environment deployment workflows to enable fast, reliable, and repeatable releases
  • Architect a centralized observability stack using Prometheus and Grafana with golden-signal dashboards across AI/ML services and infrastructure, design intelligent alerting strategies (Alertmanager, PagerDuty, OpsGenie), and lead incident response, root-cause analysis, and postmortems to improve MTTA/MTTR
  • Define and track DORA metrics to drive delivery and reliability improvements, while building self-healing, auto-scaling, and cost-optimized infrastructure for AI/ML services (LLM inference, vector databases, agent frameworks) on Kubernetes (EKS)
  • Implement infrastructure security best practices, establish platform governance standards, partner with AI/ML and data engineering teams on production-grade AI/RAG deployments, and mentor junior/mid-level engineers through design and code reviews

Target Your Resume for "Senior Staff Platform Engineer" , Zscaler

Get personalized recommendations to optimize your resume specifically for Senior Staff Platform Engineer. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Senior Staff Platform Engineer" , Zscaler

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

IT Data StrategyIT Data Strategy

Answer 10 quick questions to check your fit for Senior Staff Platform Engineer @ Zscaler.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.