MNC InsiderMNC Insider
NVIDIA logo

Deep Learning Kernel Software Performance Architect

NVIDIA

Deep Learning Kernel Software Performance Architect

full-timePosted: Jul 28, 2026Updated: Aug 27, 2026Shanghai, China

Job Description

NVIDIA is seeking Software Performance Architects to optimize GPU kernel performance for state-of-the-art data-center platforms. We build automated, data-driven workflows to detect, explain, and prevent performance regressions across key deep learning workloads, partnering closely with kernel developers, compiler teams, infrastructure, and architecture/performance groups.What you'll be doing:Performance analysis, optimization and debuggingBuild performance narratives using structured methodology: baselines, projections, controlled comparisons, and regression attribution.With the methodologies, analyze performance of GPU-accelerated kernels and key deep learning building blocks, identify gaps with baselines or projections, then optimize the kernels' performance to fill the gaps.Debug performance issues end-to-end: reproduce, isolate root causes, propose fixes or mitigation paths, and drive closure with the owning teams. Automation + regression infrastructure (Python-heavy)Develop and maintain Python-based automation for performance testing and analysis—using modern AI-assisted developer tools (e.g., Cursor/Claude Code/Copilot) to accelerate scripting while keeping code maintainable and reviewable.Design and operate performance test workflows: coverage definition, test/workload generation, automated large-scale execution (CI/nightly/on-demand), rerun rules, and reproducibility standards. Cross-team collaboration and operating modelWork with kernel developers and the compiler teams to ensure performance checks are practical, scalable, and aligned to release needs.Work with chip architecture and modeling teams to solidify the performance methodology across chip architecture generations and common Deep Learning operators such as GEMM, Attention, MoE.Partner with SWQA and infrastructure teams for execution at scale and reliable pipelines/dashboards.Following general software engineering best practices including support for regression testing and CI/CD flowsWhat we need to see:Masters or PhD degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or related fieldStrong programming ability in Python plus C/C++ with 2+ working experience (performance-oriented code reading/debugging)Solid fundamentals in computer architecture, parallel programming and performance reasoning (latency/throughput, memory hierarchy, parallelism) to be able to identify bottlenecks, optimize resource utilization, and improve throughputExperience with performance analysis workflows: profiling, measurement methodology, reproducibility, and regression triage.Comfortable working across teams and driving issues to decision/closure with clear communicationWays to stand out from the crowd:Experience with high-performance kernels or math libraries (e.g., GEMM/attention, CUTLASS-like concepts)GPU programming/perf experience (CUDA or equivalent parallel programming)Strong ML/DL workload understanding (training/inference shapes, precision modes, perf bottlenecks)Familiarity with simulators/analytical modeling or performance characterization methodology

Locations

  • Shanghai, China
  • Beijing, China

Responsibilities

  • Performance analysis, optimization and debuggingBuild performance narratives using structured methodology: baselines, projections, controlled comparisons, and regression attribution.With the methodologies, analyze performance of GPU-accelerated kernels and key deep learning building blocks, identify gaps with baselines or projections, then optimize the kernels' performance to fill the gaps.Debug performance issues end-to-end: reproduce, isolate root causes, propose fixes or mitigation paths, and drive closure with the owning teams.
  • Automation + regression infrastructure (Python-heavy)Develop and maintain Python-based automation for performance testing and analysis—using modern AI-assisted developer tools (e.g., Cursor/Claude Code/Copilot) to accelerate scripting while keeping code maintainable and reviewable.Design and operate performance test workflows: coverage definition, test/workload generation, automated large-scale execution (CI/nightly/on-demand), rerun rules, and reproducibility standards.
  • Cross-team collaboration and operating modelWork with kernel developers and the compiler teams to ensure performance checks are practical, scalable, and aligned to release needs.Work with chip architecture and modeling teams to solidify the performance methodology across chip architecture generations and common Deep Learning operators such as GEMM, Attention, MoE.Partner with SWQA and infrastructure teams for execution at scale and reliable pipelines/dashboards.
  • Following general software engineering best practices including support for regression testing and CI/CD flows
  • Performance analysis, optimization and debugging
  • Build performance narratives using structured methodology: baselines, projections, controlled comparisons, and regression attribution.
  • With the methodologies, analyze performance of GPU-accelerated kernels and key deep learning building blocks, identify gaps with baselines or projections, then optimize the kernels' performance to fill the gaps.
  • Debug performance issues end-to-end: reproduce, isolate root causes, propose fixes or mitigation paths, and drive closure with the owning teams.
  • Automation + regression infrastructure (Python-heavy)
  • Develop and maintain Python-based automation for performance testing and analysis—using modern AI-assisted developer tools (e.g., Cursor/Claude Code/Copilot) to accelerate scripting while keeping code maintainable and reviewable.
  • Design and operate performance test workflows: coverage definition, test/workload generation, automated large-scale execution (CI/nightly/on-demand), rerun rules, and reproducibility standards.
  • Cross-team collaboration and operating model
  • Work with kernel developers and the compiler teams to ensure performance checks are practical, scalable, and aligned to release needs.
  • Work with chip architecture and modeling teams to solidify the performance methodology across chip architecture generations and common Deep Learning operators such as GEMM, Attention, MoE.
  • Partner with SWQA and infrastructure teams for execution at scale and reliable pipelines/dashboards.
  • Following general software engineering best practices including support for regression testing and CI/CD flows

Target Your Resume for "Deep Learning Kernel Software Performance Architect" , NVIDIA

Get personalized recommendations to optimize your resume specifically for Deep Learning Kernel Software Performance Architect. Takes only 15 seconds!

AI-powered keyword optimization
Skills matching & gap analysis
Experience alignment suggestions

Check Your ATS Score for "Deep Learning Kernel Software Performance Architect" , NVIDIA

Find out how well your resume matches this job's requirements. Get comprehensive analysis including ATS compatibility, keyword matching, skill gaps, and personalized recommendations.

ATS compatibility check
Keyword optimization analysis
Skill matching & gap identification
Format & readability score

Tags & Categories

GeneralGeneral

Answer 10 quick questions to check your fit for Deep Learning Kernel Software Performance Architect @ NVIDIA.

Quiz Challenge
10 Questions
~2 Minutes
Instant Score

Related Books and Jobs

No related jobs found at the moment.