Architect - System Performance Verification and Analysis
- AI
- Excel
- Fabric
- triage
- Python
- C++
- Bash
- Verilog
- SystemVerilog
- Nim
- OpenAI
- LangChain
NVIDIA has continuously reinvented itself. Our invention of the GPU sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. Today, research in artificial intelligence is booming worldwide, which calls for highly scalable and massively parallel computation horsepower that NVIDIA GPUs excel.
NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can address, and that matter to the world. This is our life’s work , to amplify human creativity and intelligence. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join our diverse team and see how you can make a lasting impact on the world. As part of this team, you would be working on projects that will help make our next generation visual computing, automotive, GPU, HPC  systems better. You will get to work on high performance CPU and Memory sub-systems, Next-Gen GPUs , NOC based Interconnect Fabric etc. Make the choice to join us today.Our work spans the full pre-silicon and post-silicon lifecycle: performance models, RTL simulation, emulation platforms, and silicon bringup. We catch bugs early, influence architecture and design decisions, and help ensure that the final product delivers the performance that our customers depend on. If you are passionate about building the hardware that underpins the future of AI and computing, this is where you belong.
What You'll Be Doing:
Partner with System Architecture, Usecase Modeling, Unit Architecture/PV, and Design/Verification teams to define comprehensive performance test plans that reflect real product usecases.
Drive full-chip SoC performance verification across all real engines and subsystems, running realistic concurrent workloads that represent how the product will be used by customers in the field.
Execute performance test plans across the full verification stack: cycle-accurate performance models, RTL simulation, emulation platforms, and silicon, identifying bottlenecks and regressions at each stage.
Debug performance failures through waveform analysis, signal-level queries, trace analysis, and system-level profiling to root-cause issues in the memory subsystem, fabric, or individual engines.
Influence architecture and microarchitecture decisions by surfacing performance data and trade-off analysis to the design and architecture teams early in the product development cycle.
Develop and maintain performance workloads, test suites, and infrastructure — including testbench components, performance simulators, analysis scripts, and automated regression flows.
Leverage AI-assisted tools and automation to accelerate repetitive analysis tasks, improve coverage, and free up engineering time for higher-level problem solving. This includes using LLM-based assistants for querying results and specs, AI-driven triage of regressions, and intelligent tooling that improves over time.
Drive methodology improvements to reduce verification turnaround time, improve coverage of representative workloads, and enable earlier performance insight in the product development cycle.
What We Need to See:
B.E./B.Tech or M.S./M.Tech (or equivalent experience) in Electrical Engineering, Computer Science, or a related field.
3+ years of relevant experience in SoC or system-level architecture, performance verification, or hardware validation.
Strong understanding of SoC architecture including GPU and CPU pipelines, memory subsystem design (caches, DRAM controllers, coherency), Network-on-Chip (NoC)/fabric architecture, and high-speed IO interfaces.
Hands-on experience with RTL simulation and debug, including waveform-based debug and signal-level querying to isolate performance failures.
Solid programming skills in Python and C/C++; scripting proficiency in Bash/Python for automation and analysis. Exposure to Verilog/SystemVerilog or SystemC/TLM is a strong plus.
Strong debugging, data analysis, and statistical analysis skills — ability to synthesize large volumes of performance data into actionable insights.
Experience with or exposure to pre-silicon performance analysis methodologies, including performance models, cycle-approximate simulators, or emulation platforms.
Excellent communication skills and the ability to work effectively in a large, globally distributed engineering organization.
Ways to Stand Out from the Crowd:
Deep experience with RTL-level performance debug — particularly the ability to formulate precise signal queries and interpret waveforms to root-cause complex system-level interactions.
Background in system-level performance analysis for GPU, AI accelerators, or high-bandwidth memory subsystems, with knowledge of bottleneck identification across multiple concurrent engines.
Demonstrated use of AI and LLM-based tools (e.g., NVIDIA NIM/NeMo, OpenAI APIs, LangChain, or similar agentic frameworks) to measurably improve your own engineering productivity — whether for automated analysis, natural-language querying of data, intelligent triage, or workflow automation.
Experience building or deploying ML/AI-assisted tooling in an engineering or EDA context (e.g., regression analysis, anomaly detection, test generation, coverage closure).
Expertise in data analysis and visualization — ability to build dashboards and tooling that surface performance trends clearly to both engineering and architecture stakeholders.
#LI-Hybrid
Architect - System Performance Verification and Analysis · Nvidia