About the role#
This internship focuses on the performance engineering of large-scale AI systems. You will work alongside researchers and engineers to optimize the execution of foundation models on next-generation NVIDIA GPU architectures. The role involves low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
What you'll do#
- Develop analytical performance models for GPU kernels and inference workloads.
- Build and validate a simulator to estimate theoretical hardware performance limits.
- Identify performance bottlenecks in compute, memory, communication, and scheduling.
- Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
- Investigate PTX and SASS code generation to understand low-level execution behavior.
- Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
- Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
- Design profiling methodologies for Hopper and Blackwell architectures.
- Document findings and provide actionable recommendations for performance improvements.
What you'll need#
- Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
- Experience with CUDA programming and GPU kernel development.
- Understanding of NVIDIA GPU architecture and memory hierarchy.
- Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
- Experience with deep learning frameworks such as PyTorch or TensorFlow.
- Knowledge of PTX, SASS, and low-level GPU execution.
- Experience optimizing CUDA kernels for throughput and latency.
- Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
- Strong programming skills in C++, CUDA, and Python.
Location & details#
- Location: Sunnyvale, California, United States.
- Term: Fall 2026.
- Modality: On-site.
- Commitment: Full-time internship.
About MBZUAI (Mohamed bin Zayed University of Artificial Intelligence)
Mohamed bin Zayed University of Artificial Intelligence is a graduate-level research institution located in Abu Dhabi. Founded in 2019, the university focuses on the study and development of artificial intelligence. It operates as an educational organization with a staff of over 1,000 people. The institution maintains its primary campus in Masdar City.
How to get in at MBZUAI (Mohamed bin Zayed University of Artificial Intelligence)
Securing a position at MBZUAI often requires staying ahead of the standard application cycle. Intern Insider sends an instant alert the moment a role matching your target is published, allowing you to apply among the first before the pile grows. Applying early is a practical advantage because recruiters often review the initial wave of candidates with more attention. You can use this lead time to ensure your materials are polished and ready for submission. Reaching out to the right person is just as important as the timing of your application. Intern Insider surfaces the recruiters behind the company roles so you can reach out directly to ask about the position or a referral. Connecting with a human contact materially improves your response rates compared to sending documents into a general queue. This approach shows you have done your research and are serious about joining the university research community.


