About the role#
NVIDIA is looking for a PhD intern to join the Deep Learning Efficiency Research team. This role focuses on two main areas: efficient diffusion language models and multimodal generative models, and efficient agentic AI with hybrid inference orchestration. You will work on research that bridges the gap between theoretical exploration and real-world application.
What you'll do#
- Research, design, and implement new methods for efficient deep learning.
- Work on diffusion LLMs and multimodal models, including sampling efficiency, parallel decoding, and training pipelines.
- Develop efficient agentic AI solutions, such as hybrid inference orchestration across cloud and edge, plus routing and scheduling policies.
- Publish original research in top-tier venues.
- Collaborate with internal teams, product groups, and external researchers to transfer technology.
What you'll need#
- Current enrollment in a PhD program for Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Strong knowledge of machine learning and deep learning theory and practice.
- Experience with large language models, diffusion models, multimodal models, or agentic systems.
- Hands-on experience with large-scale model training, including data preparation and model parallelization techniques like tensor and pipeline parallelism.
- A research track record that includes at least one publication at a top-tier conference such as ICML, ICLR, NeurIPS, CVPR, or ICCV.
- Excellent communication skills.
- Bonus points for experience with CUDA, hybrid cloud-edge inference, or model optimization techniques like pruning, quantization, and NAS.
Location & details#
- This is a full-time, paid internship.
- The role is based in Santa Clara, California, with both on-site and remote work options available.
- Hourly pay ranges from 38 USD to 94 USD, depending on location, degree, and experience.
- Sponsorship is not provided for this position.
About NVIDIA
NVIDIA operates as a computer hardware manufacturer based in Santa Clara, California. Founded in 1993, the company focuses on accelerated computing and graphics technology. It produces hardware for markets including artificial intelligence, gaming, and data centers. The organization employs over 50,000 people and maintains a global presence.
How to get in at NVIDIA
Applying early gives you a distinct advantage at NVIDIA because recruiters review applications as they arrive. Intern Insider sends an instant alert the moment a role matching your target is published, so you can apply among the first before the pile grows. Getting your resume in front of a human early is often the difference between a screening call and a rejection. You can use Intern Insider to surface the recruiters behind NVIDIA roles to reach out directly. Asking a recruiter about the role or a referral materially improves your response rates compared to submitting into a general queue. It is a simple way to make your application stand out in a competitive process.



