About the role#
The Inference Research team builds efficient, scalable, and reliable serving systems for large foundation models. We work at the intersection of model architectures, systems engineering, and hardware optimization. Our goal is to co-design software, algorithms, and models to lower the cost and latency of AI systems. As an intern, you will work on distributed inference, compiler-aware optimization, and inference-time computation strategies like speculative decoding. You will focus on cross-layer optimizations, including KV cache design and large-scale serving architectures.
What you'll do#
- Design and conduct rigorous experiments to validate hypotheses.
- Co-design and implement cross-layer optimizations across models, systems, and hardware.
- Communicate project plans, progress, and results to the broader team.
- Document your findings in scientific publications and blog posts.
What you'll need#
- Currently pursuing a final year of a Bachelor's, Master's, or Ph.D. degree in Computer Science, Electrical Engineering, or a related field.
- Strong knowledge of Machine Learning and Deep Learning fundamentals.
- Experience with deep learning frameworks such as PyTorch or JAX.
- Strong programming skills in Python.
- Familiarity with Transformer architectures and recent developments in foundation models.
- Prior research experience in foundation models, efficient machine learning, or ML systems is preferred.
- Publications at conferences like MLSys or ICLR are preferred.
- Experience with CUDA programming for kernel development is preferred.
- Understanding of model optimization techniques and hardware acceleration is preferred.
- Contributions to open-source machine learning projects are preferred.
Location & details#
- Location: San Francisco, California.
- Term: Winter 2027 (January 4th to April 9th).
- Modality: On-site.
- Employment: Full-time internship.
- Compensation: Paid position.
About Together AI
Together AI operates a cloud platform for AI engineers and researchers. The company provides tools for inference, model shaping, and pre-training. Founded in 2022, it maintains its headquarters in San Francisco, California. The organization employs 415 people and serves clients ranging from AI startups to SaaS companies.
How to get in at Together AI
Securing an internship at Together AI requires speed, as early applicants are often reviewed before the candidate pool becomes unmanageable. Intern Insider sends an instant alert the moment a role matching your target is published anywhere, ensuring you apply among the first. This helps you avoid the frustration of submitting an application to a queue that has already grown too large to navigate. You can use this timing to your advantage to stay ahead of the competition. Beyond timing, reaching out to the right person often yields better results than a standard submission. Intern Insider surfaces the recruiters behind the roles at Together AI so you can contact them directly to ask about the position or a referral. This approach materially improves your response rates compared to waiting for an automated update. It is a more direct way to show your interest to the people actually managing the hiring process.



