
OpenAI
Software Engineer, Trainium
See how this job matches your profile
Sign in for an AI-powered fit score, breakdown, and a tailored resume.
Job Description
The team builds the performance-critical systems that allow OpenAI's models to run efficiently across a diverse set of AI accelerators. We work across the inference stack, from low-level kernels and compilers through model execution, to unlock the full capabilities of the underlying hardware. As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass. You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack. If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems.
Qualifications
Required Qualifications
- 3+ years of relevant engineering experience, ideally in ML systems, compilers, kernels, runtimes, or performance engineering.
- Strong systems programming fundamentals and experience writing performance-critical software.
- Experience working with GPU, TPU, Trainium, or other specialized accelerator architectures.
- Ability to reason about performance across multiple layers of the stack, from hardware and kernels through compilers and ML frameworks.
- Owning technically ambiguous problems end-to-end and learning new hardware and software domains as needed.
Preferred Qualifications
- Experience with AWS Trainium or the AWS Neuron SDK.
- Contributions to ML frameworks such as PyTorch or JAX, compiler infrastructure such as LLVM, MLIR, XLA, or Triton, or experience developing kernels for specialized accelerators.
Skills & Technologies
About the Company
OpenAI
View company profile →
Interested in this role?
Sign in or create a free account to see how this job matches your skills, apply with one click, and let our AI tailor your resume.
Sign in to applyJob Details
Employment Type
Full Time
Experience Level
Senior
Salary Range
$295,000 – $380,000
Location
San Francisco, CA
Posted
6 days ago
Country
US