OpenAI

OpenAI

Software Engineer, Trainium

San Francisco, CAPosted 6 days ago$295,000 – $380,000
Full TimeSeniorUS

See how this job matches your profile

Sign in for an AI-powered fit score, breakdown, and a tailored resume.

Sign in

Job Description

The team builds the performance-critical systems that allow OpenAI's models to run efficiently across a diverse set of AI accelerators. We work across the inference stack, from low-level kernels and compilers through model execution, to unlock the full capabilities of the underlying hardware. As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass. You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack. If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems.

Qualifications

Required Qualifications

  • 3+ years of relevant engineering experience, ideally in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Strong systems programming fundamentals and experience writing performance-critical software.
  • Experience working with GPU, TPU, Trainium, or other specialized accelerator architectures.
  • Ability to reason about performance across multiple layers of the stack, from hardware and kernels through compilers and ML frameworks.
  • Owning technically ambiguous problems end-to-end and learning new hardware and software domains as needed.

Preferred Qualifications

  • Experience with AWS Trainium or the AWS Neuron SDK.
  • Contributions to ML frameworks such as PyTorch or JAX, compiler infrastructure such as LLVM, MLIR, XLA, or Triton, or experience developing kernels for specialized accelerators.

Skills & Technologies

AWSPyTorch

Interested in this role?

Sign in or create a free account to see how this job matches your skills, apply with one click, and let our AI tailor your resume.

Sign in to apply
AI-powered resume optimization
Save and track your applications

Job Details

Employment Type

Full Time

Experience Level

Senior

Salary Range

$295,000 – $380,000

Location

San Francisco, CA

Posted

6 days ago

Country

US