Beam

Beam

Site Reliability Engineer

San Francisco, California, US / New York, New York, USRemotePosted 13 days ago$140,000 – $200,000
Full TimeRemoteUS

See how this job matches your profile

Sign in for an AI-powered fit score, breakdown, and a tailored resume.

Sign in

Job Description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our plat

Key Highlights

  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
  • Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
  • Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
  • Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.
  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it

Interested in this role?

Sign in or create a free account to see how this job matches your skills, apply with one click, and let our AI tailor your resume.

Sign in to apply
AI-powered resume optimization
Save and track your applications

Job Details

Employment Type

Full Time

Salary Range

$140,000 – $200,000

Location

San Francisco, California, US / New York, New York, US

Work Mode

Remote

Posted

13 days ago

Country

US