Company
OpenAI
Location
London, UK
Employment Type
Full time
Location Type
Hybrid
Department
Scaling
Working Model
3 days in the office per week
Relocation
Relocation assistance available
Click here- https://alumhive.com/browse-jobs/6a79b62b0190c1f838d6f041
About the Team
Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs.
The Process Management team develops the distributed OS responsible for launching, coordinating and supervising the large numbers of processes that make up modern training workloads.
The runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability and observability.
Success is measured by both system reliability and researcher velocity, enabling ideas to scale from experiments to production training runs.
About the Role
As a Training Runtime: Process Management Engineer, you will work on software that connects thousands of computers and exposes them as a unified system.
The system supports individual researchers running multiple parallel experiments as well as large-scale training runs spanning hundreds of thousands and even millions of machines and accelerators.
You will primarily work in Rust, building high-performance asynchronous systems with a strong focus on performance, correctness and scalability.
The role involves solving complex and ambiguous infrastructure challenges while improving the efficiency, reliability and performance of OpenAI's training runtime and compute stack.
In this role, you will:
- Work across the Python and Rust stack.
- Design, build and maintain software to orchestrate and monitor machine learning workloads on large-scale supercomputers.
- Profile and optimize software to support computation orchestration at frontier scale.
- Improve reliability, observability and fault tolerance for long-running jobs.
- Debug complex distributed systems issues across large clusters.
- Respond to the evolving needs of ML systems and enable researchers to work efficiently.
You might thrive in this role if you:
- Have experience developing distributed systems, not just operating them.
- Enjoy understanding how large-scale systems behave and fail.
- Care deeply about performance, correctness and reliability.
- Have strong software engineering skills.
- Are proficient in Python and Rust, or another systems programming language such as C++.
- Have solid Linux knowledge.
- Are comfortable with systems-level debugging, performance analysis and memory profiling.
- Have experience developing asynchronous and concurrent systems.
- Enjoy working in high-ownership environments with strong engineering agency.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity.
OpenAI develops and deploys AI systems while focusing on advancing their capabilities and ensuring they are developed and used safely.
The company is committed to equal employment opportunity and providing reasonable accommodations to applicants with disabilities.
Additional Information
- Background checks are conducted in accordance with applicable law.
- Qualified applicants are considered consistent with applicable employment laws.
- Reasonable accommodations are available for applicants with disabilities.
- Applicants can refer to OpenAI's relevant employment, privacy and equal opportunity policies for additional information.
Application
Overview: OpenAI Careers
Application: Available through the official OpenAI job posting.
Pay: £36,000.00-£60,000.00 per year
Benefits:
Work Location: Hybrid remote in London