About the role
Staff Software Engineer, Compute Platform
United States
Engineering
Experienced Professional
Individual Contributor
Yes
5739
Full Time
Get future jobs matching this search
or
Job Description
About GitHub
GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 companies, use GitHub to collaborate, and more than 77,000 organisations have adopted GitHub Copilot.
Locations
In this role you can work from Remote, United States
Overview
GitHub is changing the way the world builds software, and we want you to help lead this effort.
The Compute Platform team owns and operates the Kubernetes-based platforms that run GitHub's production workloads. We are responsible for Hubbernetes, GitHub's internal Kubernetes platform, and for the platform engineering work needed as GitHub moves more production infrastructure to Azure Kubernetes Service. Our work spans cluster lifecycle, fleet reliability, capacity, workload guardrails, developer experience, and the operational foundations that service teams build on every day.
As a Staff Software Engineer on Compute Platform, you will work with a distributed team to lead the technical direction of a Kubernetes platform that hundreds of internal teams depend on. You will design and implement systems that make GitHub's compute substrate easier to scale, safer to operate, and harder to misuse, with a focus on reliability, automation, operational clarity, and reducing toil.
The Compute Platform team is highly distributed, and you will thrive in an environment of remote work and asynchronous communication. You should have strong written communication skills, strong software engineering fundamentals, and several years of hands-on experience building or operating Kubernetes and cloud-native platform systems in production. Azure and AKS experience are strong pluses.
Responsibilities
Design, build, and operate Kubernetes-based platform systems that support GitHub production services.
Improve reliability, scalability, and operability across GitHub's Kubernetes / Hubbernetes fleet.
Lead technical work across ambiguous, cross-team platform problems, including AKS migration, fleet management, control-plane reliability, and capacity management.
Partner with service owners, infrastructure teams, and engineering leaders to define safe migration paths, operational standards, and long-term platform direction.
Debug and resolve complex production issues across Kubernetes, cloud infrastructure, Linux, containers, networking, observability, and distributed systems. Build automation and guardrails that improve workload quality, capacity efficiency, cluster lifecycle operations, and incident response.
Mentor engineers and raise the engineering bar through design review, code review, technical direction, and operational judgment.
Write, review, test, and maintain reliable platform software and automation. Go experience is helpful, but strong production software engineering fundamentals matter more than any single language.
Participate in on-call and incident response for the systems owned by the team, including driving follow-up work that prevents repeat issues.
Qualifications
Required Qualifications:
9+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Associate’s Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 8+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Bachelor's Degree in Computer Science or related field AND 7+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 3+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python.
OR equivalent experience.
2+ years experience with large-scale Kubernetes fleet operations, cluster lifecycle management, capacity management, or workload migration.
Preferred Qualifications:
Experience with Azure Kubernetes Service or operating Kubernetes on Azure.
Experience with platform engineering, developer platforms, internal infrastructure products, or infrastructure-as-a-service style systems.
Experience with Kubernetes controllers, operators, CRDs, admission control, scheduling, autoscaling, service mesh, or multi-cluster networking.
Experience improving reliability through SLOs, observability, incident response, automation, and operational readiness practices.
Experience leading Staff-level technical projects with high ambiguity and broad organizational impact.
Compensation Range
The base salary range for this job is USD $140,400.00 - USD $372,300.00 /Yr.
Minimum requirements
- 7+ years software engineering experience with production coding in languages like C, Java, Go, Python, or equivalent education/experience combinations.
- 2+ years experience managing large-scale Kubernetes fleet operations, cluster lifecycle, capacity, or workload migration.
- Strong software engineering fundamentals and hands-on Kubernetes/cloud-native platform production experience.
This listing was parsed by AI and may not be complete. Check the official posting on GitHub's site for the most accurate information.