About the role
Principal Production Engineer, Database Infrastructure
United States
Engineering
Experienced Professional
Individual Contributor
Yes
5711
Full Time
mail\_outline
Get future jobs matching this search
or
Job Description
About GitHub
GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 companies, use GitHub to collaborate, and more than 77,000 organisations have adopted GitHub Copilot.
Locations
In this role you can work from Remote, United States
Overview
GitHub is looking for a Principal Production Engineer to help scale our data platform to millions of developers. We are software engineers who specialize in reliability, working with technical partners, leading design reviews, writing SDKs and tooling to build against, and shaping how our data platform is used to prevent scaling problems before they reach production.
Responsibilities
Partner with product and feature teams by leading design reviews and shaping how they model, access, and scale their data so the applications they build are performant, available, and operable at GitHub's scale
Design and ship SDKs, client libraries, and the tooling applications are built on
Set reliability strategy for GitHub's database platforms across multiple systems, defining SLOs and operational standards
Write technical documentation and advocate for the health and quality of the systems the team builds
Participate in an on-call rotation and respond to incidents as needed
Develop and design plans for disaster recovery, load shedding, and regional failover
The team is highly distributed across geographies and time zones, and you will thrive in an environment of remote work and asynchronous communication.
Qualifications
Required Qualifications:
11+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Associate's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 10+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Bachelor's Degree in Computer Science or related field AND 9+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python
OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 7+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python.
OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python.
OR equivalent experience.
5+ years experience operating large-scale distributed systems in production, including participation in an on-call rotation.
Preferred Qualifications:
Excitement about building, operating, and maintaining resilient, scalable systems that impact a global community of users with the ability to break down complex systems into manageable components.
A track record of partnering with product and feature teams and materially changing the reliability and scalability of what they ship.
Ability to influence engineering decisions and proactively engage in system design conversations.
Experience running stateful services on managed cloud data stores, specifically Azure Cosmos DB and Azure SQL Database, or equivalents such as DynamoDB, Aurora, or Cloud Spanner with a focus on partition and index design, consistency and isolation tradeoffs, throughput provisioning, and hot-partition diagnosis.
Experience diagnosing and resolving application-level scalability problems: N+1 query patterns, hot partitions, unbounded fan-out, cache stampedes, and data access patterns that don’t scale.
A track record of building internal platforms, SDKs, or developer tools adopted across an engineering organization, written in production-grade Go, Python, Ruby, or Rust
Deep familiarity with the failure modes of large-scale systems, both in the application and in the platform beneath it e.g. cascading failures, retry storms, thundering herds, partial outages, throttling and quota limits, control plane outages, noisy neighbors, and the patterns that mitigate them.
Experience leading large-scale cloud migrations of live, high-traffic services e.g. dual-write and backfill strategies, traffic shifting, correctness verification, and rollback under load.
Effective communication skills and willingness to pair on problems, brainstorm in public, and enthusiastically engage with your teammates in group problem solving.
Compensation Range
The base salary range for this job is USD $160,200.00 - USD $425,000.00 /Yr.
Minimum requirements
- 5+ years operating large-scale distributed systems in production with on-call experience
- 5+ to 11+ years software engineering experience depending on education level, coding in languages like C, C++, Java, Go, Python, or Rust
- Degree in Computer Science or related field, or equivalent experience
This listing was parsed by AI and may not be complete. Check the official posting on GitHub's site for the most accurate information.