About the role
Senior Data Scientist
United States
Engineering
Experienced Professional
Individual Contributor
Yes
5719
Full Time
mail\_outline
Get future jobs matching this search
or
Job Description
About GitHub
GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 companies, use GitHub to collaborate, and more than 77,000 organisations have adopted GitHub Copilot.
Locations
In this role you can work from Remote, United States
Overview
GitHub is seeking experienced professionals to elevate our data and analytics efforts. As a Senior Data Scientist in Data Science, you will leverage your deep expertise and knowledge of data science, AI/ML, analytics engineering, and business to lead data acquisition efforts, conduct thorough review of data analysis and data quality, form hypotheses and discover insights in the data to support business stakeholders and their decision making.
The ideal candidate will contribute to the impact of our Data Science initiatives and gain deep insights into the latest advancements in AI, machine learning and data science.
Responsibilities
Lead data acquisition efforts and ensure data is properly formatted and accurately described, while adhering to GitHub's privacy policies
Mentor others in data cleaning and data analysis best practices. Identify gaps in current data sets and drive onboarding of new data sets from production systems or third-party vendors.
Resolve data integrity problems in collaboration with relevant teams to promote upstream change and long-term quality
Leverage broad and deep knowledge of data modeling techniques, AI/ML tools, programming languages and query languages to create models, conduct experiments, analyze results, evaluating the methodology and performance of team members' models and recommending improvements.
Drive best practices relative to model validation, implementation, and application, and partners with teams across the organization to identify and explore new opportunities for driving transformative solutions for our stakeholders and customers.
Develop and articulate data-driven strategies in consideration of business priorities and lead conversations with end customers and/or internal stakeholders to understand, define, and solve business problems.
Communicate complex statistics, and machine learning topics to diverse audiences (e.g., multidisciplinary teams, customers, technical and non-technical audiences)
Independently writes efficient, readable, extensible code that spans multiple features/solutions. Contributes to the code/model review process by providing feedback and suggestions for implementation and improvement.
Drive operational excellence for model deployment (i.e. performance, scalability, monitoring, maintenance, integration into engineering production system, stability)
Produce project plans to define necessary steps required for completion, leading to a measurable improvement in business performance metrics over time.
Qualifications
Required Qualifications:
Bachelor's Degree in Data Science, Mathematics, Physics, Statistics, Economics, Operations Research, Computer Science, or related field AND 5+ years experience in data science (e.g., managing structured and unstructured data, applying statistical techniques) or related field
OR Master's Degree in Data Science, Mathematics, Physics, Statistics, Economics, Operations Research, Computer Science, or related field AND 3+ years experience in data science (e.g., managing structured and unstructured data, applying statistical techniques) or related field
OR Doctorate in Data Science, Mathematics, Physics, Statistics, Economics, Operations Research, Computer Science, or related field AND 1+ year(s) experience in data science (e.g., managing structured and unstructured data, applying statistical techniques) or related field
OR equivalent experience
Preferred Qualifications:
Technical understanding of data science techniques for regression, classification, time-series analysis, experimental design, causal inference
Proficiency in programming languages such as Python or R, experience with query languages such as SQL and KQL, and with data manipulation tools like Spark and Airflow
Able to clearly communicate findings to non-technical stakeholders through storytelling and visualization with tools like Jupyter notebooks or Azure Data Explorer/ PowerBI dashboards
Compensation Range
The base salary range for this job is USD $124,000.00 - USD $329,200.00 /Yr.
These pay ranges are intended to cover roles based across the United States. An individual's base pay depends on various factors including geographical location and review of experience, knowledge, skills, abilities of the applicant. At GitHub certain roles are eligible for benefits and additional rewards, including annual bonus and stock. These rewards are allocated based on individual impact in role. In addition, certain roles also have the opportunity to earn sales incentives based on revenue or utilization, depending on the terms of the plan and the employee's role.
This position will be open for a minimum of 3 days, with applications accepted on an ongoing basis until the position is filled.
GitHub values
Customer-obsessed
Ship to learn
Growth mindset
Own the outcome
Better together
Diverse and inclusive
Manager fundamentals
Model
Coach
Care
Leadership principles
Create clarity
Generate energy
Deliver success
Who We Are
Minimum requirements
- Bachelor's degree in a related field with 5+ years, or Master's with 3+ years, or Doctorate with 1+ year experience in data science or equivalent.
- Experience managing structured and unstructured data and applying statistical techniques.
- Proficiency in programming (Python or R) and query languages (SQL, KQL).
This listing was parsed by AI and may not be complete. Check the official posting on GitHub's site for the most accurate information.