Automatic Core-Developer Identification on GitHub: A Validation Study
Options
BORIS DOI
Publisher DOI
Description
Many open-source software projects are self-organized and do not maintain official lists with information on
developer roles. So, knowing which developers take core and maintainer roles is, despite being relevant, often
tacit knowledge. We propose a method to automatically identify core developers based on role permissions
of privileged events triggered in GitHub issues and pull requests. In an empirical study on 25 GitHub projects,
(1) we validate the set of automatically identified core developers with a sample of project-reported developer
lists, and (2) we use our set of identified core developers to assess the accuracy of state-of-the-art unsupervised
developer classification methods. Our results indicate that the set of core developers, which we extracted
from privileged issue events, is sound and the accuracy of state-of-the-art unsupervised classification methods
depends mainly on the data source (commit data versus issue data) rather than the network-construction
method (directed versus undirected, etc.). In perspective, our results shall guide research and practice to
choose appropriate unsupervised classification methods, and our method can help create reliable ground-truth
data for training supervised classification methods.
developer roles. So, knowing which developers take core and maintainer roles is, despite being relevant, often
tacit knowledge. We propose a method to automatically identify core developers based on role permissions
of privileged events triggered in GitHub issues and pull requests. In an empirical study on 25 GitHub projects,
(1) we validate the set of automatically identified core developers with a sample of project-reported developer
lists, and (2) we use our set of identified core developers to assess the accuracy of state-of-the-art unsupervised
developer classification methods. Our results indicate that the set of core developers, which we extracted
from privileged issue events, is sound and the accuracy of state-of-the-art unsupervised classification methods
depends mainly on the data source (commit data versus issue data) rather than the network-construction
method (directed versus undirected, etc.). In perspective, our results shall guide research and practice to
choose appropriate unsupervised classification methods, and our method can help create reliable ground-truth
data for training supervised classification methods.
Date of Publication
2023-11
Publication Type
Article
Keyword(s)
Open-Source Software Projects
•
Developer Classification
•
Developer Networks
Language(s)
en
Contributor(s)
Bock, Thomas | |
Joblin, Mitchell | |
Apel, Sven |
Series
ACM Transactions on Software Engineering and Methodology
Publisher
Association for Computing Machinery (ACM)
ISSN
1049-331X
1557-7392
Access(Rights)
restricted