Submitted:
03 October 2026
Posted:
06 October 2026
You are already at the latest version
Abstract
FLOSS mailing lists record how newcomers learn, and earlier process-mining studies recognised learning activities in their messages with a-priori keyword models. This paper asks whether the architectures that now dominate natural language processing can perform that recognition step better, and what the language of contributions reveals about how participants mature. In Git and the Linux kernel every contribution is an e-mail that, once accepted, becomes a commit whose trailers (Reviewed-by, Helped-by, Mentored-by) name those who helped. This yields a corpus of 1,432,256 accepted patch e-mails with a learning clock for each author and a graph of helper-learner interactions. On author-disjoint partitions we compare an a-priori phase lexicon, TF-IDF with logistic regression, a Transformer trained from scratch and a Transformer pretrained by masked language modelling on 210,802 e-mails. Newcomers' e-mails (AUROC 0.789 in Git and 0.768 in Linux) and mentored work (0.688) are recognised best by TF-IDF with logistic regression. Pretraining improves the Transformer over training from scratch but leaves it 0.02 to 0.04 AUROC behind, and the a-priori lexicon is weakest (0.644, 0.575, 0.549). A fixed-cohort analysis shows that contributors' language moves towards a collective voice and explicit rationale as they mature, as the a-priori model predicts, whereas help received on the first patches is not associated with retention. A protocol for evaluating large language models on the same test sets is provided.
Keywords:
learning analytics
; mailing lists
; transformers
; language models
; NLP
; FLOSS
; newcomers
; mentoring
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.