How the scores work
What this dashboard measures, how each score is calculated, and how confidently you should read it — including where the analysis is still limited.
What this system does
A GitHub-Based Learning Analytics System for Analysing Entrepreneurial Learning Behaviour in Software Product Development turns raw GitHub repository activity into interpretable indicators of entrepreneurial development behaviour. It does not measure code quality, project success, or prove that learning happened — it surfaces behavioural evidence (commits, issues, pull requests and their text) that can be reviewed alongside other evidence.
Pipeline: GitHub activity is extracted for a repository → cleaned and stored → summarised into behavioural metrics → repository text is classified by the NLP component → both are combined into the two scores below and shown on each repository's page alongside the raw metrics they're built from.
The four behavioural dimensions
The Experimentation Intensity Score is built from four dimensions. Each is calculated only when there's enough data to support it — otherwise it's shown as N/A rather than guessed.
Engagement
Commit frequency and the share of days with recorded activity.
Evidence of: sustained iterative effort.
Regularity
How evenly spaced commits are, and the longest gap between them.
Evidence of: consistency vs. bursty activity.
Issue Refinement
Issue closure rate, comments per issue, and resolution time.
Evidence of: problem identification and follow-through.
Integration
Pull-request merge rate, reviewed-PR ratio, and cycle time.
Evidence of: structured collaboration and review.
Experimentation Intensity Score
- How it's calculated
- Each of the four dimensions above is min–max normalised against every repository currently in the database (and reverse-coded where a lower raw value is the stronger pattern, e.g. shorter inactivity gaps). The applicable dimensions are then averaged and scaled to 0–100.
- What it represents
- One comparable figure for how much iterative, structured development activity a repository shows — not a measure of code quality or project success.
- How to interpret it
- Because normalisation is relative to whatever repositories are currently stored, the same activity could score differently against a different set of repositories. Treat it as a within-this-dataset ranking aid, not a universal, externally validated benchmark.
- How to use it
- Compare repositories side by side, but always check the four dimension scores and the underlying metrics on each repository's page — they're the evidence behind the headline number, not just supporting detail.
Learning Quality Indicator
- How it's calculated
- Every classified commit message, issue, pull request and review is assigned one of five categories by a TF-IDF + classifier: Problem Identification, Experimentation, Reflection, Refinement, or None. The indicator is the percentage of classified text that fell into one of the four substantive categories (i.e. not None).
- What it represents
- An estimate of how much of a repository's written text reads as learning-oriented language — describing problems, trying things, reflecting on outcomes, or refining work — rather than routine or administrative text.
- How to interpret it
- The current classifier (Linear SVC) was trained on 100 labelled examples and measured 49.6% accuracy and 42.9% macro-F1 in cross-validation, with agreement rated “fair”. Treat individual classifications as exploratory suggestions, not certainties — use the category breakdown to sense-check against the actual repository text rather than relying on the percentage alone.
- How to use it
- As a starting point for identifying which artefacts might be worth reading directly — not as definitive proof that learning occurred.
Limitations, honestly
- Small samples. Metrics were developed and tested against a small number of controlled and public repositories.
- Limited NLP training data. The classifier was trained on a small labelled dataset, so it may struggle with unfamiliar wording, short messages, or ambiguous text.
- Project-specific indicators. The Experimentation Intensity Score and Learning Quality Indicator were developed for this study and haven't been validated across a large external population — read them as indicators of repository behaviour, not proof of learning.
Currently tracking 26 repositories. NLP classifier last validated at 49.6% accuracy on 100 labelled examples.