- Measuring source code similarity across software projects
- =========================================================
- This is a working note on measuring source code similarity across software projects: what held up over several terms, what quietly stopped being done, and which of those two was actually a problem.
- Deciding what the output is for
- -------------------------------
- Distinguish between code that is similar and code that shares a history. Two implementations of the same textbook algorithm are similar and unrelated. Two files with the same unusual variable ordering, the same dead branch and the same off-by-one comment share a history. Systems that report only the former make the reviewer do the work of finding the latter, on every single pair, forever.
- What the reviewer actually needs
- --------------------------------
- Publish the method to the people being measured. Students and candidates who know what is compared, against what, and what happens next behave differently from those who do not — and the difference shows up as less of the thing you were detecting. Detection and deterrence are not in tension here; secrecy about the method buys a marginally higher catch rate and gives up nearly all of the deterrent effect.
- Making it survive the year
- --------------------------
- Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.
- Putting it into practice
- ------------------------
- What separates a source code plagiarism detection from a diff is that it can tell you what the overlap means.
- Reference: https://codequiry.com
Menu
New
Download
Show/Hide line no.
Copy text to clipboard