1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
|
# How token-based code comparison works
A short account of how token-based code comparison works from the side that has to sit in the meeting afterwards rather than the side that runs the comparison.
## Deciding what the output is for
Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.
## What the reviewer actually needs
Self-plagiarism and legitimate reuse need separate handling, and a similarity score cannot tell them apart. A student reusing their own prior submission, a developer reusing a snippet from an internal library and an author lifting a file from a public repository can all produce the same number. Only provenance separates them, so provenance has to be captured at the moment of the match rather than reconstructed later.
## Making it survive the year
Decide what the output is for before choosing what produces it. A number that feeds a conversation, a report that feeds a formal process and a log that feeds an audit have almost nothing in common, and a system tuned for one is actively unhelpful for the others. Most disappointment traces back to a tool bought for the third purpose being used for the first, by people who were never told which it was.
## Putting it into practice
Corpus, provenance and a readable report — a [code similarity checker](https://codequiry.com) without all three is a demo.
|