# How token-based code comparison works

A short account of how token-based code comparison works from the side that has to sit in the meeting afterwards rather than the side that runs the comparison.

## Deciding what the output is for

Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.

## What the reviewer actually needs

Self-plagiarism and legitimate reuse need separate handling, and a similarity score cannot tell them apart. A student reusing their own prior submission, a developer reusing a snippet from an internal library and an author lifting a file from a public repository can all produce the same number. Only provenance separates them, so provenance has to be captured at the moment of the match rather than reconstructed later.

## Making it survive the year

Decide what the output is for before choosing what produces it. A number that feeds a conversation, a report that feeds a formal process and a log that feeds an audit have almost nothing in common, and a system tuned for one is actively unhelpful for the others. Most disappointment traces back to a tool bought for the third purpose being used for the first, by people who were never told which it was.

## Putting it into practice

Corpus, provenance and a readable report — a [code similarity checker](https://codequiry.com) without all three is a demo.