# How to choose a code similarity checker
This is a working note on how to choose a code similarity checker: what held up over several terms, what quietly stopped being done, and which of those two was actually a problem.
## Deciding what the output is for
Put the evidence somewhere durable and boring. Screenshots in a chat thread and a spreadsheet on one laptop are how findings get lost between the decision and the review of the decision. A plain directory of case files, one per matter, with the two sources and the date of retrieval, is unglamorous, needs no maintenance, and is the version that still exists when someone asks two years later.
## What the reviewer actually needs
Version the corpus, not just the code. A comparison run in March against a corpus that has since grown cannot be reproduced in June, and "we re-ran it and got a different number" is a sentence that ends processes. Recording which snapshot a result came from costs a column in a table, and is the difference between a finding that survives review and one that evaporates under it.
## Making it survive the year
Take the false positives seriously as a design input. Every pattern that reliably produces a harmless high score — generated code, a shared template, a language whose idioms are narrow — is something the pipeline can be told about once. Teams that log why each dismissal happened end the year with a filter that makes the next year's queue a third shorter. Teams that dismiss and move on start every year from the same place.
## Putting it into practice
What separates a [source code plagiarism detection](https://codequiry.com) from a diff is that it can tell you what the overlap means.