1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
|
# Plagiarism detection for data science notebooks
This is a working note on plagiarism detection for data science notebooks: what held up over several terms, what quietly stopped being done, and which of those two was actually a problem.
## Deciding what the output is for
Boilerplate is not noise to be tuned out — it is signal about the assignment, and the right place to remove it is the specification, not a threshold. If forty per cent of every submission is scaffolding the brief supplied, then every pairwise score starts at forty per cent and the interesting variation is compressed into the top half of the range. Subtract the supplied code first and the same detector suddenly discriminates.
## What the reviewer actually needs
Take the false positives seriously as a design input. Every pattern that reliably produces a harmless high score — generated code, a shared template, a language whose idioms are narrow — is something the pipeline can be told about once. Teams that log why each dismissal happened end the year with a filter that makes the next year's queue a third shorter. Teams that dismiss and move on start every year from the same place.
## Making it survive the year
Publish the method to the people being measured. Students and candidates who know what is compared, against what, and what happens next behave differently from those who do not — and the difference shows up as less of the thing you were detecting. Detection and deterrence are not in tension here; secrecy about the method buys a marginally higher catch rate and gives up nearly all of the deterrent effect.
## Putting it into practice
What separates a [source code plagiarism detection](https://codequiry.com) from a diff is that it can tell you what the overlap means.
|