Tag
This paper presents FindMyText, an open-source Python package that efficiently detects whether a given text appears within a large web-crawled corpus, using a novel fingerprint chaining mechanism for near-verbatim copy detection. It demonstrates superior performance on benchmarks across ArXiv, Wikipedia, and generic web content.