copyright-detection

Tag

Cards List
#copyright-detection

Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

arXiv cs.CL · 2026-07-14 Cached

This paper presents FindMyText, an open-source Python package that efficiently detects whether a given text appears within a large web-crawled corpus, using a novel fingerprint chaining mechanism for near-verbatim copy detection. It demonstrates superior performance on benchmarks across ArXiv, Wikipedia, and generic web content.

0 favorites 0 likes
← Back to home

Submit Feedback