Tag
This paper analyzes statistical patterns in LLM-generated text using n-gram distributions, revealing stylistic deficiencies and showing that style and semantics are not separable.
This paper investigates the high false positive rate of existing AI-text detectors when applied to patent claims, which are legally required to be clear and concise like LLM output, and proposes a logistic regression classifier based on linguistic complexity features that outperforms perplexity-based detectors on consumer hardware.