Skip to content
akaturk Academic measurement

Article detail · 2006 · conference-paper

Using Compression to Identify Classes of Inauthentic Texts

OpenAlex
Year2006
Citations38OpenAlex
Percentile%80.6
FWCI1.471.00 = world average

Data source split

  • OpenAlexOpenAlex enrichment (abstract, citations, topics)

Abstract

OpenAlex English

Recent events have made it clear that some kinds of technical texts, generated by machine and essentially meaningless, can be confused with authentic, technical texts written by humans. We identify this as a potential problem, since no existing systems for, say the web, can or do discriminate on this basis. We believe that there are subtle, short- and long-range word or even string repetitions extant in human texts, but not in many classes of computer generated texts, that can be used to discriminate based on meaning. In this paper we employ universal lossless source coding to generate features in a high-dimensional space and then apply support vector machines to discriminate between the classes of authentic and inauthentic expository texts. Compression profiles for the two kinds of text are distinct—the authentic texts being bounded by various classes of more compressible or less compressible texts that are computer generated. This in turn led to the high prediction accuracy of our models which support a conjecture that there exists a relationship between meaning and compressibility. Our results show that the learning algorithm based upon the compression profile outperformed standard term-frequency text categorization on several non-trivial classes of inauthentic texts. Availability: http://www.informatics.indiana.edu/predrag/fsi.htm.

Topics

Citations

OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.

38citationsOpenAlex · cited_by_count (cache / database)

Authors

No author information.