Extracting text data from PDF files

后端 未结 7 1861
[愿得一人]
[愿得一人] 2020-12-02 11:24

Is it possible to parse text data from PDF files in R? There does not appear to be a relevant package for such extraction, but has anyone attempted or seen this done in R?

相关标签:
7条回答
  • 2020-12-02 11:59

    Linux systems have pdftotext which I had reasonable success with. By default, it creates foo.txt from a give foo.pdf.

    That said, the text mining packages may have converters. A quick rseek.org search seems to concur with your crantastic search.

    0 讨论(0)
提交回复
热议问题