adbar / trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
https://trafilatura.readthedocs.io
Apache License 2.0
3.64k stars 261 forks source link

Here is an interesting example... any tips? #459

Open krstp opened 10 months ago

krstp commented 10 months ago

Here is an interesting example: https://www.exin.com/article/ai-compliance-what-it-is-and-why-you-should-care/

Only html2txt will yields some, however, polluted extracts. Any input on the subject would be appreciated.

adbar commented 10 months ago

Hi @krstp, indeed, the extraction algorithms fail to capture the text, not sure why. I'll leave the thread open to see if we can find a solution.