adbar / trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
https://trafilatura.readthedocs.io
Apache License 2.0
3.67k stars 263 forks source link

Duplicating sections, removing spaces between words, simple example #659

Closed nthomas-whistic closed 4 months ago

nthomas-whistic commented 4 months ago

ignore this