Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser
Apache License 2.0
154
stars
18
forks
source link
TLDR-369 class for full dedoc pipeline running #300
version
parameter from metadata extractors, structure constructors and parsed document methods;