allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
972 stars 107 forks source link

V2 of Gopher tagger #181

Closed soldni closed 3 months ago

soldni commented 3 months ago

This slightly improved version of the Gopher tagger no longer considers empty lines as duplicate lines.