allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
909 stars 94 forks source link

Can I use the dolma toolkit to process my own datasets? #127

Closed Tendo33 closed 6 months ago

Tendo33 commented 6 months ago

I got some data myself through a crawler, and I was wondering if I could use the dolma toolkit to remove duplicates.

soldni commented 6 months ago

Yes! you can use our dolma dedupe command. Please let us know if you have questions!