allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
909 stars 94 forks source link

BOS/EOS/PAD options in `tokens` cli; speed up tokenization by segmenting paragraphs. #102

Closed soldni closed 7 months ago

soldni commented 7 months ago