issues
search
allenai
/
dolma
Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
909
stars
94
forks
source link
Add preliminary Dolma v1.7 configurations, fix corner case in tokens.
#120
Closed
soldni
closed
7 months ago