allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
894 stars 90 forks source link

Is there a way to intergratge Dolma toolkit to Spark? #136

Open DangoWang opened 5 months ago

DangoWang commented 5 months ago

My single computer is not powerful enough to run Dolma :(

soldni commented 4 months ago

hi @DangoWang. I haven't investigated spark integration, but you should be able to wrap individual taggers in spark.