allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
894 stars 90 forks source link

New Progress Bar, Backoff, Batching #165

Open soldni opened 3 months ago

soldni commented 3 months ago

This PR adds three nice features to BaseParallelProcessor: