allenai / dolma

Data and tools for generating and inspecting OLMo pre-training data.
https://allenai.github.io/dolma/
Apache License 2.0
1.02k stars 108 forks source link

DNM: Patch FT Tagger #210

Open undfined opened 2 months ago

undfined commented 2 months ago

Experimental patch to allow tagging a specific field in the document metadata.