SkyworkAI / Skywork

Skywork series models are pre-trained on 3.2TB of high-quality multilingual (mainly Chinese and English) and code data. We have open-sourced the model, training data, evaluation data, evaluation methods, etc. 天工系列模型在3.2TB高质量多语言和代码数据上进行预训练。我们开源了模型参数,训练数据,评估数据,评估方法。
Other
1.21k stars 111 forks source link

Skypile-150B数据里是否包含Skypile-STEM数据? #34

Closed zgctmac closed 10 months ago

zgctmac commented 10 months ago

如题,麻烦问下开源的Skypile-150B数据里是否包含第二阶段预训练所需要的STEM数据

TianwenWei commented 10 months ago

如题,麻烦问下开源的Skypile-150B数据里是否包含第二阶段预训练所需要的STEM数据

不包含STEM。