GetTogetherComm / GetTogether

Event manager for local community events
https://gettogether.community
BSD 2-Clause "Simplified" License
472 stars 89 forks source link

GetTogether.community content is used to train LLMs #320

Open cassidyjames opened 1 year ago

cassidyjames commented 1 year ago

You can confirm that 8.4k tokens were scraped from GetTogether.community by CommonCrawl and are included in Google's C4 dataset. It's likely that other LLMs have scraped and will continue to scrape user-generated content from GetTogether.community to train their proprietary large language models.

This can be discouraged for CommonCrawl and ChatGPT with the proper robots.txt inclusion:

User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /