Sites scramble to block ChatGPT web crawler after instructions emerge

Without announcement, OpenAI recently added details about its web crawler, GPTBot, to its online documentation site. GPTBot is the name of the user agent that the company uses to retrieve webpages to train the AI models behind ChatGPT, such as GPT-4. Earlier this week, some sites quickly announced their intention to block GPTBot’s access to their content.

In the new documentation, OpenAI says that webpages crawled with GPTBot “may potentially be used to improve future models,” and that allowing GPTBot to access your site “can help AI models become more accurate and improve their general capabilities and safety.”

OpenAI claims it has implemented filters ensuring that sources behind paywalls, those collecting personally identifiable information, or any content violating OpenAI’s policies will not be accessed by GPTBot.

Read 12 remaining paragraphs | Comments

Post Views: 59

Sites scramble to block ChatGPT web crawler after instructions emerge

technology_o6swjd

You May Also Like

Despite the recent downturn in the cryptocurrency market, Y Combinator’s Summer 2022 batch has 30 crypto startups, up from 25 in its Winter 2022 batch (TechCrunch)

Fend off cybercrime with this all-in-one security bundle