Last Updated: 12 February 2025

Baidu, one of China's leading internet search companies, has taken steps to prevent Google and Bing from accessing data on its Wikipedia-style platform, Baidu Baike.
The update, which affects the site's robots.txt file—used by search engines to define which pages can be accessed—now explicitly restricts Googlebot and Bingbot from crawling Baidu Baike's content.
According to the Wayback Machine, an internet archive tool, this update likely occurred on August 8.
Prior to the update, Baidu Baike had permitted both Google and Bing to access nearly 30 million entries available on the platform, with only a few restricted areas.
This change highlights Baidu’s firm intention to safeguard its digital resources as demand for AI project data surges. In the AI development landscape, vast amounts of data are critical for training models and developing cutting-edge technologies. By restricting access, Baidu ensures that its proprietary content remains protected, limiting the ability of competitors to use its data for their AI initiatives.
This move echoes a similar decision made by Reddit, an American discussion platform, in July, where it restricted certain search engines, except for Google, from indexing its posts. Google compensates Reddit for this access, as the data is valuable for supporting AI training efforts.
Since the launch of OpenAI’s ChatGPT on November 30, 2022, major players like Google and Microsoft have been scrambling to secure more data to fuel their conversational systems.
Last year, reports surfaced that Microsoft even warned other search engine operators that they might restrict access to high-powered internet search data unless they agreed not to use it for developing competing AI services.
By taking this step, Baidu underscores the increasing importance of data control in the race to develop advanced AI technologies, marking another shift in the competitive landscape of global tech giants.