Respect for the law could literally be written into the source code of a model: If OpenAI is aggressively downloading data from everywhere, the least they could do is drop in the U.S. Code (particularly the section tied to the Computer Fraud and Abuse Act) and train the model not to violate it.
Eh, idk. I read that these big US AI companies actually were respecting robots.txt and identifying with a well defined clear User-Agent. So they were easy to rate limit or block.
Eh, idk. I read that these big US AI companies actually were respecting robots.txt and identifying with a well defined clear User-Agent. So they were easy to rate limit or block.
China-based IPs, however…