Protocols for People: AI, the Internet and the IETF

This blog is sort of an update about the meeting we had with other stakeholders on the issue of crawlers, generative AI and publishers last week (Aug 2026) but since I have been wanting to give an update for a long time, it includes a lot of background:
The gist of the story is that some publishers (publishers in this sense means anybody who generates content) are upset that crawlers (including AI crawlers) crawl websites and use the content/data available on websites to develop various digital and AI products. Publishers want to signal their preferences in a machine-readable fashion to the crawler or the bot so that a crawler knows, for example, that they don’t want the content used for AI training. So they went to the Internet Engineering Task Force to come up with a vocabulary on the definitions of these preferences.
The Internet is pulled into this conversation from at least two angles: 1) generative AI’s main means of delivery to the end user is the Internet (for now). 2) the data that is on the Internet and across a wide variety of applications also can be the engine for generative AI.
Generative AI is nothing fancy. Thanks to the marketing teams, we call it AI to make it fancy. Generative AI is just very advanced software that generates content, text and images (it will be able to do other things as the technology develops).
In order to get these generative AI models to respond to your questions, hopefully accurately, and correct your mistakes and give you suggestions on how not to go bald or not panic when a bomb has landed near your place of living and translate your texts and PDFs, these AI models need to be trained on large amounts of data.
Humans have been creating data and knowledge and news with different modes of delivery since the early days of civilization and dare I say since homo sapiens first appeared. With the advent of the Internet, much of that knowledge moved onto the web and the Internet. News organizations that had to move complained about the Internet and some argued that moving data to the Internet was the industry’s “original sin”.
Apparently Dave Perry (an editor, in 2015) said: “Newspapers primarily hung themselves by giving away their content online 20 years ago, giving people a reason to go out and buy a 14K baud modem. We are now unable to put the genie back in the bottle. So just where do you think all those free online stories come from? Elves? The fruits of real journalists’ labors are freely given and stolen away by you and our pseudo-colleagues.”
The claim of course didn’t go unchallenged even then. In the age of AI, those old regrets about moving onto the Internet have returned. Almost all these arguments against the Internet are now being put forward against AI and the use of AI. But this time some publishers (this could mean someone with a site or a media organization) want to signal to the data collector/scraper/indexer how to use their data and what to use their data for. Speaking of site operators, they are a new kind of publisher. Maybe we can even call them Internet native publishers. Their business started because of the Internet. Some of them want to continue using the Internet but don’t want to commit the original sin (this time appearing in AI).
Since at least November 2024, an unsuspecting forum (at least in my opinion) has become the battlefield for deciding how data creators should be able to signal their preferences of how they want their content to be used to the crawlers that crawl data for all sorts of reasons and one could be training or showing summaries or translating.
That unsuspecting forum is called the Internet Engineering Task Force. The IETF’s mission generally is to provide technical standards for Internet interoperability and keep us connected and secure on the Internet. Its mandate is about the Internet, it doesn’t have a mandate to intentionally design its protocols to address issues offline.
A working group was convened at the IETF called AI preferences in November 2024 (it was AI control but they changed the name because it is not supposed to create controls, it is supposed to provide a vocabulary on what these preferences technically mean -what is AI training/what is AI search).
So what is the issue? Websites get indexed on search engines through crawlers and bots… bots dance around the whole Internet, crawl websites and index them (there is a lot of engineering of course happening which I shall not torture this piece with). Data collectors also crawl and scrape data for many more purposes, such as showing you the price differences between the shoes you want to buy. In the past a good engineer got annoyed at these bots because his servers were being overwhelmed and came up with a technical solution: the website owner could signal what page could be scraped and what couldn’t be. There was a real technical reason behind it: too much crawling would bring the website down.
That ad-hoc signaling (came into effect in the 90s) became an Internet protocol at the IETF a few years ago. It made sense. It was addressing a technical issue. However, in the age of AI, publishers don’t think this binary robots.txt “crawl allowed or crawl disallowed” is granular enough. They want a “crawl me” for indexing and search but a “don’t crawl me” for AI training or other things. They want to also say how the user or the AI provider can use their content. Some of them want to completely opt-out of everything AI because they don’t like AI. Since robots.txt became a protocol at the IETF and it’s now a protocol called robots extend, it felt natural that a group at the IETF come up with a set of vocabulary to technically communicate to crawlers what the site operator wants to be done with their content (or as the publishers like to say digital assets). There are other good things about the IETF: it is transparent; anybody can attend, there is no voting per se, much less elitist than ISO and other organizations.
The stakeholders in this group that are well represented are tech/AI companies and publishers. Civil society groups and academics such as Electronic Frontier Foundation and Center for Democracy and Technology, Knight Georgetown Institute as well as crawler organizations such as Common Crawl and Alliance for Responsible Data Collection also are involved. There is also Open Future and re-Create.
The conversations since the beginning of this group have been focused on how broad these preferences should be. At some point we came up with a “Text and Data Mining” category which you could opt out of and then you would have been opted out of everything. That was a disaster, text and data mining is way too broad and is enshrined in the European copyright law and it has demonstrated how bad it can be for developing software especially if you are not a well resourced developer or a not for profit or a research organization.
Then there was the category of automated processing that was even broader. Then the group decided not to have an overarching broad opt out and work on AI training and search. Publishers wanted AI inference, RAG (retrieval augmented generation), AI use to be discussed but it was too controversial so the chairs decided to make progress on other parts and then come to this debate.
That approach worked. We managed to come up with a specific AI training definition that is about generative AI and doesn’t entangle other functions. It goes without saying that even the current definition of training is going to have an impact on access to data to train for purposes such as translation and providing a multilingual gpt. But we can come up with carve outs that address those issues. (Carve outs have problems of their own- more on that later)
The most controversial category is AI-Use. I am sure it has happened to you, when you ask your GPT to check a link on a website and fetch the documents and it says it is not allowed to. That’s because the website through technical means has told your AI that your AI cannot access all or certain materials on that website. In that situation you go to the website to download the material or take a screenshot. Imagine if you can’t go to that website for any reason. Well tough luck, you can’t access it from the GPT. And in the last meeting I heard that publishers want to send you to their website instead of your AI visiting them because they need eyeballs for advertisers and some say because traffic is important. (remember not committing the original sin?)
At the IETF, publishers want to standardize this preference (how they want their content to be used) to apply globally and some such have the goal to attach these preferences even to PDF, images etc.
so it will look like this: you upload a pdf to translate:
Your AI: sorry can’t translate your document. You: why? Your AI (if nice enough): because it has signaled not to operate AI use. AI use includes translation. In the minds of publishers, AI use should include both what the consumer/the person behind the computer that is uploading something on a GPT as well as autonomous agents and systems. I have had many arguments about this, and if they succeed to get this, you won’t be able to even upload those PDFs or screenshots.
When we raise this issue, the response is these preferences are voluntary, and the AI model can just ask the user if they want to follow the preferences of the PDF. But you can never control the deployment. And these kinds of messages have chilling effects on people. The AI developer might just be risk averse and follow the preference and block the user from uploading. The user might not upload the file for the fear of being punished. Especially if it’s a government document! For accessibility purposes, a carve out is not enough, the protocol should be designed in a way that it does not entangle those accessibility cases (read this article about accessibility issues: https://www.wired.com/story/ebooks-drm-blind-accessibility-dmca/)
The IETF standard setting room is filled with stakeholders that might be dominant but they are not the only ones who crawl and use the Internet and AI. We have not been able to engage web accessibility experts to tell us what problems can emerge from the signaling. The group seems to acknowledge that we might have an impact with this protocol on people’s accessibility. It was good progress that it was suggested that a carve out list be drafted but then we shouldn’t come up with a protocol that is so broad that needs “carve-outs” to preserve people’s access and reduce negative impact on the end user.
When I talk about Internet freedom issues and access issues, these issues are discarded repeatedly as: these signals are just preferences, they are not compulsory to follow. No wording will change, no suggestion is added. This signaling will have an adverse effect on the open source crowd if it’s not narrow enough. Of course you can say they can just use intermediaries that have the power to do negotiations and understand the preferences. But I thought we are here to disintermediate and not lead to centralization and not make the powerful even more powerful.
Another counter argument to my “impact concerns” is that during deployment we will tell the AI developers how to interpret these preferences. How is that possible? Creative Commons will go around the world and talk to the Taiwanese, Bangladeshi and other developer communities?
When issues of Internet shutdown and restrictions are raised instead of addressing the text of the definition or saying we can come up with a clarifying paragraph or we can restrict this, the crowd says these are preferences and voluntary. For example if someone is trying to translate a medical document during a war and a shutdown via an LLM and that LLM does not allow the upload of that document, then the person during the war can go and use another LLM. A person should go LLM shopping when bombs are being dropped. It is not so hard to imagine a medical doctor might not be able to go LLM shopping during an Internet shutdown and a war.
When I mention copyright overreach and speech restrictions as one of the adverse effects, I am told that I should not bring “rights” to the table. I so agree that IETF should not be involved with adjudicating rights when designing protocols and that protocols should be based on technical necessity. I have even co-authored papers about that. But unfortunately this group has put itself in a position to adjudicate rights. I go one step further. This group seems to want AI crawlers (tech companies, open source developers and others) to determine whether there is copyright or other rights that entitle the person declaring the signal, to declare that preference.
I say this because I have not heard a single technical argument on why publishers want to opt out, I have heard mainly social and legal reasons: the arguments have been information accuracy, I don’t like AI, fighting undesirable content and my business is being affected because traffic has decreased. At some point we might even have a government interjecting that well we don’t want people to use AI to criticize us or make parody so we don’t want to be crawled and be used by AI models.
Another issue is that other standards processes and organizations are looking at the IETF to come up with this standard so that they implement it in the “digital assets” (that means PDFs, JPEG etc, C2PA is specifically keen). So you will have these preferences on everything and anything and well for some users it’s not hard to just remove the preferences but the Internet is for everyone. Not everyone is able to operate on the Internet in this way. I think I have exhausted my words on this piece and if you are still reading then we have just debunked the attention span crisis or that you are really interested in my writing and my take. (Wait, maybe lack of interest and interesting materials is the attention span crisis cause!)
Sometimes the group says “these are just preferences the receiving party can just ignore”, but then when a clarification is introduced they worry that a crawler, scraper, AI model provider might interpret it in a way that allows it to do something the publisher did not want. If the technical job of a preference signal is to communicate publishers preferences clearly to the receiving party, then we shall not be worried about compliance and shall not design the protocol around the compliance expectations. We shall design a narrow and clearly defined set of vocabulary to give the signal receiver enough information to make its decision. That decision could be to follow the preference or not. That’s all.
Clarifying point: I do think that we can standardize signaling preferences. But our definitions need to be narrow enough to reduce impact on the ordinary user of the Internet (the people!)
I would like to end this blog with a quote by Yelvington “The Internet was a mammal appearing at the end of the age of dinosaurs. People with big ideas were creating big things, people with little ideas were creating little things, and people with no ideas were complaining about the Internet.”
P.S: when I say publishers I mean publishers in the room. There are millions of websites that are publishers and not in the room. So take it with a grain of salt. It is not representative!

Beware of saddling your motor cycle! (Richard Prince, Helter Skelter: Arthur Jafa and Richard Prince, Venice, 2026)
Farzaneh Badii
Recent Posts



