The Washington Post takes a closer look at Google’s C4 dataset, which is comprised of the content of 15 million websites, and has been used to train various LLM’s. Perhaps also the one used by OpenAI for e.g. ChatGPT, although it’s not known what OpenAI has been using as source material.
They include a search engine, which let’s you submit a domain name and find out how many tokens it contributed to the dataset (a token is usually a word, or part of a word).
Obviously I looked at some of the domains I use. This blog is the 102860th contributor to the dataset, with 200.000 tokens (1/10000% of the total).
Screenshot of the Washington Post’s search tool, showing the result for this domain, zylstra.org.
Molly White does a good write-up of the extremely odd and botched launch by Feedly of a service to keep tabs on protests that might impact your brand, assets or people. Apparently in that order too. When E first mentioned this to me I was confused. What’s the link with a feedreader after all? Feedly’s subsequent excuse ‘we didn’t consider abuse of this service’ sounds rather hollow, as their communications around it seem to precisely focus on the potential abuse being the service announced.
The question ‘how did Feedly end-up here?’ kept revolving in my mind. Turns out the starting point is logged in my own blog:
Machines can have a big role in helping understand the information, so algorithms can be very useful, but for that they have to be transparent and the user has to feel in control. What’s missing today with the black-box algorithms is where they look over your shoulder, and don’t trust you to be able to tell what’s right.
Edwin Khodabakchian cofounder and CEO of RSS reader Feedly, in Wired, March 2018
In a twisted way I can see the reflection of that quote in the service Feedly announced. Specifically w.r.t. the first part, using algorithms to better understand information. The second part seems to have gone missing in the past 5 years though, the bit about transparency, avoiding black boxes, and putting users in control. Especially the ‘not trusting people to tell what’s right’ grates. It seems to me Feedly users in the past days very much could tell what’s right and Feedly hoped they wouldn’t.
I do agree with the 2018 quote though, but ‘algorithmic interpretation as a service‘ isn’t what follows to me. That’s just a different way of commoditising your customers.
Algorithmic spotting of emergent patterns is relevant if I can define the context and network of people (and perhaps media sources) whose feeds I follow. For that I need to be in control of the algorithm, and need to be the one who defines what specific emergent patterns I am interested in. That is on my list for my ideal feed reader. But this botched Feedly approach isn’t that.
Adding this interesting perspective from Mita Williams to my notes on the effects of generative AI. She positions generative AI as bypassing the open web entirely (abstracted away into the models the AIs run on). Thus sharing is disincentivised as sharing no longer brings traffic or conversation, if it is only used as model-fodder. I’m not at all sure if that is indeed the case, but from as early as YouTube’s 2016 Flickr images database being used for AI model training, such as IBM’s 2019 facial recognition efforts, it’s been a concern. Leading to questions about whether existing (Creative Commons) licenses are fit for purpose anymore. Specifically Williams pointing to not only the impact on an individual creator but also on the level of communities they form, are part of and interact in, strikes me as worth thinking more about. The erosion of (open source, maker, collaborative etc) community structures is a whole other level of potential societal damage.
Mita Williams suggests the described erosion is not an effect but an actual aim by tech companies, part of a bait and switch. A re-siloing, an enclosing of commons, where being able to see something in return for online sharing again is the lure. Where the open web may fall by the wayside and become even more niche than it already is.
…these new systems (Google’s Bard, the new Bing, ChatGPT) are designed to bypass creators work on the web entirely as users are presented extracted text with no source. As such, these systems disincentivize creators from sharing works on the internet as they will no longer receive traffic…
Those who are currently wrecking everything that we collectively built on the internet already have the answer waiting for us: web3.
…the decimation of the existing incentive models for internet creators and communities (as flawed as they are) is not a bug: it’s a feature.
Staatssecretaris van Digitalisering Alexandra van Huffelen bereidt nu mogelijk een besluit voor langs dezelfde lijnen. Terecht lijkt me. Meta houdt zich zelf niet aan de AVG, en bovendien is de algemene uitwisseling van Europese persoonsgegevens met de VS geheel niet juridisch gedekt op dit moment.
De overheid moet zelf het goede voorbeeld geven bij online interactie met burgers en de omgang met eigen gegevens. Dit geldt voor Meta, voor Twitter, maar ook voor cloud diensten en de Microsoft lock-in waar de overheid zich grotendeels in bevindt. Facebook zelf niet meer gebruiken is een bescheiden eerste signaal, dat al verrassend lastig lijkt voor de overheid om helder af te geven.
Ik hoop dat de staatssecretaris de knoop snel doorhakt.
Ein datenschutzkonformer Betrieb einer Facebook-Fanpage sei nicht möglich, schrieb Kelber in einem Brief an alle Bundesministerien und obersten Bundesbehörden.
Iskander asks what about users, next to makers, when it comes to responsible AI? For a slightly different type of user at least, such responsibilities are being formulated in the proposed EU AI Regulation, as well as the connected AI Liability Directive. There not just the producers and distributors of AI containing services or products have responsibilities, but also those who deploy them in practice, or those who use its outputs. He’s right that most discussions focus on within the established system of making, training and deploying AI, and we should also look outside the system. Where in this case the people using AI, or using their output reside. That’s why I like the EU’s legislative approach, as it doesn’t aim to regulate the system as seen from within it, but focuses on access conditions for such products to the European market, and the impact it has within society. Of course, these proposals are still under negotiation, and it’s wait and see what will remain at the end of that process.
As I wrote down as thoughts while listening to Dasha Simons; we are all convinced of the importance of explainability, transparency, and even interpretability, all focused on making the system responsible and, with them, the makers of the system. But what about the responsibility of the users? Are they also part of the equation, should they be responsible too? As the AI (or what term we use) is continuous learning and shaping, the prompts we give are more than a means to retrieve the best results; it is also part of the upbringing of the AI. We are, as users, also responsible for good AI as the producers are.
Al in maart had ik in Utrecht een leuk gesprek met Martijn Aslander en Lykle de Vries als onderdeel van hun podcast-serie Digitale Fitheid. Digitale Fitheid is een platform over, ja precies dat, de digitale fitheid voor de kenniswerker.
In het gesprek hadden we het over persoonlijk kennismanagement (pkm) en de lange historie daarvan, en de omgang met digitale gereedschappen en de macht om die tools zelf vorm te geven. Maar ook over mijn werk, verantwoord datagebruik, de Europese datastrategie, Obsidian meet-ups, en ethiek. Er kwam aan het begin zelfs met veel kabaal een AWACS voorbij.
Een gesprek van een uur dat zo voorbij was. Achteraf denk je dan, heb ik wel coherente dingen gezegd? Terugluisterend nu bij publicatie, valt dat mee.