Jeff Duntemann's Contrapositive Diary Rotating Header Image

wikipedia

Review: The Commons Captured by Michael J. Bommarito II

Commons Captured CoverI’ve had an eye out for a history of Wikipedia for some time, and the other day stumbled across The Commons Captured, by Michael J. Bommarito II. And yes, it does provide a reasonably detailed history going back to Wikipedia’s origins in 2001. But that’s not really the book’s focus, as the title suggests.

In addition to the history of Wikipedia, the book provides a history of what Richard Stallman created and called “copyleft:” a system of licensing that provides intellectual property for free, with the caveat that all derivative works maintain the price tag of $0.00. It’s the Creative Commons, along with ShareAlike, generally abbreviated as CC BY-SA, now at version 3.0. (See this link for details.) I released my Free Pascal from Square One ebook under CC BY-SA 3.0. From a height, this means that you can share it, pass it around, and borrow parts of it for your own projects, but you cannot sell it. This is the license that supposedly surrounds Wikipedia and its (now) seven million articles. And it worked great for over 20 years.

Then alluvasudden AI happened, and everything changed.

That’s the book’s focus: How AI took advantage of Wikipedia’s freeness, and how the powers that operate Wikipedia broke Creative Commons and licensed Wikipedia’s content to AI companies…for money.

Really. Wikipedia’s operators created the Wikimedia Foundation as a licensed nonprofit. AI tech firms had begun training their AIs on Wikipedia’s content almost as soon as there were AIs, without paying or asking permission. Hey, reading Wikipedia is free, right? So why not? Well, because the AI companies sold access to their AIs to the public for money. And having digested Wikipedia’s seven million articles, those AIs began answering questions from AI firms’ customers using Wikipedia’s content.

The Wikimedia Foundation created a licensing system that offered a license to AI firms. Many (but not all) of the AI firms paid money—big money—for the right to train their AIs on Wikipedia’s content. Granted, it takes money for servers to host something as big and as popular as Wikipedia. But the Foundation brought in way more cash from licensing AI training than it needed to keep the servers pumping.

The real bummer is that Wikipedia’s articles are written and edited by volunteers who pocket nothing for their time and effort. And none of that license money ended up in those volunteers’ pockets. None. Also, as more people hooked up with AI, fewer people went up to read Wikipedia itself.

So what’s the end of the story? Well, the story hasn’t ended. Michael Bommarito sketched out four possible outcomes of the AI licensing issue over the coming years, but he isn’t optimistic, and I won’t summarize those here. Buy the book.

What I found most fascinating was the gray haze that still surrounds the issue of what AIs store, how they store it, and whether or not training AIs falls under the legal category of fair use. Wikipedia is not copyrighted, but copylefted, so that shouldn’t matter. But…the Foundation is selling Wikipedia access to the AI firms, who sell access to their trained AIs to a (mostly) paying public. So much for copyleft.

The book is very well-written but dense, and demands close attention. I found it puzzling that the book makes no mention of Justapedia, which is a very good Wikipedia fork that I use and may someday (as time permits) attempt to write some articles for. The book does mention Grokipedia and Everipedia, though not at length. So I’ll keep looking.

But if you’re interested in the emerging legal issues surrounding AI training, this is definitely the book to read.

Daywander: The Penny with the Upside-Down Date

1961 upside down.jpg

In 1961, when I was nine, our family went to the Illinois State Fair in Springfield. We took a tour of the town, saw Lincoln’s tomb, and had a lot of fun together. The fair was wonderful. My father ran into Larry Fine (1902-1975) of the Three Stooges at the beer tent and bought him a drink. When we reconvened toward the end of the day (my mother had carnival-ride duty) he had a gift for me: A 1961 penny with the rare upside-down date.

It wasn’t all that rare. In fact, I got one in change at the Mickey D drive-thru this morning. Now, the poor thing has clearly spent some time in a parking lot, but you can see from the photo that the date is indeed upside-down. Well, the nine-year-old I was in 1961 certainly thought so. It took me a week or so to figure out that he was not just pulling my leg, but yanking on it so hard it squeaked. And yes, when I figured it out I thought it was funny as hell.

Today was a weird day. Last night we got almost two inches of rain. Two inches. In Phoenix. In one night. Now, on average, Phoenix gets 8″ of rain a year. So last night represented 25% of our annual rainfall. Oh–they’re predicting another 2″ tonight.

I guess this is going to be a wet-ish year. Our ash trees may survive after all.

Now, has anyone else ever heard of “heat lightning”? When we were kids, we sometimes saw lightning flashing along the horizon (or at least above the nearby houses) on hot summer evenings, with no least whisper of thunder. Nobody explained why it was called “heat lightning.” Light travels farther than sound, and heat lightning is just lightning so far away that the thunder can’t keep up with the flash.

It must be important. Heat lightning has a Wikipedia page, and I don’t. However, the Wikipedia Gawds are threatening to delete the page if some acceptable peer-reviewed studies aren’t produced immediately to provide evidence that heat lightning is real and not a hoax. Don’t you dare suggest a Primary Source. Primary Sources make the Gawds drool on the floor and then start throwing chairs. The only way to escape them with your life is to run while screaming “I’m not notable!” at the top of your lungs until you’re past chair-throwing distance.

Heh. And you think I’m kidding.

The All-Volunteer Federated Encyclopedia of (Really!) Absolutely Everything

My regular readers will recall that I wrote an article in the June/July 1994 issue of PC Techniques, describing a distributed virtual encyclopedia that pretty much predicted Wikipedia’s function, if not the details of its implementation. My discontent with Wikipedia is not only well-known but not specific to me: The organization has become political, and editor zealots have various tricks to make their ideological opponents either look bad, or disappear altogether. Key here is their concept of notability, which is Wikipedia’s universal excuse for excluding the organization’s ideological opponents from coverage.

In one of the decade’s great hacks, Vox Day created Infogalactic, which is a separate instance of the MediaWiki software underlying Wikipedia and a fair number of other, more specialized encyclopedias. Infogalactic has a lot of its own articles. However, when a user searches for something that is not already in the Infogalactic database, Infogalactic passes the search along to Wikipedia, and then displays the returned results. I don’t know whether or to what extent Infogalactic keeps results from Wikipedia on its own servers. It’s completely legal to do so, and they may have a system that keeps track of frequent searches and maintains frequently searched-for Wikipedia pages in local storage. Or they may just keep them all. We have no way to know.

Infogalactic’s relationship with Wikipedia immediately suggested a form of federation to me, though Infogalactic does not use that term. (Federation means a peer-to-peer network of nodes that are independently hosted and maintained yet query one another.) The Mastodon social network system is the best example of online federation that I could offer. (It’s not shaped like an encyclopedia, so don’t take the comparison too far.) There is something else called the Fediverse, which I have not investigated closely. In a sense, the Fediverse is meta-federation, as it federates already federated platforms like Mastodon. For that matter, Usenet is also a form of federation. It’s been around a long time.

The MediaWiki software is open-source and freely available to anyone. There are a lot of special-interest wikis online. One is about Lego. (Brickipedia, heh.) For that matter, there’s one about Mega Bloks. Hortipedia is about gardening and plants generally. It’s a huge list; give it a scan. You might find something useful.

My suggestion is this: Devise a MediaWiki mod like Infogalactic’s, but take it farther. Have a “federation panel” that allows the creation of lists of MediaWiki instances for searches falling outside the local instance. A list would generally start with the local instance. It might then search instances focusing on related topics. The last item on most lists would be a full general encyclopedia like Wikipedia or Infogalactic.

Here’s a simple example, which could defeat Wikipedia’s notability fetish for biographies and a lot of other things: Begin a search for a given person (or other topic) with Infogalactic, which, remember, searches Wikipedia if its own database doesn’t satisfy the query. So if that search fails, submit the same search to EverybodyWiki, which doesn’t apply notability criteria to biographies. In fact, EverybodyWiki does what I suggested be done a number of years ago: It collects articles marked for deletion on Wikipedia, of which it currently has over 100,000. I’m tempted to post a biography on Wikipedia just to see if, when it’s deleted (and it would be) EverybodyWiki picks it up.

(As an aside: I just found EverybodyWiki a month or so ago, and I’m surprised I hadn’t heard of it before. It has more than just biographies and is definitely worth a little poking-around time.)

Now, the tough part: How would this be accomplished? I don’t know enough about MediaWiki internals to attempt it myself. There’s an API, and I’ve been surfing through the API doc. There’s even an API sandbox, which is a cool idea all by itself. Alas, there are remarkably few technical books on MediaWiki, and the ones I would be most interested in get terrible reviews. Given how important MediaWiki is, I don’t understand why tech publishers have skated past it. My guess is that few people bother to do more than custom-skin MediaWiki. (There’s a book on that, at least.) If the demand were there, the books would probably happen. If you know enough about MediaWiki modding, I’ll bet you could find a publisher.

I’m thinking about installing MediaWiki on my hosting services, just to poke at and try things on. Hell, I predicted this thing. I should at least know my way around it.

If you’ve done any hacking on MediaWiki, let me know how you learned its internals and what you did, and if there are any instructional websites or videos that I may not have encountered.