One of my favourite hobbies is ruining other peopleโs fun.
Because I spend a lot of time on Twitter, over the years it has become an increasingly annoying part of my schtick that if everyone is talking about something that isnโt true, Iโll be first in the queue to call bullshit on it.
Thereโs nothing I enjoy more than digging into the PDF to read what the controversial policy actually says, or spending five minutes figuring out the broader context for the outrageous thing the politician said, just so I can say โactually, itโs more complicated than thatโ. Itโs all because of my long-held view that we should try to post things that are true โ not things that we want to be true.
So as you might imagine, Iโm an incredibly popular person. And when I asked my partner if Iโm like this in real life too, she just looked at her feet and sighed heavily, for some reason.
But the truth is that, deep down, I know Iโm fighting a losing battle against the nonsense and the lies.
The reality is that Iโm never going to be able to fact-check the entire internet. Worse still, on a psychological level, our brains are wired so that we donโt actually care all that much about the truth: we treat things that our friends and allies say more credulously, and approach the utterances of our opponents with scepticism and suspicion.
Thatโs why I want to die inside when I tell someone the damning quote theyโre sharing isnโt real, or is missing context, and they respond by saying: โI donโt care if itโs fake, it’s funnyโฆ and besides, it feels like it could be true.โ
But my problems donโt really matter. My ordeals in being a pedantic asshole on the internet are only foreshadowing whatโs coming. As just over the horizon is a tsunami of misinformation, fakeries and lies.
The AI tsunami
Iโm talking, of course, about generative AI. Weโre seeing on an almost daily basis technological leaps that would have seemed like magic just a few short years ago.
Take AI image generation. In mere months, the technology has leapt from fuzzy DALL-E, unable to render faces properly, to the incredible capabilities of Midjourney 5, which can generate near photo-realistic images of pretty much anything.
It seems like a lifetime ago now, but it was only back in January we were sneering at AI for not being able to get fingers right. But itโs safe to say that those takes did not age well.
There is no sign of the rapid pace of improvements slowing down. Iโm sure in the not-too-distant future, weโll start seeing similarly impressive videos โ especially as AI techniques are slowly getting better at creating consistent characters across multiple generations.
So what weโre living through is a Cambrian explosion in AI capabilities. Although it’s exciting to consider the new creative possibilities, itโs scary too. Because the internet is about to be flooded with fake images, text and โ eventually โ video and audio, at a new order of magnitude to what weโre currently used to.
Worse still, thereโs nothing we can do to stop it โ because the genie is already out of the bottle.
Why? Because even if the big AI platforms were to close down, become heavily regulated, or otherwise restrict how they are used, it wonโt stop the AI bullshit tsunami. Even if the AI worriers get the โpauseโ they want, it is too late. Because generative AI doesnโt require a super-computer of specialist equipment. It just runs on the regular, normal computers that we all have on our desks, in our laps, and in our pockets.
Iโve seen this myself. I was able to download Stable Diffusion (it was a little over 4GB), an open source AI image generator, to my Mac and run it locally โ no cloud required. Itโs only a matter of time until we have apps that make generating realistic images as easy as posting on Instagram or composing a tweet.
The tools are already getting easier to use. Literally as I was in the midst of writing this essay, Adobe released a beta version of Photoshop that has built-in image generation powered by its own Firefly generative AI model.
So itโs inevitable that generative AI capabilities will become accessible to anyone who wants them.
Scale is critical. You could argue that you donโt need AI to spread nonsense online. We humans are already pretty good at it ourselves, and often all you need is to see a photo of a politician out of context, or even just confidently assert something without evidence, to see retweets and likes racking up.
But this is why I think AI is at risk of drowning us: Because it makes fakery even easier. A photo is more compelling than a tweet โ and simply seeing it glide past as you scroll a news feed is going to make it more believable, as you wonโt be forensically inspecting everything you see.
This scale could have new, emergent, downside consequences for our information ecosystem. If you canโt quite picture it, consider how in the 1990s it was technically possible to make your own website, but it was also hard to do. But it was only with the arrival of social media that posting online became easy โ as a consequence, many more people do it, and because of that scale, a number of new threats, problems and challenges emerged that still cause problems today.
Thatโs why the tsunami of fakeries is inevitable. Itโs like inventing nuclear weapons in a world where every household already has a stash of uranium piled up in their garden. Sure, some people might choose to build a nuclear power generator โ but some people might have more dangerous ideas.
Selling shovels to gold miners
So far Iโve painted quite a dark picture of the future, not least for the perma-migraine Iโm going to have screaming at everyone for posting so much nonsense.
In fact, Iโm actually pretty bullish about the good that generative AI can do too. The negative consequences will need managing, but I think this provides a potentially lucrative opportunity for important new businesses and platforms.
So letโs say you buy my thesis, that generative AI is about to flood our information ecosystem with AI-generated content at a scale we havenโt seen before. Imagine what that looks like in a few years’ time: how will we trust anything we see online?
From an impressive viral dance move, to footage from the frontlines in Ukraine, it wonโt be easy for journalists โ let alone news consumers โ to know what is real, and what was the result of a few taps on an app.
If you thought the 2020 US election was fraught and contentious, just wait until every gaffe is dismissed as a deepfake, and social media is flooded with ultra-realistic looking AI forgeries to undermine political opponents.
The only way weโre going to be able to make sense of what we see is if there is an entire ecosystem of institutions, organisations, and APIs to help information consumers sort the facts from the fiction. And the companies that build them are inevitably going to be valuable and important.
For example, an obvious way to help identify real images would be for social networks such as Facebook and Twitter to display a little โverifiedโ tick against images that have been confirmed as authentic. eBay, AirBnB, or dating apps could find such functionality useful too โ perhaps performing a quick check when a new image is uploaded to restrict the fakes. And maybe weโd even want the โPhotosโ app on our phones to determine whether the photo we saved was real too.
The technical fight against fake news
Technically, this shouldnโt be too difficult to implement. A cryptographic technique known as โhashingโ has been widely used for a long time by authorities and platforms for identifying child abuse and terrorist content. Basically, some clever maths is applied to an image to create a key that uniquely refers to it – and if you โhashโ the same image twice, you get the same key. So you can identify any given image, without needing to store the original.
The tricky part for the torrent of AI imagery is going to be maintaining the database of authenticated images. This is where I see a huge opportunity to sell shovels during a gold rush.
Itโs unlikely that there will ever be one, single, canonical database โ but I can envisage a situation where there are multiple, different databases that add up to a less chaotic information environment.
As such, I see three main types of authentication tool:
Trusted Databases will use the weight of their brand and reputation to convey their authority when determining the authenticity of an image. Itโs a role that โlegacyโ organisations with decades of credibility could perform well in.
For example, perhaps Getty Images or even news organisations such as the Press Association or the BBC could build-out publicly accessible databases of authenticated photos. If a photo is listed as real here, then you can trust that it is real โ just as if something is reported by the BBC, you can broadly assume that due diligence has been done and that the facts are accurate. In exactly the way that you canโt if you read the same information on some random Twitter account or blog.
Crucially, what could make them work effectively โ and provide new revenue streams โ is that they could go beyond simply photos they own or have licensed, and could generalise their authentication credibility to images across the web. This is because, as per the above, they wouldnโt need to actually save or host copies of images (which would be a copyright and licensing nightmare) โ they can simply store the hash, for authentication purposes.
The second type of tool is what Iโm calling Provenance Scrapers, which would work more like Google, spidering across the web and scraping through every webpage, looking for images. Using the information gathered, these scrapers could attempt to determine the provenance of any given image by looking for the earliest possible sighting of it on the web, and allow individual images to be traced around to see where they came from โ like a more sophisticated version of a reverse image search.
Trusted Hardware would be the final puzzle piece. This would be primarily something for Apple and Google to agree on. But perhaps future versions of iOS and Android could somehow store an authentication hash at source, at the moment of creation โ so that it can be proven that a given photo really was taken with the iPhone camera app, and not the result of AI.
Iโm not smart enough to know exactly how such a system would work, as it would be uniquely complicated and necessarily be carefully designed to maintain privacy. Perhaps hashes could somehow be sent blind to databases maintained by Apple and Google, which will effectively provide nothing more than a checksum.
Finally, a use for the blockchain
Now hereโs the wild part. So important could proving the authenticity or provenance of photos and videos become that perhaps it could (whisper it) provide an actual, viable use-case for the blockchain.
This is because of the blockchainโs unique selling point โ that it is an immutable, decentralised ledger, or database, that canโt be tampered with.
So, for example, the hashes of authentic (or authenticated) photos could be written to it, and checked against, and it would essentially last forever. So that even if the authenticating body disappears โ say, the BBC is blown up by a future government, or Getty Images goes bankrupt โ the records will remain in circulation.
The fact that the blockchain is essentially write-only means that, by definition, we can trust that its records haven’t been meddled with. We should then be able to effectively page back through the chain to check the provenance of any images listed on it.
So perhaps the real legacy of NFTs might turn out to be not garish drawing of apes โ but a new way of proving authenticity across the web.
Signal and noise
Whether or not the platforms I imagine emerge or not, one thing we can say for certain is that in the future there is going to be a lot more noise. The ease of generative AI, in the hands of every bot and bullshitter online, means that finding the truth is going to become harder than ever.
But I am pleased to say that there are early signs of the industry reacting. For example, the Coalition for Content Provenance and Authenticity has already signed up Adobe, Microsoft, Intel and the BBC as members. The โTrusted Databasesโ hopefully canโt be far away.
Similarly, the BBC has also recently launched BBC Verify, which alongside veteran open source intelligence analysts Bellingcat is a solid foundation on which to build fact-checking credibility.
There are already companies who have built scrapers for the purpose of things such as copyright enforcement โ which will use many of the same technological building blocks as the envisaged โProvenance Scrapersโ.
So it is still early days, but Iโm convinced: when the AI tsunami floods us with falsehoods, the truth is going to become a lot more valuable. Thereโs a big business opportunity for the companies that build the institutions that will maintain it.
This article has been tagged as a โGreat Readโ, a tag we reserve for articles we think are truly outstanding.
James O'Malley
James O'Malley is a freelance politics and technology writer, journalist, commentator and broadcaster. He has written for Wired, Politico and The New Statesman and was editor of Gizmodo UK. He has also written a number of articles about cutting-edge technology for TechFinitive.com.
To provide the best experiences, we and our partners use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us and our partners to process personal data such as browsing behavior or unique IDs on this site and show (non-) personalized ads. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Click below to consent to the above or make granular choices. Your choices will be applied to this site only. You can change your settings at any time, including withdrawing your consent, by using the toggles on the Cookie Policy, or by clicking on the manage consent button at the bottom of the screen.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behaviour or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.