AI Deepfakes

A drawing of two human heads in silhouette overlooking hovering digital text

Image copyright Adobe Stock

The puffy jacket pope image is an example of a deepfake. These are generated using "deep" learning algorithms. This is where the term “deepfake” comes from. These algorithms can extract “features” (sets of characteristics for how to arrange pixels) from very large databases of real photographs (“training data”) and recombine them realistically in new images. In the case of the Pope Francis puffy jacket image, for example, the machine learning algorithm was given enough images of puffy jackets and Pope Francis that it was able to generate a machine vision model of the visually meaningful characteristics of puffy jackets, another of Pope Francis, and combine both.

There are a number of machine vision and visual synthesis models, ways to train them, and how to optimize them for photorealistic output – this is a burgeoning and extremely active field of computer science that has been buoyed by high-performance graphics cards. In turn, the graphics computing hardware industry has directly benefited from the additional cash influx provided by cryptocurrency trading (which uses the same graphics cards for coin minting), in addition to the usual research, gaming, and large-scale computing clients which have traditionally made up their audience.

A deepfake of the pope wearing an orange long puffy coat, a cross necklace, and a skullcap with his arms extended in front of him.

Pablo Xavier (last name withheld), a construction worker in Chicago, used Midjourney to generate the deepfake pope images.

Emerging tools

Until recently, these machine learning algorithms required significant programming knowledge to implement and create new art with. Now, with online image generators such as Midjourney, AI generated graphics are more accessible to the public than ever before. This coincides with the rise of user-friendly AI text chatbots such as ChatGPT. Midjourney alternatives include:

  • Adobe Firefly
  • DALL·E 3
  • Microsoft Copilot Image Creator
  • Pareto
  • Stockimg.ai
  • ArtSmart
  • Stable Diffusion
  • Microsoft Designer

Because the computer code used by most of these tools is black-boxed and proprietary rather than open source, it is difficult to tell what precisely is different about all of them. This is ironic, as large chunks of the data processing routines they bundle are likely to be open source. This is a common point of contention with AI tools that have political power but whose private maintainers, owners, or shareholders are not held accountable for the consequences of the misinformation produced by their systems.

In many cases, these generated images are humorous, satirical, artistic, etc. In some cases, they are also misleading, which requires the discerning internet user to be vigilant for ever-shifting forms of attention and opinion manipulation. Deepfakes are not only applicable to still images, as equivalents can be found for voice or video (see, for example, this 2016 Adobe VoCo demo, or this fan edit of Star Wars: Rogue One with deepfaked characters to replace the lookalikes for the original actors from the 1970s). Recently, criminal, extortionary or threatening applications of deepfake, AI-powered tools have also emerged (Global News, 2023).

The idea is not new. People have been digitally or manually recomposing pictures for almost as long as the tools to make realistic art have existed, but this is the first time AI tools make this process available to people without technical or artistic expertise. All you need is access to the online tools made available by the relevant software developers. This has wide-ranging consequences for copyright infringement, defamation, public trust, artistic credit, or accurate record-keeping, amongst many other issues. Content providers like YouTube are slowly moving to begin to define the boundaries of the experimentation they'll platform, with an initial set of guidelines published in English on November 14th, 2023 (YouTube, 2023). For now, these seem to restrict themselves primarily to additional mechanisms for flagging, reviewing and user-initiated labelling, more than the platform taking responsibility for detecting and identifying AI-generated content (Hard Fork podcast, 2023).

Cheapfakes

Cheapfakes bypass the costly nature of high-quality deepfakes to similarly mislead with lower overheads. They can be created with standard media editing tools available on any device. These leverage the fact that media is often consumed rapidly, at low-resolution, and with little scrutiny, to mislead or manipulate people with content that is not photorealistic or accurately impersonating a recognizable figure.

The Deepfakes to Cheapfakes Spectrum, a graphic by Data+Society, used under a CC BY NC SA 4.0 international license.

The image is titled “the deepfakes / Cheapfakes spectrum."
There is a banner at the top which reads, next to the title, on the left side (deepfakes): “This spectrum charts specific examples of audiovisual (AV) manipulation that illustrate how deepfakes and cheap fakes differ in technical sophistication, barriers to entry and techniques. From left to right, the technical sophistication of the production of fakes increases, and the wider public’s ability to produce fakes increases. Deepfakes – which rely on experimental machine learning – are at one end of this spectrum."
On the right (Cheapfakes) the banner reads: "The deepfake process is both the most computationally reliant and also the least publicly accessible means of manipulating media. Other forms of AV manipulation rely on different software, some of which is cheap to run, free to download, and easy to use. Still other techniques rely on far simpler methods, like mislabeling footage or using lookalike stand ins."
Under the banner there are a series of cards labeled “technology”. From left (deepfake) to right (cheap fakes) these read: 1) “Recurrent Neural Network (RNN), Hidden Markov Models (HMM), and Long Short Term Memory Models (LTSM). 2) Generative Adversarial Networks (GANs) 3) Video Dialogue Replacement (VDR) model 4) FakeApp/After Effects) 5) After Effects, Adobe Premiere Pro 6) Sony Vegas Pro 7) Free real-time filter applications 8) Free speed alteration applications 9) In-Camera effects and 10) Relabeling / Reuse of extant video." 
Under the technology cards is a an arrow pointing both ways, colored red at left and grey at right. On the left it reads “deepfakes (more expertise and technical resources required)” and right “Cheapfakes (less expertise and fewer technical resources required). 
Under the arrow, each of the cards numbered aboves are connected to techniques and examples. These correspond to: “1) virtual performances; Suwajanajorn et al. Face2Face: Synthesizing Obama [there are 2 screencaps of Obama giving a speech from the Oval office] 2) Virtual Performances; Mario Klingemann: AI Art [there is a composite digital image made of many pictures attached] 3) Voice Synthesis; Posters and Howe’s: Mark Zuckerberg [there is a picture of Mark Zuckerberg in a living room] 4) Face Swapping; Gal Gadot, not pictured because of image content) also connected to 4) Lip-Synching; John Peele and BuzzFeed: Obama PSA [there is a picture of Obama with a US flag] 5) Face Swapping: Rotoscope; Huw Parkinson: Uncivil War [there is an image of Hilary Clinton’s and Donald Trumps’ half faces glued together, with the text in the image reading: Presidential Avengers: Uncivil War as a parody of the Avengers: Civil War movie poster] 6) Speeding and Slowing; Paul Joseph Watson: Acosta Video [there are two pictures of a debate with people in suits in a draped room side by side for comparison] 7) Face Altering/swapping; SnapChat: Amsterdam Fashion Institute [there are three portraits side by side, one of someone with short hair and purple coming out of their mouth, one of someone wearing large blue flight goggles and one black and white with someone opening their mouth wide open and contrast shifted to make their eyes look big and dark] 8) Speeding and Slowing; Belle Delphine Hit or Miss Choreography [picture of someone with long pink hair, white gloves and tight tshirt smiling posing hand on chin] 9) Lookalikes; Rana Ayyub (not pictured because of image content) 10) Recontextualizing; Unknown: BBC Nato Content [there is an image of a BBC news cast studio]

Click on the graphic for a detailed description of the spectrum from deepfakes to cheapfakes.

Resources

References

Hard Fork. An A.I. Pin Drops + Youtube's Take on Deepfakes + a Lab-Grown Thanksgiving. Podcast. November 17 2023.

Mannie, Kathryn. AI Kidnapping Scam Copied Teen Girl's Voice in $1M extortion attempt. Global News. April 18 2023.

YouTube. Our Approach to Responsible AI Innovation. Blog Post. November 14 2023.