Qiu
Let me start with a quick introduction. Welcome to the Wuzhi Society's March Wuzhi Reading Room, which we are co-hosting with the Photography Theory Shelter. First, a few words about the two communities. The Photography Theory Shelter was founded in 2024. It is an open community that starts from discussions of photography theory but is not limited to them. The Wuzhi Society was founded in 2021 by Penny, who holds a PhD in art history. It is an interdisciplinary exchange community for researchers around the world. We hope to build a community where people learn from one another and share together, and also to create an open, mutually supportive interdisciplinary think tank. We welcome researchers from all fields and at every stage of their careers. Every member can discover new knowledge, find support from peers, and share their own insights and experiences. The Wuzhi Reading Room is our monthly reading group. Each session, we invite a PhD or graduate student from a different discipline to recommend books and films on a particular theme and guide everyone through interesting knowledge from different fields, broadening the ways we look at the world.
Today's session leader is Prismo (Guosen Chen), a PhD student in philosophy at Fudan University and a co-founder of the Photography Theory Shelter. He recently gave a talk on AIGC images at our spring Prism event. In today's Reading Room, Prismo will lead us in exploring how AI-generated images challenge the nature of photography. You are very welcome to type in the chat during the talk, and we will set aside some time for discussion afterward.
I'll hand it over to Prismo.
Reading List

Thank you, Qiuqiu. I'll get straight into today's talk. It's really an attempt at a guided reading, and I'll share some of the texts we picked for this month. We originally chose four books. One of them is Art, Artificial Intelligence, and Creativity, an edited volume. Its content has little to do with photography. It deals with the relationship between AI and art in a broad sense. I especially highlighted a few essays in it, such as Karen Barad's "Posthumanist Performativity."
The second book deals with photography more directly, but it isn't directly connected to AI. It was published in 2017, before today's generative AI existed, or at least before it had much of an impact. In 2023 the author published another book called The Perception Machine. That book does discuss things like GPT, but perhaps not very directly, because it builds on her earlier Nonhuman Photography.
I had planned to say something about the second book today if time allowed, but I doubt there will be enough time, because there's quite a lot of material. So today I'll focus mainly on the third item, the latest issue of the magazine Aperture, the Winter 2024 issue. It is about AI photography and covers a very wide range of topics, and most of the articles are fairly short. I think that makes it easier to discuss, or for you to read later. The articles are also not that hard to read. They are less academic than the second, first, and fourth books. The second and fourth books also have fairly distinctive writing styles, which can make them harder to get into.
I recommended that first book, especially the first three chapters including the introduction, because I want you to think about this question: can so-called creativity be used to defend the originality or authenticity of human photography? In fact, it can't. The author of the second book, Joanna Zylinska, is coming to Shanghai to lecture in a couple of days, and her announced title is "nonhuman creativity." I can roughly guess what she'll say. She'll probably use photography as her entry point to show that, from the very beginning, photography's so-called creativity has had nonhuman qualities and nonhuman components, going back to Talbot's The Pencil of Nature. Human and nonhuman creativity are not opposed to each other. That's why I chose this book.

To address the current situation more directly, the AI photography that is closest to us now, we need to look at Aperture. The second, fourth, and first books don't touch on this directly. Some generative AI can now produce images that are indistinguishable from real photographs, and those books don't yet deal with that problem. The Aperture issue does engage with it to some extent.
The fourth book is about invisible labor, invisible digital labor. For example, its first essay argues that "Google Street View drivers are people too." Companies like Google are actually very harsh toward the workers who drive around cities doing surveys and mapping for them. The 3D maps and panoramic maps we see now are all collected by people driving special Street View cars around the streets. Interestingly, when they recruit workers, they don't use Google's name. To avoid attracting attention, they hide who the real employer is. Sometimes they also put out notices telling workers that the job may give them the chance to travel elsewhere to do this kind of surveying and map data collection, at company expense, a paid trip in a sense. But it's a little trick. By the time workers actually hear about it, the sign-up deadline for the project has already passed. Google doesn't actually want these workers going elsewhere to map for them. Google doesn't give them proper insurance or standard benefits. It hires them because it values their lived experience and familiarity with their own city. So it would never send them out to map other places, because that would make no economic sense for Google. That's an example from the first essay. To some extent, whether we're talking about AI or earlier digital images, the digital labor behind them is often hidden.
This issue comes up in the Aperture articles I'm about to discuss, especially in Crawford's Anatomy of an AI System. If any of you attended the Prism event, Zihan talked about this project then. So that's roughly the situation with the four books. Today I'll organize my talk by article, because each Aperture article has a somewhat different focus. Instead of picking a single unifying question, I've selected several of the more important articles, in the order they appear in the magazine, and I'll say a little about each.
1. WAYS OF SEEING: Trevor Paglen’s art unravels the politics of computer vision

The first article is about Trevor Paglen, whose image appears on the cover of this issue. He started thinking about AI, and how artists use it, very early. His work isn't only about what kinds of images generative AI can produce. His question is this: if photographs are no longer shaped by human seeing or taken by humans, and instead become tools automatically generated by machines and used to make judgments, how do we understand images under those conditions? And how might we even learn to see the way machines do? He has been working along these lines for more than ten years. So in a sense, he is concerned with the conditions that make this kind of generation possible before any image is generated.
In that sense, one very important part of his work is what he calls "dataset archaeology." The name comes from a project Paglen did with Crawford, the author of a later article in the issue, "Metabolic Images." The name may have been Crawford's. I introduced it at the Prism event. It means digging into the material that AI systems need in order to be trained, namely datasets. To people unfamiliar with them, these datasets may seem very neutral and so-called objective. But Paglen made a project, ImageNet Roulette, showing that such datasets are full of cultural biases, gender discrimination, and even violence. The Aperture article mentions that Paglen had viewers upload selfies and then used AI to classify them. The machine recognition used here is essentially what real-world surveillance systems rely on to judge people, such as the machine vision systems with recognition functions you encounter at borders, checkpoints, or major transport hubs. That is his first line of work: tracing the biases and artificially shaped preferences hidden in the image datasets needed to train generative AI.

His other line of work is simulating, or interrogating, how machines see. Take this screenshot, from a video work. It shows viewers a string quartet performing a Debussy string quartet. As the music goes on, the picture is gradually replaced by various AI analysis results. In the process, it shows you the difference between human vision and machine vision. The image slowly turns into quantifiable frameworks from which features have been extracted. For example, labels at the top give the person's age and emotion. In this image, it shows someone aged roughly 25 to 30, labeled "sad."
The judgments these systems make can sometimes seem absurd, especially the ones shown in the project I just mentioned. But for Paglen, the key issue isn't whether the AI's judgment is accurate. Accurate or not, these judgments and their results are already used as tools to manage people's labor and as algorithms to secure so-called public safety. All of this has already been quietly built into the infrastructure of our daily lives. So even if a judgment is wrong, it has already happened and produced real effects in real life. In this sense, these images resemble what Harun Farocki called "operational images." Such images are not made for people to look at. They are tools, produced within real, machine-readable systems, and as they circulate through those systems, they have concrete effects on people.
Paglen also rarely takes traditional straight photographs. He often uses drones, explores satellite mapping, and uses AI to generate strange, "off" images. One of his works is called Adversarially Evolved Hallucinations, made several years ago. He trained the AI himself and had it generate all kinds of images, such as vampires, or images that look somewhat abnormal.
"Paglen uses the camera less for depictive purposes than as an analogy for the extraordinary labor it takes to make something visible—whether that thing being seen is a surveillance drone, the routine computational analysis performed on photographs of our faces, or a psychological-warfare strategy that manipulates our media consumption." In this sense, Trevor Paglen is redefining what photography does in our time. It is no longer the traditional act of a living person holding a camera and photographing another living person. In fact, a more mainstream, or at least rising, view holds that photography is better understood as a kind of event, or even a "potential agent." Whether or not you take pictures, whether or not anyone picks up a camera, photography surrounds us, envelops us, and acts on us. In many ways, how people perceive the world now, and how they are positioned in social life, is governed by images that humans don't produce. This is a view Zylinska, the author of the second book, puts forward in her 2023 book The Perception Machine. The view has many sources, and she isn't the only one making it. The Israeli scholar Ariella Aïsha Azoulay, for example, takes a similar view in her book Civil Imagination. Western photography scholars now tend to see photography as an event that encompasses life and acts upon our lives. It is a potential, ubiquitous event or practice, rather than an activity defined by a single act of taking a picture.

Paglen treats photography as an analogy, a system of analogy for understanding how these things actually operate in our real lives and how they act on us. The work on the left is a video in which he makes something self-contradictory with respect to truth and falsehood. The subject was an officer in the US Air Force, though he no longer was by the time Paglen made the work. In the video, the man admits he spread a great deal of disinformation about UFOs, but later says that UFOs are actually real. Paglen uses this self-contradiction to talk about so-called hallucinations, or false information. Whether it's true or false doesn't really matter, in a sense. What matters is how false information produces real results. Or take a hallucination, in whatever sense you use the word. It might mean hallucination in the sense of truth versus falsehood, or the hallucinations that AI tools themselves produce as we use them. Either way, it can have the same consequences as true information and drive you to do something. That's why he says in the conversation that, in a sense, all photographs are UFO photographs.
That's the first article. It's a bit long, and the question it leaves us with is this. We may have reached a point where we need to recognize that photography is not necessarily an act performed by a human subject. Often it isn't even conditioned on, or aimed at, direct human activity, and it doesn't necessarily seek to represent a person in the image. In that case, how much force do the older, more immediacy- and meaning-oriented ways of identifying photography still have in helping us understand photography today? For Paglen, the image, or photography, is not a technical question. Before technology, what should be considered is politics, a political question. That political question determines how images are used, how they are defined, and how they are classified. That's roughly what the first conversation covers.
2. GENERATIVE STYLE: What FSA photography can teach us about AI

Next, the second article is about "generative style." It isn't about whether images are true or false either. On the left, the authors show AI-generated photographs in the style of Dorothea Lange. During the Great Depression, Lange was as renowned as documentary photographers like Walker Evans. Looking at the photographs alone, it's almost impossible to tell whether they were generated by AI or taken by Lange herself. So the Lange-like feel of these images doesn't come from someone imitating her style with a camera. It comes from a model the authors fine-tuned themselves. The authors call this quality "latent specificity." A model can be trained to generate images close to a particular photographer's style, without any image duplicating an existing photograph of hers. The features the AI extracts are, in a sense, latent features that humans can't access directly, hence the word "latent." Yet the data the AI has learned and its abstracted weights are not photographs produced directly by the reaction of light in the traditional sense. In a sense, the output is a visualization of data from a neural network. When we generate photographs by this means, all that remains of photography is style, a style that can be simulated and mobilized. In that case, the mechanism for producing images is no longer what we used to call an indexical mechanism. The indexical mechanism comes from the traces the world leaves on the camera. AI-generated style instead extracts latent features from existing images, features humans may struggle to extract directly, and recombines and reproduces them.
What follows from this? In a sense, I think it actually reveals something about photographic practice. If you create or shoot only to achieve a certain style, then photography really will "die." If you pick up a camera purely to take pictures that look like some master's work, to copy Webb's style or Cartier-Bresson's, then perhaps its only remaining meaning is the meaning of the experience itself. The experience of the act of "shooting" is, of course, meaningful to you. But the photograph you end up with may not be as good as what AI can produce. So in this sense, you find that photography doesn't end at producing an image. We often equate "photography" with the photograph. But a photograph is only an intermediate result within the photographic process. It cannot use its own meaning to replace the meaning of the whole photographic event or process. They are two different things, and this distinction existed long before AI. AI didn't create it. It has become clearer because AI can seamlessly generate the styles of certain masters, and can even imitate your own photographs and generate something that looks like you took it. When this happens, we see that the meaning of photography is largely not determined by the meaning a single photograph gives off. I think that's the insight this very short article on "generative style" offers.
3. METABOLIC IMAGES: Is there time to rethink artificial intelligence?

Next is Crawford's "Metabolic Images." What does "metabolic images" mean? It means the process of generating these images is like the body's metabolism. First something is eaten. After digestion, large pieces of food are broken down into small nutrients, which the AI model absorbs, and then they are expelled, like excretion. What we see is the final, excreted result. But in this article, "metabolic images" refers more to ecological issues, questions of ecology and resources.
I think the article reminds us of something. We often hear people ask what would happen if AI took over humanity or got out of control: the mad robot, the mad AI. But before discussing that, according to Crawford, we need to recognize that AI today is first and foremost a resource-intensive infrastructural technology. It cannot decide on its own "madness." It needs huge data centers, vast amounts of electricity, enormous energy consumption, huge amounts of water, and mineral resources. In this sense, the natural resources AI consumes are also part of the "nutrients" in her metabolic images. This production is not transparent to us either. Few commercial companies are willing to talk about the energy or water they consume to train their models, though they do like to talk about how many GPUs they used. So Crawford argues that before image generation is an aesthetic or cultural question, it is an ecological question, a question of everyday conditions. That's the main point of her "metabolic images." On one hand, it extracts natural resources. On the other, it extracts all kinds of data from the internet, breaks it down into color labels, turns it into a model, and spits it back out.
In this sense, the "metabolic images" metaphor is very biological, and what lies behind it is also very material. It's close to the example at the beginning of the fourth book on the reading list. The fourth book is mainly about labor: how the labor in images is erased, made invisible, and disguised after the image is produced, as if that labor never existed. I think this corresponds to what I've just described. The AI-generated image is not a transparent thing either. It has costs, and it consumes natural resources.
4. GHOSTS IN THE MACHINE: Inside the bizarre world of AI copyright law

Next is a short article on AI and copyright. In the past couple of days, ChatGPT updated its image generation feature, and many people have been using it to generate images in the style of Studio Ghibli, including Altman himself. I don't know how Miyazaki and his colleagues feel about this, but some people think it's rather unethical and may carry copyright risks.
The author of this article has a legal background. The question he discusses is how we should determine authorship in an age when AI images are appearing everywhere. If AI generates a picture, should the author be you, or the model you used? He gives an example, the photograph on the left. In 1884, the photographer Napoleon Sarony photographed Oscar Wilde, and he successfully obtained copyright for the photograph. The court ruled that it was an original work. It held that the photograph was not simply a mechanical reproduction, because it clearly reflected Sarony's creative choices. He arranged Wilde's pose, controlled Wilde's expression, and set the light source. So it was his creation. This precedent on photographic copyright is still followed today. The US Copyright Office, in handling AI-generated works, still seems to follow this line: if a work reflects a person's aesthetic judgment and creative acts, it can be considered that person's work. The problem is that this precedent seems less and less applicable to generated images today.
He then gives several cases, such as the Urantia case from the 1990s, and the "monkey selfie" on the right. A monkey pressed the shutter and photographed itself. Does the copyright belong to the monkey or to a human? The court said copyright can only belong to humans. Even though no human intervened, and the monkey simply pressed the shutter and photographed itself, a monkey cannot be an author. If you read the original article, it also gives an example of someone who used AI to generate an image in the style of Van Gogh's The Starry Night and applied for copyright, and the Copyright Office refused. The style-blending process, transferring Van Gogh's style onto the image, was done by AI, and the applicant did not directly control the image's generation, so it wasn't allowed. This standard is very awkward. In many cases, humans really are participating in the input for these images, for example by writing prompts. But the law does not recognize prompt-writing as creative, because the legal standard for creativity still dates from the nineteenth century: the artist's total control over their work. I think that standard has increasingly become a fiction since AI appeared.
Long before AI, some scholars and photographers had already pointed out that the values behind this notion of copyright, and even this notion of creativity, amount to a kind of "possessive individualism." Legal copyright determinations further reinforce the photographer's status as the sole owner and sole beneficiary of the photograph. The problem is that photography is not something the photographer does alone. In this photograph of Wilde, for instance, Wilde's own cooperation and his own image were extremely important to how the photograph came to be. So Azoulay, and many scholars of the new generation working in visual culture, have begun to argue that we can't treat the people who appear in images simply as subject matter, as a motif. They are subjects in their own right, no less than the photographer. Photography is a process completed jointly by multiple participants, and it can't be anchored to a position determined by the photographer alone. I think this, too, has become more and more apparent since AI appeared.
The author also argues that our current copyright theory can't handle these issues because of the way Americans approach them. They rely too heavily on precedent. In the end, whether AI images can be copyrighted depends on courts drawing analogies from earlier precedents to see whether there's a definite answer. So the author doesn't give a particularly clear answer about how we should decide copyright for AI images. He simply lays out the problem.
5. DEEP REALS: What is realism in the age of AI?

I find this next article especially interesting. It's called "Deep Reals." In Chinese it might be translated as something like "deep realness" or "deep reality," because it corresponds to a now-familiar phenomenon, "deepfakes." His question is simple: when images from generative AI are everywhere, how do we actually know whether a photograph is real? Notice that the earlier articles hardly discuss whether images are real. Paglen simply sets the question aside and looks at the operational effects of images. This article takes it on. The passage is a bit long, and I'll read it out, because I think he's already written it very clearly. He puts it in a very conversational way, so it hardly needs more explanation from me. He says:
The phrase "real photo" is self-contradictory. It makes sense not because it describes the ontological status of an image but because it expresses a judgment—about the image's maker, the technology they used, and, ultimately, their intentions; as if "real" meant excluding any explicit purpose or human perspective. In a sense, the "artificial intelligence" in text-to-image models works precisely this way: by averaging billions of indiscriminately ingested images, it produces a statistically based, expressible conceptual image, thereby extracting some essence of the "real" from the contingencies of experience and the biases of subjectivity. Yet it is precisely evidence of subjectivity that usually allows us to regard an image as "real"—because it indicates that someone once saw something in a particular way and found a way to share that seeing. Generated images undermine our sense of reality not necessarily because they look "too much" or "not enough" like what they depict, but because they lead us to mistake them for images that have a perspective—when they are really just synthetic representations.
—— Aperture, 2024 Winter, NO.257, 87-88
If a "real photo" just means an image not generated by AI, then what is it? Does it mean that any image without editing or post-processing counts as real? The problem is that when we call an image real today, it's often precisely because it reflects someone's perspective, because a living person went there and took the picture. Otherwise, if we held that a "real image" is one with as little human subjective intervention as possible, then AI images should be the most real of all. Behind this confusion lies an assumption: that whenever technology intervenes, we should doubt an image's truthfulness. The idea is that if a photograph hasn't been technically processed, we can trust it. This view of truth rests on a very simple premise. If technology intervened, such as AI generation or post-production tampering, we say the image is fake. If there was no technical processing, we call it real, however it was selected. That holds even when a photographer deliberately uses framing and composition to exclude content that might prevent a misunderstanding, content whose absence can reverse how the photograph's meaning is understood. Because there was no post-processing, we still call the photograph real. This view obviously doesn't hold up. The author also thinks this assumption no longer stands at all.
He then gives several more examples, such as the "AI-generated content" labels I just mentioned. When Meta labeled AI-generated content, it tagged not only AI-generated images but also many photographs that had been post-processed. So it becomes a kind of irony. If you're going to do it that way, why not label every photograph as not being a true representation of reality? And that would actually be correct, because no photograph corresponds to reality in a complete, all-encompassing sense. The author says that when this happens, we as viewers assume the platform will judge for us whether content is AI-generated, and that if a photograph is so-called fake, the platform will flag it. Then we become mentally lazy. We stop checking an image's source, its context, and its intent. As a result, we no longer go through the complex process of interpretation, perception, and judgment we used to go through when encountering a photograph.
He draws an analogy with another book, a book by Ritchin. When Ritchin wrote it, Photoshop had only just appeared, and the traditional belief that "a picture is proof," that the photograph is evidence of fact, was already being shaken. Now, of course, the shaking is far more thorough. Ritchin has just published a new book called The Synthetic Eye, which pushes the question further.

After that, the author pushes the phenomenon one step further and turns to another question: how to understand "realism." The categories of realism, both in the sense of verisimilitude and in the sense of reality, have changed, and he proposes a new dimension for rethinking them. When we used to say a photograph was real, we seemed to mean that it recorded facts that happened in reality. But now, when a photograph is taken to be real, it often seems to be because it moves us emotionally and helps form some consensus or shared belief in social life and in circulation. He uses this to draw a distinction. "Deepfakes" use computational means to create a fake image and make it look like something a person photographed or could have made, in order to deceive viewers. Its logic is that you take the deepfake as a record of reality, a record of fact. If you trust it, it has succeeded. So it has to hide the traces of the algorithm and pass itself off as a photograph or image made by a human. "Deep Reals," by contrast, are images that really were photographed, not made by AI, such as images of natural disasters or major news events. But because these photographs are so typical, so polished, so lacking in the randomness of real life, you begin to wonder whether they were made by AI. His example is the photograph Evan Vucci took when Trump was shot. Its composition closely resembles the Statue of Liberty, and it looks rather like an AI rendering. You know it's real, but intuitively you still feel it's too perfect, as if AI had generated it. So "Deep Reals" describes how, in a climate where all sorts of fabricated images circulate, photographs of real-world facts can end up being treated as deepfakes. The sense of realness they depend on has moved far from the indexical logic of earlier images. We don't believe an image is real because it records something. We judge it to be real because it looks like things we've seen before and are willing to share. "Deep Reals" describes a reconstruction of our sense of realness. AI has changed our sense of what counts as a real image. Whether or not an image is real, when you judge it, AI images are already part of your frame of reference. This means you can no longer look at anything, even so-called real things, apart from the visual logic of generative AI images. This is a new variant emerging around the category of realism.
6. THE SIMULATED CAMERA: Documentary photography

The next article is in some sense connected to the previous one. I think there's a lot of continuity in the questions they discuss. Ritchin's conversation largely lays bare how we judge true from false, but it doesn't directly ask which images can or should be judged fake. Instead, it asks whether the effects that traditional real images produced have changed, or been altered, today. For Ritchin, photography's role comes from a mechanism that allows us to share a common understanding of reality, what he calls "shared reality."
The core question of this conversation is this: now that generative AI is everywhere and deeply embedded in world culture, is there still a need for documentary photography that relies on traditional straight photography? And if so, why? From the very start of the conversation, Ritchin sets aside clichés about "the death of documentary photography" or "the death of photography," since these have been around almost since photography was invented. Look at how critics have written about photography's history, and photography seems to die every ten or twenty years. Yet it still hasn't died. "The death of photography" has just become a new cliché.
What Ritchin cares about is whether we can still rebuild trust in photographs as evidence. As he makes very clear, photographs, or photography, have never been important because they represent some natural fact. They matter because, at certain crucial moments in history, people have used them to make visible things or events that everyone needed to see, and that visibility provoked a collective response. That's why they matter. We can jointly confirm that, in the world we live in and share, certain things really happened, things none of us can deny. Based on that belief, we come together, and we gain a belief in and understanding of the "reality" we share. He gives some examples. This one is the 1972 image of the young girl Kim Phúc, burned by napalm, running. Another, not shown here, is the 2015 image of a little Syrian boy lying on a beach.
These photographs matter to Ritchin not because they show suffering or the cruelty of war. They matter because they give all of us a reference point to discuss and argue about together, and even to respond to collectively on the basis of the situation the image presents. Consider Eddie Adams's Saigon Execution. As it circulated, its meaning was distorted in a very intuitive way, and the police chief was treated very unfairly as a result. Yet in a sense, the photograph also gave people a shared reality. It sped up the end of the Vietnam War. You could say it changed the course of the war. This is why Ritchin defends documentary photography. Today, many photographs or images don't need a camera at all to be produced, and don't even need an audience. With AI, you just write a prompt and immediately get something that looks like such a photograph. But it cannot give us what traditional documentary photography does: an understanding of what our shared reality actually looks like. So under these conditions, the meaning of the word "photograph" itself has become blurred. Behind this photograph, or this thing that looks like a photograph, is there a verifiable act of photographing, a real event that was photographed?
Ritchin also specifically points out that many media outlets now use expressions like "AI-generated photographs," which he considers a sleight of hand. I basically agree, though I often can't avoid using the phrase myself. The word "photograph" should be tied to the specificity of a scene, to some mechanism of witnessing, and to an indexical logic, as the previous article put it, the idea that someone went to a place and took the picture. If we stretch the word "photograph" too far, even using it for AI-generated images that merely look like photographs, then photography's public mechanism, the mechanism by which it builds our understanding of shared reality, is at risk of collapsing.
On this basis, he asks a question. If every image can be generated, as we saw with Dorothea Lange's photographs being so easily imitated, how do we defend or redefine "documentary photography" today? He uses a very old framework, John Berger's idea of "quotation." I think this is quite similar to the now more common way of defining photographs through the "index." What is an index? If you see smoke, you know there's fire, so smoke is an index of fire. If you see footprints, you know someone walked here, so footprints are an index of feet. There is a clear physical causal link between them. If there is a photograph, there must have been light, and that light must have fallen on the photosensitive surface through the lens. By extension, someone really was standing in front of your lens.
Ritchin cites John Berger's idea that photography is actually a "quotation" of the appearances of reality. The lens is the quotation marks, and what's photographed is what sits inside them. The scope of the quotation marks must be clear. Any construction that goes beyond that boundary, however good the intentions, cannot be treated as documentary photography. This sets a very clear boundary: it must come from the "appearances" of reality. He stresses appearances because we can't say anyone has photographed the "essence" of reality. I think that idea is absurd, and so he specifically stresses that it must be "appearance," a quotation of appearances. Ritchin's emphasis on norms and on the boundaries of quotation marks comes from an ethical demand. If documentary photography is to play a role in social and public life, or if we still want it to, we have to defend its ability to provide this shared reality. If that foundation, that ability, can no longer be trusted, public discussion may fall into a crisis of trust.
At the same time, he doesn't treat AI as the enemy. He offers some fairly constructive examples. Consider people detained in war, or photographers who for some reason can't reach a war zone, when violence has really occurred, such as the abuse of prisoners. He accepts that in such cases, AI can be used to generate images that visualize moments that happened but that no one witnessed. AI can also do technical work to fill in what couldn't actually be photographed. But the precondition, Ritchin says, is that you make the scope of the quotation marks clear. You have to explain how the image came about and provide enough context and background.
In the second half of the conversation, Brian Palmer asks Ritchin a very pointed question. If AI can now fake any image, simulate any scene, and replicate any face, how can we possibly tell whether a photograph is reproducing some past fact or creating something based on biases? Palmer gives an example. Once when he was in Iraq, an editor at a news magazine called to ask whether he had photographed a bus bombing that had just happened. In fact, it would have taken him more than eight hours just to get to the site. The magazine asked whether he'd gotten a shot. Of course he hadn't. But he could have generated one with AI. When images become something determined by demand, look at what happens. I really can't get to the scene, but this photograph needs to exist. What does it become? What matters is no longer whether the photographer was there. You might not even care whether the event should have happened. What matters is finding a photograph worthy of the scene. Palmer says this is the moment when documentary photography faces its greatest test. It may be drawn entirely into a logic of demand-driven images, no longer concerned with facts and causes, but mainly with the effects it can produce.

This approach has in fact shaped how we look at the entire Global South. Race, class, and colonialism have always influenced how Western mainstream media depict the Global South. The scenes of the Global South they present often simply fix people, again and again, in "natural" symbols such as refugees, the poor, and victims. Even when these photographers go to Palestine, what they photograph are people living in constant fear under Israeli occupation. Because the photographers mostly come from the Global North, they assume Palestinians must live every day hungry and cold, so those are the images they take. This means that not only is the content of the photographs uniform, but the narrative structure they rely on is also one-directional.
Palmer then says that because this structure starts from demand, it fundamentally undermines documentary photography's potential to contribute a new understanding of, or renew, the mechanism of shared reality. Ritchin agrees with Palmer. Technological progress doesn't mean the right to look has been redistributed. It doesn't mean there has been any change in who gets to define how we look at particular groups. Technological progress alone can't make that happen. Real breakthroughs often come not from new technology, but from changes in how the power over image production is distributed.
That's why he specifically mentions a photographer, whose work I'm showing now, Occupied Pleasures. What's distinctive about this work is that the photographer doesn't depict Gaza as the familiar scene of suffering. It shows humor and joy in everyday life. Take the photograph on the left, which appears in Aperture. A man is sitting in his car smoking, with a sheep beside him, because he's taking it home to celebrate the upcoming Eid. The one on the right is even more typical: local women doing yoga. The photographer doesn't portray Gaza as the endless suffering we imagine. She finds many concrete joys and happiness, some small and some more significant. This refuses the mode of presenting Gaza as an abstract space of victimhood.
When Azoulay discusses images of Palestinian victims, she notes that some people object to fitting refugee images into certain visual templates, such as the Pietà, having so-called Palestinian refugees pose as the Pietà in front of the camera and then photographing them. Some have condemned this as a kind of bullying of local people. But Azoulay disagrees. The mother consented to being photographed with her child in this way, and the force of that consent can't be ignored. In a sense, they needed to be seen in this way. Azoulay also says that the way Palestinians in photographs look directly into the lens, returning its gaze, is a declaration that they should not be permanently fixed in the label and identity of refugees.
I found this an interesting point when Ritchin turned to the Global South near the end of the conversation. Real democracy in the image ecology doesn't mean democratizing AI so it can imitate images from everywhere equally. It means enabling those places, and the people who live there, to become the producers of these images and to speak through them. The key question is whether we have a mechanism that allows the Global South to decide for itself how to represent itself. If that structure doesn't change, however realistic AI becomes, it will be of no use.
7. CREATION MYTHS: Why we need a new image criteria

The last article I'll discuss is "Creation Myths." It introduces a key term, "negative images," meaning the invisible things used behind generated AI images. I think her approach is somewhat similar to Paglen's. Paglen does "dataset archaeology," while Nora N. Khan talks about "negative images." Almost everyone sees generated images every day now. You know they're not photographs, yet they look like photographs. You know they were generated by machines, but we don't know how to judge whether an AI-generated image is good or bad. In this era, many images may not be made for humans at all. Humans can't see them, but algorithms and machines can read them. In an era when images are changing like this, can traditional methods of art-historical criticism still be used? Hal Foster raised this question earlier. The author takes a different approach. She argues that what we need now is a new critical capacity. This capacity isn't about analyzing how an image looks on its surface. It's the ability to analyze what lies hidden in the negative space beneath the image.
We're puzzled by AI images not because they're too unfamiliar, but because they're too familiar. They look too much like things we've seen before. They're so familiar that they keep reproducing all our existing aesthetic formulas. Their apparent novelty is just a mixture of many things that already exist. But how exactly is the whole body of training data behind them used? How is the model trained? How is the data filtered? Which faces made it into the dataset? These things determine how AI sees the world and how it uses that worldview to generate images. So she argues that properly viewing an AI image must also mean viewing it within a critical system. We shouldn't look only at the output, but ask how it was generated. Hence the concept of the "negative image." Every AI image has a "negative image" that we can't see, one that isn't directly visible. It includes the original dataset, the model's training process, and the model's specific ties to particular institutions and systems of interest. These make up the negative side of AI image generation, the shadowed side that light doesn't reach, the inverted side.
I think this is a very useful perspective. Say a diffusion model learns what a "person" should look like from millions of images from the nineteenth to the twenty-first centuries. If you ask the AI to generate an image of a person without saying what kind of person, what will that person look like? By what standard does it carry this out? What kind of person should serve as the baseline for the image of a "person"? And what kinds of people are filtered out, even during dataset annotation? For example, people with facial deformities may never appear in these datasets and in the AI's output, yet many such people exist in reality. There is no place for them in the standard portrait the AI ultimately generates. This is actually the result of a choice. Everything excluded in this way is the "negative image" of the final image. We can't look only at the final output and ignore the whole process that produced it. When we look at AI images today, we can't judge them only by whether they look convincing, or only by whether they're real or fake. We have to gradually learn a new capacity, a new critical capacity: reading the "negative image" within images and recognizing the invisible layers behind them. We should ask who trained the model, whom it serves, and what values it rests on. Without an awareness of the "negative image," it's easy to treat AI images as neutral, or even so-called objective. That's the concept she introduces in this very short article.

Finally, I'll share a passage from Ritchin to close the talk. He says:
We are already in a different era. "The camera doesn't lie" is hard to sustain today even as a myth. [……] We must explain more clearly to the public the differences between different kinds of images, including how they are made and by whom. We need a broad movement to help rebuild a shared sense of reality, while in the process greatly enriching and deepening our understanding of the concept of "the real."
——The Simulated Camera: A Conversation with Brian Palmer, Fred Ritchen, in Aperture, 2024 Winter, NO.257, 101.
Q&A
Qiu
Thank you, Prismo. Two people asked questions in the chat.
First question: From a semiotic point of view, could we say that photography's relatively direct reference is actually quite recent and unusual, and that AI-generated images imitating a certain style fit better with how signs have long referred?
Answer: I don't think that's quite right. You're discussing the indexicality of AI images versus traditional documentary images and asking which is more real, deeper, more direct, or more authentic. Posed on its own, that question doesn't hold up. Any indexicality only works within an interpretive framework. That framework includes what we call "collateral knowledge," the knowledge you already have before you read a photograph and draw on when you encounter it. All of this, including the context that comes with the photograph itself, forms an interpretive environment. Discussing how a sign refers only makes sense within that interpretive environment. No particular kind of reference is more essential, more traditional, or more authentic than another. And in general, I think it's wrong, semiotically, to say that photographs simply have an "index" property. Any image can have indexical properties, even a painting. A painting can point to the painter's hand, and the hand in turn points to something else. Indexicality always depends on the conditions under which you interpret it. That's my answer.
Qiu
Thank you. The second question in the chat is: Can photography really be equated with pictures? Photography seems to be a technique or process of recording, whereas a picture might be a result, or perhaps only the picture has meaning.
Answer: First of all, photography and the photograph are definitely different. "Picture" is a broader category. It includes what you see in magazines, and paintings count as pictures too, but that's different from an image, and a photograph is definitely different from photography. Many papers today confuse the concepts of photography and the photograph and treat them as the same thing. I think they're completely different. Photography is photography, and a photograph is a photograph. Why? What kind of existence does photography have? When you're driving and hit a red light, why don't you dare floor it and run the light? Because you know there's a surveillance camera overhead. When you're doing 120 and slam on the brakes at a red light, you've taken action. But does a photograph appear? No, because you followed the rules and didn't trigger the camera. Yet in this process, does photography have an effect? Yes, it does. So in this case you can't say that only the final photograph or image has meaning. Photography's very existence becomes a potential event that is everywhere in your life. It isn't limited to a single press of the shutter. That's how I understand it. So in this sense, photography and the photograph must never be conflated.
Rong Jiang
Thank you very much, Guosen. I think it's great that he spent so much time sharing and summarizing the articles in last year's final issue of Aperture. And the question he raises is itself very meaningful: "When the shutter becomes a generate button, is it still photography?"
Having listened, I have a few thoughts to share. First, I noticed that Aperture is very "consistent," using the word "image" throughout. I believe Guosen knows why I keep stressing the importance of this word. Our language, the way we speak, actually represents our thinking, and even our slips of the tongue and of the pen reflect our thinking. I noticed that although Guosen kept using yingxiang ("image") during his talk today, he also used tuxiang ("picture") at some points. So I want to stress again how big the difference is between tuxiang and yingxiang. A tuxiang, a picture, can only be drawn. As Guosen just said, it's like the illustrations in magazines and newspapers. It originally comes from painting. What AI generates now is yingxiang, an image. Even if we have AI software generate something that looks just like a Van Gogh painting, if you think carefully, it's still an image that looks like a Van Gogh, because its medium is the image. It isn't a painting. But it can't be entirely photography either. Whether it is photography is something we can keep discussing. Personally, I think "photography" is a very large container. It has a great "plausibility"; its malleability is very strong. Whether it's still photography after the arrival of digital technology, and now after AI generation software, is something we can discuss further. But we must make sure to use yingxiang, image, consistently, and not tuxiang, picture. In fact, Flusser already settled this question in 1985 when he wrote Into the Universe of Technical Images. He proposed then that computer-generated images should be called "technical images," not pictures. So I think this point is very important.
Also, when we were talking about Paglen just now, the images we discussed were described as generated by "the photographer" Paglen, and that's somewhat problematic too. Paglen may have taken pictures with cameras, or with cameras he built himself, but I think he should be defined more as an "artist" than a "photographer." That's another issue.
Also, last weekend I asked Shore whether he would use AI to make work. Shore said he wouldn't. He said he'd rather go out into the real world to experience it and photograph it. I think Shore is a representative "straight photographer," a figure of pure photography, and he would define himself more as a photographer or an artist. But the artists now using AI software to make work should be called "image makers," zaoxiangshi. Of course we won't really use that term. We think "artist" is better. But we have to separate the two categories. As Ritchin says at the end, we must distinguish clearly between different kinds of images. Documentary photography and art photography are two categories. So I'd say Shore probably won't use AI to make work, but Jeff Wall and Gursky might. Gursky, in fact, has long used Photoshop to composite images, so I believe they might use AI to create images. That's one point.
I think in our "post-truth" era, already full of false images and false information, we actually need even more to use cameras to photograph everything happening now. That, of course, falls within photography, or photojournalism. Precisely because there are so many false images, we need more so-called real images, images that come from real life. I think what AI most lacks is this: it can imitate the past and fabricate the future, but the hardest thing for it is "immediacy," the direct experience. It doesn't have that, and it can't produce it. So it gives itself away in many ways. You can tell it wasn't photographed directly from the real world. It lacks "immediacy." So I think we need to be precise when we speak and translate. For example, "Deep Reals" was just rendered as shendu xianshi ("deep reality"), but I think shendu zhenshi ("deep realness") is more accurate. Xianshi (reality) and zhenshi (the real) are not the same. So I think many translations need to be very precise, and we need to be consistent when we speak.
Ritchin uses the term "photo-realistic image." Guosen and I have discussed this too. Personally, I don't think there's a fixed, accurate rendering yet, but roughly it should be translated as "an image that looks like a photograph." It looks like a photograph but isn't one, yet it is still an image. So after all this, what I want to stress is the difference between "picture" and "image." From now on, I urge everyone: when we talk about AI, please say "image," not "picture." Okay, thank you, everyone.
Response: Thank you, Professor Rong Jiang, for that addition. He and I have actually been discussing this question for a long time. What he's talking about is "image" versus "picture," and keeping these two words clearly distinct. Later, as I read the literature, I found that translators and scholars in China often use them interchangeably. They decide that "image" should be rendered as tupian or tuxiang in one place, or even xingxiang, splitting it across different words and mixing them together in translation. But doing so creates a lot of confusion, and I don't think it's necessary, because moving things can also be called images. Take the Paglen video I showed earlier, which Aperture also discusses. It isn't a photograph. Paglen's work isn't a "picture." It's an "image," a moving image. So I'm grateful for Professor Rong Jiang's correction. If we can accept it, reading these texts will actually become clearer. It really is quite necessary to distinguish them.
Professor Rong Jiang also mentioned picking up a camera and going out onto the street, or elsewhere, to photograph in the current environment, and that really struck a chord with me. When AI first appeared, many people asked: now that AI is here, is there still any point in taking photographs? Is documentary photography dead? My view back then was different. It is precisely because of AI that the value of those of us who actually pick up cameras and go out onto the street becomes clear. If you compare them, to use a not-quite-apt example: if you were getting married, would you want AI to generate perfect-looking wedding portraits and photos of the ceremony, or would you want to hire a real photographer, even one who doesn't shoot that well, to come and record it for you? I think very few people would choose the former.
Someone in the comments also said that the logic of AI images is "generation," while our thinking about documentary photography is "capture." The act of "capture" reminded me of a question someone once asked Shore: do you think game screenshots count as photography? Shore said he thought they did. Why? Because, he said, for him photography is an "analysis" of the existing world, not "composition," not assembling things together. Producing a photograph by analyzing the world in front of you is photography. So his answer was yes. Now, a friend in the comments says photography is "capture." I think "capture" is in a sense also a kind of analysis. You're analyzing the world in front of you, and that is itself an experience. The process and value of that experience can't be canceled out by the photograph you end up with. Even if an AI-generated photograph really does look better than one you took yourself, it can't give you the experience of actually going out onto the street and analyzing the world. That's my final addition. Thank you, Qiu, thank you, Professor Rong Jiang, and thank you, everyone.
Qiu
Thank you, Professor Jiang and Prismo. That brings today's session to a close. Thank you all very much for joining. The Wuzhi Society holds reading groups and interdisciplinary talks every month. Next month's Wuzhi Reading Room will be announced on our WeChat official account, so please follow us there and on Xiaohongshu. A recording of today's event will be available on the Wuzhi Society's Bilibili channel.
Hosts: Wuzhi Reading Room, Photography Theory Shelter
Moderator: Qiu
Organizers: Runxi Wang, Qiu, Connie
Guosen Chen