Back to course catalog
Special course

Special Course: Art and Images After Artificial Intelligence

From the basic principles of artificial intelligence to creativity, image ethics, and world models—forming your own judgments.

Course content8 lessonsInstructorGuosen ChenFormatOnline course
Special Course: Art and Images After Artificial Intelligence course poster

Course overview

Starting from the technical foundations of machine learning, neural networks, and vision and language, the course discusses AI's capacity for understanding, truth and trust in images, creativity, judgment, and world models. Technical explanation and philosophical analysis are interwoven, helping us understand concretely both generated images and the artistic debates surrounding them.

Course content

From Data to Intelligence: Foundations of Machine Learning and Neural Networks

Lesson 01

Trevor Paglen, Image Operations, Op. 10, 2018.
Trevor Paglen, Image Operations, Op. 10, 2018.

This is a course about artificial intelligence and image art, but the first two sessions are almost entirely technical. The reason is that if we don't know in what sense a machine "learns," or under what conditions it fails, our philosophical and aesthetic judgments about it can easily overstep their bounds. One important lesson Kant left us is that thought needs to know its own limits; drawing boundaries is not giving up on inquiry, but precisely a way of protecting the inquiries that truly matter. This lesson begins with a concrete question: when a machine is given an image of a handwritten digit, how does it decide whether it's a 0 or a 1?

The image must first be converted into a string of numbers; the machine then adjusts its own parameters by repeatedly comparing the gap between its prediction and the correct answer, gradually approaching a function that can classify correctly. But classical theory tells us that once a model's complexity exceeds a certain critical point, it memorizes noise as if it were pattern, so its performance on new data deteriorates sharply. Today's large language models have billions or even trillions of parameters and by rights should fail badly, yet the opposite is true: they spontaneously exhibit abilities they were never explicitly asked for, while also giving completely wrong answers where we least expect it. Why "big" is useful, and why it also fails in this particular way, still has no agreed-upon explanation. I hope this lesson can provide, for our later discussions of AI and images, a kind of intuition that can only form after accumulating enough technical detail.

Pixels and Word Meanings: Vision and Language in Machine Learning

Lesson 02

Trevor Paglen, an open hangar in Nevada photographed through a powerful telescope.
Trevor Paglen, an open hangar in Nevada photographed through a powerful telescope.

Last time we talked about how machines process images: however many pixels an image has, it is ultimately a two-dimensional grid in which pixels are adjacent to one another and can naturally be represented by continuous numbers. Language is entirely different. Pixel brightness values of 127 and 128 differ by just one unit, but there is no such continuous relationship between "cat" and "dog"; worse, the same word "apple" can mean a fruit or a company. If we want a neural network to process language, the first problem is how to turn words into numbers a machine can compute with. The core of this lesson is a very clever idea inspired by Word2Vec. In 1957 the linguist Firth said that the meaning of a word is determined by the words it frequently appears with. If two words always appear in similar contexts, they very likely have similar meanings. This means semantics can be reduced to a statistical regularity, and statistical regularities are exactly what neural networks are good at learning from data. The genius of Word2Vec is that it doesn't directly tell the network which words are similar; instead it gives the network a prediction task, and in order to accomplish the prediction, the network has to develop its own way of encoding meaning. In the end, what we want is not the prediction itself but the intermediate representations the network learned in order to make the prediction. This "byproduct" idea is crucial for understanding many of the issues we'll discuss in later lessons.

Limits and Fallacies: How Should We Think About AI and Its Creations?

Lesson 03

Othello-GPT: changing an internally represented piece position to observe the next-move prediction.
Othello-GPT: intervening in internal representations to examine how the model encodes the board state.

We've probably all had this experience: sometimes, when discussing a fairly complex problem with a large language model, its response is so good that it feels as if it really understood what I was saying; but at other times, when asked to compare two decimal numbers or fix a simple bug, its responses leave one at a loss. Faced with this contrast, one popular explanation is that AI is essentially just doing probabilistic prediction, and so it cannot truly understand anything. As a technical description this gets nothing wrong, but the problem lies in that "so." Does the fact that a system's training objective is probabilistic prediction mean that everything it has learned is also just probabilistic prediction? We devoted the previous two lessons to Word2Vec precisely to provide a counterexample here: a system designed to make predictions spontaneously developed, in order to make those predictions, abilities its training objective never explicitly demanded. These two sessions use several experiments to test this question. One is Othello-GPT: researchers had a Transformer learn to play Othello solely from sequences of game moves, never telling it what the board looks like, yet it spontaneously built an internal world model corresponding to the real board and actually used this model to make predictions. Of course, a game board is a closed system with fully determined rules, and natural language is far more complex. But these experiments at least force us to face a fact: between "it's just doing probabilistic prediction" and "it truly understands" lies a large territory that has not yet been seriously cleared. At the same time, I will also discuss some critical concepts surrounding AI-generated images, to see which critiques stand up to scrutiny and which leap too hastily between technical description and value judgment.

Images and Trust: The Question of Photographic Truth After AI

Lesson 04

Joan Fontcuberta, Sputnik, 1997.
Joan Fontcuberta, Sputnik, 1997.

By now we have probably all experienced moments like this: seeing an outrageous piece of news on social media, our first reaction is to wonder whether it was AI-generated. Yet conversely, when we scroll past a friend's photos on WeChat Moments, we don't seem to have the same suspicion. Behind this difference lies a deeper question: compared with text, why do we always tend to give photographs greater trust? Even today, when AI-generated images are so realistic, this trust has not been easily revoked. Peirce spoke of "collateral knowledge": before we receive a sign, a whole body of background knowledge is already helping us interpret it, often without our even being aware of it. Nearly two centuries of image culture, from 1839 to the present, have cultivated our collateral knowledge about the relationships among images, truth, and fact. Trusting photographs used to be an economically rational choice, because the cost of faking was extremely high, keeping the share of fake images among all images very low. But if AI drives the cost of faking toward zero, what happens to this structure of trust? Through images we discover things beyond the edges of our experience and are shaken or even outraged by them—and the premise of all this is trust. What this lesson seeks to lay out are the grounds of this trust, and the changes these grounds are undergoing today.

The Work of Philosophy of Technology: On the Creativity of Artificial Intelligence

Lesson 05

Holly Herndon and Mat Dryhurst, Starmirror, 2025; design by Patricia Klein.
Holly Herndon and Mat Dryhurst, Starmirror, 2025; design by Patricia Klein.

In 2022, an image Jason Allen generated with Midjourney won first prize in the digital art competition at the Colorado State Fair, provoking widespread anger. People's discontent seemed to stem from an intuition: this person didn't make the work with his own hands. But if we think about it a little, art history is full of works not made by the artist's own hand. Duchamp didn't make the urinal, and John Cage's 4′33″ plays almost nothing, yet both constituted art events of the twentieth century. So making something by hand has never been a necessary condition for artistic value; our discomfort must have another source. What this lesson really wants to discuss is not what Allen did, but what the AI he used did. Why do we instinctively feel that what AI generates lacks true creativity? This intuition may or may not hold up, but it needs some grounds either way. I will approach this question with the help of Yuk Hui's and Merleau-Ponty's discussions of Cézanne. Cézanne painted the same mountain dozens of times, trying to capture something he himself couldn't articulate, and to the end of his life never felt he had reached completion. Can this pursuit, sustained over decades, leave traces in the painting? Do such traces belong to an incalculable realm—one that our aesthetic sensibility is precisely able to reach?

The Problem of Judgment: *Kant Machine* and Moral Capacity

Lesson 06

Hilma af Klint, Untitled, 1922.
Hilma af Klint, Untitled, 1922.

The question for this lesson originates from a case: someone posted a genuine late Water Lilies by Monet on social media, told viewers it was AI-generated, and asked them to evaluate why it fell short of Monet's originals. The comment section filled with people seriously analyzing the flaws of this "AI work": wrong color choices, inconsistent depth of field, a stiff picture, an inability to move the viewer emotionally. Later the poster revealed that it was an authentic Monet. The episode is of course quite comical, but I don't really agree with the breezy conclusion that humans have no judgment at all and are simply blinded by labels. The problem is far more complex. The main thread of this lesson is Yuk Hui's Kant Machine. Hui tries to use Kant's critical philosophy to mark out the limits of AI's capacities, and his ultimate argument is that AI in principle cannot possess moral capacity. "In principle" here means that no matter how far AI develops, this will not change, because morality belongs to a realm he calls the "incalculable," which categorically does not belong to the territory of computation. But this argument faces a challenge I consider quite serious: if we accept that AI can emerge from simple training objectives with capabilities far exceeding expectations, on what grounds can we be sure it won't keep growing? Hui's answer to this, and the extent to which that answer ultimately relies on a value judgment that cannot be further argued for, is what I want to examine with you.

World Models: A Vision of Understanding

Lesson 07

Hilma af Klint, Evolution, No. 7.
Hilma af Klint, Evolution, No. 7.

In previous lessons we discussed whether AI "understands" and whether it has creativity and moral capacity. This lesson turns to a more concrete technical vision: JEPA, proposed by Yann LeCun. He argues that the current Transformer architecture has a fundamental problem: when processing continuous, high-dimensional signals, it must predict every detail, including those not worth predicting. Take a driving video: all we want to know is which direction the car is heading, yet a Transformer must first predict every pixel of the next frame and then extract the answer from it. This is predicting first, then understanding. LeCun wants to reverse the order: first compress the raw data into representations through an encoder, filtering out irrelevant details, and then make predictions at the level of representations. For him, representation means a kind of understanding. But this vision raises a very deep question. In the 1963 "kitten carousel experiment," two kittens received exactly the same visual stimulation, the only variable being that one could walk actively while the other could only be moved passively. Only the actively walking kitten developed normal visual abilities. Perception and action seem fundamentally inseparable. What LeCun envisions is precisely a system that can build a world model purely through observation—a spectator that takes part in nothing yet understands everything. All beings that have had world models in the past were also agents, living creatures that could be hurt and could die, so no one ever needed to ask what roles observation and action each play. JEPA makes this question, for the first time, something that can be tested on its own—and in turn it makes us re-examine ourselves: what do action, pain, and finitude really mean for understanding?

References

迈克尔·尼尔森. 深入浅出神经网络与深度学习[M]. 朱小虎,译. 北京:人民邮电出版社,2020.

王维嘉. 优美与崇高:康德的感性判断力批判[M]. 上海:上海三联书店,2020.

康德. 康德著作全集:第 5 卷:实践理性批判、判断力批判[M]. 李秋零,主编、译. 北京:中国人民大学出版社,2007.

康德. 道德形而上学的奠基[M]//康德著作全集:第 4 卷. 李秋零,译:393-472.

托马斯·福克斯. 脑:一个关系器官:一种现象学-生态学构想[M]. 王旭,译. 北京:商务印书馆,2025.

莫里斯·梅洛-庞蒂. 意义与无意义[M]. 张颖,译. 北京:商务印书馆,2018.

威廉·弗鲁塞尔. 摄影哲学的思考[M]. 毛卫东,丁君君,译. 北京:中国民族摄影艺术出版社,2017.

许煜. 艺术与宇宙技术[M]. 苏子滢,译. 上海:华东师范大学出版社,2022.

Azoulay, Ariella. Civil Imagination: A Political Ontology of Photography. Translated by Louise Bethlehem, paperback ed., Verso, 2015.

Crawford, Kate, and Trevor Paglen. “Excavating AI: The Politics of Images in Machine Learning Training Sets.” Excavating AI, AI Now Institute, 19 Sept. 2019.

Daston, Lorraine, and Peter Galison. Objectivity. Zone Books, 2007.

Firth, J. R. “A Synopsis of Linguistic Theory, 1930–1955.” Studies in Linguistic Analysis, Basil Blackwell, 1957, pp. 1–32.

Hui, Yuk. Kant Machine: Critical Philosophy after AI. Bloomsbury Academic, 2026.

LeCun, Yann. A Path Towards Autonomous Machine Intelligence. Version 0.9.2, 27 June 2022. OpenReview.

Li, Kenneth, et al. “Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task.” International Conference on Learning Representations, 2023.

Mikolov, Tomas, et al. “Distributed Representations of Words and Phrases and Their Compositionality.” Advances in Neural Information Processing Systems, vol. 26, 2013, pp. 3111–19.

Mikolov, Tomas, et al. “Efficient Estimation of Word Representations in Vector Space.” arXiv, 2013, arXiv:1301.3781v3.

Minsky, Marvin, and Seymour Papert. Perceptrons: An Introduction to Computational Geometry. Expanded ed., MIT Press, 1988.

Nakkiran, Preetum, et al. “Deep Double Descent: Where Bigger Models and More Data Hurt.” arXiv, 4 Dec. 2019, arXiv:1912.02292v1.

Nanda, Neel, et al. “Emergent Linear Representations in World Models of Self-Supervised Sequence Models.” arXiv, 7 Sept. 2023, arXiv:2309.00941v2.

Peirce, Charles S. The Essential Peirce: Selected Philosophical Writings. Vol. 2, 1893–1913, edited by the Peirce Edition Project, Indiana University Press, 1998.

Roberts, John. Photography and Its Violations. Columbia University Press, 2014.

Rombach, Robin, et al. “High-Resolution Image Synthesis with Latent Diffusion Models.” arXiv, 13 Apr. 2022, arXiv:2112.10752v2.

Sekula, Allan. “The Body and the Archive.” October, vol. 39, 1986, pp. 3–64.

Shannon, C. E. “A Mathematical Theory of Communication.” The Bell System Technical Journal, vol. 27, nos. 3–4, 1948, pp. 379–423, 623–56.

Silver, David, et al. “Reward Is Enough.” Artificial Intelligence, vol. 299, 2021, article 103535.

Steyerl, Hito. Medium Hot: Images in the Age of Heat. Verso, 2025.

Steyerl, Hito. “In Defense of the Poor Image.” e-flux Journal, no. 10, Nov. 2009.

Turing, A. M. “On Computable Numbers, with an Application to the Entscheidungsproblem.” Proceedings of the London Mathematical Society, 2nd ser., vol. 42, no. 1, 1937, pp. 230–65.

Vaswani, Ashish, et al. “Attention Is All You Need.” Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5998–6008.

Zylinska, Joanna. “We Have Always Been Artificially Intelligent: An Interview with Joanna Zylinska.” Interview by Claudio Celis and Pablo Ortuzar Kunstmann. Culture Machine, vol. 20, 2021.

About the instructor

Guosen Chen

Guosen Chen

Guosen Chen is a PhD candidate in the Department of Philosophy of Art at Fudan University's School of Philosophy and a co-founder of the "Photography Theory Shelter" community. He has taught nine seasons of the "Photography Theory Bootcamp," with more than seven hundred participant enrollments. His research interests include photography theory, phenomenological aesthetics, and AI ethics. He has held numerous conversations with renowned photographers and scholars including Stephen Shore, Jeff Wall, Thomas Ruff, and Charlotte Cotton. In his spare time he translates photography literature. He has published papers in Chinese and English in Photography and Culture, Twenty-First Century, Journal of Nanjing University of the Arts (Fine Arts & Design), Chinese Photography, Studies of French Philosophy, Art Work, Cultural and Art Research, and elsewhere.