From Data to Intelligence: Foundations of Machine Learning and Neural Networks
Lesson 01

This is a course about artificial intelligence and image art, but the first two sessions are almost entirely technical. The reason is that if we don't know in what sense a machine "learns," or under what conditions it fails, our philosophical and aesthetic judgments about it can easily overstep their bounds. One important lesson Kant left us is that thought needs to know its own limits; drawing boundaries is not giving up on inquiry, but precisely a way of protecting the inquiries that truly matter. This lesson begins with a concrete question: when a machine is given an image of a handwritten digit, how does it decide whether it's a 0 or a 1?
The image must first be converted into a string of numbers; the machine then adjusts its own parameters by repeatedly comparing the gap between its prediction and the correct answer, gradually approaching a function that can classify correctly. But classical theory tells us that once a model's complexity exceeds a certain critical point, it memorizes noise as if it were pattern, so its performance on new data deteriorates sharply. Today's large language models have billions or even trillions of parameters and by rights should fail badly, yet the opposite is true: they spontaneously exhibit abilities they were never explicitly asked for, while also giving completely wrong answers where we least expect it. Why "big" is useful, and why it also fails in this particular way, still has no agreed-upon explanation. I hope this lesson can provide, for our later discussions of AI and images, a kind of intuition that can only form after accumulating enough technical detail.
Pixels and Word Meanings: Vision and Language in Machine Learning
Lesson 02

Last time we talked about how machines process images: however many pixels an image has, it is ultimately a two-dimensional grid in which pixels are adjacent to one another and can naturally be represented by continuous numbers. Language is entirely different. Pixel brightness values of 127 and 128 differ by just one unit, but there is no such continuous relationship between "cat" and "dog"; worse, the same word "apple" can mean a fruit or a company. If we want a neural network to process language, the first problem is how to turn words into numbers a machine can compute with. The core of this lesson is a very clever idea inspired by Word2Vec. In 1957 the linguist Firth said that the meaning of a word is determined by the words it frequently appears with. If two words always appear in similar contexts, they very likely have similar meanings. This means semantics can be reduced to a statistical regularity, and statistical regularities are exactly what neural networks are good at learning from data. The genius of Word2Vec is that it doesn't directly tell the network which words are similar; instead it gives the network a prediction task, and in order to accomplish the prediction, the network has to develop its own way of encoding meaning. In the end, what we want is not the prediction itself but the intermediate representations the network learned in order to make the prediction. This "byproduct" idea is crucial for understanding many of the issues we'll discuss in later lessons.
Limits and Fallacies: How Should We Think About AI and Its Creations?
Lesson 03

We've probably all had this experience: sometimes, when discussing a fairly complex problem with a large language model, its response is so good that it feels as if it really understood what I was saying; but at other times, when asked to compare two decimal numbers or fix a simple bug, its responses leave one at a loss. Faced with this contrast, one popular explanation is that AI is essentially just doing probabilistic prediction, and so it cannot truly understand anything. As a technical description this gets nothing wrong, but the problem lies in that "so." Does the fact that a system's training objective is probabilistic prediction mean that everything it has learned is also just probabilistic prediction? We devoted the previous two lessons to Word2Vec precisely to provide a counterexample here: a system designed to make predictions spontaneously developed, in order to make those predictions, abilities its training objective never explicitly demanded. These two sessions use several experiments to test this question. One is Othello-GPT: researchers had a Transformer learn to play Othello solely from sequences of game moves, never telling it what the board looks like, yet it spontaneously built an internal world model corresponding to the real board and actually used this model to make predictions. Of course, a game board is a closed system with fully determined rules, and natural language is far more complex. But these experiments at least force us to face a fact: between "it's just doing probabilistic prediction" and "it truly understands" lies a large territory that has not yet been seriously cleared. At the same time, I will also discuss some critical concepts surrounding AI-generated images, to see which critiques stand up to scrutiny and which leap too hastily between technical description and value judgment.
Images and Trust: The Question of Photographic Truth After AI
Lesson 04

By now we have probably all experienced moments like this: seeing an outrageous piece of news on social media, our first reaction is to wonder whether it was AI-generated. Yet conversely, when we scroll past a friend's photos on WeChat Moments, we don't seem to have the same suspicion. Behind this difference lies a deeper question: compared with text, why do we always tend to give photographs greater trust? Even today, when AI-generated images are so realistic, this trust has not been easily revoked. Peirce spoke of "collateral knowledge": before we receive a sign, a whole body of background knowledge is already helping us interpret it, often without our even being aware of it. Nearly two centuries of image culture, from 1839 to the present, have cultivated our collateral knowledge about the relationships among images, truth, and fact. Trusting photographs used to be an economically rational choice, because the cost of faking was extremely high, keeping the share of fake images among all images very low. But if AI drives the cost of faking toward zero, what happens to this structure of trust? Through images we discover things beyond the edges of our experience and are shaken or even outraged by them—and the premise of all this is trust. What this lesson seeks to lay out are the grounds of this trust, and the changes these grounds are undergoing today.
The Work of Philosophy of Technology: On the Creativity of Artificial Intelligence
Lesson 05

In 2022, an image Jason Allen generated with Midjourney won first prize in the digital art competition at the Colorado State Fair, provoking widespread anger. People's discontent seemed to stem from an intuition: this person didn't make the work with his own hands. But if we think about it a little, art history is full of works not made by the artist's own hand. Duchamp didn't make the urinal, and John Cage's 4′33″ plays almost nothing, yet both constituted art events of the twentieth century. So making something by hand has never been a necessary condition for artistic value; our discomfort must have another source. What this lesson really wants to discuss is not what Allen did, but what the AI he used did. Why do we instinctively feel that what AI generates lacks true creativity? This intuition may or may not hold up, but it needs some grounds either way. I will approach this question with the help of Yuk Hui's and Merleau-Ponty's discussions of Cézanne. Cézanne painted the same mountain dozens of times, trying to capture something he himself couldn't articulate, and to the end of his life never felt he had reached completion. Can this pursuit, sustained over decades, leave traces in the painting? Do such traces belong to an incalculable realm—one that our aesthetic sensibility is precisely able to reach?
The Problem of Judgment: *Kant Machine* and Moral Capacity
Lesson 06

The question for this lesson originates from a case: someone posted a genuine late Water Lilies by Monet on social media, told viewers it was AI-generated, and asked them to evaluate why it fell short of Monet's originals. The comment section filled with people seriously analyzing the flaws of this "AI work": wrong color choices, inconsistent depth of field, a stiff picture, an inability to move the viewer emotionally. Later the poster revealed that it was an authentic Monet. The episode is of course quite comical, but I don't really agree with the breezy conclusion that humans have no judgment at all and are simply blinded by labels. The problem is far more complex. The main thread of this lesson is Yuk Hui's Kant Machine. Hui tries to use Kant's critical philosophy to mark out the limits of AI's capacities, and his ultimate argument is that AI in principle cannot possess moral capacity. "In principle" here means that no matter how far AI develops, this will not change, because morality belongs to a realm he calls the "incalculable," which categorically does not belong to the territory of computation. But this argument faces a challenge I consider quite serious: if we accept that AI can emerge from simple training objectives with capabilities far exceeding expectations, on what grounds can we be sure it won't keep growing? Hui's answer to this, and the extent to which that answer ultimately relies on a value judgment that cannot be further argued for, is what I want to examine with you.
World Models: A Vision of Understanding
Lesson 07

In previous lessons we discussed whether AI "understands" and whether it has creativity and moral capacity. This lesson turns to a more concrete technical vision: JEPA, proposed by Yann LeCun. He argues that the current Transformer architecture has a fundamental problem: when processing continuous, high-dimensional signals, it must predict every detail, including those not worth predicting. Take a driving video: all we want to know is which direction the car is heading, yet a Transformer must first predict every pixel of the next frame and then extract the answer from it. This is predicting first, then understanding. LeCun wants to reverse the order: first compress the raw data into representations through an encoder, filtering out irrelevant details, and then make predictions at the level of representations. For him, representation means a kind of understanding. But this vision raises a very deep question. In the 1963 "kitten carousel experiment," two kittens received exactly the same visual stimulation, the only variable being that one could walk actively while the other could only be moved passively. Only the actively walking kitten developed normal visual abilities. Perception and action seem fundamentally inseparable. What LeCun envisions is precisely a system that can build a world model purely through observation—a spectator that takes part in nothing yet understands everything. All beings that have had world models in the past were also agents, living creatures that could be hurt and could die, so no one ever needed to ask what roles observation and action each play. JEPA makes this question, for the first time, something that can be tested on its own—and in turn it makes us re-examine ourselves: what do action, pain, and finitude really mean for understanding?

