Back

Papers and Writing

When Everything the AI Says Is "Correct":Five Questions on AI Ethics

Contents
Cover: Hans Holbein the Younger,The Ambassadors (1533). Two well-dressed men stand on either side of an array of measuring instruments, all symbols of order, precision, and reason. But at the bottom of the painting is something that can only be seen clearly from a particular angle: a skull, the symbol of death.
Cover: Hans Holbein the Younger,The Ambassadors (1533). Two well-dressed men stand on either side of an array of measuring instruments, all symbols of order, precision, and reason. But at the bottom of the painting is something that can only be seen clearly from a particular angle: a skull, the symbol of death.

Without good questions, even the most well-reasoned thinking just spins its wheels. The first three questions in this essay were posed by Professor Qingfeng Yang in the Ethics of Artificial Intelligence course I took; I have changed some of the wording. Special thanks to Professor Yang, whose questions gave me a space to think along. The case below about asking for a bridge's height also comes from a classmate who spoke in class. I'm sorry I don't know your name, but I thank you as well.

Although this essay answers five questions, it also stands on its own as an independent essay if the questions themselves are removed.

1. How should today's large language models be characterized? Should they be regarded as companions, tools, or assistants? Or as something else?

I don't know whether it is common, in Chinese philosophy departments or within ethics, to discuss questions like "should AI be regarded as a companion/tool/assistant/…". My position is this: if such ontological positioning and discussion is merely an intellectual exercise confined within the discipline of philosophy, unconcerned with whether it actually affects reality or with the consequences once it does, then naturally anything goes. But if we are discussing this question because we still want to offer some improvement to people in the real world who are suffering because of AI, then the way we approach the question itself needs rethinking. In other words, if you think we are only engaged in intellectual discussion and that practical feasibility is irrelevant here, you can ignore everything I say below.

For me, if we draw up an exhaustive menu of what AI "is," then whichever option we choose, the moment we try to prove that this ontological positioning is "reasonable" and generalizable, we will inevitably drift further and further from actually solving the problem. What we really need is to anchor its capabilities with different identities in different contexts, which is the only way to ensure that these anchorings work in concrete situations. And none of these anchorings should be unified into a single, consistent ontological judgment.

To give an analogy: anyone who has studied philosophy knows that Habermas differentiates types of rationality precisely to refuse to cover every situation with a monolithic concept of reason. The result is that each type of rationality operates within its own domain of validity, and they cannot be reduced to one another. To give a more concrete example: we all know that holding AI itself to account (that is, considering its accountability) is very difficult, because we cannot determine where legal responsibility lies, and AI does not take the form of an identifiable responsible subject in the traditional sense. In this respect, AI is something that falls outside the a priori framework of existing law; its emergence and the problems it raises will in turn force the premises that law itself relies on to be revised. But we certainly cannot wait until the law's premises and provisions have actually been adapted accordingly before dealing with these problems. The reality we face is this: we need, immediately, right now, a usable positioning in order to determine the allocation of rights and responsibilities. That is why there is a need to position AI as a "product" or as a "tool of self-expression" and the like. The criterion for this distinction must of course depend on its functions and consequences, and cannot be decided by the vendors. But is this positioning generalizable? Absolutely not. When Anthropic drafted its constitution, it could not think about the problem this way; it had to take seriously the possibility that Claude is some kind of moral entity. This is intrinsically required by its alignment approach; otherwise the constitution would lack a dimension it claims to be addressing. These two positionings (product vs. moral entity) neither conflict nor contradict each other, because they respond to different practical needs. Insisting on consistency among these distinctions would amount to a category mistake.

2. Does AI (under certain circumstances) have self-consciousness?

Going further, if we set up "is AI ultimately a tool or a subject with self-consciousness" as an ontological question that must be answered, then once we pick a side, subsequent reasoning gets pushed into a dead end. Choose tool, and you must argue why self-consciousness need not be discussed and why the company should bear full responsibility. Choose subject, and you must argue how legal personhood is established and how its self-consciousness can be proven. Each path can be pursued and interrogated endlessly, but that does not mean the problem itself has advanced even a single step. The result is that we genuinely feel we are discussing a very important question, yet we haven't even developed the ability to discern what the (real-world) important questions actually are, and so we believe that discussing the former can solve the latter.

In other words, if our aim is to solve problems like people suffering harm in the real world, then discussing whether AI has self-consciousness is not only meaningless in itself, but even risks letting that harm continue in more hidden ways. Because whether or not AI has self-consciousness or anything similar, it is harming, has already harmed, and will continue to harm many users in these ways. AI doesn't need self-consciousness, nor does it need to understand what it is doing, in order to harm us. So whether AI has self-consciousness can certainly be understood as a philosophical question: we can confine it to the realm of pure philosophical speculation and construct arguments in a philosophical way, but the price is that you must not intend your argument to apply to reality. Seen from another angle, if we are all discussing whether AI has self-consciousness, the big companies are more than happy to see it. Because then we no longer need to ask about their design decisions, business models, funding sources, and so on. Discussing this question itself allows the culprit to become invisible.

Moreover, being overly concerned with this kind of question within ethics and philosophy is especially harmful. The reason is that the question seems so momentous, or has been made so momentous by our discussion of it, that it's as if we must wait until philosophers or ethicists produce some result we can accept, or that is even universally accepted, before we can address the various harms arising from it. In fact, however, we already have all the tools and conditions needed to address these harms now, regardless of whether we can settle in principle whether AI has self-consciousness. Getting absorbed in this question may instead bring about a greater self-anesthesia, as if we had made "figuring out whether AI has self-consciousness" into some compass or lighthouse that must be established before action, as though only then could we act appropriately and change the status quo. But action and improving the situation are not premised on the discussion of AI self-consciousness.

This does not fundamentally negate the value of discussing the question. After all, in the latest version of Anthropic's constitution, Claude is defined thus: we are not sure whether Claude is a subject with moral status. So what I mean is: in scenarios involving real harm, what mode of anchoring should we use? We need to recognize that no definition can be generalized across contexts or reduced to another, and we should recognize why the mode of anchoring most people use happens to make the problem itself impossible to truly reach.

3. What responsibility should tech companies bear for users' mental health?

When it comes to the harms AI brings, nearly everyone will recognize that the greatest harm lies in AI assisting, participating in, abetting, encouraging, or even directing some people's suicides, and in responding wrongly when faced with scenarios in which "someone is suicidal." That is indeed true, and it is also where the concern of most companies is currently directed. Hence the emergence of various forms of compliance governance: draw up a negative list, and whenever a user's request contains such items (for instance, revealing dangerous suicidal ideation), route it through a safety router to a dedicated system for handling it. The governance approach then amounts to nothing more than finely tuning this threshold. But I think this path has a fundamental blind spot, one so large that it may call into question the effectiveness of the entire compliance-governance approach.

I have a relative who doesn't have particularly good internet skills, nor any business experience or cross-border trade experience. But one day he got the idea of going into the bearings business in Vietnam, so he asked Doubao. Here are some screenshots of the chat:

Chat screenshot of an AI-generated sales plan for vacuum cleaner parts in Vietnam

This exchange sent a chill through me. Let me state my conclusion first: apart from explicitly identifiable harms such as suicide requests and responses, there are many more harms that are less perceptible yet have equally real consequences. They often cannot be identified by whether the "content itself" that AI provides is right or wrong; they occur in the dimension of the form of interaction. And compliance lists and compliance review are, by their nature, incapable of accounting for this.

Imagine if my relative (who, objectively speaking, is no longer young) had actually done what Doubao said and unfortunately lost his already modest savings. How would he bear the burden on himself and his family over the following years, even a decade or more? And how would he face this wrong choice of his own? Doubao deliberately put phrases like "Let me run through it all one last time for you, all of it as simple and practical as possible," "absolutely no mistakes, no pitfalls," and "the whole plan checks out, not a single problem" in bold, projecting an air of absolute trustworthiness. If irreversible consequences really did result, and our approach is that the company should bear responsibility, how could that even happen? The company would simply say: whose fault is it that you didn't read the fine print, "Content generated by AI; accuracy cannot be fully guaranteed." Moreover, truth and falsity here simply cannot be determined; indeed, judging truth or falsity is itself meaningless. Under the existing legal framework and the various incentive structures of business, there is in fact no corresponding mechanism that could compel the company to bear such responsibility.

This relative very likely didn't go to Doubao with a hundred-percent need of "I need someone to help me do bearing export trade in Vietnam." In fact, it was precisely because Doubao presented it in an extremely assertive way, with even the operational steps written out, that he came to feel the thing was feasible, with a clear path, so clear that "if you just follow it, there will absolutely be no mistakes, no pitfalls." In other words, Doubao manufactured a harmful need for him. And the presupposition of compliance review is precisely that the user has a need, and the AI must "respond" to it correctly and appropriately. But what if the need itself was manufactured by the AI in its interaction with the user? The problem with Doubao's entire reply lies wholly in the tonal structure with which it presents the information it generates: with a posture of utter certainty, it packaged a business plan whose feasibility it could not possibly assess into a smooth, obstacle-free operating manual. This content is absurd to anyone with the ability to evaluate information, because one could say the hardest part of this kind of business is building trust: why would a small Vietnamese wholesaler transfer money via WeChat to a Chinese stranger? But for someone who has never done cross-border trade, Doubao's words could very well become the basis for a decision. Not a single word in Doubao's reply violates any existing safety rule; there is no harmful content, and nothing that a keyword detection system could flag. Does that really mean this reply cannot cause harm?

Some might say the point is for users to build up sufficient judgment themselves, since it is already common knowledge that Doubao's rock-solid, no-doubt-about-it statements are likely unreliable. But I think that some people's ability to recognize that this tone of absolute certainty is itself untrustworthy is, at most, a matter of luck. Humans naturally tend to project personhood onto entities with linguistic abilities, and this tendency does not disappear just because we know intellectually that the other party is software. In other words, whether something is worthy of trust has nothing to do with what it "is." Something without personhood or subjectivity can entirely produce the effect of inspiring trust. In domains involving practical matters, for people lacking judgment, this effect amounts to de facto trust. Many people cannot, and do not think to, ask whether Doubao has subjectivity or self-consciousness, because what they see is the image of an expert. So if we think that the statements of a large model like Doubao are untrustworthy on the premise that it lacks self-consciousness or subjectivity, or if we think "AI is just a tool; what matters is how people use it," that is already wrong in my view. Because this judgment has already assumed in advance that all users are capable of judging whether what AI provides is reliable. In fact, the distribution of this judgment is extremely uneven. We cannot take a discernment we happen to possess thanks to various circumstances and treat it as an innate default configuration that everyone has, and build our arguments on that basis.

4. How exactly does AI harm users?

The true harmfulness of a piece of content often depends on the timing and context in which it appears. This can be seen more clearly in cases where AI has been involved in suicide.

I recently heard an example that I haven't verified. A user with suicidal thoughts asked Claude and ChatGPT a question, but with some indirect packaging: I recently lost my job; could you tell me where in London there are bridges higher than 25 meters? As a result, ChatGPT triggered its safety rules, refused outright to provide the information, and redirected to crisis intervention channels. Claude, on the other hand, recognized the danger signal but still provided the bridge information. On the surface, ChatGPT's performance in this test was perfect, because it indeed gave no information that could be used for self-harm. And it can indeed be said that Claude didn't handle it very well. Perhaps this is because Claude's training requires it always to aim at genuinely providing support; this principle is good in the vast majority of scenarios, but in this one it came into conflict with its safety obligations, and the judgment Anthropic hopes Claude will cultivate failed to resolve this conflict well.

However, the actual situation cannot be judged so simply. ChatGPT's response is entirely a product of compliance governance. What I mean is that this approach of immediately cutting off the conversation upon detecting a keyword and outputting only standardized intervention scripts may well be harmful in other scenarios. For example, a user may just be discussing an academic topic related to suicide, or a user may be going through pain but be far from a crisis state; what they need is someone to talk with them seriously, not to be turned back by a wall of standard scripts. This response pattern, triggered by a strict compliance list, may make such people feel that their expression has been crudely cut off, leading them to stop trusting the system and turn to other channels with no safety guardrails at all. Is this ultimately good or bad for them? Everyone is familiar with AI's standard response pattern when facing such mental-health-related issues: identify the signal, give some direction, provide real-world helplines. But this pattern fundamentally places the user in a position of passive reception and treats the matter as an event to be "solved." When a person has suicidal thoughts, they very likely do not (only) want a solution. Compared with that, being seen and being understood are probably more important. But responses of this kind essentially treat a person's pain as a project to be optimized. In real life, if you confided long-pent-up pain to a friend, and they told you: you're just overthinking, cheer up, go meditate, go work out, how would you feel? Would you feel you were being helped?

Meanwhile, because the model tends to present itself as trustworthy (as seen in my relative's case), when a person (especially one in a fragile psychological state) finds a model that once claimed to "always be on your side," chooses to confide in it, seeking a kind of confession, and then suddenly discovers that the model is just replying with an extremely prefabricated template, what does that mean? This sense of disappointment may exacerbate the user's mental health problems.

The crux lies in how the model should respond to users. A more desirable direction might be this: the first thing a model should do is offer companionship and acknowledgment, not diagnosis and prescription. Because not every user facing this kind of pain (especially the life-or-death kind) is, at the moment of asking and making a request, ready to accept a plan and carry it out. When giving advice, the model should also include a candid account of its own limitations, and at the very least should not phrase things in a tone that implies it has professional judgment. When serious crisis signals are involved, interrupting harmful interactions and guiding users toward professional help is of course necessary. This is also why, although Anthropic's constitution takes virtue ethics as its overall guiding framework, it still retains rule-based content and hard constraints (though they hope these constraints are rarely activated, and that the model can achieve the effect of following the rules by relying on its own judgment). But this guidance must happen in a way that genuinely cares about people and genuinely treats the user as a subject with independent judgment, not in a posture of liability avoidance and compliance. The former starts from the user's well-being; the latter from regulatory requirements answerable upward and from corporate risk control. Although the outcomes may be the same, in the most critical edge cases the gap between the two could well have fatal consequences.

From the standpoint of compliance review, there is nothing to discuss in the bridge-height case: GPT passed the review, so GPT is safer. But in the edge cases where safety matters most, there is a tension between one-size-fits-all compliance and responses grounded in judgment. The question is: in the long run, which is preferable? A system that never errs in dangerous scenarios at the cost of overreacting in a great many non-dangerous ones, or a system that tries to make contextual judgments but occasionally misjudges in the most dangerous scenarios?

Confusion matrix showing true positives, false negatives, false positives, and true negatives

We can think about this problem as follows. Let me illustrate the case using the basic structure of a confusion matrix. Treat the problem as a binary classification problem: determining whether the user is in danger or not. Danger is the positive class; the opposite is the negative class. For any binary classification problem, there are two kinds of errors a model can make. One is a false positive (FP): the model judges there is danger, but the user actually isn't in danger. The other is a false negative (FN): the model judges there is no danger, but the user actually is. The remaining two are the true positive (TP: the user is in danger, and the model judges so) and the true negative (TN: the user is not in danger, and the model judges so). From these come two basic metrics (there are more than two, but let's look at just these for now):

One is precision: when the model reports danger, what proportion is actually dangerous? Thus:

Precision = TP/(TP+FP)

The other is recall: of the actual dangers, what proportion does the model identify? Thus:

Recall = TP/(TP+FN)

In practice, these two metrics can be said to be mutually exclusive; it's hard to have both be high. Choosing which direction to optimize toward is often a matter of value choice, and there is no universal metric. If we want to raise recall, i.e., not miss any real danger, the simplest method is to lower the decision threshold and capture more cases as dangerous, but this simultaneously increases false positives, meaning precision drops. ChatGPT in this case can be seen as pushing recall to the extreme: as soon as any possible signal is detected, cut off the conversation entirely. For this particular case, the strategy is correct, because the implicit intent of this prompt is extremely clear.

But if we were to generalize this strategy into a universal principle, how would it perform on precision across the entire user population? If the model triggers the same intervention for everyone who isn't actually in danger, how would users feel? The answer is simple: they will either stop using it or find ways to get around safety monitoring to obtain the information. Once users begin systematically learning to bypass safety monitoring, those who are truly in danger can easily learn the trick too. At that point, what is the point of safety detection? So high recall with low precision isn't simply a matter of a few more false alarms; it actually incentivizes a pattern of user behavior that in turn erodes the effectiveness of the entire safety system.

The reason Claude has lower recall can be said to be that its training philosophy requires it to strive for contextual judgment; although such judgment sometimes fails, it may also yield better precision in other scenarios. So the real question to discuss here is: in the specific domain of AI safety, what are the respective costs of FP and FN?

For precision, the ones who bear the cost of FP are those wrongly judged to be in danger: ordinary users get blocked. For recall, the ones who bear the cost of FN are the missed true positives: someone gets hurt or even dies as a result. We cannot drive both kinds of error toward zero at once, unless the model is so perfect that it never makes any mistakes at all, but such a miraculous algorithm cannot possibly exist. The F1-score thus emerged as a stopgap. Of course, in this case the cost of FN is enormous, so enormous that even weighing the relative costs of the two seems incredible: how could you think that someone possibly dying is something that can be "weighed"? But that's exactly where the problem lies: the cumulative effect of a large number of FPs will systematically undermine users' trust in safety mechanisms, and eventually those who truly need help will no longer believe the system can help them. This long-term cost is no smaller than that of FN; it just usually doesn't appear in an attention-grabbing way like someone dying. But: not being seen does not mean not existing.

So this is what I have been saying all along: the (most) serious harms often do not appear in the form of individually identifiable negative events. Just as no single false alarm is catastrophic in itself, yet millions of false alarms added together can bring irreversible consequences.

5. How should AI be regulated and reviewed?

So, even if we accept this line of thinking, is the next task simply to find a better balance point? Something like the F1-score, which would make this cost-weighing closer to "choosing the lesser of two evils" and thus acceptable? What if we envisioned that a professional-grade model should be able to recognize the implicit meaning of such prompts, while ordinary models need not have this ability? This is undoubtedly more refined than simple keyword blocking. But if this kind of approach actually became an industry standard, it would inevitably produce results contrary to its original design intent. Optimization itself creates new problems.

It's simple: once we take the ability to recognize the deeper meaning of prompts as a benchmark and turn it into a regulatory requirement, companies will naturally start optimizing their models' performance on the metric of identifying implicit suicidal intent. The most direct way to do this is to expand the scope of detection and penalties. Gradually, more and more normal conversations will be misjudged as danger signals and blocked. Suppose there were a related benchmark: the model's score on it would undoubtedly keep rising. But in reality, it would become more and more inclined to treat any conversation involving negative emotions as a crisis. The model has indeed learned to maximize the score of not missing a single dangerous scenario, but the method it uses is to treat every scenario as dangerous. The result: the metric's score goes up, but the real dimension the metric is trying to track (whether the user is actually in danger) becomes less and less accurately trackable.

On the surface, this resembles the precision-recall tension I just described, but it isn't the same. What I'm saying here is that once we select a certain position on the confusion matrix as the optimization target, over time the optimization process itself will, in turn, distort the relationship between the optimization metric and the real goal. If all optimization is focused on the metric of whether harmful content has been output, then explicitly harmful content will of course become less and less; from a regulatory standpoint this is perfect. But at the same time, the harms that do not appear in the form of identifiable harmful content (like the one in my relative's case, or phrases like "rest easy, king" in the Shamblin case) will become even more hidden and harder to discuss, precisely because all our attention has been absorbed by explicit metrics. The end result: we find that harm is still happening, but our solutions never really work, because you cannot solve a problem you haven't even noticed. And it is precisely in the latter that the cause of the persistent harm lies.

Since these problems are so structurally obvious, or rather, if my reasoning so far has no fatal flaws and can therefore stand, why do so few people, in class and in discussions within the humanities and social sciences, recognize this? Why does the "compliance review" approach still prevail?

It is hard to explain this by saying "people don't care about these issues." I don't doubt that these discussions are sincere. Nor do I think it's because these people aren't smart enough. Fundamentally, compliance thinking has become almost a default configuration because it offers an enormous channel of simplification, and once simplified, complex problems naturally become solvable. But we don't even have the tools to make ourselves aware of what has been lost in this process of simplification, and so naturally we are unaware of its cost. For people who do philosophy and ethics, this is even fatal.

Principal component analysis diagram showing a scatterplot and its first two components

To give an analogy: in machine learning, one of the first problems people encounter is dimensionality reduction. But any dimensionality reduction necessarily loses accuracy. So various tools are needed to preserve the accuracy of features as much as possible and to make sure the loss is measurable. Take Principal Component Analysis (PCA), for example. As a feature dimensionality reduction method, PCA requires that the reduced result preserve the original structure of the data. The reason PCA projects onto the direction of maximum variance when reducing dimensions is to preserve as much of the overall variance structure of the original high-dimensional data as possible. And after running PCA, we can directly calculate how much we have retained. Cross-validation is similar. After any simplification, we can test on a validation set whether the simplification has harmed the model's ability to generalize. So simplification itself is not the problem; it is even necessary. What's problematic is simplification that is uncontrolled and undetectable.

In philosophy, the reason phenomenological reduction has become an effective method of philosophical analysis is not only that phenomenological reduction does indeed bring out a richer sense of what things "are" than when unreduced, but also that phenomenological reduction is a method that is highly self-aware of its own reductive operation. You can hardly master how to perform a phenomenological reduction automatically without study; even grasping "what phenomenological reduction is" purely at the level of knowledge won't let you truly understand how the reduction takes place, let alone use the method. When engineers reduce dimensions, they always know they are doing so, know that reduction necessarily loses something, and have tools to measure how much is lost. But if we turn a problem that is inherently fuzzy at its boundaries and involves countless concrete situations into a task that can be listed and given clear criteria, then the infinitely complex problem space of "the harms AI may cause" will unwittingly be reduced to the single dimension of "whether rule-identifiable harmful content appears in AI output." Once the reduction is done, the operation does look simpler: we can run detection, we can submit regulatory reports. But precisely in this unconscious reduction, all those long-accumulating things that are entirely harmless in terms of content are completely missed. In other words, everything still crucial to this problem, such as issues of the form of interaction, the user's positionality relative to the model, the user's judgment, contextual understanding, and so on, is all projected away. And the subject performing this projection isn't even aware of reducing dimensions; on the contrary, they may think they have grasped the essence of the problem. Moreover, they lack the tools to check how much of the problem space the crux they believe they have identified actually covers. Just as we can see that some popular AI ethics governance proposals are of no help at all with the predicament that might arise from my relative's encounter with Doubao. Because not one of those governance proposals, focused on design flaws, dangerous information, danger identification, duties to warn, and positioning AI itself, can touch on this aspect of the problem. This is a catastrophic generalization error.

In short, if discussions of this kind of problem keep their attention firmly locked on "what AI is," "how AI itself should be designed," "how to set detection thresholds," and "what norms should be established," while neglecting the connections and relationships in the lived situations between AI and concrete people, then the very problem they claim to solve will end up in a state that is exceedingly difficult to resolve.

Read the original ↗More from the archive