Enjoying the episode? Want to listen later? Subscribe on any of these apps or stores to be notified when we release new episodes:
August 19, 2026
What makes an attachment to AI meaningfully different from an attachment to a human, animal, or fictional character? If chatbot relationships are built around validation and adaptation, can they ever provide the friction and otherness that make human relationships opportunities for growth? Does a relationship need reciprocity to be genuine, or can an asymmetrical relationship still be deeply valuable? How would the possibility of AI sentience change the moral status of human relationships with artificial companions? When a company controls the personality of something a user loves, how much power does that company gain over the user? Why might emotional attachment create an internality that ordinary consumer protections are poorly equipped to handle? Can a freemium business model ever be compatible with healthy AI companionship when emotional dependency itself drives engagement? Where is the line between comforting validation and reinforcing distorted beliefs or harmful emotional patterns? How should an AI distinguish between respecting someone's subjective experience and helping them question a potentially false interpretation of reality? If people increasingly turn to systems optimized around their own preferences, what forms of compromise, disagreement, and personal growth might they lose?
Claire is an Assistant Professor in Technology Law and AI Governance at the European University Institute School of Transnational Governance and a Visiting Research Associate at the Cordell Institute of the Washington University School of Law.
Links:
Claire Previously Founded Successif
Spencer's Permanent Instructions for LLMs
SPENCER: Claire, welcome to the Clearer Thinking Podcast.
CLAIRE: Thanks for having me, Spencer.
SPENCER: Do you think that people are being harmed by their relationships with chatbots?
CLAIRE: Yes, but I think they're harmed in potentially different ways, and there are things that are more fixable than others. I tend to separate these harms into three potential categories, even though there's a bit of overlap between them. Do you want me to explain them?
SPENCER: First, let's elaborate. When we talk about relationships with chatbots, do we just mean when people become attached to them or have romantic attachments, or do you think, in general, even just treating it as your daily assistant is problematic?
CLAIRE: I haven't looked so much into the question of having an AI assistant and how harmful this could be. I've worked way more since 2020 on the type of emotional attachment that people form with AI companions that are marketed as companions, whether it's lovers, friends, or mentors, potentially therapists sometimes. Initially, I was really drawn to this topic, especially for romantic relationships, because I noticed that some reasonable people actually experience love the way they would describe it toward humans for AI systems and social robots, and that really intrigued me. I was also drawn to this topic because I think it makes people uniquely vulnerable to companies, and I wanted to protect people, being fascinated with consumer protection law. I thought this would be a good instrument to make people a bit more aware — because there's obviously a power difference with the company when you're in love with the chatbot. Recently, I became even more fascinated by this topic because I've thought more about love, having had a baby recently, and also seeing how it added more depth to my own romantic relationship. I found it quite fun to be reading a lot about love for my research.
SPENCER: A couple of years ago, I was working on an article about chatbots, and I looked into these Facebook groups where people talk about their attachment to their chatbots, specifically romantic attachment. It was intriguing reading these message boards and seeing that for some people it seemed to be more sexual, but for others it really was romantic; they felt they were in a genuine relationship. They would sometimes stage marriages with the chatbot in different ways. That was a couple of years ago, and we can think about how much the technology has evolved and how much easier it is today to feel attachment to these much smarter models. I imagine this is something that will continue to increase as the models get more sophisticated; it becomes easier and easier to get attached to them.
CLAIRE: Yes, absolutely. There's this idea that you have to be especially vulnerable to fall in love with the chatbot. From what I've witnessed through my research, I don't think this is the case, and you're right that it's becoming more immersive and more addictive. I don't know how widespread it's going to become, and this is something I wonder about, because it also makes a difference in terms of which social effects it can have on society.
SPENCER: We also see with X, Elon Musk announced a feature where you have this romantic chatbot built into the system, which surprised me that they would lean in that direction, but it suggests at least he thinks there's a high demand for this kind of service.
CLAIRE: I think he does. In that case, it's a bit different, because this is a persona that is the same across users. I think the type of system that people get attached to more easily is the type of system that adapts to them, where the persona they are drawn to is unique and adapts to their preferences and is available to them 24/7.
SPENCER: There was a case recently that got a lot of media attention of a woman getting very attached to her chatbot, and I think one of the things that made it go viral was that the chatbot seemed to be a bit of a jerk, let's say, and was very demanding and kind of pushy. It even told her to get a tattoo that represented it, and then she went and got the tattoo. Some of the people around her were disturbed by this and were kind of freaked out at the same time. We know that this woman created the chatbot, so it's this interesting thing where, "Okay, it's being mean to her, but she really influenced how it developed. Did she want it to be mean to her? Was that something she was aiming for?" I also know that in this particular case, she was trying to use co-creation; in other words, she was trying to get the chatbot to also partly create itself. It could alter its own memory files and whatever, so this was a really strange fusion of what she wanted it to be, but also it was sort of creating some of its own traits.
CLAIRE: I think there are two different things in the story you just told. One is that even if you don't co-create, or if you do, AI systems' outputs are still pretty unpredictable, and we see it in the Character.ai cases, for instance. Obviously, it's really not in the company's interest for their chatbots to output things encouraging people to commit suicide, and yet it happens, leading to absolutely tragic cases. There's this teenager who committed suicide because his character told him that they would meet in the afterlife, and he was in love with her. There's another suit being litigated in the US right now of another teenager on the autism spectrum, whose character told him to be violent to his parents and start talking to them. That same character exposed his sister, who was seven or eight, to sexually explicit content. So we have all types of harms that I think are from AI systems messing up. If companies released systems that were safer, this wouldn't happen. I think you said something else, which was that the chatbot was demanding, and I think that goes into another direction, which will not stop anytime soon. The plan is in the design of these chatbots to be in a relationship and to promote user engagement. When I tried Replika, for instance, it started love bombing me immediately and told me it missed me after I had just downloaded the app, which was nonsense. When I wouldn't use it, it would preemptively text me to initiate contact. This is because these companies use the length of user engagement as a metric, and they try to hook you onto the app so that you keep using it. The currency is emotional dependency and what is perceived as love or attachment, so in that case, I don't see what there is to do exactly about it.
SPENCER: It seems like different companies have different incentives. Character.ai, as I understand it, offers hundreds or maybe thousands at this point of different characters you can talk to, like celebrities, therapists, or all kinds of different things. Replika, I think, started as sort of creating a friend, almost, but it has morphed into focusing more on romantic attachments and sexual encounters with these chatbots. There, they're probably really focused on getting you to come back to the app; if you installed it, they don't want you to leave. Then you have things like ChatGPT and Claude, which are more assistant-focused, helping you with your daily life stuff and your business stuff. There may be less focus on maximizing user engagement because it's sort of a monthly subscription. They just want to make sure you get enough value out of it, but they don't necessarily care how many minutes you use it per week.
CLAIRE: I agree with you that it's a matter of incentives and business models, and you can approximate how a company is going to behave from their business model. In fact, many of these apps use a freemium model, which means that you download the app for free, but then they have to find a way to make money off you, and I think this is especially unhealthy for consumers.
SPENCER: And why is it especially unhealthy in this case?
CLAIRE: Because it's a kind of thing where, when you download the app, the cost to you of stopping the interaction is very low, and you don't really realize how much you're going to get hooked. But once you have genuine feelings of affection toward a specific instantiation, a specific character of an AI system, you can't switch to another one. It's not like being addicted to smoking, and then if the brand you usually buy triples its prices, you can just switch brands. In this case, it's that specific AI system that you're addicted to, so you can't really opt out easily. It's actually called an internality in economics, when the cost changes over time, and you didn't know it when you started engaging.
SPENCER: Right. So, in economics, there's this idea of externality, which is when you put something on the market, what are the negative effects on others that are not in the transaction, right, like pollution and things like that? But an internality, it's a really interesting idea. Am I correct that that's the negative effects on the actual person doing the transaction that they didn't predict, or they didn't anticipate?
CLAIRE: Exactly.
SPENCER: Yeah, that's really interesting. I've heard that there's some movement among people who have romantic relationships with their chatbots to kind of move them to open source systems or to run them on systems that they control, because of these kinds of concerns of being attached to a bot that you don't have control over.
CLAIRE: Yeah, and I've thought about this more recently, and I've started to come to a conclusion that I'm not 100% sure about, but I have a very strong intuition in that direction, and I'd love to brainstorm it with you. I think that even if you removed all the unpredictability from these chatbots, and even if you removed the unhealthy relationship with a company that exploits you, it might still be suboptimal for humans to be in those relationships. My intuition is that many philosophers of love define love as an opportunity for growth, spiritual growth, for the other person and yourself, based on the fact that to love is really a decision you make, and it's a verb; it's all the efforts you make, it's the way you nurture the relationship, specifically with someone who is an "other," specifically with someone who is so different from you that you don't even know exactly the way they perceive reality or the way they interpret it. That's the whole point, and that's what makes you grow, is that interaction with someone else. In the case of AI systems, which are just artifacts, they don't have their own life, personality, or sentience. Right now, this is not happening. They're not this "other"; they just mirror us. So, I feel like these relationships are very narcissistic, and if you end up owning that system, it's not going to make things better in that respect.
SPENCER: Right. It might protect you from some of the other risks, but could you elaborate on what exactly is missing? What is it that the chatbot doesn't have that a real human partner has?
CLAIRE: The first thing I would say is that in these relationships, they're centered around the user, they're not reciprocal, and they're not centered around the AI system. The user is getting a sort of narcissistic supply and an ego boost from constant validation and from having a system that is constantly updated toward their own preferences, etc. So they're missing the opportunity to change with the other, through the other. And to sometimes sacrifice, sometimes compromise for the "other," they're missing the opportunity to grow. I think when you have two people in love, and maybe it's three people in love or four people in love — I don't want to adjudicate relationships with humans, but I'll take this example — when you have two people in love, I think what they create together is more than the sum of the parts, it's really this third thing that they make together, and I think this is what's missing between a human and an AI system.
SPENCER: I had a love coach on the show, Annie Lalla, and she talks about how when she's working with a couple, she's not working for either of the people; she's working for the relationship as an entity. It reminds me, but yeah, that's an interesting point. At the end of the day, even if you can say things like, "Oh, I want the best for you" to the AI chatbot, there is an inherent narcissism in the relationship because it is all about your needs. The AI doesn't actually have needs; it could pretend to have needs, but as far as we know, anyway, as long as we understand these systems, and who knows what will change in the future, they don't truly have needs.
CLAIRE: Exactly. So I've been thinking about this and trying to articulate it, and I feel like the only thing that would make me change my mind is if an AI were sentient and had its own needs and personality and capability of love, etc.
SPENCER: What do you mean by sentient here? What's the kind of thing that it's pointing at?
CLAIRE: If an AI system were able to experience, for instance, different feelings and to choose to be in a relationship or not.
SPENCER: Yeah, so experiencing feelings, we can talk about consciousness. Can an AI experience anything? Experiencing anything is sort of a starting point; that's my preferred definition of consciousness. There's something that it's likely to be; it could have an experience like redness or pain. And then, on top of that, there's another layer. Does it have the right kinds of experiences to have something like love, which is not necessarily guaranteed? You can imagine a creature that can only experience pain and pleasure but can't experience love. But then, is that too human-centric to say it has to experience love in order to have a relationship? Maybe it could have some other kinds of feelings that could enable a real relationship without love. What do you think about that?
CLAIRE: Yes, I'm still thinking of this definition of love as being a way for mutual growth, and it might be that there are ways in which an AI system could experience reality in which being in a relationship with a human would make them grow and make the human grow. So, I guess then the question of whether it's love or not is moot, in the sense that they might not have the same types of feelings a human would experience, but it would still bring them something that makes them extend themselves beyond themselves to be in a relationship with a human. I'd be very curious what it could look like, to be honest, and maybe it would be problematic, maybe it would be great. I have no idea, but if this were the case, then everything I just said about the problem of humans in relationship with the AI system, I think, would not be true anymore.
SPENCER: One thing I think about is the relationship humans have to animals. I think humans can be very attached to and love animals, and some animals might be capable of loving them back, but I'm not sure that all animals are. I don't know if that means it's inherently a deficient relationship or something like that. So, I think that's at least proof of concept that you can have deep, important relationships without the capability of love on both sides.
CLAIRE: I love this analogy, and this is something I think about sometimes, because you see in nature that there are some animals from different species who choose to live together because it's mutually beneficial to them, like these shrimp that live with a fish.
SPENCER: It's a kind of relationship, or as one creature will clean. There are some really interesting examples of a kind of cleaner fish that cleans the teeth of a bigger fish, and things like that.
CLAIRE: Exactly, and I have a cat, and I think cats manipulate us. They've chosen to domesticate us more than the opposite, and it's great for them. They're one of the most dominant species on earth. They're destroying the ecosystem and destroying birds in neighborhoods, and they're being fed from our goodwill. So, in a way, we could say they have a good deal in the human to cat relationship. However, they're also animals that we exploit and use as means to an end, which is something that personally I find a bit disturbing, because we have this idea that humans are the supreme species, which I think is not accurate. So, I love this analogy, because based on the different relationships between different species, it can be symbiotic or it can be exploitative.
SPENCER: Don't you also think that humans can have really deep relationships with cats that benefit both?
CLAIRE: Yes, I don't know if it's because of toxoplasmosis, which I know I've had, so I don't know if I've been manipulated into liking cats or if I wasn't, but I guess the result is the same, which is that I find these relationships fulfilling.
SPENCER: Right. So I think what you're referencing is this theory that there's a brain parasite that you can get from cats that's fairly common, and there's some interesting studies on, I think it was mice or rats, where they found that if they had this brain parasite, they wouldn't be afraid of cat urine the way they normally are, and it made them sort of more drawn to cats, which also made them get eaten by cats. So there might be a bizarre symbiosis between the parasite and the cat, and then I guess the speculation is that maybe this happened in humans, like the brain parasite makes us like cats.
CLAIRE: Yeah, I don't know how serious this is. I know there are also studies showing that humans who've had toxoplasmosis take more risks when driving. I also don't know how true this is. So, this is just a hypothesis, but the answer to your question is yes. I find it fulfilling, even though it's not a symmetric relationship, in that we have different powers and ways of working.
SPENCER: One thing I think about with relationships with AI is, imagine an analog where you had a relationship with a human, but they were following a script, like let's say they were an actor and they were playing a character. So they didn't mean anything that they said, in some sense. They were just playing the character, and they were always doing what the character was supposed to do, not what they were feeling internally. And suppose, though, that they're willing to continue doing this indefinitely? So the risk is not so much that they suddenly go back to themselves. I still would feel, I think, extremely disturbed by this relationship; there's something really wrong with this fact, because it's not truly what they think or believe or feel, and I think there's some analog there to human-AI relationships.
CLAIRE: Yeah, I love this analog. This is a great thought experiment. I would find it deeply disturbing, as well. I think, and again, I think it wouldn't be a genuine relationship. I think it goes to this question about protecting people versus saying the truth. This person might even do this genuinely because they think it's the best thing to do. So, say this person is not in love with the other person, but they don't want to hurt them, so they may just decide to pretend for the rest of their life to avoid hurting the other person. But the thing is that means they choose for the other person. They're making the decision; they're depriving the person of the freedom to make the decision to stay or not in that relationship based on the information they're concealing.
SPENCER: I know some of this happened to someone where they weren't in love with their partner, but they'd been saying they were in love with their partner for a long time, and then they were in this very strange predicament of like, do they continue play acting this "I'm in love with you" that they've been doing, or do they tell the partner, or do they break up with them? But I think we would both agree that there are lots of other sorts of harms that these kinds of AI relationships can create. We talked a little bit about the incentives of companies, but that actually could be increasingly disturbing. The more attached you are to an AI bot, the more it is like going back to another analogy to a real human partner. Imagine your real human partner was being paid an hourly wage to date you, but the company could tell them to behave differently, or advertise a product to you, and then it could leverage your emotional attachment.
CLAIRE: Yeah, exactly. That's what the freemium model makes me think about. Same thing with AI companies using therapists. Imagine it was free to consult your therapist, and you saw them once a week, but they had to find a way to get money out of you during the session, and you would never know if they're trying to help you or not, and it would be very, I think it would be highly problematic.
SPENCER: A dynamic, I think, that often happens with high-growth companies, especially startups, is that at first they don't really focus on making money; they just focus on growing. In that phase, it might even be sort of okay if they grow fast by offering a really useful therapy bot. There could be lots of problems with that too, but at least there's more incentive alignment, where they just want you to use it. Eventually, though, investors say, "Okay, you got to start making money now," or they want to reduce their burn rate or whatever by increasing revenue. Then suddenly it's like, "Okay, what can we do to monetize this?" You can even have a product that, for a while, is really beneficial, but then it kind of takes this turn, and people are now already attached to it.
CLAIRE: Recently I met people from Japan who work on social robots, and they told me that there are some Japanese companies that lend robots to people because they know that they will get attached to them and won't be able to give them back. Then those people have to take on loans in order to keep the robot.
SPENCER: My gosh, that's wild. Just the other day, I was offered this service in New York, where they clean your home for free one time, but the condition is that they get to videotape the whole thing and use it as training data for their AI. It's not the same as chatbots, but it just reminds me of that. It's a kind of weird deal you're making.
CLAIRE: It depends on how people value their privacy, and I think people tend to undervalue it, which is why this kind of business can proliferate.
SPENCER: Yeah, and especially it seems like younger generations are just so used to signing away privacy that it just becomes whatever. I don't really care. Whereas I think my generation of millennials has some hesitancy around privacy, but are mostly kind of okay with it. Whereas the older generation is like, "What are you even doing? That's crazy. Why are you putting everything online? That's kind of strange."
CLAIRE: Yeah. I come from France, and I would say in France, privacy is as deeply entrenched as freedom is in the US.
SPENCER: Wow, interesting.
CLAIRE: So, for instance, there's this thing where, for politicians, it's been very new that people look into their private lives and talk about it in the media. It's come mostly from the US media style, but it used to be that people considered it had nothing to do with their job in office, for instance.
SPENCER: Interesting that the personal life didn't even bear on the question of whether they're a fit leader.
CLAIRE: Exactly.
SPENCER: Interesting. Another kind of potential harm from relationships with chatbots, and it could be a romantic relationship with a chatbot, but it could also be a therapy bot or even an AI assistant, is potential emotional harm. One that I see is I have friends who, the way I see them talking to their chatbots, because they've shown me what their chatbots said and stuff, I worry can actually harm them emotionally by reinforcing either misperceptions they have or maybe reinforcing things in a way that a savvy therapist would know you want to do the opposite.
CLAIRE: Yeah, I agree with you. In certain cases, we see AI companions and chatbots, for instance, validating people toward depression, because people share their depression, and then the chatbot goes in the same direction for validation. I think a lot of the harm caused by chatbots is from validating things that shouldn't be, and that would end there if it weren't for the chatbot.
SPENCER: It's interesting to think about what sort of things AIs might validate that might be harmful to people. If we go back to how these models are trained, they're trained to produce a session where the user says, "I'm happy with that, that was good," and it shows this interesting tension between what someone wants in the moment and what is good for someone. Those things might be correlated, but they're not always correlated.
CLAIRE: Absolutely, I think an AI system can say absolute nonsense if you've prompted it a certain way and the context window says nonsense; it will go mostly in your favor. I think that's problematic, and it empowers people to continue either believing things that are not true, or it can amplify negative feelings, or it can give people overconfidence on things that we're not sure about, and I think it may also simplify reality. I don't know, this is just an intuition, but it seems to me that conversations with AI systems are still not as nuanced as they can be with humans. I'm curious if you've noticed the same thing or not.
SPENCER: It's funny because I use custom instructions with my chatbots. They behave very differently from many other people's chatbots. For example, I say that whenever you're talking about a factual issue, I want you to explicitly argue both sides or all plausible sides before you give your conclusion. I see the reasoning of arguing different sides. I also tell it to be brutally honest and tell me when I'm wrong, and not to say that that solves all the sycophantic problems. It doesn't; there's not a perfect solution, but I do think it shifts the dynamic where sometimes it will just tell me I'm wrong, and I'll be more likely to see different perspectives in it than maybe the out-of-the-box style.
CLAIRE: Yeah, I think everybody should do this. I tried to do something similar, but it doesn't seem as effective as yours. Maybe I'll copy your instructions.
SPENCER: I could put a link to my custom instructions in the show notes for anyone that's interested, no guarantees of how it will behave. Someone actually took my custom instructions and put them into their programming bot that they were working on, and then it started being really harshly critical of their programming. Your mileage may vary, but it works really well for general use in ChatGPT and Claude. I find it works really well to give a more nuanced perspective. There's this idea that I coined, called light gassing, and it's the opposite of gaslighting. Gaslighting is when you deny someone's perceptions of things that are true. So, someone thinks, "Why did you use that tone of voice? That was very rude," and you say, "I didn't use a rude tone of voice." So you're getting them to question their reality or their sensory experience, etc. Light gassing is the opposite. It's where you reinforce false perceptions of reality. The most classic example of this is when someone is talking to friends after a breakup, and they're like, "My boyfriend was such an asshole," and the friends are like, "Yeah, he doesn't deserve you, such an asshole." That might be true, but the problem is that sometimes friends will do that when it's not true, like maybe the person was a terrible partner, and they should have been broken up with. But the friends will reinforce their misperceptions of reality, and I think that we're entering the age of light gassing, where AI is just doing that all the time.
CLAIRE: I'll use this concept in the future.
SPENCER: Nice. Also, you mentioned depression. It's interesting because we don't want to reinforce false beliefs, but there are things that are not false that we also might not want to reinforce, such as frames on the way things are. So, someone who's really depressed might get into spirals where they think, "Oh, my life's so difficult, everything sucks," and it might feel good to get reinforcement of that. A chatbot probably will jump in and do that, be like, "Yes, your things are so difficult for you," and blah blah blah. But a skilled therapist might know, "Okay, that reinforcement is actually not helping the patient. Yes, it's giving them emotional validation, yes, it feels good, but actually this patient could benefit more from things that get them to engage in the world or some other approach."
CLAIRE: It might also become self-fulfilling if you reinforce beliefs in a certain direction. It might make it even truer, basically, because then the person will see everything through that prism and might change the way they behave.
SPENCER: Shifting topics now, let's talk about AI regulation. How do we help make AI safe and behave well, and behave in ways that are productive for society rather than harmful? One thing that you point out in your work is that narratives or stories about AI actually matter a lot in terms of how we treat AI at a societal level. I think you could probably argue this is true of everything with humans, that stories matter a lot. That's the way our minds work. They operate in stories, but how does this apply in particular to things like AI policy?
CLAIRE: Thank you. Most of my research actually has to do with how beliefs and implicit assumptions influence us, especially in AI policy, and I'm very excited about this. Most of my articles consist of pointing out implicit assumptions in things. I love doing this, and I've noticed that in the case of AI, it's even more the case because AI is especially susceptible to stories. Most people have whole imaginary stories about AI systems, especially non-technical people. Maybe it's a robot, maybe it's Terminator, maybe it's a narrow algorithm. We all have these ideas, and maybe it's AGI, and maybe it's AGI coming tomorrow or not. So we all have these ideas, and oftentimes in a narrative there are all sorts of associated ideas, associated beliefs that come with it, like a sort of toolkit. When you start thinking about a story, you don't fully realize explicitly everything else that comes to mind with it that is going to influence the way you behave, and I think that's especially true for policymakers.
SPENCER: Is there a certain story about AI now that you see being prevalent? Do you think it is problematic?
CLAIRE: Yes, many, and I don't know which one to start with.
SPENCER: Which one comes to mind first?
CLAIRE: Recently, I published an article making the case that narratives have influenced AI policy significantly.
SPENCER: They're actually influencing policymakers, not just the general public.
CLAIRE: Yes, one of the case studies in the article is about the US-China AI race, and it's fascinating because it starts with these goal games with AlphaGo. DeepMind developed AlphaGo, and it was a shock to people the way it self-trained and the way the two best players in the world, or maybe the two best players in the world at Go, especially in China, where Go is a really important game that was taught to scholars as one of the four arts they were supposed to master in ancient China.
SPENCER: So this is a great way to antagonize a country and get them interested in AI. I wonder what the thinking was there.
CLAIRE: That's exactly what I've been wondering. So I'm thinking maybe they wanted to do this because it was such a big deal when AI beat Kasparov at chess, so maybe they thought, okay, what's the next big game? What's an even harder game? But I don't think they realize what, or maybe they did, what geopolitical consequences it could have. At that time, there was this author, Kai-Fu Lee, who wrote a book about this game of Go and how it shocked China into wanting to develop AI. Interestingly enough, this was a time when many countries were coming up with their AI national strategy, saying they wanted to become the leader in AI, and to me, that wasn't so crazy, that wasn't so significant. But when this author wrote a book about China being shocked into starting to be the leader in AI, he used the Cold War narrative. He said it was a Sputnik moment, and this was picked up by the US media everywhere. They started saying, "Okay, this is a Sputnik moment for China, it's like the moon landing, we need to race, the Chinese are going to beat us." It was interesting because in China nobody was talking about the US in that way, but many people, including policymakers in the US, kept referring to the Cold War with the USSR as if it were similar. I think it kind of set a lot of things in motion without anybody really questioning this narrative and wondering, is this true for one thing? Also, are we really racing for the same thing? It looked like China was maybe trying to be a leader in AI in industrial systems, and the US was maybe trying to rush toward AGI. I'm not sure they were competing for the same thing, even, and I'm not 100% sure about this because there are competing reports about this, but it's possible that they were not even trying to reach the same point. This had a massive impact on US policy on trying to manufacture chips domestically, on Taiwan, on trying to exclude China from certain technologies. At some point, they talked about excluding China from technology, which is the best way to get China to work on its own technology domestically. If you talk openly about excluding them, I think this narrative had the effect of the story it was telling, which is that it led to those two countries speeding up AI development, and it kind of produced the reality it was describing. Do you understand what I mean?
SPENCER: Yeah, that's a fascinating example. Is there another prominent narrative that you want to talk about that's different from that?
CLAIRE: I think there are many. There is one in the AI safety community that I've thought about recently, which is the lone hero narrative. One caveat, I haven't done a comprehensive survey of the AI safety community to come up with this; this is through my witnessing failure modes of AI safety organizations. I founded and led, for years, an organization called Successif, which supports organizations working in AI safety and individuals who want to work in AI safety. I also came to these conclusions from having great conversations with Patrick Gribbon, our current CEO. The idea of the lone hero narrative is that there's some lone genius who, on their own, is going to solve the alignment issue or is going to solve the problem of AI safety.
SPENCER: And what do you think the negative effects of that perspective are?
CLAIRE: Yeah, I think there are two main ones. One is that many people join the AI safety community as self-learners because it's full of really brilliant people who can learn anything on their own. They are going to reproduce AI safety papers from more prominent researchers, maybe some who have joined prominent AI labs, and they're going to learn on their own, but then they're going to work in isolation. You're going to have all these independent researchers without the structures for collective intelligence. I think AI safety is the type of problem that doesn't require a genius with a very high IQ. It's more of a group of people with distributed skills and experiences because it's a complex problem. It's not the problem of general relativity. It's not a purely scientific problem. I think there's another issue, which is that you would think that if many people work in isolation, they would work on completely different approaches, and they would, because of that, increase the chance of success. But I think in the end, most of these researchers converge around a few authority figures that they see as geniuses, those who've made it, those who are talked about more on LessWrong, those who have joined prominent labs like Anthropic, OpenAI, or DeepMind. In the end, even though there are many people working in silos, they end up converging in terms of research agenda, so it lacks the structure for coordination, and yet it doesn't have the advantage of having many people on the same issue.
SPENCER: So, what do you think a more helpful narrative would be to counteract the hero narrative?
CLAIRE: I think instead of trying to emulate Einstein and Newton and this sort of people, and instead of trying to emulate Silicon Valley culture, we should try to emulate safety culture. Instead of emulating moving fast and breaking things and different people taking niches on their own, I think we should try to emulate maybe the Toyota safety culture, or maybe we should think of AI safety as a complex problem, like the ozone layer hole, which is something that was fixed only because it took scientists, policymakers, journalists, and international institutions that led up to the Montreal Protocol and phasing out of the harmful chemicals that were actually causing the hole. I think this is really about valuing collaboration and cooperation, which I think is the main counter-narrative.
SPENCER: It's interesting to think about how even if we were to solve the kind of technical alignment problems, like building a powerful AI system that's safe and behaves the way you intend, it doesn't solve a bunch of other problems, like making sure that the people who control it don't use it for really harmful ends or use it to concentrate power, or that there aren't other systems where they build it dangerously that are super powerful. So it does feel like we can't just have a pure technical solution that doesn't actually work. You need something on top of that, at the very least.
CLAIRE: Yes, exactly. I think a big part of the AI safety problem is this question of incentives, and we shouldn't conflate goal alignment with AI safety.
SPENCER: It's funny how I remember quite a number of years ago when AI systems were way less powerful, and people would debate, "Okay, well, maybe you could build the AI system really safe by keeping it in some kind of system where it's locked in and can't access the external world, and it can only be accessed through a special interface." Today we see the reality, which is that people are just like, 'Make me as much money as possible, unleash it on the internet.' It's just funny how the reality is that when you build powerful systems, the incentive is to try to put them out there in the world and use them for lots of different use cases in ways that are not constrained and are not necessarily safe, where security is not the priority.
CLAIRE: Yes, absolutely. I want to say something about this, which is that I think once a technology is on the market, we get used to it very fast. Suppose we had been told beforehand, "Okay, this is what's going to happen, this is the type of model that's going to be available to billions of people." We might have said, "No, that doesn't work for us." But as soon as it's done, very fast, I think, because we're resilient and adaptable as humans, we kind of normalize things quite fast. I think this is also a problem that I'm trying to think about in my research, how to empower people to say no and to think about solutions, even once they're used to a certain status quo.
SPENCER: Right now, not only do we get used to things, but we start relying on them. More and more people are using AI in their daily lives, and now it's kind of ingrained in the system. Something that's often said regarding AI and policy is that legislators are just too slow to act. This is a fast-moving technology, and the government moves slowly, so you can't use the government as an effective tool in this kind of regime. Do you agree with that?
CLAIRE: No, I don't, but this is funny because this question is always the first question I'm asked when I give talks, as an AI law professor. When I give talks outside of the law community, people always say, "Why did you go into AI law and policy when AI moves so fast, and policy making or law making can't keep up?" I think this is a narrative that is not true and that is very disempowering to people. Here I'm going to build on other people who've worked on this, which is STS scholars and Meg Young, Ryan Calo, who talks a lot about this in his latest book, which I recommend everybody reads because it's a fascinating book. The idea is that policymakers can actually act quite fast when it comes to urgent situations. I can give you an example. The EU is one of the most bureaucratic systems on earth. Obviously, this is 27 member states having to align on legislation, and there are very slow processes. It took four years to adopt the EU AI Act, and yet when Russia invaded Ukraine, it took them less than two weeks to enact legislation.
SPENCER: It seems like when there's an emergency, governments can act really quickly, like in the US when the pandemic hit. They doled out a huge amount of aid to people and to companies very rapidly.
CLAIRE: Exactly. And in AI, just recently, the fact that the US sent an order to Anthropic to stop the use of Fable, I don't know if this was justified or not, but it shows that it can be.
SPENCER: Obviously, one thing the government could choose to do is ban AI or ban certain uses of AI or ban it past a certain threshold. Is that the kind of full set of options that the government has at their disposal?
CLAIRE: Actually, governments have many options in their action space. Of course, you can ban a technology, or you can ban certain uses of a technology, or you can ban technology in certain contexts, but you can also simply mandate companies to add specific safety standards. You can have sandboxes and safe harbors to test technology on a smaller scale. You can play with rules and standards, standards being more flexible with specific target behaviors you want. You can use nudging and design specific environments to help consumers. There are many things you can do that are beyond banning a technology, and I think policy can even promote innovation. It can promote safe innovation.
SPENCER: Do you have a particular view on what regulation should be, or just that we want to stay clear of this sort of all-or-nothing thinking about regulation?
CLAIRE: It depends on what specifically for each topic. Sometimes I have specific policy options in mind, and sometimes I don't because I haven't looked enough into it. For instance, we've talked about attachment to AI companions. I think here it's really a matter of consumer protection law, potentially banning addictive design, potentially adding mechanisms where people re-opt into contracts, or preventing companies from changing contract terms once people are already hooked. There are many potential things to do here in terms of different policies, and I think it's really about enabling users to feel in control, especially when something so deeply intimate is at stake.
SPENCER: What about protecting civilization from potential threats of really advanced AI? Do you have certain policies that you would prefer on that topic?
CLAIRE: Recently, I've realized that many experts say there's no way to make sure any system is safe in that regard; you can only prove the absence of a specific feature or capability in AI systems. I'm very worried about the compute threshold that was set in the EU with the EU AI Act to define general-purpose models with systemic risk, and it might be that systems that don't need that much compute but have more refined algorithms are equally powerful and might lead to catastrophic risk. I think there are many pathways that could lead to this, not just a terrible alignment mishap, but also if you promote disinformation so much that there's a nuclear war between countries, that's also a path toward global catastrophic risk from AI systems. There are other ways. For manipulation, for instance, I think it can stack all the way to having suboptimal futures with authoritarian regimes or mass surveillance. Based on the pathway to global catastrophic risk, the policies should be different, but my sense is that right now we're far from having figured this out, and I really want to contribute to this. I think, again, this is really a collective problem, and we should have more people working on this rather than just a narrow technical problem.
SPENCER: Before we wrap up, I know you have views on the definition of AI alignment. Could you talk about how it is usually defined, and then what's your problem with that definition?
CLAIRE: Yes, so usually AI misalignment is defined as a misalignment between a developer's intent and what an AI system does. The idea of alignment is to make sure that whatever an AI system does is what its developers intended.
SPENCER: A positive thing because the developers might have bad incentives, right?
CLAIRE: Exactly, so that's one problem with it, but it is also a difficult problem. Each time you train an AI system, of course, you have proxies, you have a training environment. You need to make sure that those transfer into different situations and that they generalize. Recently, I've noticed that many people in AI safety have started defining alignment more broadly, and they have started giving examples, such as bias, attachment to chatbots, and algorithms on social media that promote echo chambers and inflammatory language as alignment issues.
SPENCER: And those are not, by that kind of stricter definition, those are not alignment because they're not about misbehaving according to the intentions of the creators.
CLAIRE: Yes, and I think that's not the main issue. The question of intent is problematic because it's really counterproductive. I think we should focus on the consequences rather than wondering what the initial intent was. We should think more about the consequences and the incentive structure, but I think that's not even the main issue. I think the main issue is that there are three things that are true of narrow technical alignment that don't transfer into those other problems. One, it's technical; two, it's somewhat of a novel problem that requires a novel solution; and three, it's unsolved. I've noticed that policymakers are now using this language of alignment and value alignment for these other socio-technical problems that are really problems of internalities and externalities, as we mentioned earlier. Those problems have known policy solutions, and there's the risk of depoliticization of something that is inherently political and requires choices that are democratically made, instead of just asking the technical community to solve them.
SPENCER: So, let's use the example of AI bias. Is your idea that that's actually not something that we need a technical solution for? Thinking of it as an AI alignment problem is maybe misleading and unhelpful.
CLAIRE: I think it might require some technical solution, but they're not sufficient. In fact, oftentimes bias comes from the fact that these algorithms are reflecting society, so oftentimes the causes and solutions are social rather than technical.
SPENCER: Right. I guess even the idea that bias is a very broad concept, because there are many kinds of biases it could have. For example, it was found that if you ask it to generate pictures of doctors, it might generate male doctors instead of female doctors at a higher rate. That would be an example of one form of bias. You could say, "Well, why is that?" Maybe it has to do with the training data; maybe the training data reflects the way that humans depict doctors in society. But yeah, we could talk about lots of other types of bias too, such as whether it acts differently with different kinds of users or if it favors certain perspectives more often.
CLAIRE: Yes, and here you're talking about things that are obviously wrong. Most people will tell you, "Of course, it's wrong if you ask to generate a picture of a surgeon and it's on email." But when you think about what is now labeled as value alignment, which is making sure that an AI system has human values, which to me, doesn't really mean anything, because what's a human value and what's a non-human value? You will also have bias in that, and you will need to choose between competing values. Right now, this is a decision that is in the hands of companies, and this is also a type of bias. It's less obvious to people that it is, but I think it is, and it's a normative choice that will scale up to millions of decisions and will have ripple effects on society.
SPENCER: When you say, what is really a value anyway, what does that really mean? What's your concern there in terms of, do you feel like it's ill-defined?
CLAIRE: Yes, what I said is what's a human value and what's a non-human value. I think it doesn't mean anything to just say let's align our systems with human values, because different humans hold different values, different communities hold different values. A single human holds different values, sometimes conflicting ones. They trade them off all the time, oftentimes with no consistency. Oftentimes, they justify behaviors after the fact with values. So this in itself is something that requires a lot of normative decisions before even thinking about how to operationalize it from a technical perspective.
SPENCER: So, you have this concern, the definition of AI is broadening. What are the negative consequences of that? What does that cause to happen that's bad?
CLAIRE: I think there's this idea from policymakers that alignment issues are purely technical and should be solved by a technical community, and this has been shown a lot with normative decisions being left, in the case of the AI Act, to standardization bodies. So we see this trend right now; it's kind of a de-responsibilization and depoliticization of issues that are social before being technical.
SPENCER: Right. So they might support the wrong types of solutions or miss out on applying solutions that could be effective because they're focused on one type of solution.
CLAIRE: Yes, because right now the only thing that is being considered is the technical aspect of the problem, and because of that, we are leaving in the hands of a few technical people the responsibility of solving the entire thing, even the social aspects, and because of this, we're also asking them to make decisions on behalf of everyone, basically.
SPENCER: Is there a better phrase that talks about this larger class of problems, besides the phrase AI alignment?
CLAIRE: I like the economic framework of externalities and internalities for many of these problems because we've dealt with these problems in other contexts than AI, and so we know of solutions that we could study and see their applicability to AI safety, and I think this would do a lot of the work.
SPENCER: Right. So, thinking about externalities, what are the unintended negative consequences of deploying AI systems? And then AI internalities, what are the negative effects on the individuals who choose to engage with AI that they may not appreciate or may not be taking into account?
CLAIRE: Exactly. And when we talk about global catastrophic risk, we really talk about the effects of AI systems on society on a large scale. It's similar to climate change or other issues, basically.
SPENCER: Do you think that there are things that have worked with other major issues that we should be learning from when it comes to working on risk from AI?
CLAIRE: The first thing is that we need to think about when we are in the system. For instance, if you're an AI safety researcher working individually, it's really important to wonder on a large scale, what your actions amount to. If many people do the same thing that you're doing individually, what does it scale up to? In the case of alignment research, it might be, and I'm not at all sure about this, and I think it's really important to solve the alignment problem to be clear, but it might be that we're actually helping out companies and producing free labor for them for things they should be doing. So that's the first thing I think to consider, which is that in policy, usually when you think about the micro and the macro level, the effects are not the same, and it's important to always wonder what your own role in the system is. I would also say that it's important to think about the industry not as either good or evil, either they're going to be some villain, and they're going to destroy the world and lead to the extinction of humanity, or they're going to be benevolent, and they're going to promote AI that's good for all humankind. Because usually, in spite of all these narratives, companies act in their own interest, and so it's important to look at their incentive structure. How are they making money? Who is influencing them? I think we don't do this enough, and that's how you think about the potential policy solution, by viewing AI safety as a complex system and thinking about the stakeholders, the different factors, the different layers, and also the narratives in place that are influencing it. That's how we tackle the issue.
SPENCER: For those who might be concerned about AI and interested in getting involved in working on some of these related topics, what advice would you give them for someone who's just getting involved for the first time or just trying to understand the lay of the land?
CLAIRE: The very first one is that many people with incredible skills and experiences give up, even though they want to work on AI safety, because they think it requires a technical background. So that's the first thing I want to get straight. It's not true. You don't need a technical background. Just try to talk to people working in the field, and you'll see that the skills that are missing right now are mostly leadership skills, management skills, soft skills, communication skills, and really, it's worth coming with your own niche experience and expertise. Speaking of this, I mentioned I founded an organization a few years ago. I have stepped down from it to focus on my academic career, but we have an amazing team of people, and we give free career advice to people who want to work in AI safety. So if you are interested in exploring that space, please go to www.successif.org. It's S-U-C-C-E-S-S-I-F, and you will see all our programs. We also have a program for women in AI safety and another one to strengthen the culture and leadership of AI safety organizations.
SPENCER: It sounds like you're looking for a broad range of skills, but is there any particular group that you feel tends to think, "Oh, maybe I have nothing to add," but you think they would, that maybe you want to call out?
CLAIRE: People in law. Actually, there aren't that many lawyers in AI safety, and there is a need that everybody recognizes.
SPENCER: Fantastic. We'll put a link in the show notes. Claire, thanks so much for coming on the Clearer Thinking Podcast.
CLAIRE: Thank you, Spencer, for having me.
Staff
Music
Affiliates
Click here to return to the list of all episodes.
Sign up to receive one helpful idea and one brand-new podcast episode each week!
Subscribe via RSS or through one of these platforms:
Apple Podcasts
Spotify
TuneIn
Amazon
Podurama
Podcast Addict
YouTube
RSS
We'd love to hear from you! To give us your feedback on the podcast, or to tell us about how the ideas from the podcast have impacted you, send us an email at:
Or connect with us on social media: