Cybersecurity & Tech

Lawfare Daily: Consent in the Age of AI

Renée DiResta, Kate Klonick, Elissa Redmiles, Jen Patja
Friday, July 17, 2026, 7:00 AM

Discussing technical proposals aimed at preventing AI systems from generating exploitative content.

Lawfare Senior Editor Renée DiResta sits down with Senior Editor Kate Klonick and Elissa Redmiles, an assistant professor of computer science at Georgetown University. They examine the people who create AI-generated sexual content and whether prominent technical proposals can actually prevent AI systems from generating exploitative content.

For further reading:


Please note that this podcast discusses sexual violence and the harms of image-based sexual abuse. Listener discretion is advised.

To receive ad-free podcasts, become a Lawfare Material Supporter at www.patreon.com/lawfare. You can also support Lawfare by making a one-time donation at https://givebutter.com/lawfare-institute.

Click the button below to view a transcript of this podcast. Please note that the transcript was auto-generated and may contain errors.


Transcript

[Intro]

 

Elissa Redmiles: "Is this media fake?" is different than, "Is this media of me?" And that's where I don't think the existing kind of report and attestation piece solves the kind of search problem or prevention of the content being made.

Renée DiResta: It's the Lawfare Podcast. I'm Renée DiResta, senior editor at Lawfare, and I'm joined by Kate Klonick, also senior editor at Lawfare. And our guest today is Elissa Redmiles, an assistant professor of computer science at Georgetown University.

Elissa Redmiles: So I think there's a lot of attention that would need to be paid to, like, what qualifies something as a sufficient safeguard, and how do we set people's expectations appropriately for how much that's gonna protect them individually versus, kind of, mean something about mitigation in a specific product.

Renée DiResta: Today we're talking about the people who create AI-generated sexual content and whether prominent technical proposals can actually prevent AI systems from generating exploitative content.

[Main Podcast]

So I, I wonder if you might wanna start by telling us a little bit about the research that you've been doing. I know that most conversations about AI-generated imagery begin with a focus on detection, but, you know, you've been arguing that that really misses some central facts. I'd love to hear you just start by telling the, the listeners how you think about this problem.

Elissa Redmiles: So when it comes to AI generation of non-consensual intimate imagery, sometimes called revenge porn, one of the main harms is on the use of someone's likeness without their consent to create sexual content of them.

And in particular, as a computer scientist, I was looking at different kind of claims that companies were starting to make. For example, like, DALL-E has a claim about using advanced techniques to prevent photorealistic generations of real people. Civitai, which is a website for hosting different kind of model objects that got found to have a lot that make specific people started making some claims that you could opt into preventing models from being shared that could generate your likeness.

And as a computer scientist, that made me wonder, can we actually do that in practice? And what are the flaws that are gonna kind of pop up if we try to automate that type of protection?

Renée DiResta: We actually just saw this from Meta. We saw some interesting decisions about not only opting in, but in fact encouraging. I got a push notification from Meta AI telling me I could go create content of people just by “@-ing” their Instagram handle and then, and then it would go and it would generate content using the person's public Instagram account as the base layer. And this was a very interesting, it was a very interesting product decision, I thought in light of the potential for abuse. I'm curious what you thought about that call.

Elissa Redmiles: Yeah. I share your concerns about that choice particularly given that we've seen with Grok there's been a lot of instances of- nudification or at least sexualization of individuals by using a very similar affordance where you, like, @ the individual's account, or it's just kind of a reply to content that they've shared. And so when I saw the Meta release, I was kind of immediately thinking about that kind of situation and was concerned that people may not expect that their content can be modified.

And I think this goes beyond nudification, right, even for various job or reputation concerns people may have. And actually, in particular, we're seeing people being asked to kind of set social media to public, right, for various kind of immigration or border reasons. And so that may put people in a really challenging position. The other thing I would say is, you know, even after, say XAI has mentioned trying to implement new technical controls, we're still seeing many cases of nudification.

And in a report that we'll have coming out soon with the Center for Democracy and Technology, we took a look at all of the kind of technical ways that foundation model providers are trying to prevent this generation, and we really don't have reliable means of doing that right now. And so, you know, while I'm sure that Meta's intent is not to allow for that, there's definitely a risk with a feature like this.

Renée DiResta: One of the things that I have followed in your work over the years is that you write a lot about how sometimes abuse isn't necessarily focused on the substance of the content, but whether or not the person consented to its creation or distribution, right?

And that might mean that there is a nude image that the person actively made themselves and chose to upload voluntarily, or a person who is, as you know, not nude, but put into, you know, a potentially dicey image situation and that did not consent to that and doesn't have the same access to having it addressed because the, the image generator will in fact generate it. So can you tell us a little bit about how you see the consent problem and its differences from the content classification problem?

Elissa Redmiles: I think the idea with consent that we've seen even before use of AI for non-consensual intimate imagery was so large, we had done work with, as you said, victims who had not taken content but maybe had it made of them using secret cameras, other approaches, as well as folks who had taken content of themselves either for recreation or for commercial purposes.

And what we saw was for folks who are sharing this kind of content commercially some people viewed the re-sharing of their content or now the editing of that content as, like, a commercial threat, like theft. Like, basically you've taken something that I make commercially and you're trying to make money off of it yourself. Whereas others saw it as both theft and a kind of sexual harm, privacy harm, victimization in the same way that folks who were not commercial actors would have felt about their content being disseminated or themselves being depicted.

And I think that's one of the reasons why the idea of likeness or being identifiable as a real person and whether or not that person got the opportunity to consent is so compelling to me in this space. It's also compelling to me given kind of how we think about pornography in the past and people being able to provide consent forms and that being kind of a huge distinction or, like, point of process at least in the U.S. So I think for both of those reasons it resonates with me.

Kate Klonick: So just to be clear, NCII and why it's not called revenge porn is because the idea, anymore, like, that's kind of like it's like, it's like a, you know, not a politically correct term among those who kind of study it anymore. And I've always understood that to be kind of true because the idea of pornography is that one does it consensually, and inherently, like, pornography is kind of thought of as consensual.

Whereas NCII, non-consensual image, intimate images, is essentially, like, inherently the opposite of that. And so it does, a nd also because revenge is not, some often we find it in revenge situations, but revenge isn't the only way that, like, non-consensual images show up, intimate images show up. So I just kinda, is that correct? Is that kind of the definitions that we're operating under?

Elissa Redmiles: Yeah, absolutely.

Kate Klonick: Yeah, so I wanna kind of pull at a string here, which I think is super interesting from a legal perspective, which is this question of focusing on consent. So from, like, a legal perspective, I think there is something really interesting about kind of this model that you bring up. One is kind of a theft of property, so to speak, like a theft of something that you can commercialize or, or something like that, like a content kind of theft. And then there is, you know, this privacy-based regime.

So those kind of, like, end up in different buckets in the law, I would say. But in both there's, like, an evidence issue, right? Trying to show or demonstrate that consent has been given either explicitly or implicitly and kind of to track that.

One of the ways I know platforms have dealt with this in the non-consensual image context is to just believe the person who says, "This is my image, and I didn't consent to this, and so please take it down." Is that as feasible going forward? Did you do any kind of research on that? Do you see just kind of a, a believe it, like, in terms of just how platforms should react to this, like, is that actually going to be an actively useful framework going forward?

Elissa Redmiles: Yeah, that's a great question. So, I think there are a few issues here. So one is once the content is created, what do we do about it, right? And so if that content is shared on a platform that has a reporting stream and people are able to report, then certainly kind of attestation is one approach. I think the one concern we sometimes have with attestation is that there have been cases of, like, widespread sort of activist-based reporting of, say, sex workers' content for takedown. And so that becomes, like, a question of how you balance that.  But typically, there are other signals when that's happening where you may be able to kind of balance things out.

So I think for content that is up, that's, like, a reasonable approach. I think the problem, though, becomes that this content often spreads to platforms that are specifically for hosting non-consensual imagery or are just not interested in doing takedowns. For example, thinking places like 4chan or a bit more hidden.

The other issue that can come up is people discovering this content has been made in the first place, and that's a big issue for the AI generation of this kind of imagery is how are you gonna go and detect that you're depicted? And that kind of comes back to this question of, like, can computational approaches do detection of you or any real person in generated media in a reliable way?

And that's a different question than the one we usually ask, which Renee brought up, which is, like, “Is this media fake?” Is different than “Is this media of me?” And that's where I don't think the existing kind of report and attestation piece solves the kind of search problem or prevention of the content being made.

Renée DiResta: So technically speaking then, when we think about the technology for detecting consent, is that even, you know, it sounds even weird to say that. Or how do you credential an image as in, in a way that is, like, machine readable? We can do this, of course, right? This is, this has come up a bit. But we, we're seeing quite a lot of, a lot of platforms struggling with labeling regimes, with detection regimes. There's a lot of different things that they're being asked to build simultaneously. How should we think about essentially credentialing content as consented imagery that is not NCII?

Elissa Redmiles: Great question. So I think about this in these kind of two components. The first is how do I recognize that this image that has been generated or posted is an image of you? And how do I do that in a way that is privacy preserving, preferably? And then the second piece is how do I know whether that is an image that you consented to being created?

So, for the first piece of, like, is this an image of you one of the first things that may come to mind for computer scientists is facial recognition, which are, right, the technologies when you go through TSA now, they often want you to stand in front of a camera. They're using what's called 1:N facial recognition, where they're searching for you in their database of faces. And essentially what they're doing is they're comparing the image of your face that they just took, doing a similarity score check against the other images in the database, and presumably your face or name should come up as the most similar one.

When it comes to generated images, a few things change. The first thing is that the underlying distribution of pixels, so all the little points that makes up the picture that distribution is different in a generated image than in a photograph. And there actually has not been pretty much any research on how well facial recognition algorithms will identify individuals in generated photos.

And the other thing that comes up is that facial recognition was originally created for these kinds of border-type cases where you wanna be, like, perfectly sure this is the person that you think it is. When it comes to something like NCII, I think the concept is a bit fuzzier, right, in terms of, like, can a reasonable person recognize the individual in the image?

And so how well something like facial recognition will map to human perception is an open question, especially because it depends, like, is this you think it's you in the photo? Is it someone who knows you well in real life thinks it's you in the photo? Is it just, like, your online fans think it's you? All of those can add noise to this kind of machine process. So that's one aspect.

I think in terms of, like, attaching consent credentials to an image certainly there are different approaches like we've seen for AI labeling to do kind of cryptographic watermarking that attaches content to an image. All of that relies on the platform where the content is being posted, keeping that metadata and respecting it. And so that is, like, an open question that one might have after the generation phase.

Kate Klonick: So, I think this is really fascinating that you did all of this kind of research between the various different types of ways that we move from something like TSA. When you say 1:N, like, I think that, like, you know, most people are so creeped out by things like TSA and its incredible quick speed and its incredible accuracy, mostly because they think that, like, the TSA is running your face past a database of all known human beings on the planet, and it's not. It's, like, the flight list, like, the manifest, right? Like, it's not in any way that, like, kind of, which I think is actually kind of a super important point to make.

And so I think that there's this really fascinating delta between our ability scientifically and through qualitative, I don't know, I guess you could say proof or mathematical proof, whether something is AI-generated or whether it's real. But actually, that doesn't end up mattering for the vast majority of, of things.

So, like, authenticity in both one sense is so crucial to certain types of, like, facial recognition and certain types of, like, of AI generation stuff, but it's actually not on point to, like, kind of this problem with non-consensual intimate images because the damage is really just the idea that it's you, right? That it's a reputational damage, it's a normative damage, it's kind of a s- breaking, like, you know, being shamed or, or kind of, you know, embarrassed by something that you're caught on doing or that is revealed about you.

How did you kind of capture that sociological component of this? I mean, you're a computer scientist, so I'm just really curious. Like, you, you live in the mathematical proof, the authenticity proofs of, like, these things. So how does one measure, how if you, you know, if you're designing a study, how does one measure that other component?

Elissa Redmiles: So one of the things that we do here is we often try to look at fields that measure human perception. So take psychology, for example. There's a well-studied task called the familiar facial recognition task that's used to diagnose things like prosopagnosia. So it's been studied quite a lot.

And basically the task is I'm gonna show you an image, I'm gonna ask you, what is the name of the person in the image? And you're gonna reply either with the name or with identifying details, like, "Oh, this is the person who played Wolverine in the most recent movie." And so that's an example of a task that we use to see, like, can someone recognize? And then we control for all sorts of confounds, things like, are you of the same gender and ethnicity as the person who we're asking you about? Are you familiar with them in real life or not? Et cetera, et cetera. So we do that kind of work to capture human perception and compare it with computational metrics.

But that doesn't get at the deeper part of what you're talking about, which is things like, even if this is a cartoon of me, is that something that would cause harm? And how do we think about that from, like, a recognition standpoint? Because certainly I can recognize a cartoon of Albert Einstein even if it's not photorealistic.

And so in those cases, we do one of two things. One is we create surveys that we call vignette surveys. So we basically create a scenario for someone to imagine themselves in. So for example, we would say, "Imagine that a stranger created a video showing you as a cartoon having sex, would you find that acceptable or not acceptable? And then we have, like, a rating scale, right? And survey methodologists like these kinds of scenarios because they kind of approximate people's feelings. They're not precisely real, but they're, they're closer. That's one option.

The other option is we interview those who run, like, support organizations for victim-survivors or work with them to speak with victim-survivors themselves to understand case studies that are a bit more nuanced. Like cases where there's cartoons, cases where things are not quite as clear-cut as maybe we think about in a kind of like photorealistic nudeification case.

Renée DiResta: You've also done a lot of work looking at the motivations of the creators. You had a really fascinating paper where you talked about how how people who create these images think about their role and what they're doing, particularly in cases where they make them for themselves and don't share them. And you have a lot of really interesting nuanced dynamics to your thinking about the creation process. I wonder if you could share a little bit about that with the audience.

Elissa Redmiles: Yeah, absolutely. So we started studying communities of folks who use AI to create sexual content because we were curious the extent to which those communities were creating norms that sort of separated out non-consensual creation. We were also curious how they thought about issues like accidentally creating NCII. For example, if a particular model has been really heavily trained on a particular celebrity individual, it's possible you would write sort of a generic prompt and get back someone who looks a whole lot like a celebrity who you know.

And so we started going into these online communities, some of which have hundreds of thousands of members, to understand, like, how are people getting into this? What are they making? Is it abusive or not? How are they doing governance? What are their norms? And that helps us both see the type of content being created, it helps us understand what are the pipelines for creation.

Because I think we all talk a lot about, like, nudification websites, right? I upload an image, I get back one that's undressed.  But there's a whole wide world out there of folks using what we call open weight models, things that have been uploaded on Hugging Face, elsewhere. And they are either modifying them or using them just as they are to create their own content, either locally or in the cloud.

And I think before we did this project, the sort of prevailing wisdom was like, "Oh, you need some computer science background, maybe even some machine learning training to, like, do this." But what we actually find in these communities is they're actually serving as, like, upskilling centers, so they're kind of training each other. If, like, a model is too big to run on someone's computer, then they're gonna, like, compress it down so that they can run it. And in fact, one of our participants said, "You know, I'm actually not interested in making sexual content. I have this other content I wanna make, but the most active technical support I could find was in these communities for making sexual content."

And of course, not all of that content is abusive but of the 28 people we did interview across a couple of different communities, three of them talked to us about making non-consensual intimate imagery and several other people talked about getting requests for making it that they turned down.

And the moderators who we discussed with, to greater or lesser degrees, were trying to moderate this type of behavior, but they were actually struggling a bit because people come to these communities for kind of an uncensored place for sexual expression, something they didn't feel like was easy to find, places to talk about sexual content creation or preferences. And so when there are kind of rules put in place, sometimes the moderators get pushback that's like, "Excuse me, you're not supposed to kind of yuck my yum. Don't tell me what to do."

And so they actually said that they found two things useful to justify why they had rules. The first was the terms of service for the platform that they were hosting the community on. So the terms of service have some mention of non-consensual intimate imagery. And because people so want to be in the community if they bring up, "Hey, the community could go down if you share this or if you take these requests," that really helped.

And the second was the new Take It Down Act. The idea that there might be legal consequences for sharing really, like, motivated the moderators and was also a good justification to the community.

Renée DiResta: That's interesting that you bring up these notions of, you know, perceptions of censorship versus permissiveness within the community. I know there's also an emerging legal debate over whether generative AI outputs receive First Amendment protection, either through the rights of users to see content or the rights of developers to, you know, build models and the sort of expressiveness of code. This is a long-running debate, but you know, this is obviously unsettled law. We're starting to see it play out in a few different cases. There was the Garcia versus Character Technologies one.

But right now, I think there's an interesting question. I am not a constitutional lawyer. Kate can, like, jump in and then, and pick this up, but one thing that I think about a lot is, does that conversation leave out the person who is being created? When we talk about the prompter's expressive interest, the model company's editorial or design interest, where does that leave the person whose likeness is being rendered without consent?

Elissa Redmiles: Yeah, I guess from my side, this is something I think about a lot because in computer science, the main, one of the main privacy threats that people address with computer science research is something called membership inference, and so the idea of this is, like, there are some images used in the training data. Can I attack the model and recreate the exact image from the training data, or can I extract it?

And this is an important problem for various reasons, but I often talk to colleagues about my frustration that it is not the whole problem because my ability to reconstruct the exact source photo is maybe not my concern when it's a harm around my likeness. The issue is, can you reconstruct an identifiable version of me?

And one of the challenges we've had is with membership inference, you can define this mathematically because, like, that training data object existed at some point, and so we can, like, write some math about what we're trying to do. With the likeness, we need some sort of mathematical definition of, like, what is a likeness, and that's a much more, as we s- talked about, a human perception kind of thing. Like, can we achieve that? I'm not terribly convinced we can, and yet it's an important problem to work on, to evaluate for, et cetera. So that's where I see this coming up

Kate Klonick: Okay, so I love this question. I think that this is where kind of Elissa’s amazing research, and kind of a lot of the work that you do, Renée, and some of the stuff that we've talked about, like, kind of offline and off the podcast about deepfakes and the implications and what the legal answers to gen AI is going to be, I think that this is kind of where, like, this is going to maybe be the fore of where it's decided, because I think that these are such obvious places where vulnerable individuals can really be exploited, but also it is the place where people who have the most economic stakes can be exploited And the combination of those things motivates people.

Renée DiResta: Well, yeah, this was I think Meta pulled their weird tag in the Instagram person you wanna make a picture of thing because one of the creator guilds protested, as I recall, right? Complained about it.

Kate Klonick: Yeah. So I don't wanna kind of go all the way back to kind of, you know, to bore everyone in this pod who did not turn in for Kate Klonick's history of IP law and internet history. But this is essentially what you saw in kind of, in my, in my version of events, th- this is what you saw with the early internet, that we're kind of speed running through as an is- legal issue now for gen AI, which is that you have property-based rights being the most enforceable and the most economically viable rights that you can enforce in courts. And so you see them at the fore, and you see huge interest groups going to bat and taking down the technologies, or trying to take out the technologies that do this.

And so obviously I'm thinking of, because this was, like, happening when I was growing up and a teenager, and so I was, like, on both the side of, like, this this group of lawsuits ending right as I got to law school, and being a person who, like, downloaded illegally tons of stuff from Napster, obviously like the Napster-Kazaa peer-to-peer sharing kind of issues where you had huge interest groups like the RIAA and the Motion Pictures Association of America kind of decide to bring suit because their, their actual content was getting ripped off.

Now, that is a solely copyright kind of IP- like, squarely an IP type of right, right? The, and that was the, the question there for a lot of people was like, do we change how we distribute IP rights and what we think of as IP and how we do this in the age of the internet?

That's a little different when you have generative AI, which is both mathematically from, like, the authenticity point that, like, Alyssa was talking about before, mathematically new material that happens to have a likeness. Like, it is pixel by pixel a completely different image than a photo of Renee. You know, if you had a generated image of Renee that looked exactly like Renee looks right now in her Georgetown office, like, you know, and then you had this image of Renee, like, that we asked the, the machine to produce, whatever model we wanted to use, and if it was, like, almost one-to-one, they would still be different images from a, from a, from a, from a, like, I guess, like a granular- mathematical perspective.

But, and this is kind of what I think is so interesting, I was getting at before, phenomenologically the picture plays the same role, so s- sociologically it plays the same role. It doesn't actually matter. And so the effects of the picture, the harm that it creates, right? Doesn't end up totally mattering, and this is, I think, it, the thrust of essentially what the debate is going to be.

Like, do we end up going down, and I mean, when we, I say, like, the law. Does the law end up going down this really particularized route of, like, mathematical authenticity, or is it going to go to kind of down this humanized route of, like, how it categorizes the harm and, like, whether or not the generation of that harm exists through AI or something else, like we are going to recognize this harm?

And that has different First Amendment balances. So if you tend to recognize generative AI as, like, or you were a proponent of recognizing gen AI as having its own First Amendment rights because it is generating something completely new and technically authentically different, you're just going to be able to block all, a lot of, not all of, but potentially block a lot of the regulation and potentially litigation, regulation and litigation that is going to kind of be filed for, like, against, against some of these models and companies that are running the models for the outputs that they produce, that are exactly the type of outputs that Elissa’s work is kind of talking about.

If you decide to take a, a more kind of harm-based approach, a more dignitary privacy-based approach where it doesn't really matter, like, where the image comes from, just that it exists and it can be traced to the generation from, like, either a model or a human, then you have, like, a whole different kettle of fish and kind of a whole different set of legal solutions that you're going to put in place there.

So I don't wanna kind of take this over with a legal discussion, but I do think that this ends up being kind of the payout of a lot of Elissa’s work if we c- you go to, like, put it to policy. And I think it's super interesting just because it reveals the limits of, like, what authenticity gets us, and it reveals the power of things like likeness, right?

Likeness being something that no one has been able to perfectly quantify or understand in either cognitive neuroscience or IP or whatever for years. I mean, like, literally, you know, how trademark likeness is decided is through survey. It's through the exact type of research and kind of, like, qualitative and quantitative stuff that, like, Elissa does or, like, you know, social scientists do.

And so I don't know. I just kind of, I like bringing this up just because I think it, it gives us kind of also context for, like, some of these ideas are new and new at scale, and some of them are, like, taking a new twist on a really old discussion and a really classic discussion that we're kind of just getting to see a new valence of.

Renée DiResta: I think this question of what should the regulated act be, I'm kind of curious to hear your thoughts on this, Elissa. Is it, is it the creation? Is it the distribution? Is it the implied threat maybe of putting a person into a very particular type of sexual scenario? Is it profiting from the ecosystem? What, where, where do you see the, the lines around policy on this front? Or maybe regulatory law is a better term than policy for this one.

Elissa Redmiles: So in this space, I'm often kind of torn between my background in computer security and focus on kind of anti-censorship, protection of people's communications, and how we can best address harm. So I think certainly regulation that addresses sort of the pipeline that funds this would be most effective if done in a way that is careful about addressing non-consensual intimate imagery specifically, and not sexual content in general.

And I think that's always kind of a tension, at least I've seen from doing work on the security side in this space, is just when companies, say, react to a regulation by going really broad in their prohibitions, sometimes we find that just kind of pushes everything to the corners of the internet where we have the least ability to kind of take control.

So I always think about it in terms of a, a balance there. And because this is dual use technology particularly that used to kind of create things, I think one of the biggest areas for computer scientists to do research is how do I identify when something's sort of purpose-built for creation of non-consensual intimate imagery, whether that's a website with very clear kind of keywords and functionality, or a model that's very clearly kind of trained or fine-tuned on particular individuals. And there are folks doing research in both of those directions. I think addressing those tools and both who owns and profits from them as well as who uses them is important.

The other thing that I think about, maybe more so from corporate responsibility, I think I'll defer to you all on regulation, is kind of we have seen the use of mainstream platforms like Google, Facebook, et cetera, as single sign-on providers. So you know, you click and then you can easily sign into the platform for various notification services. And when that happens, it gives this kind of like acceptability stamp to the platform, even if it's not the intent of the single sign-on providers. Same when we see those platforms like running ads or having applications in their app stores that can do this or are dedicated that way.

And so I think there's a lot of kind of corporate governance points in terms of just like keeping an eye on where your brand is popping up and what you're promoting. And we've seen lip service for doing that. We've also seen some failures after the lip service. And I think a lot of those failures are because this is a dual use kind of thing, you have to really like test the actual application. Like you need a, we usually like generate an image with a face covered. We try to do some double checks that it's not identifiable, and we use that to test what the system can do.

And I think that often that's kind of not what's happening, just like the form is being read by whomever is reading the form or whatever is reading the form, and that's where we've seen a lot of failure. So I think whether that's a regulation thing that gets addressed or just sort of a corporate responsibility thing, that would be important.

Kate Klonick: Yeah, so I wanna actually kind of cut in, 'cause I know Renée might have thoughts on the same question. And she's also done a ton of work. I mean, Renée, you've also done, like, so much work on kind of how all these different systems are running, kind of capturing the idea of how they each respond to different types of prompts, how Grok is responding to prompts, how Meta is responding to these same types of prompts. I'm curious how your research, you would describe it as different or the same as some of the stuff that Elissa’s done, and also if it has pointed you towards a different k- policy kind of suggestions.

Because I do think that Alyssa's point, which is like, it kind of depends on, like, what it is exactly that the model or the actor or the whatever is doing. Like, it is one thing if it is, like, a jerk ex-boyfriend using Grok to, like, to do this thing, and it's another if you're talking about the model itself. And then there's another thing if you're talking about kind of, like, literally some type of, like, six guys in a trench coat in Russia deciding to kind of run a shop that does this for individuals off of Grok, right? Or something, like, to that extent.

And so kind of I'm just, like, wondering, like, both of you have had kind of these qualitative conversations and interviews with people and done a lot of the qualitative experimental work yourself. So anyways, I, I just kind of will start with, like, Renée, I'd love to hear that.

Renée DiResta: Yeah, it's a real big challenge. I used to think that corporate responsibility was, was something where you would just never see this from large foundation models, including Grok, because they would just prevent it. And then we had that debacle with donut glaze, right? And on, on Elon's, on Elon's platform, where for, I think probably most of the country was aware of what happened with that, but just to restate it in case they're not, it was basically that people began to realize that spicy Grok would generate all sorts of images of people in very compromising situations.

And with X it's, it's interesting because you could just ask Grok to generate the image right under the person. So this was an extraordinary vector for harassment, and many of those images got millions and millions of views. Ashley St. Clair, an influencer who also had a child with Elon Musk, is currently suing him. I don't think we have really seen that case move very much. I haven't seen any updates on it in a while, but you know, she sued because of what happened. I think he's fighting about venue at the moment.

But this is where this, you know, beginning to realize that corporate responsibility isn't gonna hold. You know, so you just see these, these rather extraordinary things where, like, morals go out the window and…

Kate Klonick: I, I just wanna say for, for, for listeners that the donut glaze thing is you would essentially ask… It, it was a way of routing around things-

Renée DiResta: Yes.

Kate Klonick: That were. So I just wanna say, like, because I do not think that most people are going to be fa- that are not very online are going to be familiar with this kind of absolutely horrid specifically for the generation of child sexual explicit material-

Renée DiResta: Right

Kate Klonick: That, like, how bad this was. You could, instead of asking for a pornographic, a specific type of, like, thing to be all over somebody, you could ask for donut glaze to be all over someone, which was the approximate look and style of, like, what that thing would be to create kind of a very explicit image that could be taken into various contexts and just, like, you know, having a child covered in this or anything else. So it was very, it was a very kind of huge controversy in the world of people who follow this type of thing.

Unfortunately, I don't think it actually, kind of, made the mainstream news at all. I mean, both because it's incredibly explicit and because, I don't know, people tend to not care about these things at the margins, and, like, Elon Musk has such a high threshold for the thing, like, of all of the absurd things that he does, it's like half of them don't even rank anymore.

Renée DiResta: This one was interesting because it did lead to immediate regulatory action in Europe, right? So you had a lot of immediate calls for you know, various types of the regulators in Europe requested specific data from X. That sort of pushed it into the news a little bit more than I think it otherwise would have been, in part because you began to see coverage of Elon saying, "This is censorship." You know, being asked to provide data on exploitation on his platform was censorship. Requests that, you know, or people pointing out that maybe this is actually a terrible thing you know, were recast as censorship.

I actually wrote an article about this for Lawfare for anybody who wants to track it down. We'll put it in the show notes 'cause then you can kind of see the, the specifics of this particular story. But I think the challenge on the regulatory front is not wanting to cross bounds into free expression, as we've talked about. You know, where, where are those lines?

You know, you mentioned the child exploitation content. That's much more where I have done my work, not in adult NCII so much, as, as where there is no consent. There is no consenting person in that, in that exchange, right? The, the child is either being exploited or an image that is potentially derived from models that have a perception or inadvertently were trained on that stuff, that's where, that's where you start to see that happen.

So there are some very granular laws that come into play when the child content becomes the focus of the thing that are a little bit different than the potential that it might be voluntary content creation and uploading. So I think that's where for me a little bit of my work has been focused much more on the former.

I think I, I wanna talk about the technological barriers to creation, though, because this was the other thing, right? If you have corporate responsibility and the corporate responsibility relates to them putting guardrails on their model there is this dynamic where people will go and do things with open source models. And now with Claude Code it will very easy, you know, it'll just tell you how to set up an open source model on your computer. This used to be something where it was a little bit more gated by competence, and now anybody can do it.

So each theoretical bound related to ethics or user competence, those, those have all been sort of eroded. So I'm curious, Elissa, when we think about this creation stack or this distribution stack or the payment stack for, for sites that are that are making money off of this, how do you think about the different roles that each of them play, and, and what would you like to see as policies for, for that stack?

Elissa Redmiles: So think of this as kind of a landscape, right? And as you said, we have sort of the tip of the iceberg is like the, the nudeification websites and services. I go to a website, I upload an image, I click some buttons, I get something back. I think those are in some ways the easiest to some degree. We take down the website. It's pretty clear what the website is. Now, is it gonna pop back up a million times? Yes.

And so that gets to what underpins that website. And what underpins that website as far as we can tell from the research we've done in the communities that I mentioned, where some people were building these kind of tools, as well as looking at like how-to guides that are hosted across the web, is that they're using, as you said, an open weight model. They have kind of a system prompt, and they've built this such that people can put in whatever they want, but they've kind of specialized it to perform well at, at nudeification.

And so, one of the things we find in the report that I mentioned we'll have coming out is there's no mention of open weight model providers doing any kind of monitoring of downstream use of their services. And in fact, if you look at like Stability AI, which builds Stable Diffusion, which is one of the open weight models that's used very heavily on the website called Civitai that I mentioned that has been used for modifications of people, you'll see that I believe they reported zero reports to NCMEC during a particular one-year period. During the same period OpenAI, for example, had far, far more reports.

Kate Klonick: NCMEC being the center that runs basically the database for PhotoDNA and other things that automatically use hash- hashes to take down child sexual abuse material and things like that.

Elissa Redmiles: Yes. And so, you know, as NCMEC themselves sort of says, like the number of reports maybe tells you more about how well you're monitoring than it does about how much abuse is happening. And when researchers have done research, they've found several different open weight models being used quite prevalently and in ways that are observable because they're kind of publicly on the web.

So that's kind of one piece of the puzzle. The other piece of the puzzle that you're getting at, Renée, is can we do anything in the model itself, like such that the model couldn't be used in this way? And what we find is there's like three different approaches you could take at the model level technically right now. None of them work very well.

So one set of approaches people have proposed, like what if I remove certain kinds of content from the training data? Will that prevent that idea? So let's say I take all the images of children out of the training data. Can the model now not make a child plus sexual content, otherwise known as child sexual abuse material? If you could do this perfectly, then it might help in a closed weight model, so in one w- that you can't modify. But even if you could do it perfectly, it wouldn't help in an open weight because you can just reintroduce the concept, and we can't do it perfectly.

So in our own research, we find it kind of makes it go from three prompts worth of difficult to 12 prompts worth of difficult in order to generate, just 'cause you still have a few examples. There is a bunch of research on ways to kind of make this harder, do anti-tampering, kind of prevent fine-tuning. None of it's ready for deployment. We don't see any platforms using it.

And so most of the protections that we see from foundation model providers are for the products that they build themselves on their foundation models, and they're really those, like, input-output filters, like let's try to check that your prompt is okay, let's try to check that your output is okay. And as we've discussed, those are circumventable. They're also using AI models to do that filtering, and AI models are unreliable. So right now it's a very, kind of, gappy ecosystem.

The last thing I wanted to mention was we've talked about kind of like, the EU response in terms of regulating, and I know there's been this proposed update to the, the AI Risk Act in terms of kind of saying you can't have a model that can do this unless you've deployed certain safeguards.

And I think one of the things we've found in our research is you hear a lot about the safeguards that foundation model providers have implemented, but we find very little transparency information on how well those safeguards are evaluated. And even when we interview foundation model providers, we find a lot of kind of nebulousness about the kind of threat model that they're trying to address. So are they only worried about direct generation of an image, or as you mentioned, are they worried about code being generated that can go and be used to make the image? Are they worried about instructions? Are they worried about API users, et cetera?

All of that's a little fuzzy. The answers to questions we've had about likenesses and sort of what fits the definition of NCII have been fuzzy. So I think there's a lot of attention that would need to be paid to like what qualifies something as a sufficient safeguard, and how do we set people's expectations appropriately for how much that's gonna protect them individually versus, kind of, mean something about mitigation in a specific product

Kate Klonick: Yeah, so this is actually kind of getting at something that I want to ask both of you, which is that, like, to kind of bucket a little bit the different ways that we've talked about this, I, tell me if, for both of you, if this seems right or this seems wrong, like, I feel like we're kind of talking around this idea that there's different ways of attacking this problem.

There's kind of like the post-hoc kind of legal reaction, litigation type of way to kind of create a post-hoc reaction and kind of a tort-based like kind of solution set for an individual that is harmed by this type of stuff. And, and ideally kind of like an disincentive to do it in the future for other people, for tortfeasors who might do this type of thing again to other people. There's obviously, we didn't talk about this, but criminal liability you could create around that to a certain extent, like the NCII in state kind of laws that have been created over the last few years.

And then there's, I think you just said it really well, Elissa, there is like the, and, and Renée kind of gestured at it when she was talking about like the model-based versus the open source versus the guardrails versus the regulatory sc- like world of things that we could open up. There's like how do you get at this problem when, one, information is free? You can kind of pass all of this along. These models are free. Like is there, it seems like Whac-A-Mole on steroids. Like you just can't possibly kind of do this.

And then there has been, you know, for better or worse, all of this call for transparency around these things, which lets out from a security standpoint a lot of the secrets that we're like learning and refining to like make these products and models safer.

So one of those, speaking of the donut glaze kind of thing to go back to that, was something that I had heard about a year and a half ago when I was interviewing people at OpenAI about like their, how their, what their red teaming process looked like. And that was not popular knowledge at the time, and I remember being like, "Wow, that's horrifying," and also like, "Oh, that's not gonna stay a secret just at OpenAI's red team for a very long time," and of course it didn't.

And so I'm kind of just like very interested from both of you going forward. You have things like red teams, which for those unfamiliar, our red teams are essentially like specialty people who are either in specialty areas, like you're either in nuclear development or poison control, or you're in, like, you know, some version of, you know, I, I don't know, some version of harm prevention, and you get asked by the models to essentially run a bunch of bad actor questions, the red hat, like, kind of idea. Like, you put on a bad actor hat, and you test the machine with as much stuff as you can throw on it. And that is kind of then guardrails and safety mechanisms are, in theory, brought in by the company. Like, that is a corporate responsibility idea that we could put into, into regulation and force them to have that and be responsive to it.

But that doesn't get rid of the Whac-A-Mole problem, and maybe, you know, keeps going. So like, Renee, you asked this question of Alyssa, but I wanna ask it of you, and then, like, Alyssa, I kind of wanna see if you just agree with that framework that I presented from your research. Renee, do you think that the-- where do you think the, like, the choke point of this is? Is it all of the above?

Renée DiResta: I think there's different voluntary things that the different actors can do immediately. Red teaming, by the way, it's not clear from, for outsiders what the legal frameworks are, right? There are certain things that, you know, you, you can't do, and reasonably you can't do. You know, when Meta makes an announcement about its thing and you wanna know if you can, you know, if, if, if it will generate for under 18 accounts, that's not something that you can really do ethically as an outside tester, right?

So these questions around what is the formal Legal framework for authorized AI safety research. What should it be for the people within the companies? What should it be outside of the, outside of the companies for maybe specifically designated individuals? I think it's very hard to know how to, how to do that, 'cause just saying, "I was red teaming," is not a legal defense.

As far as the, you know, responsibilities in the chain, you know, so first of all, I don't have very much faith that open source communities are not, you know, there, there are these subsets that are actively there for just this sort of content, right? That is the, that is the motivation.

So with that, then, I think you look at questions around distribution. I've seen, you know, law enforcement go after certain types of monetization sites. Sometimes that's hard because they live outside of the country. Bellingcat did a really amazing investigation into one of the sites and the name is escaping me right now. Maybe Elissa remembers it. But like sort of trying to track down where this operator was and then trying to see what sort of legal rules might apply to that person where they are and how, how to think about that issue.

There's the hosting and cloud services have something of a role to play here, but then people get very uncomfortable about the idea that they're scanning. And again, this doesn't, this doesn't adequately address issues of consent and whether a person has that in, you know, those images up in their drive because it is theirs, they made them, they're an adult, and this is fine.

So you do have a whole lot of different component parts here, and a lot of it is going to be this patchwork of there will be, I think, the mainstream companies that will adhere to laws and, and be good corporate citizens. And then you'll have these sort of niche outsiders that are going to continue to, to keep that, that market active.

Elissa Redmiles: Yeah, no, I do agree, Renée. And one thing that I might add, and I think this relates to the kind of, you know, should we require red teaming, as well as to the, the continued proliferation, is red teaming can vary in quality and it could vary in, like, the extent to which it covers the threat models, right?

Kate Klonick: Right.

Elissa Redmiles: So like we, part of why we do the research with perpetrators is so that we can make sure when we're doing tests that we're representing how those people are actually doing things. Sometimes that's not happening.

And so, I think the question kind of becomes if we find that we can't currently safeguard a particular system, or, you know, the odds of harm are, you know, X out of Y, how do we sort of trade off what we wanna release or allow kind of under like an EU AI Risk Act type concept versus what we're able to protect. Like, if we're only able to protect non-person generations of video, how does that kind of change what we wanna release?

And I don't think I alone have the answer, but it becomes kind of a societal question. And so I think the thing I worry about is we're often looking for like a technical Band-Aid where there just may not be one or there isn't one yet, and that becomes kind of a question of like, do we try to not let new stuff out? Is that feasible? Do we try to put more resources into deterrence messaging, primary prevention, doing like social norms shaping? Those are things that I worry about.

Renée DiResta: I'd love to maybe close with something like, do you think there is one intervention that is technically possible now but largely neglected by companies that you would like to see implemented?

Elissa Redmiles: Yes.

Renée DiResta: Tell us about it.

Elissa Redmiles: I would love to see implemented more widely is deterrence messaging. So this is something there's like three studies out there looking at how deterrence messaging can, which are basically messages that try to say, "Hey, the request you have appears to be for content that might be illegal to share," let's say, “under the Take It Down Act.” These have been used for child sexual abuse material before. There's been research on using them for non-consensual intimate imagery. The research shows a lot of promise when they're deployed.

However, none of the seven foundation model providers we looked at, and none of the ones we inter- interviewed were doing deterrence messaging for NCII. Two had it on their roadmap, but nobody else did. And I think there's a lot of research to be done in terms of not just kind of deterrence messages for generation, but also like for bystanders. How do we think about that within these communities where people are making sexual content that's not abusive? How do we think about it?

But those kind of social norm shaping efforts can be very promising. Yeah, and I should maybe be clear when I say deterrence messaging, I mean when you attempt to generate content on a particular, say, foundation models website or what have you, not necessarily when you're sharing content in encrypted messaging. I would not prefer that, but yes.

Renée DiResta: So I think what this conversation has really made clear is that AI-generated sexual abuse is not just a problem of bad content slipping through imperfect filters. It's really a problem of systems that are built without consent at the center trying to tack it on now, legal rules that often intervene after the harm is done, and then safety regimes that put an awful lot of burden on the target as opposed to the, the creator sometimes.

So I think the, you've really brought out a lot of the very nuanced, challenging issues here. I think maybe some people who are new to the conversation have always thought that, "Oh, well, they'll just ban it, and then it'll stop." So I appreciate you walking through the nuance with us. I think the questions are not just whether a model can generate this material but who is responsible for thinking about potential guardrails, what kind of research is necessary to test the risk, and then how we protect expression without turning content creation into a shield for, for exploitation.

So I wanna thank you for joining us, and you have a CDT report that you mentioned. We will put that in the show notes, and yeah, really appreciate you being here.

Elissa Redmiles: Thank you to you both for a great conversation.

[Outro]

Renée DiResta: The Lawfare Podcast is produced by the Lawfare Institute. If you wanna support the show and listen ad-free, you can become a Lawfare material supporter at lawfaremedia.org/support. Supporters also get access to special events and other bonus content we don't share anywhere else. If you enjoy the podcast, please rate and review us wherever you listen. It really does help.

And be sure to check out our other shows, including Rational Security, Allies, The Aftermath, and Escalation, our latest Lawfare Presents podcast series about the war in Ukraine. You can also find all of our written work at lawfaremedia.org. The podcast is edited by Jen Patja. Our theme song is from Alibi Music. And as always, thanks for listening.


Renée DiResta is an Associate Research Professor at the McCourt School of Public Policy at Georgetown. She is a contributing editor at Lawfare.
Kate Klonick is an Associate Professor at St. John’s University Law School, a fellow at the Brookings Institution, Yale Law School’s Information Society Project, Harvard Berkman Klein Center and a Distinguished Scholar at the Institute for Humane Studies. Her writing on online speech, freedom of expression, and private internet platform governance has appeared in the Harvard Law Review, Yale Law Journal, The New Yorker, the New York Times, The Atlantic, the Washington Post and numerous other publications. For the 2023-2024 academic year, she was a Fulbright Schuman Innovation Scholar in the European Union where she was a Visiting Professor at SciencesPo and University of Amsterdam researching and writing about the Digital Services Act and Digital Markets Act.
Elissa Redmiles is an assistant professor of computer science at Georgetown University.
Jen Patja is the editor of the Lawfare Podcast and Rational Security, and serves as Lawfare’s Director of Audience Engagement. Previously, she was Co-Executive Director of Virginia Civics and Deputy Director of the Center for the Constitution at James Madison's Montpelier, where she worked to deepen public understanding of constitutional democracy and inspire meaningful civic participation.
}

Subscribe to Lawfare