Peter Bauman: Hello everyone, and welcome to this Le Random podcast. I'm your host, Peter Bauman, the Editor-in-Chief at Le Random. And today we are continuing our Friday podcast episodes as companions to our Monday editorials. And this Monday we released a very special piece with Ian Goodfellow on the birth of GANs. And today we are releasing the full conversation. So Monday was basically part one. This is now the full conversation with the legendary machine learning and deep learning engineer Ian Goodfellow, who was also the inventor of GANs.
So besides inventing GANs, Ian is just a legend. He has nearly 400,000 citations on Google Scholar. He literally wrote the book called Deep Learning with Yoshua Bengio and Aaron Courville. This is honestly one of the interviews I've been looking forward to the most. Ian has been someone that I wanted to talk to for years. He's had this singularly outsized impact on image making in the 21st century, and it really hasn't been explored. He doesn't give many interviews. He hasn't spoken extensively in public in years, and he's never given an interview in this much depth about culture, and specifically his impact and GANs' impact on culture. So this is an incredibly important story, how GANs were born and their impact, and I can't wait to help tell it. And thank you so much for listening. Let's begin.
Yeah, maybe just to start. So it's been about 11 years since the paper. It was the summer of 2014. And there was even this retrospective last year at NeurIPS 2024, where they looked back at the legacy of GANs after 10 years. It was interesting because, of course, it focused on the extensive technical legacy and the development of deep learning over the last 11 years, but not that much was made about the cultural impact.
Ian Goodfellow: Yeah, the cultural impact has been quite large, and I think not quite as explored as maybe the technical side. So yeah, I'm excited to get into some of that today.
Peter Bauman: And one other area that we're beginning from is this idea that multimodal media generation is important, even historically and art historically, and in the relationship with art and technology, even as important as the advent of the digital or the internet. It has been compared to this new paradigm for media, and even moving forward, basically the next century. GANs played a really significant role in that initiation, or in the development of deep generative models, and especially turning them from something more theoretical to tasks that are actually based on synthesis and generation. So maybe we can start with just, can you explain briefly what GANs are and what they do, in a very simple, maybe broad way? And then we can talk about the story of how they were developed.
Ian Goodfellow: Yeah. And I think that turn to focusing on synthesis and generation of media is also something really crucial to touch on. And I think some of your questions get at it, and the story there might be more interesting than some of the questions land on. But yeah, so to start with explaining what GANs are. GANs are a solution to the generative modeling problem. The generative modeling problem is, broadly speaking, you have a data set containing a lot of examples of something. For GANs, we usually like to use photographs as the example of the something. So maybe you have a lot of photos of faces. Those faces all come from some probability distribution, and we want a learning algorithm to learn the probability distribution of realistic faces and then generate more faces like them.
GANs do that by having two different neural networks compete in a game, in the formal game theoretic sense. One of those neural networks is a generator network, and you can think of that one as like an artist. The other network is the discriminator network that you can think of as like an art critic, except it's not really a stylistic art critic. It just looks purely at realism.
In the game, the generator produces a photo of a face, and then the art critic looks at either real photos or fake photos coming from the generator. And for each individual photo that the art critic sees, they have to estimate the probability of whether the photo is real or fake. And the generator gets a learning signal based on the probability coming from that art critic. So the generator tries to drive the art critic's probability up toward one. The art critic tries to drive their probability on fake samples down to zero and their probability on real samples up toward one.
We can actually analyze this with game theory and show that the Nash equilibrium, so the point at which neither player can improve their strategy, is for the artist to learn to produce perfectly realistic photos of faces and for the art critic to have to essentially do a coin toss on every photo, because all the photos will be perfectly realistic. And back in 2014 that sure seemed theoretical to most people. I think in 2025 we're now seeing that AI fakes can be good enough that it's really hard to tell.
Peter Bauman: Yeah. What's so interesting, too, is that GANs had this rich technical and also cultural impact. But the story of them is also really interesting and also, I think, unique in that it happened all quite quickly. The idea happened quickly, the paper happened quickly. I wonder if we can talk about the story of how GANs came to be 11 years ago. There's a now famous story out there about a night. It was in Montreal. Is that right? I believe even the date, I got the date, was May 26th, 2014. So again, not too far from when the paper came out. But yeah, so just beginning with that origin story, it was this grad student send-off in Montreal at a bar. What do you remember from that night still?
Ian Goodfellow: Yeah, that date sounds correct to me. It was about two weeks before the deadline for NIPS, the conference that's called NeurIPS now. I had not been planning to write a paper for NIPS that year. I was working on writing the deep learning textbook instead.
I think one thing that's important to understand is, in terms of the speed that everything happens, everyone involved in this story had been working on neural nets and ideas related to generative modeling for years beforehand. We weren't trying to solve the generative modeling problem per se, but a lot of us had thought that using generative modeling as a practice exercise might help neural nets to solve supervised learning, where you show the neural net a face and you ask it, for example, is this person smiling, frowning and so on? And by 2014 it had turned out that you could solve the supervised learning problem without having to do a warm-up using generative modeling. But all of us had had a lot of practice doing generative modeling for several years beforehand. So we all had ideas floating around in our head that clicked into place on this night. It's not that the ideas came out of nowhere on this particular night.
I had just not been planning to write a paper. I'd actually been working from my girlfriend's house, writing the textbook. I had handed in my thesis. I was basically done with grad school, and so my lab mates hadn't even really seen me very much for a little while. This party was a going away party, both for me and for my colleague Razvan Pascanu, who was going off to DeepMind. Since it was their first time seeing me in a while, some of my fellow students were asking me for programming advice. These nights at the bar often ended up with friendly academic debate and sometimes friendly software engineering strategy debate. And in this case, we ended up in a pretty big disagreement about how to solve the problem that they were asking me for help with.
So they were saying that they wanted to make a generative model that had essentially a generator that would spit out photos. And then they wanted to check if the average brightness of all the pixels in a batch was correct. And then they wanted to check if the correlation between all the pairs of pixels in the batch was correct. And then they wanted to check if the product of all three pixels, for all triplets of pixels in the batch, was correct. Just checking pair-wise correlation won't get you very far in terms of getting a realistic photo. But they thought maybe if we bring in triplets of pixels, we can start to get interesting structure in the image. The problem is there's a whole lot of triplets, right? If you have a thousand pixels, which isn't even very many pixels in an image, you're already talking like a million triplets. They were asking for advice for how can we try to write a program that is able to process so many triplets.
I was walking them through the math of, well, here's how many pixels are in your image. You're going to process not just one image, but a batch. Here's how much memory is on the GPU. You're not going to be able to program your way out of this. There's just too many triplets. And so what I told them instead was, how about instead of trying to make a big long list of all the triplet features, what if you have a neural net learn the features that you should look at and have a list of like a thousand features that you think describe what you're looking for in the photo. And that was essentially the embedding space of the discriminator network is what I was describing. So have a discriminator network, keep looking at the photos, and the last hidden layer of the discriminator network gives you a set of a thousand features that you can track, instead of having, for a good sized image, billions or trillions of triplet features.
And they were essentially saying, it's hard enough to train one neural network, you can't train a second neural network in the inner loop of training the outer neural net, which is a fair objection. That's a lot of why GANs are not the state of the art in generative modeling today. This idea worked well enough that it got the ball rolling on image generation, but it did eventually get superseded by ideas that didn't involve this layer of complexity. But at the time we didn't know of a great alternative yet.
I think the context of being at the going away party and having just a little bit to drink. I think a lot of the time having your inhibitions lower just a little bit is useful for trying out an idea that you wouldn't try if you were in the lab analyzing things really coldly. There's a lot of ideas that you might shoot down just a little bit too early. My fellow students really didn't want to try it out. I did want to try it out, and I actually had quite a lot of code lying around from a previous paper that made it pretty easy to glue together. We had a previous paper on just the supervised part, where I could take the supervised model and I use that as the art critic part, the discriminator. And then I could essentially copy paste that and turn it upside down and make that the generator. And then the only new piece of code I really had to write was a little bit to say, flip the learning all around and follow it backwards, and that becomes a learning signal for the generator.
So when people hear that I wrote this really, really fast, it's not that I wrote it from scratch. It's that I had a really good code base already set up with all of the instrumentation and tooling to make it easy to make changes really fast.
The other thing that was important was I got incredibly lucky that the first time I ran it, it worked. A lot of the time you code up a new idea and there's lots of different settings you have to set on a neural net. When you update the weights, how big should the update be? How big should each layer of the neural net be? How many layers should there be? Just dozens of things like this. And a lot of the time you have to launch a computer cluster and search through dozens of different combinations of settings before you get anything reasonable. I was really lucky that just taking the best settings from supervised learning from an earlier paper and just copy pasting everything, even though I was taking it and flipping it upside down for the generator, just worked on the first try.
The first thing that I trained was taking little images of handwritten digits from a data set called MNIST that's really overused in machine learning research, but it's also very fast to train on. And the very first time that I trained on it, it worked. And for anybody that was used to the previous generation of deep learning generative models, Boltzmann machines, just the ease with which that happened made it really obvious that this was something different. With Boltzmann machines, you spend quite a lot of time struggling to get them to produce recognizable MNIST digits, and this just went ahead and spat out MNIST digits in like two minutes. From that, I was excited and went ahead and emailed the lab more or less right away.
Peter Bauman: Were those initial MNIST digits, which were, I guess, the very first GAN images, were they the ones that you used in the paper, or are they lying around somewhere?
Ian Goodfellow: I don't know whether they're the ones I used in the paper. I could try to figure that out by diffing them. I do have them. They're just in my Gmail outbox. I could forward them to you.
“With Boltzmann machines, you spend quite a lot of time struggling to get them to produce recognizable MNIST digits, and this just went ahead and spat out MNIST digits in like two minutes.” — Ian Goodfellow 13:50
Peter Bauman: Yeah, sure. That would be really interesting to see. Yeah, because I guess if those were the GANs you created the very first night, obviously those would be the first ones. Even though, of course, they would look very similar to the ones on the paper.
Ian Goodfellow: Well, it's a grid of random digits. So if it's a different set of digits, it probably... One of them would start with 7, 5, 9, and another one would start with 3, 0, 4. So it should be easy to tell if it's a different grid. I can check, and I can send you the original ones.
Peter Bauman: Yeah, sure. That would be amazing. Part of the lore of that story in that night is also that I think you stayed up all night to do that. You got home from the bar, and then... What time would that have been?
Ian Goodfellow: No, not all night. My girlfriend hadn't come to the going away party. I think she had work the next day, and I was only writing the deep learning book, so setting my own hours. I was awake after my girlfriend had gone to sleep, but I only stayed up an hour after coming home because it was very fast to get it working. Later, I did end up pulling an all-nighter while writing the NIPS paper. And there the funny story is my girlfriend was talking about, are you working too hard? Is this good work-life balance? Not saying I shouldn't, just checking in and evaluating, are you making good choices here? And I was like, yeah, it's definitely worth it to pull an all-nighter for this. This is going to be bigger than Maxout. And now everybody that hears that story is like, what's Maxout?
Peter Bauman: Yeah.
Ian Goodfellow: And so Maxout was a paper I wrote that was the most cited paper of ICML 2013, which we had thought was a huge deal up until GANs. And then clearly GANs were the much bigger deal. But the all-nighter was a week and a half later, closer to the deadline to submit to NIPS.
Peter Bauman: Right. Because I think, like you said, it was about two weeks or 12 days from the NIPS deadline. And so is...
Ian Goodfellow: The very first night I got the MNIST samples. I think the next day I sent them to the lab. I hadn't been going to the lab in person. I went into the lab and said, I had been telling everybody I've handed in my thesis, I'm not doing research anymore, I'm writing a deep learning book. But I've got this idea, who wants to write a NIPS paper with me? And a bunch of people in the lab were like, you have my sword, you have my bow. And then we all worked really hard for two weeks and shipped the paper.
Peter Bauman: And that two weeks that you worked together, were you going in in person and meeting every day? Or was it mostly a collaborative effort?
Ian Goodfellow: Yeah, once I realized I was writing a paper, I was back to going into the lab. I was back in grad student mode. And yeah, the paper was really very, very collaborative. I guess during the Test of Time Award talk in 2024 last year, I was too sick to go and do the talk, but my co-author David Warde-Farley did the talk. We really emphasized that it takes a village to do this thing. And so first, all of the co-authors each made important contributions. But beyond that, the lab itself was the environment that it took to be able to do this. People like Frédéric Bastien, who was a lab employee, developed and maintained tools like Theano, the software library that we used for all of our research. He actually had been working on a feature for Theano that we realized we needed to be able to finish the paper, and he rushed and finished it in time for us to be able to write the paper. This was a massive collaborative effort involving people on the paper and people not on the paper who are credited on the Theano publications rather than on the GAN publication.
Peter Bauman: From what I understand is that basically after that paper came out, this kicked off an intense competition to get this going at higher scales and maybe different and higher resolutions. So how quickly after you presented it did you get the sense that this would be something that the community was interested in? I guess you knew it before you even presented it, but was it right away that you presented it? Was there a noticeable interest?
Ian Goodfellow: There was noticeable interest, but it wasn't the intensity that came later. I guess a lot of things in machine learning get out of date really fast. A lot of the time, things are more like the example of Maxout that I gave, where it's the most cited paper of ICML 2013. You see a lot of people use it in their papers the next year, and then people move on to something like ResNets pretty soon. So there's a pretty short lifespan of something being state of the art.
What was a bit more surprising about GANs is that they stayed really popular for at least five years. Even now, they haven't totally died. There's still adversarial training as a component of a lot of image generation systems. The longevity was more surprising than the popularity, if that makes sense. A lot of things become popular briefly and then get replaced. So pretty early on I could tell, yeah, this will definitely be the flavor of the year. I didn't realize that it would last longer than that.
At a certain point where I realized that it had pretty big traction was LAPGAN, which actually was a bit slow. That was when there was really high-quality, high-investment work following up on it. So LAPGAN stands for Laplacian Pyramid, where they scale up the image to several different resolutions. That actually took about a year after the first GAN paper.
Peter Bauman: That was Soumith, right? And Facebook?
Ian Goodfellow: Yeah, he's the second author. The first author is Emily Denton.
Peter Bauman: Yeah. It's interesting that it maybe took a few months, but I know by that winter of 2015 Dr. Fei-Fei Li had a course at Stanford where GANs were at least mentioned. And I think that helped get some attention to them. And then, yeah, Soumith and Facebook and Emily Denton worked on them. And then by early 2015. And then, I guess unbeknownst to people at the time, but Alec Radford was also working at them at Indico in early 2015.
Ian Goodfellow: Yeah, it definitely kicked off a lot of different channels of interest right away. Yeah, DCGAN, that became public November 2015. And that was over a year later. That was the corner of the hockey stick, where that was when it did really take off. I think LAPGAN was when I was like, yeah, this is definitely going to go somewhere. And then DCGAN was like, okay, this is where it actually is going somewhere.
Peter Bauman: Yeah, it was remarkable that right after Alec posted the DCGAN paper, how much different artists started picking up on it and how much attention it got, and that you really could see even from looking back that it sparked a lot at the time on social media as those things were coming.
So you mentioned you're working on your book a few times. And I think it's an interesting way to frame deep generative models in general, because there's a diagram in your book that basically... The way that the diagram presents deep generative modeling is as this culmination of deep learning. And it seems to be presented that way in the book. And just to frame the impact of GANs and deep generative modeling, do you see deep generative modeling as this culmination of deep learning?
Ian Goodfellow: That's the way the industry has gone, but that's not what we intended with the book. If you read the intro to the book, we actually say there are important topics that we aren't covering, like reinforcement learning. I think at the time that the book came out, like 2016, a lot of people viewed deep reinforcement learning as the culmination, and the fact that we weren't covering how AlphaGo worked in detail. I think a lot of people viewed AlphaGo as the culmination of AI at the time. None of the three authors of the book had enough RL expertise to write a good chapter on RL. And so that was actually viewed as a defect of the book.
The order of the book is more based on two principles. One is if people start reading front to back, we want them to go in the order of prereqs. So we start with all the math fundamentals that they need to be able to understand the math later on. And so partly generative models are at the end because they have the most prereqs. And then the other thing is we try to order it in terms of least likely to get outdated to most likely to get outdated. So at the start of the book, there's linear algebra, which is not going anywhere. And then at the end of the book we're like, here's PixelCNN, we're probably going to be out of date by the time we print it. And so we're moving from general to specific. And by the time you're talking about deep generative models, you have to be pretty specific.
No, I don't think generative models are necessarily the culmination of deep learning. There's a lot of other interesting things you can build on top of deep learning. The way the industry has gone, it's been a little bit surprising that almost everything powerful we have built has had a generative model in it for several years. But when we wrote the book, it felt more like we're all generative model experts, and we're biasing it toward what we know. Turned out we knew something pretty powerful. But we actually felt like, oh, if we knew RL, we'd put an RL chapter in there too. It would have been arbitrary which one came first and which one came second.
Peter Bauman: It's also interesting at that time. There were these two early deep generative models. There was GANs and VAEs. I guess that could also be considered a deep generative model. But before that, for years and decades, they had been theoretical, and they couldn't generate anything. We said at the beginning, they were used for things like feature extraction and feature learning. Yeah, it was really VAEs and GANs that started changing that.
I'm interested also in just what you think GANs can and can't do on a more theoretical level. So one way that I know from the art world that GANs have been criticized. I know back in 2020 Mario Carpo wrote something for Artforum, but this criticism has been around. Many people have said it. I think maybe his is just one of the more well-known ones. But they're usually reduced or flattened to something just like mimic machines or style transfer, and usually set in a derisive way, especially when talking about art, where those are pre-modern notions. And I just wonder, do you think that that's an accurate representation and understanding of what the technologies can do?
Ian Goodfellow: Yeah. And I did a talk at a machine learning and art workshop where that was essentially my thesis. I try to be a little bit humble and stay in my lane. I don't have any art credentials or any philosophy credentials. But if you look at the dictionary definitions of words like creativity and imagination, those words are subtly different, where I would say imagination just means being able to conjure up a sensory experience in your mind that isn't driven by your senses. Can GANs create an image that isn't actually there being driven by camera sensors? Yeah, certainly. That's what they're designed to do.
Creativity seems to require more novelty than what is there in GANs. GANs are, by design, meant to go for a very specific novelty where they're told, we want you to draw a new sample from the same distribution, but we very intentionally don't want you to try to invent things that wouldn't be in that distribution. So you explicitly don't want it to be Picasso. Picasso is out of distribution. That's more creative than the specification for GANs. So I'd say that they have imagination, but not creativity.
Then there's a lot of interesting stuff about, do you have to use GANs in the way that they're described in the training process, where all you do is you train them and then you draw samples from the distribution from them? We've seen people take their decoder and use them to make paint brushes and things like that. Then they can become really interesting tools for creativity, where the human is the creative one and the GAN is providing a function that's really different than just drawing samples from the distribution. But it's still not the GAN that's really being creative.
The part that's philosophically different is when you mess around with materials and the materials give you an interesting effect, are you calling the materials creative? I believe the process has given you an interesting effect there. I'd say the same thing is happening with GANs. When you get an interesting paint brush out of a GAN, I don't feel like the algorithm had intentionality there. It's more like you train this thing and it turns out that if you crack its brain open, you found an interesting paint brush in there. But the learning algorithm wasn't trying to make an interesting paint brush.
“So you explicitly don't want it to be Picasso. Picasso is out of distribution. That's more creative than the specification for GANs. So I'd say that they have imagination, but not creativity.” — Ian Goodfellow 28:29
Peter Bauman: When you've said that they can do more than mimic, what you mean is when humans intervene with them, humans can mold them to do something more than mimicry. But on their own, what they fundamentally do is mimicry. Is that what you're saying?
Ian Goodfellow: Yeah. When you start poking around inside and say, I found this really cool brush, that's the human being creative.
Peter Bauman: Yeah. That's also really interesting. That's also something that Mario Carpo, a couple of years later, returned to when he returned to a criticism of GANs. And I don't mean criticism, really, in a negative way, or just of AI art, which is almost nearly synonymous with GANs for quite a few many years. What he then was talking about is that the machines can do interesting things, or the technologies can do interesting things, but they need that human creativity. You would agree that they can be separated?
Ian Goodfellow: I think what gets more interesting is when you get to the era of text prompt-driven conditional models, which GANs didn't really participate in that much. There is, for example, GigaGAN, which is large GANs that have been trained to go from text to images. In the generative model framing, those are still basically just output an image from the probability distribution of images conditioned on that text.
But we're also seeing that those models do get trained with reinforcement learning and human feedback. So companies like Midjourney ask users to rate images. You can actually earn more time on the server by rating images. And so then at that point, what I'm saying about this is all purely probability is no longer true. It's also probability and a reward estimate. And then the question of how much creativity is involved does start to get a little bit harder, because then at that point the model is placing its bets about what will a human like. So yeah, I'd say once the conditional text generator is trained with RL, you do get subjective enough that there's an interesting philosophical debate to be had there. That's not what we saw GANs doing a lot of the time. I don't feel like I, personally, am the right person to have a strong opinion on it.
Peter Bauman: Well, no. I think what's interesting is that often people who do make these claims about the technologies, they don't understand fundamentally how the technology works. They're not always fair or accurate in the way that they speak them or criticize them. So I think, yeah, in this way, I think you're exactly the right person who knows.
Ian Goodfellow: Yeah. I guess what I'm telling you is I understand the tech. There's a lot of stuff that's more about philosophy and art. And what I can tell you is, here's two pieces of tech. One, I can say, is definitely not creative. One where I'm like, oh, now I feel like, let's call a philosopher.
So the one that's definitely not creative is just, it doesn't have to be, again, any generative model. You just say, here's a bunch of faces, train it with a purely probabilistic criterion to make more faces from the same distribution. I'd say that's imagination, but not creativity.
Now, the next piece of tech is we say, let's have a person describe what face they wanted to make, and then have a person rate how artistically appealing they found the face that it made, and train it with reinforcement learning as well as the generative criterion. At that point, I'm now like, this is pretty similar to how a human artist learns from an art teacher. It's actually hard for me to argue, just on the basis of what the neural net is doing, that it's not creative.
I don't tend to anthropomorphize neural nets a whole lot, but I do like to try to say, what are the actual processes that we do, and do machines do the same processes or not? If the neural net is able to generate something that is not the same as what it was in its training distribution, and there's something that people appreciate about it, that is starting to look to me a lot like it fits my conception of human creativity. At that point, I don't necessarily know that I'm like, okay, pin the blue-red boot on it, but I'm definitely at the point where I'm like, I want to at least talk to an artist about this before I conclude yes or no. If that tells you, here's a technical person, we just found the boundary between where they're like, oh, definitely not creative, and like, oh, now it's hard to say. That's where the wall is for me.
Peter Bauman: Yeah, that's really interesting. So a couple of artists' takes that are pertinent here are Lawrence Lek and Mario Klingemann. I don't know if you're familiar with either of them.
Ian Goodfellow: I've heard of Mario Klingemann. It's not like I know him or anything.
Peter Bauman: Well, he created an autonomous or semi-autonomous artist called Botto. Do you know about Botto? It's similar to what I think you were describing, where the machine creates images and then, based on human reinforcement, can change them. Yeah, you could argue it does have some creativity. But I think where they get into some tricky territory is they think of it as being fully autonomous, whereas I would probably think of that as being more partially autonomous. And then Lawrence Lek talks about fully autonomous, maybe AI artists, as something that would have the agency to choose to be an artist. So maybe if it was a smart satellite that decided to stop surveilling on Earth and just decided to make art from its images or something like that. So there would need to be some agentic choice from the machine. It couldn't be maybe programmed to do that like Botto is. But yeah, I see fully autonomous and partially autonomous artists or things like that.
Ian Goodfellow: Yeah. I find having some creativity in what you do as quite a lower bar than being a fully autonomous artist. Yeah, I see a pretty big gap between those two things.
Peter Bauman: Yeah, fair. So you think what you described earlier was maybe the very beginnings of some partial autonomy or creativity in an artist. So you would think that we might... That even is somewhat theoretical. So we're probably quite far away from full autonomy in an artist. Is that what you think?
Ian Goodfellow: I think pretty far from full autonomy as an artist. To me, that would imply an artist is expressing their own feelings, they have some goals about what they want to communicate to other people about their own life. AI agents just don't have that independent existence. I just don't feel like there's enough personhood for AI for them to even have the possibility of being an autonomous artist right now. Creativity, I feel like there's some sliding scale. If you're purely generating more samples from the same distribution, I feel like that's zero. I feel like you can safely move beyond zero with some of the image generation stuff that we're doing already.
Peter Bauman: I also want to maybe get your opinion on how impactful you think that these technologies will continue to be moving forward, because there is a debate that, is this all hype in a few years? Are we still going to be talking about these things? Some people, I guess, do think that AI is still hype or that deep learning is hype or that there needs to be some transition to neurosymbolic AI or something like that. So there's that anti-hype. And then, of course, on the other hand, there's the hype that generative AI is this new internet, or on par with the new internet or a new digital. Where do you stand maybe in between those poles in terms of hype, and how relevant you think this technology will be in something like 50 years, on your time frame?
Ian Goodfellow: Well, no, I think it will keep going. There have been waves of more and more people paying attention to neural nets. Neural nets were essentially dead for most of the '90s. Andrew Ng is known as a big deep learning guy now. So when I first took his class as an undergrad, the intro to AI class, I went up to him at the end of it and I said, well, wait a minute, why were there no neural nets in the class? And he said, oh, they don't work. Don't bother learning about them. It was actually pretty much exactly that month that Geoff Hinton went and presented deep belief networks at NIPS 2006. They got a new state of the art on MNIST, beating support vector machines, which were the big academic thing at the time.
I think everybody has their moment where they feel like, oh, this is what put deep learning on the map. And for most of the world today, I think that's ChatGPT. For me, it's deep belief networks beating support vector machines in 2006. So for everybody, it's like, oh, things are slowing down because ChatGPT-5 was a disappointment. That's such a small speed bump compared to where things have been going. Look at the time scale from the '90s. If you zoom out that far, if you look at the stock market zoomed out since the '90s, then you freak out about a little oscillation from day to day. I don't really see this as any significant slowdown right now. I think that these technologies are going to continue growing in capability and importance.
Peter Bauman: Yeah. And what's also so interesting about GANs is that they have this direct part in the story of what a lot of people think put deep generative models or generative AI on the map, which is text to image, because GANs were really a big part of the development of text to image. And when artists started combining GANs with other technologies like CLIP, those are some of the very first widely available text to image models, and they were very artist-driven.
Ian Goodfellow: The generative model story and how it turned into media, I think that's a pretty useful one to tell. And I think one of the questions we had in the doc was about how did I first hear about generative models or something like that. So it's worth knowing generative models have been around for a long time. Yoshua was training generative models of text, like neural language model type things. GPT, I guess it wasn't a transformer. So GPN, GP Neural Net. GP Neural Net Zero, way back in the early '90s, I forget exactly when. So there were neural nets that generated text. Those were not particularly well known in academic circles worldwide.
When I studied AI at Stanford as an undergrad in the 2000s, most of what I'd heard of for generative models were structured graphical models, where there's quite a lot of those in the deep learning book, and those have died since I wrote the book. I think most people that study AI now don't learn them at all. But quite a lot of my undergrad education was about structured graphical models. One of the main books that I read when I was learning was reading this huge, thick book called Structured Probabilistic Models by Daphne Koller and Nir Friedman. Daphne Koller was this genius that got her PhD at a super young age and then became faculty super young at Stanford. And she was still at Stanford teaching the course when I was there.
And so at the time, people thought of generative models as these things where you would use them to represent gene expression data. And they would be really, really bespoke for different kinds of data. So you might say, here's this gene that affects this enzyme, and it interacts with this other gene that represents the thing the enzyme acts on. And you model directly how those two things interact with each other. And you draw a graph in terms of nodes and edges. So you'd have this thing that looks like it's a really complicated spider web. Today, our generative models are these really nice grids. Text is one token influences the next token influences the next token. So it's just a straight line in a grid, or like images, just a two dimensional grid. Back then, you would have these big spider webs of just, you've got thousands of genes, and each one is connected to five random genes next to it. So these were things called Bayesian networks or Markov random fields. And so they'd have these models, and then they would learn probability distributions over them. And then they could solve really small inference problems. If we changed this one gene, what would happen to this other gene? They mostly avoided doing things like try to generate the whole graph all at once.
I took that class with my friend Ethan Dreyfuss. He's actually the person that told me about deep learning originally. So back then, generative models were really this hard core science thing that you would use for protein folding. I guess it still is now for AlphaFold. But when we took the class, Ethan and I were talking, we were like, someday people are going to use this to make video. In 2008 we were talking about that. I wasn't thinking it would be me. But I took the class and I was like, someday this will be used to generate video. It's not going to stay used for genes forever.
And then eventually Ethan told me about deep learning. I went off to study deep learning with Yoshua. And then by 2014, there was a video of GANs. And at that point, that was like, I mentioned earlier, there were lots of pieces coming together at the party. Some of those pieces were, since taking Daphne's class in 2008, I'd had this idea in my mind of, someday people will use generative models to make video. We didn't do video in the first GAN paper, but we got there eventually. So I very much had been thinking there will be generative models for video eventually.
If you look at my talk on GANs that I did at NIPS. We didn't get invited to talk at the NIPS main track where the talks are reviewed. We got invited to talk at a NIPS workshop, which is the less prestigious, unreviewed talk. I was saying basically we want to do video conditioned on actions so you can make simulators and basically let agents think about their environment. Imagine what would my environment look like if I took specific actions? But a lot of what we were doing with the GAN was trying to steer things really strongly toward generating media and having imagination for agents.
And one thread was we were trying to lift the generative models world out of this reason about how one gene affects one protein world, and really strongly into the GPU-driven generate video world. And then the other thread was the deep learning community was already very much in the let's look at images world. They didn't need to be pulled that way. They were doing really small images. And so part of what we did was we used convolution to try to push bigger images. We didn't get very far in the first paper because we only had two weeks. But a lot of what we did there was we pushed people more toward generation.
The community had been really stuck on looking at samples and evaluating how likely they are. ICML 2013, that conference where I said I wrote the most cited paper, Maxout, there was actually a speaker there, Ruslan Salakhutdinov, that came to the Deep Learning Workshop. And during his talk at the Deep Learning Workshop, he actually got up and said, if somebody shows you samples in a paper, you should just regard it as quasi fraudulent. What you really want to look at is the probability that a model assigns to samples in the test set. And that's the only way you can tell the generative model is any good.
He is right that there are a lot of problems in evaluating a model by looking at the samples that come out of it. For one thing, you could just memorize training set samples and then show samples in the training set. And it's actually pretty hard for a human observer to tell. If you have a really big training set and then somebody shows you, hey, look, there's a really high quality photo of a face. For all you know, it's just a training set sample. So you have to have evaluations that automate the inspection of samples to detect problems like that. And it's actually been pretty hard to develop metrics that look at samples.
But the problem with what he's saying is that just looking at test likelihood doesn't tell you anything about how good the samples look. There's a paper called A Note on the Evaluation of Generative Models that came out, I think about a year after GANs, that actually demonstrates these two things are totally separate. You can have models with great test likelihood and terrible samples. You can have models with great samples and terrible test likelihood. And for GANs, they don't even compute a test likelihood. So it's not that their test likelihood is terrible. They won't measure it for you. So you can't even run the metric.
So for all kinds of things you might care about, like speech synthesis, chat bots, image generation, art, you can make things that do that task, but just aren't measurable in the way that the community was doubling down on in 2013. So in 2013 the community was saying, you really must measure the probability that your model assigns to test data. And in 2014 we showed up with this shot across the bow saying, no, we're not even going to measure that. We're going to make this algorithm that doesn't even let you measure that if you want to come in as a third party and try to do it. We're going to double down the other direction on let's just generate really good images. Our images are so much better than anybody else's that we're going to get attention for doing it.
It's pretty funny now. If you look back at the first GAN paper, you wouldn't think those images are good enough that they got attention. But they were. Everybody else's images were super blurry and things like that. And so a lot of what we did there was we really pushed things toward the samples were the product rather than the probabilities were the product.
Peter Bauman: Yeah, that's a really interesting way to think about that early impact and how GANs played this role in that time of proving that deep learning worked and that it could do something. And what it could do is it could now produce. You've now demonstrated this compelling use case. And then yours was the first paper. But then this kicked off this really, really long line of future research. And I wonder if you can just talk about some of the subsequent research that carried GANs on from your invention and continued that trajectory that you set off. What were some of the key drivers after you developed them?
Ian Goodfellow: So I guess there's LAPGAN, was one of the first papers that showed that you could make much larger images than we did. And they did that by generating at multiple scales. Generate an image, scale it up, and then add details to it.
There was DCGAN in late 2015. In our first paper, we actually only had one convolution layer. So DCGAN was able to be more serious about having multiple convolution layers, really taking a convolution network and turning it upside down completely and having a fully convolutional generator. They were able to get much better images. That's when GANs really started to take off. They were covered in Scientific American, and that's when there was a lot more attention coming into the field. After that, there was really steady growth in the quality of images and the size of images coming out of GANs.
In terms of images, the endpoint of where GANs ended up being really good was the different models of StyleGAN from NVIDIA, where they use style transfer techniques at each stage of generation and transfer styles onto the image to improve it at each step. And they can do things like render the same car from different viewpoints by transferring a view style onto it.
There's also been a lot of work on things like style transfer as an actual feature of the GAN rather than an internal implementation component. So there's a model called CycleGAN that can learn to transfer styles across photos without ever needing to have supervised examples of the style transfer. So for example, if you want to turn horses into zebras, you can train on a bunch of horses and a bunch of zebras, and it will learn to turn a horse into a zebra. That's important because it can be really hard to get aligned supervised examples for supervised style transfer. Imagine trying to go and get a zebra to stand in exactly the same position as a horse. You just can't do it. And so being able to learn to do the transfer in an unsupervised way is important for being able to do the task at all.
And then there's been a lot of work on GANs for video as well. I guess there's also been stuff about controllable representations for GANs so that you can actually change the hidden layer and get what you want in a predictable way.
Peter Bauman: Were you surprised at all that... So you were thinking about GANs and video in 2008, and still in 2014 is when you created them. But were you surprised at how quickly you were able to, or researchers were able to, achieve video from this technique?
Ian Goodfellow: No, not really. I guess we had been working on video in Yoshua's lab as an input to classifiers since 2010, at least. I guess other labs had been doing video in other contexts too.
I guess one thing to understand about GANs is they are really designed for speed and efficiency. A lot of how I got into deep learning originally is that I was a hobbyist video game programmer before I got into deep learning stuff, specifically. I mentioned that my friend Ethan Dreyfuss told me about deep learning. That was partly because he knew that I knew how to write code that runs on GPU. He had heard of Geoff Hinton's deep belief networks and was interested in whether it would be feasible to implement them on GPU. So he heard of deep belief networks. Geoff was running them on MATLAB on CPU. Ethan realized that they were extremely parallelizable and that they would be more efficient to run on GPU. He knew that I knew how to write GPU code for games and came and asked me, do you think we could run this on GPU? This was late 2007, early 2008. And we ended up building the first CUDA machine at Stanford in order to run them on GPU.
I guess basically with GANs, I was trying really hard to make something that would be media generation friendly. They do everything in one pass. There aren't any funny back and forth access patterns that would cause cache misses. There's not any Markov chain iteration like you see with Boltzmann machines or diffusion algorithms. There's not any one token at a time, waiting for one token to generate before you generate the next token, things like you see with GPT style models. So GANs are very much designed to do things like video really fast. That's one of their key advantages. And that's part of why they didn't get outdated for as long as they did.
Peter Bauman: Well, yeah, it's really interesting that you had that in mind for so long. One thing that is related to culture in GANs is there's this idea that even artists have. I think it's a growing idea with artists that this idea between the distinction between artist and engineer is really not that important or is nothing. And that if you really take that to its logical conclusions, that you could argue that the greatest artists of this generation are engineers, are machine learning engineers. Some artists have made that argument, like Mat Dryhurst. And part of the reasoning for that is because they argue that the greatest creators make protocols, and then others can downstream take advantage of that protocol creatively, and that machine learning engineers are some of the greatest protocol makers of this generation. So I just wonder what you think about that view that engineers and artists, the distinction is not really that different. And do you also see yourself as an artist?
Ian Goodfellow: I do see myself as having artistic tendencies, but I actually see that as partially separate from the engineering work. I was actually interested in being a fiction writer. A lot of what I was interested in exploring in fiction was writing psychologically realistic prose. There's a lot of modernist fiction that writes in a stream of consciousness style, where it's intended to be, we see how a person's mind actually operates instead of looking in from a third-person perspective. I was also just really interested in figuring out more about how do minds actually work. Somewhere along the way in my education, I ended up going further on the figuring out how minds work thread than the writing fiction thread, and ended up putting most of my creative energy into machine learning research, and then found that I was actually feeling creatively fulfilled by doing the machine learning research and didn't feel as much like writing fiction. Mostly, it is a lot of emotional energy to do the research.
But I don't personally see the engineering work as art. I think art involves having some message, some emotional expression, something that can actually be conveyed to other people. I feel like the protocol is more of a medium. I know, like Andy Warhol said, the medium is the message. But when he said that, I think he means more of like, when the artist selects the medium for a particular piece of art, the artist is sending a message in the selection of the medium for that work, not that the person who invented a particular medium is sending a message. We don't remember the person who invented canvas as a great artist, even though they certainly have made a contribution to the art world. It's just a different contribution.
There's different degrees of artistic thought going into work in machine learning. Some algorithmic work requires more thought about what does it mean to be intelligent and so on. Some is more just like, well, let's analyze this and scale it up. Some of my work is a little bit more on the more artistic side, and quite a lot of my work is more on the let's print out a big report and analyze it side. Certainly coming up with the idea for GANs, I had a bit more of an artistic feel to it than a lot of my days where I'm like, let's set a debug trace and print out what crashes.
Peter Bauman: It can depend maybe based on what you're doing, but you have felt creative fulfillment in your engineering work. Yeah, that's really interesting, to the point where you didn't feel like another creative outlet was as necessary. You understand the creativity involved with some machine learning engineering. Are you surprised at all by the amount of creativity that GANs have unleashed that have been beyond your control, but that you were directly maybe the first domino to fall, perhaps, or one of the early dominoes?
“We don't remember the person who invented canvas as a great artist, even though they certainly have made a contribution to the art world. It's just a different contribution.” — Ian Goodfellow 1:00:33
Ian Goodfellow: I think as AI has gotten more powerful, there is more thought about how much should we open source, because there is more of a risk of misuse in ways that could really harm people. But in 2014, the kinds of things that we released seemed more likely to be used for artistic expression than for really harmful uses. So yeah, GANs have been used more or less as I hoped they would be. I think you can see from my first talk on GANs at NIPS, I did think that they would be used for simulation more than they have been, and that's not turned out to be a thing that was easy to do. So there's been fewer engineering uses than I expected, but the artistic uses are more or less what I expected.
Peter Bauman: Did you have text to image in mind when you were thinking about GANs? Was that already something that you thought, oh, well, this will come, or, oh, this will be useful, or, oh, GANs will be able to help achieve this?
Ian Goodfellow: I don't remember how much I thought about text to image, specifically. It's always hard to remember what did I think when. I'm trying to get a handle on what was I thinking at the time. I did look at my slides from NIPS 2014 recently, and I do talk about text to speech, specifically.
I think one thing to keep in mind is that back then there weren't good data sets for text to image. Today, there's quite a lot of money in the field, both to train large models and to gather large data sets. In 2014, I was probably thinking more about what's feasible for an academic to gather, and it would have been a lot harder to gather the data set needed for text to image to work really well. In 2015 or so, we did see a lot of the other direction of image to text for captioning. My guess is that it was probably a little bit hard to foresee getting the resources for text to image to work really, really smoothly for a long time.
Peter Bauman: Yeah. It happened pretty quickly after the paper. I know people like Soumith and Alec and Elman. Yeah, these are some of the early...
Ian Goodfellow: Well, I guess what Elman was working on was not related to GANs.
One thing that really has surprised me a lot is how much the industry is willing to pay for generative modeling. I mentioned how much effort I put into making sure that GANs are extremely sample-efficient. And for example, that's why they were able to make videos. Things like GPTs, I just would not have expected that a company would be willing to spend what they spend for that. I would have thought that you would have to demonstrate more utility before spending that amount of money to train that generative model. A lot of how I personally got to GANs was assuming that it would just not be economically viable to train the kinds of models that we have today. It's not that the kinds of models we have today were impossible to think of. The transformer part is new, but that's an architectural component. But the overall generative model algorithm we have of generating one token at a time, that's been around since the '90s. Yoshua was already doing that, and we just thought it was prohibitively expensive.
Peter Bauman: Yeah, it's interesting that one thing that was also mentioned in that NeurIPS 2024 talk was the GANs and the legacy of scale, and that they really demonstrated the power of scale. That's interesting now to think about how that got out of hand in a way, because companies were willing to spend billions of dollars on this, on training at this scale. So maybe that's what was surprising, especially at that time before there was really any, or there was much, commercial interest in AI at the time. These were still back in the university research lab. How do you feel about that legacy of scale today? And I think a lot of the criticisms that AI and deep learning get is that it requires all the scale and all these resources, and that to get this data and the scale you often have to use unscrupulous means. And how do you think about the dangers and the flip side of some of this, with mostly the positives that we've talked about?
Ian Goodfellow: Yeah. I don't personally think that GANs are the cause of the current emphasis on scale. I actually feel a little bit like GANs are a relic of the previous era of being careful about cost. But you are right that GANs were a place where we saw that there were big returns to scaling them up. But when we scaled them up, they were still relatively cheap. We are talking, generally, the experiments run on a single machine, not run on a whole data center.
I remember around 2019 we started to see results with things like GPT-type images beating GANs, and then GAN researchers messaging me saying, look, this GPT beat our GAN results that we published, but look what they spent to do it. They're running a whole neural net per pixel. And in the GAN literature, we're running a neural net per image. What the hell? So it was surprising to me that people were willing to make that big of a leap in how much they would spend for relatively small improvements in the output.
I think where we really saw a change in spending habits is with the idea of scaling laws. People liked that there was this, I guess business type people liked that there was this function where you could predictably get more performance by putting in more resources. I'm a little bit surprised at how popular the scaling laws are, because what the scaling laws tell you, in my opinion, is actually not really what you would want. The scaling law is basically telling you at every point across the scaling curve, you get diminishing returns, and the diminishing is pretty strong. I feel like a lot of people saw it and said, hey, at every point across the curve, you get returns to scale. And I'm like, well, what did you expect? Yeah, of course it'd be returns to scale. What bothers me is that the returns diminish pretty fast. And people still decided, let's just push to the right really hard across it, because we've now seen this curve that we think will go forever. I think that's where the increased spending came from.
But I feel like the GAN philosophy would be more of, can we shift that curve down, or can we try to change the tilt of that curve somewhat to not have to spend nearly as much? The GANs are more like, choose an amount that we're willing to spend and then be clever to get really good results within that budget. The new approach is more like, well, let's not use artistic creativity about the algorithm. Let's just take a really dumb algorithm and pump money into it.
Peter Bauman: A more brute force approach.
Ian Goodfellow: Yes.
Peter Bauman: Earlier, you demonstrated your skepticism about a fully autonomous artist. Is that related to also how you think about AGI? Because I guess that's one of the big things over the last few years in Silicon Valley, is to what extent people believe in AGI that seems to...
Ian Goodfellow: I wasn't trying to imply anything about is AGI soon or far. I just mean, out of the systems we have now, the systems we have now are not AGI. So the systems we have now are not similar enough to AGI for them to be considered an autonomous agent. No, I think it's very hard to predict how long until AGI.
I can talk a little bit about how hard it is to predict how long until things. I've been really wrong in both directions. So in 2007 I was working with Andrew Ng again in his lab, which is also closely tied to a startup called Willow Garage. So I was working both places, Andrew's lab and Willow Garage, on household robots. And this was maybe before Uber existed, definitely before Uber was well known enough that I had heard of Uber. But the idea was basically Uber for household robots. Our goal was someday you have an app where you summon a robot and it comes and cleans your house and builds your IKEA furniture for you. And in 2007 we thought that was 10 years away.
I'm the person where any large group I'm part of, I'm always questioning its beliefs and values and so on. And I was not questioning that one. I felt like we were on track. Larry Page came to Willow Garage to judge the intern contest, and we successfully had a robot waiter serve him. We didn't use a real drink because we didn't want the robot to accidentally crush glass and spill it on him, but it gave him a plastic drink bottle. And I was like, okay, not bad for 10 years off from target. And clearly we did not end up with Uber for robots within 10 years of then, even though I had felt like that was very easily on track.
And then in, I guess I'd say, 2011, we were working on trying to use deep learning to solve supervised learning, where you look at images and you say, what's in the image? And in Yoshua's lab we were working on really small images. We were trying to look at things that were like 32 pixels by 32 pixels and say, is this an airplane? Is this a car? And it wasn't going all that well. And I felt like it was years and years until it was solved. And then in 2012 there was Ilya, Alex and Geoff did ImageNet, and we just really did not see that coming. We'd been focused on things like warming up with generative modeling before proceeding to supervised training. The generative modeling part was too expensive, so it was forcing us to use small models. We didn't realize that if you use big models in pure supervised training, you could get a lot further. And we were using other data sets instead of ImageNet. ImageNet just had a lot more examples, so you could get further with that.
So I feel like I've been very wrong in both directions. It's a lot easier to tell where are things eventually going. You can tell that capabilities will keep increasing, but it's hard to know exactly when and in what leaps. And that's why you get people shocked by things like, oh, GPT-5 wasn't the leap we expected, and things like that. But you do tend to see long term things do keep getting more and more capable over time.
Peter Bauman: I guess it's one of the things that is the legacy of GANs too. And it's something that you've talked about before. And it was also mentioned in that 2024 look back at NeurIPS. And that's about the legacy of GANs and unsupervised learning. So you said in Wired in 2017, you said, "In the process, these two neural networks can push AI towards the day when computers declare independence from their human teachers." What was the role that GANs played ultimately in that progression to more unsupervised learning?
Ian Goodfellow: Well, I think we've ended up with unsupervised pre-training being a really big part of modern AI, and it just wasn't prior to GANs. I don't know how much you can say GANs are the reason that we have pre-trained transformers as a big part of bots and chatbots. But I do think GANs are a big part of why generative modeling was a big part of the conversation. I guess it's worth pointing out Alec Radford invented GPT, and he'd been working on GANs immediately prior to that. You could interview him about what's the causality there.
Peter Bauman: Yeah, that link is really interesting, because I do think that the GANs story with Alec Radford is really interesting. Because he did go on to do GPT and Jukebox and these other multimodal realizations of deep generative models. But it's interesting that the very first thing that he gravitated towards was GANs.
Ian Goodfellow: Yeah. So I guess one thing that has shaken out a bit differently than I expected was, GANs are very sensory, right? I had been imagining that we would use unsupervised learning to discover the world in the same sequence of steps that humans and animals do. Humans spend a significant amount of time being pre-verbal, spending a lot of time seeing and hearing the world. You can actually hear in the womb before you're even born. Then after you're born, you can see. You spend a lot of time moving around, sticking things in your mouth, but not really talking very much. Then eventually you start talking, and you don't really start learning a whole lot from massive amounts of text and things like that until you're quite a bit older, right?
The way the AI industry has gone is we bypass all of that, and we just go straight to tokens now. And just from the moment of creation, the GPT is just getting high-level cognitive tokens and is bypassing forming these visual representations of the world. So I would have thought that you needed to spend the time on the sensory motor stuff to form concepts of objecthood and so on, or that you wouldn't be able to have enough of a concept of things that you're referring to to understand the tokens very well. But it seems like you're able to get pretty far in a world of pure tokens.
So that's the one thing that really surprised me about the pre-training world that I expected versus the pre-training world that we got. I thought we would need something, not necessarily the GAN algorithm, but like, pre-train on images, videos, audio, and then move on to learning a lot about knowledge from text. And it's been surprising that you can just go straight to pre-trained on text and actually seem to derive meaning from it without knowing what it refers to, without having an external reference.
Peter Bauman: That does seem to be so interesting about how deep generative models have developed, that text and language seems to be, in some ways, the winner, and that it can now be used in this multimodal way to create all these different things. Maybe this is the last thing we can talk about, but I'm interested in that platonic... Are you familiar with Phillip Isola's platonic representation hypothesis? Do you know?
Ian Goodfellow: I've met Phillip Isola before, but several years ago when we were both working on GANs stuff more. I'm not familiar with that particular term.
Peter Bauman: Neural networks trained with different objectives on different data and modalities are converging to a shared statistical model of reality in their representation spaces.
Ian Goodfellow: Oh, that's pretty interesting, actually. So he's saying, lots of neural nets trained in different things all come up with similar internal representations, basically.
Peter Bauman: Right. And so that might, in some ways, explain why these vision models are learning the same representation as text models. Or that is what it seems to be saying. I just wonder what you think about that. And then also just, you said something before about how AI can imagine worlds, or eventually will be able to imagine the world in more realistic detail and realistic images and sounds. And it seemed like you were thinking about a multimodal future. And you said that this will encourage AI to learn about the structure of the world that actually exists. Do you see any relationship there between that idea with Isola's platonic representation hypothesis, that these models are learning some underlying structure of reality? How do you think about that?
Ian Goodfellow: I'm only briefly seeing what he's arguing here, but he seems to be going a step further and saying that it's okay if you don't have overlapping experiences, that maybe if one model has only vision and text and another has only vision and audio, that they can still learn the same representation of the same thing. Yeah, that's pretty interesting.
One other thought I have about this that strikes me as interesting about his hypothesis is I worked in machine learning security for several years. And neural nets don't learn just the same good representation. They also learn the same mistakes. So when you attack a neural net and try to intentionally fool it and you find inputs that they misclassify, different neural nets will misclassify those adversarial examples in the same way. So even though those input points are just random garbage, it may look like random garbage to you or me, but the same neural net...
Imagine I'm trying to fool a face recognition system to get access to a secure facility, and I want it to think that I'm the director of the facility or something. I might be able to train my own neural net to recognize faces and then make weird glasses I can wear that make my neural net think I'm the director of the facility. If I then wear the same glasses and go to the facility, their neural net has a pretty high probability of being fooled in exactly the same way as my neural net, because of this convergent representation property.
So it's interesting. What he's describing is a known property that's considered a security risk in the machine learning security literature. Except there we're looking at two different neural nets trained on different examples from the same domain. So two different neural nets trained on vision. But I hadn't heard of there being convergence across, like, vision and text. So that's pretty interesting to me.
Peter Bauman: Yeah, it reminded me of some of your ideas. And then, well, yeah, this was a real treat. You're somebody who will be part of the story of AI art history. GANs will be in AI art, just art history textbooks moving forward, and it's going to be part of art, whether people maybe realize it quite yet, but it will be. It was really meaningful to be able to talk to someone who's had such an impact on art and ask them about some of these questions and ask them firsthand about that impact. I really thank you so much. It means a lot.
Ian Goodfellow: Oh, yeah. Thank you. I really like your timeline. I'm looking forward to seeing your other work. I hope we can be in touch and maybe talk again someday, maybe hopefully meet in person sometime. But yeah, I really appreciate at least today and the time.
Peter Bauman: Oh, yeah. Very welcome.
Ian Goodfellow: All right. Thank you.
Peter Bauman: All right. Thanks, Ian. Bye.
Ian Goodfellow: Take care. Thank you.