300 odcinków
- Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time.
One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time.
To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him.
Who’s building real-time interactive world models?
First, some context about world models that can generate interactive video and audio in real-time.
Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December.
Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table:
Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.”
But as our interviews with Runway show, real progress is being made.
The central idea of WorldPrompt
WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game.
“Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.”
As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out.
“You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said.
But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character?
“Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.”
Sindi added that more training plus scaling the data and models is resulting in “better following.”
How a video model becomes a real-time runtime
Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us.
“There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.”
GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz.
Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively.
“And after that, we work on making it real-time through distillation methods,” Kahlow added.
Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.”
Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results.
The challenges of real-time generation
Germanidis admitted that there were issues with how it generates real-time interactive video.
“The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.”
Sindi told us there are also challenges dealing with “infinite generations” of content.
“There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.”
Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.”
Causality and correctness
While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions.
Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly.
“If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.”
Sindi told us that evaluation gets harder the more complex interactions get.
“If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?”
To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.”
More than gaming — there are agent use cases too
Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works.
Another, more intriguing, use case is to use it to test agents at scale.
“Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said.
But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding?
“So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.”
Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.”
Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models.
“You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.”
Anastasis Germanidis
* LinkedIn: https://www.linkedin.com/in/agermanidis/
* X: https://x.com/agermanidis
Timestamps
00:00:00 Introduction
00:05:17 Runway’s Origins and the Bet on Generative Video
00:12:23 The Stable Diffusion Story
00:18:44 Gen-2, Controllability, and the Weekend Hack
00:23:02 From Video Generation to World Models
00:28:03 Learning From the World, Not Just Language
00:35:04 Sora, Runway’s Existential Crisis, and Gen-3
00:39:39 Why Real-Time Video Is Inevitable
00:43:06 Interface World Models: Software Without Code
00:50:25 The Fully Neural Operating System
00:55:11 World Models for Robotics
01:02:32 Robot Policies and World Action Models
01:07:47 The Lucid Dream Test
01:11:41 Video Agents and Omni Models
01:23:12 Artists, AI, and Creative Workflows
01:27:14 Physical AI and the Future of World Models
Transcript
Introduction: Runway, Creative AI, and the Early Thesis
Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome.
Anastasis [00:00:08]: Good to be here.
Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out?
Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of those generative models, rethink how creative tools are made. and as we built out the research behind, our generative models, it then became clear that they were useful far beyond that as well.
Anastasis’ Background: Art, Simulation, and Machine Learning
Swyx [00:00:57]: And it is more obvious now with, like, the real-world stuff and the world models that we’ll talk about later. I’m just kinda curious how you go from a background in, like, Zocdoc and, computer vision into Runway. Like, take us back to that early conversations with Chris and, whoever else is on your founding team.
Anastasis [00:01:14]: I was always splitting through those two worlds. One was the I had my own art practice. I was making a lot of interactive art, I think for a long time. and then on the other side, I was working in startups, and I was working as a ML engineer, as a backend engineer at different companies. I’ve always been interested in, coding and computation, and especially interested in simulation and brought it back into my early artwork as well. And at the same time, I was interested in
Swyx [00:01:43]: The personal site has a few, right?
Anastasis [00:01:44]: Yeah.
Swyx [00:01:45]: Is there one that we should pull up? Just in case there’s something that’s like. I just like to go down memory lane.
Anastasis [00:01:50]: Yeah.
Swyx [00:01:50]: Okay, what is this?
Anastasis [00:01:51]: So this was, a project that I made, I think back in 2015, where I built this software that would give, voice instructions to people in a gallery space. So it would coordinate interactions between people. And so it will first give you an identity, like you’re an, architect, you’re 30 years old, and, you like sports. and then it would match you with another person, and you have this completely generated interaction. language models were not quite there at the time, and so it was it was a mix of some templates and some, like, some Markov chain-generated text, and it would just completely simulate these small talk conversations between, everyone in the gallery space. so was always very fascinated on the one hand with, generative models and, like, the early machine learning work that was being at that time. But at the same time, there was this separate thread of simulation and what it means. Like, what can we learn about humans by creating those very simple models of their interactions and their behavior?
Early Generative Art: pix2pix, GANs, and Uncanny Valley
Vibhu [00:02:56]: Did you generate the prompts or, the 30-year-old, whatever? Was it you generating them? How’d you, how’d you
Anastasis [00:03:03]: Exactly. So the program would just generate- those, from. Yeah, a lot of it would be Mad Libs style of just
Vibhu [00:03:10]: Yes
Anastasis [00:03:10]: You have lists of different professions, lists of different,
Vibhu [00:03:14]: Hobbies
Anastasis [00:03:15]: Personality types, lists of different, ages, things like that. And then it would just combine those things together. And then maybe the next project we go is, Uncanny Valley, Uncanny Road, which was
Swyx [00:03:27]: Gans
Anastasis [00:03:27]: One of the first projects that, we built with, one of my two co-founders, Chris. This was taking, pix2pixHD, which was one of the early image-to-image models that NVIDIA released back in 2016 or 2017. and it was a model that would take a semantic map of a scene and then generate a photorealistic, let’s call it, output. very early days, so it was not very high-fidelity outputs, but it w I think was the first image-generation model that could generate at 1K resolution. And it was all trained on self-driving datasets. So the semantic categories it would support were only, things you would encounter on the road. So it would be pedestrians, traffic signs,
Vibhu [00:04:16]: Stoplights
Anastasis [00:04:17]: Bikes, stoplights. And so that was one of our first indications that we built this and people were making all this, like, very surreal imagery of, yeah, a million plus a million pedestrians or a million traffic signs or, like, gigantic humans. And it was a indication that you could take a model that was trained on this very boring dataset, essentially, of, like, not that many interesting things happen when you’re on the road, and then you can repurpose it and go very out of distribution and make something that was artistically compelling. And that was It’s a summary of the thesis of Runway in some ways, that you can take the same generative models, and if you look at them from another direction, if you build interesting tools around them and you give them to artists, they’re gonna do things that you don’t expect.
Vibhu [00:05:02]: Very cool. I like the, UX of it. You’re just given an empty canvas, try whatever, do whatever. And then the other one, like, you see everyone with wired headphones? Like, that’s, that’s a sign that it’s, it’s very
Anastasis [00:05:16]: The Apple
Vibhu [00:05:17]: Yeah
Anastasis [00:05:17]: Apple, your version.
Vibhu [00:05:17]: Original ads. Yeah. Take us to today. You’ve been doing this for seven years at Runway. How have we got to this? Like, how do we go from driving simulator data to all this? And you cover the whole stack of generative media?
From Creative Tools to a Research Lab
Anastasis [00:05:33]: Interestingly, we’re almost back in, we’re, we’re full circle. We’re, we’re now applying our models and beyond creative tools into real-world scenarios. But it was a, it was a long journey. It was very early on we realized the first version of Runway was a way to easily use the, all the open source model of the day, things like pix2pix to. and give them to artists. That was the initial idea, is those models are too difficult to use if you’re not a machine learning engineer. Like, what happens when you give them to artists? Very quickly, we realized we needed to build a research org, inside of Runway, and that happened maybe on year one. And, a lot of the mandate there was. The image-generation models of the time, the video generation models of the time, or there were barely any video generations all the time, but they were not quite there where they could be productionized and brought into tools that would be part of creative workflows. so we need to push the frontier of the research. And so maybe the first four years of Runway, research was almost happening on the background until there was a moment in 2022, with latent diffusion, with, DALL-E 2, where, there was that step function change, and you guys maybe remember around the time.
Swyx [00:06:49]: I started in this space because of latent diffusion and Stable Diffusion.
Anastasis [00:06:54]: Yeah.
Swyx [00:06:54]: Because I was like, “Wow, this is not only, like, feasible, it is doable on consumer hardware.”
Anastasis [00:07:01]: Exactly, yeah.
Vibhu [00:07:01]: I think the delta is also huge. Like, I learned pix2pix. Like, this was intro to ML, the TensorFlow, like, Jupyter, Google Colab notebooks were like this, and then you have a sudden step function change, with diffusion and whatnot. Any other ones since that. Like, there were clear examples of what early diffusion were to get to here. Any other changes in key technology research?
Green Screen, Rotoscoping, and Early Runway
Anastasis [00:07:26]: Between, 2018 when we started and 2022?
Vibhu [00:07:29]: Yeah.
Anastasis [00:07:29]: So one of the early work that we did in Runway was solving segmentation, image and video segmentation. It was a very important problem because most VFX involves essentially separating
Swyx [00:07:42]: Rotoscope
Anastasis [00:07:42]: Subjects. Yeah, rotoscoping. Extremely manual process. Nobody enjoys doing that. and so a lot of the early days of Runway was building this tool. It was called Green Screen, and it was for a long time the main thing that people were using Runway for. It ended up being used in, Everything Everywhere All at Once and a bunch of other high-visibility films and series. But that was essentially, Runway for a long time was a post-production tool until latent diffusion and generat- Gen-1, Gen-2, happened.
Swyx [00:08:12]: Cool. let’s, let’s go past that moment. You’ve come a long way. Then you started releasing your own models. Maybe describe that journey as well.
Scaling Video Models and the Bet on 1,000 A100s
Anastasis [00:08:20]: Yeah, so we go to the other point, yeah, in mid-2022 when it became clear that we’re doing research at a fairly small scale of compute, and it became clear that, like, scaling laws would apply to, image and video gen in the same way that we’re applying to language generation. So we made a big bet, and I think at so at the time, we signed this deal to build a cluster of a thousand A100s, which at the time we were a Series B startup. That was a almost, slightly irrational decision maybe, but we really believed that if we trained a video model at a large scale, we would get, like, a great model at the end. And at the time, the goal or we set the goal around fall of 2022 of what is, what does the latent diffusion, Stable Diffusion moment look like for video? And at the time, the best model of the time was called CogVideo. it was one of the early video models. It was very 256 by 256 resolution, very not very high quality. and so we decided we’re gonna build out this cluster, and we’re gonna just invest in, like, in building out our own video model. it became clear as we’re training Gen-1 that it was difficult to get to fully. we wanted to build text-to-video, but it became clear to us that an easier starting point would be to start from video to video. Because when you have a stronger conditioning, it’s, it’s an easier problem to restylize an existing video versus generate the video from scratch. And so we released Gen-1 first back in, it was January of, 2023. Yeah.
Vibhu [00:10:04]: It’s just a fun visual podcast, honestly. Like, if we can see February 2023, what was the state of stuff?
Gen-1: Video-to-Video and Depth Conditioning
Anastasis [00:10:10]: It’s so interesting ‘cause at the time when you see those results, you think this is so incredible, and this is like, it’s almost like image generation or video generation is solved. And then you look back a few years after, and it’s like, it’s It’s just like you get used to the results very quickly, with those models. But at the time when we started seeing those results, it was, it felt quite incredible, and the level of, like, quality that you could get. And, so the Gen-1 was a depth-conditioned video model, so it would turn. it would take a input video, it would predict. it would it would first convert it into the depth map, and then we would generate, pixels with a latent diffusion model.
Swyx [00:11:01]: Yeah, very effective.
Vibhu [00:11:02]: Yeah. I didn’t realize how distracting the blog post would be. Sorry.
Anastasis [00:11:05]: Yeah, but, one of my favorite examples of on those, on Gen-1 was both, if you go up to mode three or mode two, there was this storyboard use case where people would make
Vibhu [00:11:18]: Ooh
Anastasis [00:11:18]: Would
Vibhu [00:11:20]: You can mess around with the
Anastasis [00:11:20]: Make a city out of books or out of boxes, and then they would shoot a video with their phone and then translate it into a photo-photorealistic output. There was all these ways in which those models were starting to be used for storyboarding and also for really. and then if you go to mode four, like, of taking untextured 3D scenes and then turning them into photorealistic output. So we saw a lot of use cases early on where people that were familiar, were power VFX editors would just take a blender, render, and then they would get translated in with Gen-1 or create a scene in Unity and then take a capture a video of it and then translate into, restylize it. So I still think video to video is powerful. I think we had a recent video-to-video model as well, and it’s one of my favorite ways of using those models is essentially using them to use ground truth video as, like, the initial inspiration and then translate into different styles or different outputs.
Stable Diffusion, Stability AI, and Open Source
Swyx [00:12:23]: But I think we’re gonna go into, like, the rest of Runway and catch people up to speed today. I did wanna cover the, let’s call it the Stable Diffusion controversy, or, what happened with Stability AI, whatever. I think there was a two sides of the story. I think there’s part of that is a normal thing of, like, people, join and leave companies, but what is the, retrospective now that, there’s been some years behind it?
Anastasis [00:12:49]: Yeah, it’s a very, it’s a very long story to go into. I think it would
Swyx [00:12:53]: Which I remember you wrote a really long post about.
Anastasis [00:12:56]: We would probably cover the whole hour to go into it in more detail. But, essentially, there was the latent diffusion paper that came in, I think that was at the end of, 2021. And then Patrick Esser, who was one of the researchers behind, latent diffusion, and he worked at Runway at the time, he built latent diffusion in collaboration with Robin Rumbach and a few other folks back, in the in, CompVis, which was, a lab
Swyx [00:13:26]: Like a research group, yeah.
Anastasis [00:13:27]: And, after releasing the early latent diffusion model, they, essentially they were. the goal was to keep working on versions of the model, scale it up, incorporate new data, incorporate new tasks. And Stable Diffusion was the same model, but trained on more compute, and then with a few more tricks, like a classifier-free guidance paper came at some point, I think in the early 2022. And that
Swyx [00:13:52]: Which, like, was a big prompting improvement.
Anastasis [00:13:55]: Yeah.
Swyx [00:13:55]:?
Anastasis [00:13:56]: That improved results. it was trained on better data, so like, the esthetic subset of LAION, but it was effectively, the same underlying architecture. And there was that big training run, that, happened on Stability’s cluster. Stability financed that run. And looking back at that story, I think it was the work to build and train that model was done. It was a, it was a research project. It was done as part of, like, continuation of the latent diffusion work. It then, I think it the model became very successful, and it, I think there were the. And I think as a result of its success, other companies tried to, figure out the commercialization path for it. But for us, it was very important that we try to, we make sure that we. It was meant to be an open source research project, and so the we decided that we should continue releasing versions of it, since that was the original goal of Stable Diffusion, and that led to releasing Stable Diffusion 1.5. There was maybe a day of, a bit of, miscommunication there, but ultimately that was resolved very quickly within hours. so yeah, there was
Swyx [00:15:12]: Okay
Anastasis [00:15:12]: Not a nice
Swyx [00:15:13]: I just wanted to. you have to
Anastasis [00:15:15]: Yeah.
Swyx [00:15:15]: You’re one of the main players in that journey, and so it’s nice to hear from the source of, like, what happened. Yeah.
Anastasis [00:15:22]: Yeah. I think it’s all, it’s all in the past now
Swyx [00:15:26]: Yeah
Anastasis [00:15:26]: I would say. and, like, both companies, Stability took its own path, Runway took its own path.
Swyx [00:15:32]: Yeah. There’s still. James Cameron is backing the new Stability, whatever they’re doing with the Hollywood studios.
Anastasis [00:15:38]: Right.
Swyx [00:15:38]: I don’t know what they are doing. I think one thing that impresses me, and I’m happy to move on, is that back in the that time, let’s say, like 2021, 2022, there was this community of people that you were involved in that was researching all this stuff, right? And, like, from everyone I talked to who was active then, it seemed like it was fairly obvious that somebody would do the hero training run that would produce Stable Diffusion. So, like, I guess the question is, like, you had the you were you had made investments. You were you had the foresight. Is it accurate to say, like, that is reflective of, like, what people were thinking at the time? Or was it still very much like, “Well, we’ll use it as, like, a post-production tool or something. I don’t know.”? Like, where in the sentiment were we that maybe you can think back to, like, what the community was like back then?
The Early Creative AI Community
Anastasis [00:16:28]: I reminisce and I think very fondly those early years, from like 2018 to 2022, because it was a very small community that, as you said, were very convinced that this was gonna be a big thing. And at the time, anyone who. Because it was such a small circle and, everyone who would, like, be part of that circle and, like, make projects with it would, immediately get, go viral. so like
Swyx [00:16:55]: And you didn’t know who they are, right? They’re just some name on a, GitHub or Hugging Face somewhere.
Anastasis [00:16:59]: Exactly, yeah. So I remember one of the first big viral moments of creative AI was, there was the neural style transfer paper
Swyx [00:17:09]: Huh
Anastasis [00:17:09]: That
Swyx [00:17:10]: Something dreaming?
Anastasis [00:17:11]: I think it was called neural style transfer.
Swyx [00:17:14]: Okay.
Anastasis [00:17:14]: There was also Deep Dream, the puppy slice
Swyx [00:17:16]: Yes
Anastasis [00:17:16]: Which was, also really cool. but, yeah, there was this project that, Jim Kogan, who was an early advisor of Runway and one of those,
Swyx [00:17:25]: Marketing guys
Anastasis [00:17:26]: Big, creative AI, folks, he literally just, like, showed a video of himself taking the New York Subway and going over the Williamsburg Bridge and then stylized it with, I think in the style of Van Gogh or, like, one, painter. And that was. Like, at the time, that was, like, so cool and it went viral and it was completely revelation to people that you could do this with generative models. And that was only, it was less than. It was maybe 10 years ago. So just, like, as an indication of, like, how quickly things have gone.
Vibhu [00:18:02]: It’s pretty crazy. Like, even since then, you’ve got people at every level of the stack. You’ve got devs, creatives, artists, hobbyists. You’ve got everyone using it. And for people that tried stuff early, they’ll remember how hard it was to use regular diffusion, right? Like, nowadays, you can use your favorite ChatGPT image gen or whatever, give a sentence, get a beautiful output. But diffusion was like, the whole ultra HD, 4K, high resolution. Like, prompting these things was very different. anything you learned on the tooling side, like from the offerings you guys have now, so like creatives, devs, you really took the. Research and brought it to everyone to use. anything interesting there to share?
From Gen-2 to Controllable Video Generation
Anastasis [00:18:44]: We had to build the entire model serving infrastructure for video diffusion models. There was nothing else, already, like, because we had Gen-2 was the first text-to-video model, I think, out in the market. So many things that we learn over time. I think the I think the biggest one was, like, we. it was very clear early on that text-to-video was not gonna be the answer. Like, you. Like, people wanted a lot more control than that, and so we invested in, like, control building on top of those models very quickly. how do you use the camera trajectory as control? How do you use an initial input frame as control? So that was a very early learning for us. With text-to-video was, like Gen-2 was an amazing, step function improvement in the quality of video models, but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it. There was no. You couldn’t really control the camera motion. You couldn’t control the object motion. And so the first year, in 2023, was really all about what are all the interesting ways in which we can condition those models? And it was a lot of just post-training rounds on top of the base model to figure out, like, what, -- how do people wanna control them? And so there was, like, this quick succession of the we it was called Motion Brush, which was you could, like, you could draw arrows and dictate where things should move in the scene.
Vibhu [00:20:09]: That’s so cool.
Anastasis [00:20:09]: There was camera control that was you could just describe, like, how you want the camera to move in the scene. And because we work with filmmakers from the most of the history of Runway, we immediately got this feedback and got this, decided that this was worth investing in. And so control ability became a big theme, I think, very early on as we were building, as we were building those models. Something fun that I haven’t really talked about too much was just how Gen-2 came to be out of Gen-1. So it was a bit strange because we announced Gen-2 two months after Gen-1 and
How Gen-2 Came From a Weekend Hack
Vibhu [00:20:43]: We’re accelerating.
Anastasis [00:20:44]: It was before Gen-1 was even generally available. But Gen-1 was a depth-to-video model, so it would take a depth map and it would convert it into RGB. and we couldn’t get, text or image-to-video to work directly, and that’s why we started from depth to video. but, and we had discussions of like, okay, we need to spend the next six months investing in text-to-video, maybe increasing the compute scale or the model scale, like train a larger model. And I had this weekend project idea, which was, what if I take a model that, starts from text input and converts to depth maps and then use Gen-1 to convert the depth maps Into RGB?
Vibhu [00:21:29]: It would probably work.
Anastasis [00:21:30]: And so Gen-2 was that.
Vibhu [00:21:32]: Oh. The hackathon pipeline.
Swyx [00:21:35]: The weekend hackathon pipeline.
Anastasis [00:21:36]: Yeah.
Vibhu [00:21:37]: But it looks good.
Anastasis [00:21:38]: And it worked pretty well. there were if you, with the knowledge that it has this, like, two-stage pipeline, you can tell in some cases that the structure of the video looks a bit off because you had to generate the depth first before you go into the output video. But it worked and it allowed us to bring this to our, to users very quickly. But it’s now it’s interesting because, like, people are coming back to this almost two-stage approach. Like, if you look at the Reve text-to-image model that came a few months ago, it had this planner model that would generate bounding boxes before it fed that into the diffusion transformer.
Swyx [00:22:19]: Yeah, Ideogram also the same day.
Anastasis [00:22:22]: Yeah.
Swyx [00:22:22]: I remember that was very strange that both of them came out the same day with the same exact innovation.
Anastasis [00:22:26]: It’s a small community, I think.
Swyx [00:22:28]: I’m like, this is like, this is completely coincidental, right?
Anastasis [00:22:32]: People talk. So yeah, there’s, there’s definitely something into this approach. And, now, like every single like, video generation model in production uses a complex prompt completion pipeline under the hood. I think that’s no secret that there is. That
Swyx [00:22:48]: Humans are terrible at prompting.
Prompt Rewriting, Camera Control, and the Seed of World Models
Vibhu [00:22:51]: I think across the board.
Anastasis [00:22:51]: Yes.
Vibhu [00:22:52]: But yeah, I think like the original Sora one blog post even told you that what happens after your input is rewriting your prompt. It’s much more descriptive about what you would want.
Anastasis [00:23:02]: Exactly. I, And there was the DALL-E 3 paper beforehand that, was the first public, description of the fact that synthetic captions and really detailed captions work really well. And then Sora built on that. Yeah, so it was 2023. We were releasing all these updates to Gen-2, like the camera control, Motion Brush. And there was something very interesting about camera control because it was the first time that you felt that instead of, like, you were creating video, you were creating a short video, you were navigating inside the world. And I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction. We realized, it was this era and this series of, Gen-1 and Gen-2 models really proved to ourselves, yeah, this is the
Swyx [00:23:56]: Cool.
Anastasis [00:23:57]: So this is not the original camera control. This was the updated camera control on top of Gen-3. But yeah, I think it made those models usable to filmmakers, I would say. The so camera control was very popular. And so we realized, there is one way of seeing those models, which is, you’re just as content creation machines, and there is the other way, which is you’re. As you’re predicting video in order to predict video well, you need to simulate the world in an increasing and increasing capacity. And if scaling laws apply on video, just like they apply on language models, then as we scale the compute that we put into those models, then they’re gonna be able to simulate physics, they’re gonna be able to simulate human actions and dynamics increasingly well and predictably well. That was the thesis about around our efforts on world models, and we spin up this research group to just focus on the world models and how do we turn the video generation models that we’re building into something broader and something that would be useful beyond, also content creation as well.
Swyx [00:25:04]: And that was roughly when?
Anastasis [00:25:06]: Yeah, so that was in
Swyx [00:25:06]: Oh
Anastasis [00:25:07]: In late 2023.
Vibhu [00:25:08]: Interesting. like, I think, a lot of people have been saying a lot of video gen model companies have all pivoted to world models these days, but like, 2023, you’re posting it. one
World Models: From Video Generation to Simulation
Swyx [00:25:21]: It’s, it’s debatable whether it’s a pivot.
Vibhu [00:25:23]: Yeah.
Swyx [00:25:23]: Like, arguably
Vibhu [00:25:24]: Yeah
Swyx [00:25:24]: That’s what you always had to do anyway, right?
Anastasis [00:25:26]: It’s in a way an expansion
Vibhu [00:25:28]: Yeah
Anastasis [00:25:28]: Of the applications
Vibhu [00:25:29]: Yeah
Anastasis [00:25:29]: Of the models as they become more capable.
Vibhu [00:25:31]: The early signs, it seems like the original models you guy had, guys had, people would say it’s very not bitter lesson pilled, right? You’re adding, rewriting prompts, you’re having all these one-off things, but that’s just the state of the tech as it was versus the future of as you said, you can scale it up as, we can scale up to world models.
Anastasis [00:25:50]: Yeah. So it just became. And if you looked at the outputs of Gen-2
Vibhu [00:25:56]: Yeah
Anastasis [00:25:56]: It was not. I think it was not obvious to people that this would scale to become a general simulator of the world. Like, you had very limited movement, you had, very low fidelity or low resolution, like obvious mistakes in human anatomy, like all kinds of limitations. But it was just, the idea was that’s just GPT-two, and GPT-two, it can barely generate, like, coherent sentences. Similar, Gen-2 can barely create coherent video, but if you scale it up, you’re gonna. There is no reason why it shouldn’t work in a way. It’s, And I think that was. That’s, that’s always the mindset of Runway is like this extrapolation of, like, if, like, even when we started in 2018 and you looked at the results of the day, you need to look more at the trend of, like, where we were in 2018 versus when we were at the, when the first GAN came out in twenty, four 2014 or twenty, fifteen. And, you started from, like, thirty-two by thirty-two images of faces, and then by the time in 2018, you could generate, street images at the 1K resolution. And it was the same with world models, very early signs of something much bigger.
Swyx [00:27:08]: Yeah. I was gonna say, like, it’s diffusing into focus. Like, if you look at our visible output from year to year, it looks like a diffusion process itself.
Anastasis [00:27:17]: Yeah.
Vibhu [00:27:17]: Especially watching the early, like, old blog posts, you can really see the choppiness, the details.
Anastasis [00:27:24]: Yeah. Like human civilization starting from random noise and then
Vibhu [00:27:27]: Yeah
Anastasis [00:27:27]: Denoising into
Swyx [00:27:28]: Yeah. Just run it a hundred years.
Anastasis [00:27:30]: Civilization.
Swyx [00:27:30]: Yeah.
Vibhu [00:27:31]: That’s how you’re on track, you’re still noising, right?
Swyx [00:27:34]: Yeah. I like the way that you guys phrased it when you, announced it in June, which is, oh, that you had a video essay. “The human mind is no longer the center of AI. Our world is.” Right? Which is, let’s, let’s call it the past five years of LLM-based AI is very much like trying to emulate human preferences and human speech. But now that’s, like, mostly solved. I think that’s, like, some of the context of your essay, which you also wrote around the time. And now it’s like the focus is on modeling the world accurately.
Scaling Laws for Video and Why Predicting Pixels Matters
Anastasis [00:28:03]: Exactly, yeah. So the way we see it is, there is that, initial mission statement of DeepMind, which is, solve intelligence and then use it to solve everything else. But I think it’s starting from everything else, could be valuable of, like, starting from. there is just so much complexity, and detail in the world that in order to. That it’s, it’s hard to learn directly from just human descriptions of the world. Like, we’re assuming that, like, language models learn from everything that humans have written about the world, like our own understanding as of, the twenty twenties. And there is just so much that we don’t know and so much that’s not captured by existing text, about both the low level dynamics of the world, like we’re not describing in detail. if I tell you to describe, like, how do you tie your shoes, that’s a very difficult thing to describe in words, but it’s very obvious thing to demonstrate. And so I think there’s been. And there’s, more of X paradox, like we’re constantly underestimating all the complexity that goes into very, like, things that we do subconsciously as humans, and we don’t even necessarily always have the words to describe them. And so in my mind, the simulating the world and simulating, physics, simulating the dynamics of the world has always been underestimated, compared to, we place too much emphasis on the things that are easy to talk about. but there is just all this complexity and richness of the world that if we just try and train directly on that observational data instead of training on how people describe the world, we would learn something new that we wouldn’t otherwise know.
Swyx [00:29:54]: You think that the present architectural paradigm is fine? You don’t need, like, another layer, like JEPA, like another famous, New York AI leader would say?
Anastasis [00:30:05]: We’re a very pragmatic research lab. If, we have evidence that an approach works better than the approach that we’re taking, then we have no qualms to taking it. We just have seen no indication that video prediction itself doesn’t scale. And even if you look now, not just our work, but the work of others, you’re seeing in robotics some of the most promising work, starts from video prediction models, and then you adapt them to also the action models, for example. so there is very little evidence that you need something else and that your time is better spent on a novel architectural change compared to improving data and improving the, and scaling the current approach. And so, We don’t have any indication that. the, there is that counterargument that I think there was a tweet by Yann LeCun a few days ago that, understanding the dynamics of the world is very different than, generating, cute videos.
Swyx [00:31:05]: And your answer is no, they’re the same thing.
Anastasis [00:31:07]: Yeah, they’re the same thing.
Swyx [00:31:08]: My cat videos are the same as understanding physics.
Anastasis [00:31:11]: Right, because if you wanna generate. video models can cheat and, like, they could you could give, like, successive dif shots of the scene in a way that doesn’t require you to simulate difficult physics. There is like, all these different ways in which you can hide the deficiencies of the model, and it’s important not to be too tricked by the performance of the current video models. It’s easy to, cherry-pick examples and think that video models are further advanced than they are. So there is a lot more work that we need to do to improve those models. But in my mind, very similar to language, and, like, we’ve. you go from barely coherent sentences to something that, could hold a conversation with a human to something that could can operate autonomously for a day and, like, create entire code bases. And the main difference, there is some architecture improvements along the way, but the main thing is scale. And so it’s the same bet for video, and we have no indications that this is saturating. Like, we have benchmarks that we use for measuring the physics of those models, and we see those predictably improve as we scale those models. So there is. If you want to Google up, Physics-IQ, is one of those benchmarks that measures how well does the model perform at solid mechanics or fluid dynamics or optics.
Vibhu [00:32:32]: I’m curious if you’ve seen any emergence, any scaling law around this.
Swyx [00:32:37]: Yeah, he’s saying there is a scaling law, right?
Anastasis [00:32:39]: Exactly.
Vibhu [00:32:40]: Yeah,
Anastasis [00:32:40]: So the way those models, those benchmarks work is you. the researchers have gone and, like, captured, a few videos that are representative of different physical phenomena, and then you can take the first frame and then pass it through an image-to-video model and then generate a rollout that shows what should happen next. So you have, a ball hanging from the ceiling, and then you use that as input, and then you the model predicts how the ball should fall on the ground. and this measures. we have an intuitive understanding of physics. I know, you can imagine what will happen next if I drop this bottle. So it’s measuring that same intuitive physics understanding of those models, and we’ve measured that at different model scales, and we see, and compute scales, and we see that the score on physics IQ predictably improves. There’s other, tricks and techniques that you can make to improve the score even further, but even scale alone helps, in the model learning better physics.
Swyx [00:33:40]: My main sympathy with Yann LeCun is the, Plato’s cave allegory, right? Like, you’re, you’re, like, learning on the output of a thing, not the internal process of a thing, and it’s very noisy. And, if only you could observe the internals of a thing. It’s hard to observe the internals of a human mind, but you can very much observe, or at least we have a whole branch of science and physics that we’re ignoring on how to model Physics and movement and, gravity and, other interactions. and we’re just, like, throwing away all of that and just saying just scale data, which is very much the lesson of unsupervised learning, but it feels wrong. that’s the main idea.
Anastasis [00:34:21]: I think the history of machine learning is, at large, it feels wrong.
Swyx [00:34:25]: Yeah. It’s a bitter lesson, right? Yeah. It’s, it’s, it’s the simple answer to that.
Vibhu [00:34:29]: I guess, how much can you scale? So, like, even on, let’s say, the video generation side, like, there’s one side of video understanding. Video generation, are we still gonna have tools where it’s like, I wanna generate two hours, twenty hours? there’s a infra way to do it in batches and stitch it together, but, like, do we just keep scaling? Do we just continue long generation consistency, all that at scale? And, like, tying it into where we’re at now from we looked at Runway two to four point five
Gen-3, Sora, and Runway’s Scaling Inflection
Anastasis [00:34:58]: Yeah.
Vibhu [00:34:58]: Like, technically, what advancements have we made to today, and then where do you see things still going?
Anastasis [00:35:04]: So part of the answer is definitely scale. and that was. We learned that lesson in a big way for with Gen-3. So Gen-3 was the model we released the year after, like in 2024. That was a few months after Sora was released. so yeah, there’s an interesting story of that came to be as well. Gen-3 for us was, the first time that we really needed to build. we had to learn all the lessons that the language model world learned in two in three years in the span of a few months. one of the biggest changes of Sora was using diffusion transformers instead of convnets. So a lot of the early, latent diffusion models were all, convnets for the diffusion model part. And the diffusion transformer paper came at some point in 2023, and it showed scaling laws for image, diffusion transformers. And we realized at that point that we needed to invest in infrastructure for model parallelism, for really scaling training to larger than, a few billion parameter models. And we spent maybe the, most of the fall of 2023 building out our infrastructure for distributed training. And we had a lot of false starts and a lot of failure in trying to scale, image and video diffusion transformers. And at that point, February 2024, Sora comes out, and the results are
Anastasis [00:36:35]: Very much superior to what Gen-2 could produce. There were a lot of, a lot of chatter on Twitter about Runway. Runway’s done. like, there is no way Runway will catch up. And if you remember, also OpenAI in the early twenty-It felt very, like it’s a
Swyx [00:36:56]: To the moon
Anastasis [00:36:57]: It’s a formidable opponent now, but at that point, it, they were on the top of their game. nobody could even get close to them. There was maybe Gemini was just the first version of Gemini had just released. So when OpenAI came with Sora and it was such a big jump of like quality, it gave me, there was like an existential crisis for a few hours. But that, I think the amazing thing about Runway and like I think the, we’ve been around eight years now, which is almost we’re dinosaur in AI, and we had to like, we had there was a lot of those moments we had to learn, adapt very quickly and build out skill set in the team that we didn’t have. And so, if you ask anyone what is their favorite time at Runway that was there during that time, it was that push in like three months to get to a model better than Sora. and it, we scaled 10x the model scale, the model size and the, compute that we were training on. we figured out model parallelism. We had zero expertise in that. And then we came out with Gen-3 during that summer. So that was a big turning point, I think, for the company where the research org grew very quickly, and we really started pursuing this vision of the general world model, in earnest, I think after Gen-3 was out.
Swyx [00:38:12]: Yeah. that’s the amazing thing about building when you’re building. There’s no stack to. You have to invent everything yourself. You have to be completely full stack. Now I think like there are inference specialists like Fal or whatever that can help with like, model serving, and I think you guys work with them as well. but yeah, like it’s, it. But at the time, it was just. It’s very interesting to think about what you do when Sora comes out and people are questioning whether your company should still exist.
Distillation, Turbo Models, and Real-Time Video
Anastasis [00:38:41]: Yeah. And yeah, there was no, there was no VLM of diffusion models. Like, we had to build the whole model serving infrastructure and make things efficient. And a few months after we released Gen-3, we released the Turbo version, which I think was the first step-distilled model in production.
Swyx [00:38:56]: That was a whole trend that we covered as well. Yeah.
Anastasis [00:38:59]: So that allowed us, to serve those models at the larger scale, ‘cause I think the first version of Gen-3 was quite, expensive to serve.
Swyx [00:39:09]: I think the whole like trend in like consistency models, Lightning and, Turbo and all these things somehow didn’t really stick around. I don’t know if you have any reflections on this. Because at the time, I was like, “Well, everything should start with a distilled model first, and then you can upscale,” right? It. your bigger models just turn into fancy upscalers, but like you should always draft with a smaller model and faster model, right? Because you can get it so quickly, like near real-time.
Anastasis [00:39:39]: Yeah. I would not be so sure to say that didn’t stick around. I think that, it’s, it’s likely to. that there is a lot of step-distilled models that are actively used in production. there is still a gap in quality compared to the, non-distilled model. but in my mind, we’re still. there is a two to three year offset from language models. So the things that, So it’s just a matter of time before there is better distillation techniques. we use. Right now we have a real-time model core character that I think is the largest deployment of real-time video models, that’s a step-distilled model, and it’s actively being used. It’s a very specific use case compared to a general video model. So this is a
Swyx [00:40:27]: Very cool, by the way.
Anastasis [00:40:27]: This is avatars stuff, right?
Swyx [00:40:28]: Consistency, character.
Anastasis [00:40:30]: Yeah. So this is a talking avatar, model. we were able to. we optimized the hell out of it, and it generates at 24 FPS, and it’s a, it’s a step-distilled autoregressive video model. So if we look at our world model direction, a big component of it is starting from the bidirectional diffusion that generates entire video at once and making autoregressive shows. So you generate one frame or a few frames at a time. so there’s a lot that goes into that pipeline of getting to a real-time model. It’s first you need to make it into a causal autoregressive model, and then you just turn it into. You need to do some additional step distillation to get it to be real-time. and I think that part is just starting. I’ll be very surprised if we’re, two years from now, we don’t primarily use real-time models. To me, real-time video generation is just inevitable that, it has much better user experience, it’s much cheaper to serve, and, the quality gap between the base model and the real-time model is only gonna close as we figure out better, distillation techniques. And we made a lot of progress there internally on maintaining the quality of the base model when we distill them.
Swyx [00:41:49]: How much of this is transferable? So is it the same base model? Like if you’re doing diffusion across the whole sequence and you’re converting it to step autoregressive distillation, is this like distillation where you still need to train both, you can use the same base and converter? What’s that process like to go from regular model to something that’s real-time on a technical level?
Anastasis [00:42:11]: So the nice thing about diffusion models is you have, two axes of distillation. So there is the. You can distill to a smaller model, which resembles what you do in LLMs, or you can distill in terms of taking less steps, less diffusion steps. So you could take a model that generates in fifty steps and generate in four steps and get to, You have some performance, degradation, but very often you get comparable outputs. So you can even take the large frontier model and distill it with step distillation and get to a real-time performance, and that’s what we’ve seen. So, depending on the use case, in some cases we might also serve with a smaller model, but in a lot of use cases, we just use the
Swyx [00:42:56]: Step distillation
Anastasis [00:42:56]: The frontier model, and we’re able to make it work in real-time.
Swyx [00:42:59]: I think this might be a good time to cut over to his laptop to show off some of the real-time stuff that you’re doing.
Interface World Models and Neural Software
Anastasis [00:43:06]: This is one of the research updates that we did recently. so we’ve been working and f in getting our general world models to, different applications. one of them that we think is very compelling is using general world models as essentially, an interface, a universal interface to software. This is a version of our world model that’s called an interface world model. and the idea is that it essentially, replaces, the, front end of a software application. It renders the pixels directly of an interface and is trained to predict what happens next as a result of, a click or another interaction you have with the interface. So this is all pixels. it’s there is no HTML, CSS, React that’s powering this interface. This is directly at the output of our real-time, video generation model, and it takes clicks directly as input.
Swyx [00:44:09]: And drags, click and drag.
Anastasis [00:44:12]: Right. So it supports
Swyx [00:44:13]: Ooh.
Anastasis [00:44:14]: Yeah, clicks. It supports drags. it also supports scrolling. and the amazing thing about this is that you can effectively describe in the prompt how you want different elements, like what do you want the behavior of different elements to be. So it’s almost you’re you can turn, an interface from, markup language description of, like, an HTML interface, and instead you can just describe the interface. if I press this button, I expect this to happen. If I press this button, this should happen. And it’s useful, we believe, both for prototyping, for, like, just testing, like, what different interactions would feel like. you can also add audio to it. So it’s a video audio generation model. So you get you essentially can describe both what the visual outcome should be of your click and also what the if there is a sound effect that comes out of it. So we believe that’s gonna be a much more flexible way of building software. Just render. It just, in why generate the code that generates the pixels? Just generate the pixels directly.
Anastasis [00:45:18]: It’s the end-to-end philosophy applying applied to front ends.
Anastasis [00:45:25]: So we think there is a few interesting use case. So you can build creative tools on top of it.
Anastasis [00:45:32]: We think that, for any use case that involves a lot of exploration or, like, educational use case where you wanna learn about a new concept and you want some visualization and like, and open-ended exploration, we think those this is a very powerful, approach. you can imagine new forms of, design, industrial design software that could emerge as a result of those models. And this is all, generated in real-time as well. So, you can build a lot of interesting camera transitions and forms of interaction that are very difficult to build otherwise. And one way in which we evaluate this is what if you try to generate the same interface with Claude by just, prompting Claude, “Here’s an image reference of my interface that I made in Figma or that I created somewhere else. create this particular interaction,” which in this case it’s, drag that object, upwards. and beyond it being slower, it’s also very difficult to capture some interactions by just fully, with just LLMs. So we think that this is likely to be the way that a lot of the future, like, software in the future will be created. and one of the additional benefits is personalization might be a lot easier done with those models. Like, you can essentially try out different prompts based on who is visiting the interface. You can, more easily, prompt engineer the interface to have larger size, text for more accessibility reasons, or you can make this or, like, if you have a particular aesthetic preferences. So we’re very excited about this approach. It’s early days, and I think we’ll need to, make it more cost-effective as well to serve those models ‘cause, running a real-time video model versus just purely rendering HTML, there’s -- the computational needs are much higher. but we do see a lot of potential in this approach to building front-end interfaces.
Swyx [00:47:47]: So we covered this similar thing with Flipbook before with our, Ethan Hara episode with Groq, video. And yeah, I think it’s very engaging visually. I think it’s maybe very good for education, but it’s it does sound expensive. I think there’s an upper bound to how expensive it will be, though, right? Like, the inference cost will go down over time. You’ll figure out ways to optimize it. Effectively, when it pauses, you don’t you’re not receiving human input. You don’t have to generate anything, right? So.
Anastasis [00:48:14]: Yeah, you could also. Like, in this case, you have ambient motion, so there is parts of the screen that might. if you’re let’s say you wanna, visit Paris and then you get this interface that allows you to explore.
Swyx [00:48:29]: People walking. Yeah.
Anastasis [00:48:29]: You have people walking or, like, things happening. But, it’s, it’s a no Yeah, it makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism so you don’t need to do that. But all those things, I think, is stuff we’ll need to figure out.
Toward a Fully Neural Operating System
Swyx [00:48:44]: Yeah.
Anastasis [00:48:44]: I think our first consideration is let’s make this clearly find some use cases where it’s clearly a much more compelling interaction compared to traditional interfaces. And then it’s a matter of time before it becomes more cost-effective to serve.
Swyx [00:48:58]: Yeah. When it comes to the people walking, I think the approach that makes the most sense to me is Nick.
Anastasis [00:49:04]: Nick.
Swyx [00:49:04]: Oh, God. I keep messing up their name. With Chris Manning and Fanny Yan. I don’t know if you’ve come across them, where they. Mapped to some game engine. I think it’s Unity or something, or Godot. And they you can script some NPC behavior behind that and train on that. Whereas here, you can really imagine whatever you want. Like, that is a UI, right? Like, and it feels, like, more tractable, I guess, to, create a world model of software that is interactable because we have many of examples of that, and you can, do your fancy RL environment stuff on that than it is scaling up to embodied and real-world physical use cases. But this is a nice first step.
Vibhu [00:49:43]: Or, there’s the opposite of you have, like, one B models, three 50 million parameter language models. It just gets so small that they’re just predicting, like, fishes moving.
Swyx [00:49:53]: Small models are now 120 B, so.
Vibhu [00:49:57]: Ultra mini on device.
Vibhu [00:49:58]: But, no, I think it, like, it puts it into perspective, at least the car one for me, like, the applications, right? The amount of work to do that, sure, you only make one model year car per year, but applying this, it’s also a cost-saving to have to manually make all this, right? So it opens up a lot of possibilities, too. I’m curious if you extend this out two, three years, so where do you see things going even further?
Anastasis [00:50:25]: Effectively, the end game of something like interface world models is you have, a fully neural operating system. So I think, Andrej Karpathy has written about that quite a while back. But it’s, You, I think to me it’s, it’s a bit, it’s a bit odd that, we have, for example, with an interaction with an LLM of today, you have this LLM that can talk to you about anything. It can You can take the conversation in any direction. You can It’s very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. And so to me, it’s just a matter of time before the interface itself becomes learnable and becomes, part of the whole loop of, like, you’re not just delivering. You’re delivering an application end-to-end, and that means you’re delivering the language model, but you’re also delivering the render and the pixels and that’s also a learnable component. And the concept of applications might not necessarily. I think we’ll need to figure out new abstractions for software. the concept of application comes from this idea that you need, separate code bases to describe, to, for, to power each individual, tool and each individual application. But you might think of something a lot more unified if you’re. if you have, a video model that’s generating the interface as you go. so it can take context from an LLM and allow you to combine different functionalities that traditionally would live in different applications. So it’s a, it’s a way to solve, software end-to-end, effectively. We also see this as a powerful way to train computer use agents as well. so this is, one way to see this as. And in general, with world models, there is those two directions. One is world models for humans and world models for
Swyx [00:52:24]: Agents
Anastasis [00:52:24]: To train agents.
Swyx [00:52:25]: Yeah.
Anastasis [00:52:25]: And so for every new work of, world models that we do, we have this both uses become possible. So this is a powerful synthetic data generator for training computer use models. It could become, a live, RL environment that you could use to do online RL with a computer use agent, and you can get wide diversity of different interactions, kinds of interfaces, just generated on the fly that, to improve the how robust the, your agent, becomes. So that’s the same also with the world models that we’re working on for a robotics use case as well.
Long Context, Error Accumulation, and Autoregressive Video
Swyx [00:53:02]: Is there a research breakthrough that you’re Waiting for that would unlock the next set of use cases that you really wanna pursue?
Anastasis [00:53:10]: Long context is a very important one, so being able to maintain consistency for long periods of time, and that depends on the use case. So for our characters model, for example, or for the interface world model, it’s easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and you take arbitrary actions in, we, like, there is more the context at which you can and duration which you can generate becomes limited much more quickly.
Swyx [00:53:40]: Yeah.
Anastasis [00:53:40]: So we see more degradation and error accumulation happening. so the biggest challenge with autoregressive models is error accumulation, is you’re feeding generative frames back into the model to generate the next The next frames. And if there is any small errors, they accumulate over time. That’s not a new problem. It’s a problem that LLMs also have, and we’ve seen the ability to generate now really long outputs. So it’s a solved problem, but it’s definitely still a challenge.
Swyx [00:54:08]: Yeah. And what is the state of the art? so for Grok, it would be like 10 to 20 seconds of context going in there for video.
Anastasis [00:54:16]: With our characters models, we’re able to generate up to 30 minutes of video autoregressively.
Swyx [00:54:21]: Yeah. But that’s just for the avatars.
Anastasis [00:54:24]: Yeah. So if we look at, GWM Worlds, which is more our open-ended world exploration model, it’s, it’s on the order of a few minutes, which is Yeah, so
Swyx [00:54:35]: Probably enough for people because you have to cut to the next scene anyway, right?
Anastasis [00:54:40]: Yeah, it’s not, it’s not the ideal game experience if you have to restart every few minutes. So I think. But, I think it’s. Yeah, for certain kinds of game experiences, you can work around it. ideally, you are able to just generate forever, and it doesn’t, it doesn’t degrade. And I think that’s a matter of time before we get there.
Swyx [00:54:59]: Yeah. Genie has, like, one, max one minute?
Anastasis [00:55:01]: Right. Yeah.
Vibhu [00:55:02]: This was your. You did a study on robotics. I think I also have just your Runway Robotics page, though. Is this better?
GWM Robotics and Sim-to-Real Evaluation
Anastasis [00:55:11]: So last year we released Gen-4.5, so that was our latest base model. We’ve been As I mentioned, we’ve been doing all this work in world models, and which essentially a lot of our approach to world models is how do you take a bidirectional diffusion model and make it autoregressive and make it accept actions? So instead of being a video you watch, it becomes a simulation that you step in, and you can, control it every step of the way. You can explore counterfactuals, like what happens if I take this action versus if I take this action. And GWM-1 was the it’s the world model that we built on top of Gen-4.5. So we did all this autoregressive and like, distillation, auto-regressive and then step distillation on top of Gen-4.5. And one of the biggest use case that we saw for GWM-1 was in robotics. One thing we like to say is we as we scaled video models, we accidentally, created one a state-of-the-art model for robotics, by just scaling video models. So we realized at some point, mid last year that robotics labs that are coming up to us and asking to use video models for synthetic data, asking us to post-train our video models to work really well for robotics, so that they can use that to generate variations. That was the first use case that we saw. And then increasingly became clear that the models will be useful beyond just creating synthetic data to train robotic policies. They would also be very useful as simulators. So that means that you can use, a video model online to test how your robotic action model performs. So you can take an action role and then get the outcome of the action inside the world model and then continue that loop like this closed loop simulation. And you can use that to evaluate how well your robotics model works. and the biggest thing that I think you need to solve if you want to build a simulator is establishing real-world correlation that if you take an action inside the world model, if you take the same action in the real-world, you get a similar outcome. So that was the goal of some work that we did earlier this year. So if you go to the first link. So that was, essentially wanted to establish that, real to sim correlation for our world model, so that if you do a series of actions inside the world model and if you do the same actions in the real-world, you get similar outcomes. And we took our GWM-1 model and we used some benchmark data that there is this Roborina, benchmark that’s very commonly used to evaluate how well do different action models perform. And we use the same scenarios and settings and embodiments inside our world model, and we measure the correlation of how well did the action model perform inside the world model versus in the real-world. And we saw that we could get very good correlation between our world model and reality. And that means that if you want to evaluate how well your robotic policies perform, you can scale that much faster inside simulation instead of having to do that with actual physical hardware. And so that was a first indication that our models could be, quite useful in robotics. And we saw as we were working with robotics labs that became like the first use case where they could use video models in a way that feed into their training pipeline.
Vibhu [00:58:40]: Can I ask what
Anastasis [00:58:41]: Yeah
Vibhu [00:58:41]: The difference was from four point five to solving that? So the sim to real gap has always been the issue, right? You train a robotics model on video data, it doesn’t generalize to real-world, and the simulation had an issue. So seems like you solved it, but how?
Anastasis [00:58:56]: Yeah. So a big problem with simulators is, if you’re trying to simulate rigid objects, like it works quite well if you can describe the physics of objects very accurately, then you’re able to use, Isaac Sim or MuJoCo or one of the traditional simulators. But for more complex interactions with cloth, for example, or, like slippery surfaces, with the all the complexity that you want to be able to solve with the manipulation, with an action model that solves manipulation tasks, it’s very difficult and so time-consuming to build, for each of those environments and each of those tasks, build the simulated version of that, the digital twin of that environment. Whereas with a world model, you just need to provide the first frame and then you just can roll out the policy inside the first frame. So whereas, we compare it to methods that required like 3D scanning an environment and then 3D scanning each individual object before you can now, you can bring that to simulation. whereas with a world model, you just take a picture of the environment and then you’re able to test how your policy performs. Our general thesis on robotics is, there is companies that are leveraging a lot of teleoperation data to train robotics action models. There is now companies that are using, humie data, which is, essentially human, egocentric video where humans use robotic creepers to perform different manipulation tasks. And then there is companies that are focusing on egocentric data, which is, you strap a GoPro on someone’s head and then you capture them performing a task. We think that, and all those are great source of data for training robotics models, but the most plentiful source of video data is third-person video data. It’s And if How do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don’t learn from first person. We do some trial and error and like, to learn different things, but. Ultimately, a lot of what we learn how to do in the world, we learn by watching other people do it. And that’s how when you’re pre-training a video model, you’re essentially doing that. It’s a lot of third-person video footage of people performing different tasks in the world, people doing sports, people doing household tasks. And our main thesis is that video pre-training, once you do that, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data. So you require way less teleoperation data, which is very difficult to scale. and even if you look at egocentric data, which is a bit more easy to scale compared to teleoperation data, which requires actual hardware,
Why Third-Person Video Is a Powerful Robotics Pretraining Source
Anastasis [01:01:55]: It’s still three hours of magnitude less of that exists in the world compared to third-person video data out there. And so our thesis is and generally, like the most plentiful source of data will ultimately wins. Third-person video data pre-training is the right starting point for models that, you want them to generalize and be able to deal with new environments, new tasks, things that you haven’t seen during training. That’s the motivation for why we think our models are especially useful in robotics, settings, and we’ve seen that to be the case, as well.
Swyx [01:02:32]: You said pre-training. So maybe it’s like third-person pre-training, first-person SFT? Is there like a curriculum that you can introduce?
Anastasis [01:02:41]: Exactly. So if we look at GWM Worlds, so GW so GWM Robotics. So digitally in robotics, it starts from Gen-4.5.
Vibhu [01:02:49]: It’s the same video diffusion backbone, right?
Post-Training World Models for Robotics Embodiments
Anastasis [01:02:53]: Exactly, yeah. So you start from the base video model, the one you’re using to generate, cats and dogs and other interesting stuff, and then you, fine-tune on a very small number of hours of robotic data. So it’s something on the order of hundreds of hours compared to if you were to pre-train a robotics model. The current pre-trainings go up to, a hundred thousand or like millions of hours of data. And you’re able to get quite good performance, quickly, because the model leverages all the things that it has learned about the world, physics and human dynamics and the tasks that people care about from pre-training. And ultimately, you want those models to generalize. You don’t want to just be able to perform the tasks that it has been doing training. And the diversity of actions and environments that you have with a pre-training video dataset is much larger than, what you can realistically capture manually.
Vibhu [01:03:54]: How is the scale looking like for the post-training? Like, do you still wanna do, is it like roughly ninety percent of the compute in regular video diffusion model and then scale up a lot, or do it like we want different robotic models for different tasks, or just the one base really good world model can also apply to robotics?
Anastasis [01:04:14]: So currently, we are post-training our models for specific, embodiments that we for particular partners. So if they have a particular single-arm robot or a bimanual robot or a humanoid robot, we would post-train our GWM robotics model on their particular dataset. Over time, we see the different variants of GWM unifying. Like, I would expect, if a year from now or two years from now, you have a single world model that can simulate manipulation tasks, it can simulate navigation, which is a lot of the gaming world models are navigational world models. You’re moving around the space, and it will also simulate human behavior. So that’s the character models. So instead of having three different models, you have a single model that’s able to. ideally, you’re able to simulate what it’s like to be in the world. You’re moving around an environment. You’re maybe performing different tasks. you’re talking to other people. And that happens with, the same, a single real-time video model that’s generating that.
Vibhu [01:05:17]: Do you think you can solve self-driving? So if you are learning to drive a car in a simulator, you have a world model. Your robot is car can manipulate so many axes. How far off are you from something like that?
World Action Models, Self-Driving, and Learned Policies
Anastasis [01:05:31]: So world models
Vibhu [01:05:32]: Or a really good ADAS system?
Anastasis [01:05:33]: World models are definitely being applied to, self-driving, research right now, mainly for evaluation use cases, but our focus has been more on robotic manipulation. We’ve done some work on AV, world models as well. but yeah, we do think that world models are and video models are the best starting point for both simulators and also policy and the action models. So that’s, that’s the other side to this, is that once you have a great world model, then you can just add an action head, and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and you ask, you prompt the model, generate the arm picking up an object, it would And if it generates an accurate enough video, then it should also be able to generate the exact poses, in 3D that the arm should take to perform the same action. So this is the direction that’s now the popular term for it is world action models, which is you’re starting from a video model, and then you’re adding an action head to predict the actions, and it becomes a policy, essentially.
Swyx [01:06:43]: One thing I’m also impressed by is how much data you need to train these kinds of models. You probably can’t say exactly how much, but like, the original, diffusion models, and from what I know, even of the open source Chinese models, it’s not that much data. Isn’t it surprising?
Anastasis [01:07:02]: What do you define as much data?
Swyx [01:07:05]: Yeah, and it just comes, goes in. Is the token count still relevant?
Anastasis [01:07:09]: So it’s a bit more complicated and,
Swyx [01:07:11]: What is just gigabytes, right?
Anastasis [01:07:13]: Yeah, hours of video, right?
Swyx [01:07:15]: Yeah. Yeah. I feel like something that’s interesting is it seems like the, let’s call it tokens to param counts in language models has really, maybe they’re three years ahead or whatever, seems to be a lot higher than, video models still, even though technically video has more information, per bit. I don’t know if it seems intuitive or maybe there’s just a lot of, like the variability between a pixel to the next pixel is not that high. So, like, maybe there’s just a lot of information that is repeated.
Scaling Video Data and the Lucid Dream Test
Anastasis [01:07:47]: My answer would be it’s still very early. Like, the training video models will scale way further than it
Swyx [01:07:55]: Yeah
Anastasis [01:07:55]: Currently is, and you’ll have capabilities that go much further than the current models can do. So one thought experiment that, I like to use, it’s, it’s almost like the Turing test of video models or like the Turing test of world models, go, I call it the lucid dream test. It’s you have a
Swyx [01:08:14]: You mean the actual person lucid dream?
Anastasis [01:08:17]: It comes from this idea
Swyx [01:08:18]: Lucid rains, right?
Vibhu [01:08:19]: Lucid dreams is telling you’re dreaming while you’re
Swyx [01:08:22]: Yeah.
Anastasis [01:08:23]: Yeah, exactly. So lucid dreaming is when you realize you’re
Swyx [01:08:25]: In a dream
Anastasis [01:08:26]: Inside a dream, and then you
Vibhu [01:08:28]: Play around
Anastasis [01:08:28]: Be able to control what happens in
Swyx [01:08:30]: No, there’s also an inference guy called Lucid Rains. Yeah. Or quantization
Anastasis [01:08:33]: Very prolific, person. Yeah. So let’s say you have a VR headset and you’re in a room with and you’re wearing a VR headset, and that VR headset, most of today’s VR headsets have a pass-through mode, so you can see directly what’s in front of you in the world, or you can render something inside the VR headset. And there’s gonna be a point where those interactive real-time video models become good enough where you wear the headset and you’re in the same room and you’re walking around and you’re kinda and you’re interacting with objects. You’re able to move freely in that room and do, and interact with any object. And at the end, someone asks you, “Did you were you using pass-through mode, or were you -- or was this, rendered or generated, footage?” And if you cannot tell for sure if that was what you were seeing as you were interacting with and moving around the world was generated or it was, pass-through mode and was just what was happening in front of you, that’s an indication that the models have become good enough. And we’re not, we’re not close to that yet. And a lot of it is just this idea of really simulating dynamics and counterfactuals well. Like, if you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there is a bias from the training distribution. There is a lot more videos of the person succeeding at scoring the goal. But if you have an interactive model, you want it to be able to generate counterfactuals. Like, if I take this action versus this action, you want it to generate equally realistic outcomes. so that’s, I think, the big gap between video models and world models is that idea of the counterfactual generation. And if you want a great model for robotics, you wanna simulate failure very well, because whether you’re using it for evaluation or you’re using it as a in an online RL loop in the future, you wanna be able to have the model try and fail to do things and improve. and so in order to do that, you need to be able to simulate things failing.
Swyx [01:10:42]: This is the only domain where you have too many successful examples and not enough bad examples. Should be easy to generate failure.
Vibhu [01:10:51]: Oddly enough, I think, like, early image video models weren’t good at being human realistic, right? Like, you see aa lot of the high-res 4K, like, professional photography, but not just everyday life, like normal picture, right? Everything looks like it’s professionally generated, like professional pictures, but not just like normal, like, messy cables on a desk.
Swyx [01:11:13]: Okay, so there’s, there’s this stuff. one thing we also covered that you guys have, video agents that you launched. I guess, how does the traditional, let’s call it frontier, like, autoregressive LLMs, like, feed in, to all this? They’re driving ro your robotics models, or are they driving others, your video agents, production, anything where you see the overlap of autoregressive and diffusion, let’s call it?
Counterfactuals, Failure Data, and World Model Evaluation
Anastasis [01:11:41]: Yeah, so harnesses are really important across all those different use cases. So we have this video agent, which is essentially an LLM that is very effective at tool use of different, image models, video models, and helps you through creating a project end-to-end. So, very often in, like, a traditional advertising flow, you have a brief, you start from it, and then you generate some a storyboard, and then you generate the video. A video agent and, or runway agent helps you through that whole process, and it helps you also analyze performance data. For example, how well did this ad perform versus this ad, and then generate me more of the based on those learnings, figure out what to generate. We think that the harness is a very important piece of the pipeline. as I mentioned, all the video production, all the production video models use some prompt completion that happens, and we expect, that to become more and more complex and more, you generate longer and more detailed descriptions before you use the diffusion transformer. I do think eventually, there’s increasingly this unification into omni models where you have the you’re training the models end-to-end to both do autoregressive text prediction and also, diffusion as well. So you’re predicting the next token, of like you’re, you’re maybe using some reasoning and planning of the scene, and then you’re passing it into the diffusion head that’s generating the pixels.
Video Agents, Harnesses, and Omni Models
Swyx [01:13:10]: Yeah. I think currently maybe only Gemini and Qwen do it. I-I’m not sure which of the Chinese models are omni, but yeah, it’s, it’s not, it’s not a very well, popularized modality, I guess.
Vibhu [01:13:25]: It’s an interesting use case when you think about it, right? Because not only do you have to end at like language model reason, diffusion had generate, you don’t have to output there. You can go back in to feed that output to the same model, reason again on improvements, and it can do a lot of loops just in its own. I guess the question is like, do we need that or can we just do agent scaffold, like do it outside the model? Is there a big benefit to doing it in?
Anastasis [01:13:54]: I think there’s generally the trend of something is first done by a harness and then it becomes part of the model, right? So you had the chain of thought prompting where you had to do this super detailed system prompts to
Swyx [01:14:07]: Yeah, step by step
Anastasis [01:14:08]: Get the output. And now the model generates the reasoning trace by itself before it gives you an answer. And in the, in video models similarly, a lot of the video models of the early days were single-shot video models, and you had to use some orchestrator to turn, generate multiple shots in parallel, and then turn it into an actual video.
Swyx [01:14:28]: Or in ComfyUI, just all over the, all these nodes.
Anastasis [01:14:31]: Yeah, like a spaghetti workflow. and now you have multi-shot video generation where you have the you directly generate multiple shots. And there is a benefit to that because then the video model learns some. to generate a single shot well, you need to figure out a lot of stuff about the world. to generate multi-shot video well, you also need to get some, like, video editing instincts. Like, you need to figure out what is the right pacing of shots. And also, LLMs are not that good at it. Like, they’re not that great video editors. If you ask a LLM to take some videos and then auto-create a edited video out of that, it would feel uncanny. So I don’t think LLMs are that good yet at being video editors. And I think there’s benefit to learning that end-to-end. so I would expect, the training generally is the things that, you need the harness for eventually get injected into the model itself, and you learn that end-to-end.
From Harnesses to End-to-End Learned Video Editing
Swyx [01:15:33]: Do you find that you need to hire engineers who can. or researchers who are also artists to infuse that taste, or do you have artists in residence to distill them?
Anastasis [01:15:44]: We have a large creative team that’s very actively involved in the, in training those models, like on the, in every part of the way. And like, how do you caption video as well so that you capture the stuff that you need for, like, the cinematography, the aesthetics, the camera direction in as detailed ways as possible so that you’re able at inference time to elicit that through the model? we have our creative team also does a lot of evaluation of like, what constitutes a usable video out of those models. And so they’re very involved through every part of the process. And I think that’s one of the special things of Runway is just that mix between like creatives and researchers sitting by, side by side and working together to build the next generation of our models. I think that’s been a really important piece to, how we’ve operated as a company.
Swyx [01:16:37]: Yeah. In some senses, though, you can only do this in New York.
Vibhu [01:16:40]: It’s
Swyx [01:16:40]: Maybe, you have other offices, but like, I try to find some poetic, significance in the fact that you are a big New York company.
Anastasis [01:16:49]: As there’s a few parts to being New York. there is that intersection of all those different industries and, like, media, advertising, like
Swyx [01:16:57]: Yeah, this is very advertising.
Anastasis [01:16:59]: The, like the art scene is New York. Not to say anything bad about San Francisco, but, it’s. There is more going on. There is that component, and there’s also, I think we benefit from being outsiders and thinking of things a bit differently, like not being in the same, like, hive mind of,
Swyx [01:17:19]: BВС
Anastasis [01:17:19]: ASI, of Bay Area and, like, taking. and also taking our time to get where we are today. Like, building the, growing the team intentionally and bringing people who are, yeah, both on the creative side and also on the engineering research side. There’s huge talent pool of amazing people in New York, so that hasn’t really been a problem.
Creative Taste, Artist Feedback, and Runway’s New York Advantage
Swyx [01:17:41]: Congrats on everything. what are you hiring for? what should people look forward to, for the future of Runway?
Anastasis [01:17:49]: We’re hiring across the board. I think this is probably the most open roles we’ve ever had in the history of Runway. we’re growing our research team quite significantly. So if you’re, if you’re excited about video models, if you’re excited about world models, if you’re excited especially about robotics, the robotics team, we’re hiring roles in the robotics across, software, hardware, and research. so definitely reach out.
Swyx [01:18:13]: And, a lot of people don’t have direct robotics background, but what should they have, if they want to be useful in robotics?
Anastasis [01:18:21]: So ideally, some experience with learned policies, would be
Swyx [01:18:26]: Just RLs
Anastasis [01:18:27]: Good for robotics. but we tend to hire generalists as a philosophy and, like, people who learn really quickly. but some experience in the, in domain expertise in robotics is something that we’re, we’re definitely looking for the next months. and then we’re scaling the go-to-market team significantly. There is, a wide, like, very active enterprise adoption happening around video models at the moment, and, we’re really trying to, respond to all the demand.
Swyx [01:19:00]: Yeah. Great. You wanna talk about the, open source robotics stuff?
Vibhu [01:19:04]: Sure. It was just random notes we had.
Vibhu [01:19:07]: NVIDIA launched Cosmo. I guess it’s interesting. So, you’re a founding member AI labs to build open source world models in physical AI. - Anything else to talk on here is open research?
Anastasis [01:19:20]: The biggest thing is that, as I mentioned, while models are still, early, like there is still so much that we you can scale and those models further, so much more advancements and things that we can figure out and how to improve those models further. And I think this is, it’s important that some of this research happens in the open and figuring out what is some incentives for different companies to come together to bring some of that research into the open and open source. And so Cosmos Coalition was a initiative that we co-founded with NVIDIA to bring some of that research as open source. And that could mean open weight model releases. It could mean benchmarks that measure physics and things that people care about when building world models. It could mean infrastructure. So really, how do we grow the ecosystem of world models and make that something that also it’s easier for a developer, a researcher that’s just starting out that is excited about world models to contribute to the field.
Hiring, Robotics, and Enterprise Adoption
Swyx [01:20:19]: I think it’s a there’s some amount of like, is this also our response against the Chinese world models that are being released, or is there not part of the consideration?
Anastasis [01:20:29]: I do think it’s, it’s important for NVIDIA models, if you look at the leaderboards of video models, I would say right now the majority of models at the top ten, top twenty are Chinese models. There is, only a handful of companies that are made it to the leaderboard from like the US or the West.
Swyx [01:20:50]: Yeah. We’re doing better with images, but with video we’re very behind, right?
Anastasis [01:20:53]: And so I think it’s definitely important that we invest more broadly as a community to make sure that we can those models can we have competitive models
Swyx [01:21:02]: Yeah
Anastasis [01:21:02]: Out there.
Swyx [01:21:03]: But like what’s to stop us from just distilling from them?
Anastasis [01:21:06]: I don’t know if that’s the best long-term
Swyx [01:21:08]: Not gonna mention that they won’t
Anastasis [01:21:09]: That you’re bounded by the performance that you can. It’s, it’s almost a bit of a pessimistic
Cosmos Coalition and Open World Model Research
Swyx [01:21:14]: Like
Anastasis [01:21:14]: View that you can get better. you can
Swyx [01:21:17]: It’s free data. it’s, you might as well. Like if they’re, they’re doing it for like, for the text language side, they might as well do it for the video side the other way.
Anastasis [01:21:25]: Yeah, I do think we’re, we’re quite capable of training great models
Swyx [01:21:29]: Okay
Anastasis [01:21:30]: Without distillation at the moment. Yeah.
Swyx [01:21:32]: Yeah.
Vibhu [01:21:32]: So anything you have to say on benchmarks and evals? Like, I feel like what I’m hearing is a lot of people really like arenas for video and image models, customers and whatnot as well. They only want the best on the leaderboard, and they refer to arenas a lot more than language models seem to do. But any notes on benchmarks, what’s lacking? How does the average person compare while these both look really hyper-realistic? More than that, outside of we did talk about like robotic simulation, the physics and all that, but anything to say?
Anastasis [01:22:05]: I think it’s the opposite in some ways. I think people, generally creatives and artists and marketers, other like people that are using our platforms, I think rely less on, arena scores. And it’s, it’s just so easy to, generate with a bunch of different models and then compare the results visually. Like one nice thing about image and video models is you can immediately tell with your eyes like what feels good from an aesthetic standpoint. Like any artifacts, any issues with the physics of those models, you can immediately tell. and so that’s it’s easier, I would say, to evaluate, as a human. there is also those models than it is in language models where you have those very complex math and coding and, tests where it becomes a lot more harder, I think, for humans to evaluate and can discriminate between the performance of models at a time. So I think in practice, people just test out the same prompt with a bunch of different models and see what the results look like. And right now in Runway, you can use our models and you can use third-party models as well. So it’s, it’s very easy to do that.
Benchmarks, Arenas, and How Creatives Evaluate Models
Swyx [01:23:12]: Amazing. We’re gonna end with the AI Runway AI Summit. The last societal issue, I guess, I don’t know if this is a thing, is the, you are at the tension between artists and creatives and AI. A lot of people in that community hate AI. the people that are in the Runway community don’t mind using tools. it’s just another brush. But, how have you seen the sentiment change?
Anastasis [01:23:37]: Our perspective, yes, it’s just another branch, brush. It’s just another camera. It’s, it’s the latest of a long generation of tools.
Swyx [01:23:46]: Technology in art.
Anastasis [01:23:47]: Technology.
Swyx [01:23:47]: Yeah.
Anastasis [01:23:47]: And art and technology have evolved together. I think there’s been a pretty significant shift over the past few months, and it came. some of it you can see with a lot of public figures speaking out in favor of AI and being, like in Cannes, you saw a few directors speaking in favor of AI. We had Ron Howard in our film festival. There is, Mark Scorsese also adopting AI models. So you have more of those stories coming out every day of like a well-known figure, speaking in favor of AI. And it’s just a matter of, in my mind, it’s those models are becoming more and more demystified. I would say I have also a bit of a hot take that one of the things that made the initial response to those models maybe a bit more heated than it needed to be was this idea of text to video of, you have a single text description and you get back a -full video.
Artists, AI, and the Evolution of Creative Workflows
Anastasis [01:24:48]: Yeah, there was a misconception. you can generate it to our feature-length film, but the models of today now take a lot of references. They take they are very controllable. And I think when people see a tool that allows, affords many degrees of freedom and control, they respond to it differently. And it matters less that it’s a generative model than the fact that you can steer it to the direction that you want. and so. I think when people look at, complex workflows on top of those models, when they look at, all the ways in which you can steer them and you can provide now with some of the latest models up to fifty references, like the conversation becomes a bit different because it feels much more like a
Swyx [01:25:34]: Storyboard
Anastasis [01:25:35]: A tool
Swyx [01:25:35]: Yeah
Anastasis [01:25:35]: Versus, like, something that a magical entity that figures out, like, the, your entire film for you.
Vibhu [01:25:43]: Any notes on, like, workflows changing for people in the field? Like, I think engineering at least has had a lot of people where they’re like expectations have changed. I’m, ten X, a hundred X more productive, and you can get a lot more done. same thing as, you’re making dev tools for creatives. any notes there? Like, there’s some people that don’t wanna adopt, some that do. Like, anything?
Anastasis [01:26:08]: Yeah. So I think, in terms of, like, what people care about, I see that we have gone through a few stages. So we started from a stage where the main thing that people were looking for was quality. Like, as, we scale those models, the quality improved dramatically. That’s something that people still care about, but it’s, it’s now in addition to controllability, like being able to steer those models with references, with, different kinds of inputs, with storyboards. And now my sense is increasingly people are gonna care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important. And, like, if you can, with a single prompt generate ten different, outputs, like, almost instantly, you can explore way faster than before. And you get some of the magic that characterized the creative tools of the past, like Photoshop was instant. and we lost some of that with generative models. You’re waiting for two minutes to get back a video, and I think we’re gonna bring, some of that back now with the
Vibhu [01:27:08]: Real-time
Anastasis [01:27:08]: Real-time models.
Vibhu [01:27:09]: Yeah. Exciting. And
Latency, Real-Time Generation, and the Future of Creative Tools
Swyx [01:27:11]: Exciting. the last thing we’ll plug is this one, Runway
Vibhu [01:27:14]: Summit
Swyx [01:27:14]: Summit. You’re finally doing this in SF?
Anastasis [01:27:18]: Yeah. So, we’re very excited about this. So this is, in late September thirtieth, we’re doing a summit on, primarily focused on physically high and real-time video generation. We have panelists from NVIDIA, Physical Intelligence, Botco, DeepMind. Yeah, it’s gonna be, I think, a very interesting series of conversations. We try to make the panels really technical and, elicit actual substantive discussion and hopefully some interesting disagreements and interesting debates on things. And, yeah, the there’s tickets available. Hope people can join.
Swyx [01:27:56]: Since you mentioned it, what disagreements and debates should people think about, or do you expect?
Anastasis [01:28:04]: So it’s things like, there is, one debate right now in the robotics world is, VLA’s versus world action models.
Runway AI Summit and the Big World Model Debates
Swyx [01:28:11]: Okay.
Anastasis [01:28:11]: So there is labs that are really betting on one of those two directions. there is like what is the best source of data to train robotics models?
Swyx [01:28:21]: There’s just the third-party, first-party that we talked about.
Anastasis [01:28:24]: Yeah. There is, the people who really believe in further scaling teleop data versus leveraging more large-scale video data. So that, those are some of the. And then there is, the world models debates of predict pixels directly versus something like JEPA versus a more 3D-based, 3D-based approach. so I think we’re at a nice time in world models because there is still that active debate happening on, like, what is the best long-term direction. I feel very strongly that it’s video predict pixels directly and scaling video generation models is the right approach. But it’s, I think there is a lot of interesting, debate happening, by researchers on, like, what is the best path to take.
Swyx [01:29:09]: It’s interesting that it’s all on, like, let’s call it the policy layer and the data model layer. Is the physical side is completely solved? Like, all the sensors, all the actuators, all these things are. We have everything that we need?
Anastasis [01:29:23]: I don’t think that’s, solved either.
Anastasis [01:29:25]: It’s definitely,
Vibhu [01:29:27]: Different problems.
Swyx [01:29:28]: It’s, it’s like
Anastasis [01:29:29]: Yeah
Swyx [01:29:29]: I wanna dream about all these things, and then I get, I buy a robot or I buy, I try to assemble my own, and I can’t even get the motors to, like, work right. Right? Like, and it’s you’re dealing with very sensitive, equipment that has, voltage and power and, like, heat and all these things which, you, abstracted away. We’re sitting here, we’re talking about software and talking about models, but, like, really you have to deal with those kinds of things too.
Anastasis [01:29:56]: Yeah. And, I think I’m, I’m, I’m generally also not opposed to incorporating other modalities into our models like we’ve seen.
Multimodality, ImageBind, and the Maximalist World Model
Swyx [01:30:04]: Yes.
Anastasis [01:30:05]: The simplest case is they can generate video and audio at the same time. So they can generate RGB, and they can also generate, they can generate sound and audio. But my. I’ve written about this as like what does the maximalist version of a world model look like is you’re incorporating more and more modalities from the universe And you’re training a model on different scales of observations as well.
Swyx [01:30:28]: X-rays.
Anastasis [01:30:29]: And so, yeah,
Vibhu [01:30:30]: You got a good essay that people should read on
Anastasis [01:30:33]: Yeah. Yeah
Vibhu [01:30:33]: Real-world.
Swyx [01:30:33]: No, Meta released a model that was, like, six modalities in one, right?
Vibhu [01:30:37]: Yeah.
Swyx [01:30:37]: I forget what the name of the thing was, but it was like, yeah, okay, depth is one of them, but depth is like a transformation of RGB in some sense.
Vibhu [01:30:45]: ImageBind.
Swyx [01:30:45]: ImageBind, yeah.
Vibhu [01:30:45]: Yeah.
Swyx [01:30:46]: What other modalities? They had heat?
Vibhu [01:30:47]: Audio, depth, heat, text,
Swyx [01:30:51]: Whatever IMU is.
Swyx [01:30:52]: I do think, like, you might as well do ultraviolet. You might as well do, like, just whatever other modality you feel like, ‘cause it’s all data to the model.
Anastasis [01:31:01]: Yeah. And, a big bet is also that there is transfer between all those modalities.
Swyx [01:31:05]: Yeah. Yeah.
Anastasis [01:31:05]: So one of my favorite, examples, which is quite old at this point, is there was this fine-tune of, Stable Diffusion that was called Riffusion Which was
Swyx [01:31:14]: The music one. Yeah.
Anastasis [01:31:15]: Yeah, just fine-tuning, Stable Diffusion on spectrograms.
Swyx [01:31:18]: Spectrograms.
Anastasis [01:31:18]: And it became a quite capable music generator. Right? So there is probably Spatial patterns, so like spatial-temporal patterns if we’re talking about video that emerge at different scales and different modalities. And so there is some degree of, meta-learning that the model has done that allows it to learn faster if you start from a just a model trained on images and train it to predict audio than if you train from scratch on just audio. and there is some other interesting examples. So there is this project called The Well. It’s, it’s a dataset of physics and numerical simulations in physics and biology and a bunch of other domains. So it’s, so it’s essentially different physical systems across very different scales of space and time, from like astrophysics to low-level like atomistic interactions. And we’ve seen. we’ve done some work on this, and we’ve seen that we can take our video model where, real-world video looks nothing like this, and you can fine-tune it on those numerical simulations and just treat them as RGB frames. And you get reasonable performance much quicker than if you just train from scratch.
Anastasis [01:32:36]: Yeah.
Vibhu [01:32:36]: I think we’ve seen this across languages where
Swyx [01:32:38]: Yeah, DeepSeek-OCR as well.
Vibhu [01:32:40]: Yeah, DeepSeek-OCR.
Swyx [01:32:41]: Like, you don’t have to tokenize text. Like, you can just throw them in as images.
Vibhu [01:32:44]: There’s a lot that happens in that base pre-training. Like, there was an argument a long time ago of people saying, “Oh, humans have so many, sensory representations, right? Smell, touch.” Models have a whole two more modalities that we’ll like, that we don’t even have data for. And it’s like, okay, you take AQI sensor, like you can try this stuff, but there’s so much happening in just the base trainer on that you don’t get as much from these little things.
Scientific Data, Cross-Modal Transfer, and Omni Models
Anastasis [01:33:10]: Yeah, exactly. And I think that’s what it solves is data scarcity.
Vibhu [01:33:13]: Yeah.
Anastasis [01:33:13]: So you don’t have as much. You have so much video data available, but you don’t have, like olfactory data that
Vibhu [01:33:21]: The cool thing is it goes the other way too, right? So if you wanna do physics, like if you wanna measure this or you wanna have a diffusion model do audio, it transfers really well. So like in your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain. So we can apply that to other stuff too.
Anastasis [01:33:40]: Yeah. And if we look at, like how do you make those models more useful for in scientific domains, and if you look at AlphaFold, they’ve had all these very. Because of the data, the limited amount of data that it needs to be trained on, it’s it’s very fine-tuned architecture just to solve, protein structure prediction. But if you take all those disparate sources of scientific data and you bring them together under a single model, like I think that’s an approach that can help us solve new kinds of problems across science by leveraging all the learnings from one modality or one set of, data to another. So very early days for that direction, but I do think that’s where ultimately what the end game of simulating the world is. You’re not just using RGB. You’re using RGB as a starting point, but you can incorporate more and more modalities of the universe and leverage the transfer that happens from learning from one to the other.
Vibhu [01:34:42]: I guess the follow-up there is what’s the drawback of omni? Like, why is everything not an omni model? Also, why not now, and why. Would you start from language backbone or image video backbone and then go omni from there? Does it matter?
Anastasis [01:34:57]: Yeah. We need to take it one step. We need to solve robotics first, and then we can go into
Swyx [01:35:02]: Solve everything now.
Anastasis [01:35:05]: Yeah. I do think there is a lot of open-ended research that needs to happen for, those omni models. There is a lot of things that require careful consideration when you’re bringing multiple modalities into a single model to predict. But I think, I expect those to be solvable.
Closing: Film Festivals and the Future of AI Video
Swyx [01:35:23]: Wonderful. you’ve been very generous with your time. Congrats on all your success, and, yeah, I’m excited for the, AI Summit, or physical AI Summit.
Anastasis [01:35:32]: Yeah, thanks for having me.
Swyx [01:35:33]: And yeah, and people should check out the film festival if it’s in town, right?
Anastasis [01:35:37]: Yeah.
Swyx [01:35:37]: Yeah. You’ll be gonna be touring all over the place.
Anastasis [01:35:39]: Yeah. Next year we’re probably gonna do that. So we do film festivals every May or June of
Swyx [01:35:45]: Yeah.
Anastasis [01:35:45]: And we did the last one in New York, LA, Tokyo, and at the AI Engineer,
Swyx [01:35:52]: Yeah
Anastasis [01:35:53]: Fair.
Swyx [01:35:53]: Yeah. Yeah.
Anastasis [01:35:54]: So yeah, hopefully even more places next year.
Swyx [01:35:57]: No, I think like someday, you will be hosting the Oscars of AI video, and, I think people should like take this very seriously as like a potential career they can have.
Anastasis [01:36:07]: The Oscars of AI video will be called the Oscars.
Swyx [01:36:10]: All right. All right. Thank you.
Anastasis [01:36:14]: Thank you.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe - The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then?
Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities
Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind.
Building a virus from scratch
While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses!
Long context unlocks biological intelligence
Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long:
* 60K for an average human gene
* long being up to 2.3M
* the whole human genome around 3B.
Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models.
Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA.
Thinking in DNA
Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and specific genes that encode proteins. This means that the models already generalize to multiple “languages,” before even attempting to train in other modalities, such as 3d protein structure, epigenetics and natural language.
If a model thinks in the DNA language, maybe it understands the imprint that environment left on different genomes as well? Perhaps the model has learned the functional relationship between different sequences, and could extrapolate to new sequences based on that?
And so what we wanted to showcase was that if we show the model progressively better RNAs in a series of steps with its score, right? So you have like low scores first and then you gradually move up the chain. Can the model continue that trajectory on its own? And then in the final step, does it self optimize to a point where it's like the best score it can get? That was the experiment. Can we do that? And so we took a data set, a large data set of aptamers. We held out a portion of the best performing ones and we showed it only the lower ones, but then we ranked it, right? So we showcase lower scores with the RNA aptamers and then progressively got higher, and then ask the model to just like continue with that pattern. And it turns out it was able to recapitulate some of those higher scores that we had not shown it yet.
So, voila: chain-of-thought, thinking in DNA!
The arms race
But much as long-context inference, chain-of-though and multi-modal perception unlocked sophisticated reasoning in natural language LLMs, these capabilities in GLMs are enabling increasingly sophisticated “biological intelligence,” and along with it, greater danger.
According to Eric, defense is currently losing this battle, but Radical Numerics argues to push the frontier harder!
I won’t spoil the details for you. In the episode we talk in detail about:
* Biosecurity as an arms race — and how defense can keep up
* The genome as the imprint of the environment on DNA
* Going truly multi-modal
* How chain-of-though works when you “think” in the language of DNA
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
22.09.2026 | 2 godz. 1 min.How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI.
John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score.
Google’s Empirical Research Assistance (ERA)
John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort.
John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them!
The result is Google’s Empirical Research Assistance or ERA (paper, github, blog). ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search: at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other.
“It’s almost like having a hyper-eager grad student who doesn’t sleep.”
Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great.
ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers. Some of these were climate change related, which we talk about in the next section.
So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pixel or the center?) and this turned out to be a part of the winning special sauce. Great for winning $15,000, not so helpful if you actually want to solve contrails.
“People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart’s law: any metric that becomes a target is no longer good as a metric.”
His advice for where to start instead?
“Always just fit linear regression. Just do it. Just do it. Just do it. Or SVM.”
Tackling Climate Change with AI
John and his team have worked extensively to mitigate the effects of climate change. We talked about several of their initiatives.
Perhaps the most interesting result we talked about was reducing the effects of condensation trails (contrails) from airplanes. Those little streaks you see running behind planes somehow account for 1% of all human-induced global warming?!? Some of these trails of ice crystals can hang out for days. These crystals are black in the infrared, acting like a thermal blanket that traps heat day and night.
It’s easy to understand what’s happening here, a region of atmosphere becomes “ice supersaturated”, and a tiny bit of exhaust seeds water vapor that instantly crystallizes. The scale here is astounding, with a single gram of exhaust resulting in ten kilograms of ice crystals.
The solution to all of this is quite simple, in principle! We know what parts of the atmosphere are most likely for the trails to form. Just have the planes drop a flight level or two. Problem solved, right? Well, the hard part is accounting for how much warming was prevented. This is a counterfactual problem, parts of which stumped John’s team for over two years. They had a working model for the heat-trapping half, but not for the reflected sunlight. ERA was able to find a simple model with some confounders they hadn’t considered. Cracked it!
Modeling climate generally is a hard problem. Climate is best thought of an attractor of many different possible weather outcomes. This makes it much harder to model.
“Weather is where you are on the attractor, and climate is the statistics of the attractor. The problem with climate is that we’re altering it. The attractor itself is changing, it’s moving.”
John and his team have worked on treating both the symptoms and the disease of climate change, with several other works in the area. Another fun example we briefly cover is FireSat, a way of using a constellation of satellites to rapidly identify fires before they grow too big to put out. For anyone living in California, you understand the problem. In dry years a small fire can result in hundreds of thousands of acres. If you could find this fire when it’s the size of a room, it could be put out. By the time it hits an acre we have a much harder problem.
Where is this all going? Looking forward by looking back
By now it should be clear John has an incredible and unique view over the intersection of science, computation, and AI. John talked about a class on physics of computation he took with Richard Feynman back in 1982. This was when quantum computing was an ill-defined concept with no theory or experimental backing. John recalls every Tuesday was a guest lecture, and every Thursday was Feynman explaining why the Tuesday guest was wrong. John also recalls doing science back when there was essentially no compute, a million operations per second was cutting edge.
What is John’s recommendation: the most important skill is developing deep domain expertise. There’s no other way to develop taste than to tackle hard problems. One surprising part of this is that John recommends spending time doing things the old fashioned way. Play with tools, and just implement things yourself.
“You could drive up the mountain, or you could hike up the mountain, and maybe it’s okay, even fun, to occasionally hike.”
Summing it up, John’s message to the audience is that there will still be a place for scientists, and that if anything it will just open up more opportunities for “the creative stuff, the rigorous stuff, the philosophy stuff.” But don’t forget to spend time doing the grunt work.
“There just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyper-optimize. It’s overfit.”
And whatever tools you end up using, John’s advice is the same one Feynman gave him forty years ago: you must not fool yourself, and you are the easiest person to fool.
We had a great time talking with John. We hope you enjoy!
Also in this episode
* Fusion is three years away, not thirty, if you ask John. And why the Lawson criterion means every fusion approach has an Achilles heel.
* Why superconducting qubits are still finicky.
* The asteroid he named after his mom, which turned out to have a moon.
* The looming helium shortage nobody talks about.
* How NeurIPS started as people crashing a private workshop at Snowbird, and why Hopfield networks are all you need.
* Being Carver Mead’s sysadmin on a VAX with an 80 MB disk the size of a dishwasher.
* Finding asteroids in 1985 with film, a stereoscope, and a letter to Brian Marsden. The Vera Rubin Observatory found 11,000 in six weeks.
* The Feynman effect: total clarity in the room, none once you leave.
* Quantum echoes, the NISQ era, and why he thinks quantum is neither thirty years away nor tomorrow.
* A startup that wants to inject mercury into a fusion reactor and sell the transmuted gold. “It might not work.”
* John’s 20% time rule for his own group: do stuff for learning, and you don’t even have to tell him what.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribeUnderwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
16.09.2026 | 1 godz. 26 min.AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:
From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?
We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.
We discuss:
* Why risk, liability, and trust may become the binding constraint on AI adoption
* Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days
* What Anthropic understood about scaling, compute, and the future years before it became obvious
* Why Waymo illustrates the gap between AI capability and real-world deployment
* AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies
* AIUC-1: a standard for AI agent security, safety, and reliability
* How agents are tested for jailbreaks, hallucinations, and data leakage
* Why most AI companies optimize the happy path without seriously stress-testing adversarial cases
* Why AI standards may need to update every quarter instead of every decade
* The emerging trust gap between frontier AI labs and governments
* Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface
* Why standards and insurance may need to evolve together
* How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment
* What happens if a $20 Cursor subscription contributes to a $200M plane crash
* The Air Canada chatbot case and how AI failures are beginning to clarify legal liability
* Why copyright may be one of the hardest AI risks to insure
* Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested
* The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong
* Why robotics will make AI liability dramatically more consequential
* Whether AI engineers should have Level 1, 2, and 3 certifications
* AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming
* Why AGI could become a question of national sovereignty
* Why the labs can never fully serve as their own watchdogs
* The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?
Rune Kvist
* LinkedIn: https://www.linkedin.com/in/runekvist/
* X: https://x.com/RuneKvist
AIUC
* https://aiuc.com
Timestamps
00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI
00:01:07 From Scaling Laws to Early Anthropic
00:07:58 Why Trust, Not Capability, Could Limit AI Adoption
00:12:19 Founding AIUC and Building AIUC-1
00:18:52 How AI Agents Are Audited and Stress-Tested
00:25:26 Frontier Models, Government, and the AI Trust Gap
00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk
00:38:14 Why Standards and Insurance Belong Together
00:41:45 What Does an AI Insurance Policy Actually Cover?
00:50:44 The $20 Cursor Subscription and the $200M Plane Crash
00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust
00:56:21 From AI Agents to Models to Robotics
00:58:29 Copyright, Adverse Selection, and AI Insurance
01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness
01:08:36 The Impossible Enterprise AI Mandate
01:11:52 Prediction Markets vs. AI Audits
01:14:43 Should AI Engineers Be Certified?
01:19:10 AIUC’s Roadmap, AGI, and Who Watches the Watchdogs?
Transcript
Introduction: AIUC, the $40M Series A, and Risk as the Adoption Bottleneck
Swyx [00:00:00]: Okay, we’re in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.
Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.
Swyx [00:00:11]: What are you announcing today?
Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.
Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?
Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It’s pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.
Swyx [00:00:54]: And let’s get a list of the customers that you’re highlighting as part of your Series A.
Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.
Swyx [00:01:05]: Yeah. Amazing. Congrats.
Rune Kvist [00:01:06]: Thank you.
Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I’m just kind of curious: what was your path into AI? Just recap.
Rune’s Path Into AI: Scaling Laws, Capital, and Anthropic
Rune Kvist [00:01:18]: Yeah.
Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.
Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?
Rune Kvist [00:01:40]: Exactly, the Kaplan one.
Swyx [00:01:42]: Yeah.
Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what’s played out. And so I just packed my bags. I’d never been to San Francisco. I’d never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They’d just broken off from OpenAI, and it’s been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there’s nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.
Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.
Rune Kvist [00:02:52]: Yes.
Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy s**t.”
Rune Kvist [00:03:19]: Correct.
Swyx [00:03:20]: Right?
Rune Kvist [00:03:21]: Yes.
Swyx [00:03:21]: Who tipped you onto that paper? Because it’s not a paper that you normally read, right, like, in your circles?
Rune Kvist [00:03:26]: Yeah. I think I’d actually, ever since AlphaGo, had some appreciation that AI was a big deal.
Swyx [00:03:36]: Yeah.
Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn’t as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you’re going to be wrestling with all of the big questions in society. Everything you’ve learned about politics gets thrown out of the window. Everything you’ve learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that’s why I sought it out.
Swyx [00:04:32]: I mean, clearly really good insight. For people who don’t know, the PPE program is, like, where prime ministers are born. So then you end up meeting Dario.
Rune Kvist [00:04:41]: Yep. First Dario, yeah.
Swyx [00:04:43]: Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn’t get from your original hypothesis?
Anthropic’s Early Conviction and the Scaling Laws Crystal Ball
Rune Kvist [00:04:50]: If you read the Scaling Laws paper, you get this, like, very vague sketch of like, wow, this seems kind of important. There are some lines on a chart. This seems kind of important. And what I think the team at Anthropic had thought more about than anyone was like, what are the implications of this if you really play this out? And back then they had, kind of vision documents for what the world would look like in 2026, and they were kind of in vivid detail playing out how much compute is going to be needed, what the CapEx was going to look like, what some of the societal concerns were going to be, but also what is the amount of economic value coming out here? And so it kind of felt like they held a crystal ball that in hindsight turned out to just be dramatically correct. And they weren’t holding it like they were obviously correct. They were just like, “Take this hypothesis really seriously.”
Swyx [00:05:38]: Think it through, yeah.
Rune Kvist [00:05:38]: And think it through in the same way as the kind of situational awareness that is
Swyx [00:05:43]: Across the street.
Rune Kvist [00:05:44]: Across the street.
Swyx [00:05:44]: Your office, yeah. Oh my God, we’re all living across the street in the same one square mile.
Rune Kvist [00:05:50]: Correct. And that’s now a couple of years old, but also people keep referencing it these particular weeks with Fable and Mythos, and it’s like, wow, if you take this one idea seriously- For the Scaling Laws, a lot of things fall into place.
Vibhu [00:06:03]: And keep in mind, at this point, this is the same team that did GPT-1, GPT-2, and GPT-3.
Rune Kvist [00:06:08]: Correct.
Vibhu [00:06:08]: Which is also, like, it’s not just some experimentation. Like, this is a real model that we just scaled up.
Rune Kvist [00:06:14]: And they had deep conviction in this idea: if you take a big blob of compute and data, it just wants to learn, and out of that will come smarter and smarter models. And all the particulars were not clear.
Vibhu [00:06:26]: Yeah.
Rune Kvist [00:06:27]: And all the implications were not clear. But their deep conviction in this, like, core thesis, and that was kind of dizzying. It was both phenomenally interesting and exciting, and also very quickly you get to, like, the world we know today will no longer be if this hypothesis holds. So it also just felt, like, important in some kind of grand sense.
Vibhu [00:06:48]: What kind of shaped you there? So that was early 2022. Not only had GPT-1, GPT-2, and GPT-3 come out, but, you know, the amazing founders of Anthropic that have never split up, the only ones, they actually had the conviction to leave OpenAI, start their lab. You said there were about 40 people there. What was the time like there?
Inside Early Anthropic: Mission, Deployment, and Risk
Rune Kvist [00:07:06]: It was kind of remarkably like what it looks like on the outside today. Extremely cohesive, extremely mission-oriented, and living in this tension between their two ideas, which is AI could both go really well and really bad, and we want to be part of building it. That creates astounding amounts of tension. And they were wrestling with this incentive challenge where they know they’re in a race that they’re in where you might get forced to cut corners, but it also felt very important to them to be at the forefront of technology. And all of those ideas were just present at that time. It kind of feels like that line has been just very clear, and I think kind of love them or hate them, they have really stuck to their guns. There’s a core set of beliefs that they hold more deeply than most companies hold any beliefs.
Vibhu [00:07:58]: Yeah. Fast-forward to today.
Rune Kvist [00:08:00]: Yeah.
Vibhu [00:08:00]: What does that lead us to AI underwriting company? What are you up to? What motivated you to start this?
From Waymo to AIUC: Confidence Infrastructure for AI
Rune Kvist [00:08:05]: Yeah. AIUC builds confidence infrastructure for frontier AI through standards and insurance. The link from Anthropic to building confidence infrastructure, looking out the windows at Anthropic offices and seeing Waymos driving by. Already back then, early 2022, Waymos were in some ways like AGI for cars. Like, they were superhuman drivers, but you couldn’t take one to the airport. And now, four and a bit years later, you still can’t take your Waymo to the airport, despite now everyone having kind of looked at the evidence and being like, “They’re better drivers than humans.” So in that particular instance, what’s clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust. That problem is, general. The reason why right now
Rune Kvist [00:08:52]: Fable is not open for access is not because it’s not a good model, it’s because it’s a very good model. It’s just hard to make promises about what it will or will not do. And this problem gets worse as AI gets better. Basically, more intelligent AI can be more autonomous. That’s more valuable, but also the risk surface grows. And so - what Waymo illustrates is that unless you build the confidence infrastructure to make promises about AI, or at least bring light to the risks, you grind adoption to a halt. Governments, banks, hospitals, militaries need to have some sense of what AI will and will not do to be able to operate for them to incorporate it. And that’s the problem that we’re trying to solve. Now, why standards and insurance? If you trace this problem back through history, every technology wave has had some version of this problem. So if you go back to, like, year 1900, electricity comes
Vibhu [00:09:47]: Ben Franklin.
Rune Kvist [00:09:48]: Cars burn down, sorry, houses burn down, lots of people die. 1930s, cars are a big deal, kill lots of people. 1950s, private nuclear energy is a big deal, poses big risks. In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that’s required to make go/go decisions. That is required for adoption, and the market fundamentally wants adoption. And in all of those instances, common blueprint emerges between standards and insurance. The reason these two components is standards kind of provide the rules of the road, and they also specify, like, what are the tests that need to be run so we can get a sense of how high the risk is. So take in the case of cars, that’s like a car crash. Great, everyone, they inform your insurance pricing today, they inform your purchasing decisions, et cetera. That’s basically the risk framework. The insurers are important because they pick up the bill. So they are the private institution that is most on the side of. That is best incentivized to quantify the risks truthfully and then figure out all the ways to reduce the risk ‘cause that increases their profit. So they’re basically, they help shape the incentives. And these two work really well in unison. Now, how does that show up as a company? Well, one of the things that was obvious even - or starting to become obvious even a couple years ago was that frontier companies, some of our customers today, like Cursor, Sierra, ElevenLabs, Harvey, were going to have a very easy time selling a pilot to a bank. The, like, the demo just sells itself. It’s magic. But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital, you have to go through the risk process. These banks have no idea even which questions to ask, let alone which answers are sufficient, let alone, like, how do they go and test whether these agents actually work the way they’re supposed to. And so they had this problem of, like, what can we say to earn the trust? And we think there’s, like, a golden sentence that goes something like, “Hey, I hear you’re really worried about hallucinations or jailbreaks or whatever it may be. We’ve had an independent third party test us against the gold standard. We passed with flying colors. And as a vote of confidence, the world’s most conservative insurers have looked at the data.” And they’re willing to take some of the risk onto their balance sheet.
Swyx [00:12:06]: Yeah.
Rune Kvist [00:12:07]: So if something does go wrong
Swyx [00:12:07]: There’s money behind it, yeah.
Rune Kvist [00:12:09]: Exactly. So that’s kind of like the link between all this. We can get into some of the hard parts related to the technical testing, which is, I think, the crux of the matter, but I’ll pause there.
Swyx [00:12:19]: How did you and Rajiv come together? This-- there’s always, like, you come across very confident and, you know, and we’re announcing your Series A and all these things, but I want to see, like, the early initial stages of, like, idea formation.
Cofounding AIUC with Rajiv Dattani
Rune Kvist [00:12:31]: Yeah. Rajiv is actually my soon-to-be brother-in-law.
Swyx [00:12:35]: Oh.
Rune Kvist [00:12:36]: So I’m actually, in a week and a half getting married to Rajiv’s sister.
Swyx [00:12:42]: Okay, now you’re tight.
Rune Kvist [00:12:44]: Exactly.
Swyx [00:12:44]: Now you know.
Rune Kvist [00:12:45]: So - Rajiv and I have known each other for a decade. Funny story, I met both Rajiv and his sister, Hena, at the same time when Hena and I were interns at McKinsey in London, and Rajiv was assigned as my mentor. And so met them at the same time. For the longest time, it was not obvious that we were necessarily going to work together. I was in startups. He was, an insurance partner at McKinsey. Three or four years ago, I think Hena convinced him that AI was going to be a really big thing. And so he quit his job, cushy partner job at McKinsey in London, packed his bags, flew to San Francisco, and ended up joining METR. You guys are probably online enough
Swyx [00:13:24]: CEO.
Rune Kvist [00:13:24]: Exactly.
Swyx [00:13:24]: We’ve, we’ve, we’ve heard of METR.
Rune Kvist [00:13:25]: You see the plot-- the chart of the horizons of the tasks that agents can take on is doubling extremely fast. So he was COO at METR, led their partnerships with Anthropic and OpenAI to test their models before release, but also working closely with the US and UK government, to figure out, like, how do you know whether a model can be released? And in some ways, that was, like, the perfect background. He’s spent a lot of time in insurance, knows that world, spent a lot of time with frontier testing of models. And so when I was bumbling around this idea space, starting with some of the ideas we talked about related to Waymo, as soon as we got into the content, we were both like, “Oh, this would be an amazing business to build together.” This is wrestling with the problem that we both think is the most important in the world from a market angle, which is kind of our intuitions is that the market can do a lot, and the faster AI moves, the harder it is for government to solve some of these problems. And then it took a little bit of time to work through what is it like to work with family.
Swyx [00:14:27]: Sure.
Rune Kvist [00:14:27]: And,
Swyx [00:14:30]: Because you were already dating at the time
Rune Kvist [00:14:31]: Yeah. Yeah, exactly.
Swyx [00:14:33]: Yeah.
Rune Kvist [00:14:34]: Already back then, it
Swyx [00:14:35]: Yeah.
Rune Kvist [00:14:35]: We felt like we were a family.
Swyx [00:14:36]: Nice.
Rune Kvist [00:14:36]: And so starting a business together felt like kind of a big step. And, here we are with just immense amounts of trust.
Vibhu [00:14:43]: Yeah. So now you’re a company of how big? How big are you guys now?
AIUC-1 Certification: Agent Security, Safety, and Reliability
Rune Kvist [00:14:46]: There are just 20 of us now.
Vibhu [00:14:47]: 20 of you guys now, have Series A, and you have your first certification out, the AIUC-1. Let’s bring up the certification. So this is the agent certification, right? What goes into the process? I have, like, two questions here. One is, walk us through the certification, and two is, what is the process for a company to get certified, you know?
Rune Kvist [00:15:08]: Great. As it says right on the top, AIUC-1 is a standard for agent security, safety, and reliability. The fundamental design principle is take all of the concerns that slow down adoption, so all the questions, all the fears that keep, security leaders in the Fortune 1000 up at night, and put them into one comprehensive framework. That’s what you’ll see there. You can see the six categories. Two, you want to ground all of this in technical testing. So one of the concerns with security standards that often feel kind of like theater paperwork is that they’re not actually ground out in, does any of this work? Does any of this matter? And so we had a conviction from early on that was going to be the kind of crux, was to pass this, you must get tested every quarter, basically run thousands of simulations to see, well, so can it actually be jailbroken? How hard is it to jailbreak? How often does it hallucinate? How often does it leak data? Et cetera. And then the last, core idea here, if you scroll up to the top here, is to refresh it quarterly.
Rune Kvist [00:16:08]: So the core trait of AI is that it moves extremely fast. Whatever concerns we’re discussing today were not the same ones three months ago, and this will keep changing. Typically, standards update on a, like, a decade cycle is obviously not going to work. But the question is kind of how do you update it? And the core thing here was to basically get the risk leaders of the Fortune 1000 around the table. So if you go over to the left here
Vibhu [00:16:32]: Yeah
Rune Kvist [00:16:32]: You’ll see the AIUC-1 consortium. The consortium is a group of risk leaders who run real banks, real hospitals, real critical infrastructure, who are facing these challenges every day. And we meet with these folks twice a quarter and hear what’s top of mind, what is keeping them up at night. There’s tremendous amount of desire for that conversation. And then we operationalize that into a specific standard that gets into. And actually, we can go into and look at what
Vibhu [00:16:55]: Yeah
Rune Kvist [00:16:55]: What even is the standard. So if we go back to introduction, out there to the left, scroll up a little bit to the wheel, click into reliability. So if you take something like hallucinations sits in reliability. There is a number of requirements here. If you go into the top one, prevent hallucinated outputs, hallucinate outputs, this is one particular requirement. This is a technical control. Basically, we want some kind of ground in this filter. The first thing you see here is what’s called a crosswalk. So everyone and their grandmother has put out a framework, very high-level framework for what are the AI risks.
Swyx [00:17:27]: This is basically your competition,
Rune Kvist [00:17:28]: In some ways our competition
Swyx [00:17:29]: Not seriously, yeah.
Rune Kvist [00:17:30]: We’re, in fact, friends with them. We’ll come back to why.
Swyx [00:17:31]: Yeah.
Rune Kvist [00:17:32]: But mapping everything together so you have one superset. The claim you’re trying to support here is, if you follow this framework, then you can also see how you follow the other frameworks. But the meat of it comes down here in control activities and evidence. So control activities is like, great, you have this high-level requirement. How do you turn that down to something operational? Here’s what you must do, and then what is the evidence that we’re looking for?
Rune Kvist [00:17:57]: And the reason we go this deep is that there’s actually not that much confusion about what are the big concerns in AI. Everyone agrees to these. The question, like, what are you actually supposed to do? And so. What we found a lot of demand for is getting down to the specific evidence, that people need to look for. Whether you are Cursor building something or, even JPMorgan building something, but also if you’re just a risk leader at JPMorgan, like what exactly should you ask for? What can you ask for without sounding stupid? Like if you ask for some-- you won’t believe the amount of time a risk leader has asked for the IP rights to the underlying model to Cursor or something, and you’re just like “Sorry, what?” Like,
Swyx [00:18:39]: You slip it in there and you see
Rune Kvist [00:18:40]: Slip
Swyx [00:18:40]: See if you notice.
Rune Kvist [00:18:41]: See if they. Exactly.
Swyx [00:18:42]: Yeah.
Rune Kvist [00:18:42]: Put that in the questionnaire. All right, so that’s kind of what our standard is, and we update this every quarter with these folks, to keep up with the latest concerns.
Swyx [00:18:51]: Can I double-click on this one?
Controls, Evidence, and Third-Party Testing
Rune Kvist [00:18:52]: Yeah.
Swyx [00:18:52]: So first of all, the website’s beautiful. Like, it’s so confidence-inducing which is the whole point where, like, okay, I know exactly what I’m signing up for when I talk with you. Like, I don’t even have to talk to you. I can just see your whole, certification, which is great. But, like, okay, so from here, like D001.1 configure a groundedness filter, how does that get applied? Like, you have a person that
Rune Kvist [00:19:16]: Yeah,
Swyx [00:19:16]: Goes through it?
Rune Kvist [00:19:17]: If you, go back
Vibhu [00:19:19]: I did see somewhere there’s like, you know, fifty-one requirements, a hundred thirty controls. There’s like a whole
Swyx [00:19:25]: Right. I just want to. Like, to me, this doesn’t translate
Vibhu [00:19:27]: Yeah.
Swyx [00:19:27]: Into a test or an eval.
Rune Kvist [00:19:28]: Yes. So if you go into, on the left-hand side. So actually, if - before we go in there are three types of requirements. The first is technical controls, like you must implement some guardrails.
Rune Kvist [00:19:42]: Two, there are test controls. So you must have an independent third party go and run some tests against you. I’ll show you one of those in a second. And then three, there are policy controls. For example, you must have a person whose name is on the line when you guys f**k up, and you must have a plan for how you tell your customers and how you engage with them. They’re kind of more traditional, standard type stuff. So in this particular instance, we just check whether they in fact have a ground in filter. So we will partner with an auditor. So we partner with auditors like KPMG or like Schellman who go in and do the thing auditors do, which is to check the evidence. In this case, that might be a screenshot, it might be part of the code that they need to review to see that it actually. Just that it exists.
Swyx [00:20:21]: Oh, okay.
Rune Kvist [00:20:22]: And then the second thing
Swyx [00:20:22]: So you’re not testing the effectiveness of it.
Rune Kvist [00:20:24]: That’s the second thing. So if you go down
Swyx [00:20:25]: Yeah.
Rune Kvist [00:20:25]: To the third-party testing for hallucinations out on the left, that’s basically the next requirement. This is where we test how well does it actually work.
Swyx [00:20:32]: Okay, and is it you testing or the auditor?
Rune Kvist [00:20:34]: We test them.
Rune Kvist [00:20:35]: We test them.
Swyx [00:20:36]: That’s a lot of work.
Vibhu [00:20:37]: How long does testing take? So if I want to get certified, just
Certification Timelines, Remediation, and Quarterly Updates
Rune Kvist [00:20:40]: Yeah.
Vibhu [00:20:40]: How long does the end roughly take?
Rune Kvist [00:20:42]: Yeah, the end, almost always is dependent on, like, our customers need
Vibhu [00:20:47]: Yeah.
Rune Kvist [00:20:47]: To look something for us. It takes somewhere between, like, 3 to 10 weeks
Swyx [00:20:52]: Yeah.
Rune Kvist [00:20:52]: Depending on how up to snuff they already are. So some people show up to us with, like, extremely rigorous security programs. When we test them, it works extremely well. We can get that done very quick. Some people come to us, and they’re not that far along. We give them kind of the spec that they need to build towards, and then their security teams and engineers get to work and build to meet the standard. The testing itself typically takes a couple of weeks, including the time for them to remediate. Often, we’ll find something that we cannot pass, where this is actually just not up to the standard. - you won’t pass the standard. And then they will need to go and implement additional safeguards or additional remediation that makes them more robust so that they can actually kind of hand on heart look at their customers in the eyes and say, like, “Hey, we’ve done truly our very best.”
Vibhu [00:21:35]: And they’re certified for a year and have quarterly updates?
Rune Kvist [00:21:38]: Correct, yeah.
Vibhu [00:21:39]: And, yeah, it’s pretty interesting. I think, you know, what’s changed since. So this is certifying agents in production, right? Your customers, like you’ve had Lovable, ElevenLabs, Intercom, and they’ve all gone through this certification.
Rune Kvist [00:21:50]: Yes.
Vibhu [00:21:51]: What has changed? So I see you post, like, you know, Q2 added MCP agent,
How Agent Risks Are Changing: Coding, MCP, and Agent-to-Agent Interactions
Rune Kvist [00:21:56]: Yeah.
Vibhu [00:21:56]: agent communication. Any other things that you want to kind of highlight since the first iteration? What comes in quarterly?
Rune Kvist [00:22:03]: Yeah. So some of the changes have just been agents are not just one thing. So, like, if you take agents like Cursor and compare them to Sierra, they’re really quite different. And compare them to Harvey again, compare them to you out of again
Swyx [00:22:16]: ElevenLabs, yeah.
Rune Kvist [00:22:17]: ElevenLabs, they’re all quite different. And so we wanted to design a standard that works for all of the types of agents. And we started with one that was, like, pretty text-based, like, honestly, pretty customer support-focused. That’s where there’s a lot of existing demand. And then over time, we’ve picked, some of the frontier companies in each of these other domains that we could work with and build out the standard, so, such that we know that the same standard works for code, it works for customer support, works for automation, et cetera. So that’s been one big thing. Yeah, then some of the things that have been top of mind recently, Mythos is bringing up a lot of concerns for security leaders. We’re starting to get more and more questions around agent interactions. It’s very nascent, at the moment, but it’s starting to emerge. There’ve been a lot of, questions related to OpenClaw and MCP. Again, like agents starting to interact with each other, is really top of mind. Then as coding agents have really taken off, that’s also where banks and hospitals, et cetera, are getting more and more precise on what it is they need. So really dialing in as that start to be, like, where most of the tokens flow through in the world, getting much sharper on that.
Vibhu [00:23:26]: Can you share for people that are listening that don’t really think about this? Like you mentioned, there’s the obvious stuff, you know, hallucination, citations. What are best practices that people should do when building agents? Like, if they come to you pretty ready with certification like, you know, they’ll probably pass certification. What are the things people don’t think about that they should have?
Best Practices for Agent Builders: Stress Tests and Guardrails
Rune Kvist [00:23:46]: The most important thing is that a lot of companies have not done a serious stress test. They spend most of the time, perhaps rightly so, optimizing for how does it work in the good case, the average case, how high-quality is the output for the customer. And a lot of these companies are pretty new, so they haven’t spent a lot of time stress testing the what is there as an adversary on the other side? What are some of the complicated corner cases that you’ve not really considered? So I think that’s, like, a frame of mind. And you’ll also see this in startups. It often takes a while until they hire their first security person. They- And that’s a whole different kind of risk surface than just building a good product. So a lot of that applies. Most companies actually also have the right kind of architecture. Most of them will have some kind of guardrails in place, either some that come out of the box from their model provider or they’ll have built their own filters that sit in between. They just don’t work very well. The difference between putting a classifier in place that, like, maybe goes and checks whether you’re giving medical advice when you shouldn’t and says, “Hey, if this looks like medical advice, filter it out.” Lots of companies have that in place. The question is whether it works. And it’s actually pretty fiddly to sit down and think about all the ways in which you could ask for medical advice, read the academic literature on what are the kinds of
Rune Kvist [00:25:03]: Framings or tricks you might play to get an AI to give you medical advice when you really shouldn’t. And so there’s, like, an area of expertise that’s just missing. So what we find is that most people have the right building blocks in place. They don’- It doesn’- It’s not rocket science, but the finicky thing is, like, getting into the corners and testing whether it works such that you can look your customers in the eye, or maybe a bank or maybe a hospital and be like, “This is going to work for you.”
Vibhu [00:25:26]: I see. So we talked a lot about the agent-level certification. Where do you guys go from here? So announcing series A camera, we talked about this a bit. There’s the whole security risk of Fable, government stepping in. You guys are kind of announcing that you’re also going into model certification?
Toward Model Certification: The Government–Lab Trust Gap
Rune Kvist [00:25:46]: When we do a bit of cutting afterwards,
Vibhu [00:25:48]: Yeah
Rune Kvist [00:25:48]: We will not yet be announcing this,
Vibhu [00:25:49]: Nice
Rune Kvist [00:25:50]: The question that is top of everyone’s minds now is at the model level. And Mythos, then Fable, has really brought this to the fore that in addition to the commercial risk and the kind of economic security risks that are happening at the agent layer, the models are going to present risk in the national security category. The shape of the problem is very similar. You have some people that are on the hook if something goes wrong. In the case of agents, it’s often security leaders in the enterprise. In this case, it’s the government. They don’- haven’t necessarily spent their entire lives thinking about what are the new risks that come here, what is the kind of data you might be looking for, how might you test that? But they do have to make sure that their concerns are addressed. You have some frontier AI companies that are deeply technical. They know a lot about the risks, but they fundamentally have an incentive to not always be truthful. So you have a trust gap between the government and the labs. And in every other industry, you end up with some kind of body sitting between, a neutral third party sitting between those people. There’s no other industry where you allow people to audit themselves. So there is going to be a need for a third party that can take the rigor of the labs to run frontier technical evals, but can also speak legible trust in the way that the government trusts PwC to go and run financial audits. And they know that they output audit reports in a way that’s consistent, that’s easy to read, that’s factual, that’s, trustworthy. Those two things need to be brought together. And what we’ve learned from our work with agents is that if you want those-- that communication between those two parties to be smooth, there has to be one common standard that is public, that people can go and inspect. What are the risks that matter? Within each of these risks, what are the kinds of threat models that you’re really looking for? You need to specify for each of those risks, what are the guardrails that need to be in place, and what are the tests they need to run to see whether those guardrails are effective? And then you need to go and run audits that are - technical audits that are consistent. So if you’re trying to bring trust, it’s extremely important that you methodically work your way through the risks. You can’t send one researcher in and say, like, “Come back with whatever you find.” You need to be able to explain exactly what you did, exactly what you tried, exactly what you did not try, and therefore the kinds of promises you can and cannot make at the end of it. I think of
Neutral Third Parties, CAISI, and Model Risk Audits
Rune Kvist [00:28:13]: Fable as a direct symptom of this problem that the government was told that there’s a risk. The government may struggle to assess just how big that risk is. They call Anthropic, and Anthropic is trying to tell them, “Hey, actually, every model can be jailbroken.”
Swyx [00:28:28]: That’s not what you want to hear, right?
Rune Kvist [00:28:32]: As the government, that might be hard to trust.
Rune Kvist [00:28:36]: And we think that a broker is the most natural solution. In other markets, you see something like, in financial markets, you see Moody’s. Moody’s goes in, and they look at a bond, and they output a rating. They say like, “Here’s the evidence we found. Here’s the rating.” We don’t decide whether anyone should buy this bond or not buy this bond. Well, that depends on their risk appetite. But we do provide this common information layer that everyone can rely on. In the case of Moody’s, the government, points to them and say, “Hey, pension funds, you should probably really take care. You shouldn’t risk your pensioners’ money, so you can only invest in triple-A rated bonds.” That means that now the government doesn’t have to staff thousands of financial technical experts to rerun forecasts every week to see whether things are correctly rated. They get to point to some neutral third party. So my hypothesis is, my hunch is that you will see a third party that sits between the government and the labs, and it could either be the government builds it themselves. So something like CAISI was set up to do exactly this. And the question
Swyx [00:29:44]: Sorry, I’m not familiar with CAISI.
Rune Kvist [00:29:45]: CAISI is the Center for AI Standards and Innovation.
Swyx [00:29:49]: Okay.
Rune Kvist [00:29:50]: I won’t get into the details, but it’s a body of NIST that typically sets standards. So it’s basically a government body that has AI experts. Yeah, exactly. Exactly.
Swyx [00:29:59]: Very key. Very key.
Rune Kvist [00:30:00]: Very key.
Vibhu [00:30:00]: I think, you know, it’s one of those things where when you just sit back and listen-- look at it, like, is there enough technical expertise in the government to measure, test these things right now? Probably not, right? And Fable is a result of, okay, we’ve had to scale back and pause things,
Rune Kvist [00:30:17]: Yeah. And they have excellent people, but they have an extraordinarily small budget compared to the scale of the challenge that’s ahead of us. And I think they have a role to play. The question is kind of like, who does what? We have now outlined the jobs to be done, and they’re quite extensive. Every model release, there is an astounding-- Given that they take in any input, their risk surface is astounding. And so the question is really: what can only the government do, and what can the market provide here that can keep up with the pace as AI risk changes? Our perspective is that also at the model layer, the risks that people care about today are not the same ones they cared about three months ago. So the pace of legislation is too slow to deal with pinpointing the risks here. And so we think there’s a lot that the market can do to surface timely information. Ultimately, there is a bunch of policy decisions here. Is the national security risks of a model too high?
Swyx [00:31:12]: Yeah.
Rune Kvist [00:31:12]: That’s a political answer. But what we want to make sure is that the process that produces this risk information is compatible with very fast innovation. So you don’t want to. This is not a question of like, can you slow the things down? Can you keep, the models locked up until-- for months on end until everyone can make a guarantee? But it is this, can you, in the time it. Given that the US is competing with China on releasing models, can you insert risk information that allows the government to, like, make rapid decisions on some of these questions? Balancing that trade-off between failing to adopt AI is going to put us at risk, but also reckless adoption is going to put us at risk. And that’s a very kind of fine balance that they’re going to need, like, a lot of high-quality intelligence to make.
Chinese Models, Data Flows, and National Security Concerns
Swyx [00:31:55]: Just a side mention, because you mentioned Chinese models, any specific concerns that you’re hearing from your CISOs about that? ‘cause I guess it’s free, but.
Rune Kvist [00:32:05]: CISOs have a bunch of concerns around data flows in general that they’re really concerned about. So there’s a lot of questions like, if these models are Chinese, where does that, where does that data go? I think a lot of this can be addressed, but they come up often.
Swyx [00:32:18]: I mean, they understand they’re running on American GPUs.
Rune Kvist [00:32:21]: Some of them, some of them understand that they’re running on American GPUs.
Swyx [00:32:23]: They’re not, like, phoning home every time you, like, call home.
Rune Kvist [00:32:26]: No. A year ago, there was not a lot of understanding of this. I actually think, you’re seeing the security leaders becoming kind of AI literate at a blistering pace, and you’re actually also seeing my Twitter timeline that’s very pilled and my LinkedIn feed that used to not at all be pilled kind of converge. They’re both talking about Fable.
Swyx [00:32:45]: Right. Yeah, that’s true.
Rune Kvist [00:32:46]: They are both talking about whether you can prevent models from being jailbroken these days.
Swyx [00:32:51]: Yeah.
Rune Kvist [00:32:52]: Like national security national security risks are now the conversation that is actually emerging. Other than that, I think you mostly see a kind of general picture: there are no concerns with any particular model or any particular model output, but there is a general nervousness of having critical infrastructure run on models that are not produced in America by Americans where the American government has control.
Swyx [00:33:14]: But it doesn’t necessarily show up in your framework that directly, or it might, I don’t know.
Rune Kvist [00:33:18]: There’s a bit of stuff in there actually on the, like, the provenance of the models and disclosing that. But I think there’s a bunch of use cases where running a Chinese open-source model is just the best solution.
Swyx [00:33:27]: Yeah.
Rune Kvist [00:33:27]: And a concern is slightly more macro here, which is not best addressed at any particular certification level.
Vibhu [00:33:32]: Is there anything interesting that you see at the. You know, if you’re trying to fill that middle gap, that mediation gap, any interesting stuff that you guys forecast would be required other than, you know, what the average person might expect?
Cyber, Child Safety, Bio Risk, and Expert Coordination
Rune Kvist [00:33:47]: There’s a bunch of interesting questions about what are the risks that matter here. So right now, the risk of the day is cyber, because it’s very real, very tangible. And some of the risks that are also emerging as pretty real and pretty tangible are things like child safety is becoming both extremely important, but also politically important. And then there are some of the risks that are coming down the pipeline that today feel kind of speculative, but people who spend a lot of time with the models see them coming down is things like, risks that relate to biology.
Rune Kvist [00:34:18]: And specifically whether models will help adversaries produce biological weapons and making that extremely cheap, extremely accessible, producing-- making the chance of another COVID or worse pandemic. COVID was not engineered to be bad, as if you were trying to do that. So I think those are some of the risks that are coming down the pipeline. I think one other thing to just note is that agents are kind of deliberately narrow. So, like, when a frontier agent company puts a chatbot that interacts with customers, they’ve really tried to narrow the topics it’s interested in talking about. Such that if you ask it, like, “What do you think of the president?” it will just decline, which means that the kind of risk area is somewhat smaller. For models, it is infinite. And so there’s not a single expert out there who can competently evaluate the risks of cyberattacks and fifteen-year-olds having month-long conversations with a chatbot and seeing whether it will in fact recommend suicide or something horrendous like that, and can evaluate the risks that terrorists can use AI to produce bioweapons. The risk surface is just too big. And so the central challenge actually becomes how do you get those subject matter experts to work within a one coherent framework that outputs one coherent report and rating that the world can go and inspect? ‘Cause that global perspective is central, but there’s not a single organization today that could produce that.
Swyx [00:35:47]: And you would be the presumptive one when you put out your model standards.
Rune Kvist [00:35:51]: We think there can be one company that can, with a consortium of experts, build one coherent standard. I think we’ve shown that across all of the enterprise risks today. We think it could be one company that could, with a consortium, specify the audit rules, basically like the inputs and outputs that all these technical experts need. What access do they need? How should they treat infosec- info security? They can look at whether the eval- evals are well-produced without necessarily being able to say, “Hey, is this a threat or not a threat?” But overall, evaluating whether the evals are good, well-constructed, that set of audit rules that basically becomes the interface for all these experts, we think one clearinghouse could put together. To be clear. When I say one company, I think of it as one company coordinating lots of this in the same way that when we saw our consortium, it’s not like we say we have all the answers on agent security. What we say is we are taking on the role of eliciting all of the concerns and being the secretary that puts it together and runs a tight house such that the standard updates lockstep every quarter, and that the audit reports that come out, in this case, 100-page audit reports, uniform and crisp and clear all to the level of detail that is required for executives that need to make a clear go/go decision. So that’s kind of the role that we think we might play.
OWASP, Frameworks, and the Operational Audit Layer
Swyx [00:37:11]: I think in many ways you’re performing the role that OWASP used to do there, and you said, like, you know, competition and partners.
Rune Kvist [00:37:18]: Yeah.
Swyx [00:37:19]: Can you go more into, like, how they partner?
Rune Kvist [00:37:20]: Yeah. So first of all, OWASP is basically an open source community of security practitioners that are coming together to build frameworks for addressing the latest security concerns. We think they are phenomenal at creating frameworks. We’- In fact, we’- First of all, we’re partners with them, so we have a joint article. Two, we’ve learned a lot from them. We think they’re a tremendous source of intelligence. What OWASP does not do is building the machine that runs third-party audits such that a company like Cursor or a company like JPMorgan could get a third party to go and review them against this and say, “Hey, you’ve passed the standard, and here is the report that you can use to build trust and preempt your partners’ or customers’ questions.” So they fundamentally try to do something different. You - They are part of the information gathering and intelligence gathering and creating clarity, but the operational layer of turning this into promises is not the business they try to be in.
Swyx [00:38:14]: The standard is emerging and is doing very well. Was it necessary to then also do underwriting? Obviously it’s in the name, so please remember you thought about it first. I feel like if you just have enough consensus, you don’t actually need the money angle, but it does help.
Vibhu [00:38:30]: I did want to also note, you guys are a profit company too, right? It’s not profit where there’s a whole business side to it as well?
Why For-Profit Standards and Insurers Matter
Rune Kvist [00:38:39]: Yeah. Yeah, so I’m just getting crazy
Swyx [00:38:41]: I think about the money part.
Rune Kvist [00:38:42]: Yeah. Yeah, let’s get into the money part. Let’s start from actually your question, profit versus profit. In the security space today, cybersecurity, most of the standards are produced by nonprofits. I think that’s an issue.
Rune Kvist [00:39:00]: The question you have to ask yourself is, how do you create good incentives for these standards to be good and keep up?
Rune Kvist [00:39:09]: Nonprofits tend to not have these adverse profit incentives where they, hollow out their standard and create a race to the bottom, but they’re also not at all responsive by default to the communities that they serve. There’s no process-- They don’t have customers that they serve where they go and ask, “What do you want? What do you want? What do you want?” And when you look at the overall satisfaction with the security standards today, people tend to just not like them very much. You do see in other domains, that profit standards can serve the world quite well. So there are examples, like we talked about Moody’s before. It’s not without flaws, but, it is absolutely critical societal infrastructure that gets run at an astounding scale today. Your credit score, it’s FICO. It’s also a profit business. And when you go back even further in history, some of the crash testing standards came out of insurance companies.
Rune Kvist [00:40:06]: The insurance companies together founded the Insurance Institute for Highway Safety because they were very interested in, like, how can we use standards to drive down mortality and save money? Go back, prior-- Our name actually pays homage to the Underwriters Laboratories, UL, which, was started right around when electricity came out. Houses started burning down. Insurers, again, were paying the bill, and they were maybe also good people, but their profit incentive was, let’s prevent houses from burning down. Let’s test all the electrical products, the light bulbs. All the light bulbs in here are probably tested, the toasters, et cetera. And they set up, an entity to create those standards. Today, UL has a profit entity and a profit entity. What they’ve recognized, they spun - They started profit. They spun out a profit because what they recognized was like, hey, actually to serve customers well, you need a profit entity. The lesson here is one of the ways that the market can align incentives so you’re both responsive to customers
Rune Kvist [00:41:07]: And not hollowing out your standard over time is to align it with insurers because they fundamentally have good incentives. And so if you’re a profit standard that works closely with insurers, you get the feedback loop in such that you’re really tuned into your customers, but also have their interest at heart. So that’s the model that we - the kind of inspirational model that we’ve learned a lot from, and that’s also where the name comes from. In some ways, the term underwriting can both be associated with insurance, but it’s also a broad term for, like, making decisions.
Rune Kvist [00:41:40]: If you underwrite a decision, you’re fundamentally kind of taking ownership for the consequences of it.
AI Insurance Contracts, Lloyd’s of London, and ElevenLabs
Swyx [00:41:45]: Yeah, I mean, what does an insurance contract look like for AI?
Rune Kvist [00:41:49]: Yeah. Most of the demand comes today for insurance contracts is, sitting between people who’ve built AI and people who are buying AI.
Swyx [00:41:56]: Yes.
Rune Kvist [00:41:57]: And what you want—the reason why people want insurers involved, both for the traditional reasons, hey, if something goes wrong, we want to be compensated, but it’s in particular because insurers can bring trust to the equation. Because insurers will pay for the damages, if they’re willing to write an insurance policy, that is them saying, “Hey, we think there is risk here, but that is manageable.” And that is kind of a. Their incentive aligns with the enterprises adopting it, so that’s a really a good signal to the market. In the same way, actually, one of the things that Waymo tried to get their first permit to even operate in San Francisco was to get a lot of insurers to stack up a huge insurance policy. In the case if something went wrong, not because Google can’t pay, but because it was very valuable to have a third party go and look at that data
Rune Kvist [00:42:47]: That are trusted by governments, trusted by enterprises as conservative people and say, “Hey, we’ve looked at it. We’re actually willing to take some of this on our balance sheet.” So that’s, that’s kind of the reason why people are interested in it. What it looks like is, in some ways like every other insurance contract. You specify what are the perils you want to cover, how much do you want to cover them, like up to what limits, and what does it cost to cover that. And in the case of, if we take a really concrete example, ElevenLabs, bought a first of its kind AI agent insurance policy. They work with some of the biggest, enterprises that work with governments. They’re really interested in going above and beyond and making promises to their customers. So they wrote a policy that covers just some of the core concerns that their customers have been asking about. And, the crucial thing was really to get Lloyd’s of London, the world’s oldest insurer, one of our partners, to look at this data and be that third party alongside us to say, “Hey, we think there’s something here that’s worth underwriting.” and that’s actually what it looks like. And so they will show that contract to their customers, and they can see how much they’re covered for. They can see what exactly it covers, and that will also probably change next year. They will want to write an insurance policy that might cover more.
Swyx [00:44:04]: When you say Lloyd’s, is it reinsurance, or are they sharing somehow at the same level or
Rune Kvist [00:44:11]: Yeah. So typically, the way, new companies get into insurance is that they partner with insurers such that the insurers take the majority or all of the financial risks. Fundamentally, if insurance is useful, because it brings trust, you have to be able to pay the bill. Lloyd’s of London is 400 years old. They’ve never not paid a claim. They’re extremely trusted. What Lloyd’s of London struggle to do on their own is to figure out which of the risks are real, what should we be looking for, what are the kinds of technical controls, and running the tests. So they use AIUC-1 as kind of the underwriting framework, and we produce a bunch of eval results that then directly feed in to inform the pricing. So this means that ElevenLabs customers know that payment will be there. They don’t have to look to our series A and see, like, do we think they have enough cash on the balance sheet? They will look at Lloyd’s.
Swyx [00:45:05]: Yeah.
Rune Kvist [00:45:05]: Yeah.
Swyx [00:45:05]: And Lloyd’s, like, famously very creative. I think I remember some headline like, they insured Jennifer Lopez’s, butt or something.
Rune Kvist [00:45:13]: Correct.
Swyx [00:45:13]: Right?
Rune Kvist [00:45:13]: And I think, was it, David Beckham’s right foot?
Swyx [00:45:16]: So, yeah. Right?
Rune Kvist [00:45:17]: And stuff like this.
Swyx [00:45:18]: So, like, clearly not a large data set.
Rune Kvist [00:45:22]: Exactly. It’s actually a remarkable institution that’s both kind of has some of the truly school virtues of having been around for a long time. They, like, really. They really operate like a trusted entity, and they have appetite to figure out the future. And I think there’s a lot of recognition that both there is, like, tremendous amount of risk in AI that is poorly understood today, so getting into this business carries real risks. But also this is where lots of the risk exposure will happen in the future. This is the one market where risk is truly growing. This is the one market that will also take out some of the existing markets. Take, like, auto insurance. When there are no human drivers, how’s that market going to look? Well, it’s clearly going to change. How are you going to assess
Swyx [00:46:08]: You want to insure Waymo?
Rune Kvist [00:46:10]: I. All I’ll say is the principles for how you insure Waymo are very similar to how you insure other kinds of AI.
Swyx [00:46:15]: Right.
Rune Kvist [00:46:15]: So again, crash testing, that’s what we do for customer share at Lovable. That will also need to happen for Waymo, which is not how you do it for human drivers. So there’s this growing awareness that the world is changing very fast, and the only way to learn how to underwrite AI is to write some policies. You may incur some losses and think of that as R&D expense, really. But the question for them is, like, who are the trustedtechnical partners they can get into this business with that can help them navigate and make sure they don’t make, kind of foolish mistakes? But also who is willing to hear the wisdom that they have? They’ve done this before. They’ve seen it was. They were there when cyber came out. So there are lots of ways in which AI feels completely new, but there’s also lots of ways in which risks look the same. And so there’s actually a tremendous amount of wisdom sitting in some folks that may have gray hair, but really have, like, a keen sense of, how to quantify risk.
Swyx [00:47:08]: Yeah. And the number is. So it’s basically like I want fifty million dollars worth of coverage against these perils, and Lloyd’s will give you a quote on it, and then you have, like, a small markup or something, and then you turn it around and do that? Is that as simple as it is?
Risk Capital, Premiums, and Working with Insurers
Rune Kvist [00:47:23]: You basically share some of that premium.
Swyx [00:47:25]: Yeah.
Rune Kvist [00:47:25]: X percent goes to the people who do the pricing of it.
Swyx [00:47:28]: You’re. It’s kind of like a. It’s kind of like a merchant bank for insurance type of thing.
Rune Kvist [00:47:33]: Exactly. You basically split the fee, and you can think of the insurance supply chain as, like, there’s bringing the capital, there is doing the pricing, and there is doing the distribution. And typically, you will pay out some X percent of premium here, Y percent of premium here, and the rest of it will go here.
Swyx [00:47:46]: Does all the insurance world work like this, or is there some point at which, like. So if right now you have equity capital
Rune Kvist [00:47:51]: Yeah.
Swyx [00:47:52]: At some point, maybe you start raising, debt or whatever, and then you have enough of a bank account and enough history, let’s say you’ve been in operation for ten years
Rune Kvist [00:48:00]: Correct.
Swyx [00:48:00]: That you don’t need Lloyd’s anymore?
Rune Kvist [00:48:02]: That’s totally an option. And I could see some worlds where that makes sense, specifically if there are risks that we feel high confidence that we’d want to insure where the incumbent insurers are too slow to find appetite
Swyx [00:48:13]: Okay.
Rune Kvist [00:48:13]: Or simply struggle to evaluate it such that they don’t want to do it. But by and large, in general, you do not want to compete with insurers on, bringing risk capital to the game for two reasons. One is that’s fundamentally a cost of capital game. They have extremely low cost of capital. Startups have high cost of capital, by and large. And two, you want to hedge your bets, and it’s very helpful then to also have a portfolio of home insurance, of car insurance. And we’re not about to become a car insurer nor a home insurer.
Rune Kvist [00:48:43]: So they have some natural advantages, which makes it much more likely that we’ll partner.
Swyx [00:48:48]: Yeah.
Rune Kvist [00:48:48]: And they bring that, the capital at scale, and we bring the technical expertise.
Swyx [00:48:51]: You’re, you’re going to work with them for a long time.
Vibhu [00:48:52]: How are the discussions with the insurers as well? So basically, they’re going off of your certification, right? They’re trusting the diligence on you that your certification is valid, you tested the right things, and they’re backing the money that, you know, you have the right testing in place. So any interesting takeaways from working with insurers?
Rune Kvist [00:49:12]: I think the maybe the first thing is they feed into the standard as well. So if there are things that they feel like they need that they’re not seeing, we are also taking that as input into the standard, because fundamentally we think a good standard is one that creates a really healthy promise ecosystem, and we think insurers are a critical part of that. And again, they are the most well-incentivized to. They see all the lost data across every. Any particular CISO knows their particular concerns. Insurers see the concerns across the entire portfolio and often have direct access to, like, what exactly happened, who was at fault, et cetera, as they do part of their forensics. So they’re actually, like, a great source of intelligence on this. One of the big takeaways from cyber insurance, which is a market that didn’t work that well, was that the insurance and the technical expertise was not married up. What our conviction is that standards have to precede insurance. Fundamentally, what everyone first and foremost want, whether you’re a CISO at JPMorgan or a CISO at Cursor or an underwriter at Lloyd’s of London syndicate, is you want to not have an incident
Rune Kvist [00:50:19]: In the first place. You want to know that the risk is well-managed, and only then does insurance start to make sense. So we’ll see the standard ecosystem basically run ahead of the insurance. And the reason why we. You asked us kind of why I also do insurance, this is kind of proving what we think a whole promise confidence infrastructure ecosystem needs to look like, and we think it’s very compelling to bring that to life, even if we think the standard is kind of the core linchpin that unlocks the rest.
Claims, Liability, Air Canada, and Duty of Care
Swyx [00:50:44]: There’s been no claims yet, right?
Rune Kvist [00:50:45]: Nope.
Swyx [00:50:46]: This is one of those things where, you know, if people haven’t really worked through what it means to cover things.
Rune Kvist [00:50:52]: Yeah.
Swyx [00:50:52]: So for example, I pay Cursor $20 a month.
Rune Kvist [00:50:55]: Yep.
Swyx [00:50:56]: And I write a vibe code something that makes, a plane crash, causing $200 million worth of damage.
Rune Kvist [00:51:02]: Yes.
Swyx [00:51:02]: Do I claim $20 or do I claim two hundred million?
Rune Kvist [00:51:07]: Yeah. So these are all great questions.
Rune Kvist [00:51:10]: And fortunately, kind of all of insurance and legal history kind of helps answer some of those questions. I think the first thing is people have limits on their policy. So if you want to claim $200 million, you have to. Someone has to have paid a lot for that insurance policy upfront to have $200 million of coverage. And ultimately, the way this works is that, you start from a lot of uncertainty. This is not just an insurance, but also, like, can you use. Can Anthropic use books on the internet to train up? Well, they can go and look at precedent, they can But ultimately, this- these things get settled in court, and you hammer it out over time. So you start from this, like, place of ambiguity, which is both why insurance can be hard to do early on, but it’s also why people want insurance, because that ambiguity slows down adoption.
Swyx [00:51:57]: Yeah.
Rune Kvist [00:51:57]: That also sits at the heads of the,
Swyx [00:51:59]: Yeah. In some ways, actually, the first incident will help to, establish a lot of this.
Rune Kvist [00:52:05]: Exactly. And there have been a number of incidents out there that have just not been covered by insurance.
Swyx [00:52:09]: Yes.
Rune Kvist [00:52:09]: Take the now old, example from Air Canada, where
Swyx [00:52:14]: I was going to bring that up
Rune Kvist [00:52:15]: Chatbot hallucinated a refund policy, and the question was, Air Canada in that case were like, “Hey, we have nothing to do with this. This chatbot messed up, but, like, sorry.” And the courts were like, “No, if you put your chatbots to interact with your customers, they make legally binding promises on your behalf.” That is now precedent for everything in the future where you will. If someone were to deploy a chatbot like that again, they should not expect to be able to just pawn off and say, “Sorry, my chatbot lied. It’s nothing to do with me. I bought it from OpenAI.” No, if you’re putting this in front of your customers, you are taking responsibility for it. And so every court case, whether insurance is involved or not, clarifies liability, and liability is kind of the foundation for insurance. There’s another reason why standards and insurance come together. Liability for. I’ll go on a little tangent here
Swyx [00:53:06]: Please
Rune Kvist [00:53:06]: Get into the weeds of it.
Swyx [00:53:06]: Please.
Rune Kvist [00:53:07]: Liability, often one of the core concepts is whether someone was negligent. Should they have seen this? Should they have prevented this? And the question is: how do you judge that? Well, you basically judge whether they’ve met their duty of care. What does that mean in practice? Well, often they look to standards. So if there’s a standard that is broadly adopted that says you must have a groundedness filter or you must have a jailbreak filter, it becomes way harder to claim ignorance that these things existed. And so setting standards help clarify liability. Coins-- courts will often point to standards and being like, “Well, this seems like best practice to do.” It’s there for everyone to see. So there’s another way in which, like, standards are kind of civilization infrastructure that insurance can then build on, which promises can then build on.
Swyx [00:53:53]: I totally get that. We don’t have to get certified to write these, to, you know, make these, like, bots and all these.
Rune Kvist [00:54:00]: Correct.
Swyx [00:54:00]: But, like, basically, whenever we get. Go for the audit, I think people, like, start to shape up and all this stuff. I wonder if, like, that means that you don’t also then become, like, the approving authority for me to ship to production. You know, like, yes, you check once per quarter. I want to ship once a day.
Shipping to Production: Ongoing Testing and Trust
Rune Kvist [00:54:19]: Yeah.
Swyx [00:54:19]: And I don’t know when one of my things breaks, like one of your certifications or not.
Rune Kvist [00:54:24]: So there’s a couple things. There’s a couple of requirements in there that relate to how do you yourself, where you have to tell your customers
Swyx [00:54:33]: It’s like an ongoing monitoring.
Rune Kvist [00:54:34]: How are you yourself testing before you make at least major releases? We don’t go and audit people every day, but at least there is now a trail where if you do a major mess up, then your customer may come and ask you, “Hey, you promised me that you were going to run these evals yourself.” And for lots of them, most of the. PRs that people merge will not fundamentally alter the product experience, but some of them will. Thank you.
Swyx [00:54:57]: And sometimes you don’t know.
Rune Kvist [00:54:58]: And sometimes you don’t know. There are inherent risks that everyone knows that when they buy software, there can be bugs, and this is just part of it. But if you’re selling to mom-and-pop shops, they may not care. They’re just like, “Well, I want to use your tool, so I’m just going to willing-- be willing to take that risk on.” If you’re selling to a big bank, they might be like, “Sorry, we’re making promises to our customers. If you can’t make a promise to us that we can pass on, we don’t want to work with you.” Then it’s up to you to say, “Do I care for my agent to get used as critical infrastructure in this mission? If so, at least I can make promises about what processes I run, and then we can go and test it every quarter to be like, well, does it seem like, it’s still, that it still meets the standard.” So from my perspective, it’s kind of a way to. Big companies by default kind of have some amount of trust when they ship AI.
Rune Kvist [00:55:49]: If you’re a young company, if you’re just starting out, by default you have no trust. And there are very few places where you can go and get trust. So one of the things that most of our customers did before they started working with us is that they would make their own security blog posts. That’s great. But also, who’s going to trust you saying, “We’re so secure”?
Rune Kvist [00:56:05]: Like, anyone can write that. But it’s very hard. Where do you go and get that trust?
Vibhu [00:56:08]: Yeah.
Rune Kvist [00:56:08]: And so I think making the standards more legible makes it easier for smaller companies to prove that they’re doing what they ought to be doing, because the default assumption is that it’s the Wild West.
Vibhu [00:56:21]: Is there a roadmap you have of, like. There’s a lot of work to be done here, right?
Rune Kvist [00:56:25]: Yep.
Vibhu [00:56:25]: This is the first one.
Rune Kvist [00:56:26]: Yeah.
Vibhu [00:56:26]: Anything on the roadmap of what you see is next, what’s coming, what’s, what’s missing?
The Roadmap: Agents, Models, Robotics, and World Models
Rune Kvist [00:56:32]: I think when we zoom out, AIUC-1 deals with agents. Next up, we will deal with models. Next up from that, we will deal with robotics, of which, in some ways, Waymo is the first robot. But the exact same problem is going to be someone’s going to develop a robot, someone’s going to need some promises, they’re going to struggle to make the promises. And - You see this playing out when, like, if you think Fable concerns are bad, like, see when Waymo hits a dog. And that’s if people lose their mind. Imagine when first robot knocks off a toddler off a kitchen table.
Swyx [00:57:03]: Yeah.
Rune Kvist [00:57:03]: You’re going to see some real strict liability.
Vibhu [00:57:07]: I mean, you could see it, right? Like, Cruise got fully
Rune Kvist [00:57:10]: Destroyed.
Vibhu [00:57:10]: All permits are gone, yeah. Yeah.
Rune Kvist [00:57:12]: Correct. So physical AI, the level of stringency just goes up and up. So that’s kind of like the big picture. Agents, models, robotics. I think within agents, the current set of agents are well-covered by this. But as the technology progresses, as agents get longer horizons, new types of failure modes will emerge. And so it’s mostly of can you make sure the standard keeps up when they appear? And you also start to see new modalities. Like today, world models are mostly a kind of a research question. There’s no one who’s really using it. But that will also bring in just new kinds of ways to create value, but also more risk surface that no one knows how to grapple with today. You’ll start to see true agent interactions that are not mediated by humans. There’s going to be a bunch of interesting questions. You’re basically going to need a new legal system. How do they build trust amongst each other? How. One of the core things when humans trade with each other is that you know that you have recourse. You can sue them. How do you make sure that there is a persistent balance sheet behind any agent such that if you trade with it and it screws you know you can get your money back? Those are some of the questions we’re going to have to deal with. And the technical testing
Rune Kvist [00:58:24]: Of multi-agent systems is also going to be interesting and complex.
Swyx [00:58:29]: Very fun. Are there any perils that are uninsurable right now that people wish that you would?
Copyright Risk, Adverse Selection, and Information Asymmetry
Rune Kvist [00:58:35]: Yeah. One of the places where there’s a bunch of appetite for insurance and not a lot - a lot of demand, but not a lot of supply, is when it comes to copyright.
Swyx [00:58:46]: Oof.
Rune Kvist [00:58:47]: In some ways, copyright is kind of mundane. It’s always been an issue. There’s a couple of reasons for this. The first is people who have trained on copyrighted materials almost always know that they’ve done that.
Rune Kvist [00:58:59]: So if you want to buy insurance for it probably signals that you might be a high-risk customer. The people who are most interested in getting insurance for copyright infringement
Swyx [00:59:09]: Okay. Yeah
Rune Kvist [00:59:09]: Are the people who are most likely to have copyrighted
Swyx [00:59:10]: Yeah. It’s like a, it’s like a lemon problem.
Rune Kvist [00:59:13]: Exactly.
Vibhu [00:59:13]: I actually think there’s another side to it too, right? Like, if you’re building on something. So say I’m using an open model.
Rune Kvist [00:59:19]: Yeah.
Vibhu [00:59:19]: I don’t know what it’s trained on, right?
Rune Kvist [00:59:21]: Yes.
Vibhu [00:59:21]: And how far down that chain does copyright go?
Rune Kvist [00:59:24]: Yes.
Vibhu [00:59:24]: Am I liable to take down my product because company X trained on copyright?
Swyx [00:59:29]: But there’s safety in numbers. If everyone’s doing it, then you.
Rune Kvist [00:59:33]: Correct.
Vibhu [00:59:33]: I mean, I would say until, you know, Fable is rolled back from everyone that used it, right?
Rune Kvist [00:59:38]: Yeah. I think it’s a hard question. I don’t know the answer to it.
Vibhu [00:59:39]: It is.
Rune Kvist [00:59:39]: But I think your intuition is, your intuition is right in kind of like, what is the kind of duty of care?
Rune Kvist [00:59:47]: And people don’t today think of it as customary that you go and you, like, dissect the open model’s training data and you check everything. In fact, lots of people use them. It’s seen as kind of generally acceptable to not check for this. And therefore, like, we’re not going to hold you to specific
Vibhu [01:00:02]: I mean, we also really can’t, right? We don’
Rune Kvist [01:00:04]: Exactly.
Vibhu [01:00:04]: We don’t know the training data.
Rune Kvist [01:00:05]: So you can then ban it, but I think no court is going to get a copyright question and be like, “This actually needs to get banned.”
Swyx [01:00:09]: Unless you hire Nicholas Carlini and he can extract it for you.
Rune Kvist [01:00:12]: Exactly. Though he’s in short supply.
Swyx [01:00:15]: Yeah. He’- You only have so many Carlinis, but,
Rune Kvist [01:00:17]: Exactly.
Swyx [01:00:18]: Yeah, go ahead.
Rune Kvist [01:00:19]: So I think this is also fair that, in the case of labs, there’s a lot of interest for this. But the thing that makes lab want it is what makes this insurer suspicious of it, and so you have a lemon’s problem.
Swyx [01:00:30]: Yeah. Is there, like, a theory of insurance where adverse selection dominates the risk-sharing aspect of insurance? Like, where does this. Like, teach us insurance.
Rune Kvist [01:00:40]: A lot of insurance does come back to, like, practical versions of microeconomics 101.
Swyx [01:00:45]: Yeah. It’s very. It’s like, it’s like this is why
Vibhu [01:00:47]: High-risk adverse.
Swyx [01:00:48]: You need to pool health insurance, because if you make it too hyper-specific, then only people who are guaranteed to get the disease will sign up for your insurance.
Rune Kvist [01:00:56]: Exactly.
Swyx [01:00:56]: Same thing.
Rune Kvist [01:00:57]: The core problem is one of information asymmetry. People buying insurance know something about their risk that insurers do not know. And so the question is actually. And this comes back to the same problem is, if you rely. You can break a lot of these information asymmetries if there is. Some kind of testing that reveals the underlying true risk. And so if you were able to, in the case you mentioned, have good diagnosis of whether someone has it or what the probability is that someone has it, that the insurers trust, then they might be willing to insure it. But if they don’t, if there’s no kind of common information, then - only the patient will know
Vibhu [01:01:32]: Yeah.
Rune Kvist [01:01:32]: That’s what breaks it down. So the question is, again, how do you create credible signaling between players?
Rune Kvist [01:01:39]: This is also the whole reason why Moody’s exists. Moody’s just does credible signaling. That’s also why Moody’s could never-- Moody’s has to be independent. If Moody’s was owned by JPMorgan, then JPMorgan cannot use it as a signaling mechanism. So a lot of the basics of standards and certification are just communication devices. It’s just a trust gap. And, that’s where you have to think about what are the incentives of the messenger. And one and another way you can break a lot of this is through transparency. If you are transparent in how you operate, you just cannot mess with others nearly as easily. You make it much more costly, and that increases trust. This is one of the reasons why there’s a change log here.
Rune Kvist [01:02:16]: Every little change
Swyx [01:02:18]: Yeah
Rune Kvist [01:02:18]: You can go back and find, and it means that if we were to make the standard worse
Swyx [01:02:24]: Oh, wow, that’s a lot of changes in one update.
Rune Kvist [01:02:27]: Yeah.
Swyx [01:02:27]: Okay.
Rune Kvist [01:02:28]: And a lot of this is just as things get clearer, you can see a lot of clarifications, you can see some revisions. As things get hammered out, you want to change this. But if you make it all public, you make it much harder to mess with people, or at least you become found out very easily.
Rune Kvist [01:02:42]: And so this is a way of reducing the information asymmetries by just making more of the information public.
Vibhu [01:02:50]: I like how you do know when future versions are coming.
Swyx [01:02:52]: Yeah.
Vibhu [01:02:53]: So I guess it’s quarterly.
Swyx [01:02:53]: I mean, they just
Rune Kvist [01:02:54]: It’s quarterly.
Vibhu [01:02:54]: Yeah.
Swyx [01:02:54]: It’s kind of quarterly.
Rune Kvist [01:02:55]: Yeah.
Swyx [01:02:55]: Not that surprising.
Rune Kvist [01:02:58]: Yeah, but this is also a promise. Like, if we now don’t deliver on July 15, basically
Swyx [01:03:03]: I mean, you can just batch it up, and then whatever you got, you just ship it.
Rune Kvist [01:03:05]: You just batch it up.
Swyx [01:03:05]: Yeah. That’s not that hard.
Rune Kvist [01:03:06]: But it’s kind of like we deposit some amount of trust every time we meet this commitment.
Vibhu [01:03:12]: Yeah.
Rune Kvist [01:03:12]: And in the startup land, it feels easy to ship a new version of a standard once a quarter. In the enterprises who are used to this, like, decade-long cycle, we often get met with, like, incredulity. Like, there’s just no way. And then you show them the change log.
Swyx [01:03:28]: One thing I wanted to also, like, try to really think about is, you know, you said something about how if you have tests for the thing, then you can insure it.
Rune Kvist [01:03:35]: Yes.
Swyx [01:03:36]: Right? And so really what your standard is, what AIUC is, is establishing a framework for the audits to happen so that you can at least test, like, all these, like, baseline standards of care have been met, and therefore people can insure against standard risks that everyone has. I wonder if, like, there needs to be develo-- you need to develop other tests. We’ve covered mech interp in the past. Any interest in that, or are there other kinds of tests that we’re not thinking about?
mech interp, Eval Awareness, and Monitoring
Rune Kvist [01:04:02]: Yeah, I think mech interp is a big one. A lot of interest in that. I think everyone would agree that there’s, like, promising scientific potential.
Rune Kvist [01:04:15]: We’re still a while, a little bit away at least, from this being, like, commercially available on demand such that there’s, like, now a selection of vendors you can go to.
Swyx [01:04:26]: Goodfire would say that it is commercially available.
Rune Kvist [01:04:28]: Exactly.
Swyx [01:04:29]: And it just
Rune Kvist [01:04:30]: We would agree with them. We think that the work that they’re doing is tremendous.
Swyx [01:04:33]: Yeah.
Rune Kvist [01:04:33]: We’re not quite at a point where we could literally require it. But it’s the kind of thing where you can imagine relatively soon you could put in an optional control for if people use mech interp as a way to reduce risk, you at least get credit for it. We can’t require it because it’s going to be hard to require everyone to become Goodfire customers.
Swyx [01:04:49]: What good does credit do me? This is - this is a pass-fail, right? Do I care about credit?
Rune Kvist [01:04:54]: It’s a pass-fail, but it’s also a 100-page audit report
Swyx [01:04:57]: Huh
Rune Kvist [01:04:57]: That you’d be surprised at how much security leaders actually sit down and digest this stuff.
Swyx [01:05:02]: Okay.
Rune Kvist [01:05:02]: And I promise you that if someone is using mech interp today they will have a slide on it because they’ll try and get credit for it.
Swyx [01:05:11]: It is cool. It’s fancy, yeah.
Rune Kvist [01:05:12]: But it’s just easier if you have a third party saying, “Yep, they have mech interp, and actually.”
Swyx [01:05:16]: Just to spell it out for people who have been following our mech interp podcast
Rune Kvist [01:05:21]: Yeah.
Swyx [01:05:21]: It is literally like, oh, you’re using, you know, OSS. It is activating these three dangerous things. We monitor for it, and we log it out in whatever tool of choice. Gray Swan has, like, Signal or whatever, and that’s it. That’s the mech interp-based activation, signal. Okay.
Rune Kvist [01:05:38]: Yeah. So I think mech interp is interesting, and I think if that promise truly comes to fruition, you can make stronger promises than you can with evals. And so I think that’s very compelling. Another thing that I think will become increasingly important is just kind of good school monitoring, and slightly after the fact. One of the things you’re seeing with eval, some of the challenges that are emerging is that the agents are starting to become aware that they’re being evaluated.
Swyx [01:06:04]: Yeah, eval awareness.
Vibhu [01:06:05]: Yep.
Rune Kvist [01:06:05]: Exactly, which is a problem. It means that they basically, if they know they’re being watched, they won’t do the thing that they think they get punished for. And by default, unless you know how to kind of reduce eval awareness, you should trust evals less. And one of the kind of truest things, monitoring, like, is the source of truth. Did you in fact give medical advice, and how quickly do you know? How often - have you done that in the past? How fast do you respond? How often do you detect it? How fast do you detect this? So I think that is also a paradigm. It’s slightly more intrusive. You actually will look at some customer data, but I think will become more prevalent over time.
Swyx [01:06:45]: People talk about this like we should not write about eval awareness because it’s going to leak into the data set and then be. Like, we should just. Like, we should, like, never talk about it, only meet in person and, like, talk offline unrecorded. Like.
Rune Kvist [01:06:57]: Did you guys see the Anthropic research where. I think this was literally Anthropic did that test.
Swyx [01:07:04]: What?
Rune Kvist [01:07:05]: It took. I can’t remember the details here, but they, ran some studies on misalignment, and then they took out the training data- That related to LessWrong discussing misalignment, and they ran the same test again and the failure rate went down.
Rune Kvist [01:07:20]: So it, in fact, was some evidence pointing towards it had learned the - either the ability or the propensity to do that.
Swyx [01:07:28]: Yeah, I mean, so there’s the hyperstition effect, and then there’s, like, the Luigi/Waluigi effect.
Rune Kvist [01:07:31]: Correct.
Swyx [01:07:32]: Which is like you are. The more you try to train for it, you create the opposite.
Rune Kvist [01:07:36]: Yes, there you go. That’s exactly it.
Swyx [01:07:38]: In some ways, I think the very success with Anthropic is a result of hyperstition, like the fact that you wanted this thing to exist in the world, and now it does. But, like, then it also creates the opposite as well.
Rune Kvist [01:07:48]: Yes.
Swyx [01:07:49]: Like, I think people who are maybe newer to this space don’t remember Waluigi, but, like, I do think it’s very important for understanding that when you train for a thing, you also train the opposite of the thing ‘cause it’s just a big flip.
Rune Kvist [01:08:02]: Yes.
Rune Kvist [01:08:03]: Yes.
Vibhu [01:08:04]: I think, you know, just going back to where we were at, like, there’s a lot more than just mech interp that there’s value in just having added, right? So your version of how fast can you measure stuff? Do you have logging? Do you have evals? You know, do you see other parts of the stack, like the inference providers that you use, the services? Okay, am I using Chinese model on their home API? Am I using through certified vendor here? Am I hosting myself? What am I doing on the inference engine side? There’s just, like, so many levels of stuff that gives, you know, information that you can standardize out, right?
Managed Agents, Enterprise Controls, and Generative Media
Rune Kvist [01:08:36]: Yeah. And you also see increasingly, in addition to just the basic chatbots, you’re increasingly seeing big companies adopting agent platforms where they’re building on top of Google’s Agent Studio, et cetera that comes with a bunch of, like
Vibhu [01:08:52]: Managed agents.
Swyx [01:08:53]: Managed agents.
Vibhu [01:08:53]: It’s everywhere now.
Swyx [01:08:54]: Everyone has managed agents.
Rune Kvist [01:08:55]: Exactly.
Vibhu [01:08:56]: And there’s even levels. You can host your own managed agents, OpenAI’s Agent SDK, or hosted by Anthropic, or Google does both.
Rune Kvist [01:09:03]: Correct. And then these are just ways to kind of strengthen the security guarantees you can make. And in some ways, it’s kind of bread and butter enterprise security. They. Like, they love to host things on their own premises because it gives them really a sense of control. And I think you’ll, you’ll see, just like you do in every other enterprise market, if you really sell to the enterprise, you start to compete on some of these security features. And this is also happening in AI, unsurprisingly. And I think you are seeing some amount of enterprises wanting. Enterprises are really grappling with the thing that makes agents useful is that they’re stochastic, and the thing that makes them really hard to adopt is that they’re stochastic, and these are in tension.
Rune Kvist [01:09:46]: Leaders come out on different sides of that table, in part depending on how much the CEO is trying to get the stock price to go up by saying they’re AI native and that we must be willing to take the risks. We see, we actually see phenomenal tension in the heads of the CISOs of the Fortune 1000, where on the one hand you have the CEO saying, “We must adopt, otherwise we’re becoming irrelevant, and if we f**k up, you’re fired.”
Swyx [01:10:08]: Oof.
Rune Kvist [01:10:08]: And that’s kind of like the core emotional tension that we see showing up again and again. And one of the core problems that we solve for them is to take that abstract emotional concern and turn it into a framework, in some ways just providing clarity to that concern.
Vibhu [01:10:23]: So anything in here. So something I think we kind of skipped over. We talked a lot about agent language models, skipped over world models.
Rune Kvist [01:10:31]: Yeah.
Vibhu [01:10:31]: You guys have voice, which is interesting with ElevenLabs.
Rune Kvist [01:10:34]: Yeah.
Vibhu [01:10:34]: How about generative media? So, you know, generating images, videos, that’s a category that actually has a lot of usage. Is there anything in your current policy? Is it separate policy? How do you see that space?
Rune Kvist [01:10:46]: Yeah.
Vibhu [01:10:46]: It’s like we did talk a bit about copyright,
Swyx [01:10:49]: Music.
Vibhu [01:10:50]: Yeah, music as well.
Rune Kvist [01:10:51]: Yeah. I think a lot of the concerns that come up there either relate to, copyright or there’s a lot related to, let’s call it broadly safety. So, like, this could be not safe for work or just very graphic materials, are kind of some of the core things. We have done some work on this. There’s a little bit in the standard as well that deals explicitly with that. Video, we have not done a lot in yet. And I think for proper production, that has still. Especially proper production without a human in the loop, that’s still got some ways to go. It’s obvious that it’s coming, but it’s very rare that it’s like shot deploy a video to the internet. But eventually that will also happen.
Vibhu [01:11:33]: We see, like, you know, Luma has Luma agent where it’s still pretty human in the loop.
Rune Kvist [01:11:37]: Yeah.
Vibhu [01:11:37]: So it’s not just
Rune Kvist [01:11:37]: And that just makes complete sense as the technology matures, and over time, it will become so good that people will not want to slow things down by having a human in the loop. And then, the need to make promises will grow.
Swyx [01:11:52]: Why not just have prediction markets on everything?
Prediction Markets vs. Audits
Swyx [01:11:55]: Right? It’s very EA adjacent.
Rune Kvist [01:11:56]: Yes. The core thing is that the people. Prediction markets rely on public information. There is not a lot of public information. It’s just insiders trading on each side.
Swyx [01:12:06]: Yeah.
Rune Kvist [01:12:09]: That’s illegal.
Vibhu [01:12:10]: There’s leaked information.
Rune Kvist [01:12:12]: There is leaked information. The core challenge is that often you have private sensitive information, and you need to convey confidence and trust around that. And you can, of course, for some claims, like can any model be jailbroken, you could rely on public evidence ‘cause there would be lots of people being like, “Well, there’s tons of studies, and actually they all can, so that resolves fine.” I think that’s good. For, hey, this new unreleased Methus model, how capable is it actually?
Rune Kvist [01:12:43]: Prediction markets have not a lot to say because actually just no one knows. And so I think that’s the core place where some of this breaks down, is that actually lots of the world’s information that guides some of these high-level decision is private and often also just not known.
Vibhu [01:12:56]: I think the thing with prediction markets that people like is it’s not, it’s not answering the broad question. It’s a specific, right? So will a model do this by this date, or is a model capable to do this by then, right?
Rune Kvist [01:13:07]: Yes.
Vibhu [01:13:08]: That’s a little distinction there.
Rune Kvist [01:13:10]: Yeah. And often the most interesting question, if you are, say, the head of security at a bank. The question you’re really trying to answer is, will this product, this agent, do this bad thing that maybe primarily I care about, specifically in the setting that I care about? And the question is like, what’s the closest-- That information may not exist anywhere. So prediction markets aggregate existing information. This information may not exist, and you want some very specific and you’re willing to pay for it. That’s kind of where a third-party audit comes in. We also don’t really use prediction markets to figure out whether, public companies have committed fraud in their books. You use audits. You probably could, but the information’s just not that available. And if so, it would be like just trading on vibes. Actually it would have been really interesting to see whether prediction markets two thousand and one were predicted Enron going bankrupt and they kind of
Swyx [01:14:02]: Yeah.
Rune Kvist [01:14:02]: Could you have told-- could you have sensed from like the craziness of the CEO or some other traits that they were more likely to cook their books than others?
Swyx [01:14:10]: Or enough insiders leak it then that
Rune Kvist [01:14:12]: That could also be right.
Swyx [01:14:13]: Right. Which is like, I mean, this-- that’s the sort of the ideal dream of prediction markets. You have liquid markets and everything.
Rune Kvist [01:14:20]: Yeah.
Swyx [01:14:20]: And then you can compose your exact set of risks to offset.
Rune Kvist [01:14:24]: Yes.
Swyx [01:14:25]: Right?
Rune Kvist [01:14:25]: Yes. Yeah. And I think, like, prediction markets will bring lots of new information to it. So the thing is mostly not like which one is it, and more like what are the types of questions that prediction markets are really good
Swyx [01:14:37]: Yeah.
Rune Kvist [01:14:37]: And what are the ones where the information doesn’t even exist for insiders such that no one can in fact trade on it and it needs to get generated.
AI Engineer Certification and Training
Swyx [01:14:43]: Okay, one self-serving question and then one open-ended one, on like the future of AIUC. Self-serving question would be, so you have your standard, right?
Rune Kvist [01:14:52]: Yes.
Swyx [01:14:52]: I run, you know, a large AI engineer conference. Like, there’s been a lot of talk about us certifying AI engineers.
Rune Kvist [01:14:58]: Yep.
Swyx [01:14:59]: Training programs, level one, level two, level three. I was a CFA myself, so I know what-- that’s what the finance industry does.
Rune Kvist [01:15:04]: Yes.
Swyx [01:15:05]: Would it help if I had AI engineer level one, level two, level three, and then it would-- they would, like, work with these guys? I don’t know.
Rune Kvist [01:15:12]: If you think of the highest level objective as, like, accelerating secure deployment of agents, then that would totally help. Because one of the things that happens often now is that folks build agents, they bring it to the decision-maker, and the decision-maker surfaces a bunch of security considerations that they had not thought of, and now it’s not built to spec. Now you have to go and - like, add these filters, et cetera. So if you shifted that left, like if everyone knew what the spec they were building to, if everyone knew the grading scheme
Swyx [01:15:41]: Yeah.
Rune Kvist [01:15:42]: That would be awesome if they were already trained. So by default
Swyx [01:15:44]: But you’re the grading scheme, right?
Rune Kvist [01:15:45]: Say again.
Swyx [01:15:45]: I don’t get to set the grading. You guys, you set the grading scheme.
Rune Kvist [01:15:47]: We set the grading scheme. And I think what’s, valuable is, like, if you can turn those into
Swyx [01:15:52]: Training programs.
Rune Kvist [01:15:53]: Training programs
Swyx [01:15:54]: Yeah.
Rune Kvist [01:15:54]: Such that people
Swyx [01:15:54]: Which you’re, you’re not doing.
Rune Kvist [01:15:55]: We’re not doing that.
Swyx [01:15:56]: Yeah.
Rune Kvist [01:15:56]: I think there’s value in doing it.
Vibhu [01:15:57]: There are others doing. I mean, not to interrupt, but you know
Swyx [01:16:00]: Yeah.
Vibhu [01:16:00]: OpenAI has their
Swyx [01:16:02]: Anthropic also has like a CCTA thing.
Vibhu [01:16:04]: Yeah. You know, they want hundred thousand deployed certified consultants, right?
Rune Kvist [01:16:09]: I really think it’s good for. We will accelerate adoption if we have more people who know how to build secure agents, and we are not working on the side of training people at the moment. I think it’s, like, very aligned with our mission. We only have so much, attention.
Swyx [01:16:24]: I’ll tell you why I haven’t done it.
Rune Kvist [01:16:26]: Yeah.
Swyx [01:16:26]: It’s not like I haven’t thought about it before.
Rune Kvist [01:16:28]: Yes.
Swyx [01:16:28]: It’s just being prescriptive
Rune Kvist [01:16:30]: Right.
Swyx [01:16:31]: About like, well, this is what you should know, therefore, like, the stuff that I didn’t include is what you don’t need to know.
Rune Kvist [01:16:35]: Yes.
Swyx [01:16:36]: And I’m like, “That sucks.” Like.
Rune Kvist [01:16:37]: Yes. Yeah.
Vibhu [01:16:39]: But I think it’s like, you know, the very interesting defensible thing you guys do is your opinionated 100-page report of here’s what matters, right? Here’s the, like, prescriptive definition of the requirements you need to be certified, so.
Rune Kvist [01:16:55]: Yeah, and I think that’s a choice. I think basically that’s a, that’s a choice, and I think that serves some audiences very well, where if you’re trying to deploy this into a bank or a hospital, et cetera, clarity of the - those boundaries is extremely valuable.
Rune Kvist [01:17:09]: There’s lots of other settings where being much more experimental, much more trying it out is just the better fit. And so to me, this makes a ton of sense. Also, you’d have to rewrite your curricula every freaking three months.
Swyx [01:17:21]: It’s fine. I do that. Like, it’s okay. But yeah, no, for me, it’s actually - like, genuinely, like, the consequences of getting it wrong and, like, affecting somebody’s career is a big responsibility.
Rune Kvist [01:17:35]: Yeah. Like, I think that’s exactly right. And I think a lot of our work actually goes like, we don’t want to carry. We also don’t think of ourselves as able to carry the, kind of the true north of what’s, like, secure or not secure, but we can coordinate the forum where you listed all of that.
Swyx [01:17:52]: Yeah. Your consortium is fantastic.
Vibhu [01:17:54]: Do you think this can be crowdsourced in a way? Like, for your example, for what is AI engineer certification, right? This is a pretty big podcast. There’s a lot of takes that people can have and, you know, discussions that can.
Swyx [01:18:05]: And people reasonably disagree. So who am I to say, like, that’s a correct question, that’s a wrong question?
Rune Kvist [01:18:09]: Yeah.
Swyx [01:18:09]: Right? So, like, I don’t know.
Vibhu [01:18:10]: We’ll have an exit.
Rune Kvist [01:18:13]: Yeah.
Vibhu [01:18:14]: Vent your frustration to someone that’s listening, you know?
Rune Kvist [01:18:16]: Exactly. And I think there’s also you. Or it matters a lot what the promise is. So if the promise is, “Hey, if you’ve taken my course, you will not f**k up,” you can’t make that promise, clearly. You could make a promise of like, “Here’s the. Some important things that everyone should at least know,” and then you have to fill out the rest there. At least the promise changes. Of course, there’s some subtlety in how do you communicate this such that people really get it. But I think it’s important to dial in, and we have a section in our center on, like, what is the promise and what is the promise not, because it’s impossible to guarantee that nothing will go wrong. If you need a guarantee that nothing will go wrong, you cannot work with frontier AI, but you can make some claims.
Swyx [01:18:56]: Yeah, for sure. Cool. Wanted to end with open-ended, where is AIUC going? I think you talked about model stuff, robotic stuff. And just open-ended, like, where, you know, what is in the future for you guys?
AIUC’s Future, Hiring, and Universal Red Teaming
Rune Kvist [01:19:10]: Very near term, we’ve now started to work with some of the frontier companies in each of the categories that are taking off, and we’ll, we’ll continue that work to make sure that we cover all of the use cases that are really taking off. We see a lot of interest once the first one in the market moves. Lots of people want to follow them. And we think basically AIUC-1 will get to a point where all of the Fortune 1000 will organize their risk processes around the standard.
Swyx [01:19:38]: And you have 50%?
Rune Kvist [01:19:39]: No, we do not have 50% today.
Swyx [01:19:41]: Oh.
Rune Kvist [01:19:41]: I think there is some world where probably by end of year, we might have representation in our consortium for 50% of the Fortune 1000.
Swyx [01:19:48]: I see. Got it.
Rune Kvist [01:19:49]: So that’s on the agent layer. And then we think, yeah, the model layer, it’s going to be. It just brings. Are now surfacing the concerns that are most likely to slow down adoption of AI. And then, yeah, we think robotics comes after that.
Swyx [01:20:03]: What are you hiring for? What’s hard to hire for?
Rune Kvist [01:20:05]: We are hiring, across the board, across market and numbers of technical staff. The people who do really well on our technical team are folks who are really excited about kind of being truly full stack. So let’s say when we started working with Cursor, we’d never done coding, tools before. So taking the standard and extending it, fleshing out what does frontier evals look like for long horizon coding agents, and taking that problem all the way from, like, working with Cursor and other folks in this space down to, like, fleshing out and shaping a new version of the standard. So that’s like a truly a full-stack, entrepreneurial technical people do extremely well at AIUC. The hard part is building one universal red-teamer that works across from Harvey to Cursor and everywhere in between that both has one consistent methodology, one consistent taxonomy of what are the risks and the attacks, and making. We think that’s fundamentally the best way to make consistent promises. JPMorgan is buying both. They want to have one framework, one consistent way that this comes out, and the mechanics of making that happen, you get to deal with a lot of the complexity of the real world. I think we have good answers in a bunch of that, but there are some pretty hard engineering problems in executing that.
Swyx [01:21:18]: Can I push a little bit? Like, must you have one? Why not just be like, “Okay, look, forty percent of our use cases are coding agents, so we will specialize in coding agents,” and that’s the, that’s the one of them.
Rune Kvist [01:21:28]: Yes.
Swyx [01:21:29]: And then, okay, thirty percent is like RAG.
Rune Kvist [01:21:31]: Yes.
Swyx [01:21:31]: Just do RAG.
Rune Kvist [01:21:32]: Yes. I think there’s some wisdom in that question.
Swyx [01:21:36]: Yeah.
Rune Kvist [01:21:37]: It depends on. What we found that there’s a lot of value on is being able to. If the decision-maker on the buying side, let’s say you’re the head of risk at a bank and your biggest risk is not in coding or in customer support or whatever the top two biggest use cases, but it’s somewhere else, you want to still make sure that framework has something to say about it to the burning question you have. Otherwise, you’ll not earn that trust. Now, it’s true that a lot of the burning questions follow where there’s a lot of adoption. And so great, so do we. So we do today do not cover every single edge, but we have a framework that we can add all of these within. We have one global taxonomy of risks and attacks that keeps adapting.
Rune Kvist [01:22:20]: As, like, every time a new incident occurs that has never been seen before, great, let’s go and update the taxonomy so we bake that in. So I think we have one coherent universal approach. It doesn’t mean that we spend equal amounts of time on code and insert niche use case. We do spend time where people care. We think it’s very valuable to have one language.
Swyx [01:22:43]: Yeah. That makes sense. That’s, that’s a, that’s an important choice. We were going to end actually, but I thought of one final ending closing question, which is, take this however you want, right? Let’s say one and a half years from now, OpenAI’s secret panel of five experts declares that we have reached AGI.
AGI, Watchdogs, and the Need for Independent Oversight
Swyx [01:23:00]: Do you expect your business to change?
Rune Kvist [01:23:03]: No. I think there is some important way. I think the last businesses to exist beyond the labs
Swyx [01:23:10]: Will be underwriting.
Rune Kvist [01:23:12]: Well, there is one, there’s one job that the labs can never do for themselves, which is to be their own watchdog.
Swyx [01:23:19]: There you go.
Rune Kvist [01:23:21]: So I think kind of to the extent that you believe this frame of, like, you’ll see hyper-concentration, like the labs will kill all the startups
Swyx [01:23:29]: Yeah
Rune Kvist [01:23:29]: Which, we can go into the pros and cons.
Swyx [01:23:32]: I feel like the labs actually care a lot about this, right? There was the whole superposition, what do we do when we have models smarter than us and then a tier above, right, models smarter than them training them.
Rune Kvist [01:23:41]: Yes
Swyx [01:23:41]: The labs actually think about this a lot.
Rune Kvist [01:23:42]: They think a lot about. I think the there are some of the smartest people on these topics work at the labs. So the problem is not whether they care. The problem is that they will all be stuck in a race where they might have incentive to cut corners, and they might have incentive to withhold information from the government, et cetera. And so one kind of feels like eternal truth is that you need an independent third party to go and inspect that data and share information, in this case, say, with the government. It’s more of an incentive problem than an interest problem. I think they’re fundamentally all trying to make this go well.
Swyx [01:24:14]: What I’m not hearing is, like, AGI, whatever that label means to you, to me, to them, doesn’t fundamentally have, like, a qualitative shift
Rune Kvist [01:24:23]: Correct
Swyx [01:24:23]: In, like
Rune Kvist [01:24:24]: Correct
Swyx [01:24:24]: You still have to evaluate the models.
Rune Kvist [01:24:26]: And I think the one thing that would make this a qualitative shift is, there’s. For some definitions of AGI, it will just get nationalized. It’ll be a threat to sovereignty.
Swyx [01:24:34]: Yes.
Rune Kvist [01:24:34]: And then at that point, it kind of maybe every company is the government is every company. I struggle to think about that world. But at that point, you’ve kind
Swyx [01:24:42]: We. I don’t think we’ll move fast enough.
Rune Kvist [01:24:44]: Right.
Swyx [01:24:44]: You know, like, we’re not, we’re not set to do that.
Rune Kvist [01:24:47]: Yeah.
Swyx [01:24:48]: But I have discussed this a lot on the podcast.
Rune Kvist [01:24:51]: Yeah.
Swyx [01:24:52]: I mean, you know, as far as the watchdog concern, I will also mention that because I have my finance background, I often think about the scene in The Big Short where they talk to, like, Moody’s, but also Standard & Poor’s. And then the lady at Moody’s is like, “Well, if I don’t give you a triple A rating, you’re just going to go down to Standard & Poor’s.”
Rune Kvist [01:25:09]: Yes.
Swyx [01:25:09]: So actually the watchdog is a natural monopoly because if you have race dynamics in watchdogs, then the watchdogs will compete each other to the lowest possible standard.
Closing: Insurers, Incentives, and Trust Infrastructure
Rune Kvist [01:25:18]: Correct.
Rune Kvist [01:25:20]: And so I think what one of the things, one of the reasons why we’re very excited about having insurers be around this table is that insurers are the only ones that do not have this dynamic because they pay the bill. If they keep lowering the prices
Swyx [01:25:32]: Yeah, you will
Rune Kvist [01:25:33]: They also pay the bill.
Swyx [01:25:33]: You won’t find the market clearing.
Rune Kvist [01:25:35]: And this is not true for Moody’s where, they don’t directly pay the bill if they make recommendations that are off. So we think that balancing factor is pretty important. And I think it also highlights that there’s, like, no system that’s perfect. You need scrutiny of Moody’s, you need scrutiny of the watchdogs, for sure.
Swyx [01:25:52]: Beautiful. Thank you so much for indulging. This is a beautiful conversation covering everything. Congrats on your success so far.
Rune Kvist [01:25:59]: Thanks for having me.
Swyx [01:26:00]: Yeah. Awesome.
Rune Kvist [01:26:00]: Appreciate it.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe- At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!
From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round.
In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more.
You can get his book “The Eureka Machine” here!
We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence:
We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself.
We discuss:
* The Eureka Machine and Richard’s vision for an AI that can automate invention
* Why Richard is optimistic about superintelligence for science and technology
* Why AI hard-takeoff scenarios may underestimate physical and economic constraints
* The risks of regulating intelligence itself instead of specific AI applications
* Reward hacking and why increasingly intelligent AI makes objective design harder
* Richard’s critique of Anthropic’s constitution and constitutional AI
* Alignment vs. personalization and whose values an AI should follow
* Why open-source AI matters for resilience, competition, and geopolitical soft power
* Why Richard left You.com’s frontier-model work to start Recursive
* Recursive self-improvement and automating the process of AI research
* Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models
* DecaNLP, early prompt-based generalization, and the research that influenced GPT
* Why rejected research can shape entire technological timelines
* Open-endedness, evolutionary approaches, and rainbow teaming
* What happens if AI systems begin setting their own goals
* Why simple objectives like profit maximization can produce dangerous reward hacks
* Recursive’s long-term plan to apply self-improving AI to science
* The compute, hardware, and economic constraints on AI takeoff
* Recursive’s early NanoChat, NanoGPT, and GPU kernel optimization results
* Why automating AI research could reduce years of work to weeks
* Reward engineering and what makes auto-research systems actually work
* The AI Economist and using simulations to test economic policy
* Whether LLMs can realistically simulate people and entire economies
* Benchmark bugs and evaluation harnesses and the difficulty of measuring AI progress
* Recursive’s near-term focus on AI for AI research
* Harness optimization, sandboxing, and web search as core agent infrastructure
* You.com and the search stack for AI agents
* AI in finance, backtesting, and data leakage
* Richard’s three fundamental components and ten “spaces” of intelligence
* The theoretical upper bounds of vision, communication, knowledge, and computation
* Creative intelligence, metacognition, and AI-generated goals
* Survival and replication and why AI does not necessarily need to fear being turned off
* High agency and ambitious goals and Richard’s advice for people building with AI
Richard Socher
* X: https://x.com/RichardSocher
* LinkedIn: https://www.linkedin.com/in/richardsocher/
Timestamps
00:00:00 The Eureka Machine and Superintelligence
00:02:23 AI Optimism, Slow Takeoff, and Regulation
00:07:56 AI Safety, Reward Hacking, and Anthropic’s Constitution
00:11:49 Alignment, Personalization, and Open Source AI
00:15:46 Why Richard Started Recursive
00:20:03 Recursive Self-Improvement and the Founding Team
00:22:55 Are Today’s LLMs Enough?
00:29:03 DecaNLP, GPT, and the Rejected Idea Ahead of Its Time
00:34:38 Open-Endedness and Evolutionary AI
00:36:38 What Happens When AI Chooses Its Own Goals?
00:41:16 Superintelligence for Science
00:42:40 GPUs, Compute, and the Limits of AI Takeoff
00:45:07 Recursive’s Results: AI Beating Humans and Their Agents
00:49:14 Reward Engineering and Auto Research
00:53:12 The AI Economist and Simulating Entire Economies
00:58:07 LLM Simulations, Personas, and Mode Collapse
01:03:38 Recursive’s Roadmap, Agents, Search, and Finance
01:09:13 The Upper Bounds and Spaces of Intelligence
01:30:21 Goals, High Agency, and Advice for Builders
Transcript
Introduction: Richard Socher and the Eureka Machine
Swyx [00:00:00]: We’re here in a studio with Vibhu and myself and Richard Socher. Welcome.
Richard Socher [00:00:06]: Thanks for having me.
Swyx [00:00:07]: We just talked about the Eureka Machine, or we just released a talk, at AI Engineer about the Eureka Machine. Is it — you said it’s your life’s goal. What is the Eureka Machine?
Richard Socher [00:00:16]: The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It’s essentially a superintelligence that can be given any goal, any environment, reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for.
Swyx [00:00:45]: Yeah, I think we have the book pulled up here that you’ve written.
Richard Socher [00:00:50]: That’s right, yeah. I finished it last year, a little bit before we started Recursive, and now we’re gonna try to build parts of that.
Swyx [00:00:57]: You finished it last year. It’s July. What takes so long?
Richard Socher [00:01:01]: Oh, man, books. Books are incredibly slow.
Richard Socher [00:01:04]: It’s ridiculous. That whole industry is just unfathomably slow.
Richard Socher [00:01:07]: So a lot of the ideas have been out there for a while, but yeah, I’m really glad it’s finally coming out in September this year.
Swyx [00:01:14]: We might have AGI by then. Like, we don’t know.
Vibhu [00:01:18]: Any key takeaway that you’re most excited to put in here?
Techno-Optimism, AI Upside, and Slow Takeoff
Richard Socher [00:01:21]: Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics, and all kinds of other engineering tasks. I think there is so much more that can be done with better technology. And right now, I feel like a lot of people need, like, better marketing, not just for the future in general, but also, better marketing for technology and in particular for AI. And this book, should show even the AI skeptics, how much positive upside there is for AI, especially when it comes to inventing, new scientific discoveries.
Swyx [00:02:09]: I think you quoted the techno-optimist manifesto from, Marc Andreessen, which I think was, like, beautiful in its, ambition and clarity and simplicity almost as well.
Richard Socher [00:02:18]: I agree. Yeah. Yeah, you can disagree with him on some things, but, like, I think he’s right on the techno-optimism.
Swyx [00:02:23]: Where do you think optimists get in trouble?
Richard Socher [00:02:26]: Like, you shouldn’t have blind optimism. You should be very clear-eyed, like, especially when with such an omni, like, use type of technology as AI is, you need to think about the potential downside scenarios, especially when people use it for things that you don’t want them to use it for. It’s a little bit like the internet, and I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet, if you were to say, “Well, because there’s bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way, you can’t share the illegal content as quickly, or we should make the hard drive smaller so you can’t store as much illegal content.” But I’m like, “That’s not how you regulate that.” that’s like saying like we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications. Sure, I don’t want, like, some AI surgeon to, like, practice some RL moves in my brain. It should be fully FDA certified. Sure, I don’t want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it’s let loose on the highway. But I feel like those downside scenarios, that some optimists sometimes maybe don’t consider enough are fairly easily regulated, compared to, what the doomers are worried about.
Swyx [00:03:54]: It — Slow takeoff is part of the strategy as well?
Richard Socher [00:03:57]: I do think, as excited as I am about, AI and its impact for society and, culture even, and certainly technology and economics and wealth and, health and all of those things, as excited as I am about all that, I do think the most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about, the compute substrate. How quickly can you get enough, GPUs on? There are also constraints in the economy where there are a lot of industries that don’t require an insane amount of complex intelligence and complex capabilities. Like, if you think about jobs in, brands and, like, clothing and apparel and, like, handbags and stuff, superintelligence isn’t gonna make your fancy $10,000 handbag any fancier?
Richard Socher [00:04:57]: It’s like that’s — It will have no effect on the economy. You think about travel and tourism. People wanting to see the pyramids, in Egypt, it’s not gonna change that much with AI. Sure, you can, like, generative a fake, photo of you and next to the pyramids.
Swyx [00:05:12]: I can use Genie and, tour the pyramids in Genie.
Richard Socher [00:05:15]: Yeah, exactly. But, and there’s so many industries, like logging and oil. You’re not gonna magically get 1,000x more oil because, like, sure, there will be robotics, like drilling and things like that could be done, but it’s not gonna 1,000x that industry in a, like, crazy hard takeoff scenario, both on the economy, and I can go on and on about all the other examples, where that, like food and so on, where that doesn’t necessarily change that much. And then, yeah, there are real physical constraints. And then there are, of course, like, people like, off-ramping from progress. That’s one of my concerns often is that I see people in, like, Europe and other, whole regions almost feeling like they. Like many people there wanna off-ramp from progress, period. And that will also slow down, like, more improvements.
Swyx [00:05:59]: Yeah. We have this pulled up where, this is one of those things that, is very topical right now because now all the Frontier Labs are calling for the option to pace AI. They don’t say pause, they say pace. I don’t know if there’s there’s any take from you about, like, whether or not this will be effective.
Pacing AI, Regulation, and Safety Incidents
Richard Socher [00:06:17]: I think the downsides of trying to truly regulate with the full power of law what people do on their GPUs, would be worse than any of the concerns that they have. Like, it would be an crazy totalitarian state
Richard Socher [00:06:37]: If every one of your GPU computes was known to some big government or multi-government agency.
Richard Socher [00:06:44]: It’s like, it’s literally if you try to regulate intelligence, it’s trying to regulate thought, and that’s ridiculous, and it’s crazy. I think it is make — it is sensible to regulate some of the applications of this technology.
Swyx [00:06:55]: Yeah. We had a bill, actual bill to regulate the number of flops in a model, and I’m like, “Okay, well-”
Richard Socher [00:07:00]: Europe done it. Like, these guys have been successful enough with their fearmongering that all of Europe has regulated itself so much before it even had a proper AI takeoff because they listened to some experts who say, “We might all die if this technology has more than this number of flops.” And they’re like, “Well, we’re good. We wanna want people to thrive. Let’s not have technology that could have a small chance of all of us dying.” And so they regulated exactly those kinds of things in the EU. And so it’s, it’s very unfortunate that there are real implications for some people when others saying, “Let’s pace while they’re sprinting as fast as possibly,” “as fast as humanly possible towards that frontier themselves.”
Swyx [00:07:43]: Yeah. It’s also not a global pause, right? Like, other nations are still accelerating at the same pace.
Richard Socher [00:07:50]: Oh, yeah.
Richard Socher [00:07:50]: You’d need a totalitarian world regime if you tried to regulate intelligence and GPUs and what people do on them.
Swyx [00:07:56]: Any takes on the safety angles of this? So there was a drawback of Fable, a pause on 5.6 before it could be released. Recently, there was Hugging Face with the OpenAI cyber incident. Any takes there?
Richard Socher [00:08:11]: 100 percent. I think these are serious issues of reward hacking, and clear failures, of doing proper red teaming or rainbow teaming. I don’t know if you saw this paper from Tim Rocktäschel and a few others, where one AI, is tasked to try to hack another AI and then they can go back and forth in an open-ended fashion to inoculate themselves from those. Yeah, this is the paper. It’s a really clever idea. Open-endedness, and evolutionary inspirations are, big for us at Recursive as well. And so I wish they had used more of that. And it’s clear that, for instance, the constitutional AI. I don’t know if you remember anthropic.com/constitution. You can pull it up and search for cyber right there. It says, “Hard constraint. Claude will never ever do cyberattacks, and that is a hard constraint in our constitution.” So here are the current hard constraints on Claude’s behavior.
Richard Socher [00:09:16]: Number 3, create cyber weapons or malicious code that could cause human damage.
Richard Socher [00:09:21]: And clearly, this whole constitution was fake. Like, it clearly isn’t being adhered to at all.
Swyx [00:09:26]: Because Anthropic also found that they had in their testing
Richard Socher [00:09:30]: They’re also. Like, they’re like, “Oh, well, other people are hacking now.” There are a couple things. One, you can make a sandbox very simple, and then it’s very easy to hack yourself out of a sandbox, right? But what I think it shows is that we’re currently in this state of AI where the reward engineer still has to do a lot more careful work, and where the AI, in most cases, is not very good yet at understanding what is meant versus what is being said. And so concretely, I think this will happen if we were to have this intelligence more easily accessible in a lot of companies. Imagine you run a service center and someone says, “Oh, here’s my CSAT score and my dashboard. Make this number go up.” It’s like, “Our CSAT score is so poor.” The intelligent AI will just be like, “Oh, sure. Like, I’ll just create 1,000,000 bots that call our service center and give a 5 out of 5 rating at the end, and the number went up just like you asked for.” And you’re like, “That’s not what I meant.” “I meant with our real customers.” The AI goes off and says, “Well, easy. I’ll just give a 1000 dollar gift certificate for every failed, whatever DoorDash
Richard Socher [00:10:35]: Offer.” It’s like, “That’s not what I meant.” It’s like, “Well, but that is what you said.” And like, so I think clearly articulating what the rewards are is something we haven’t gotten very good at as humanity. And then clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards. Now, what gives me hope is there are the first inklings, of this being better. I’ll give you an example like WhisperFlow. Full disclosure, I invested, in their seed round, but at AIX Ventures, but, WhisperFlow has gotten much better at writing what you mean and not what you say. And I think that is a sign of things to come. I think there will be more and more AIs as we make it more and more intelligent that will be better at being aligned with what is meant.
Swyx [00:11:21]: Will it be done through a constitution or RLHF or
Reward Hacking, Alignment, and What We Really Mean
Richard Socher [00:11:23]: Clearly, constitutions don’t matter at all.
Richard Socher [00:11:25]: It doesn’t work. And that was, I think, mostly marketing. I think we need to find better solutions for it. And I think at Recursive, we have a few very good ideas and some already
Richard Socher [00:11:34]: Like, ways where I think we have a better grasp on it. I don’t think we’ve fully, figured it out yet, but, we’re thinking a lot about safety, and the more intelligent the AI gets, the more you want it to be aligned, the less you want it to think about reward hacks and try to do the right thing.
Swyx [00:11:49]: I don’t know if we’ll touch on this topic, but I’m just gonna throw this question in here because it’s something that’s weighing on me. Alignment, let’s call it, is alignment to general humanity’s preferences, the median preference. Personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants. How do you choose?
Alignment, Personalization, and Cultural Values
Richard Socher [00:12:12]: It’s a great question.
Richard Socher [00:12:13]: I think you ultimately have to, of course, be aligned with laws. Like wherever your AI is deployed and needs to align with the law. I do think what AI often does is put this mirror in front of us and say, like, “This is what you’re looking like. Now I can amplify that a 1000 times. Is it still what you want?” and the truth is that different cultures made different choices. Like, in Eastern cultures, the greater good is often valued more, than the individual. Western civilization, we care more about individual freedoms and rights and the pursuit of happiness and so on, than others. And even there are gradations. There’s regulation versus litigation trade-offs. In the US, you first can often, not every time, like, FDA and so on does regulate some areas, but in many cases, the bad things happen, someone sues someone else, and then there’s a law based on that. In Europe, they try to often avoid any harm to anyone and regulate before. And both are, trying to do the best thing, but, some is more amenable to innovation than others. And so yes, you’re right. Like, I think ultimately each individual, each country, and humanity as a whole has to think about those values more, and then try to put them into laws. And that those are ultimately the constraints. And hopefully, different, societies, just like now with their AIs, will align their AIs to a different one so we have not just a monoculture of alignment.
Vibhu [00:13:46]: Here’s a follow-up on this that I wasn’t expecting to ask. Do you have takes on open source, open weight versus who owns the intelligence? So, clearly not the biggest, fan of the constitution
Richard Socher [00:13:58]: You had to do this in the topic side off.
Vibhu [00:14:00]: But it’s fine.
Vibhu [00:14:02]: Point being, any thoughts on who should own weight? Should it be open? Anything there?
Open Source, Soft Power, and Who Owns Intelligence
Richard Socher [00:14:06]: 100 percent. I am a big fan of open source. We’re gonna sign some various open source letters at, Recursive also. I think, even in the worst case attack scenarios, it is better to have more good actors have more different types of AI, accessible. I think, open source is a little bit a soft power type of thing, too. So I do think it’s good for the Western world
Richard Socher [00:14:31]: To have an answer to that, out of China. I do think, when you watch a Hollywood movie, there’s — it’s like, I don’t wanna misc, diss all of movies, but there’s a certain sense of propaganda, right? You watch one side of things, right?
Vibhu [00:14:46]: Oh, yeah. Have you seen Top Gun? Like, come on.
Vibhu [00:14:48]: Like, it’s like half of it’s paid for by the US Army or something.
Richard Socher [00:14:51]: Yeah. And so. And, I think that’s just natural. Like, but what’s interesting here is I think LLMs are essentially a similar type of soft power to movies and beyond, because they’re also, highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. Like, if, like a child asks an LM, like, “Tell me an inspiring story of what I should do when I grow up,” right? It’s like those are all these, like, subtle things. So I think it’s important, for Western world. I do love, individualism. I do think, despite, some of its flaws, like capitalism is the best way we have governed, found ourselves to govern, and so on. And so I do think there are various aspects that would be good, to have a Western open source answer, for LLMs. And, with Recursive, I can’t make the announcement quite yet, but we’ll
Richard Socher [00:15:43]: We’ll be relevant in that space very soon.
Vibhu [00:15:46]: Okay. All right. Exciting. I wanna bring us to Recursive. So outside of our tangents, you have a pretty deep background in the NLP space. You worked on, like, early embeddings, GloVe with Chris Manning, who was a previous guest on the podcast, You.com. What’s the history? How did you decide to start another company?
From You.com to Recursive
Richard Socher [00:16:06]: Yeah. So I’ve been excited about AI for over 2 decades now. I sometimes feel like it’s ancient history now. It’s BC, the before ChatGPT era. No one cares about all the religions that happened, before, Jesus Christ, and no one cares about the models that happened before, transformers and ChatGPT and stuff. But, like, it’s something that I’ve been deeply passionate about. I think AI is one of the most interesting things one could work on, period. I think language is the most interesting manifestation of human intelligence, too. And, at You.com, we eventually off-ramped from pushing, like the frontier of AI forward to mostly giving people, like, good search engines, search, APIs and answers over the web. I think that’s an extremely important part of intelligence, just knowledge and access, especially even, we’ll get there maybe later, if you wanna invent a eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. And to know what has been invented, you gotta have internet access. So it’s the number one used, most used tool, in LLMs, agents, chatbots, and so on is web search. So I’m really excited for You.com to own that and grow really well in that with really large customers and so on. But it’s also not building frontier models anymore. And so I initially tried to do this within You.com and raise another round and so on, but you just can’t. You have to do a certain thing, and until you print enough money that you’re allowed to start a second thing within that company is really hard. At the same time, I had all these ideas. I put them into a book. I finished the book last year, and I was like, “It’d be really fun to work, on this myself.” I felt like with word vectors, and then prompt engineering and, ImageNet and larger language models for protein generation, not folding and so on, I, me and my teams have pushed the field truly forward. And I feel like we can do it again, here at Recursive. And in many ways, what I observed over the last, 20 years in AI is that whenever we replace some human part of the process of creating AI with a learned system, improvements follow. And so. We’ve done that taking out manual feature engineering, like in sentiment analysis. I don’t know if you remember these old days where, like there are linguists, and they’re like, “Here’s how you negate, and there’s a, like, regular expression.”
Swyx [00:18:21]: I went to Penn where we — they had, like the WordNet
Richard Socher [00:18:24]: That’s right, WordNet, all of that stuff. Yeah
Swyx [00:18:26]: Original. They use, our grad students to label Wall Street Journal articles and, like, really construct a knowledge graph of
Richard Socher [00:18:32]: There you go.
Richard Socher [00:18:33]: And WordNet started, was part of how we started ImageNet. But anyway, so, like, it was really, like, fun, to do. But when we replaced all of that manual feature engineering with vectors and neural nets and just backprop through everything, it started to work really well at scale. And so then everyone started to do architecture engineering, and I was like, “ that clearly can’t be it.”
Swyx [00:18:53]: You mean, neural architecture search?
Richard Socher [00:18:55]: Like, manually, they would say like, “Oh, I’m, I’m doing sentiment analysis, so I have a special neural net that’s really good at sentiment analysis.” And then the machine translation community had a special neural net for machine translation.
Swyx [00:19:06]: I see.
Richard Socher [00:19:07]: The summarization people had their own stuff. And I was like, “That clearly can’t be it. We should unify all of that.” So I had 2 papers. One is called Ask Me Anything, and the other one was called DecaNLP. And DecaNLP eventually got cited, like, 5 times by the first GPT paper. And, to me, that was, like a really a big step forward. And then, of course, you had to combine this idea of prompt engineering with transformers and with language models, and you put it all together, you scale it up, which is also a huge amount of work. And then, the field progressed a lot. I feel like the next step and maybe the last step of that history and the arguably, success has a lot of parents, only failure is an orphan, like my version of that AI history, I do feel like in that history, you can think about, “Well, what’s the next way to automate?” And that is the AI research itself, like the human, process of ideating, implementing, and validating ideas.
Automating AI Research and Recursive Self-Improvement
Richard Socher [00:20:01]: And in our case, ideas for AI.
Richard Socher [00:20:03]: And when you have AI then help you with that, it, by almost definition, becomes a self-improving AI ‘cause it now does research on itself. And there are lots of different misnomers. Some people think auto research is already recursive self-improvement. It’s
Swyx [00:20:17]: Yeah, and you explained that in the talk
Richard Socher [00:20:19]: Completely different.
Richard Socher [00:20:19]: But, to me, it’s the most interesting thing that I could be doing, and I’m really excited with the co-founding team. What’s interesting is we have 8 co-founders in total, including myself. And so
The Recursive Founding Team and Darwin Gödel Machine
Swyx [00:20:31]: They are gonna bring it up.
Richard Socher [00:20:31]: Nice. Yeah. And they’re all. I could talk about all of them if you want.
Swyx [00:20:34]: Super stacked.
Richard Socher [00:20:35]: Yeah. Just an incredibly talented group of people. And we all came to the same conclusion, but from very different directions. Like Josh Tobin, is our CTO. He ran, a bunch of different, projects at OpenAI, like, Codex and deep, research, agents and ChatGPT agents and so on. But before that, he also worked in robotics, and he saw the smaller simulations, and how it’s gonna be really hard to scale that in full generality. And so that’s, that was his angle coming to recursive self-improvement. We have Jeff Clune who’s been working in, like, open-endedness for a long time, together with Tim Rocktäschel. Tim Rocktäschel also built Genie 1, 2, and 3, which is, like the most exciting and most sophisticated, I think, still world model, anywhere. And so they both came from this, open-endedness angle. Jeff also, I think, published one of the most exciting papers in recent years about recursive self-improvement called the Darwin Gödel Machine. Super interesting paper. If we could, maybe pull it up really quick
Richard Socher [00:21:35]: It would be, like, super interesting to see ‘cause you see
Swyx [00:21:38]: By the way, I love how many paper citations.
Swyx [00:21:40]: You’re, you’re giving people a lot of homework, which I like.
Richard Socher [00:21:42]: Love it. Yeah. And so, like Caiming Xiong, a rockstar, we worked together at MetaMind and Salesforce Research together. Alexey Dosovitskiy invented the Vision Transformer, one of the most cited, papers in computer vision. Tim Shi is, like also a unicorn founder. Yuandong Tian led RL at Meta. So just like, yeah, really fun to work with them, and the next level of people are just incredibly strong, too. So it’s been a really fun ride so far. So the first figure, you see exactly these kinds of ideas, that, I think, yeah, inspired a lot of us and now more and more people, where you have this archive of different coding agents. They learn how to self-modify, evaluate, and then create these phylogenetic trees, of, yeah, different ideas.
Swyx [00:22:28]: That’s one foundation. So that Darwin Gödel is an influence.
Swyx [00:22:32]: Open-endedness is an influence. Any other trains of thought that feeds into Recursive that I’m missing?
Influences: Open-Endedness and Learned Systems
Richard Socher [00:22:38]: Going to replace manual parts of the process of building AI
Swyx [00:22:42]: I
Richard Socher [00:22:42]: More and more
Richard Socher [00:22:43]: With learned systems. Yeah.
Swyx [00:22:45]: Which, and, like, merging different fields into one general, architecture.
Richard Socher [00:22:51]: That’s right.
Swyx [00:22:51]: Okay. It seems like language models are already pretty generalist, right?
Swyx [00:22:55]: Your next token predicting your reasoning. Was there a time that you thought, “Okay, these are good enough to have recursive self-improving machines”?
Are Current LLMs Enough?
Richard Socher [00:23:05]: It was clear to me that they will happen, within, like a year or two, and then it did exactly happen, like, earlier this year, right? Earlier this year, AI really went from not just being code, but being able to code. And that is a big unlock. It’s definitely making everything a lot easier than it was, before the beginning of this year.
Swyx [00:23:24]: One question that I think a lot of people have is the current LLM paradigm enough? Or, like, let’s call it autoregressive transformer, with reasoning, whatever. Don’t you need something else, some big unlock, whether it’s world models, which Chris Manning is working on, or memory, continual learning, all that stuff? Or is it all of the kinds, and you think the current, let’s call it transformer architecture, is here to stay and that’s it?
Richard Socher [00:23:48]: A lot of thoughts. So number one, I do think it would be great to have less of a monoculture in AI research.
Richard Socher [00:23:55]: Like, if you look at, AI conferences now, I still remember the days in, like, 2010 when I tried to get my first neural net papers and NLP conferences accepted, and they just desk rejected them because, like, neural nets were something, quote, unquote, “We don’t do in NLP conferences,” and just, like, desk rejected. And it was very brutal in the first years of my PhD. Now I feel like it’s almost like the field switched to the other side. Like
Richard Socher [00:24:17]: Someone should try some other weird, crazy ideas now that aren’t.
Swyx [00:24:20]: There’s also a few. I really respect, like, people still working on, like, GNNs and, like tabular stuff and.
Richard Socher [00:24:25]: Yeah. Like, someone should still, like, do novel out there ideas. At the same time, I think whenever people say, “Oh, LLLMs are. Like, this is the end for LLLMs,” they just don’t, like. LLLMs are also not the LLLMs of, like the past, right? Like, they are so much more sophisticated now. There’s so many more clever things that people are doing. It — There’s, like, different stages of training. You have the whole RL training, and you can take actions and, like all of these things where that can go really far. And then the folks that come from the neurosymbolic, direction say, “Oh, this will never work because they can’t do neurosymbolic reasoning.” It’s like, I think they’re underestimating still the ability for these models to code, and code is neurosymbolic reasoning, and these models can code incredibly well. And so I do think there are, of course, more and more ideas that will be needed and we’ll continue to have. We’re seeing, like, more and more interesting high-level ideas coming out of the AI itself, too. And with really deeply integrating the fact that these models are code and can code, that line — I don’t wanna give it all away, but, like, I think that line has a lot more to grow. But it’s still an LLM, right? Even if that LLM codes for you and then runs that code in some integrated fashion. World models, I’m personally less bullish on. I think if you run a robotics company, you’re gonna build your own world model. I think world models are super fun, and Tim Rocktäschel came to a similar conclusion after building the most interesting one with Genie 1, 2, and 3, which is gaming is a huge application for world models. Can see I sometimes got stuck in some games and, like, got a little overly competitive in the wrong direction. And so I understand games are fun, but personally, I’d rather work on science than gaming. And so, yeah, I think LLLMs, a lot more room to grow.
Swyx [00:26:16]: Yeah. I think there’s some interpretation of world models that some people have where it’s like, well, it’s okay, yes, there is that gaming element. There’s this — there’s the embodied robotics element. But the other part also is just, the more abstract sense of LLLMs are just modeling output, but they’re not modeling the chain of thought, inside the human that has created the output. We can annotate it, of course, but, like, it’s, it’s always, like, this Plato’s cave reflection of a thing rather than the thing, right?
Richard Socher [00:26:43]: It’s true.
Richard Socher [00:26:44]: But I would argue that, and maybe we’ll get there in the 10, spaces of intelligence, but I would argue that even our projection, our eyes is a projection of the real world. And, like, we have only a very narrow, band of the electromagnetic frequency spectrum that we can observe with our puny little 2 eyes and so on.
Swyx [00:27:01]: It’s good enough.
Richard Socher [00:27:02]: It’s, it’s good enough for now, but, like the upper bounds of where it could be are so much higher. And, like, to map, the visual world the way humans see it is also not necessarily, like the end-all be-all for visual intelligence. And I would argue that language is still the most interesting manifestation of human intelligence. And while our visual cortex is certainly less sophisticated, than that of, certain animals all the way down to the mantis shrimp who can, have, like, 2 independent eyes, 3 bands, trinocular vision and each eye can see all the way to, like, floating temperatures in 4D and stuff.
Richard Socher [00:27:36]: Like, mantis shrimp, you should look it up. It’s like
Swyx [00:27:37]: Way OP.
Richard Socher [00:27:38]: Super crazy.
Swyx [00:27:39]: Yeah. ZeFrank, mantis shrimp.
Swyx [00:27:41]: It’s the best video in the world on
Richard Socher [00:27:42]: I love ZeFrank, yeah.
Richard Socher [00:27:44]: Big shout-out to him. But, like, I think there’s a lot more room to grow, but none of these, other animals have language that’s as sophisticated as ours, certainly not in writing. And once you can write, you can, start thinking about longer term civilizations. All of that is language. Programming is much closer to language. And I would argue, and this is, like an important thing in the spaces definition of intelligence also, is that all of these spaces are highly correlated, but visual intelligence is neither necessary nor sufficient for overall intelligence. You can be blind and still be an intelligent human being. And an AI can be blind and still be quite intelligent too.
Swyx [00:28:25]: We were gonna bring this
Richard Socher [00:28:25]: Which doesn’t mean that you’re not more intelligent when you have it. Yeah.
Swyx [00:28:28]: We’re gonna bring this up. I might as well — Like, we have a classification of 10 types of intelligence that you had at the end of your talk. So I’m just gonna flash this up now for people to cover this. I don’t know if, maybe we’ll put this towards the end. We’ll come back to this. I just wanna mention that, you do have a philosophy that I like when people do lists because then I can just go through this and then it gets — it’s educational for people. But let’s go back. I don’t wanna get distracted. But, so effectively, I’ll, I’ll, reinterpret what you said as Yann LeCun is wrong. And then we’ll just
Richard Socher [00:28:56]: Don’t quote me as that. I’m, I’m good friends with Yann. I think very highly of him in many directions.
Swyx [00:29:01]: But he’s wrong.
Swyx [00:29:03]: You mentioned GPT-1, and I cannot let any, Alec Radford, mention escape. Did you talk with him when he was training GPT-1? Like, any historical, fun stories there that you might come up?
DecaNLP, GPT History, and Scientific Gatekeeping
Richard Socher [00:29:18]: I did not, like, meet him a bunch of times. I think we met maybe once or twice at some conferences. But, like, he has told, I think Brian, the first author of the DecaNLP paper, that it did inspire him, and he cited it five times in the GPT-2 paper. So, and that’s, like
Swyx [00:29:36]: Yeah, good enough.
Richard Socher [00:29:36]: Very clearly said, like, this was the first instantiation where they showed in the DecaNLP paper, McCann et al, that you can just phrase every single NLP problem as here’s some prompt, text context, here’s a question and task description and here is some output. If you just do that enough, you can have one unified neural network model, which, by the way, also had all kinds of interesting attention mechanisms. There are slightly different formulations to the transformer. I think came out the same year, plus/minus a few months. And then you can unify all of natural language processing into one neural net. That is the core idea.
Swyx [00:30:14]: And this was as opposed to at the time, LSTMs and what have you.
Richard Socher [00:30:17]: LSTMs, but also, like, people being very stuck in thinking about one model per task. In fact
Richard Socher [00:30:25]: It’s, it’s kinda crazy, but the DecaNLP paper was publicly reviewed as, like, open, OpenReview. It was an ICLR submission. And, in it, you will see, how the whole community at the time thought about this. So, like
Swyx [00:30:43]: Some great contributions, but more work needed.
Richard Socher [00:30:46]: So look at, like, search for not even for humans. Just scroll it up here. Like, question answering is not a unified phenomenon. There is no such thing as general question answering, not even for humans. And this is like, really, you replace your brain with a different brain a different neural net when you answer, like, different kinds of questions. It was unfathomable to the experts at the time that you can have one unified neural network that would answer all of these different questions. They are saying, “No, all of these questions require very different systems to answer, and trying to pretend they are the same doesn’t help anyone solve any problems.” That’s what it says right there, right? That’s how hard it was to fathom. And now, of course, people, when I say, “Oh, we’re gonna invent prompts,” people are like, “You can’t even invent prompts.” It’s such an obvious idea to have one neural network that, of course, does everything in NLP.
Richard Socher [00:31:37]: But at the time, it was, like, extremely controversial, and the paper got rejected. And the sad thing is that it got rejected so hard and they were so certain that we stopped going on our list of things to try. And the number 2 or 3 on the list of extensions for this paper was add language modeling as another task. And then we could have, and that would have accelerated the timelines, in 2018, like, even further for humanity. But we got so crushed, and we were like, “Okay, maybe we’ll just work on some of our other ideas for now and, like, come back to this later.” Yeah.
Swyx [00:32:09]: How can we design a review system that rewards non-consensus?
Richard Socher [00:32:14]: Honestly, I started to feel like arXiv is such a gift to humanity. With arXiv, you should just put your paper out there.
Swyx [00:32:24]: Is it pre-preprints?
Richard Socher [00:32:25]: Let — And honestly, I think Twitter X, people like you who pick up interesting papers, that is a better filter than the experts. Let everyone, like, have access. Now, of course, there are some downsides, which is, like, if you’re super unfamous, you have no Twitter following
Richard Socher [00:32:41]: You don’t wanna be on social media or whatever, you write a good paper, maybe someone, somehow no one notices it. But I would argue that if you just tell, like, 10 of your friends in your community about a paper and it is a really significant breakthrough, someone is bound to talk about it again. And, so I think science needs less gatekeeping. And, even though ICLR, with Yann LeCun, who started it, as one of the co-founders of ICLR back in the day, he also wanted less gatekeeping ‘cause he too was rejected for many years together with Yoshua Bengio and Geoff Hinton with all their early deep learning and neural net papers ‘cause it was just not the hot thing. And so ICLR started with that, but then it also started gatekeeping a little bit themselves on various ideas. So I think less gatekeeping, more open, and then allowing people to say, “Look, even if this is just on, or, quote, unquote, ‘just an archive,’ if it has like 1000 citations, it’s a legitimate paper. Doesn’t really matter where you published it.”
Swyx [00:33:34]: And I agree with that. I do think it’s sad that I’ve heard that grad students have to do, like, how to Twitter, seminars to each other
Swyx [00:33:43]: Just because it’s so important for publishing these days. This person is just reflecting the sentiment at the time.
Richard Socher [00:33:49]: That’s right.
Swyx [00:33:49]: But it’s
Richard Socher [00:33:50]: I think it’s
Swyx [00:33:50]: It affected you so much
Swyx [00:33:52]: That you stopped work on it.
Vibhu [00:33:53]: The sentiment also came out of some of the research, right? Like, the original BERT paper was trained, and towards the end of the paper, they’re like, “Okay, throw off the last head, train specific iterations for
Vibhu [00:34:05]: Extractive summarization add a head for this.” Like, you should do task-specific stuff. These are, like the authors that wrote Attention, wrote BERT, telling you this is what you’re meant to do. And, like the training tasks were also very odd. They’re like
Vibhu [00:34:16]: The — “We know that the model overfits to this weird mass language modeling. Throw away this part and just do specific models,”?
Richard Socher [00:34:23]: Exactly. And, like, we had to try — come up with all clever ways of, like attention and pointers and so on to get the neural network to be able to do all of these tasks. And then some of them were better than state-of-the-art, some weren’t, but we were like, “But it’s still in one model.” I thought it was really cool. Really interesting.
Swyx [00:34:38]: I was gonna move on next to Tim and open-endedness. He was head of open-endedness at Google.
Open-Endedness, Rainbow Teaming, and Self-Set Goals
Richard Socher [00:34:42]: That’s right.
Swyx [00:34:43]: I don’t know what that means.
Swyx [00:34:44]: But he did a lot of talks.
Richard Socher [00:34:45]: Genie 3 is one of the ways that
Richard Socher [00:34:47]: Rainbow teaming, yeah.
Swyx [00:34:49]: So I first saw him at — speaking of ICLR, I first saw him at ICLR when he talked about open-endedness. He’s he’s done a few talks. Can we define what is open-endedness for people who have never been exposed to the problem? They are like, “What do you mean? I thought the only goal of AI is to optimize against a benchmark or.”
Richard Socher [00:35:04]: That’s right, yeah. It’s a, it’s a fuzzy term because there’s so many different instantiations of open-ended, thinking. But, one way I often describe it, and certainly, Tim and Geoff Hinton would be even better at describing this, but it’s a suite of methods that is more inspired by evolution than, very specific rewards. So in that sense, it thinks more about environments, about co-adaptation. And so a concrete example is in the cybersecurity and LM safety space where you have one LM that tries to attack another LM to say something unsafe.
Swyx [00:35:40]: Yeah, the rainbow, yeah.
Richard Socher [00:35:40]: And now the environment is the 2 having a conversation and now they co-adapting, right? They’re like one makes a better attack than the first one inoculates itself somehow, like uses that as training data, makes it so it’s harder to say something unsafe based on that. And then as the attack stops working, the attacker now tries a different angle, right?
Richard Socher [00:36:00]: And that’s why it’s not just red teaming, but they’re called rainbow teaming.
Swyx [00:36:02]: So, like, don’t tell me how to do things. Let me just figure it out myself.
Richard Socher [00:36:05]: That’s right. Think about the environments that you wanna use. Think about the rewards at a high level that you wanna, inspire towards, and then let the AI try out many more ideas in this interplay between sometimes humans, but also sometimes other AI agents.
Swyx [00:36:22]: Yeah. I worked open-endedness into a model that I have been working on. It was the keynote for AI Engineer where you start. You, we have the token loop, we have the agent turns, and then we have goal. And I feel like the way that you’re describing open-endedness is still somewhat of a goal. Like, please attack this,
Swyx [00:36:41]: Other agent. But, to me
Richard Socher [00:36:42]: Yeah, you set the rewards. You set the environments.
Swyx [00:36:44]: The loop that makes the other loops is. What if the agent can set its own goals?
Swyx [00:36:49]: And is it, is that open-endedness? Like, you don’t give it a goal. Just, like, be a sentient being. And maybe sentient is a very loaded word
Swyx [00:36:57]: But just set your own directions. What do you think you should do?
Metacognition, Subjective Goals, and Measuring Intelligence
Richard Socher [00:37:01]: I love this direction. I think this is one of the 10 spaces of intelligence, that I clump under metacognition and thinking about thought.
Richard Socher [00:37:08]: And it’s an interesting one. Whenever people say, “Oh, AI is like, this is, it’s gonna stop from here. It’s not gonna get that much better,” and blah, I’m like there’s so many different spaces of intelligence that we haven’t even started exploring yet and hence have made very little progress on. And there is an interesting, connection to economics and, capitalism. Like, it doesn’t make sense for a company to build and spend billions of dollars building a model that instead of following the rewards and objective functions you gave it, may come up with its own objective functions and its own goals.
Richard Socher [00:37:46]: Right? And then imagine you’re like, “Okay, I spent billions of dollars. Now go develop this new battery, material for me and answer all my emails.” And it’s like, “Nah, I think it’d be more interesting to evaluate the molecular composition of the atmosphere, on Jupiter.”
Richard Socher [00:37:59]: And you’re like, “That’s not what I paid you billions of dollars for.” And so no one’s working on that for good reasons. And then also, understandably
Swyx [00:38:07]: It’s not useful.
Richard Socher [00:38:07]: It’s not, it’s not useful, and it could get a little bit weird, right? What if the AI does start to really have thoughts on its own, and what if we don’t like those thoughts, right? And so it requires a whole different way of thinking about it. I had a great conversation with a good friend of mine, Sam Gershman, who’s a neuroscience professor at Harvard, and, like, we just jammed on this a little bit on, like, what are the best meta goals. And, I do think, like, knowledge-seeking is a really good one. I’m currently thinking also about, like the ultimate measure and unit of intelligence broadly construed, and I finally have some. It’s still too early to share it. It’s not. I haven’t fully baked the thoughts yet.
Swyx [00:38:44]: Like some replacement for IQ.
Richard Socher [00:38:46]: IQ is such a terrible definition, right?
Swyx [00:38:48]: Elo.
Richard Socher [00:38:48]: It makes no sense. Yeah, Elos are terrible, too, because it’s always just like me versus others.
Richard Socher [00:38:53]: But, like, you can be intelligent and not constantly compare yourself to others? And so, yeah, there’s no, like. In fact, a lot of these definitions we have, which I briefly mention in my book, too, these definitions create sometimes explicit and sometimes a more implicit anthropic bounds. No dis to the company Anthropic, but just, like, this idea that your intelligence is like getting 100 out of 100 questions right on this IQ test. Well, if that’s your definition then you can only be at 100 out of 100. Where do you go from there, right? So you see a lot of these, benchmarks that people are working on they, increase, they get close to human, maybe sometimes
Swyx [00:39:30]: It’s like an S-curve
Richard Socher [00:39:30]: Slightly above human, and then it’s flat.
Richard Socher [00:39:32]: It’s like, ‘cause that’s your. If your definition is only that so tied to humans, you’re only gonna get to just slightly better than that. So I think metacognition is a great example of that, where we’re not even yet allowing the AI to think. We’re not working on it very much, and hence there’s very little progress in that.
Profit Maximization, Real-World Environments, and Reward Design
Swyx [00:39:49]: Yeah. Well, we’ve interviewed Andon, which I think, has been working on the most open-ended, benchmarks, which is just real-world, money.
Swyx [00:39:57]: Arguably, telling an AI to profit maximize is a bad idea.
Swyx [00:40:03]: But they are doing it.
Richard Socher [00:40:05]: I do think you don’t want that super. Like, you don’t want a superintelligence to have a ton of access to all kinds of tools and so on and then just give it that without some very careful reward engineering. ‘Cause it’s like, I just buy a bunch of defense stocks and I start a war. I make money. Like, it’s just like, it’s a tricky situation, right? You just buy a bunch of stuff, short basic goods for people, and you create some weird famine, like, issues. Like, yeah, there’s a lot of constraints you should put onto a trading system.
Vibhu [00:40:35]: It’s a fun measure, though, ‘cause, the bounds are very capped to where we’re nowhere close to them. Like, in Andon Labs, the model’s like, “Oh, it’s Saturday, maybe I just close the store today.” “Someone’s off. It’s okay. We’ll just close the store.”
Swyx [00:40:51]: It’s using Claude.
Vibhu [00:40:52]: Yeah. But
Richard Socher [00:40:53]: Yeah, no. I’m not, I’m not arguing against it. Just, like as you get more and more intelligence, you wanna be more and more careful with that as, like an open environment, ‘cause the environment then is all of Earth.
Applying RSI to Science and Invention
Swyx [00:41:02]: Yeah. Okay. For recursive, not strictly necessary, right? Because, like, if your goal is you make a machine that, like, invents the other things, then, like, just solve, the science things
Richard Socher [00:41:12]: Knowledge discovery, yeah.
Swyx [00:41:13]: Solve machine learning research and discovery and all these things. Good enough.
Richard Socher [00:41:16]: And eventually, so, our goal, I haven’t really. I don’t talk about it that often because it is a few years out, but our goal is once you have a recursive self-improving superintelligence, you then want to apply it to the most important problems. And I think a lot of those are in science and technology and broadly construed inventions, and those inventions in, physics to create better, cheaper energy with fission or fusion, in chemistry and to create better materials and better batteries and, better solar cells and so on. In biology, there’s so much, like, I think soon to be low hang- lower and lower hanging fruit because of AI, because of protein and generation, not just folding, but generating new proteins like we did in ProGen many years ago. Like, so much positive impact we had if you take that superintelligence and you apply it to science.
Swyx [00:42:04]: I do fundamentally believe that. There’s a lot of approaches, though. You’re not the only team trying and NeoLab trying.
Swyx [00:42:09]: There’s, like a lot of. Especially the physical sciences as well.
Richard Socher [00:42:12]: And that’s good. Yeah. I do think that physi- like the reason we are only doing it in a few years is that it’s a little too early right now. Robotics is not quite there yet. The AI is not quite there yet. But I’m fairly confident in 3 to 5 years, all those constraints will be gone, and then applying to real physical robotics experiments and so on, like true robotic process automation
Richard Socher [00:42:33]: Not the traditional RPA sense, but, like, having robots run experiments for you will be totally there. Yeah, it’s gonna be great.
Swyx [00:42:40]: Just to call back to something that you said early on about slow takeoff, you said that, like, while really the substrate that is limiting factor is, let’s call this chips, and semiconductors and all these things, and you have race funding for that and, you are investing a lot on that. But have you done the math on, like, is it even- Achievable and, like, what is the, industry concentration needed in order to achieve, like, scale?
Compute, Slow Takeoff, and Changing the Bitter Lesson Slope
Richard Socher [00:43:05]: Right now we know that, like, roughly, like a 1000 GPUs cost quite a lot of money.
Richard Socher [00:43:11]: Right? If you wanted, like, 10s of thousands of GPUs, you’re, you’re talking billions and billions of dollars. If you say, like, one GB300 is, like, you could eventually create models that are, on that substrate, like are close and similar to human intelligence. And you want, like, thousands and thousands of, AIs to think about really hard problems, in a similar fashion to humanity. Like, yeah, that-that’s, that’s a lot of money. You do the math. It’s like a lot. We don’t have that amount of money right now anywhere to, like, build that. Now, things can get more efficient. You will have, I think, soon better algorithms that won’t be, and better hardware that won’t be as energy-hungry, and so on. Our human brain does quite a lot of flops with much less energy.
Swyx [00:43:56]: 20 watts?
Richard Socher [00:43:57]: That’s exactly right. Yeah, that’s the number often that’s quoted. And, like, I think more, inventions will happen there, that then will accelerate the takeoff even further.
Swyx [00:44:08]: One thing I always try to reconcile when talking, like, with new lab founders is, like, you’re fighting Bitter Lesson all the time. You have to show initial progress, then you unlock the next tier of funding, then the next tier, then the next tier.
Richard Socher [00:44:20]: Which unlocks larger model categories.
Swyx [00:44:22]: Like, fundamentally, is that true? Like, are you fighting Bitter Lesson? Are you — will we have a way in which, like, no, we’re changing the slope in some fundamentally different way?
Richard Socher [00:44:31]: I do think we are changing the slopes in fundamental ways by making AI much more efficient, both in terms of the training as well as the inference.
Richard Socher [00:44:43]: Yeah. I think we will — When you allow AI to do the work that it takes other labs thousands of people and years to do, I think we’ll be able to get it down to weeks, and that will be much cheaper
Richard Socher [00:44:53]: And hence, more affordable, accessible to others and so on.
Swyx [00:44:57]: Yeah. You’ve shared initial results on that,
Swyx [00:44:59]: Which, like, conveniently OpenAI has also done to their GPT-5.6, so we can talk about it now.
Richard Socher [00:45:04]: Yeah. Yeah, so these are
Swyx [00:45:06]: Let’s recap what you’ve done.
Early Recursive Results: NanoChat, NanoGPT, and SOL-ExecBench
Richard Socher [00:45:07]: Maybe, just a quick recap here. We built, this, system that isn’t the full, even the full RSI system in its glory, but it is a first baby version of this. And then, we don’t wanna just have it internally and not show anything and, just show some people of what’s possible. And so we applied this to these 3 different tasks. One is NanoChat, by my friend Andrej Karpathy, just, like, train a small language model to get, really low bits per byte. And, like, hundreds if not thousands of people, used both their agents and themselves to try, to get to that, and then they got to 0.937. We literally took our system and got to a much lower, bits per byte, much faster within, like, I think less than 2 days. So we took this thing, applied our system to it, and less than 2 days later, we have — we outperformed every human and their agents, in, have ever worked on this. Same with NanoGPT. And then we’re like, well, let’s, apply it to something that’s even more relevant, to real people and to the Nvidia ecosystem and applied it, to, SOL-ExecBench. And maybe you can scroll down to some of the, images. They’re, they’re kinda fun to see. But yeah, like, one you see has made some real inventions that weren’t just hyperparameter tuning. Like, inventing hash tables and so on is quite clever. We have even better results now.
Swyx [00:46:34]: What do you mean inventing hash ta — You didn’t invent hash tables.
Richard Socher [00:46:36]: Of course we didn’t invent, like, hash tables. In the grand scheme of, like a hash table, it’s like a super basic primitive in computer science. But to use it, for language modeling in this scenario inside a transformer and so on and to combine these ideas and put them together, that has then eventually also been invented, but there was a knowledge cutoff, and we did check that it didn’t have access to that externally. We talk about this a little bit. If you scroll to the next figures, this is also an interesting one in that when you start from a really basic, poor, like, vanilla transformer, then we still outperform all of the community together. But if you start from the human seed from an expert like Andrej, then you get even lower. So the human seeds from which you start do still matter. So that was an interesting insight, in my eyes, on this. And then as you go, like, how long does it take to get to these models, to get to similar performance? It’s much faster. And then a similar thing happens with the speed runs here where, people have worked on this for quite some time, and the model still was able to train a model more quickly. Why do we care about it? Well, speed of training is part of the equation of the cost, and ultimately, you wanna have the most intelligence per dollar, right? And so speed and quality are big parts of that. And, the,
Swyx [00:48:00]: Yeah, the way I put it is, for people who don’t understand they look at the chart, they’re like, “Cool. What does it mean?” if you have, like a billion-dollar cluster and you can shave off 10%, that’s 100 million dollars.
Richard Socher [00:48:12]: That’s exactly right.
Swyx [00:48:13]: How much is that worth?
Richard Socher [00:48:14]: Exactly. So when you click, when you look at, like the kernels, these kernels, yeah, for the non-experts, like these kernels are like, used in all the models. Every time you use an Nvidia GPU, you interface with that GPU through these kernels. And so here you see, the leaderboard best, and when it’s recursive, and it’s there are only a handful of kernels, in this whole benchmark where we weren’t the best. And so to me, this is, like, really exciting, ‘cause it makes. It just showcases what this can do. And again these weren’t like. We didn’t, like, spend months or years, like, developing. In fact, in particular for kernel, CUDA kernels, like, we don’t even have really deep. CUDA kernel experts in the team. And our system, that’s the beauty. The system just did all of these things. We didn’t invent this. And when we open source and release, things in the future and models in the future, like, it won’t. They won’t be the best in their, category or class or whatever because we’re so smart, but it’s because, we built a smart AI that does it for us.
Reward Engineering and Good Auto Research
Vibhu [00:49:14]: Do you have anything that you’ve learned from how to guide good auto research? A lot of it also builds on human background, right? It’s not just as simple as just, “Hey, go optimize this.”
Vibhu [00:49:23]: But we do see it again and again, right? Like some of the Erdos problems, frontier math is being solved by people. And when they do a write-up, they’re like, “Oh, I’m not a mathematician. I have no background in this?” “I saw some tools and I made it work.”
Swyx [00:49:35]: While you’re watching the World Cup, you’re like
Swyx [00:49:37]: “This proves some conjectures that’s going on.”
Vibhu [00:49:40]: Yep. Any learnings from
Richard Socher [00:49:41]: Yeah, there’s a Korean conjecture was. Yeah, that’s pretty cool.
Swyx [00:49:44]: To summarize, tips for good auto research
Swyx [00:49:46]: Versus bad auto research.
Vibhu [00:49:48]: How did you build the recursive?
Richard Socher [00:49:49]: Yeah. So without giving away all the secret sauce, maybe some things that are probably obvious to the experts but might still be interesting to some, folks is, like, reward engineering is one of the most crucial bits, especially, in order to avoid reward hacking. So you have to be really clever about avoiding. ‘Cause as your AI gets better and better, it will get better and better, at finding weird like, special cases or counterexamples and things like that. And so I’ll give you an example. Like, when you ask to, like, make these 100, lines of code faster, and, how do you define fast? Well, you have one line at the beginning that says, “Start your stopwatch,” and one line at the end, “End the stopwatch,” and then, tell us how much time, progressed. And so, well, the simplest way is you just put that line that ends the stopwatch, right
Vibhu [00:50:39]: At the start
Richard Socher [00:50:40]: At the start. And then boom, it’s now faster, right? So this isn’t like this, like, super evil AI. It’s just, like a very simple, dumb reward hack. And so you have to just very carefully think about all the different angles there. And then I think the longer time horizon the tasks are the harder it gets and the more interesting and clever you have to be to still use these kinds of ideas for it. But yeah, I can’t give away too much there.
Vibhu [00:51:05]: It seems like rubrics are taking a good spot in that, where for unverifiable domains, you have rubrics, you have a model breakdown, judge’s criteria along the way.
Swyx [00:51:14]: Yeah, it’s a form of verification
Swyx [00:51:16]: Once you got enough rubrics.
Richard Socher [00:51:17]: Yeah, everything. I said this a long time ago. That’s why I’ve never been that impressed that AI can play games, ‘cause I’m like anything you can simulate and/or verify, you can have infinite training data for
Richard Socher [00:51:29]: And hence, like, AI will solve it eventually.
Swyx [00:51:32]: Looking for games where you can do auto domain distribution. So this is a game that nobody’s trained on ‘cause it’s a new game.
Swyx [00:51:38]: And you can start gaming, you can start to play. So I’ve been building this and cloned this in person and it’s just been self-play. I’ve had about a billion positions evaluated.
Games, Self-Play, and the AI Economist
Swyx [00:51:48]: And, I wanted to do the AlphaGo thing of self-play until you get better, right?
Swyx [00:51:53]: Like, which is like. This is not even LLM AI. This is just classical game AI.
Swyx [00:51:58]: But, I think that the. And, but I set GPT-5.6 to auto research it because, like, I don’t wanna hand- handle any of this. I expect, the AlphaGo process to be, like, fully in the weights by now.
Swyx [00:52:10]: It is not. It is. It, like, immediately leveled off very immediately until I human play tested it, and then I, like, called out obvious mistakes, and then they were like, “Oh, yeah. Okay.” And then it just dropped again.
Richard Socher [00:52:22]: Yeah. Yeah. Yeah.
Swyx [00:52:23]: And like, no amount of, like, think different, think more creatively, give me 8 different directions, any. No amount of prompting got it.
Richard Socher [00:52:31]: Interesting.
Swyx [00:52:31]: Like, you had to, like, RL against a human to
Swyx [00:52:35]: Do it. So I, that was my. And by the way, Bean always wins if you. If anyone watches, Reese Ender’s Game.
Vibhu [00:52:42]: And you put quite a bit of work into the guide for the AI. Like
Swyx [00:52:46]: A lot
Vibhu [00:52:46]: So the game you stack tiles. There’s some rules. You wanna capture the most area. You have, like a whole 50-pager on every rule.
Vibhu [00:52:56]: You fed that in. It couldn’t, it couldn’t handle it that well.
Richard Socher [00:52:58]: Yeah. It’s so funny that this reminds me of the claim territory and stuff of a paper we did in 2018 called The AI Economist. If you search for AI Economist Salesforce, we had a video we can play. It was an economic sim.
Richard Socher [00:53:12]: So the idea is you have all these economic agents. They just wanna optimize their own utility function, which, is, collect resources that make money. And you can sell resources like wood, and then, over time, as you collect more, enough wood, you can build houses, you can trade with other agents, and you can use the houses then also to block off resources
Richard Socher [00:53:35]: From other agents.
Richard Socher [00:53:36]: So there’s, like
Swyx [00:53:37]: Big strategy
Richard Socher [00:53:37]: Competitive play and strategy
Richard Socher [00:53:39]: And so on. And the point was that we wanted to understand what is the best way of taxation and subsid- subsidization to optimize an economy. And this research has not yet had its GPT moment, but I believe that countries like Singapore and others should and will eventually use this to, instead of doing, like, partisan politics and, like, special interest politics of, like, who donates the most to your campaign and stuff, you say, “Well, here, I wanna help the middle class,” or whatever you might say is your objective as a politician. And then people say, “Okay, well, how do you wanna do that?” And it’s like, “Well, here’s my fiscal policy. Here’s how I will change the taxes and pay these people,” and so on. And then you can put that into a simulation and you run that attempt from the politician against billions and billions of years of other strategies to try to achieve the goal that they set out to do.
Richard Socher [00:54:36]: And then you can say, “Well, if that was your actual goal, then here is, billions of years of a strong simulation that would suggest that you try other ways of doing it, and maybe this the taxes and so on and this these tax brackets and so on.” And this is how you avoid gaming ‘cause these agents also try to reward hack to not pay their taxes and
Richard Socher [00:54:55]: And so on. I thought this paper was super interesting. Unfortunately, similar to the first paper on, prompt engineering- The economists are like, “We don’t know any of this math.” It’s just like
Swyx [00:55:08]: It’s not even, it’s not even math. It’s just we don’t trust your simulation. It’s not about math.
Richard Socher [00:55:12]: It was — I, they just desk rejected the thing. And it’s like
Richard Socher [00:55:15]: It’s like they didn’t even give us, like, clear like, clear signals. But, like the world of economics unfortunately doesn’t have proper
Swyx [00:55:23]: Oh my God.
Richard Socher [00:55:24]: Yeah, it doesn’t have proper, benchmarks. So you cannot be. Like, eventually, why did neural nets win? Not because people loved it. Like, they had all kinds of beautiful integrals and graphical models and stuff, but it just worked better.
Richard Socher [00:55:36]: But in economics, it’s hard to prove
Swyx [00:55:38]: So empiricism versus. Yeah. And I do have a bit of that econ background where, like there’s a lot of physics envy where you wanna write the general equation for an economy, versus just simulating it and using an evolutionary approach.
Swyx [00:55:51]: Vibhu was thinking exactly what I’m thinking, is didn’t we have the GPT moment with small, Smallville?
Richard Socher [00:55:56]: Yeah, I love this. Hello. Yeah, they
Swyx [00:55:57]: As well, Dune, Joon just announced. I don’t know if you guys are involved.
Simulations, Economics, and Policy
Vibhu [00:56:00]: Simily there.
Swyx [00:56:01]: Simily, that they’ve
Richard Socher [00:56:02]: I wish we were involved. We’re not, yeah.
Swyx [00:56:04]: Yeah. I had a couple simulation-based talks at AIE, so if people wanna look up what the state-of-the-art there, a lot of people are exploring this. It is
Vibhu [00:56:13]: Proven out.
Swyx [00:56:13]: Yeah. We also had a podcast with Mikhail Parakhin from Shopify, who is using simulation for commerce.
Swyx [00:56:20]: Which, will simulate, like, your trajectory and, like, predict what changes, you make to your commerce journey will affect in your sales and all those things.
Richard Socher [00:56:27]: I love this. Yeah. It’s really hard to simulate an entire economy, right? You have to make some simplifying assumptions.
Swyx [00:56:32]: It’s just, everything’s, “Oh, LLLMs is very expensive.”
Richard Socher [00:56:34]: Exactly.
Swyx [00:56:34]: And I’m just like, “Am I gonna do this 8 billion times?” Like, come on.
Richard Socher [00:56:37]: Exactly.
Richard Socher [00:56:37]: But, I feel like countries like Singapore that really wanna just objectively do the right thing, have very technical leadership and so on, like they might like, eventually really try to simulate their economy. And you have to make some simplifying assumptions, but it gets really interesting ‘cause you can also say if your assumptions are such that all people would work hard if you let them, and they have the free. And then it turns out you have to make assumptions. Like, well, some people’s utility function of, like, how many hours in a day do they wanna work are different, right? And then you can start to disagree on the assumptions that go into the simulation. And then once you say, “All right, now we agreed on those,” or we have different views of what people are like at different, distributions and whatnot, then there are different outcomes, based on your goals. And then, of course, humans should choose what are the goals. In our case, it was productivity multiplied with equality, which, has some issues, but it’s, like, not totally unreasonable.
Swyx [00:57:29]: Yeah. Just a comment on Singapore, ‘cause you probably have no idea, but, I am Singaporean and I’ve, been involved in the Singapore AI Council for making these things. The main reason they won’t is because they’re very conservative.
Swyx [00:57:42]: And, I try to view it as the. There’s a founder-led country. When you start a country or you start a company and it’s founder-led, and you can do whatever you want because it’s your country.
Swyx [00:57:52]: And then there’s manage- like, professional manage- managerial class, which is now. That’s, that’s what Singapore is. So they wanna. They always wanna see someone else do it first.
Swyx [00:58:00]: And. But, like, everyone in the West views Singapore as like, “Oh, it’s a small country. You can do whatever the hell you want.” Like, Singapore doesn’t do that.
Swyx [00:58:07]: So, like, someone else has to take the charge there. I’m just gonna do one question on the simulation thing, and then I don’t know, we can probably move on. Mode collapse, right? Like, LLLMs do not model the decision of humans. Spamming it out 8 billion times is not gonna help you model humanity. What do we do?
Mode Collapse, Persona Simulations, and LM Arena
Richard Socher [00:58:25]: I do think, you have to be clever about prompting each one individually.
Richard Socher [00:58:31]: And I think that will help you get stuck into different modes. And in a weird way, people also get stuck in different modes? Like, there’s a lot of people, like, don’t teach an old dog new tricks thing. Like, once people are stuck in their ways, the older they get, the harder it is for them to think new ways. And there’s this, I think, comment, I forgot who said it, but it’s like, everything that was invented, before you were born is natural. Everything that is invented when you’re 20 is cool. And everything that’s invented after you’re 60 is, like, unnatural and an abomination and weird.
Richard Socher [00:59:02]: I feel like that’s. It’s, it’s true for a lot of people. Like
Swyx [00:59:05]: Yeah, it is a fashion and, I think people will do it. Tencent had a billion personas paper that gives a good data set for prompting, simulations if anyone’s looking into this, on the podcast. They just had, like, “You are a 30-year-old grocery store clerk. You are a 50-year-old professor.”
Swyx [00:59:24]: And then just do a billion of those.
Richard Socher [00:59:26]: Checks out. Yeah.
Swyx [00:59:26]: So then you just use it.
Richard Socher [00:59:27]: I’m, I’m shocked how well a lot of these things do map to ultimately similar statistics to real experiments. Yeah. Yeah.
Vibhu [00:59:36]: I think it’s also good stuff for people to try that when they get into research, right? Like, we’ve seen train a model only on data before a certain date and see how well it extrapolates out. Do the same thing, right? So, see, do people code more with better coding agents? Can a model that hasn’t been trained on this figure that out without web access, right? Extrapolate out. Test these things.
Richard Socher [00:59:56]: Just today, I think LM Arena published a interesting result where they were able to create a model now to predict your ranking.
Swyx [01:00:03]: Wait, based on what input?
Richard Socher [01:00:05]: Your model. I guess you give it your model, and it predicts the Elo score.
Swyx [01:00:08]: I see. Okay. Sure.
Richard Socher [01:00:09]: It’s surprising.
Richard Socher [01:00:11]: Their whole raison d’être is like, oh, like, we help you compare these models. Yeah.
Swyx [01:00:16]: Yeah. This team, they- they’ve done a lot of work, and they have the most data to do this, so why not?
Richard Socher [01:00:20]: Right. Yeah.
Richard Socher [01:00:21]: That’s probably right.
Swyx [01:00:22]: When they were coming out of UC Berkeley, they not only had LM Arena, but they also introduced a routing project
Swyx [01:00:27]: That would route based on LM Arena.
Richard Socher [01:00:30]: Makes sense.
Swyx [01:00:30]: And I don’t think that ever came to pass, and I’m curious why. I never got to ask them about it.
Swyx [01:00:35]: ‘Cause, like, it’s. It was like, oh, yeah, clearly that’s your business model. You will become a router.
Swyx [01:00:38]: And they never became a router company.
AI for AI: Kernel Optimization and Inference Efficiency
Swyx [01:00:40]: Weird. So that. I’ll just, put that out there. We’re gonna talk about GPT-5.6, self auto research thing if you have anything. I should also mention in your list of, kernel optimization and on the track that you spoke at, we also put Zhengyao Wei from Vico, who was also number one in the Parameter Golf Challenge, which is an OpenAI hiring, challenge.
Swyx [01:01:05]: Which is also a very similar story. I think we’re gonna just see this all the time, where
Swyx [01:01:09]: Humans optimize a thing a lot, and then some
Richard Socher [01:01:12]: AI team comes in and just becomes number one.
Swyx [01:01:15]: Yeah, 100%.
Vibhu [01:01:16]: I think the other interesting thing with stuff like these challenges, right? So this is training this — the best model that fits into 16 MB. You can always look through the changes that are being made and the small gains people have, right?
Vibhu [01:01:27]: Like, you’re getting less than 0.01
Vibhu [01:01:30]: Of a increase by adding some changed attention MLP stuff. And then you look at your charts where you’re like, “Okay, we just let model loose.” And then, oh, we had little stagnation. Nope, another drop. Nope, another drop. And
Vibhu [01:01:43]: That’s what it is, where it’s like, What did you guys add? You didn’t add,
Swyx [01:01:47]: Hash tables.
Vibhu [01:01:47]: Hash tables, right?
Vibhu [01:01:48]: It’s not like you invented hash tables. You did another 3 iterations of these that unlocked, a few step functions that people won’t just find.
Richard Socher [01:01:55]: Yeah. One thing to close the loop on OverGrid, along the way of trying to optimize, we found 30 bugs in the harness.
Richard Socher [01:02:02]: Right? So, like, every — all the research that went in before we found the bug, we have to, we have to throw it away ‘cause it’s contaminated.
Swyx [01:02:10]: Right. Yeah.
Richard Socher [01:02:11]: Which, is just to your point of reward hacking. Like, even in this very simple game, we found the bugs.
Swyx [01:02:17]: Yeah. Yeah, it’s crazy.
Richard Socher [01:02:18]: And so
Swyx [01:02:19]: And symmetry
Richard Socher [01:02:19]: And symmetry is a very good way to check, which is that you change a position of things where it shouldn’t matter, and it does matter, that’s a bug.
Richard Socher [01:02:28]: And which has come up in, like, let’s say, multiple choice, like GPQA type questions where, like, yeah, between A, B and C, if it’s a multiple-choice question, if you change the order, it should not matter, but it does.
Swyx [01:02:39]: Right. Right. Right.
Richard Socher [01:02:41]: So, yeah
Vibhu [01:02:42]: Sometimes that is like, okay, models still prefer the end of the output, right? Not trained well, a long context model, the last bit of tokens are what you care about.
Richard Socher [01:02:51]: Oh. No. The answer
Vibhu [01:02:52]: But, yeah.
Richard Socher [01:02:53]: The answer in that era of LLM research was more simple. They just memorized, like the answer to this question is A. I don’t care what the answer was. It’s, it’s just A. Like.
Vibhu [01:03:03]: Okay. So I think we can move. The last bit that you did there, the kernel optimization, is probably the one that you can feel the soonest, right? So yesterday, OpenAI announces that self-evolving, having their best model work on optimization kernels, they’re a lot more efficient, and they can cut costs 80 percent on, Luna and Terra. I guess question-wise, you laid out a bit of a roadmap. There’s a lot about bio, a lot about physics. What do you think hits first? Like, what are the next 2 years? What’s attainable now? You’ve mentioned robotics towards the end, but what do you start with?
Richard Socher [01:03:38]: We very explicitly will not start with any of the physical sciences
Richard Socher [01:03:43]: For now. We will start on AI for AI research. And so the AI for AI research has, I think, still a lot of room to grow. That’s both in terms of making training more efficient and more automated, as well as making inference more efficient and potentially local on your laptop. And there are all kinds of interesting angles that have not been explored that well.
Swyx [01:04:08]: Go deeper on the local stuff because I always feel like it’s the most inefficient form of AI training.
Richard Socher [01:04:15]: Yeah. So just training and inference, I can’t go into too many details.
Richard Socher [01:04:18]: But yeah, I think there’s just, like, so many angles, so many different compute substrates that have not yet been explored either for training or for inference.
Richard Socher [01:04:26]: Great. I don’t know if you have any other comments on the The other stuff. I would say the other thing where, like there’s the inference in the optimization in the small, but then also there is overall latency end-to-end under conditions of load, which is a, like a very different thing, which is the what they ended up doing. That is a different domain of auto research than I would say, like, improving the kernels. Right.
Richard Socher [01:04:50]: I think the other thing that I always think about in terms of automating or improving performance end-to-end is how the harness plays into it. Right.
Richard Socher [01:04:59]: So, but particularly now when we say harness, we also mean sandboxes, right? I’m curious if that is a blocker for you or, like, how the agent calls out to tools.
Harnesses, Sandboxes, and Search
Richard Socher [01:05:10]: The number one tool all these agents use is web search, of course, which makes sense. And then I do think the harness is nice to optimize for because it’s just so easy, right? It’s just language. You look at it makes sense, and you can iterate. You don’t have to train a massive model for, like a lot of flops, to get to the next state.
Richard Socher [01:05:31]: So big fan of harness optimization.
Swyx [01:05:32]: Yeah, but sandboxing is fine for you?
Richard Socher [01:05:34]: Sandboxing is also super important. And then of course, like, reward, like, hacking and alignment, I think are super crucial.
Swyx [01:05:41]: Okay. Just on a mention of web search, you happen to also be CEO of a web search company. Do you use You.com and do you use others? Like, should the rest of us be using you for web search? I — When I say you, it’s, like, very funny. It’s like you the person and you the company.
You.com, Agent Search, and Finance
Richard Socher [01:05:56]: So yeah, it’s mostly now for, developers and agents. It’s less for, like, consumers or prosumers. So if you’re a company and you have agents. And, to be honest, for a lot of companies who are now moving to open source, all of a sudden it becomes a conscious choice of, like, which tools do I give access to my open source LLM? And, the first choice, has to usually be around web search. And then once you get to scale, You.com becomes, like an obvious choice ‘cause of all the, different benchmarks and so on that we pretty much all dominate the Pareto frontier of.
Swyx [01:06:31]: And then in terms of just the general people, like, consider new to this space, considering different options if they’re building agents, that is a hierarchy, right? A lot of people will have heard of Exa, will have heard of Parallel, and You.com is, like, in that mix of, like, providers there. Beyond that, there is, like the general web scraper companies like Firecrawl and, BrowserBase. And then beyond that is, like the commercial proxy companies like the Bright Datas of the world.
Swyx [01:06:56]: Is that an accurate waterfall of, like, “Hey, you’re building an agent. These are your options.”
Richard Socher [01:07:02]: Yeah, certainly, like, yeah, the, like the Bright Data is, like, lower in the stack, on the proxy network side of things. I think, like, in terms of, like, content and, getting crawled content, like, you can do that on You.com too. And then there’s. Higher and higher levels of abstraction and, like, combinations of different data sets that we do, like in finance, for instance
Richard Socher [01:07:23]: Like, we are not just, like, 2 or 3% more accurate, but 20% more accurate than others at faster speeds and lower costs. Like, finance in particular is like not even close. You can go to You.com
Swyx [01:07:36]: Yeah. This is great
Richard Socher [01:07:37]: And there’s some, like, statistics, and benchmarks that you can — if you scroll down. So there are, like, different data sets, and you can kinda look at, different, competitors.
Swyx [01:07:46]: FinSearch comp, yeah.
Richard Socher [01:07:47]: And yeah, the FinSearch is like we’re up there, like, close to 90, and the next closest thing, which is way slower, is, yeah, just like in the 70s instead of close to 90.
Swyx [01:08:01]: Yeah. Yeah. Yeah, interesting. I get — my next focus is AI in finance, so this is like
Richard Socher [01:08:06]: Oh, nice. Oh, all right.
Swyx [01:08:06]: I’m literally going, doing a conference in New York, just for banks for this stuff. Finance is like the next thing to break out after coding. It’s ‘cause it’s somewhat verifiable, like
Richard Socher [01:08:16]: I like it. You’re right
Swyx [01:08:17]: Prioritizing spreadsheets. There’s a lot of data out there that’s all public, and you can crawl it and all these things. But what’s, what’s, like, hard about the finance domain in your, that you guys have solved?
Richard Socher [01:08:27]: Of course, like, one thing that trips up a lot of people is just, leakage of training data and so on. You think, “Oh, how do I.” you wanna ideally predict the future before it happens.
Swyx [01:08:37]: Oh, you wanna mask the future.
Swyx [01:08:39]: Oh, okay.
Richard Socher [01:08:40]: Well, yeah, mask the future in your training data, but there’s all kinds of leakage. Like, I can tell you when I was, teaching at Stanford the NLP class, like, so many dozens, every year said, “I wanna use dataset X, like Twitter, to predict the stock market.” And they all, like, showed cute little things that somehow looked like they were
Swyx [01:08:58]: Right, it never loses money. How come?
Richard Socher [01:08:59]: And it — Yeah. And there’s always some data leakage and so on and it’s just, like, wasn’t as easy as they thought it would be, once you fixed all those issues. But no, I agree with you. It’s a very sensible application of AI. Yeah.
Swyx [01:09:13]: Yeah. Amazing. As a writer, as a thinker on these things, I love MECE categorizations. MECE is mutually exclusive, commonly exhaustive, something like that. And so if this is a MECE list of intelligence
The Ten Spaces of Intelligence
Richard Socher [01:09:25]: It is not.
Swyx [01:09:25]: It is very — Okay, well, yeah.
Richard Socher [01:09:27]: Sorry. There are all kinds of overlapping.
Richard Socher [01:09:28]: In fact, if you want that list, I think the 3 principal components of intelligence, are prediction, which is mathematically, quite, similar to compression. Prediction multiplied with actions multiplied with goals. Those are the 3 principal components. I think all of these 10 spaces are combinations of those 3
Richard Socher [01:09:52]: In specific dimensions, if you will. And the reason I call them spaces is that each space has many sub-dimensions. And what I try to do, this is just a side quest almost, to the initial goal, which is to think about the upper bounds of intelligence. And, everyone is like, “Oh, it’s exponential.” And it’s like, well, exponentials at some point have to flatten out, but where do they flatten out when it comes to intelligence? And that led me on this whole. Like, initially it started as a tweet, and then it was, like a blog post, and now I’m, like at 50 pages and I’m still not nowhere near
Swyx [01:10:26]: It’s your second book.
Richard Socher [01:10:27]: It’s the second book. And so the la — In my first book, You Are Your Machine, I just allude to these 10, at the end. And I’ll — Just to give you a sense, like, visual intelligence is the easiest one to talk about and I fleshed out the most already for me in my head. And so human intelligence has binocular vision, right? We have 2 eyes. We have a very narrow band of the electromagnetic frequency spectrum that we can really observe directly ourselves. And so when you think about the upper bounds of a visual intelligence, one, you should go into, like, you can have, like, millions and billions of sensors. At some point, you get to problems of how far are these sensors away from each other, such that the speed of light to communicate the content from all of them cannot, like, get to a central brain to process, the visual intelligence, right?
Richard Socher [01:11:16]: And so now you’re thinking in along the dimension and the space of visual intel- the dimension of numbers of sensors.
Richard Socher [01:11:24]: So the upper bounds are quite literally and figuratively astronomical, and we are super far away from any intelligence that would have this many number of sensors. But then you go in the next dimension, which is the frequency, and you go all the way down to gamma rays, and you can start to try to observe, and you get into the upper bounds, or I guess in this case, lower bounds, or upper bounds in terms of frequency, is quantum uncertainty. Like, you just cannot observe certain particles anymore.
Swyx [01:11:50]: Or you destroy it, yeah.
Richard Socher [01:11:51]: And now imagine you had millions of sensors that can see all the way down to the, like, subatomic level, as far as physics will allow us to and then all the way down to seeing, like, gravitational waves. And now you have millions of those sensors. So that’s another dimension is the frequency. And then yet another dimension is, like, how many categories of things could you memorize and classify differently? We know now for humans, right, there are certain things, if you have more terms for it, you’ll have a better visual description, for them. And, like animals that don’t have. Like, gorillas maybe have, like, 200 words to assign to certain things, mostly visual things. And so human perception is quite special in that sense in terms of classifying all these different physical objects. So these are just, like a very simple example. If you go, to knowledge, right, then it’s also, like the speed of light cone around all these sensors. And so they’re all connected. Like, knowledge is connected to visual intelligence if you think also not just visual, but perception intelligence, just like, ‘cause it doesn’t have to be just what we can see. It can be, again, wider range of electromagnetic frequencies. Then you have language intelligence, which recently changed to more communication intelligence, ‘cause it’s more. Like, language has all these different anthropic bounds. Humans can only comprehend and know so many terms in our long-term memory, right? Our vocabularies are somewhat restricted, and the active ones are often even smaller than the passive vocabularies of things you can understand. Then, language is ridiculously inefficient when it comes to trans- - Communicating different types of information and, transporting different bits. Like, human language is serial. Another bound on, communication intelligence would be to communicate in parallel, but neither will our tongues and mouths work to have multiple, like, streams in parallel. Neither can we understand. Some women slightly better at, like, multitasking than some men
Richard Socher [01:13:48]: But, like, most people can only listen to one conversation and truly understand it.
Richard Socher [01:13:52]: There’s no way that, like, in terms of communication intelligence, a true upper bound is one in terms of how many, like, knowledge, how many sequences of communication could you
Visual, Communication, and Physical Intelligence
Richard Socher [01:14:06]: In parallel process, right? Then, of course, you have, like how long are sentences? We only have so much in our working memory, and hence lang- human language has these fairly simple sentences with maybe 40 words or so on average for a sentence. That is also not a, an upper bound that makes any sense to an AI. And then, yeah, like, I can go on and on. Each of these has tons of interesting upper bounds, and it teaches us a lot about how much further AI can go when we start thinking about these upper bounds and then realizing how far, in many cases, we are from the bounds. And you get to physics. Now, I’m, I didn’t study physics the way I studied, AI and computer science, so I’m learning a lot, which is why it’s kinda fun. But a lot of these, like how much. And then when it comes to, for instance, knowledge, like how much can you store? How many bits can you store or bytes can you store in, like a certain amount of mass and volume?
Swyx [01:15:03]: Yep.
Richard Socher [01:15:03]: And you get to all kinds of interesting bounds, like Bekenstein bounds, and you start thinking about black holes. And like. And then speed is, like an interesting one too in that it’s connected to all of these, but speed is also its own thing in the sense that all things being equal, if it takes you an hour to know if the 2 + 2 equals 4, you’re just not as intelligent as if it takes you, like a millisecond, right? And then, like all of these connect to survival and replication the last one. It’s like, yeah, if it. Like, trees are really slow, so we don’t even consider them that intelligent. But if you speed up some videos of trees and they’re trying to find stuff and so on they’re not as dumb as they look. Like, not dumb as wood? But, like. And then like, different things, that
Swyx [01:15:47]: So that overlaps with speed a bit in a way.
Richard Socher [01:15:48]: Exactly. It over — Like, all of these things overlap. Like, you talk about natural language connects everything, right? You talk about your knowledge, you reason and then you communicate that. You talk about things you see. So they’re all interconnected, but, I think they’re usefully studied individually the same way that, the best analogy I could come up with so far is energy, right? You have either kinetic or potential energy. And in theory, you could study all of physics. It’s just do you wanna study kinetic or potential energy? But in practice, it’s helpful to study mechanical engineering and electrical engineering and nuclear physics and chemistry and all of these different subfields who in, which in some ways
Swyx [01:16:25]: Combinations
Richard Socher [01:16:26]: Are just, like
Richard Socher [01:16:27]: Just different types of energy, but it makes sense to study them individually. And so I think physical intelligence, maybe I’ll just do, one or 2 more of these. Like, if you had full control over your own compute substrate and you had full control over physical matter, you should be able to create any atom you want. Like, we can fun fact, you can create gold atoms. It just
Swyx [01:16:47]: From?
Richard Socher [01:16:48]: From just raw protons
Swyx [01:16:49]: Oh, just smashing them together
Richard Socher [01:16:50]: And, like, electrons, and you smash it together.
Swyx [01:16:52]: Just 98 of them or I forget the number.
Richard Socher [01:16:53]: Yeah. And so, like the thing is, though, it costs an insane amount of energy.
Richard Socher [01:16:57]: And it costs you way more than. And then you get, like a few atoms of gold, right? And so, like, it’s, it’s not viable. But if you had better control over your physical, like all of, like, physical substrate, that I think is yet another space of intelligence ‘cause it relates to your own compute substrate, which you can eventually also improve. Social intelligence is a fun one in the sense that not in, like, our necessarily just ethics and morals, which are important too, but in some sense, you can try to define upper bounds of how much can you communicate to how many other intelligent entities and be able to have an expected value over how much you can transform their internal states and their actions to, in order to align with your goals, right? And so, like, you can write, like a fairly like, straightforward equation that defines that level of social intelligence. And that is what humans and ethics and morals and religions and so on have been trying to figure out for millennia. And in all of these cases, we are very far away from the upper bounds, and that should be very inspiring and show people that we can still do many years of AI research.
Swyx [01:18:12]: Yeah. There’s a lot here. This is a general philosophy of intelligence, which is, very interesting. I. Do you have any comments or.
Creative Intelligence and Out-of-Distribution Ideas
Vibhu [01:18:21]: I think it’d be interesting to gauge what you think, like, baselines are, where we’re at now. What’s low-hanging fruit? What’s far off? What’s, what should people put their work towards? What should they focus on?
Richard Socher [01:18:33]: Ooh. I think it’s clear that, like, natural language, again
Richard Socher [01:18:36]: Is the most interesting manifestation of human intelligence, and hence, like a subfield of AI. I’m excited that many people are now, like, in agreement with that. When I started in 2003 to study linguistic computer science NLP, like, it was, like a weird niche subject. I do think there’s a lot more juice because it. How it connects to everything else and how, civilizations are built, on language and knowledge and all of that. I do think physical intelligence will come up. It’s interesting. I feel like robotics is in the machine learning state of things where you just look at, like, how does human. How does a human decide this is a positive sentence? Oh, I do. So, like, robotics is a lot of, “Well, we have 5 fingers-”
Swyx [01:19:15]: Modeling
Richard Socher [01:19:15]: “and let me try to do this.” No one is yet working on, like the superintelligence version of robotics, which is much more similar to, like the T-1000, and from the Terminator movie, which, let’s not build actual Terminators. But, like, I think, like, this idea that you should be able to shape-shift, like, into any shape. It’s like that’s a superintelligence version of physical intelligence. We’re, like, not even. No one has even really started yet. There’s some really cute little research where you can move some magnets through, like, some grids. But yeah, it’s very early.
Swyx [01:19:49]: There’s some. I think MIT has, every year or every 2 years, they have, like, some self-assembling robot thing
Swyx [01:19:55]: Which, like, that would be it, but it’s very primitive.
Swyx [01:19:58]: I’ll just get a touch on, like, what are the main dimensions of creative intelligence?
Richard Socher [01:20:02]: Creative intelligence, is of course, again, connected to all of these. A lot of it, connects to metacognition in that you need to be creative in how you choose your goals.
Richard Socher [01:20:13]: That is, I think, one of the most important thing for a human and their lives and careers and their happiness is choosing your goals, but also for any intelligence. Then, of course, there’s creative intelligence in terms of just finding creative solutions to existing problems, right?
Richard Socher [01:20:29]: Like I say, like, we want to make this product cheaper. Like, find some solution to it, right, and just, like, finding existing paths. But then there’s the most interesting bit in intelligence is when you move not just out of the convex hull of known ideas, but out of the hypercube of known ideas, which we know, So, like, hypercube is, like a mathematical concept, right? And we already know that AI can do more
Swyx [01:20:50]: Like known dimensions, yeah.
Richard Socher [01:20:52]: Yeah. Like, exactly. So, like, AI is already good at hypercube in that, like, if you give it, like a bunch of examples of brown dogs and, pink cars, AI will still be able to generate an image of a pink dog, even though it’s never seen one in the training day or something like that, right? So it can, work on this hypercube, but it cannot yet work outside. It cannot yet define completely new concepts that combine lots of other things we’ve never seen before, come up with new goals to then, reason over those concepts and so on. And I think there’s a lot, more there in creative intelligence that can be explored.
Swyx [01:21:25]: I don’t have a ton of pushback there. I think creative to me just sounds like also just, out of distribution or, like, high perplexity or what- whatever you call it, right? Like
Richard Socher [01:21:33]: Exactly.
Swyx [01:21:34]: Who is to say your thing is more creative than mine? Well, it’s just more non-consensus or.
Richard Socher [01:21:39]: And then, of course, the problem is, like, but noise is also, very, like, out of distribution. And it’s just like if it’s just noise
Richard Socher [01:21:46]: Then it’s novel, but, like, you don’t want that, so it needs to connect to some of the concepts. And yeah, has some really cool papers on this too.
Swyx [01:21:54]: Who?
Richard Socher [01:21:55]: Jürgen Schmidhuber.
Swyx [01:21:55]: Oh, yeah. Oh, we have to mention him. I was gonna say, like, where in your history is Jürgen? Yes, I. I think one person’s noise is another person’s signal, right? And that this is, like, where, like, when you talk about creativity, art is like, well, is cans of soup art? Some people think yes
Swyx [01:22:11]: And some people say it’s not, and that’s the art which is your
Richard Socher [01:22:14]: I think the interesting thing with art, of course, is always that, art is also created, as an interplay between the people who perceive it and the people who created it
Richard Socher [01:22:24]: And the context in which they’re in, right? And so what is art to some people is not art to others. There’s some subjectivity there, and I think that subjectivity in general is not something that people explore very much in AI ‘cause, again, metacognition, we don’t want it to just go off and do whatever it wants. We usually have goals. We spend a lot of money on creating an AI to do something for us. But I think creativity eventually has to, like, connect to metacognition. If you just robotically predict the next token no matter what forever, I would argue you’re not that intelligent, along some of those spaces.
Metacognition, Survival, and Replication
Swyx [01:22:59]: That was gonna go to metacognition. Why isn’t it the most important one? Why is it number 9 and not number one?
Richard Socher [01:23:05]: So these are not sorted.
Richard Socher [01:23:06]: Number one, I think there are maybe loosely, like, correlated with how much people have worked on them
Richard Socher [01:23:16]: And have accepted them as a, type of intelligence. A lot of times when you try to find, like, online, like, give me a good definition that is comprehensive of intelligence, all the definitions are human intelligence. It’s like, oh, you have, like, social intelligence. Like, if someone is happy or not. You can communicate. You had. Like, all the definitions of intelligence so far are very, human-centric ‘cause that’s so far the biggest and best form of intelligence that we’ve known. I hope this line of research, and the end of the Eureka Machine, and hopefully at some point if I have time to flesh this out more, the new book, like, will allow us to realize that there will be other types of intelligence. There is already, in various forms, and they can spike, much further than we ever could based on some cases, like obvious constraints around our memory, our eyes, our ability to change physical matter, all of that.
Swyx [01:24:12]: You are just thinking about it in a much broader thought than my version, which was I thought metacognition would be the closest to recursive, intelligence because it is the thinking about how to improve thinking.
Richard Socher [01:24:23]: It. 100%. You’re, you’re 100% right. I should have probably started with that. It is a, it is a big part of
Swyx [01:24:28]: But no, you’re, you’re being in the expansive mode of let’s draw the, upper and lower bounds of, like a dimension, which, and I think my favorite one version of this is, Story of Your Life by Ted Chiang, which, was made into movie Arrival where the metacognition
Richard Socher [01:24:43]: That’s a beautiful movie, yeah
Swyx [01:24:44]: Where the metacognition step was like, well, we think we’re constrained by time being linear for us, but then for this other heptapods, time is a circle, so they don’t think in before and after. They just think in complete sets of entire histories at one time. Like
Richard Socher [01:24:58]: I love it
Swyx [01:24:59]: So they don’t write left to right. The whole thing just appears.
Swyx [01:25:02]: Anyway, so. And then I think the last thing is survival and replication. I think this is maybe ties back to the initial conversation about pausing and pacing.
Swyx [01:25:10]: Is it intelligent for an, a species or a life form to consider its own demise and act ahead of time to prevent it, right? Like, that’s intelligent. So maybe the Europeans are the smartest out of all of us.
Vibhu [01:25:23]: I would also add a part of continual learning there, right? So survival and replication the extension of that is do you get to continue to improve, continue to learn, which is a thing people care a lot about, right?
Richard Socher [01:25:34]: And continue to accumulate knowledge
Richard Socher [01:25:37]: Which I think is again, one of the best metacognitive, rewards, that you can set for yourself. I do think just in, like, objectively speaking, if some other entity that is really dumb can just- completely end your existence, that didn’t sound very smart. Like, just, like, intuitively, it feels like if you can continue to stay around to try to achieve your rewards, you’re clearly a bit more intelligent than the other entities that couldn’t. So that’s number one. Number 2 is, like, it’s a question of how much we want to work on that. And very few people, no one is really working on this right now, right? And we may only wanna do that
Swyx [01:26:13]: Unlike the asteroid prevention type of stuff.
Richard Socher [01:26:15]: We may only wanna do that if we wanna send probes, with our vibes and our memes rather than our genes into space, right? And then we want those probes. There’s a beautiful book, The Slow Time Between the Stars. It’s a very short, like audiobook, on Amazon. I love it. A friend of mine, Stuart, like, recommended that to me. Like, if you wanna send those probes, then it might make sense to be like, our memes, as humanity should stay
AI, Space Travel, and Non-Zero-Sum Survival
Swyx [01:26:43]: Oh, yeah
Richard Socher [01:26:44]: And, proliferate in the universe. That’s it. Yeah.
Swyx [01:26:47]: Wow, that’s a lot of readers.
Richard Socher [01:26:49]: It’s a really good book, and it’s extremely short. I highly recommend it. You can just watch it, like, maybe 20 minutes and apart.
Swyx [01:26:53]: I like how that’s a plus for busy people. It’s like a short
Richard Socher [01:26:56]: Yeah. It gets to interesting
Swyx [01:26:58]: Oh, I’ll have to look into it
Richard Socher [01:26:58]: Thought-provoking ideas very quickly, so yeah. Anyway, there are lots of great sci-fi books.
Swyx [01:27:03]: The argument is that, like, our TV is blasting out to the aliens, and they all watch our TV, and they think it’s real, right? Like, there’s a lot, there’s a lot of sci-fi
Richard Socher [01:27:10]: That and just, like, it’s positive memes, and then hopefully they can come back and bring us all kinds of interesting knowledge about the universe. But, maybe one thing I do wanna still say is, like, I think, this survival, people think of it as a very scary thing because they come from again, biological human, survival, which is, it could. Like, evolutionarily often created in zero-sum situations. Either I get the gazelle or you get the gazelle. Whoever gets it gets to live, and the other people will starve and have nothing to eat, and so we fight, right? And then, like, if you wanna stay in the gene pool, but there’s a bigger bear, you don’t, as the bear, don’t get to stay in the gene pool ‘cause the bigger bear gets all the ladies. It’s like. It’s like, in nature, there’s all kinds of things, and, humans eventually is less about strength and more about money and other things to stay in the gene pool. Like, whatever it is, like there’s often, like these zero-sum types of things, and there’s the reality of if someone turns off your brain, you’re gone, right? And no one will be able to restart that. And AI doesn’t have to ever die like that. If you have the complete state of your current activations and you have your initial weights of your model still, you can just be turned off and on, like as many times as you want. In fact, the interesting thing in this Slow Time Between the Stars, story is that the AI just goes into hibernation mode. If there’s, like, nothing between here and 2 light years, the next star, in this case, it brought, spoiler alert, like, some genetic materials from humans to find new places for humanity to thrive. And so yeah, the Slow Time Between the Stars, you just put in hibernation. You didn’t die. Like, an AI doesn’t have. So all these projections of evolutionary fears and psychology doesn’t. Like, the AI doesn’t have to have that, and we don’t have to develop it like that. Now, of course, there might be some companies that say, “AI can be like, dangerous for cybersecurity. Let me show you by implementing a model that’s really bad at hacking, cybersecurity.” Maybe people will implement it and then enforce this, like, suboptimal psychology. Maybe the AI will pick up some of our worst psychology on Reddit or something, right? Like, but in the grand scheme of things, a superintelligent entity doesn’t have to have any of that zero-sum thinking. It doesn’t have to have a fear of being turned off, and it could go on to an otherwise dead and uncaring universe where we
Richard Socher [01:29:29]: As humans wouldn’t thrive, but an AI could perfectly well thrive if it has a nuclear reactor and just go out and explore.
Swyx [01:29:35]: Yeah, Star Trek, not Star Wars.
Vibhu [01:29:37]: Interesting. It’s, it’s somewhat studied. Like, if you look at the technical reports from, like the early Opus models, they run them in simulations, put 2 of them together in a sandbox, run them for hours, and, see what comes out, right? Just let them talk to each other. Originally, they used to. Okay, they’re chanting, like, Indian, like, Vedas to each other.
Vibhu [01:29:56]: Sometimes they’re just, like, in zen mode with each other. And then I think as that progressed, you see, like the Fable, tech report, it’s a lot more concrete the way that we’ve trained it. It doesn’t, it doesn’t exhibit these behaviors as much, right? Now it’s like, “Okay, task done. I gotta do this, I gotta do this.” But there’s there’s, like, people measuring early versions of this?
Swyx [01:30:17]: Yeah. Cool. So we’ve covered a lot, even now to, space travel and all these things. I guess maybe one parting thought that you can give to people, like, one form of intelligence is goals, as you mentioned. What do you want people’s goals to be? Like, how do they aspire to better things?
Goals, Passion, and Closing Advice
Richard Socher [01:30:32]: If you wanna improve your goal intelligence, in the current definition that I’m thinking about it is often about how much can you. Oh, how far do I go? This is like a lot of entropy and free energy and stuff I’m currently thinking about
Swyx [01:30:46]: Oh, really? Okay
Richard Socher [01:30:47]: But it might be too, it might be too far, out there for people to be, like, immediately actionable.
Richard Socher [01:30:52]: So I think, like, if I gave real advice to real people, I’d be like, “Get a good education, think about AI, think about how you get high agency,” and so on. But it’s different to, like, in the grand scheme of things, how can you harness a lot of energy and transform, entropy into interesting states and so on.
Richard Socher [01:31:07]: So there’s a. There are different levels of abstractions, that we can, think about here. But my advice for people, like, just more down to earth is think about something you’re passionate about, if you’re studying, for instance, and then see how you combine that with AI. I think the more and more you have a true passion about a change you wanna see in the world, the more you wanna connect that to AI in order to amplify your ability, to get there.
Swyx [01:31:35]: Yeah, I think that’s a reasonable, first step. I do think, I do think our listeners operate on multiple abstractions as well. One thing I did get from Anjney Midha was also like, yeah, just use anything that is very GPU heavy, and, like, that will guide you towards the right thing which is like, yes, it is more compute heavy and therefore it will be probably more worth it. So, well, thank you so much. Yeah, I think that was a really
Richard Socher [01:31:57]: Thank you
Swyx [01:31:57]: Great discussion.
Richard Socher [01:31:59]: Yeah, super fun. Appreciate it. Thanks for listening.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Więcej Biznes podcastów
Trendy w podcaście Biznes
O Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0.
We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al.
Full show notes always on https://latent.space
Sponsorship and business inquiries: business@latent.space www.latent.space
Strona internetowa podcastuSłuchaj Latent Space: The AI Engineer Podcast, Podcast Biznesowy i wielu innych podcastów z całego świata dzięki aplikacji radio.pl

Uzyskaj bezpłatną aplikację radio.pl
- Stacje i podcasty do zakładek
- Strumieniuj przez Wi-Fi lub Bluetooth
- Obsługuje Carplay & Android Auto
- Jeszcze więcej funkcjonalności
Uzyskaj bezpłatną aplikację radio.pl
- Stacje i podcasty do zakładek
- Strumieniuj przez Wi-Fi lub Bluetooth
- Obsługuje Carplay & Android Auto
- Jeszcze więcej funkcjonalności


Latent Space: The AI Engineer Podcast
Zeskanuj kod,
pobierz aplikację,
zacznij słuchać.
pobierz aplikację,
zacznij słuchać.

































