Background

Ryan Greenblatt – What happens once AI can automate AI research?

Dwarkesh Podcast11 de agosto de 20267952
Compartir episodio:Descargar

Descripción del Episodio

Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/HuggingFace incident. In my opinion, he's one of the most interesting thinkers on the future of AI.

Had him on to discuss/debate recursive self-improvement. This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.

I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.

If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.

We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.

We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.

And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.

The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!

Watch on YouTube; read the transcript.

Sponsors

* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh

* Jane Street’s back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn’t tell me what the chip actually does. So that’s the challenge: reverse engineer the circuit and figure out the chip’s purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they’re also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh

* Cursor and SpaceX recently released Grok 4.5, and I’ve been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkesh

Timestamps

(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?

(00:16:52) – Is AI progress bottlenecked by human expert data?

(00:34:02) – Flat token prices suggest scaling has been slow

(00:39:47) – Skills AI can’t train on: does it even need them?

(00:48:07) – Aligned to whom?

(01:09:18) – Recent incidents of AIs colluding and deceiving humans

(01:19:38) – What could possibly go wrong? A concrete scenario

(01:48:02) – From reward hacking to takeover



Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

Episodios Recientes

Ryan Greenblatt – What happens once AI can automate AI research?

Más podcasts de Sociedad y Cultura

Ver toda la categoría →
We!

We!

By shows

<p>The #1 U.S. How To podcast for emotional growth, WE is a movement for reconnection—heart to heart, soul to soul. Hosted by Joe Mittiga, WE explores emotional healing, mental health, personal growth, relationships, and spiritual awakening in a world that feels increasingly disconnected. Each episode reveals why you feel the way you do—and the emotional, mental, and spiritual growth that leads to healing, peace, purpose, and deeper human connection. Through honest conversations and real-life experiences, WE helps you better understand yourself, others, and your place in a changing world. Healing doesn't happen alone. It happens when we remember WE.&nbsp;<br><br></p><p>Visit <a href="http://www.wepodcast.global"><b>www.WEPodcast.global</b></a> to learn more<br><br></p><p>Download my books free at <a href="http://www.joemittiga.com"><b>www.JoeMittiga.com</b></a>.<br><br></p><p><b>Voices of WE</b></p><p><em>Healing humanity, one story at a time.</em></p>

The Permanent Questions With David Brooks

The Permanent Questions With David Brooks

By shows

What makes us human? How can I find purpose in life? How can my relationship survive a rough patch? What do I say to a friend who’s suffering? Does any of this mean anything at all?&nbsp; It’s easy to lose sight of fundamental human concerns amid the churn of the news cycle and the demands of every day. Each week on The Permanent Questions, host David Brooks seeks to return them to our attention through thoughtful conversation with the world’s deepest thinkers.

Simulacro

Simulacro

By shows

Marcos Oliveira, uno de los padres de la inteligencia artificial, desaparece de forma misteriosa en Canarias y lo único que saben es que iba buscando una puerta. El escritor Santiago Álvarez recorre la última ruta que hizo el científico para intentar resolver el misterio pero, cuanto más lo busca, más perdido se encuentra en un laberinto entre la realidad y la simulación.<br /><br />Simulacro es un thriller psicológico especulativo creado por <b>Julio Rojas</b> (autor de 'Caso 63') y protagonizado por Isak Férriz, Mónica López y Alberto San Juan.<br /><br />Producido por El Extraordinario para Turismo de Islas Canarias y financiada con fondos Next Generation de la Unión Europea.

Cisne Rojo

Cisne Rojo

By shows

<p>Sara es la principal accionista de Summer-Petersen, una importante compañía de servicios financieros. Es la ejecutiva más agresiva. La que no tiene miedo. Si se trata de negocios, la información es poder y Sara ve lo que otros no ven. Hasta hoy. Este lunes, reaparece en su vida Miranda Celis, una analista de grandes datos y experta en predicciones basadas en algoritmos. A partir de una sola llamada, Miranda conectará recuerdos del pasado con el anuncio de trágicos e irreversibles eventos futuros. Si Sara piensa que todo es estable, que las personas que ama, que la ciudad que habita,&nbsp;y que todo&nbsp;lo que la rodea estará mañana, hoy Sara deberá pensar otra vez.</p><p>&nbsp;</p><p>Protagonizada por Blanca Lewin e Ignacia Baeza. Creada y escrita por Julio Rojas, el creador de Caso 63. Producida por&nbsp;<a href="http://emisorpodcasting.com/" rel="noopener noreferrer" target="_blank">Emisorpodcasting.com</a></p>

Franco Escamilla Canal Oficial

Franco Escamilla Canal Oficial

By shows

¡Bienvenidos al canal oficial de Franco Escamilla! Suscríbete para escuchar un nuevo episodio cada día. Martes: La Mesa Reñoña Miércoles: Los Amos Del Universo Jueves: Toy Aburrido Viernes: Psico y Psico Domingo: Cabareteando Recuerda que no tengo otro canal de audio, no aceptes imitaciones.

Café Francés ~ Aprende francés!

Café Francés ~ Aprende francés!

By shows

<p>Café Francés es el rincón preferido de hispanohablantes para aprender francés de forma estructurada y paso a paso.&nbsp;</p><p><br></p><p>Aprenderás frases útiles, diálogos reales, vocabulario esencial, gramática y todo sobre la cultura, costumbres y vida en Francia. Cada episodio se conecta con el anterior e incluye ejercicios practicos.&nbsp;</p><p><br></p><p>Además, tú también formas parte de esta aventura! Conversa conmigo y propon nuevos temas en las redes sociales. Juntos creamos una comunidad apasionada por el idioma y la cultura francesa.&nbsp;</p><p><br></p><p>Bienvenue à Café Francés – on y va !</p><p><br></p><p>Sébas</p>

La Cultureta

La Cultureta

By shows

Rubén Amón, Rosa Belmonte, Guillermo Altares, Isabel Vázquez JF León y Sergio del Molino hablan sobre cine, música, libros, series y mucho más...

Espías: clave secreta

Espías: clave secreta

By shows

Hay guerras que no aparecen en los libros de historia. Se libran en silencio, en los despachos más discretos y en las calles más concurridas. Sus protagonistas no llevan uniforme. Sus victorias no se celebran. Y sus derrotas, a veces, cambian el mundo. Espías: Clave Secreta es un podcast de Vicente Vallés sobre los hombres y mujeres que operaron en la sombra durante el siglo XX. Topos que pasaron décadas infiltrados en el corazón del enemigo. Operaciones diseñadas para engañar a ejércitos enteros. Traiciones que costaron miles de vidas. Y espías que, en secreto, evitaron una guerra nuclear. Cada episodio, un caso y cada caso una historia que parece imposible y que, sin embargo, ocurrió. Créditos&nbsp; Dirección y guion: Vicente Vallés y Cristina de Juan&nbsp; Narración: Vicente Vallés&nbsp; Diseño de sonido: Andrés Moraleda&nbsp; Una producción de Onda Cero Podcast y Antena 3 Noticias&nbsp;