08/19/26

Creating ‘world models’ for robots + An AI math shakeup

When you ask an LLM like ChatGPT or Claude a question, the model goes through its massive amount of training data and guesses the answer by mathematically predicting the word most likely to appear next in a sentence.

This model, experts say, will not work well for technology designed to navigate the physical world. Something like a robot that works in a warehouse will instead require a “world model” that can understand spatial surroundings, like the stuff we walk by or bang into.

But what is a world model, exactly? And how do you train AI to recognize what the real world looks like? Host Ira Flatow checks in with tech journalist Joanna Stern, who’s seen the early days of these models up close, even in her own home.

Then, we check in on the math world, where frontier AI models have made meaningful progress on decades-old problems. Mathematician Emily Riehl gives us the big picture on how significant these results actually are.


Donate To Science Friday

Invest in quality science journalism by making a donation to Science Friday.

Donate

Segment Guests

Joanna Stern

Joanna Stern is a tech journalist who writes newsletters and creates videos for New Things Media.

Emily Riehl

Dr. Emily Riehl is a professor of mathematics at Johns Hopkins University in Baltimore, Maryland.

Segment Transcript

IRA FLATOW: Hey, I’m Ira, and you’re listening to Science Friday. When you ask ChatGPT or Claude a question, the model goes through its massive amount of training data, makes guesses on the answer, mathematically predicting the next most likely word to appear in a sentence, for example.

This model, experts say, will not work well in a world that’s not based on textual information. Instead, they argue you need a completely different approach, a world model that can understand the physical world around us, stuff we walk by or we bang into, and enable humanoid robots to work in warehouses and improve self-driving cars. It’s an area of AI development that’s raised billions over the past few years, with big players like Amazon and Google’s DeepMind showing more and more interest. So what does a world model mean, actually, and how do you train the AI to recognize what with the real world looks like?

My next guest has seen the early days of these models up close, even in her home, and is here to tell us how ready for prime time they are. Joanna Stern is a tech journalist who writes newsletters and creates videos for New Things Media. Welcome to Science Friday, Joanna.

JOANNA STERN: Hello. Thank you for having me, Ira.

IRA FLATOW: You’re welcome. OK, tell me more about why tech companies want this different approach to AI.

JOANNA STERN: Well, if you think about the physical world and the type of AI we are thinking we’re going to get in the physical world, best examples would be self-driving cars and humanoid robots or robots of any kind. You can see why just studying the text of the world, and even just studying basic images of the world, is not going to be enough for a car to get you from point A to point B, or it’s not going to be enough for a humanoid robot to step into your house and start doing the laundry.

IRA FLATOW: So a company like Amazon that has to move a lot of stuff around could really benefit from this kind of robot.

JOANNA STERN: Sure. And Amazon has over a million robots already functioning in some of their warehouses and in their partners’ warehouses. And so the big difference there, though, in terms of what Amazon’s talking about in robotics, is that those robots are largely industrial work type robots. They have a very, very regimented, simple environment that never changes.

Think about it. If you are taking things from one box and putting them into another box, the box size never really changes. They’re usually located in the same exact spot. And so this is really simple to program robots. That’s very different than, say, a self-driving car where, sure, you’re still going from point A to point B, but a bicycle just went in between that road that wasn’t there before, or they just set up a road closure sign that wasn’t there before. And so these things are constantly changing in our physical world.

The home is even worse, right, Ira? I mean, I don’t about your home, but my home is changing every day. Someone moves something in the refrigerator. Somebody moves a plate. My kids have put their toys all over the ground.

IRA FLATOW: And so is video very helpful? Is that what we’re talking about here?

JOANNA STERN: Well, so there’s some debate about how it’s going to be best to train these robots. One way is through something called teleoperation, where the robot actually is in the home. There’s cameras and sensors on the robot, and somebody is controlling that robot with a virtual reality headset someplace else. They’re wearing arm controllers, and they are controlling basically as a puppet the humanoid robot that’s in the house.

And I’ve seen this demoed a number of times now. So the robot isn’t doing it autonomously. And so this idea of teleoperation is that we are training and gathering data of the movements of the robot. We’re also gathering video and sensor data from the robot. And that, over time, can start to train.

There’s another idea that’s becoming pretty popular, and that’s we use video data. We take lots of video off YouTube. We take lots of video from humans doing this in the real world, and we train it. And in fact, last month, I did a story on a company called Micro AGI out of Germany, and they have a new app called Shift. And you can say, I want to earn money by filming myself doing things in my house, doing the dishes, vacuuming.

IRA FLATOW: They’ll pay me to do these things?

JOANNA STERN: They will pay you to do these things. The catch is you have to wear an iPhone on your head, some sort of smartphone on your head, or now they’ve started to make hats with embedded cameras.

So I tried this. I tried it for about three weeks. I will say I was not very successful at it. Turns out you have to really– the footage has to show both hands in the frame, and you’ve got to really show it very carefully. But look, I did make some money folding my laundry, putting my dishes in, which is more to say that, look, I would say everyone in the house was happy that I was finally contributing.

IRA FLATOW: Well, I actually watched your video. I saw your video doing that, and it looked very interesting. And of course, as you say, to train the robot, it has to see physically what you’re doing with your hands so that it could tell the robot what to do with its hands.

JOANNA STERN: Exactly. The big thing, and I kind of challenged the CEO on this. I said, well, look, I filmed five hours. You’re only paying me for three hours. He said, well, yeah, because only three hours is useful to me. Only three hours showed my hands both in frame and my hands doing something useful that they think they can then go train the robot on. And in fact, they sent me the footage where they apply the 3D modeling to my hands. And you can see, OK, the robot has to see these weird kind of like almost skeleton like lines in your hands doing things for it to start to make sense of what is happening.

IRA FLATOW: If I were to go backwards engineer things, if I were a Luddite, I might say, you know, instead of a robot trying to do this, my 15-year-old will do this for next to nothing. Don’t have to teach anything. Why go with a robot on this?

JOANNA STERN: It’s true, it’s true. But look, I will say, there’s many times in the week where I have a big pile of laundry that I need to fold, and I would rather do something different. Last summer, I actually had a company come to my house and set up their laundry folding contraption, and that wasn’t a humanoid robot. That was really these two robotic arms.

You can picture– you know when you go to an arcade and they have the claw machine that you can put the money in and it will pick up a stuffed animal or some other toy? The robotic arms looked a little bit like that. And so these two robotic arms sat over the table, had some cameras. It was a big sort of contraption set up. And these two robotic arms just on their own were folding the laundry.

And it didn’t work great, but it was really great to see, because I could see, first of all, it improve over time, and it gave me a sense of where we are at right now. I mean, it was only folding t-shirts. The company now says it folds more things. Sometimes it would take five minutes to fold one t-shirt. This gives us a good idea of where we’re at in this moment.

And look, this was last year and things have improved now. But this is the progress we need to see and why this data is so important to these companies and why there are companies specifically popping up that just collect the data to sell to the robotics companies.

IRA FLATOW: Right. And that brings me to– I’m glad you brought that point up, because I’m sure there are privacy experts who are not thrilled about people sending thousands of hours of footage inside their own homes.

JOANNA STERN: Totally. And I thought about that, obviously, before embarking on the experiment I did in that video with that company, Micro AGI, though, they gave me a pretty long list of the ways they’re protecting privacy. One of them being that they blur out any identifiable information they might see in your footage.

So, for instance, they sent me back a video where some writing on a t-shirt was blurred out. And I said, why is this blurred out? They said, well, we blur any text that we see. OK, it’s was my kid’s Pokemon t-shirt, but that’s fine. But these companies are taking steps towards that. But look, you’re agreeing to do this. And if you think the trade off of putting some cameras on your head and showing them your house is worth the $20 they might pay you an hour, well, that’s your decision to make.

IRA FLATOW: How do you see this panning out? Do you think this model is going to catch on like the large language models did or do you think this is just the beginning? I mean, folding a shirt for five minutes has a long way to go.

JOANNA STERN: Well, that’s one of the reasons I love covering this space right now. I feel like we saw this big breakthrough in large language models. We’ve seen how that really disrupted society and technology and all of the things we see playing out with AI right now.

And there’s a lot of hype about how this is coming right now. People should not believe the hype about a humanoid robot coming to their house right now. It is not coming right now. I don’t think it’s coming in the next five years. What we’re watching play out is this need for more data and this need for better– we haven’t talked about it here, but what we need is better and safer robots too. I don’t want to allow a 100 pound robot in my house that can’t stand upright, that has a risk of falling down.

So there are a lot of things that have to happen in the next number of years. I think it will happen. I can’t give you a time frame, but I can tell you it’s not happening this year or next year.

IRA FLATOW: All right. That’s a good place to end it, Joanna. Good luck in your research.

JOANNA STERN: Yeah, I’m going to keep on. I’m keeping on this topic. I’m going to be with the robots for many, many years. And they’ll remember when they’re finally working. They’ll say, this was the woman who told the world.

IRA FLATOW: So when the singularity comes, they’ll keep you in mind.

JOANNA STERN: Yeah, that’s right.

IRA FLATOW: Joanna Stern, tech journalist who writes newsletters and creates videos for New Things Media. After the break, we’re adding up the drama happening in the AI math world. Stay with us.

[MUSIC PLAYING]

Now we’re turning to news in the math world. In the last month, AI models have disproved a handful of conjectures. OpenAI announced that it had solved 10 long standing problems with an advanced, unreleased model. Some mathematicians said the results were impressive, while others said they were overblown and accused the announcement of sloppy attribution. Play nice in the sandbox, folks.

So what do mathematicians make of all of this? We thought it was a good time to check in with our mathematical referee, Dr. Emily Riehl, professor of mathematics at Johns Hopkins University, who’s joined us before to walk us through the AI math world. And she’s back with us from Baltimore, Maryland. Welcome back, Emily.

EMILY RIEHL: Thanks for having me.

IRA FLATOW: Nice to have you. All right, let’s get into this. When we last had you on in March, you talked about AI also. Are things progressing slower or faster than you thought?

EMILY RIEHL: I think most mathematicians would say this summer has been very surprising at the pace of improvement in the models. I think the first one that really caught my eye was a disproof of something called the unit distance conjecture, which was a 1946 conjecture of Paul Erdos, celebrated Hungarian combinatorialist.

I would say that there were relatively few areas where AI had made important contributions, and that is changing. AI is still not contributing to all research areas in mathematics, but in certain areas, it is certainly making progress. It seems to do best in areas where the problems are easy to state, if not easy to solve, and maybe as less skilled in the areas where the problems are harder to understand the statements of.

IRA FLATOW: Like I said, OpenAI announced their new unreleased model also solved 10 long standing problems. What’s been the reaction from mathematicians and what do you think of them?

EMILY RIEHL: So the announcement came in the form of a 250 page PDF entirely written by AI, so not with any human mathematical input that we can tell, together with some further verification of each of the 10 solutions in the form of something called a lean proof. Lean here refers to a computer proof assistant that can automate some of the refereeing process.

Maybe the first thing to say is that when a new mathematical breakthrough happens, it usually takes the community some time to respond. And if the new breakthrough involves a delicate mathematical argument that might take dozens or even hundreds of pages, then it can take sometimes even a year, or even more for the community to really verify the correctness of the solution and also understand the impact down the field.

IRA FLATOW: I’m trying to think of what a mathematical prompt to AI would look like. I mean, it must be huge.

EMILY RIEHL: You might think that, but one of the surprises is that often these prompts are really short and relatively naive. A recent example is the Dinitz-Garg-Goemans conjecture, which is a question in graph theory. And here the prompting was four extremely naive prompts telling the model which problem to work on, and to keep working until it found a counterexample.

So what do we make of this? I mean, one phenomenon is that there are a lot of folks out there who love math but are not getting paid to do math full time, who are kind of rediscovering their love of math by pushing the frontier of research knowledge in this way. You could imagine anybody who was aware of that problem could write that prompt and then generate a solution. So in one sense, this is expanding the number of people who are contributing to the mathematics research enterprise.

And we’ve seen this with the Erdos problems too. The Erdos problems were sort of famous for attracting non-specialists who could contribute in a productive way. And those non-specialists are superpowered now with the latest AI models.

IRA FLATOW: Did AI have a particular strength that allowed it to do this?

EMILY RIEHL: There’s been some discussion of that. So there’s some mathematical problems that essentially involve finding a needle in a haystack. And it does seem that AI, for whatever reason, is good at finding unusual mathematical objects with unusual properties. It’s not as good as explaining to us how it found them. The sort of interpretation work for some of these counterexamples has been done by humans.

IRA FLATOW: Yeah, my math teacher always said, show your work. [LAUGHS] Speaking of the latest AI models, are you getting a clearer picture on what the relationship between mathematics and AI models will look like in the future?

EMILY RIEHL: I think there is a lot that is currently in flux. One open question is how much AI can actually help us understand mathematics, as opposed to just tell us what is true and false. Right now, it still seems like there’s a lot of work for humans to do to map out the mathematical universe.

So you can think of the mathematical universe as some sort of vast building of unbounded expanse that mathematicians have spent millennia both exploring and renovating. This building is so large, there’s no architect who has seen the floor plan for everything. There’s no contractor who’s directing everybody to do the work. Instead, there are specialists in different areas who are interested in different problems, who are going off down some dark corridor and trying to open the doors that are stuck to get inside.

So when there’s a door that’s stuck, a problem that we don’t immediately how to solve, it’s not clear immediately how to get through. Maybe it just needs a shove in the right corner. Maybe it needs somebody to find a key. Maybe it needs somebody to create a new key. And what the AI models are doing right now is helping us get in some of these doors that mathematicians were having trouble breaking through before.

But what mathematicians still have to do once they’ve opened up a new door is see if they can turn on any lights in there, maybe renovate a bit to make sure the ceiling doesn’t collapse, see whether this door leads to just a closet, or maybe a whole entire new wing of the building where there will be new mathematical discoveries around the road. And right now, at least, AI is not contributing to those sort of larger enterprises of expanding the mathematical universe.

IRA FLATOW: Well, Emily, it’s always fun having you coming on the show. Thank you for taking time to be with us today.

EMILY RIEHL: Thanks very much.

IRA FLATOW: Dr. Emily Riehl, professor of mathematics at Johns Hopkins University in Baltimore, Maryland. This episode was produced by De Peterschmidt. I’m Ira Flatow. Thanks for listening.

Copyright © 2026 Science Friday Initiative. All rights reserved. Science Friday transcripts are produced on a tight deadline by 3Play Media. Fidelity to the original aired/published audio or video file might vary, and text might be updated or amended in the future. For the authoritative record of Science Friday’s programming, please visit the original aired/published recording. For terms of use and more information, visit our policies pages at http://www.sciencefriday.com/about/policies/

Meet the Producers and Host

About Dee Peterschmidt

Dee Peterschmidt is Science Friday’s audio production manager, hosted the podcast Universe of Art, and composes music for Science Friday’s podcasts. Their D&D character is a clumsy bard named Chip Chap Chopman.

About Ira Flatow

Ira Flatow is the founder and host of Science FridayHis green thumb has revived many an office plant at death’s door.

Explore More