{"id":8217,"date":"2026-08-10T11:00:00","date_gmt":"2026-08-10T11:00:00","guid":{"rendered":"https:\/\/kk.org\/thetechnium\/?p=8217"},"modified":"2026-08-04T17:46:52","modified_gmt":"2026-08-04T17:46:52","slug":"worldbuilding-with-spatial-intelligence","status":"publish","type":"post","link":"https:\/\/kk.org\/thetechnium\/worldbuilding-with-spatial-intelligence\/","title":{"rendered":"Worldbuilding with Spatial Intelligence"},"content":{"rendered":"\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/kk.org\/thetechnium\/files\/2026\/08\/spatial-intelligence.png\"><img loading=\"lazy\" width=\"650\" height=\"364\" src=\"https:\/\/kk.org\/thetechnium\/files\/2026\/08\/spatial-intelligence.png\" alt=\"\" class=\"wp-image-8218\" srcset=\"https:\/\/kk.org\/thetechnium\/files\/2026\/08\/spatial-intelligence.png 650w, https:\/\/kk.org\/thetechnium\/files\/2026\/08\/spatial-intelligence-300x168.png 300w, https:\/\/kk.org\/thetechnium\/files\/2026\/08\/spatial-intelligence-500x280.png 500w\" sizes=\"(max-width: 650px) 100vw, 650px\" \/><\/a><\/figure>\n\n\n\n<p>The next stage in the development of AIs is to give them spatial intelligence.<\/p>\n\n\n\n<p>Our current, smartest AIs are masters of words. They have been trained on zillions of words. Their education consists of the knowledge we have written down into new books and journals. They are smarty pants, the classroom genius that has read everything. But not only have the AIs read most books, they actually remember everything they have read. The LLMs today have a PhD level of knowledge in literally every subject, which makes them powerfully book smart.<\/p>\n\n\n\n<p>But they often lack common sense, and when they are given a body, as in a robot, they flail, flounder, and stall because they have no embodied intelligence. They don\u2019t know about reality. Operating in the real world takes a different kind of intelligence than book smartness. The smartness of a body as it moves in the world requires an intuition about gravity, and lightning fast visual perception, and an awareness of three dimensions \u2013 up, down, front and back \u2013 and many other basic responses that we humans have learned over millions of years of evolution and is baked deep into our reflexes. That kind of embodied spatial intelligence has yet to be trained into our AIs.<\/p>\n\n\n\n<p>Many labs are trying to give AIs this missing spatial intelligence. Every major robot company is working on some version of this research. The prize for succeeding in equipping a robot with an embodied intelligence is monumental; we would finally have robots in our homes, offices, factories, and everywhere. They could get around as well as we can, fold a t-shirt, cook a burger, bathe an invalid. The AIs would do to physical tasks what they have done for intellectual tasks.<\/p>\n\n\n\n<p>But in addition to unleashing the robot world, spatial intelligence would also unleash something else: worldbuilding.<\/p>\n\n\n\n<p><strong>World Models<\/strong><\/p>\n\n\n\n<p>Spatial intelligence would give us world models: AIs that have real world knowledge, not just book knowledge. Instead of being trained on words and descriptions of reality, as they are now, they would be trained on reality directly. They would witness the bounce of a ball instead of a description of a ball bouncing.<\/p>\n\n\n\n<p>The primary bottleneck restraining the arrival of this world model is the lack of sufficient quantity of quality data. Large Language Models (LLMs) worked because the internet had already digitized language: libraries of books, all journals and newspapers, and years of public conversations and personal blogs. Petabytes of digitized text existed and were vacuumed up as training material for this model based on language. There are no equivalent sources of digitized petabytes of data derived from reality. If you are trying to model the world you need tons and tons of data about water moving in all its ways from waves, to splashes, to sprays, to streams and drips. You need data about clouds and wood as building material, and the way clay squishes, and traffic moves, and fabric falls, and balls bounce. You need real data from every corner of life, just as we have text about every corner of life.<\/p>\n\n\n\n<p>The nearest deposit we have of this kind of reality data is the video on YouTube. YouTube has never disclosed how many hours of video they have but it is widely estimated to be in the billions. Given the approximately one million hours of video uploaded every day, the range of human activities YouTube captures is fairly large. There are endless hours of sports play, cooking in kitchens, people working at physical jobs, moments of everyday life, including millions of hours of accidents and improbable events, which are even more valuable when training a model. In addition, there are billions of hours of CCTV security camera footage, and the video recordings from car cams on the roads. These are also being used to train world models.<\/p>\n\n\n\n<p>Many startups are racing to create foundational world models, but I think the first ones to succeed will likely be the platforms in control of this data bank. In the US, Google (owner of YouTube), and in China, ByteDance (owner of Douyin\/TikTok), or those working in partnership with them. Importantly, the best source for a robotic spatial intelligence data will be the continuous experience of robots themselves. As robots work, they rapidly accumulate very good data about the real world that they are scanning. The more robots that have been turned on, in more locations and occupations, the more and better data they collect. Even limited, lame, or poor robots can collect good data, which gives great incentive to get robots out in the world. It will probably be economically shrewd to lose money on early robot models in anticipation that the data they gather will be worth more later in improving newer versions. This incentive might be so strong that robots are sold below cost to you as long as you keep using them.<\/p>\n\n\n\n<p><strong>AR and XR<\/strong><\/p>\n\n\n\n<p>The second significant source of world modeling data will come from smart glasses. Smart glasses have transparent screens you look through as well as cameras that look out. Wearing them you basically see the world that a robot sees. The cameras in the glasses scan the world ahead of you, helping the chips inside to render a digital version of the scene ahead which is laid over the real scene, so that you see a merged version of both. This enables software to render smart annotations to the real scene, whether they are navigation aids (follow the blue arrows on the ground), or fictional fantasies (follow the blue fox running in front of you). The result of melding AI generated scenes with real scenes is known as Augmented Reality (AR), or Mixed Reality (XR).<\/p>\n\n\n\n<p>The crucial step for AI is that the cameras in the glasses in the millions are constantly scanning and re-scanning the world, feeding huge amounts of data about the real world into AIs to be digested and processed. Huge amounts of 3D AI is needed to perceive and \u201cunderstand\u201d the layout of the world, to recognize where something is, or to recognize <em>what<\/em> it is. All the deep situational awareness we expect from a pair of smart glasses requires world modeling. We can\u2019t have AR, or XR without cheap, ubiquitous AI. Conversely, there probably is no 3D AI without cheap, ubiquitous smart glasses to train the models on. We might also expect some AI companies to sell smart glasses at a loss because the data they are generating might be more valuable than the cost of the hardware.<\/p>\n\n\n\n<p><strong>Three Stages of Digitization<\/strong><\/p>\n\n\n\n<p>I like to divide the digital world into three stages. In the first stage, we digitized information and made it machine readable. Machines could read, recall, search, and share all the information of the world. That happened during the dotcom era, and the owners of the machines became the dominant cultural gatekeeper. In the second stage, we digitized the relationships between humans. Machines could see, recall, search and process who was friends with whom, who dated whom, who worked for whom, who liked, who disliked, who swiped, who watched, who voted up \u2013 the entire realm of social relations was now machine readable. And the gatekeeper of those machines of the social networks dominated. We are now about to enter the third stage, where the entire physical world is machine readable. Once millions of people start wearing smart glasses, scanning, perceiving, ingesting, the world 24 hours non-stop, every bit, every action, all physical phenomena are digitized and available to be read by a machine (AI). That would enable us to process, search, manipulate the real world in the same way we did with information.<\/p>\n\n\n\n<p>At the most trivial level we could search the entire world for example; find me a park bench with an unobstructed view to the west, where the sunset light in mid-December glints off the windows of a tall building in the shape of a teardrop, and it would search the known world for that configuration. But it could also search the world for visible evidence (vibrations, rust stains, cracks) that a bridge is in danger of failing. This real-time constant scan of the world, combined with an AI world model, generates a digital twin of the world. We can probe it, search it, for patterns, but we can also run what-if simulations and scenarios in it. This is the Mirrorworld.<\/p>\n\n\n\n<p>This full-strength digital twin applies not just to the whole world but to all its parts. If enough people scan a building with their glasses as they work in it, and the workers maintaining it physically scan its behavior, these all together create a digital twin of that building. That 3D digital twin then becomes a tool for managing the building. The digital twin of the building \u2013 its mirrorworld version rendered by spatial AI \u2013 can generate predictions of what the building might do next, scenarios of what could be done with it as is, and reminders of what needs to be done next to keep it going. Plus the digital twin serves up an augmented reality for any visitor to the building.<\/p>\n\n\n\n<p>The world AI model in smart glasses does three things at once. They display virtual worlds and virtual annotations. You put them on, and virtual smart things are added to what you see. So the generative powers of AI will create these visuals in the same way they generate fictional video clips. The 3D models can generate arrows, advertisements, virtual characters, avatars, text annotations, anything. At the same time, the second task the same AI world model performs in the glasses is to scan the real world in order to cast the virtual layer exactly. The generated images in AR and XR will match the lighting of the scene you are in, and the virtual annotations will geometrically fit into the scene perfectly. If your friend is going to appear as a 3D avatar sitting in the chair next to you, the AI needs to construct this synthetic melding of both worlds with a precise perception of the room you are in. So the cameras in the glasses are constantly scanning the world in order to make the display of the virtual parts believable. Thirdly, the cameras scan the world in order to gain more, and more up-to-date, data to improve the model\u2019s intelligence. They scan not just to render a synthesized view, but to keep getting smarter. In this way, augmented reality (the mirrorworld) will train the AI world models.<\/p>\n\n\n\n<p><strong>The Internet of Things<\/strong><\/p>\n\n\n\n<p>The Internet of Things was a vision promised by digitization. The idea is that eventually every object, every artifact, every thing, is added to the internet. Each item in your home, your toaster and your washing machine, even your shoes, are connected to the internet, and become smart. All the objects in an office and factory are added. Every item on the shelf of a store is added, with their prices dynamically changing as they age, or are in demand. To achieve this internet of things, we would have to add a tiny chip inside each artifact produced, and maybe add power. Everything would get its own IP address. This did not happen, as plausible as it sounds.<\/p>\n\n\n\n<p>However, spatial AI can create a version of the Internet of Things using the data from the mirrorworld of AR or XR. Smart glass scans everything in their view, so in the goodness of time, the entire world will be scanned and digitized. The outside of buildings, the inside of buildings; public spaces, your bedroom; monuments, and the things in your refrigerator. The nature of spatial intelligence is that it perceives the whole view; it can semantically structure what it sees, distinguishing a person here, a car there, a hat on a person\u2019s head, a window in a building, the sill of that window, one pane of glass in that window, a sticker on that pane of glass. The AI can extract all of these patterns from the mess of reality. In other words, the AI is capable of identifying every object and part of every object in the world. It identifies singular things not by giving it a unique number (like an IP address) but in relation to other things. So that window sill is the sill in the window that is 5 up and 4 windows across from the main door of the building that is next to the post office in the city along the harbor at the mouth of the river, etc. That window sill, or that door, or that chair is connected to everything else by some semantic relationship that can be extracted by the AI. To the extent that we can reach the AI via the internet, which is always on, everything the AI sees is therefore on the internet. When systems are scanned and included in the mirrorworld, they also appear in a virtual internet of things.<\/p>\n\n\n\n<p>The virtual Internet of Things \u2013 a system that knows about all objects and places \u2013 is a powerful tool. You could use it to simulate a factory, a city, a company, a household, or a store.<\/p>\n\n\n\n<p>As chips become free and batteries ever better, many, many physical objects really will get their own IP address as well. Those chips will transmit internal states of things that scanning won\u2019t reveal. But even these animated objects will be fitted into the model of the world that is created by constant scanning. A fully developed world model will be able to semantically parse the world, and just as an AI knows every page of every book, the spatial AI will know every room of every building, and every object in every room.<\/p>\n\n\n\n<p>Similarly, this same spatial intelligence will be able to semantically parse the entire visual world we have recorded in video and movies. It will be able to superhumanly comprehend, remember, grok, search, find, and reconstruct the content in every frame in every video ever made. Find me all the moments when a white rabbit is pulled out of a magician\u2019s hat. Where do tweezers appear as a close up in movie history? How has the shape of bathrooms changed over time?<\/p>\n\n\n\n<p>This semantic knowledge of the world, both real and fictional, will unleash thousands of new services and products we have not imagined yet.<\/p>\n\n\n\n<p><strong>Worldbuilding<\/strong><\/p>\n\n\n\n<p>For instance, simple worldbuilding will become one of the new superpowers we get from spatial intelligence. Once the world\u2019s detail is machine-readable the way its language already is, a director can generate a coherent world instead of filming one. Solo individuals will be able to create a feature-length film in their bedrooms. In the same way as a young talented J.K. Rowling could singlehandedly create a deep satisfying wizarding world in immense detail, young directors will create movies and games with deep satisfying details and drama with little additional help. Of course most of these will be unwatchable with an <a target=\"_blank\" rel=\"noreferrer noopener\" href=\"https:\/\/kevinkelly.substack.com\/p\/an-audience-of-one\">Audience of One<\/a>, but this is true of text creations as well. Most novels, fantasy stories, or sci-fi, and books in general are not worth reading; most AI co-generated movies and games won\u2019t be worth it either. But occasionally, one will be brilliant. And it will be marvelous and something that could not have been created via the old-fashioned way.<\/p>\n\n\n\n<p>Worldbuilding is a major part of what the best fiction does. The author creates a world and brings you into it. In the past, to do it well was such a demanding task that few individuals could excel at it; it was often a group collaboration. Science fiction movies and games harnessed the best worldbuilding at great cost. But as it becomes easier and automated, worldbuilding should become a common endeavor. A what-if question could be answered by building a world. What if our product was in every car? What if everyone were given some stock at birth? What if we put cameras everywhere that anyone could watch? And then you build out an entire world to see what happened.<\/p>\n\n\n\n<p>Worldbuilding is a type of simulation; but instead of trying to ensure the dynamics adhere to reality, you can feed it alternative rules. Try stuff. This fiction can be entertaining, but also useful. Worldbuilding assisted by AI could easily become a standard way of designing and managing complex projects. You build out an entire world based on your idea, and then immerse yourself in it to evaluate it.<\/p>\n\n\n\n<p><strong>The Trust Problem<\/strong><\/p>\n\n\n\n<p>The greatest friction slowing down the arrival of this mirrorworld, worldbuilding, and spatial AI is the tension that constant scanning and privacy surveillance brings to this tech. Already, many people are upset by the cameras embedded in current smart glasses. Terms like glasshole are applied to those who leave the cameras on in public. The degree of trust required to allow a company to record not only everything we see, but to let others record everything we do in public is beyond the limit for most people today. It seems unlikely we\u2019ll change our attitudes about being constantly scanned any time soon.<\/p>\n\n\n\n<p>I think we will change our attitude, gradually. The main driver will be the benefits we get. I bet we eventually will become oblivious to being recorded in public. Residential and commercial buildings have cameras around the property, cities film public spaces, and cars \u2013 especially self-driving cars \u2013 film everything around them. Having cameras on people\u2019s glasses will not feel so out of place. Additionally, the kind of scanning done for AI might come to be seen as very distinct from a traditional recording. In fact, the streams of images that are used for training will probably not be saved, in the same way that the text for training AIs is not saved once ingested. The smart glasses may be on, capturing images, but not saving the stream for later reviewing. Furthermore it may be recognized that having an AI watch everything is different from humans watching. Billions of people have become comfortable in having AIs read all their email on gmail; it is not the same as humans reading all your email, and there are benefits to Google \u201creading\u201d it all. There may be ways to anonymize our own captured behavior, so that we can benefit from its digitization, but not be concerned that our privacy is compromised. Technically this is possible; but it requires trust in corporations to execute this process reliably.<\/p>\n\n\n\n<p>AIs in general have trust issues, so the advent of spatial intelligence will depend on how we resolve our trust in large corporations. AIs will remain black boxes, hard to understand and hard to predict. Spatial AI, smart glasses, and robots will share some of this uncertainty, as well as the challenges of maintaining a sense of privacy.<\/p>\n\n\n\n<p>However, the benefits of spatial AI and world models will be huge, and that helps us overcome our fears. From these mysterious models will come real working robots in the millions, engines of wow generating movies and games by solo individuals, a new social media of convincing avatars in immersive 3D presence, augmented mixed reality, and a thousand other things that exceed my meager imagination.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The next stage in the development of AIs is to give them spatial intelligence. Our current, smartest AIs are masters of words. They have been trained on zillions of words. Their education consists of the knowledge we have written down &hellip; <a href=\"https:\/\/kk.org\/thetechnium\/worldbuilding-with-spatial-intelligence\/\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[],"tags":[],"_links":{"self":[{"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/posts\/8217"}],"collection":[{"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/comments?post=8217"}],"version-history":[{"count":1,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/posts\/8217\/revisions"}],"predecessor-version":[{"id":8219,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/posts\/8217\/revisions\/8219"}],"wp:attachment":[{"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/media?parent=8217"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/categories?post=8217"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/kk.org\/thetechnium\/wp-json\/wp\/v2\/tags?post=8217"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}