InventDB
Research

InventDB Brahma: an intelligence that learns from its own experience

InventDB Brahma is our long-term research toward a human-like intelligence. It learns how its world works from its own experience, starting from a small, counted set of building blocks, and it writes what it finds as rules a person can read. Below we set out what we mean by intelligence, why Brahma does not use a language model, and what it has shown so far in tests against neural networks.

How Brahma learnt in its first seventeen days, in two and a half minutes, with music. Everything in the film comes from Brahma's own recordings: the room where it began, real drivers on a Los Angeles freeway, and a nursery where it first meets warmth, blows and breaking. Every number on screen comes from our research notes.
In this article

We put Brahma in a room of moving balls and let it learn how they move. Then we made the room 50% larger and did not tell it. Brahma kept predicting where each ball would go next, and was right 89 to 94% of the time. A small neural network trained on 14,100 moments of the old room, and so only ever on one room size, fell to between 7 and 41%.

Brahma's predictions stayed right because the rules it wrote for that room are stated in terms of the room's size, so they still applied when the room grew. Nobody gave it those rules. It found them in its own experience, and a person can read them. A network trained on rooms of several sizes would be a harder comparison, and we have not run it yet. The result still shows the two things this research is about: learning how a world works from very little experience, and knowing it well enough to keep predicting correctly when it changes.

Brahma's room with five balls and a dashed ring where it expects each ball next, beside the rules it wrote for where each ball goes
This is Brahma's room in its live view, and each dashed ring is where it expects a ball to be next. The panel on the right lists the rules it wrote for the room, which are right on about 92% of the 5,249 sightings it remembers and learned from, and the weights it had measured by pushing for four of its five balls. The held-out results are in the table further down.

What we mean by intelligence

We take intelligence to be how cheaply a mind turns new experience into skill at things it has never met, rather than how much it already knows. This is François Chollet's definition, skill-acquisition efficiency, and it is close to what psychologists call fluid ability: the capacity that builds knowledge, as opposed to the knowledge itself. Chollet's own test of it is ARC. Brahma's live mind has not yet solved any of the 18 kinds of one-row ARC puzzle it has met, and full ARC grids are kept as held-out tests for later.

On this view a score on its own is a weak signal. Enough built-in knowledge or training data can buy high skill at almost any fixed task. To see intelligence, you have to count what a system started with and what experience it was given, and then watch how little it needs to learn something new. We try to report every result that way. The parts we wrote by hand, the bridge that turns each world into records, the planner and the executor, sit outside that count, and we name them where they matter.

In people, much thinking happens without words

Language seems to be mainly a tool for communicating. The brain's language network stays largely quiet during non-verbal reasoning, and some people with severe aphasia can still do algebra (Fedorenko, Piantadosi and Gibson, Nature 2024). Babies have expectations about objects, number and space long before they have words, and Elizabeth Spelke calls these expectations core knowledge. Babies also find out how the world works by experimenting on it, as Alison Gopnik's work on the child as a scientist describes.

We use this evidence to decide what Brahma has to learn to do. A human-like intelligence first learns how its world works from its own experience, and what it later learns from teaching and text builds on that. It predicts what will happen, is surprised, and learns from its own mistakes. It finds out what it cannot see by acting on the world. It consolidates what it learned in sleep, builds new concepts out of older ones, and knows what it knows and what it does not.

Any learner improves fastest when something outside it checks its answers. Language models show this clearly: asked to correct themselves with no outside signal they rarely improve (Huang et al., 2023), while feedback from tests, tools and verifiers does help them. Rich Sutton argued in 2001, in "Verification, the key to AI", for knowledge that an AI can check for itself because it is stated as a prediction about its own experience. Brahma keeps a rule only after checking it against moments it has not seen.

Why Brahma does not use a language model

Language models are remarkable systems, and InventDB's products use them. For this research they answer a different question from ours. A model trained on trillions of words starts from a prior too large for anyone to audit, so its skill cannot tell you how much it learned and how much it was given.

Our question is how a mind comes to know its world before it has words, and how little experience that takes. Answering it needs a learner whose starting point we can count, so Brahma uses no pretrained model, embedding, knowledge graph or corpus anywhere.

What Brahma is built to become

The goal is a mind that learns its world from its own experience, sees and hears that world through senses it learned to use, finds out what it cannot see by acting, judges what it knows and what it does not, makes its own concepts, and can be spoken to and say what it has found. Human language comes last on that list, because it is mainly how a mind reports what it knows. Brahma is a long way from all of this today: it has no language yet, and it cannot act for a purpose of its own.

How Brahma starts and how it learns

Brahma starts from 41 basic operations and a small language for quantities, and we count and log every addition we make later, with the result that needed it. Beyond those it has only its own experience of its worlds, with no labels, rewards or examples to copy. The equation test further down is the exception: there it is handed readings and asked for the formula behind one of them.

It learns by predicting what comes next and checking each prediction against what happens. It works on the moments it predicted wrongly. For each, it searches for short programs built from its operations that account for those moments, prefers the shortest program that explains the most, and keeps a rule only if it also holds on moments it has not seen. Pieces that keep proving useful are named and reused, and those names are the words it makes. Each rule it keeps comes with the conditions under which it holds.

In the room, at the table, in the workshop and on the highway, a bridge we write for each world hands Brahma and the networks the same records of each thing: where it is, where it was a moment ago, its size and its colour. On video, Brahma sees through a small neural network that it trains from scratch, without labels, on what it watches. Everything Brahma knows about how its world behaves is kept as rules.

Its nearest relatives are program-synthesis learners that grow a library of their own abstractions, such as DreamCoder (Ellis et al., 2021). Brahma differs in posing its own questions from its prediction errors in a world, rather than solving tasks it is given.

How we test it

We test Brahma in worlds where the true answer is known: a room of moving balls, a simulated table, a workshop built on the MuJoCo physics engine, the I-PHYRE puzzles, the CLEVRER videos, a highway in highway-env, and a set of 119 physics equations. A moment is one step of a world's clock, the same step for every learner in that world: one tick of the room, a twentieth of a second in the workshop and a tenth of a second on the highway.

We run two kinds of neural network beside Brahma, both built for the same worlds and given the same records. Each is a plain network of two layers of 64 units, about 5,400 numbers. A neural world model learns to predict what happens next and drives the same planner that Brahma's rules drive. A deep Q-network learns the value of each move from the score. The networks get as much experience as Brahma or more, and we count it. Stronger rivals exist for several of these tests, among them networks that infer hidden properties from recent history, object-based networks, model-based agents such as Dreamer and symbolic regression systems, and we would welcome them on the same tests.

Who chooses the moves. When Brahma acts in these tests, a planner we wrote uses its rules to choose each move. At each moment the planner tries every pair of first two pushes, runs the model four moments ahead, and gives the first push of the best future. The neural world model drives the same planner, and given the true physics the planner scores 96% on the table, which is the ceiling for any model it uses. Brahma's part is its knowledge of the world; it does not plan.

Brahma was developed with these worlds in view, and most of these tests were run again after changes to it, so the results here come from recent builds rather than a first attempt. Every build and every result, including the ones that went against Brahma, is dated in our research notes.

What it has shown so far

Most of Brahma's results are a range across several lives. A life is a separate run grown from an empty memory: there were nine lives in the room, six at the table, in the workshop and on the highway, and three on CLEVRER and the equations. Where a network was trained more than once, its range is given too. On the table and in the workshop, a score is the share of each round the puck spends where it should be, over 20 rounds.

TestBrahmaNeural world modelDeep Q-network
Experience needed to predict a room right four times in five783 moments on averageAbout 9,500 moments (4,800 to 14,100 over three runs)Does not apply: it predicts nothing
The room made 50% larger without telling anyone89 to 94% right7 to 41% right, trained on one room sizeDoes not apply
Simulated table: keep the puck at home90 to 96%, never having seen the table96%, after 20,000 moments of the table96%, after 200,000 moments of the game
A puck four times heavier, without telling anyone91% in five lives, 83% in one83%83%
A bigger table, the goal moved and a heavier puck86 to 88% in five lives, 72% in one81%65%
Park the puck in a corner, a goal none of them practised78 to 92%93%77% at first, 90% after 50,000 moments of practice
Something unseen keeps knocking the puck88 to 95%94%96%
Workshop: three look-alike pucks, bring the heaviestThe heaviest every round, in all six lives30%; it cannot weigh by design, and a blind guess gets 33%Not run
Workshop: keep the puck at home, about 14,000 moments each92 to 93% in four lives, 81% and 5% in two86% (92% with 20,000 moments)82%, practising that game (91% with 200,000)
CLEVRER videos: what happens next, held out90.1% after about 12 videos, the same as the plain rule that things keep going72.6% after 10 videos, 90.5% after 300Does not apply
Icy highway, without telling anyone: crashes in 36 drives3 or 4 in four lives, as on a dry road; 5 and 10 in two5, against 2 on a dry roadNot run

It learns from very little experience

In a room of moving balls, Brahma's lives predicted where each ball would be at the next moment, right four times in five, within 750 moments of their own experience in eight lives of nine and 1,050 in the ninth, 783 on average. 750 was the first point we measured. A small neural network we built for the same room needed between 4,800 and 14,100 moments, about 9,500 on average. Part of the reason is Brahma's language: "where it will be is where it is plus its speed" is a short program in it, so a few hundred moments are enough to single it out.

On CLEVRER, a public set of videos of colliding objects, Brahma found from about a dozen videos that things keep going as they were going. That rule answers 90.1% of the held-out "what happens next?" questions. A neural network that saw the videos through the same vision network as Brahma answered 72.6% after 10 videos, 82.1% after 30 and 89.1% after 100, and needed 300 to reach 90.5%. Both learners look through a vision network that first trained itself, without labels, on frames from all the watched videos, so these counts are what each learner's model needed after that. The plain rule that things keep going also scores 90.1% when we write it by hand, so what Brahma achieved here is finding that rule by itself from about 12 videos. None of the rules it has found about collisions scores higher yet. The test is our own harness on 137 held-out questions with two choices each, which differs from CLEVRER's official scoring.

A drawn robot watching a CLEVRER video of coloured objects, with the held-out question and the scores of Brahma and the neural network
A drawn robot with Brahma as its brain watches a held-out CLEVRER video and is asked whether the yellow and red objects or the blue and red objects collide next. Brahma's rules answer 90.1% of these questions after about 12 videos, and the network answers 90.5% after 300.

It copes when the world changes

We then used a simulated table that Brahma had never seen. Our planner played it once with the rules Brahma found in 14,100 moments of another room and once with a neural world model trained on 20,000 moments of that table, and a deep Q-network chose its own moves after 200,000 moments of playing the game. With either network the puck spent 96% of each round at home, and with Brahma's lives 90 to 96%. The same planner given the true physics also scores 96%, so the networks and Brahma's best lives are at the ceiling.

Then we made the puck four times heavier and told no player. Five of Brahma's six lives scored 91%, level with the planner given the true physics, and the sixth scored 83%. Both networks scored 83%. Brahma's rules say that how far a push moves a thing depends on a number that belongs to the thing. Our executor uses that rule to weigh the new puck in play and keeps the weight from one round to the next, while the neural world model, which has no input for weight, plans as if the puck were the old one. On a bigger table with the goal moved and a heavier puck, five lives scored 86 to 88% and one scored 72%, against 81% for the neural world model and 65% for the Q-network.

A drawn robot pushing a puck on a blue table, with Brahma's rules for the table and the weight it measured for the puck
The drawn robot pushes a puck at the simulated table in this recording from 28 September. The panel on the left shows the rules Brahma learned in another room and uses at this table. Brahma estimates this puck's weight at 3.56, and its true weight is 4.

On a simulated highway in highway-env we made the road icy and told no driver. Four of Brahma's six lives crashed no more than on a dry road, 3 or 4 times in 36 drives, and the other two went from 3 to 5 and from 6 to 10; across all six lives Brahma went from 22 crashes to 28. The neural world model went from 2 crashes to 5, and highway-env's own driver, which knows the true settings, went from 2 to 3. With 36 drives each, differences of one or two crashes are within chance. On a dry road the network is the safer driver, with 2 crashes against Brahma's 3 or 4.

Three lanes of simulated highway, driven by highway-env's own driver, the neural network and Brahma, with crash counts beside each
This is the highway on a dry road, recorded on 28 September. Over 36 drives the median of Brahma's six lives crashed 3 times, the neural network crashed 2 times and highway-env's own driver crashed 2 times. A car given no command crashed 32 times.

It works out what it cannot see

In the MuJoCo workshop, three pucks look alike but weigh different amounts, and the task is to bring the heaviest. A routine we wrote nudges each puck twice, with the same nudges for Brahma and for the network. Brahma's own rule for weight converts how far each nudge moved a puck into a weight, and with those weights the heaviest puck was picked in every round, in all six of its lives. Over all rounds, two nudges gave weights up to 62% off, which was still enough to tell the pucks apart. The neural world model is given no input that carries weight and no memory of the nudges, so it cannot weigh by design; it chose the heaviest 30% of the time, against 33% for a blind guess. This shows what Brahma's rule for weight adds. A network that remembers the nudges would be the fair rival, and we have not built one yet.

Two simulated tables with three look-alike pucks; Brahma's weights for them beside the true weights, and question marks on the neural network's table
Brahma weighs three look-alike pucks in the MuJoCo workshop in this recording from 27 September. In its own unit, with the lightest puck set to 1, its weights for this round were 2.11, 5.31 and 1 against the true 2, 5 and 1. The network has no way to estimate weight, so its pucks are marked with question marks.

In a separate run of nine lives in its room, we gave Brahma a record of how much each push sharpened its measure of a ball, and a word for how unsure it is of each ball, but we did not tell it which ball to push. Every life's own search found that the useful push is "the ball I am least sure of". From then on it knew a newly arrived ball's weight to within a tenth after about 7 pushes, where pushing at random took 19 to 32. A routine we wrote by hand still does it in 3 to 5. In I-PHYRE, a set of physics puzzles, it found that falling things gain 0.0334 units of speed each moment, in the picture's units; the engine's gravity in the same units is 0.0333.

Given about the same amount of experience in the workshop, around 14,000 moments each, four of Brahma's six lives kept the puck at home 92 to 93% of the time without ever practising that game, and the other two scored 81% and 5%. The neural world model scored 86%, and a Q-network that spent its experience practising that game scored 82%. With more experience, 20,000 moments for the network and 200,000 for the Q-network, they reach 92% and 91%, level with Brahma's best lives.

Three simulated people at three MuJoCo tables: Brahma 93%, the neural network 86% and Q-learning 82%
Brahma, the neural world model and the Q-network play at three identical tables in the MuJoCo workshop, each with about 14,000 moments of experience, in this recording from 27 September. The panel at the top right shows Brahma's rules for this table, written partly in words it made. These workshop games have not yet been run again on the current build of the workshop.

It finds laws in raw numbers

Given readings with their units, Brahma recovered 89 or 90 of the 119 equations in the Feynman physics set in each of three lives, writing each as a short program. We count an equation as recovered when Brahma's formula matches a hundred readings it never saw to within a millionth. For this test we gave it pi, the square root, the sine, the cosine and the exponential, and a search that uses the readings' units as a physicist does; we count all of these as additions. When this work began it recovered 25 of the 119. AI Feynman, a system built for exactly this task, reports all 100 of its main equations and 18 of its 20 harder ones on its own version of the set; the public copy we used holds 119 of those 120.

It makes its own words

Brahma's records give each thing's place now and a moment ago. In its first room Brahma found that the difference between them, the quantity we call speed, kept explaining its mistakes, and it made a word for that quantity. In its world of numbers, given only the step to the next number, it built addition, then multiplication out of its own addition, along with other ideas nobody told it about, such as even and odd numbers and the minimum and maximum of a list. An answer key told us afterwards which human idea each one matches. The language it uses for physical quantities has arithmetic built in, and that is counted as part of what it starts with. On 25 September, 121 of the 207 rules the live mind held used words it had made.

Counts from the live mind: laws found, concepts it named, kinds of problem solved, and problems it set itself and solved, above a list of its concepts that match human ideas
These are the live mind's counts on 4 October: laws found in its worlds, concepts it named, kinds of problem solved, and problems it set itself in sleep and solved. The list below the counts shows concepts it built that an answer key matched to human ideas.

The mind that is running now

One Brahma mind has been running since 19 September. We stop it for each new build and restart it from its saved memory, and the latest build went live on 4 October. Later that day it had run 23.7 million episodes, each a short run of one of its worlds, and it had slept more than 120,000 times; sleep is an offline pass in which it re-checks its rules and sets itself problems. It had made 7,935 concepts, of which 119 were in use, because it discards most of the concepts it makes. None of the comparisons with networks draws on this mind: each comes from lives grown from an empty memory.

Brahma's live mind: the concepts it made clustered at the centre, and the kinds of problem its worlds pose drawn as diamonds around them
This is the live mind on 4 October, running the latest build. The circles in the middle are the concepts it made, and the diamonds around them are the kinds of problem its worlds pose, green where it found the law.

Where it is behind

  • Asked to park the puck in a corner, a goal none of the players had practised beforehand, Brahma's lives scored 78 to 92% against the neural world model's 93%. Once the Q-network practises each new task for 50,000 moments, it reaches 90% or better on every one.
  • When something unseen keeps knocking the puck, the Q-network scores 96% with no practice at all and the neural world model 94%, against 88 to 95% for Brahma's lives.
  • Results vary between Brahma's lives. On the table two of the six lives scored 90 and 93% where the networks scored 96%, and in the workshop one life failed the one-puck games.
  • On the dry highway the neural world model crashes less often than Brahma: 2 times in 36 drives, against 3 or 4 in five of Brahma's lives and 6 in the sixth.
  • On CLEVRER Brahma does no better than the plain rule it found, and on IntPhys 2, a test of photorealistic rendered video, Brahma and the network are both at chance.
  • On the Feynman equations a dedicated system recovers more than Brahma does.
  • Brahma needed far less experience than the networks in the room and on CLEVRER, and about as much on the table and in the workshop. It needs far more computing: it searches millions of short programs for each rule, and a life in its room takes 15 to 25 minutes on a 32-core server running six lives at once, where the neural world model trains in 7 seconds and the Q-network in 53.

What it cannot do yet

Brahma has no language: it cannot be told anything in words, and the words it makes are its own labels, which it cannot yet exchange with a person. It cannot act for a purpose of its own. Asked to keep a puck near a ring and left to choose its own pushes, it succeeds 3% of the time, which is what doing nothing scores, while our planner with the same rules gets 90%. Its vision network works on simple rendered scenes and is at chance on photorealistic video.

What comes next

In the current step we give Brahma a body in a simulated workshop and ask it to predict where a puck, a ball and a block will be up to two seconds ahead, more accurately than a neural network and more accurately than the guess that nothing moves. On a trial build that we have not yet kept, its nearest life of six already does this for a puck and a block at every horizon we test, and for a ball up to a second ahead. Things that bounce off the edge of the table are what stands in the way. Seeing photorealistic video is still open work. The last step is for Brahma to understand what it is told in words and to say what it found.

Brahma is one of three research programmes at InventDB. InventDB Sparkle is our inference engine for open models, now in production in the platform, and InventDB Cluster runs one database across many machines. All three are listed under Research.

If you work on world models, open-ended learning or how minds learn, we would like to compare notes. Write to us at contact@inventdb.com.

Sources

  • Chollet, F. (2019). On the Measure of Intelligence. arXiv 1911.01547.
  • Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology 54(1), 1 to 22.
  • Fedorenko, E., Piantadosi, S. T., Gibson, E. A. F. (2024). Language is primarily a tool for communication rather than thought. Nature 630, 575 to 586.
  • Spelke, E. S., Kinzler, K. D. (2007). Core knowledge. Developmental Science 10(1), 89 to 96. Spelke, E. S. (2022). What Babies Know: Core Knowledge and Composition, Volume 1. Oxford University Press.
  • Gopnik, A., Meltzoff, A. N., Kuhl, P. K. (1999). The Scientist in the Crib. William Morrow.
  • Huang, J. et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024.
  • Sutton, R. S. (2001). Verification, the key to AI.
  • Ellis, K. et al. (2021). DreamCoder: Bootstrapping inductive program synthesis with wake-sleep library learning. PLDI 2021.
  • Mnih, V. et al. (2015). Human-level control through deep reinforcement learning. Nature 518, 529 to 533.
  • Hafner, D. et al. (2023). Mastering Diverse Domains through World Models. arXiv 2301.04104.
  • Udrescu, S.-M., Tegmark, M. (2020). AI Feynman: A physics-inspired method for symbolic regression. Science Advances.
  • Yi, K. et al. (2020). CLEVRER: Collision Events for Video Representation and Reasoning. ICLR 2020.
  • Li, S., Wu, K., Zhang, C., Zhu, Y. (2024). I-PHYRE: Interactive Physical Reasoning. ICLR 2024.
  • Bordes, F. et al. (2025). IntPhys 2. arXiv 2506.09849.
  • Todorov, E., Erez, T., Tassa, Y. (2012). MuJoCo: A physics engine for model-based control. IROS 2012.
  • Leurent, E. (2018). highway-env.