Everything around us can be understood, at least in some sense, as information encoded in different forms: mechanical waves, matter, protein sequences, text, electromagnetic radiation...
Our senses have evolved over millions of years to decode part of that information and transform it into electrical signals that our brain can process, understand, and use to interact with the environment.
But obviously, there is a huge amount of information around us that our senses cannot directly decode.
We cannot see ultraviolet light. We cannot hear many frequencies that other animals can hear. We do not directly perceive the Earth's magnetic field, while some animals seem to use it for navigation.
Probably, we never needed these signals enough for survival, so evolution did not give us the biological machinery to perceive them.
This made me think about artificial intelligence.
When we process text using large language models, what we are doing is, in a very simplified way, surprisingly similar to what our senses do.
A model cannot directly understand letters or words. First, we transform them into tokens, and then those tokens into mathematical representations that the model can process.
Very roughly:
The original word, at some point, disappears.
What remains is a mathematical representation containing information that the model can manipulate.
In some sense, our senses do something similar.
The ear receives mechanical pressure waves and transforms them, through the cochlea and the auditory system, into neural activity that the brain can process.
The important thing is not the original format of the information.
The important thing is finding a representation in which the system can work with it.
Thinking about this led me to another idea.
The world seems to operate with an enormous number of degrees of freedom, but maybe most of those degrees of freedom do not matter most of the time.
Reality needs, however, what I like to think of as escape doors.
Let me explain.
Imagine that we build a model of water around room temperature and normal atmospheric pressure.
Most of the time, under those conditions, water behaves as a liquid.
If our model only needs to describe this regime, perhaps we can ignore many of the other possible states of the system.
The model could work extremely well.
But then imagine throwing a piece of red-hot metal into the water.
Suddenly, part of the water may evaporate.
Or reduce the temperature enough, and it freezes.
The system had other possibilities available. Its state changed because the conditions pushed it into another region of what was possible.
Most of the time, perhaps we only need to understand a small region of the complete space of possibilities.
The rest contains those escape doors: unusual states, transitions and rare events that normally do not matter, until suddenly they do.
A more complex example appears in proteins.
When we study the dynamics of a protein and how it behaves when interacting with a ligand, such as a drug, the complete number of possible microscopic configurations is enormous.
If a molecular system contains (N) atoms, its Cartesian coordinates alone require variables.
A system containing 100,000 atoms already has around 300,000 positional degrees of freedom.
And yet, a protein does not explore every possible combination of those coordinates equally.
Chemical bonds restrict movement.
Angles and torsions restrict movement.
Steric interactions restrict movement.
Energy barriers restrict movement.
As a consequence, a huge high-dimensional configuration space may contain a much smaller region where most physically relevant configurations actually live.
And this is where things become interesting.
Maybe a protein contains thousands of atoms, but an important biological process is dominated by only a few collective motions.
A domain opens.
A helix rotates.
A binding pocket expands.
A loop moves.
A ligand changes position.
The complete microscopic description is enormous.
But the behaviour we care about may be much simpler.
And this brings me to a question I keep thinking about:
Can we explain 80% of the world using 20% of the information?
I do not mean these numbers literally.
I have no idea whether there is anything close to an 80/20 rule for reality.
It is simply a way of asking a deeper question:
Does most observable behaviour live in a much smaller space than the complete space of possibilities?
If the answer is yes, then perhaps we can explain large parts of the world using much less information than its complete microscopic description would suggest.
Maybe complexity lives in a high-dimensional space, while behaviour lives in a much smaller one.
This idea becomes especially interesting when thinking about molecular dynamics.
Molecular dynamics is a computational method used in chemistry, physics and biology to study how molecular systems evolve over time.
Very roughly, we know the positions of the atoms, calculate the forces acting on them, and then use Newton's equations to move the system forward in time:
We do this again and again using extremely small timesteps.
And that is part of the problem.
Biological systems may contain hundreds of thousands or millions of atoms, while the timestep required to simulate them is incredibly small compared with many of the biological processes we actually want to observe.
So I started wondering:
What if we did not need to evolve the system in its complete coordinate space?
What if we could reduce the dimensionality of the system and express most of its behaviour using only a small fraction of the original variables?
Imagine taking the complete molecular configuration and compressing it into a much smaller representation:
where (x_t) represents the complete molecular system and (z_t) represents a much smaller latent description of it.
Instead of simulating every degree of freedom directly, perhaps we could model how this smaller representation changes over time.
Then, when necessary, we could reconstruct the molecular configuration again:
In other words, compress the system, evolve it in the compressed space, and then return to the original representation.
If that latent representation really captures the important degrees of freedom, perhaps some processes could be simulated much faster.
But there is an obvious problem.
The rare events.
Suppose a protein spends 99.9% of its time exploring a small group of common conformations.
A reduced representation might describe those states extremely well.
But perhaps the remaining 0.1% contains a rare conformation that opens a hidden binding pocket.
Or enables a ligand to enter.
Or triggers a conformational transition.
Or changes the biological activity of the protein.
Then the apparently insignificant part of the state space becomes the most important part.
This is what I mean by the world needing escape doors.
A compressed model can work extremely well inside the region it understands, but reality can move into states that the reduced representation considered irrelevant.
And perhaps this is one of the central problems when trying to compress complex systems.
The challenge is not simply finding the smallest possible representation of reality.
It is finding the smallest representation that still preserves what matters.
This idea goes much further than molecular dynamics.
Maybe intelligence itself, biological or artificial, is partly the ability to discover these compressed representations.
Our senses already discard enormous amounts of information.
Our brains do not model every photon, every air molecule or every microscopic interaction around us.
They extract structure.
Objects.
Movement.
Faces.
Danger.
Patterns.
Meaning.
Artificial intelligence seems to do something related.
It transforms enormous amounts of raw information into internal representations where useful structure becomes easier to manipulate.
And science does something similar.
We replace billions of microscopic variables with temperature, pressure or concentration.
We replace complicated molecular trajectories with collective variables.
We replace enormous descriptions of systems with models that capture only the behaviour we care about.
Maybe understanding the world has always been, at least partly, a problem of dimensionality reduction.
Not removing information randomly, but discovering which degrees of freedom matter and which ones can be ignored.
And maybe one of the most interesting questions is not why the universe is so complex.
Maybe it is:
Why is so much of that complexity compressible?
Because in principle we could imagine a universe where understanding one billion variables required tracking all one billion variables independently.
But that does not seem to be the universe we live in.
Nature has patterns.
Symmetries.
Correlations.
Structures.
Scales.
Regularities.
And because of those regularities, a small number of variables can often tell us something about systems that contain an enormous amount of microscopic information.
Perhaps that is one of the reasons science is possible at all.
And if that is true, then building better models of the world may not only be about collecting more data or increasing computational power.
It may also be about discovering better representations of the information we already have. Perhaps intelligence is, in part, the art of finding the right compression of reality.
What do you think?

