Suppose there were a spectre in the universe called “Laplace’s Demon.”1 It possesses infinite computational power and knows the momentum and position of every atom in the universe. You might think that such a being must possess the highest wisdom in the universe, able to see through.
The reality is completely contrary to our expectations. Mathematics and computer science tell us a brutal truth: infinite compute means that it has no need for what humans define as “intelligence.” In the eyes of Laplace’s Demon, the world contains no “apples,” no “countries,” and no “love”—only a collection of atoms moving tediously according to physical laws. It does not need to summarise patterns, because it can brute-force everything directly.
A striking paper titled From Entropy to Epiplexity recently appeared on arXiv. Several researchers from Carnegie Mellon University and New York University re-examined classical information theory and proposed a deeply counterintuitive conclusion: if you want to learn a genuinely useful system of knowledge from the world and build what we call higher cognition, you must be a computationally bounded observer.
Although the paper uses rigorous mathematical formulas and cryptographic concepts throughout to discuss large language models (LLMs) and machine learning, it actually reveals a profound algorithm for life.
I. Shannon’s blind spot and the curse of television static
In classical information theory, information is measured in bits. Claude Shannon2 tells us that the essence of information is “the elimination of uncertainty.” The less predictable and more chaotic an event is, the more information it contains. This is what we call information entropy. By this standard, we are mercilessly bombarded with enormous volumes of information every day.
Some may say that we need more information to find the “right” path through life. But imagine an old television with no antenna plugged in, its screen filled with dense static. The image is completely random: you can never predict whether the next pixel will be black or white. According to Shannon’s theory, and the later complexity theory of the Soviet mathematician Kolmogorov, that field of static contains an enormous amount of information—nearly the theoretical maximum.
Yet no normal person would stare at television static all day. There is no “meaning” in it whatsoever.
To fill this gap, the paper’s authors introduce two new core concepts: Time-bounded Entropy and Epiplexity.
Time-bounded Entropy refers to unpredictable content that, given limited time and compute, appears entirely random and impossible to make sense of. Most of life’s trivia, emotional venting on social media, and the stock market’s high-frequency daily fluctuations all constitute extremely high Time-bounded Entropy for an ordinary person. No matter how much effort you put into tracking them, you cannot extract any reusable pattern. You are simply spending your life calculating something akin to a cryptographic pseudorandom sequence.
By contrast, Epiplexity can be translated precisely as “structural complexity.”3 It represents the higher-order patterns, long-range dependencies, and underlying logic hidden in data that can be extracted with finite compute.
Confronted with the world’s complexity, ordinary people often become trapped in front of the static. They greedily scroll through short videos and consume fragmented gossip, their brains running at full speed to process extremely high Time-bounded Entropy. They believe they have acquired a vast amount of information, yet their minds remain empty.
Experts adopt an entirely different strategy toward the same world. With extreme restraint, they filter out random noise and devote all their precious compute to extracting Epiplexity. In his classic Fooled by Randomness, Nassim Nicholas Taleb used probability theory to reveal a similar principle. Taleb has said that he barely looks at social media and does not even read newspapers.
If you observe your portfolio every day, you will see countless tiny fluctuations. In Taleb’s view, 99% of those movements are noise and only 1% is signal. Investors who constantly check quotations are wasting compute trying to capture the noise of a random walk. Taleb argues that the quality of information falls sharply as the frequency of observation rises. This high-frequency, disorderly information is exactly what the paper calls Time-bounded Entropy.
Only when you lengthen the observation window—from once a day to once a decade—does the fine-grained noise automatically vanish. What remains, the structural trends that truly change one’s destiny, is Epiplexity. Experts understand this deeply. They actively give up their fixation on instant feedback and keep their distance from the outside world, forcibly reducing the burden that high-entropy information places on the brain. Such restraint is the necessary cost of extracting underlying structure.
The paper contains a brilliant experimental finding. Researchers compared language-text data with high-definition image data and found that text contains extremely high Epiplexity. Although a high-definition image—one from the CIFAR-5M dataset, for example—occupies a huge number of bytes in a computer, more than 99% of its information is irregular pixel-level random noise. Text has a completely different structure: the arrangement of every word embodies strong logical relationships and abstract concepts. A passage from a classic may occupy far fewer bytes than a low-resolution landscape photograph, yet contain hundreds or thousands of times as much Epiplexity.
This also perfectly explains why today’s AI revolution has been led by pretrained large language models. Models pretrained on massive bodies of text can develop astonishing cross-domain generalisation, even solving complex logical reasoning problems zero-shot. Models that only look at images find this extraordinarily difficult.
Life works the same way. Dense, deep reading and systematic, rigorous thought are essentially efficient ways to extract content rich in Epiplexity. Chasing nothing but brief, fast sensory stimulation amounts to gulping down enormous quantities of useless random entropy.
II. The curse of computation and the miracle of “emergence”
The paper’s most compelling insight is its complete exposition of the central value of being “compute-bounded.”
We often complain that our memory is not good enough and that our brains process information too slowly. We imagine that if our brains could remember everything and calculate instantly like supercomputers, nothing could stand in our way. With rigorous mathematical proof, the paper proves that only computational constraints can force a system to produce the phenomenon of “emergence.”
True intelligence is born precisely from limitation.
Consider Conway’s Game of Life. Across its vast two-dimensional grid, black and white cells live or die according to only a few extremely simple rules about their neighbours. Laplace’s Demon would have no difficulty predicting its future. It would merely follow the basic rules and calculate, step by step, the microscopic state of every cell. In the eyes of a supercomputing system, the whole world contains nothing but tedious zeroes and ones.
The human brain faces an absolute computational bottleneck. We cannot calculate the complex evolution of tens of thousands of cells in a second. To predict where the Game of Life is heading, people were forced to invent an entirely new vocabulary of higher abstractions. We observed combinations of cells with particular shapes and named them “gliders,” “static blocks,” and “oscillators.” We went on to summarise the fixed speed at which gliders move and the macroscopic laws governing their collisions.
In this remarkable process, the machine with infinite compute sees only local rules, while computationally bounded humans perceive “structure.” These combinations of higher-order concepts, distilled to overcome our lack of compute, are precisely the Epiplexity we need to acquire.
The same phenomenon can be mapped onto experiments with Elementary Cellular Automata. Researchers asked large language models to learn the evolutionary rules of different automata. A rule such as Rule 15 is too simple: its images consist entirely of periodic, repeating patterns with no learning value. Rule 30 produces images so chaotic and full of pseudorandom information that a model can exhaust its compute without achieving anything. Only a complex rule such as Rule 54, poised at the edge of chaos, both contains change and conceals a macroscopic geometric logic. In learning it, the model extracts high Epiplexity.
This behaviour by large language models vividly demonstrates what genuine learning is.
If you try to memorise every detail of your work, every word every client has spoken, and every minute change to every line of code, you are merely downgrading yourself into an inefficient mechanical hard drive. True experts calmly accept the physical limits of their brain capacity and processing speed. They actively abandon mechanical enumeration of microscopic variables and instead wrestle with the macroscopic patterns and underlying laws behind things.
You cannot remember every leaf, so you invent the concept of a “tree.” You cannot calculate every price movement, so you summarise “cycles” and “mean reversion.” Physical limitation does not imprison human intelligence; it is the central force driving humanity toward higher cognition.
III. The pain of factorisation is a shortcut to reshaping the brain
Classical information theory has another long-revered assumption: the total amount of information is unrelated to the order in which an observation is decomposed. Observe element A and then element B, or B and then A, and the total information you ultimately obtain should be identical. Mathematically, this is called the symmetry of information.
Feedback from the real world completely shatters this beautiful illusion. Through a demanding chess experiment, the paper’s authors broke through this theoretical filter and revealed a third secret of cognitive advancement.
The researchers rigorously trained AI models on professional records from tens of thousands of chess games in the Lichess dataset. They deliberately arranged the data in two radically different ways:
- Forward logic. The model first saw a long sequence of moves—White plays E4, Black plays E5, for instance—and only at the end saw the final board position.
- Reverse logic. In a deeply counterintuitive arrangement, the model first saw the final frozen board position and then had to predict and infer the preceding complex sequence of moves.
The results were striking. For the model, the difficulty of prediction soared under the second, reverse-training method. Classical theory says that the two datasets contain exactly the same absolute amount of information; only the order has changed. In actual training, however, reverse prediction forcibly made the model extract richer Epiplexity.
The difference became fully apparent in a subsequent test on an unfamiliar task. The researchers asked the two models to do something neither had ever seen: score an unfamiliar chess position and assess which side held the advantage, in a so-called centipawn evaluation. The reverse-trained model demonstrated extremely strong out-of-distribution (OOD) generalisation and scored far above the first, forward-trained model.
Why? The answer lies in the asymmetry of computation. Difficult reverse inference forces a model to abandon superficial statistical memorisation entirely. Deep inside its neural network, it must construct a profound internal representation of the overall position on the board.
Moving from cause to effect is often a flat highway. Given the moves one by one, you can flow naturally toward the final position through simple accumulation of rules. It is like reading a detective novel narrated chronologically: you effortlessly reach the ending.
Moving from effect to cause is a rugged mountain road. Seeing the final board, you must enumerate countless possible historical paths in your mind and work backwards through the tactical intention behind every move. Because its compute is limited, the model cannot brute-force the reconstruction. Difficult reverse inference compels it to abandon surface-level statistical memorisation and build, deep within its neural network, a profound internal representation of the overall position, piece values, and high-level tactics. This internal representation is precious Epiplexity.
When learning any new skill in real life, we face the same choice between two radically different paths.
Forward learning is like sitting in a classroom, gliding along a ready-made derivation laid down by others and listening to executives share the secrets of their success. Everything feels perfectly smooth. You feel that you understand it all, yet no deep cognitive circuit has formed in your brain. Such knowledge is fragile: the moment it encounters an unfamiliar cross-domain problem, experience accumulated in the forward direction collapses.
True deliberate practice must contain this reverse, extraordinarily difficult “factorisation.” Take a wildly successful business case, stripped of any background hints, and infer in isolation the life-or-death choices its founder originally faced. Take a leading industrial product and reverse-engineer its core design logic from scratch.
This road is covered in thorns. It rapidly consumes mental and cognitive resources and leaves you profoundly frustrated. Precisely for that reason, it can break your existing neural connections to the greatest possible degree and convert cold information into living Epiplexity inside your brain. Day after day, experts deliberately create this “intense discomfort” in their mental training.
The comfort zone contains nothing but low-grade entropy. Only through immensely demanding reverse decomposition can you refine the gold of wisdom.
IV. Closing the country and an algorithm pulling itself up by its bootstraps
We often hear a confident claim: if a person does not actively engage with new information from the outside world, and does not maintain an intense hunger for the latest developments, they can never make a new cognitive leap. The Data Processing Inequality of classical information theory expresses the same view in cold mathematical language: applying a fully deterministic computational transformation to existing data can never create even the slightest additional amount of information out of nothing.
But AlphaZero, developed by Google’s DeepMind—or, one might say, 0-shot learning—is a powerful rebuttal to this iron law.
At the beginning of training, AlphaZero was given only the most basic rules of chess. It rejected the massive archive of game records left by human masters and received no external knowledge. Inside a closed system, it simply played tirelessly against itself. A few days later, it had evolved extraordinarily deep new strategies that astonished the entire human chess world. It even invented sacrificial openings that human players had not conceived of in hundreds of years.
According to the Data Processing Inequality, with no injection of new external data, the total information in AlphaZero’s system should have remained zero. Where, then, did the enormously complex strategic thought housed in that vast neural network, with its tens of millions of parameters, suddenly come from?
The paper’s researchers offer an incisive theoretical explanation. When we marvel that AlphaZero has learned “new knowledge,” what we mean lies entirely outside information quantity in Shannon’s sense. Its essence remains Epiplexity.
The Data Processing Inequality has a hidden, fatal premise: it assumes an observer with infinite compute. For a god of infinite computation, the rules of chess and the optimal solution to chess are equivalent in information content. Once the rules exist, the optimal solution is already destined to be there.
Observers in the real world have extremely limited compute. The rules may be simple, but deriving high-level strategy from them requires crossing an immensely wide computational gulf. By investing an enormous amount of “computation,” a resource-bounded system can forcibly transform highly deterministic basic axioms, apparently carrying no incremental information, into structured information of enormous practical value.
The act of computation itself continually creates new knowledge.
This also perfectly explains why Synthetic Data can make modern large models smarter, rather than producing “garbage in, garbage out.” A model continues training on apparently unoriginal data that it generated itself. From the perspective of classical theory, this attempt to lift itself by its own bootstraps is absurd. Under a computationally bounded framework, however, the model uses inference and generation to make latent deep structure explicit.
Project this principle onto human history and you will find that the greatest thinkers often had similar experiences. Consider Isaac Newton. To escape the severe plague in London, Newton shut himself away at the remote Woolsthorpe Manor. For more than a year, he completely severed contact with the outside academic world. With no new information coming in, relying only on a few basic physical intuitions and very simple mathematical axioms, Newton calculated furiously inside his own mind and ultimately constructed the vast, rigorous system of classical mechanics and the fundamental theorem of calculus.
Wang Yangming was the same. The History of Ming records:
(Yangming) was exiled to Longchang, a desolate place without books, and daily worked through what he had learned before. Suddenly he realised that the investigation of things and the extension of knowledge should be sought within one’s own mind, not in external objects. He sighed, “The Way is here.” Thereafter he believed without doubt. His teaching centred on extending innate knowledge. He held that after Zhou and the two Cheng brothers of the Song, only Lu Xiangshan’s simple and direct approach continued the transmission from Mencius, while Zhu Xi’s Collected Commentaries, Questions and Answers, and similar works were unsettled views from Zhu’s middle years. Scholars followed him in great numbers, and the world thus came to speak of the “Yangming School.”
Most people in modern society place excessive faith in the magical power of “acquiring new information.” They suffer from a severe Fear of Missing Out (FOMO), as if forcing themselves to scroll through hundreds of information-packed industry reports and listen to dozens of podcasts with avant-garde opinions every day would automatically make them smarter.
This is an illusion that puts the cart before the horse. Genuine barriers of knowledge depend intensely on deep internal computation. Once you have worked hard to master sufficiently strong first principles, the action you most need is to close the door decisively, cut off the raging stream of random entropy from the outside world, and use your own brain to wage a ruthless game against itself.
Infer furiously, calculate repeatedly, and collide violently. Think through how a minimalist rule changes under different extreme scenarios. This apparently tedious, deterministic internal reorganisation can absolutely temper astonishing structural intelligence within you.
Qualitative change in knowledge always happens through profound internal computation, never through restless external search.
Conclusion
Return to the thought with which this essay began. Information science in the real world is, in essence, a rigorous science of allocating extremely limited cognitive resources.
Because our physical lives and mental compute are extremely limited, we must firmly refuse to waste precious compute on random events that look lively but are in fact high-entropy. News headlines, short-term stock-price movements, and gossip merely exhaust your Time-bounded Entropy. Your goal is to find knowledge tested by time, with deep logic and long-range dependencies, and relentlessly extract its Epiplexity.
Because our physical lives and mental compute are extremely limited, we must completely abandon the attempt to become a human camera that remembers every detail, and bravely embrace our “lack of compute.” Because you cannot remember everything, you are forced to seek the macroscopic patterns behind things, summarise laws, and create abstract concepts. Limitation is the strongest catalyst for emergent intelligence.
Because our physical lives and mental compute are extremely limited, we must deliberately and courageously choose the reverse and extraordinarily difficult path of reasoning, forcing the brain to undergo a genuine foundational upgrade. Be wary of forward knowledge that has been chewed up and fed to you. Decompose, reverse-engineer, and infer causes from effects. Amid the profoundly taxing discomfort, the neurons in your brain are undergoing substantive rewiring.
Because our physical lives and mental compute are extremely limited, we must stop the endless intake of external information. Leave large stretches of empty time. Use the first principles you already possess to perform deep calculations and self-play inside your mind. Through deep internal computation, you can create an entirely new system of knowledge that belongs to you.
The expert knows with absolute clarity that they are only a mortal vessel with limited compute. They never aspire to become the omniscient and omnipotent Laplace’s Demon. In this barren universe filled with endless noise, they simply devote themselves, unwaveringly and without distraction, to the ultimate survival algorithm: extracting Epiplexity.
-
For readers without a background in physics, “Laplace’s Demon” is a famous thought experiment in the history of science, proposed by the French mathematician Pierre-Simon Laplace in 1814. Laplace imagined an intelligent being that knew the precise position and momentum of all matter at a given instant and possessed extraordinary power to process the data. Under the causal laws of classical mechanics, the universe’s entire past and future would then be determined for it. The concept is the ultimate expression of mechanical determinism: a clockwork universe whose every development can be foreseen. The development of quantum mechanics in the twentieth century shattered this fantasy. Heisenberg’s Uncertainty Principle established at a fundamental level that a microscopic particle’s position and momentum cannot both be measured precisely. Even such a “demon” could therefore never obtain the initial data needed to infer the future, invalidating the omniscient model scientifically. ↩
-
Yes, Anthropic’s AI product is named after Shannon. Personally, I think Company A is a company with taste. ↩
-
For readers familiar with etymology: Epiplexity, literally, means “complexity on top of existing complexity.”
- The prefix Epi- (ἐπί): Greek for “upon,” “added,” or “outer.” It commonly indicates a higher dimension, a superimposed layer, or something derived on top of a foundational structure.
- The root -plexity: From the Latin plectere, “to weave” or “to entwine.” It shares an origin with complex and refers to multiple elements woven together into a whole that is difficult to separate.