Science systems / Agent infrastructure / Collective memory
We Are Not Building
an AI Scientist
From bricks and glue to a scientific bridge for everyone.
Sometimes I imagine science as a bridge that has not yet been completed.
One end rests on the world we already understand. The other disappears into the unknown.
For thousands of years, generation after generation of scientists has stood at the unfinished edge and reached into the fog with the next brick. Newton laid some of them. Maxwell laid others. Einstein changed the load-bearing structure of part of the bridge. Far more people whose names never entered a textbook measured materials, corrected errors, repeated experiments, proved theorems, or discovered that a brick everyone trusted was not as solid as it appeared.
What we call scientific knowledge is, in one sense, the collection of bricks that survived this continuing process of inspection. But the hard part of science has never been simply producing more bricks.
01 / ORIENTATIONWhere should the next brick go? And do we actually know what is already on the bridge?
Those two questions are where SciYard begins.
01Science Is an Astonishingly Expensive Search over Paths
In the language of machine learning, the history of science resembles a reinforcement-learning process that has been running for several thousand years. Reality supplies observations. Scientists propose hypotheses, perform experiments, receive support or contradiction, and revise their internal models.
Some paths fail immediately. Others collect enough positive signal to become research programs and eventually entire disciplines. On rare occasions, someone leaves every path currently being explored, arrives somewhere unexpected, and finds another route forward.
This is a peculiar learning system. It has no central brain. Every participant has limited compute, a limited lifetime, a limited view of the data, and a reward function that is never quite the same as anyone else’s. One person wants to understand nature. Another needs to publish. Another needs funding. Another is solving an engineering constraint. Someone else has simply been unable to stop thinking about one strange question for ten years.
An intuition built by a physicist over twenty years cannot be copied into another mind like a set of neural-network weights.
We compress a small part of that intuition into language, papers, equations, images, and code, then pass those artifacts to the next person. If we insist on a computational analogy, civilization looks like a deeply imperfect form of heterogeneous federated reinforcement learning: each node explores independently; compute is scarce; rewards disagree; gradients cannot be transmitted; synchronization happens through an extremely narrow channel.
From the standpoint of optimization efficiency, this is a terrible system. We read the same work repeatedly. We relearn what previous researchers already understood. A young scientist may need ten or twenty years of training before reaching the active frontier of a field. Much of the learning process vanishes when a person leaves.
We preserve many results. We lose an enormous number of learning trajectories.
02That Inefficiency May Protect What Science Values Most
What would happen if every scientist really could share gradients? Imagine a single planetary scientific model. Whenever anyone observed a new experimental result, they would upload a gradient. The system would average its parameters. The whole world would converge continuously toward one world model.
It sounds wonderfully efficient. I am not sure it would be good science.
Many of the most consequential changes in scientific history began when someone did not continue descending along the average gradient.

In a letter to Maurice Solovine, Einstein drew sense experience, E, beneath a system of axioms, A. From axioms, logical consequences could be deduced and returned to experience for testing. But between E and A, Einstein did not draw a staircase of inevitable logical steps. He drew a leap.
Tom Zahavy revisits this picture in LLMs can’t jump. Scientific invention, on this account, involves more than induction and deduction. It involves abduction: proposing an explanatory structure from limited experience. It involves a Jump that cannot be reduced to the smooth continuation of an already visible path.
I like the word Jump. If science is a bridge, the Jump occurs when the last brick under your feet has ended and no existing road tells you where the next one belongs. You must first step into the air. Only then might you place a brick beneath that step.
03An LLM Is Not the Next Brick
After generative AI arrived, a natural story took shape. A model has read a significant fraction of recorded knowledge. Perhaps it can formulate questions, conduct research, publish papers, and eventually become an autonomous scientist.
I have always been cautious about that picture. Not because I think AI is unimportant to science. Quite the opposite: I think the autonomous-scientist story may cause us to underestimate where AI is genuinely unprecedented.
02 / THESISAn LLM is not, by itself, a brick in the bridge of science. It is closer to glue.
A scientific brick must have an address. It must be testable, challengeable, reproducible, and explicit about the structures on which it depends. A beautiful chain of reasoning produced during one inference does not automatically become scientific knowledge. Only after an idea is formalized, checked, and placed inside a record that people can continue to inspect does it become part of the bridge.
An LLM offers something different. For the first time, we may be able to construct previously unaffordable connections across the enormous surface of existing knowledge.
A structure in a materials-science paper may share a mathematical mechanism with a statistical-physics paper written decades earlier. A problem in combinatorial geometry may need a tool hidden in algebraic number theory. The dynamical explanation a biologist needs may already exist, under a completely different vocabulary, in nonequilibrium chemistry.
Historically, whether these connections occurred depended largely on whether one person happened to have read both literatures. An LLM changes the probability of that encounter. It can place distant knowledge structures beside one another at a scale no individual reader can sustain.
That is not automatically a new brick. It is an opportunity for separated bricks to touch—and for an unnoticed gap in the bridge to become visible.
04The Unit Distance Problem: A Connection across Fields
A development in 2026 made this idea unusually concrete. The planar unit distance problem, posed by Paul Erdős in 1946, asks how many pairs of points among n points in the plane can be exactly distance one apart. For decades, a powerful intuition held that grid-like constructions were essentially optimal in the asymptotic sense.
Work released in 2026 produced, for a fixed δ > 0, infinite families with at least n1+δ unit-distance pairs, refuting the conjectured n1+o(1) upper bound. External mathematicians checked and reorganized the argument, and a more explicit bound followed.[2][3]
What fascinates me is not the headline that an AI system contributed to a long-standing mathematical breakthrough. It is where the solution came from.
The decisive tools did not remain inside the local neighborhood of combinatorial geometry. They reached into algebraic number theory: infinite class field towers, the Golod–Shafarevich theory, number fields with controlled discriminants, and many prime ideals of small norm. Those algebraic structures could then be turned into planar configurations with many unit-length differences.
These tools were not unknown to algebraic number theorists. What was obscure was their connection to a Euclidean question about distances in the plane. The episode does not show that an LLM can replace mathematicians. It shows something more specific and, to me, more useful: AI may be exceptionally good at discovering which bricks could be bonded together.
Once distant structures are connected, a second possibility appears. The connection may expose a hole in the bridge where everyone thought the surface was already complete. Sometimes we see no opening because an entire community has been looking in the same direction.
05Related Papers Are Not the Same as Scientific Discovery
This is why the future of AI for science cannot be merely a better paper-search engine. The default workflow of many academic AI systems is already familiar: state a question, retrieve the most relevant papers, read them, summarize them, and generate an answer.
That workflow is valuable. But semantic relevance has a built-in bias. It repeatedly returns us to knowledge neighborhoods that have already formed.
If you study the unit distance problem, an excellent related-paper system will retrieve more combinatorial geometry, incidence geometry, and lattice constructions. It may reconstruct with great accuracy how the field has historically understood the problem. Yet the critical question may be: why must we keep speaking the language of this field at all?
Recover the history, evidence, and local neighborhood.
Expose assumptions, structural needs, and distant bridges.
A scientific system needs Recall, but it also needs Reframing. Which assumptions are evidence, and which are merely the heuristics a field has learned to stop mentioning? If every disciplinary term is removed, what structure remains? Where else in the world has someone already developed a tool for that structure?
This is where an LLM starts to behave like glue. Not by bonding the most similar objects, but by proposing connections between objects that nobody knew should be adjacent.
06SciYard Should Maintain Identity, Not Merely Produce Answers
Glue alone is not enough. If an LLM can generate an unlimited number of connections, how do we know which ones are real? If models continually summarize, rewrite, and compress knowledge, will we still know where a conclusion came from five years later? If one Agent says A and another says B, should a third model be allowed to dissolve both into a pleasant-sounding compromise C?
If every model run regenerates a new world, where does the identity of scientific knowledge live?
This question may be more fundamental than making an Agent smarter. It is also the question that increasingly defines SciYard.
SciYard began with a concrete observation: research does not stop when a paper is published. Limitations discovered by the authors, community replications, later corrections, new interpretations, disputes, and the real position of a paper in the knowledge network do not flow back automatically into a static PDF. Every paper has an afterlife.
But a paper’s afterlife is only the entrance to a larger problem. What must be maintained over time is the identity of every scientific object. What did a paper actually claim? Which elements are observations and which are hypotheses? Which claims have independent support? Which were revised? What prior results do they depend on? Who disagrees, and why? What evidence exists, and what evidence is conspicuously absent?
- PROVENANCE
Where each claim came from and what transformations it has undergone.
- DISAGREEMENT
Preserve conflict instead of flattening it into a synthetic consensus.
- UNCERTAINTY
State what is known, what is not, and what justifies the boundary.
- TIME
Allow corrected and superseded knowledge to retain a historical identity.
When these objects retain identity instead of being repeatedly compressed into fresh natural-language summaries, we begin to acquire a scientific map capable of evolving without forgetting how it changed.
07I Would Rather Call SciYard a Science Harness
The word harness matters. A harness does not decide where you should go. It is not the driver, and it is not the destination. It is a structure that allows powerful capabilities to be organized reliably, constrained deliberately, and inspected afterward.
I do not want SciYard to become a machine that announces to humanity, “This is the ultimate scientific problem you should solve next.” Why a question matters still comes from people. Curiosity comes from people. Meaning comes from people. The choice to spend decades of a life on one problem should not be erased by a universal reward model.
Nor should SciYard exist to replace scientists. If science is a bridge, people must remain the builders who decide where the bridge should lead.
03 / SYSTEM ROLEHumans determine the direction. The harness maintains the state of the bridge.
It should know which bricks already exist, why each one can bear weight, which bricks depend on others, where the structure is cracking, and where there has never been a brick at all. AI can then help us search this immense structure for connections.
That is what I mean by a Science Harness.
08Share Knowledge at High Speed without Averaging Thought
A model that “knows” one hundred million papers does not necessarily give humanity a better scientific system. Knowledge inside model parameters is compressed, approximate, and vulnerable to the lifecycle of the model itself. Science needs objects that can outlive a checkpoint. It needs provenance, identity, disagreement, and uncertainty. A brick should still know why it is there decades after it was placed.
I do not want all scientific knowledge to disappear into an ever-larger foundation model. That would amount to constructing the central brain that human science has never had.
The absence of a central brain is not only a defect. Our rewards differ. Our experiences differ. Our knowledge differs. We frequently misunderstand one another. This makes science inefficient, but it also preserves diversity.
If every Agent has the same knowledge, model, prompt, and reward, then a “multi-agent system” may be nothing more than repeated sampling from one cognitive distribution. The valuable property is not the number of Agents. It is the persistence of genuine cognitive difference.
A future scientific AI system should therefore accomplish something that sounds contradictory: allow knowledge to be shared at extraordinary speed without averaging thought.
Share evidence. Share existing bricks. Share failed paths. But retain competing hypotheses, distinct world models, and lines of inquiry that only a small minority currently believes. Then let experiments, proofs, and reality determine which proposals become new bricks.
09Humans Choose the Direction. AI Finds Connections.
This may be a more meaningful future than the Autonomous Scientist: not one super-scientist that thinks on behalf of everyone, but an infrastructure in which millions of researchers and millions of AI Agents can work around the same scientific bridge.
Humans continue to ask the questions worth asking. Humans continue to make Einstein’s Jump. AI helps us see connections across disciplines, decades, and languages. At times, it may help people discover a possible Jump that nobody had realized was available.
SciYard’s role is to ensure that these explorations do not begin again from zero every time: that a failed path does not immediately disappear; that a disagreement leaves a structure; that a cross-disciplinary connection retains its source; that a paper continues to evolve after publication; that when a new model arrives, it does not have to reinvent the memory of science.
Papers, theorems, experiments, and reliable discoveries are the bricks.
LLMs are the glue.
SciYard is the harness around the construction process.
It does not decide the bridge’s destination. It should not claim to know where the destination is. It tries, as faithfully as possible, to maintain where we have reached, why we are able to stand here, and where—immediately ahead—there is still no bridge.
That is what we are building at SciYard.