Logo
Overview

Autoresearch, and the first map

2026.09.22
3 min read
The broken clay tablet on a museum stand: rows of cuneiform above a circular map ringed by a band of ocean

The Babylonian Map of the World, 700–500 BC, via Wikimedia Commons.

The earliest surviving map of the world is 8cm circle with a few lines representing the Euphrates river and cities along it. It encompassed the known world of the creator, bounded by the bitter river of the unknown beyond.

I want to make a napkin math argument that our current understanding of the universe is closer to this map’s author’s than what will come in the next decades.

It’s impossible to measure the growth of human knowledge of nature, but if we estimate based on the number of publications across different fields, in the next decade, we will approximately double the total production of knowledge that has existed.

Evidence

0.050.10.250.51251030todayAbove this line the next decade out-produces everything recorded before 2026Everything recorded up to end-2025 = 1×32.1 by 2035×2.70 by 2035×1.93 by 2035×1.56 by 2035arXiv AI papers (cs.AI)arXiv preprints, all fieldsScholarly works with a DOI (Crossref)Biomedical papers (PubMed)20002005201020152020202520302035cumulative output ÷ everything recorded up to 2025 (log scale)

Running totals as a multiple of each record’s total at the end of 2025. Data: arXiv, Crossref, PubMed.

Maps of the known world and cosmoscosmosEarthCopernicus1543Herschel1785CfA slice1986COBE1992Planck2013–18DESI2026Babylonian Mapc. 600 BCEPtolemyc. 150al-Idrisi1154Waldseemüller1507Tharp & Heezen1977Seabed 20302026050100150200250300350cumulative scholarly works(millions, Crossref)2× the dated Crossref works to 2025, the level the next decade must reach1665, first scientific journalForecast for 2026–2035, +155 million(80% range 140–181)Dated works registered with Crossrefto 2025, 167 million101103105107109items produced per year(log scale)manuscripts copiedper year (Latin Europe)printed books peryear (Europe)scholarly works peryear (Crossref)AI research papers per year(OpenAlex, to 2023)time axis stretched 27× after 1900600 BCE1 CE500100015001900195020002025

Time axis compressed before 1900. Data: Buringh & van Zanden via Our World in Data, Crossref, OpenAlex.

The next step in the exponential growth of knowledge is automation in agentic data science tools within more and more highly accurate simulation environments. In 2026 building is easy, but developing agents to generate knowledge is still hard.

I spent the better part of this spring creating agents that do experiment design, data source selection/preprocessing, run regressions and analyses, and interpret the results into a knowledge graph.

My friends this week are concerned that ICLR 2026 received 50k+ abstract submissions, more than all the previous years combined. Many are rightly skeptical of spam, and we cannot lose scientific rigor, but a reality may not be far off where we have 500k+ abstract submissions, and our knowledge of nature expands at a rate that we previously could not have imagined.

Autoresearch and simulation

A 3dgs of a stegosaurus skeleton from the AMNH, fully articulated and walking (without hand-created animations)
A jet engine with processed turbofan, articulated and running based on engine schematics
Sony CFS-43, photogrammetry of a radio cassette player and articulations of its door, keys, knobs, and selector
Sharpa hand picking up and unfolding sunglasses then putting them on a face
A Sharpa dexterous hand scooping to cup a delicate flower in its palm
Dual robot arms lifting a marker into a Sharpa dexterous hand, uncapping it, and writing Hello World on a whiteboard

I recently co-developed SceneAgent with Tianxing Fan and Hank Yang at the Harvard Computational Robotics Group to improve creation of simulation environments. I think this will be an important step toward autoresearch in simulation in the future.

Consider the delta in human understanding from the earliest surviving map, 2,600 years ago to today — and then imagine that same amount of growth occurring over the next 10 years, and again the next several years after.

Imagine being the first cartographer 2,600 years ago — gathering evidence from travelers, sorting truth from fable, sending out explorers. Expeditions returned starved, robed only in a wealth of dust, from the edge of the known. The map of the world would have to wait for the next king, the next empire, maybe.

Autoresearch—especially in connection with digital twins like the sort that we build—will be one of the most disruptive technologies in the coming decade. Training agents to follow the scientific process or in statistical analysis is less hard than using agents to create much more highly detailed digital twins of the universe—a factory, plasma accelerator, or the world—where we can run experiments much faster than in the real world.


Data and Sources

Map images via Wikimedia Commons, scaled: the tablet photograph by O. S. M. Amin (CC BY-SA 4.0), the Planck map by ESA and the Planck Collaboration (CC BY 4.0), and the DESI map by the DESI Collaboration and DESI Member Institutions/DOE/KPNO/NOIRLab/NSF/AURA/R. Proctor, image processing M. Zamani, NSF NOIRLab (CC BY 4.0). The others are public domain or CC0.