The Museum of Models is a small project with a large claim: it has collected the answers 371 AI models gave to the same set of questions. That is the whole pitch. No leaderboard, no win-rate chart, no Elo. Just a room full of answers, arranged so you can walk from one model to the next and read what each one said.

The number is the story. Three hundred and seventy-one. Not three hundred seventy-one models that matter. Three hundred seventy-one models that exist, or existed, or shipped as a checkpoint someone could call. The archive treats each as a specimen worth keeping, which is a strange and useful editorial decision in a market that mostly wants a winner.

Why a museum is the right metaphor

Benchmarks are scoreboards. They compress a model into a number and throw away the thing that produced it. MMLU, GPQA, SWE-bench, AIME: each is a ruler, and rulers are only as good as the assumption that the thing being measured is one-dimensional. Anyone who has actually used two frontier models side by side knows the assumption is false. The same prompt yields answers that differ in tone, in hedging, in willingness to say “I don’t know,” in how they format a table, in whether they cite a source at all.

A museum does not rank. It preserves. That is the gap the Museum of Models fills, and it is a gap the benchmark industry has no incentive to fill. Labs publish the numbers that flatter them. Aggregators publish the numbers that drive clicks. Nobody publishes the raw answers, because raw answers are unflattering to everyone.

371 is a warning about the long tail

Here is the part the AI economy should sit with. If a casual project can enumerate 371 models, then the model supply is not a handful of frontier labs. It is a long tail: fine-tunes, merges, quantizations, community forks, regional models, abandoned checkpoints. Every one of those was, at some point, somebody’s release. Most will never be evaluated by anyone, which means most of the “AI capability” in the world is invisible to the benchmarks that shape policy, procurement, and press coverage.

That has consequences. Regulators writing model-evaluation rules tend to write them for the models they can name. Procurement teams buy on the leaderboards. If the real distribution of models is 371-deep and the measured distribution is 12-deep, then the entire evaluation apparatus is sampling the head and ignoring the body. The Museum of Models is a reminder that the body exists.

What a fixed question set actually tests

There is a methodological question the project raises without answering: what were the questions? A fixed prompt set is a lens, and a lens has a shape. If the questions skew toward reasoning, the archive rewards reasoning models. If they skew toward instruction-following, it rewards the tuned ones. The value of the archive is not that it settles anything. It is that it makes the question set visible, so a reader can argue with it.

That is more than most leaderboards offer. A benchmark publishes a score and hides the items. A museum publishes the items and lets you score them yourself. For anyone doing model selection, that is the more honest artifact. You can read two answers and decide which one you would actually ship.

The economics of keeping everything

Archiving 371 models’ outputs is cheap. Archiving their weights is not. This is the quiet tension in any “museum” framing for AI: the exhibit is the answer, but the artifact is the checkpoint. Answers are kilobytes. Weights are gigabytes to terabytes, and they come with licenses, gated repos, and deprecation notices. A model that answered a question in 2024 may be unreproducible in 2026 because the endpoint was retired, the weights were pulled, or the license forbade redistribution.

So the Museum of Models is, whether it intends to be or not, an argument about preservation. The AI industry is generating a historical record at a rate no archive can keep up with, and it is doing so on infrastructure that is designed to be replaced. The answers survive. The models mostly do not.

The exhibit is the answer. The artifact is the checkpoint. Only one of them is being kept.

What this means for builders

If you are choosing a model, the useful move is to stop reading leaderboards first and start reading answers. Find a prompt set that resembles your workload, run it against the candidates, and read the outputs. The Museum of Models is a proof that this is tractable at scale. You do not need a vendor’s eval harness to compare models. You need a fixed set of questions and the patience to read.

If you are building on top of models, the archive is also a warning about lock-in. The model you build against today is one of 371, and the number is going up. Abstract the model call. Keep your prompts portable. Assume the endpoint will change. The museum exists because models come and go, and the ones that go are mostly forgotten. That is not a reason to stop building. It is a reason to write down what you built against.

The outstanding question is whether anyone maintains this. A museum with 371 specimens is a snapshot. A museum with a submission process, a versioning scheme, and a way to re-run old prompts against new models is infrastructure. The first is a Product Hunt post. The second is something the field actually needs, and nobody with a leaderboard has an incentive to build it. Watch whether the archive grows, and whether the answers stay reproducible. That is the test of whether it is a museum or a scrapbook.

{/* TODO: confirm the specific question set used in the Museum of Models archive, and whether the project documents its prompts, licenses, or re-run process */}