A developer going by plicerin has rebuilt Activision’s River Raid from its original Atari 2600 ROM, one 6502 instruction at a time, in JavaScript. The port is not an emulator. Every routine of the 1982 cartridge was rewritten by hand, working on the same 128 bytes of RAM at the same addresses. To prove it, a small 6502 core runs the real ROM beside the port and compares every frame byte for byte. The result: 0 differences across 45,000 compared frames of play.

That number is the story. Not the nostalgia, not the browser play button. The method.

What the port actually proves

The verification claims are specific and checkable. 300,000 river-generator states checked across sections 1 to 48. 31,840 pixels of the opening frame identical to a Stella capture. 4,096 bytes in the whole cartridge, code and graphics included. The port reproduces the original’s quirks rather than smoothing them: collisions come only from the TIA’s hardware latches, read once per river block, with no collision arithmetic anywhere. Flying into a ship, helicopter or jet still scores it as you go down. Touching a house runs the same code path as touching a fuel depot. When a new block has only just scrolled in, a collision at the very top gets filed under a seventh block that does not exist, and the flag lands in the first object’s position byte.

The port keeps all of it. That is the correct decision, and it is the decision most reimplementations get wrong.

The AI angle is the oracle, not the game

Here is the part worth extracting. The interesting thing in this project is not that a human can port a 2600 game. It is that the author built an oracle: a second implementation, running the original artifact, producing a comparable output stream, frame after frame, byte after byte. The port is graded against ground truth at a granularity fine enough that a single wrong instruction shows up in seconds.

Compare that to how most AI-generated code gets evaluated today. A model writes a function. A test suite passes. The suite was written by the same person who prompted the model, covers the happy path, and says nothing about the seventh block that does not exist. The River Raid port does the opposite. It does not ask whether the output looks right. It asks whether the output is identical to the reference, 45,000 times in a row.

Carol Shaw’s original is a useful stress case precisely because it is hostile to approximation. The river is never stored. Each 32-line block is generated just before it scrolls into view from a 16-bit random number generator seeded at $A814, the same seed every game. The whole playfield is derived, not saved. Any reimplementation that drifts on the RNG, on block boundaries, or on the collision latch produces a river that looks plausible and is wrong. A screenshot comparison would miss it. A frame-by-frame oracle catches it.

Why this matters for AI tooling

The AI coding tools shipping right now are optimized for the first draft. They are very good at producing code that compiles, runs, and appears correct. They are much weaker at producing code that is provably equivalent to a reference, because most teams never build the reference. The oracle is the expensive part. The generation is cheap.

The River Raid project is a working example of the discipline that fixes this. Pick a reference implementation. Instrument both. Compare at the finest granularity the domain allows. Treat every divergence as a bug in your code, not a rounding error. That is how you get from “the demo works” to “0 differences in 45,000 frames.”

There is a hardware lesson buried here too. The cartridge is 4,096 bytes. Code and graphics share that budget. The river is generated on the fly because there is nowhere to put it. Modern AI workloads run on the opposite assumption: memory is cheap, so store everything and retrieve it. That assumption is starting to strain. Inference cost, KV-cache pressure, and context-window economics all push toward the same question the 2600 answered in 1982: what can you derive instead of store? A 16-bit RNG and a seed beat a lookup table when the budget is tight.

The oracle is the expensive part. The generation is cheap. Most teams skip the part that would actually tell them whether the code is right.

The unglamorous work

Nothing here is AI-generated, as far as the writeup says. That is the point. The project is a demonstration of what verification looks like when someone cares about correctness at the instruction level, and it lands in the same month that every major lab is arguing about how to evaluate model outputs. The labs are building benchmarks. The River Raid author built an oracle. The second is harder to game.

For AI builders, the takeaway is narrow and practical. If you are using a model to port, translate, or reimplement existing code, you need a reference to diff against, not a vibe check. The reference can be the original binary, a golden output file, or a captured trace. What it cannot be is your own test suite written after the fact. The 2600 port got to zero differences because the author had something to compare against that he did not write.

The port is playable in a browser. The controls are arrow keys, space to fire, enter for game reset. Fuel drains every frame. The gauge runs from E to F. Fly over a depot to refill, or shoot it for 80 points and lose the fuel. Three jets to start, another every 10,000 points, up to nine. After a crash the river rewinds to the start of the section and replays it exactly, because the cartridge saved its random numbers there.

That last detail is the whole argument in miniature. The original saved its RNG state so the river would be reproducible. The port saved the same state at the same address so the comparison would hold. Reproducibility was a design constraint in 1982 and it is a verification strategy in 2026. The difference is that one team shipped it in four kilobytes, and most teams today still do not ship it at all.