What does the 8 October preprint claim?

arXiv’s abstract page for 2610.11794 shows a 8 October 2026 publication timestamp for v1. The title frames Memento 3 as model-based recursive self-improvement through reflective rulebooks, building on the authors’ earlier Memento series.

The abstract says a single-model agent clears every level of all 25 public ARC-AGI-3 games, reaches mean Relative Human Action Efficiency 100.0, and uses 44% of the human action count. For other world-model coverage on this site, see Odyssey-3 and DreamTrue’s robot world-model preprint.

A separate Atari Pong case study in the abstract reports a learned feedback controller winning 21:0 across three evaluated episodes with different openings, without further LLM inference at test time. That is the authors’ case study, not our measurement.

Memento 3’s headline is methodological: learn an explicit, code-shaped model of world rules while keeping the underlying LLM frozen. That split is why the authors can talk about reflective rulebooks without claiming a new frontier weights release. Read the arXiv object for the algorithm; treat any ARC-AGI-3 leaderboard screenshot as author-reported unless the public board row matches.

Daytime view of the UCL Portico Building colonnade and dome under a blue sky. No identifiable people appear in the frame.
UCL Portico Building, CC BY-SA 3.0 via Wikimedia Commons (File:UCL Portico Building.jpg). Campus context for the UCL-affiliated authors; it does not depict ARC-AGI-3 levels or Atari Pong controllers. Photo: LordHarris at English Wikipedia / Wikimedia Commons. CC BY-SA 3.0 · Cropped and resized.

How does Code as Model actually work?

The paper’s core design keeps a natural-language rulebook as semantic memory—revisable hypotheses about objects, dynamics, and goals—and compiles that rulebook into executable code for prediction and planning.

Updates follow a loop of observation, reflection, rule revision, compilation, and verification. The abstract and introduction say new code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces observed transitions.

A population extension maintains multiple world models in parallel that share interaction evidence. N=1 recovers the single-model setting. We are summarizing the paper’s description, not endorsing completeness.

Figure 1 in the paper maps numbered natural-language rules onto branches of a simplified transition function that a planner can query. The accompanying text stresses that the rulebook states what the agent believes while execution fixes how those rules are applied.

The introduction situates Memento 3 against earlier Memento work on episodic memory, stateful reflective learning, and procedural skills: this paper moves the external memory object from policies/skills to explicit world-model hypotheses.

Code-as-model means the hypothesis about the environment is stored as editable programs or rules the system can revise after feedback, rather than only as opaque weights. The practical promise is inspectability; the practical risk is that buggy rule code fails loudly in ways weight updates might have smoothed over. Ask what verifier checks a candidate rule before it enters the rulebook.

Red brick and grey stone exterior of the British Library under an overcast sky. No identifiable people appear.
British Library exterior, CC BY-SA 3.0 via Wikimedia Commons (File:British Library Exterior IMG 1173.JPG). Nearby research campus context in London; not a screenshot of Memento 3’s rulebook or compiled world-model code. Photo: Deror_avi / Wikimedia Commons. CC BY-SA 3.0 · Cropped and resized.

What is frozen, and what is allowed to change?

The authors emphasize that the underlying LLM parameters stay fixed while the external rulebook and executable change. That is the sense in which they call the process recursive self-improvement without fine-tuning.

That design choice matters for cost and safety reviews: you are not merging gradients into the base model each episode, but you are still trusting LLM judgments inside the verification loop. Related framing on agent memory is in our agent memory design guide.

The paper discusses version-space ambiguity: many programs can fit the same history and disagree off-distribution. Replay rejects inconsistent programs; it does not uniquely identify the true rule.

The authors also report that preferring concise rulebooks and implementations acts as a practical minimum-description-length bias under the fidelity and replay constraints. That preference is a heuristic inside their loop, not a proof of uniqueness.

How should readers treat the ARC-AGI-3 leaderboard claim?

The authors compare against Claude Opus 5 with the ARC Prize Standard harness and report a 59.3-point higher mean RHAE for their agent. That comparison is theirs, on their evaluation setup.

ARC-AGI-3 public-game clears are impressive if replicated. Until independent runs land, treat RHAE 100.0 and the 44% action-count figure as author-reported results from the preprint.

We did not download code from CatalyzeX or ancillary links, and the abs page’s code finder is not a substitute for a reviewed release.

ARC-style benches change versions and private holdouts. If the paper cites a public board, open that board and match the submission name, date, and score. If the citation is only an internal table, keep the number in the authors’ column. Either way, ARC transfer to your production tools is unproven.

What practical implication exists for agent builders?

If you already externalize memory, Memento 3 argues for storing explicit world rules—and compiling them—rather than only storing episodes or skills.

The verification gates (fidelity + exact replay) are the load-bearing safety story inside the paper. Teams copying the slogan without those gates are not copying the method.

Pong’s zero-LLM controller is a narrow control result after learning, not a claim that every domain yields a classical controller.

Compared with end-to-end RL on the base weights, an external rulebook is easier to inspect and roll back—but only if teams keep the verification gates. A rulebook that is edited freely without replay checks is just another mutable prompt file.

Frozen-LLM plus evolving memory is an architecture pattern you can prototype with durable skill files and agent memory design habits: version the rulebook, unit-test it, and never let the frozen model silently overwrite audit trails.

What did we not test?

We did not run Memento 3 on ARC-AGI-3 or Pong, inspect private environments, or verify Huawei/UCL compute setups. All benchmark figures above are quoted from the preprint we opened.

Common questions

Is Memento 3 a shipped product?

No. The dated object we reviewed is an 8 October 2026 arXiv preprint. Product packaging, licenses, and support are not established by the abs page alone.

Does the LLM get fine-tuned?

The authors state the underlying LLM remains fixed while the rulebook and compiled code are revised. We did not inspect training scripts beyond the paper’s claims.

Are the ARC-AGI-3 scores independent?

No. They are reported by the paper’s authors. We did not reproduce them.

THE TAKEAWAY

What to remember

Memento 3 is an 8 October preprint about learning code-shaped world rules without fine-tuning the LLM—chase the algorithm and verifier design; keep ARC rows labeled as author-reported until the public board agrees.

Sources & further reading

  1. Memento 3: Model-Based Recursive Self-Improvement… (abs) ↗
  2. Memento 3 HTML ↗
  3. arXiv API query 2610.11794 ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories