What shipped in Seahaven 0.6.0?

PyPI’s JSON API lists seahaven 0.6.0 with upload_time_iso_8601 2026-10-08T21:06:41.202687Z, authored by Steve Cosman, MIT licensed, and requiring Python >=3.14. GitHub’s Kiln-AI/Seahaven repository shows a push on 8 October 2026 and an MIT SPDX license. Those timestamps are the concrete release event for this article.

The README defines Seahaven as a synthetic-world framework for agent evals and RL: you recreate the tool surface an agent would see—REST APIs, sandboxed SQL, search, custom formats—while each run gets an isolated instance. The name nod to The Truman Show is explicit in the README; it is branding, not a film-rights claim.

For other October agent-evaluation tooling already on Ai Lookout, compare Microsoft ThinkingBox stateful benchmarks and how teams evaluate AI agents.

How do the worlds stay stateful and resetable?

Each world instance owns an independent SQLite database. Fixtures freeze starting states with names such as small_startup, agency, or big_co so many runs begin from the same seed. Reproducibility is stated as the combination of fixture, controllable clock/time, and random seed.

Grading targets world state: every row the agent changed is logged, so evaluators can score outcomes instead of reading chat transcripts for vibes. That design choice is the practical center of the release for teams drowning in ungrounded agent demos.

Worlds are composable. The README’s example is a company world that includes StripeAPIWorld and ShopifyAPIWorld packages so payment and storefront tools share one agent-visible state document rather than three disconnected mocks.

Concurrent serving is part of the same story: the README says a process can host hundreds of instances. That matters for GRPO-style loops and nightly eval sweeps where spinning a shared staging database would contaminate every run. The change log gives graders a structured diff even when the agent’s natural-language summary is incomplete or optimistic.

Stacked shipping containers at the Port of Rotterdam under a bright sky. No people appear.
Container stacks at the Port of Rotterdam, 26 August 2017, AgainErick. CC BY-SA 4.0 via Wikimedia Commons. Isolation metaphor for parallel world instances; not a Seahaven screenshot. Photo: AgainErick. CC BY-SA 4.0 · Cropped and resized.

How do OpenEnv and MCP fit?

seahaven serve hosts a world as an OpenEnv environment. The README claims each process can host hundreds of parallel instances and thousands of requests per second; those throughput numbers are vendor-stated and were not re-measured here. Any OpenEnv client—or publication to Hugging Face environments—can drive the server.

seahaven mcp serves one world to an MCP client so a human can poke the same tools from an editor or chat app before wiring RL loops. A web console on serve lets you open instances, call tools, and inspect state in a browser.

Authoring is intentionally plain Python: tools are functions, tests use pytest, and seahaven new writes a project skeleton including AGENTS.md aimed at coding agents that will maintain the world. seahaven check is described as returning exact fixes for mistakes.

Built-for-agents authoring is not a slogan alone. seahaven new emits AGENTS.md pointing at the installed docs version, and seahaven check is described as returning exact remediation text. The practical test for a team is whether a coding agent can extend a world without forking undocumented private APIs—the README claims extension seams are first-class and ships an XML-RPC example as a worked pattern in broader project materials referenced from the docs tree.

Seahaven README comparison of environment options
NeedSeahavenProduction/stagingHand mocks
Realistic tools and dataYes (design goal)YesOften no
Stateful across tool callsYes (SQLite instance)YesOften no
Private instance per runYesNo (shared)Yes
Evening light on Scourie Pier with calm water and distant hills. No people appear.
Evening at Scourie Pier, 8 August 2010, Peter Moore. CC BY-SA 2.0 via Geograph/Wikimedia Commons. Harbor context for the Seahaven name; not Kiln facilities. Photo: Peter Moore. CC BY-SA 2.0 · Cropped and resized.

What should builders verify before adopting it?

Confirm Python 3.14+ in the environment that will run CI and RL workers. Teams pinned to older interpreters will need an upgrade path before Seahaven is more than a laptop experiment.

Map your real tools to world tools carefully. A synthetic Stripe surface that diverges from production semantics will train agents on the wrong affordances. The StripeAPIWorld example is a starting pattern, not a guarantee of API fidelity.

Runtime isolation for coding agents remains a separate layer from synthetic eval worlds. See AWS Strands Box when the question is where untrusted agent code executes, not how you grade its tool use in a mock company.

Also decide who owns fixture realism. Product managers who invent cheerful demo companies will under-train agents for messy inventory or refund edge cases. Pull anonymized shapes from production—table names, status enums, failure modes—into fixtures, then keep secrets out of the SQLite seeds you commit.

A DEC VT420 text terminal with a German keyboard on a desk. No people appear.
VT420 terminal photographed by Jacek R. CC BY-SA 3.0 via Wikimedia Commons. Tooling/terminal metaphor for agent-facing interfaces; not Seahaven’s web console. Photo: Jacek R. CC BY-SA 3.0 · Cropped and resized.

What should readers do next?

Install from PyPI into a 3.14+ environment, read the MIT LICENSE, and run the quickstart until you can reset a fixture and observe a state diff after a tool call. Only then wire OpenEnv into an RL loop.

If you already use Kiln for prompt/eval workflows, test the documented connection path from Kiln to a Seahaven world so scenario authors and RL engineers share one world definition.

Finally, separate Seahaven from skill-sync tools such as TeamAI. Seahaven answers where the agent practices and how you score state; TeamAI answers how a human team shares skills across harnesses. Most organizations eventually need both layers, but they fail for different reasons and should be piloted on different weeks.

  1. Pin seahaven==0.6.0 (or newer) and record the PyPI upload time in your runbook.
  2. Author one fixture that matches a real customer-support or checkout path before scaling parallel instances.
  3. Keep OpenEnv throughput claims labeled as README statements until you load-test on your hardware.

Common questions

Is Seahaven a hosted product?

No. The primary materials describe an open-source Python package you run yourself. Kiln the app can connect to Seahaven worlds, but the framework’s MIT release is the news object here.

Why does it require Python 3.14+?

PyPI metadata for seahaven lists Requires-Python >=3.14. That is a hard environment gate before you chase OpenEnv or MCP demos.

Does Seahaven replace production systems?

The README’s comparison table says the opposite: production or staging can be realistic and stateful but not privately instanced per run; hand mocks are private but often not realistic. Seahaven aims at the middle.

THE TAKEAWAY

What to remember

If your agent evals still grade prose, Seahaven’s concrete offer is fixture-reset worlds with SQLite state and OpenEnv/MCP adapters—so you score what the agent changed, not how confident it sounded.

Sources & further reading

  1. Kiln-AI/Seahaven README ↗
  2. Seahaven LICENSE (MIT) ↗
  3. seahaven on PyPI ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories