What did H2O.ai publish on 7 October?

Jonathan McKinney’s H2O.ai blog post is dated October 07, 2026. The lede says the 4-billion-parameter open model is #1 on JevBench’s official Composite Score, outscores Jev itself 72.5 to 71.5, and beats every 12B, 26B and 31B open model on the board. It states Apache-2.0 availability on Hugging Face, stock vLLM serving, and that the current release reads images.

JevBench is described as Benchmark Heaven’s public leaderboard for decision models: a record plus a fixed answer set in, a probability for each answer out. The blog’s table—read 7 October 2026 on board revision notes cited in the footnotes—places H2O-Lightning-4B v1.1 first among open-weight systems, with Jev 1.13.0 shown as an unranked reference row.

This is an open decision-model drop, not a hosted Foundry SKU. For Microsoft’s parallel hosted scorer, see Microsoft-Decision-1 in Foundry. For a small open Jev-shaped alternative, see WaterSheep.

Bar chart of JevBench v1.6.1 open-weights composite scores with H2O-Lightning-4B v1.1 at the top.
Chart file assets/jevbench_composite_2026-10-07.png from the Apache-2.0 h2oai/h2o-lightning-4b Hub repository. Vendor/board visualization; not an Ai Lookout rerun. Photo: H2O.ai / h2oai/h2o-lightning-4b. Apache License 2.0 · Cropped and resized.

How does Lightning-4B answer typed questions?

The model card describes a Qwen/Qwen3.5-4B fine-tune that answers choice, yes/no (noul), and score questions about a record and about images that come with the record. Each decision is one forward pass and one output token; cost is framed as input tokens only. Serving instructions pin unmodified vLLM 0.30.0 and a standard-library shim (h2o_lightning_shim.py) that exposes POST /v1/systemone.

Image input is part of the current public line. The blog’s note 6 says the board measured v1.1 as a text model, while image input—with its own measurements—is on the card. Up to four images may ride in one request as data URIs. A request with no image takes the text path that was scored on the board.

An adapter/ directory is described for running the decision adapter beside the base Qwen3.5-4B writer on one GPU. Demo numbers in the card—1,000 claims in 63 seconds, 192 vendors in 145 seconds, Freedoom played to an exit—are H2O’s logged runs on one H100. We did not replay them.

JevBench v1.6.1 open-weights composite rows as printed by H2O on 7 October
SystemCompositeIntelligenceCalibrationNotes
H2O-Lightning-4B v1.172.560.090.0Vendor/board table; not our rerun
Jev 1.13.0 (reference)71.563.690.6Unranked hosted reference on the board
Quyet-1.0-Large (31B)71.473.490.0Leads raw intelligence in this excerpt
High-speed photograph of a water drop impact crown on a wet surface. No people appear.
Water drop impact photographed 4 June 2016 by Nikk. CC BY 2.0 via Wikimedia Commons (File:Water drop impact.jpg). Illustrative for the H2O name; not a company lab and not a model artifact. Photo: Nikk. CC BY 2.0 · Cropped and resized.

Which scores stay in the vendor column?

Composite 72.5 is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost on JevBench’s official weights. The blog is clear that 31B models lead on raw Intelligence (73 vs 60) while the 4B wins on the blended score through calibration, cost and speed. Estimated cost is about $0.021 per 1,000 decisions; raw median latency cited is 29 ms before the board’s self-hosted adjustment.

The Hub card we opened on 11 October includes a 10 October accuracy / latency / composite animation that keeps Microsoft’s accuracy numbers, corrects latency onto one H100 basis, and prints JevBench composites beside them. The footnote says Microsoft’s page used the board’s adjusted latency for self-hosted rows and an 85 ms figure for Decision-1 that H2O replaces with JevBench’s measured API median of 460 ms. That paragraph is H2O’s rebuttal, not an independent timing study.

For another open decision path with a free API angle, see HAL-X THX-01. For a physics-inspired open scorer, see Vega 0.8B.

Digital caliper resting on a wooden surface showing a measurement readout. No people appear.
Digital caliper photographed 25 January 2020 by Jacek Halicki. CC BY-SA 4.0 via Wikimedia Commons (File:2020 Suwmiarka cyfrowa.jpg). Illustrative calibration tool; it does not depict JevBench scoring or H2O-Lightning weights. Photo: Jacek Halicki. CC BY-SA 4.0 · Cropped and resized.

How do you run the published checkpoint?

The card’s commands download revision v1.2.1 with hf download, serve with vllm serve, then start the shim on port 8741 (or bash serve.sh). Warm-up is a /v1/systemone noul probe. Real requests can carry several questions about one state. Those are documentation commands; we did not allocate a GPU.

The Hub API object lists pipeline_tag image-text-to-text, library transformers, base_model Qwen/Qwen3.5-4B, license apache-2.0, and tags decision-making, calibration, jevbench, multimodal. Downloads and likes on the object we opened are interest counters, not quality proofs.

Hosted API offerings on Benchmark Heaven’s separate API board are explicitly unranked on the open-weights composite. The blog names Sage 1.3.0 and Liquid AI’s d1 as higher on that other view. Do not collapse the two boards.

What remains unverified?

We did not download the safetensors, start vLLM, or call /v1/systemone. We did not rescore JevBench, rebuild H2O’s H100 latency chart, or verify Microsoft’s 36-benchmark accuracy table. We did not open Benchmark Heaven’s raw JSON beyond the figures printed on H2O’s pages.

Image-input quality claims beyond the card’s own notes are not independently checked here. Demo reel wall-clock numbers are H2O’s. Freedoom play is a controlled demo, not a robotics result.

This article is an evidence review of the 7 October blog, the Hub card, the Hub API object and the Apache-2.0 LICENSE as opened on 11 October 2026. It is not a first-hand bake-off against Decision-1, Jev or Quyet.

Common questions

Is 72.5 an independent Ai Lookout score?

No. It is JevBench v1.6.1’s official Composite Score as reported on H2O’s 7 October blog and model card, read against Benchmark Heaven. We did not rerun the board.

Does the model generate explanations?

The blog and card describe Jev-class typed decisions: state and a bounded rubric in, probabilities out, in one forward pass and one output token. Natural-language replies in the demo reel are attributed to the base Qwen3.5-4B path on the same GPU, not to the decision adapter.

How does this relate to Microsoft-Decision-1?

Microsoft’s Command Line post lists H2O-Lightning-4B among appendix models and, on the page we reopened 11 October, calls it the latency runner-up at 2.5× slower than Decision-1 in Microsoft’s harness. H2O’s 10 October card chart disputes Microsoft’s latency basis. Those are competing vendor narratives; we ran neither harness.

THE TAKEAWAY

What to remember

Use the 7 October H2O blog and Apache-2.0 Hub card for access and the JevBench composite snapshot; keep Microsoft’s accuracy/latency tables and H2O’s rebuttal chart as attributed vendor claims.

Sources & further reading

  1. H2O-Lightning-4B: #1 on JevBench, and Ahead of Jev Itself ↗
  2. h2oai/h2o-lightning-4b model card ↗
  3. h2oai/h2o-lightning-4b Hub API object ↗
  4. h2oai/h2o-lightning-4b LICENSE (Apache-2.0) ↗
  5. Microsoft-Decision-1 Command Line post (latency wording check) ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories