What did Base Compute publish on 8 October?

The Hugging Face post presents Superfluid as open-source under Apache-2.0 and installs via a curl installer or cargo install superfluid. Example serve lines pull a Hugging Face GGUF tag (for example unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M) or an MLX community directory.

GitHub’s basecompute/superfluid README matches that framing: one scheduler, one durable session log, and one API surface in front of llama.cpp, MLX, and baseRT worker processes. Workers are installed on first use from their upstream releases rather than vendored into the daemon.

Local GGUF serving is crowded; for a recent quantized coding checkpoint aimed at the same machines, see Underdog’s Saluki 27B GGUF. Superfluid is infrastructure, not a new weight release.

Superfluid’s pitch is local multi-agent serving: several agent loops sharing one machine’s model servers without each spinning an isolated full replica. That is an orchestration and scheduling story as much as a model story. Read the 8 October materials for the process model and APIs they document, and keep any tokens/sec comparison charts in Base Compute’s column until you reproduce them on your GPUs.

Open laptop on a desk with papers and a coffee cup nearby. No people appear.
Desk workspace still life released as CC0 by www.Pixel.la Free Stock Photos via Wikimedia Commons, dated 5 April 2015. Contextual home-office photo; it is not a Superfluid dashboard screenshot. Photo: www.Pixel.la Free Stock Photos / Wikimedia Commons. CC0 · Cropped and resized.

How does multi-agent scheduling claim to help?

Base Compute argues that data-center throughput servers were not built for one person sharing a laptop with several agents. Superfluid’s listed behaviors include preempting batch work so interactive chat is answered at the next scheduler tick, writing every token to a log before display so crashed workers can resume, and computing a shared system prompt once across many requests.

A headline comparison in the post says that with eight agents on 7.5k-token documents, a chat question arriving after two seconds saw Superfluid answer in 0.36s versus more than 10 seconds for llama-server on the same Mac—described as about 30× faster for that scenario. A table also reports TTFT for a chat request behind eight agents as 1.5 s on Superfluid versus hundreds of seconds on llama-server, Ollama, and mlx_lm.server.

Those figures are labeled as measurements with superfluid/tools/suite on an Apple M5 Pro with Qwen3.8-27B at 4 bits. They are not AiLookout timings.

Shared KV caches, admission control, and per-agent fair queues are the usual levers in this design space. Confirm which of those Superfluid actually implements, what happens when two agents thrash the same hot weights, and whether pin-to-GPU policies survive a host reboot. A scheduler that looks good on two chatbots can still starve a long tool-using agent.

Closed laptop beside a small desk lamp on a wooden table. No people appear.
Laptop and lamp still life by Tatiana Lapina, released as CC0 via Wikimedia Commons (File:Laptop and lamp on table (Unsplash).jpg), dated 3 February 2015. Illustrative local-machine context; it does not show llama.cpp workers or Superfluid fleet nodes. Photo: Tatiana Lapina veila / Wikimedia Commons. CC0 · Cropped and resized.

What client and agent paths are documented?

The post shows an OpenAI Python client pointed at http://127.0.0.1:8453/v1. It also says Anthropic SDK and Ollama clients work when aimed at the same server address.

superfluid launch claude --model … is documented to start the server if needed and point Claude Code at it. The same launch helper is listed for Codex, pi, OpenCode, Hermes Agent, Cline, OpenClaw, and the Ollama CLI.

Desktop coding harnesses remain a separate product layer; DeepSeek’s recent desktop Harness release is one example of an app that still needs a backend—see DeepSeek Harness v0.2.

Prefer the documented client SDKs and agent adapters over scraping example notebooks. If the server speaks an OpenAI-compatible surface, note which extensions are proprietary. Compatibility claims should be tested with your existing agent harness’s streaming, tool-call, and cancellation behaviour—not only with a hello-world completion.

What about multi-device fleets?

The post describes a head server that listens for nodes, issues a fleet token, and places whole sessions on joined machines over a link encrypted with that token. Nodes may run different runtimes for the same model—BaseRT on a Mac beside llama.cpp on Linux, in the authors’ example.

If a node dies mid-stream, the write-up says the session is replaced and no token is delivered twice. That failover language is Base Compute’s; we did not join a fleet.

Docs are pointed at superfluid.sh. The GitHub README marks Superfluid pre-1.0 and warns that protocols and on-disk formats may change between releases.

How should readers read the license and maturity?

Apache-2.0 on the server is a permissive software license for the daemon. It does not change the license of the model weights you serve. Keep open weights versus open source distinct when you publish a stack diagram.

Pre-1.0 means expect breaking changes. Production home-lab users should pin versions and back up session logs independently of the marketing latency table.

Shared prompts and long agent histories still burn context; scheduling does not replace context compaction for agents.

Check whether the open components and the hosted control plane share one license. Preview-quality schedulers often change queue semantics between tags; pin a release and record the config hash beside any benchmark you publish internally.

What did we not test?

We did not run the install script, serve a GGUF, launch Claude Code through Superfluid, or reproduce the M5 Pro benchmark table. This article reports the 8 October Hugging Face post and the GitHub repository pages we opened on 10 October 2026.

Common questions

Does Superfluid replace Ollama?

It presents itself as a serving daemon that can speak Ollama’s API while also fronting llama.cpp, MLX, and BaseRT. Whether it replaces Ollama for you depends on your runtime mix and whether you need its scheduler.

Are the 30× and 1.5 s TTFT numbers independent?

No. They are Base Compute’s published measurements on specified hardware and model settings.

Is a whitepaper available?

The GitHub README says a whitepaper is forthcoming and gives a temporary citation note until it is published.

THE TAKEAWAY

What to remember

Superfluid is Base Compute’s 8 October local multi-agent serving bet—evaluate scheduler fairness and API compatibility on your harness before you consolidate model servers.

Sources & further reading

  1. Introducing Superfluid: the local-AI native LLM server for multi-agent inference ↗
  2. basecompute/superfluid ↗
  3. Superfluid documentation site ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories