Announced 10 Oct 2026 · Sources checked
What shipped on 10 October?
The public announcement is the DEV Community article “I Distilled a 397B Model Into a 4B One for $1.60…,” published 10 October 2026 at 13:59 UTC for the Hacktoberfest Open-Source AI Challenge. The companion repository SoumyaEXE/trailtruth was created the same day (08:02 UTC) and last pushed at 13:55 UTC. LICENSE is MIT.
The product pitch is narrow on purpose: fetch every live National Park Service alert, classify each into go / caution / no_go trail decisions, and print one weekly card per watched park so hikers stop doom-scrolling alert feeds. That is a decision-model workflow, not a chat trail guide.
It sits beside other 10 October open decision releases such as WaterSheep’s Jev-compatible scorer and earlier hosted scorers like Microsoft-Decision-1, but TrailTruth’s domain is NPS alert JSON rather than general ticket routing.
How does the distillation pipeline work?
ml/README.md names the teacher as Qwen/Qwen3.5-397B-A17B (non-thinking) on Tinker and the student as Qwen/Qwen3.5-4B with LoRA rank 32. The teacher labeled all 627 alerts with valid JSON on the first pass, at about $0.95 according to the author’s spend table. Training is cookbook-style SFT on the last assistant message for 96 steps / 3 epochs.
Gold decisions start from NPS category labels mapped to go/caution/no_go. The author documents one adjustment: Park Closure alerts that only name facilities (visitor centers, parking lots) become caution rather than no_go—23 of 627 alerts. decision_acc_raw reports the unadjusted mapping. SFT targets use the teacher’s JSON with decision overridden to gold when they disagree (235 alerts).
The student is trained and served with a short production system prompt; the teacher used a longer rules block (~340 extra tokens). A base_guided ablation gives the 4B base that long prompt without LoRA weights. Splits are stratified ~80/20 by category with seed 42 (501 train / 126 test).

Which numbers are in eval.json?
Ai Lookout hashed ml/results/eval.json. It records generated_at 2026-10-10T07:42:58Z, test_size 126, teacher_source teacher, and per-model rows. Selected author figures:
The DEV post additionally states McNemar p = 0.015 for the LoRA versus base decision-accuracy lift and p = 0.007 for 28/33 versus 17/33 no-go recall. Those p-values are the author’s; we did not recompute them.
| Model | decision_acc | closure_f1 | $ / 1k alerts |
|---|---|---|---|
| Teacher 397B-A17B | 0.6508 | 0.9231 | 1.4982 |
| 4B base | 0.6429 | 0.8889 | 0.1282 |
| 4B + long prompt | 0.6746 | 0.8444 | 0.2409 |
| 4B + TrailTruth LoRA | 0.746 | 0.8842 | 0.1314 |
| Keyword heuristic | 0.5794 | 0.8738 | 0.0 |

What should readers not over-claim?
The README is unusually explicit that the LoRA beats the teacher on decision_acc partly because it was trained toward NPS-derived gold while the teacher was not. Against the base model with the teacher’s long prompt, the DEV post says the gain is not significant (p = 0.12). The teacher still leads closure_f1 and hazard self-agreement.
Famous alerts such as Chickamauga barred owl and Kīlauea closures are in the train split by the author’s note, so they are not evidence from eval.json’s test rows. Latency through Tinker’s hosted sampler is about 2.4 s p50 for every model; self-hosting the adapter is required for a real latency win.
Those caveats match the discipline we apply to other vendor benches, including HAL-X THX-01 ticket scores: report the harness, do not promote a single percentage as universal skill.

What else is in the open stack?
Beyond ML scripts, the repo includes a Temporal workflow (fetch → classify → compose → publish) with crash/resume demos, a FastAPI service, and a Next.js BoardUI dashboard for printable cards. download_adapter.py can export PEFT weights for vLLM, SGLang, or llama.cpp. NPS_API_KEY and TINKER_API_KEY are required for live runs; a heuristic fallback exists without keys.
End-to-end pipeline cost is stated as about $1.60 in spend.json references inside ml/README.md (teacher labeling ~$0.95, LoRA training ~$0.35–0.41, evaluation ~$0.25). Those are the author’s Tinker bills, not Ai Lookout receipts.
Who should try it?
Teams studying cheap structured distillation, outdoor-safety classifiers, or Temporal-backed agent jobs get a full reference implementation with splits and eval scripts. Park visitors should still read official NPS notices; TrailTruth is an assistive weekly card, not a ranger.
- Reproduce eval.py on the pinned split before trusting the 0.746 figure.
- Read derive_gold facility-closure adjustments before comparing to other alert taxonomies.
- Self-host the 4B adapter if latency—not just token cost—is the reason you distilled.
Common questions
Did the 4B model beat the 397B teacher fairly?
On the author’s decision_acc versus NPS-derived gold, yes (0.746 vs 0.651). The README states the student was trained toward that gold and the teacher was not, so the comparison is not a pure capability bake-off.
Are weights on Hugging Face?
The repository documents a Tinker sampler path and download_adapter.py for PEFT export. Ai Lookout did not find a separate public Hub adapter id in the README files we opened; start from the GitHub results paths.
Is this affiliated with the National Park Service?
No affiliation is claimed. Alerts are fetched from the public NPS API; labels and decisions are the author’s pipeline.
What to remember
TrailTruth is a complete, MIT-licensed 10 October case study: distill park-alert decisions into a 4B LoRA, keep eval.json honest about gold-label quirks, and wrap the job in a Temporal weekly card. Reproduce before you hike on the percentages.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





