Laya-CoreML: A Local AI That Decides in ~5 Milliseconds on the Neural Engine

By Prahlad Menon 4 min read

This is now the fourth post in a thread I didn’t plan to write: Kev (an open decision model), Laya (fast, local, calibrated), Keel (a decision model doing a real job inside a coding agent), and now Laya-CoreML — the same Laya model ported to run on Apple’s Neural Engine, deciding in about 5 milliseconds, fully offline.

If Laya was “fast, local decisions,” this is “fast, local decisions on the chip in your Mac — and here’s the honest power bill.”

What it is

Laya-CoreML takes the open-weight Laya typed-decision model — the one that answers choice / ordinal score / yes-no (“noul”) questions with calibrated probabilities in a single forward pass, no generated tokens — and ships validated ports to Core ML targeting the Apple Neural Engine (ANE).

The pitch: run a real AI decision agent on your Mac that answers in ~5ms, uses a fraction of the energy of standard setups, and never touches the network.

The Snake demo is the hook

The repo’s centerpiece is delightful: a real Laya model plays terminal Snake, itself, live. The screen shows, in real time:

  • the model’s move probabilities (up / down / left / right)
  • score, length, latency
  • a visible cycle-safety layer and safety interventions

Across three uncapped 600-step episodes it sustained ~49–50 decisions/sec with zero deaths (two safety interventions total). It’s the clearest visualization I’ve seen of a System 1 model doing what it’s built for: making a bounded, typed choice, many times a second, with its confidence exposed.

pip install 'laya-coreml[demo]'
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
laya-coreml-snake --model ./models/snake

Download once, then play offline — no PyTorch, Transformers, or MLX needed for inference.

The benchmarks — and why the honesty matters

Here’s what makes this repo trustworthy. It does not oversell. The measured M3 Max numbers, one 91-token question:

MetricCompiled MLX FP16Core ML ANE FP16Core ML ANE W8
P50 / P95 latency6.94 / 7.39 ms4.98 / 5.31 ms4.88 / 5.23 ms
Mean system power61.39 W30.75 W27.39 W
Energy / decision0.4288 J0.1540 J0.1344 J
Speed gain1×1.39×1.42×
Energy gain1×2.78×3.19×

And then, in plain text: “the requested 10× improvement was not achieved.”

That line is worth pausing on. A lot of “local AI destroys the cloud” posts quietly bury the caveats. This one leads with them. The speed win over an already-fast MLX baseline is a modest ~1.4×. The real story is energy: ~2.8–3.2× less joules per decision, from moving work onto the Neural Engine instead of GPU. For anything battery-powered or always-on — an agent making thousands of routing/guardrail decisions — that’s the number that matters.

The calibration bug they caught (this is the good stuff)

Buried in the README is a detail that tells you these are serious people. Laya’s decisions are only useful if the probabilities are honest — that’s the whole point of the proper-scoring-rule training I described for Laya. But calibration temperatures can go wrong:

The shipped choice:11+ bucket temperature is 0.1006, which would sharpen logits ~10× and report a coin flip as near-certainty.

So they clamp fitted temperatures to [0.5, 5.0] before use, keep the raw values available for inspection (agent.temperature_raw), and emit a RuntimeWarning naming every clamped bucket at load. That’s exactly the kind of guardrail that separates “cool demo” from “thing you’d put in production” — and it rhymes with the calibration-temperature-fitting step in Laya’s own fine-tuning loop.

The honest limits

Also refreshingly stated up front:

  • ANE bundle: 96-token total limit (question + options + state). Longer inputs error out — use aac6fef/laya-multilingual-coreml for the 1024-token general model.
  • Requires Apple Silicon, macOS 15+, Python 3.11–3.13.
  • Snake’s decision-rate numbers include rendering serialization; single-question latencies exclude load/warmup. They’re careful to say which is which.

Why this closes the loop

Four posts in, the arc is clear:

  • Kev — you can build an open decision model.
  • Laya — you can make it fast, local, and calibrated.
  • Keel — you can put it to work as a permission-checked router.
  • Laya-CoreML — you can run it on the Neural Engine at ~5ms and ~⅓ the energy, offline, with the calibration bugs actually caught.

This is the same lesson as our RF-SRC memory-optimization work: small, purpose-built models that emit calibrated probabilities and run on hardware you already own beat a giant generalist for a huge class of structured decisions. Laya-CoreML just proves it down to the joule — and, notably, proves it without inflating the claim.

No cloud, no lag, no hype. Just a tiny model making honest decisions on your chip.

👉 github.com/mizorewww/laya-coreml