Typed questions in, calibrated answers out, in milliseconds.
One binary, models pulled by name and a TypeSafe-compatible API, the way Ollama runs LLMs.
Website · Models · Results · Docs · Releases · Hugging Face
curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "I was charged twice this month and want a refund."15 model families, from millisecond encoders that run on a CPU (laya, nli, gliclass, von) to
decoders that take a GPU (winnow, kev, decider, nimble, jeb, jeeves, cygnet and more).
Their accuracy, calibration and speed on our own GPUs and CPUs, with the raw data, are at
ollaya.dev/results.
Weights always come from their authors' own Hugging Face repositories, pinned to a commit and verified by sha256. Ollaya never re-hosts them.
Created and maintained by Mert Cobanov (@mertcobanov). Apache-2.0. Ollaya is an independent project, not affiliated with Ollama or TypeSafe.