Menu

Ollaya · Run decision models locally.
📰
618

Ollaya · Run decision models locally.

Hacker News·14 days ago
#aJVe4nBm
Reading 0:00
15s threshold

Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware. Download Browse models Real output: decider:2b answered in 178 ms on an RTX 4090. Fast Drop-in compatible Open models Your data stays yours Platforms Fast Decisions in milliseconds. A decision model answers in a single forward pass, with no token-by-token generation. On your own GPU, a five-question request to Laya takes about 10 ms, end to end through the HTTP API. Every model, one scale · median latency, lower is better laya:multilingual 8.1 ms laya:en 9.6 ms gliclass 14.7 ms nli 20.4 ms decider:0.8b 155 ms decider:2b 190 ms TypeSafe Jev hosted API 236–276 ms Ollaya: median of a five-question request through the HTTP API on an NVIDIA RTX 4090 (laya in fp16, the others in fp32). Jev: median request latency of the hosted API in third-party benchmarks ( AbdelStark/jev-benchmarks , nibzard/decision-model-benchmark ), which includes the network.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More