Objective checks (regex / exact tokens), not writing quality. A 135M model is allowed
to fail — that is the measurement. Pick models, then run. Estimate uses your
last tok/s if we have one.
Speed (tokens/s, sustained decode, suite wall) and accuracy (pass rate on objective tests) from runs
in this browser. Numbers stay on this machine. Charts use the latest suite per model.
Write a benchmark in JavaScript
The editor is eval()’d in this origin, then each check runs on the
model’s decoded text.