Skip to content

SLM Lab

POST /api/v1/chat/completionsliquid/lfm-2.5-2.6b

Load a 135M-parameter model into your browser over WebGPU and watch it answer on your own hardware. When the local model is out of its depth, escalate the identical prompt to AnyRouter and compare throughput, latency, and quality side by side.

SmolLM2 135M Instruct
135M params · Apache-2.0 · runs in this browser on WebGPU
114 MiB
Repo
HuggingFaceTB/SmolLM2-135M-Instruct
Quantization
q4f16
Weights
huggingface.co
Cache
Browser Cache API
Ask the on-device model
Runs entirely in this browser. Nothing is sent anywhere until you escalate.
0 / 400 characters

First run downloads about 114 MiB from huggingface.co and caches it here.

Same prompt, both answers
Local throughput versus the AnyRouter API, on one page.
Local run

Local answer appears here

Run it on your GPU, then escalate to liquid/lfm-2.5-2.6b to see the difference.