Laya on WebGPU
Laya is an open-source "System One" decision model: give it a text (a support ticket, an email, a chat message) plus typed questions (choice / score / yes-no) and it returns calibrated probabilities in a single forward pass. No text generation, nothing to parse.
This page runs the multilingual checkpoint (322M params, 100+ languages) entirely in your browser with ONNX Runtime Web. The encoder runs on WebGPU (falls back to WebAssembly), so your text never leaves your machine.
How to use: click "加载模型 / Load model" (one-time ~1.3 GB download from Hugging Face, cached afterwards), pick a preset or edit the questions JSON, then click "决策 / Decide". "测速 x20" measures warm latency. On an Apple M4 Pro: ~150 ms per 3-question decision on WebGPU vs ~800 ms on WASM.
Requirements: a desktop browser with WebGPU (Chrome / Edge 113+). Browsers without WebGPU fall back to the much slower WASM path. Needs several GB of free memory.
Credits: model by ConvAI Innovations / Nandakishor M, github.com/NandhaKishorM/laya (Apache-2.0). Weights mirror: huggingface.co/Steven10429/laya-multilingual-webgpu. Base checkpoints are close to chance zero-shot on hard benchmarks; fine-tune for real use.
| Published | 2 days ago |
| Status | Released |
| Category | Tool |
| Platforms | HTML5 |
| Author | StevenLi-phoenix-work |
| AI Disclosure | AI Assisted |
Leave a comment
Log in with itch.io to leave a comment.