Add ONNX CPU builds (fp32 + int8 variants) with benchmark results
Browse files- README.md +29 -1
- onnx/SHA256SUMS +5 -0
- onnx/benchmark_results.jsonl +26 -0
- onnx/raya-int8-blockwise.onnx +3 -0
- onnx/raya-int8-emb.onnx +3 -0
- onnx/raya-int8-mixed.onnx +3 -0
- onnx/raya-int8.onnx +3 -0
- onnx/raya.onnx +3 -0
README.md
CHANGED
|
@@ -29,7 +29,7 @@ Innovations or TypeSafe.
|
|
| 29 |
## Quick start
|
| 30 |
|
| 31 |
```python
|
| 32 |
-
import laya # pip install laya (
|
| 33 |
|
| 34 |
raya = laya.Agent("TextCortex/raya", device="cuda") # or "mps" / "cpu"
|
| 35 |
|
|
@@ -52,6 +52,34 @@ Raya was trained on three routing questions: the minimal choice above, a detaile
|
|
| 52 |
3-level difficulty score. Use one of those. Option order does not matter because options were shuffled in
|
| 53 |
training. Raya serves through Laya's Jev-compatible HTTP server (`POST /v1/systemone`).
|
| 54 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
## Benchmark: 3-tier routing on real multilingual prompts
|
| 56 |
|
| 57 |
**Test set.** 563 first-turn prompts from WildChat-1M (shards never used for training), 14 languages,
|
|
|
|
| 29 |
## Quick start
|
| 30 |
|
| 31 |
```python
|
| 32 |
+
import laya # pip install laya (tested with laya 0.3.7 and 0.3.20; identical results)
|
| 33 |
|
| 34 |
raya = laya.Agent("TextCortex/raya", device="cuda") # or "mps" / "cpu"
|
| 35 |
|
|
|
|
| 52 |
3-level difficulty score. Use one of those. Option order does not matter because options were shuffled in
|
| 53 |
training. Raya serves through Laya's Jev-compatible HTTP server (`POST /v1/systemone`).
|
| 54 |
|
| 55 |
+
## CPU inference (ONNX)
|
| 56 |
+
|
| 57 |
+
No GPU? `onnx/` has ONNX Runtime builds of Raya that run through Laya's own `ONNXAgent`, which gives the same answer format as `laya.Agent`. Use a **512-token** input budget, which is what Raya was trained on.
|
| 58 |
+
|
| 59 |
+
```python
|
| 60 |
+
from huggingface_hub import snapshot_download
|
| 61 |
+
from laya.onnx_agent import ONNXAgent # pip install laya onnxruntime
|
| 62 |
+
|
| 63 |
+
FILE = "onnx/raya.onnx" # see the table below
|
| 64 |
+
path = snapshot_download("TextCortex/raya", allow_patterns=["rl_agent_config.json", "tokenizer/*", "encoder/*", FILE])
|
| 65 |
+
raya = ONNXAgent(path, onnx_path=f"{path}/{FILE}")
|
| 66 |
+
raya.cfg["max_len"] = 512
|
| 67 |
+
print(raya.system_one({"prompt": "Schreibe eine professionelle E-Mail …"}, {"route": ROUTE})["answers"]["route"]) # ROUTE as above
|
| 68 |
+
```
|
| 69 |
+
|
| 70 |
+
**Which file?** If you're not sure, use `raya.onnx`. At the same token budget it matches the PyTorch model exactly (0 of 563 choices changed, probabilities within 0.0001), on any CPU. On CPUs with VNNI int8 instructions (Intel Cascade Lake or Alder Lake and newer, AMD Zen 4 and newer), `raya-int8-blockwise.onnx` keeps accuracy and was about 10–15% faster on our Intel i9-13900. The other int8 files are faster still, but lose about 2 accuracy points.
|
| 71 |
+
|
| 72 |
+
| File | Size | Accuracy | Choices changed vs PyTorch | Intel i9-13900, VNNI (p50 / p95) | AMD EPYC 7502P, no VNNI (p50 / p95) |
|
| 73 |
+
|---|---|---|---|---|---|
|
| 74 |
+
| `raya.onnx` (fp32) | 1.23 GB | **81.2%** | 3 / 563 | 38 / 297 ms | 60 / 487 ms |
|
| 75 |
+
| `raya-int8-blockwise.onnx` | 0.89 GB | **81.5%** | 5 / 563 | **34 / 282 ms** | 127 / 975 ms |
|
| 76 |
+
| `raya-int8.onnx` | 0.92 GB | 79.0% | 53 / 563 | 24 / 200 ms | not recommended |
|
| 77 |
+
| `raya-int8-emb.onnx` | 0.35 GB | 79.4% | 47 / 563 | 24 / 201 ms | not recommended |
|
| 78 |
+
| `raya-int8-mixed.onnx` | 1.05 GB | 79.2% | 41 / 563 | 31 / 242 ms | not recommended |
|
| 79 |
+
| PyTorch reference (1,024-token budget) | — | 81.0% | — | 42 / 421 ms | 89 / 526 ms |
|
| 80 |
+
|
| 81 |
+
All rows use the same 563-prompt benchmark below, with the minimal routing question and single requests on 8 threads. The fp32 and block-wise rows include the switch from a 1,024- to a 512-token budget, which accounts for 3 of their changed choices. Per-tensor int8 (the last three files) overflows on CPUs without VNNI; on the EPYC it dropped to 54–80% depending on settings. 16 threads was slower than 8 on both CPUs. `onnx/SHA256SUMS` has checksums and `onnx/benchmark_results.jsonl` has the raw results.
|
| 82 |
+
|
| 83 |
## Benchmark: 3-tier routing on real multilingual prompts
|
| 84 |
|
| 85 |
**Test set.** 563 first-turn prompts from WildChat-1M (shards never used for training), 14 languages,
|
onnx/SHA256SUMS
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
98acb9f2bffb87fa2513cce596b6f279b1793683a0e4ac062853b64b3569fb0e raya-int8-blockwise.onnx
|
| 2 |
+
dd51b72f58843e7f8926513f912e20f0ffed417a5536856c1483067b1f03c836 raya-int8-emb.onnx
|
| 3 |
+
67b88d62b07c17db797473118e44b5372755d24c7550b60463e47d3a3fedadda raya-int8-mixed.onnx
|
| 4 |
+
4b16144bfbb371ccc98701f8c356893bbdc1f413c6f46f3b240a4ab96b1fecba raya-int8.onnx
|
| 5 |
+
521c19623384a00a539d47bf79f1f114ef8dc7775ede9ece76fcf61dabb653ac raya.onnx
|
onnx/benchmark_results.jsonl
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 53.0, "p95_ms": 508.9}
|
| 2 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 49.4, "p95_ms": 400.8}
|
| 3 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 26.7, "p95_ms": 248.2}
|
| 4 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 27.0, "p95_ms": 253.9}
|
| 5 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 36.6, "p95_ms": 309.1}
|
| 6 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 42.7, "p95_ms": 379.1}
|
| 7 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 41.7, "p95_ms": 421.2}
|
| 8 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 37.8, "p95_ms": 296.8}
|
| 9 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 24.2, "p95_ms": 199.5}
|
| 10 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 24.0, "p95_ms": 201.2}
|
| 11 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 30.6, "p95_ms": 241.7}
|
| 12 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 33.7, "p95_ms": 281.9}
|
| 13 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 56.6, "p95_ms": 539.5}
|
| 14 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 43.3, "p95_ms": 367.9}
|
| 15 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 34.3, "p95_ms": 232.8}
|
| 16 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 27.7, "p95_ms": 209.7}
|
| 17 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 34.6, "p95_ms": 252.7}
|
| 18 |
+
{"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 35.8, "p95_ms": 277.7}
|
| 19 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 105.7, "p95_ms": 844.3, "note": "equivalent build of the same recipe, produced on that node"}
|
| 20 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 59.6, "p95_ms": 486.6, "note": "equivalent build of the same recipe, produced on that node"}
|
| 21 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 16, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 78.0, "p95_ms": 499.6, "note": "equivalent build of the same recipe, produced on that node"}
|
| 22 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8135, "changed_vs_pytorch": 4, "p50_ms": 148.3, "p95_ms": 1192.1, "note": "equivalent build of the same recipe, produced on that node"}
|
| 23 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8135, "changed_vs_pytorch": 4, "p50_ms": 127.2, "p95_ms": 975.3, "note": "equivalent build of the same recipe, produced on that node"}
|
| 24 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 155.7, "p95_ms": 1333.2}
|
| 25 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 88.7, "p95_ms": 526.4}
|
| 26 |
+
{"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 16, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 79.5, "p95_ms": 466.2}
|
onnx/raya-int8-blockwise.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:98acb9f2bffb87fa2513cce596b6f279b1793683a0e4ac062853b64b3569fb0e
|
| 3 |
+
size 933944851
|
onnx/raya-int8-emb.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:dd51b72f58843e7f8926513f912e20f0ffed417a5536856c1483067b1f03c836
|
| 3 |
+
size 370123937
|
onnx/raya-int8-mixed.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:67b88d62b07c17db797473118e44b5372755d24c7550b60463e47d3a3fedadda
|
| 3 |
+
size 1097707917
|
onnx/raya-int8.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4b16144bfbb371ccc98701f8c356893bbdc1f413c6f46f3b240a4ab96b1fecba
|
| 3 |
+
size 959947611
|
onnx/raya.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:521c19623384a00a539d47bf79f1f114ef8dc7775ede9ece76fcf61dabb653ac
|
| 3 |
+
size 1290246243
|