cderinbogaz commited on
Commit
8654b9c
·
verified ·
1 Parent(s): ce95ef8

Add ONNX CPU builds (fp32 + int8 variants) with benchmark results

Browse files
README.md CHANGED
@@ -29,7 +29,7 @@ Innovations or TypeSafe.
29
  ## Quick start
30
 
31
  ```python
32
- import laya # pip install laya (Raya was built and tested with laya 0.3.7)
33
 
34
  raya = laya.Agent("TextCortex/raya", device="cuda") # or "mps" / "cpu"
35
 
@@ -52,6 +52,34 @@ Raya was trained on three routing questions: the minimal choice above, a detaile
52
  3-level difficulty score. Use one of those. Option order does not matter because options were shuffled in
53
  training. Raya serves through Laya's Jev-compatible HTTP server (`POST /v1/systemone`).
54
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55
  ## Benchmark: 3-tier routing on real multilingual prompts
56
 
57
  **Test set.** 563 first-turn prompts from WildChat-1M (shards never used for training), 14 languages,
 
29
  ## Quick start
30
 
31
  ```python
32
+ import laya # pip install laya (tested with laya 0.3.7 and 0.3.20; identical results)
33
 
34
  raya = laya.Agent("TextCortex/raya", device="cuda") # or "mps" / "cpu"
35
 
 
52
  3-level difficulty score. Use one of those. Option order does not matter because options were shuffled in
53
  training. Raya serves through Laya's Jev-compatible HTTP server (`POST /v1/systemone`).
54
 
55
+ ## CPU inference (ONNX)
56
+
57
+ No GPU? `onnx/` has ONNX Runtime builds of Raya that run through Laya's own `ONNXAgent`, which gives the same answer format as `laya.Agent`. Use a **512-token** input budget, which is what Raya was trained on.
58
+
59
+ ```python
60
+ from huggingface_hub import snapshot_download
61
+ from laya.onnx_agent import ONNXAgent # pip install laya onnxruntime
62
+
63
+ FILE = "onnx/raya.onnx" # see the table below
64
+ path = snapshot_download("TextCortex/raya", allow_patterns=["rl_agent_config.json", "tokenizer/*", "encoder/*", FILE])
65
+ raya = ONNXAgent(path, onnx_path=f"{path}/{FILE}")
66
+ raya.cfg["max_len"] = 512
67
+ print(raya.system_one({"prompt": "Schreibe eine professionelle E-Mail …"}, {"route": ROUTE})["answers"]["route"]) # ROUTE as above
68
+ ```
69
+
70
+ **Which file?** If you're not sure, use `raya.onnx`. At the same token budget it matches the PyTorch model exactly (0 of 563 choices changed, probabilities within 0.0001), on any CPU. On CPUs with VNNI int8 instructions (Intel Cascade Lake or Alder Lake and newer, AMD Zen 4 and newer), `raya-int8-blockwise.onnx` keeps accuracy and was about 10–15% faster on our Intel i9-13900. The other int8 files are faster still, but lose about 2 accuracy points.
71
+
72
+ | File | Size | Accuracy | Choices changed vs PyTorch | Intel i9-13900, VNNI (p50 / p95) | AMD EPYC 7502P, no VNNI (p50 / p95) |
73
+ |---|---|---|---|---|---|
74
+ | `raya.onnx` (fp32) | 1.23 GB | **81.2%** | 3 / 563 | 38 / 297 ms | 60 / 487 ms |
75
+ | `raya-int8-blockwise.onnx` | 0.89 GB | **81.5%** | 5 / 563 | **34 / 282 ms** | 127 / 975 ms |
76
+ | `raya-int8.onnx` | 0.92 GB | 79.0% | 53 / 563 | 24 / 200 ms | not recommended |
77
+ | `raya-int8-emb.onnx` | 0.35 GB | 79.4% | 47 / 563 | 24 / 201 ms | not recommended |
78
+ | `raya-int8-mixed.onnx` | 1.05 GB | 79.2% | 41 / 563 | 31 / 242 ms | not recommended |
79
+ | PyTorch reference (1,024-token budget) | — | 81.0% | — | 42 / 421 ms | 89 / 526 ms |
80
+
81
+ All rows use the same 563-prompt benchmark below, with the minimal routing question and single requests on 8 threads. The fp32 and block-wise rows include the switch from a 1,024- to a 512-token budget, which accounts for 3 of their changed choices. Per-tensor int8 (the last three files) overflows on CPUs without VNNI; on the EPYC it dropped to 54–80% depending on settings. 16 threads was slower than 8 on both CPUs. `onnx/SHA256SUMS` has checksums and `onnx/benchmark_results.jsonl` has the raw results.
82
+
83
  ## Benchmark: 3-tier routing on real multilingual prompts
84
 
85
  **Test set.** 563 first-turn prompts from WildChat-1M (shards never used for training), 14 languages,
onnx/SHA256SUMS ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ 98acb9f2bffb87fa2513cce596b6f279b1793683a0e4ac062853b64b3569fb0e raya-int8-blockwise.onnx
2
+ dd51b72f58843e7f8926513f912e20f0ffed417a5536856c1483067b1f03c836 raya-int8-emb.onnx
3
+ 67b88d62b07c17db797473118e44b5372755d24c7550b60463e47d3a3fedadda raya-int8-mixed.onnx
4
+ 4b16144bfbb371ccc98701f8c356893bbdc1f413c6f46f3b240a4ab96b1fecba raya-int8.onnx
5
+ 521c19623384a00a539d47bf79f1f114ef8dc7775ede9ece76fcf61dabb653ac raya.onnx
onnx/benchmark_results.jsonl ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 53.0, "p95_ms": 508.9}
2
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 49.4, "p95_ms": 400.8}
3
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 26.7, "p95_ms": 248.2}
4
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 27.0, "p95_ms": 253.9}
5
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 36.6, "p95_ms": 309.1}
6
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 4, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 42.7, "p95_ms": 379.1}
7
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 41.7, "p95_ms": 421.2}
8
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 37.8, "p95_ms": 296.8}
9
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 24.2, "p95_ms": 199.5}
10
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 24.0, "p95_ms": 201.2}
11
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 30.6, "p95_ms": 241.7}
12
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 8, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 33.7, "p95_ms": 281.9}
13
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 56.6, "p95_ms": 539.5}
14
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 43.3, "p95_ms": 367.9}
15
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7904, "changed_vs_pytorch": 53, "p50_ms": 34.3, "p95_ms": 232.8}
16
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-emb.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.794, "changed_vs_pytorch": 47, "p50_ms": 27.7, "p95_ms": 209.7}
17
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-mixed.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.7922, "changed_vs_pytorch": 41, "p50_ms": 34.6, "p95_ms": 252.7}
18
+ {"cpu": "Intel Core i9-13900 (AVX-VNNI)", "threads": 16, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8153, "changed_vs_pytorch": 5, "p50_ms": 35.8, "p95_ms": 277.7}
19
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 105.7, "p95_ms": 844.3, "note": "equivalent build of the same recipe, produced on that node"}
20
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 59.6, "p95_ms": 486.6, "note": "equivalent build of the same recipe, produced on that node"}
21
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 16, "file": "raya.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8117, "changed_vs_pytorch": 3, "p50_ms": 78.0, "p95_ms": 499.6, "note": "equivalent build of the same recipe, produced on that node"}
22
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8135, "changed_vs_pytorch": 4, "p50_ms": 148.3, "p95_ms": 1192.1, "note": "equivalent build of the same recipe, produced on that node"}
23
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "raya-int8-blockwise.onnx", "max_tokens": 512, "n": 563, "accuracy": 0.8135, "changed_vs_pytorch": 4, "p50_ms": 127.2, "p95_ms": 975.3, "note": "equivalent build of the same recipe, produced on that node"}
24
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 4, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 155.7, "p95_ms": 1333.2}
25
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 8, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 88.7, "p95_ms": 526.4}
26
+ {"cpu": "AMD EPYC 7502P (AVX2, no VNNI)", "threads": 16, "file": "pytorch", "max_tokens": 1024, "n": 563, "accuracy": 0.8099, "changed_vs_pytorch": 0, "p50_ms": 79.5, "p95_ms": 466.2}
onnx/raya-int8-blockwise.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:98acb9f2bffb87fa2513cce596b6f279b1793683a0e4ac062853b64b3569fb0e
3
+ size 933944851
onnx/raya-int8-emb.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dd51b72f58843e7f8926513f912e20f0ffed417a5536856c1483067b1f03c836
3
+ size 370123937
onnx/raya-int8-mixed.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:67b88d62b07c17db797473118e44b5372755d24c7550b60463e47d3a3fedadda
3
+ size 1097707917
onnx/raya-int8.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4b16144bfbb371ccc98701f8c356893bbdc1f413c6f46f3b240a4ab96b1fecba
3
+ size 959947611
onnx/raya.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:521c19623384a00a539d47bf79f1f114ef8dc7775ede9ece76fcf61dabb653ac
3
+ size 1290246243