CPU latency numbers for laya-multilingual β€” the model card only has T4

#12
by htqmike - opened

Edit (Sep 27): two corrections. The "state length" column used to show tokens per question (state plus the question's own head): 71 / 261 / 927. The actual states are 37 / 227 / 893 tokens; the latencies are unchanged. And the 1392.5 ms_per_case I called consistent with these numbers is for the English checkpoint, not multilingual. Details in the thread below.

Hi β€” I spent a few days measuring laya-multilingual on CPU and wanted to share the numbers, since the model card only publishes the Tesla T4 table and a 322M non-autoregressive model is exactly the kind of thing people will try to run without a GPU.

Full harness and raw JSON: https://github.com/taiheqi718-art/laya-cpu-benchmark (Apache-2.0, same as the weights)

Setup: Intel i9-14900HX (24 cores / 32 threads), 32 GB, Windows 11, laya 0.3.5, torch 2.14.0+cpu, FP32, Chinese input, device="cpu".

Steady-state latency:

state 1 question 10 questions 50 questions
37 tokens 50 ms 401 ms 4,105 ms
227 tokens 307 ms 2,150 ms 9,103 ms
893 tokens 1,315 ms 10,199 ms 40,131 ms

Each question adds its own 33–37-token head (instruction, options, special tokens) on top of the state.

Against the published 32.8 / 72.3 / 337 ms, at a 227-token state that is 9.4x / 30x / 27x.

Three things came out of this that might be worth folding into the card:

1. The repo's own CPU numbers exist, they are just invisible.

research/results/cpu_51_language_sweep.json records CPU timings at 4 threads, but none of them is in the model card, and "case" is only definable by reading research/scripts/bench_local.py. The one per-case figure, 1392.5 ms_per_case, is for the English checkpoint (part B, 400 cases); for multilingual there is only wall time per 100-case language in the MASSIVE sweep. Anyone reading the card sees 32.8 ms and nothing else. A one-line CPU row would save people a lot of surprise.

2. The batching economics do not survive the move to CPU.

On T4, 10 questions cost 2.2x a single question β€” near-free batching, and a genuinely attractive property. On CPU, cost per question is flat: at a 227-token state it is 307 ms/q at n=1 and 182 ms/q at n=50. Batching buys ~20-30% from 1 to 10 questions and nothing after that.

That follows from the design β€” agent.system_one calls build_sequence once per question, so the state is re-encoded per question and compute scales as questions x sequence_length. A GPU at batch 1 is mostly idle so batching fills it; a CPU is already saturated. Worth a sentence in the docs, because "decompose your judgement into ten cheap sub-questions" is good advice on GPU and bad advice on CPU.

3. Thread scaling saturates around 8.

threads 1 2 4 8 16 24
latency 6,606 4,026 2,159 1,701 2,031 1,675 ms

24 cores buy 3.9x, and essentially all of it arrives by 8 threads. For anyone deploying, three 8-thread processes will beat one 24-thread process by a wide margin.

A methodology note that cost me two discarded sweeps: on a laptop, a short benchmark measures turbo boost rather than the machine. Running one config continuously from idle, the median goes 1,135 ms over the first 15 s and settles at 2,654 ms after about 90 s β€” 2.34x, over a window longer than most benchmark runs take in total. It is not thermal throttling (a 180 s sustained run drifted -0.9%); it is package power budget, and burst capacity takes roughly four minutes of idle to come back. So all the numbers above are steady-state medians, and the sweeps use randomized order plus repeated control configs to detect drift. Two of my sweeps produced clean-looking monotonic tables that were entirely artifacts, and only the controls caught them.

Everything is reproducible from the repo, and the results JSON records CPU model, thread count and library versions so contributed runs are self-describing. Mac/MPS, Ryzen and ARM are all missing if anyone wants to fill them in.

One question while I have your attention: the package hard-codes FP32 on CPU and disables autocast there (agent.py: elif self.device.type in ("cpu", "mps"): self.dtype = torch.float32). Community INT8 ONNX exports exist at ~325 MB. Is a supported quantized CPU path something you have looked at? That is the single change that would most move these numbers, and it seems like the difference between "runs on CPU" and "deployable on CPU".

Thanks for open-sourcing the weights β€” none of this would be measurable otherwise.

Useful numbers β€” the CPU side of this model is exactly what the card is missing, and the shape of
your scaling is the interesting part.

One thing to add to the analysis rather than compete with the table, since our boxes differ: the
state is not the only thing being processed, and past a point it is not what dominates.
Every
question costs up to head_max_len tokens before any state is read β€” the instruction text, plus
one [MASK] slot and up to 48 tokens per option. So your 50-question column is paying for 50
option blocks, and the state only gets max_len - head_max_len and whatever is left after them.

Two consequences that are easy to hit on CPU:

  1. Questions are not free even with a short state. Your 71-token row already shows it:
    1 question 50 ms β†’ 50 questions 4,105 ms is roughly linear in the question count, not in tokens.
  2. Raising max_len past the real state length buys nothing and costs tokens. I ran the ten
    hardest MASSIVE languages at max_len 1024 / 768 / 512 / 384 / 256 and got bit-identical
    accuracy at every one
    , because the utterances never reach the cap. On CPU the practical recipe
    is to size max_len to your actual states and leave head_max_len alone.

For reference on a different box (Intel Model 198, torch 2.14.0+cpu, threads=4): in a 51-language
sweep the English checkpoint at max_len=512 averaged ~28 s per 100-case language, and
multilingual at max_len=1024 ~12 s. The multilingual model is the smaller encoder despite the
larger window, so those two rows are not comparable to each other, let alone to yours β€” the harness
is a whole-language loop rather than your per-call one.

One more that cost me time: the two checkpoints do not share a max_len. laya ships 512 and
laya-multilingual 1024, so a benchmark that holds the token budget fixed across both is measuring
two different state windows.

β€’
This comment has been hidden

Thanks β€” the head point is right, and it caught a labeling error in my table.

The "state length" column was usage.input_tokens at n=1, which counts the state plus the question's own head. With the questions in my harness that head is 33–37 tokens (instruction, a [MASK] and the text for each option, [CLS]/[SEP]), so the actual states were 37 / 227 / 893 tokens, not 71 / 261 / 927. The latencies don't change, only the labels, and I've corrected the post and the repo README. It also sharpens your first point: at the 37-token state the head is almost half of every sequence, so even a nearly empty state doesn't make questions cheap.

Your note that the two checkpoints aren't comparable made me re-check the upstream figure I quoted, and I had the wrong row. The 1392.5 ms_per_case in cpu_51_language_sweep.json comes from part B, and part B in that file only has the English checkpoint (557 s over 400 cases). For multilingual the file only has part A, the MASSIVE sweep: 7.7–10.3 s per 100-case language at 4 threads. So "consistent with what I measure" was a comparison against the wrong model. That's corrected as well. Everything in my table is laya-multilingual (1024 / 256), so the 512-vs-1024 split doesn't enter it.

On max_len, though, I don't think a larger value costs anything. build_sequence slices the state to max_len - len(head) - 1, and collate_items pads to the longest row in the call, not to max_len. When nothing is truncated, the input ids are identical at every max_len, and so is the compute. Even the CUDA fast path only rounds the length up to a 16/64 bucket, capped at max_len. That's also why your MASSIVE runs came out bit-identical: nothing reached any of the caps, so all five settings saw the same inputs. I checked this against 0.3.5 (what I measured) and current main (0.3.20). The only way a smaller max_len saves time is by truncating the state, which can change the answer. If you did see a latency difference between those settings on untruncated inputs, I'd like to know which path it was on.

Sign up or log in to comment