Independent measurement · LiteRT-LM web · @litert-lm/core 0.17.1

Kernel 0112

One of the 153 GPU programs Google's LiteRT-LM web engine compiles, and one of the 65 it runs for every token it generates (a token is a word or part of one). It runs 40 times per token and carries 13.7% of the GPU time in a per-kernel profile on the reference machine (one 7-token run). The page runs it as shipped, and again with the same work spread over eight times as many threads.

Checking WebGPU…

No model weights and no dataset are downloaded; the page loads its fonts from Google Fonts. Nothing you run leaves your browser. The page makes its own random matrices; a run takes a few seconds.

Reference runApple M2 Max (30-core GPU) · Chrome 146 (inferred) · 2026-09-22 · node harness · one timestamp per run · 576 cold and 24 hot samples · 24 matrices
1.65×faster when each dot product is split 32 ways instead of 4, same 64-thread workgroups, and the same answer to within 16-bit float rounding; the same harness reads 0.84× on Chrome 131
original kernelchanged kernelsame, weights already in cache (hot)

Mean of two harness runs' medians, each variant timed forward and again in reverse order after a discarded warm-up, with the browser's timestamp rounding switched off. At the original's thread count the bigger-workgroup control reads the same as the original (60.6 against 60.9 µs), so workgroup size alone does not move the original's number; at 49,152 threads the repository README records a small secondary workgroup-size effect. Chrome 146 is inferred for these two records, which carry only the adapter string apple/metal-3 and no browser field; a later single-dispatch record of the same kernel on the same machine, out/microbench-0112-2026-09-24T07-36-19-655Z.json, gives 1.638× with its browser recorded as Chrome/146.0.7680.153. The node harness behind these numbers times one dispatch at a time; on Chrome for Testing 131 on the same machine it reports the reverse in a committed record, 78.1 µs for the original against 93.1 µs for the 32-way split, while this page's batched method still favours the split there, which is a console observation with no committed record. On the reference machine this page's own method lands above or below the harness figure from session to session, and on a contended or battery-powered GPU it reads lower, below 1.0× in one reading on a throttled machine on battery, also a console observation with no committed record, with every variant still matching the reference; the repository README records the range.

What is being measured

Generating one token with a language model means multiplying the model's stored numbers, its weights, by the current input, layer after layer. This kernel does one such multiplication: a matrix of 12,288 × 1,536 weights, each stored in 2 bits, times a vector of 1,536 inputs. That is 4.5 MiB of weights the kernel reads each time it runs, and it runs 40 times per token. Byte counts on this page are MiB (220 bytes) and GB/s is decimal; figures quoted from Google's model card are the card's own MB.

Because the work is dominated by reading those 4.5 MiB, the yardstick used here is effective weight-streaming bandwidth: the bytes of weights the kernel reads per unit of kernel time, shown as GB/s, which is not a measurement of memory traffic, since a byte served from cache counts the same. For scale, the Apple M2 Max is rated at 400 GB/s. The fastest kernel measured in this project, a 4-bit one in a single run, reaches 195 GB/s, about half of that. The original kernel reaches about 77 GB/s here (60.9 µs, the two-run mean cold median), and 80 GB/s inside the running model (one 7-token profile run).

To read 4.5 MiB the kernel launches 12,288 threads: four for each group of four outputs, each thread computing those four outputs over a quarter of the 1,536 inputs. That is few threads for a chip with thousands of execution lanes, which typically hides memory latency by keeping many reads in flight. The variants below change one thing at a time: the size of the thread teams (workgroups) the engine uses for this shape, and how many threads share each dot product.

Matrix of 12,288 by 1,536 two-bit weights multiplied by a 1,536 input vector gives 12,288 outputs; below, the original 16 by 4 workgroup, a 64 by 4 control with the same total threads, and the 2 by 32 workgroup that dispatches eight times as many threads.
Top: the multiplication this kernel performs. Bottom: the work is split into workgroups, small teams of GPU threads. Left, the original: teams of 16 × 4, four threads per output slice, 12,288 threads in total. Middle, the control: teams of 64 × 4, four times bigger, still 12,288 threads; it runs no faster. Right, the change: teams of 2 × 32, still 64 threads each, but 32 threads per output slice and 98,304 threads in total; it runs about 1.65× faster on the reference machine in Chrome 146 (inferred from the adapter string), and slower in a single-dispatch run on Chrome 131. The 2 × 32 shape is the one Google's own kernel 0113 already uses on the transposed shape (12,288 → 1,536): the engine uses 16 × 4 for this shape and 2 × 32 for the transposed one, so what separates 0112 from the faster configuration on this adapter in Chrome 146 is the configuration chosen for its shape; how the engine chooses is not established.

The four variants

VariantWorkgroupWorkgroupsThreadsWhat changed in the source
Original16 × 4 = 6419212,288Nothing, apart from trailing spaces removed from three lines. Google's kernel as captured.
Bigger workgroup, same threads64 × 4 = 2564812,288The workgroup-size constant and the matching shared-memory array size. A control: it isolates workgroup size from thread count.
16-way split4 × 16 = 6476849,152Each output's dot product is split across 16 threads instead of 4: the loop stride, the shared-memory size and the reduction depth follow from that one number.
32-way split2 × 32 = 641,53698,304The same, 32 ways.

How it is measured

  1. Random weights, made here. The page generates 24 different 4.5 MiB matrices of random bytes in your browser (108 MiB; if the GPU cannot hold them it retries with 8 and the results panel says so). No model weights are downloaded and no results leave your browser; the page's fonts are fetched from Google's font servers. The real model's weights are not needed to measure speed: a memory-bound kernel should take the same time whatever the values, and the bytes are unpatterned because some GPUs compress patterned textures.
  2. A reference answer on the CPU. For the first matrix, plain JavaScript computes the 12,288 outputs. Every GPU variant must match it to within 16-bit float rounding (0.05 on outputs of size about 2) or it is marked wrong, and a wrong variant's speed does not count.
  3. Cold and hot. "Cold" rotates through all 24 matrices (108 MiB, far more than the reference chip's caches hold), which makes cache reuse from one run to the next unlikely, as in the real model where 40 different matrices stream past. "Hot" reuses one matrix, small enough to stay in the chip's cache. The gap between them shows how much a variant is limited by memory rather than by arithmetic.
  4. GPU timestamps, in batches. The GPU records the time before and after each batch of 48 back-to-back runs. Browsers may round these timestamps to 100 µs for privacy; batching keeps that rounding between about 3% of the slowest bar and 6% of the fastest. Where timestamps are unavailable the page times batches of 480 runs from JavaScript instead and says so in the results panel.
  5. Interleaved order after a warm-up. The GPU is slow for whatever is measured first. A discarded warm-up of the original kernel (16 batches) runs before any timing, then every variant is timed in order and again in reverse, and the two positions' samples are pooled: eight batches per variant per position, median reported. Other tabs, thermal state and power mode all move the numbers; run twice if the first looks odd.

What the numbers mean

The kernel

The source below is what @litert-lm/core 0.17.1 emitted at run time when it created this shader (the WebGPU API receives every shader as plain text). The engine that emits it is published by Google under the Apache License 2.0; the text is reproduced here for debugging under the same terms, with attribution and with trailing spaces removed from three lines. The variants above are produced from it at run time by the substitutions described in the table. What the source shows: weights kept in a texture, 2-bit fields unpacked with floating-point floor arithmetic, and a shared-memory reduction with barriers, no subgroup operations.

Kernel 0112 source, 3,519 bytes of WGSL as shown (3,525 as captured)

  

What the kernel costs in a whole token

Across a whole generated token (one traced run and one 7-token profile run), nine weight-reading kernels move about 741 MiB in 277 runs. On the reference machine that is an effective weight-streaming bandwidth of 81 GB/s as an overhead-adjusted estimate (78 GB/s raw; the per-kernel measurement moves each dispatch into its own timestamped pass, which adds about 1.5 µs per pass, and the adjusted figure subtracts it), a fifth of the 400 GB/s the chip is rated for. Four of them, this one included, take 48.5% of the GPU time, and each launches only 6,144 to 12,288 threads for 4.5 MiB.

Deeper K splits were then applied inside the running model on the reference machine to this kernel and the three other quantized matrix-vector kernels, as the engine creates them (0112 and 0105 to 2 × 32, 0099 to 2 × 128, 0113 from 2 × 32 to 1 × 256): decoding went from 57 to 71 tokens per second, 1.20× as the median of the per-repetition pairings (1.18 to 1.33×) and 1.25× as the ratio of the condition medians (17.65 → 14.12 ms per token), over five repetitions of each of six patch conditions, 30 runs, with no GPU errors. The per-repetition pairing is the aggregate the shuffled design supports: the condition order is reshuffled inside every repetition, and pairing within a repetition uses that blocking while the ratio of condition medians does not. Google's model card reports 73 decode tokens per second for this same web bundle on a newer M4 Max, measured over 256 decode tokens after a 1,024-token prefill with a context length of 2,048 tokens, while the runs here time the last 67 to 69 tokens of a 76- to 78-token reply to a one-sentence prompt, so the two differ in context length as well as hardware and are not comparable (its 160.2 native figure is for a different, larger build). The generated text was identical apart from one word, which is consistent with a near-tie where a 32-way sum rounds differently from a 4-way one. In an earlier sweep (out/patch.json), 0112 patched alone left the text unchanged; in the 30-run sweep every patched set changes that one word, including the one that leaves 0105 unpatched. The in-model runs were made on this one GPU, all 30 in Chrome 146.

Four of those conditions leave one kernel unpatched, which measures each kernel's marginal contribution with the other three already patched: by the ratio of condition medians, without the 4-bit kernel 0099 it drops from 1.25× to 1.13×, and without the 4-bit 0105 to 1.15×, while without this page's kernel 0112 it is still 1.23× and without 0113 1.22×. So within the four-kernel patch the two 4-bit kernels make the largest marginal contributions, and 0112 and 0113 smaller ones. Those last two cannot be ordered against each other: dropping 0112 costs less than dropping 0113 under the ratio of condition medians and more under the per-repetition pairing, so which of them ranks last follows from the choice of aggregate. The 1.65× above is the isolated kernel's speedup in Chrome 146 (inferred from the adapter string; 0.84× on Chrome 131); over its 40 dispatches per token that predicts about 0.96 ms per token saved. Patched alone, in an earlier two-repetition sweep with a fixed condition order (out/patch.json, whose unpatched baseline of 18.4 to 18.9 ms makes it not directly comparable with the 30-run sweep), 0112 saved 0.690 and 0.885 ms per token, near that prediction. Leaving it out of the four-kernel patch costs 0.240 ms per token by the condition medians (per repetition −0.235 to 1.960 ms), so with the other three already patched its marginal contribution is smaller than its stand-alone saving in that earlier sweep; whether that reflects overlapping gains or the difference between the two sweeps is not established. 0099, a 4-bit kernel that also serves the 1.5 MiB q and o projections, moves the most in four of five repetitions.

This page's reference numbers are from an Apple M2 Max; the repository README holds the Colab T4 run. Copy yours with the button above.