Performance
Every number here comes from a script you can run against the npm package. Nothing is extrapolated.
npm install @kanunilabs/datagrid-core
curl -O https://kanunilabs.com/benchmarks/datagrid-pipeline.bench.mjs
node datagrid-pipeline.bench.mjs 1000000
1,000,000 rows × 9 columns, Node 24.14.1, Intel Core i7-11390H (4 cores, laptop). Each figure is the median of three runs; the small numbers are the fastest and slowest run. "First query" is the first timing of each run, which pays for the per-column caches described under Collator ranks and Search elimination. Your machine will differ. The benchmarks page has the chart, the method and what these numbers leave out.
1,000,000 rows
| Operation | Row-based | Columnar | Columnar, first query | Ratio |
|---|---|---|---|---|
| Sort, number column | 226 ms 208 ms – 309 ms | 101 ms 92 ms – 101 ms | 101 ms 68 ms – 130 ms | 2.2× |
| Sort, text column | 2.2 s 1.9 s – 2.3 s | 90 ms 80 ms – 106 ms | 1.8 s 1.6 s – 1.9 s | 24.2× |
| Sort, date column | 351 ms 344 ms – 355 ms | 92 ms 86 ms – 96 ms | 96 ms 69 ms – 102 ms | 3.8× |
| Sort, two columns | 433 ms 433 ms – 433 ms | 208 ms 202 ms – 218 ms | 245 ms 237 ms – 248 ms | 2.1× |
| Filter, equality | 159 ms 159 ms – 186 ms | 18 ms 17 ms – 19 ms | 22 ms 22 ms – 23 ms | 8.9× |
| Filter, 3 clauses (AND) | 206 ms 188 ms – 267 ms | 60 ms 55 ms – 63 ms | 82 ms 73 ms – 106 ms | 3.4× |
| Search, all columns | 737 ms 714 ms – 737 ms | 65 ms 50 ms – 69 ms | 413 ms 366 ms – 438 ms | 11.3× |
| Filter + sort | 205 ms 193 ms – 206 ms | 40 ms 38 ms – 44 ms | 41 ms 40 ms – 50 ms | 5.1× |
100,000 rows
| Operation | Row-based | Columnar | Columnar, first query | Ratio |
|---|---|---|---|---|
| Sort, number column | 10 ms 9.4 ms – 11 ms | 5.7 ms 5.5 ms – 5.8 ms | 7.9 ms 7.0 ms – 9.5 ms | 1.8× |
| Sort, text column | 179 ms 174 ms – 184 ms | 5.2 ms 5.0 ms – 5.5 ms | 147 ms 134 ms – 150 ms | 34.5× |
| Sort, date column | 20 ms 20 ms – 21 ms | 4.9 ms 4.5 ms – 5.0 ms | 5.0 ms 4.5 ms – 5.0 ms | 4.1× |
| Sort, two columns | 21 ms 20 ms – 22 ms | 11 ms 10 ms – 12 ms | 11 ms 10 ms – 12 ms | 1.9× |
| Filter, equality | 16 ms 16 ms – 18 ms | 1.8 ms 1.8 ms – 1.9 ms | 3.6 ms 3.5 ms – 3.8 ms | 9.1× |
| Filter, 3 clauses (AND) | 20 ms 20 ms – 20 ms | 8.2 ms 7.4 ms – 8.9 ms | 9.5 ms 7.4 ms – 11 ms | 2.5× |
| Search, all columns | 80 ms 70 ms – 82 ms | 5.9 ms 4.9 ms – 6.1 ms | 17 ms 17 ms – 18 ms | 13.5× |
| Filter + sort | 29 ms 24 ms – 31 ms | 9.5 ms 9.0 ms – 11 ms | 21 ms 20 ms – 23 ms | 3.0× |
Where the speed comes from
Columnar encoding. Columns become typed arrays — Float64Array for
numbers, dates and booleans; a string dictionary plus Int32Array codes for
text. Row objects are never copied to the worker; only the buffers are
transferred, which is a zero-copy handover.
Dictionary evaluation. A contains or equality filter on a text column is
evaluated against the column's distinct values, then reduced to an integer
scan. Work scales with cardinality, not row count.
Collator ranks. Text sorting compares precomputed integer ranks instead of
calling Intl.Collator inside the comparator. Most of the gap on the text sort
comes from this. The ranks are built by sorting the column's distinct values
once, and kept until the data changes, so the first sort on a column with many
distinct values is much slower than the ones after it.
Radix sort. Numeric sorting uses an LSD radix sort over an order-preserving 64-bit key — eight linear passes with no comparator callbacks at all. A comparator sort of a million rows costs roughly 20 million JS callback invocations; this costs none. Multi-key sorting runs one stable pass per key.
Search elimination. If the search term appears in no dictionary entry of a
column, that column is skipped entirely. And since String(someNumber) can only
contain a fixed set of characters, a term with letters never scans numeric
columns. In the benchmark, a search for "Marmara" scans one column of nine. The
lower-cased dictionaries this needs are built by the first search and reused
after that.
Quantized scrolling. The virtual window only changes when a row or column boundary is crossed, so most scroll events never reach React at all.
The worker
Above ~20,000 rows the pipeline moves off the main thread. Below that the main thread is genuinely faster — latency beats parallelism for small data.
The worker runs the same pipeline function the main thread would, so results are identical by construction rather than by convention. A suite of golden tests asserts that both paths produce byte-identical row orderings across every operator and data type.
If no Worker global exists (SSR, tests, restrictive environments) or the
worker fails to start, the grid degrades to the synchronous path instead of
breaking.
Rendering
Rows and columns are both virtualized. A 100,000-row grid keeps ~25 rows in the DOM; columns scrolled out of view are replaced by a single spacer element rather than being rendered.
Measured in a real browser at 100,000 rows: the scroll container reports a 2,800,032 px scroll height while the DOM holds 21–30 row elements, and the rendered block's transform lands exactly on the row grid.
The first-query encode
Before the first sort at large row counts, the grid encodes the columns into
typed arrays and transfers them to the worker. That encode has to run on the
main thread: it calls your valueGetter and calculateSortValue, and functions
cannot be sent to a worker.
The figures in this section come from an earlier measurement on a wider 15-column dataset, not from the 9-column script above. At 100,000 rows the whole pass is 169 ms and unnoticeable. At 1,000,000 rows × 15 columns it is about 3.1 seconds of work — but it is not one 3.1-second freeze. The encode yields to the browser between columns, so the longest single task is one column:
| 1,000,000 rows × 15 columns | |
|---|---|
| Total encode work | ~3.1 s |
| Longest blocking task | ~0.55 s (the widest text column) |
| Tasks | 15, with a paint opportunity between each |
The total work is unchanged — the same rows still have to be read — but the page keeps painting and responding while it happens instead of going white.
Yielding is skipped below 100,000 rows, where the whole encode is a few milliseconds and the scheduling round trip would cost more than it saves.