Files
plainleaf/website/Space Lua/Lua Integrated Query/Grouping.md
T
Matouš Jan FialkaandGitHub 6cbb61c7db [LIQ] Add new aggregate functions, aliases, and queryable aggregate registry (#1891)
* [LIQ] Add new aggregate functions, aliases, and queryable aggregate registry

* Extend with 13 new built-in aggregates: `product`, `string_agg`,
  `yaml_agg`, `json_agg`, `bit_and`, `bit_or`, `bit_xor`, `bool_and`,
  `bool_or`, `stddev_pop`, `stddev_samp`, `var_pop` and `var_samp`.

* Introduce `aggregate.alias` API allowing users to define custom
  aliases for any aggregate. Standard aliases (`every`, `std`, `stddev`
  and `variance`) are now defined via this API rather than hardcoded.

* Add `index.aggregates` queryable collection so users can discover
  all available aggregates directly from LIQ queries.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix config pass through query path so custom aggregates work

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Preserve `LuaTable`/`LuaFunction` values in aggregate config storage

`config.set` uses `LuaNativeJSFunction` which calls `luaValueToJS` on
all arguments. This converted the aggregate `LuaTable` to a plain JS
object and wrapped `LuaFunction` callbacks in JS functions that also
converted their returned values via `luaValueToJS`. The result was that
state returned by initialize (a `LuaTable`) got converted to a plain JS
object before being passed to `iterate`. Therefor Lua operations like
`table.insert` on that were failing because they expected a `LuaTable`
and not a plain JS array.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix formatting

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Improve aggregate functions descriptions, fix `sum` divergence

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Align `product` with `sum`

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: extract `alias` from `LuaTable` via `rawGet` in `aggregates()` registry

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Rename `alias` in `aggregates()` to `target` for clarity

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Add a null guard at the top of `jsToLuaValue`

This preserves `null`/`undefined` as-is (both map to Lua nil) and
prevents them from falling through to the `typeof` "object" branch.

For this PR it means that null `target` in our `aggregates` entries will
correctly show as empty/`nil` in query results rather than `{}`.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Documentation reflects recent changes

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: make `sum`/`product` return null on empty input; stop `LIQ_NULL` leaks

* `sum(`) and` product(`) now return null when no rows match (matching
  Postgres semantics) instead of returning 0 and 1 respectively.

* Query result columns that hold null are internally preserved using
  a `LIQ_NULL` sentinel so that column keys survive in `LuaTable`
  storage.  This sentinel was leaking into Lua code as "userdata"
  through three read paths:

  * `luaIndexValue`: `rawGet` returned the sentinel directly to Lua when
    accessing table fields,

  * `rawget` (stdlib): the builtin `rawget` function exposed the
    sentinel without converting it back to `nil`,

  * `createAugmentedEnv`: string interpolation unpacked table values via
    `rawGet` into local variables, making the sentinel visible in
    template expressions like `${var}`.

  All three now convert `LIQ_NULL` to `nil` at the read boundary,
  keeping the sentinel internal to table storage where it belongs.

* Update affected test expectations accordingly.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Remove duplicated LIQ_NULL hazard, add guard for all builtin aggregate `iterate`s

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: `array_agg` preserves NULL positions

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Add symbol guard to `json_agg`

`JSON.stringify(Symbol(...))` in an array produces null by accident.
That is a JS implementation detail we **MUST NOT** rely on. Explicit
null push makes intent clear and avoids surprises if the `Symbol`
representation ever changes.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Add symbol guard to `yaml_agg` (ditto)

`js-yaml` has no knowledge of the `LIQ_NULL` symbol. Passing null makes
it emit YAML null (or `~`), which is the correct YAML representation of
a missing value and matches standard `json_agg`/`yaml_agg`
NULL-inclusion semantics.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Add intra-aggregate ordering null guards

Without this, `LIQ_NULL` sort keys would fall through to `valA < valB`
which is always false for `Symbol`s which is breaking the `nulls
first`/`nulls last` contract...

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Ditto, but for `order by` null comparisons

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Guard `luaTypeName`, `luaTypeOf` and `luaToString` against `LIQ_NULL` sentinel

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Guard presentation layer against `LIQ_NULL` sentinel leaking as visible text

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Fix: Evaluate extra args per-item in `executeAggregate`; add new aggregates

Extra arguments (2nd, 3rd, etc.) to aggregate functions were evaluated
against the outer query environment where the object variable is not
bound. This caused multi-argument aggregates like `covar_samp(data.y,
data.x)` to fail with nil reference errors. This commit addresses this
by evaluating extra args per-item inside the iterate loop using the item
environment so all arguments resolve correctly.

We also add few common aggregates:

- `covar_pop`, `covar_samp`, `corr`: population/sample covariance and
  correlation coefficient using online co-moment algorithm.

- `quantile(value, q, method)`: general quantile with interpolation
  methods: lower, higher, nearest, midpoint and default linear.

- `percentile_cont(value, q)`: continuous percentile (linear)

- `percentile_disc(value, q)`: discrete percentile (lower)

Note: `percentile_cont` and `percentile_disc` share the `quantile`
implementation through `ctx.name` at initialize time.

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Update docs

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Make the ordering for quantile aggregates explicit

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Update docs

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

* Improve docs

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>

---------

Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>
2026-03-19 09:37:09 +01:00

4.3 KiB

#maturity/experimental

The group by and having clauses extend Space Lua/Lua Integrated Query with SQL-style grouping and aggregate filtering.

After group by, each result row has two fields:

  • key - the group key (single value or table for multi-key)
  • group - a table (array) of all items in that group

The field names used in group by are also available as bare variables in having, select, and order by. Use #group to count items per group.

Note

Note having can only reference group key fields, key, group, aggregate expressions like #group, and aggregate functions like count(). To filter individual rows, use where.

Examples

All examples below use tags.tag.

Group by single key

Group all tags by name:

${query [[ from t = tags.tag group by t.name limit 5 ]]}

Group by multiple keys

Group tags by name and parent together:

${query[[ from t = tags.tag group by t.name, t.parent limit 5 ]]}

Filter groups by count

Only show tags that appear more than 2 times:

${query[[ from t = tags.tag group by t.name having #group > 2 limit 5 ]]}

Find unique tags

Tags appearing exactly once:

${query[[ from t = tags.tag group by t.name having #group == 1 select key ]]}

Filter groups by key value

Only show the group where name is "meta":

${query[[ from tags.tag group by name having name == "meta" ]]}

${query[[ from t = tags.tag group by t.name having t.name == "meta" ]]}

Multi-key having

Groups by name and parent, keep only page-level tags with more than 1 entry:

${query [[ from tags.tag group by name, parent having parent == 'page' and #group > 1 ]]}

where before group by

Filter to page parents first, then group by name:

${query [[ from tags.tag where parent == 'page' group by name ]]}

where, group by and having combined

Filter to page parents, group by name, keep groups with 2+ items: ${query [[ from index.tag 'tag' where parent == 'page' group by name having #group >= 2 ]]}

select name and count

Project each group into a table with name and count: ${query [[ from index.tag 'tag' group by name select { name = name, count = #group } ]]}

select with multi-key

Project both key parts and count:

${query [[ from index.tag 'tag' group by name, parent select { name = name, parent = parent, count = #group } ]]}

Full pipeline: where, group by, having and select

Filter, group, filter groups, then project:

${query [[ from index.tag 'tag' where parent == 'page' or parent == 'task' group by name having #group > 1 select { tag = name, total = #group } ]]}

Order groups by count

Sort groups by size, largest first:

${query [[ from index.tag 'tag' group by name order by #group desc ]]}

Top tags with having, order by, and select

Tags with 2+ occurrences, sorted by count, projected: ${query [[ from index.tag 'tag' group by name having #group >= 2 order by #group desc select { tag = name, count = #group } ]]}

Top N groups with limit

${query [[ from index.tag 'tag' group by name order by #group desc limit 3 ]]}

Full pipeline with limit

Top 5 tags with 2+ uses, showing name and count: ${query [[ from p = index.tag 'tag' group by p.name having #group > 1 select { tag = name, count = #group } ]]}

Multi-key with explicit object variable

Full pipeline with p = binding and two group keys:

${query [[ from p = index.tag 'tag' where p.parent == 'page' group by p.name, p.parent having #group >= 2 order by #group desc select { tag = name, parent = parent, count = #group } ]]}

Access key directly

For single-key grouping, key holds the value directly: ${query [[ from index.tag 'tag' group by name having key == 'meta' ]]}

Access key table for multi-key

For multi-key grouping, key is a table indexed from 1: ${query [[ from index.tag 'tag' group by name, parent having key[1] == 'meta' and key[2] == 'page' ]]}