* [LIQ] Add new aggregate functions, aliases, and queryable aggregate registry * Extend with 13 new built-in aggregates: `product`, `string_agg`, `yaml_agg`, `json_agg`, `bit_and`, `bit_or`, `bit_xor`, `bool_and`, `bool_or`, `stddev_pop`, `stddev_samp`, `var_pop` and `var_samp`. * Introduce `aggregate.alias` API allowing users to define custom aliases for any aggregate. Standard aliases (`every`, `std`, `stddev` and `variance`) are now defined via this API rather than hardcoded. * Add `index.aggregates` queryable collection so users can discover all available aggregates directly from LIQ queries. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix config pass through query path so custom aggregates work Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Preserve `LuaTable`/`LuaFunction` values in aggregate config storage `config.set` uses `LuaNativeJSFunction` which calls `luaValueToJS` on all arguments. This converted the aggregate `LuaTable` to a plain JS object and wrapped `LuaFunction` callbacks in JS functions that also converted their returned values via `luaValueToJS`. The result was that state returned by initialize (a `LuaTable`) got converted to a plain JS object before being passed to `iterate`. Therefor Lua operations like `table.insert` on that were failing because they expected a `LuaTable` and not a plain JS array. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix formatting Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Improve aggregate functions descriptions, fix `sum` divergence Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Align `product` with `sum` Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: extract `alias` from `LuaTable` via `rawGet` in `aggregates()` registry Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Rename `alias` in `aggregates()` to `target` for clarity Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add a null guard at the top of `jsToLuaValue` This preserves `null`/`undefined` as-is (both map to Lua nil) and prevents them from falling through to the `typeof` "object" branch. For this PR it means that null `target` in our `aggregates` entries will correctly show as empty/`nil` in query results rather than `{}`. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Documentation reflects recent changes Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: make `sum`/`product` return null on empty input; stop `LIQ_NULL` leaks * `sum(`) and` product(`) now return null when no rows match (matching Postgres semantics) instead of returning 0 and 1 respectively. * Query result columns that hold null are internally preserved using a `LIQ_NULL` sentinel so that column keys survive in `LuaTable` storage. This sentinel was leaking into Lua code as "userdata" through three read paths: * `luaIndexValue`: `rawGet` returned the sentinel directly to Lua when accessing table fields, * `rawget` (stdlib): the builtin `rawget` function exposed the sentinel without converting it back to `nil`, * `createAugmentedEnv`: string interpolation unpacked table values via `rawGet` into local variables, making the sentinel visible in template expressions like `${var}`. All three now convert `LIQ_NULL` to `nil` at the read boundary, keeping the sentinel internal to table storage where it belongs. * Update affected test expectations accordingly. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Remove duplicated LIQ_NULL hazard, add guard for all builtin aggregate `iterate`s Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: `array_agg` preserves NULL positions Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add symbol guard to `json_agg` `JSON.stringify(Symbol(...))` in an array produces null by accident. That is a JS implementation detail we **MUST NOT** rely on. Explicit null push makes intent clear and avoids surprises if the `Symbol` representation ever changes. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add symbol guard to `yaml_agg` (ditto) `js-yaml` has no knowledge of the `LIQ_NULL` symbol. Passing null makes it emit YAML null (or `~`), which is the correct YAML representation of a missing value and matches standard `json_agg`/`yaml_agg` NULL-inclusion semantics. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add intra-aggregate ordering null guards Without this, `LIQ_NULL` sort keys would fall through to `valA < valB` which is always false for `Symbol`s which is breaking the `nulls first`/`nulls last` contract... Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Ditto, but for `order by` null comparisons Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Guard `luaTypeName`, `luaTypeOf` and `luaToString` against `LIQ_NULL` sentinel Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Guard presentation layer against `LIQ_NULL` sentinel leaking as visible text Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Evaluate extra args per-item in `executeAggregate`; add new aggregates Extra arguments (2nd, 3rd, etc.) to aggregate functions were evaluated against the outer query environment where the object variable is not bound. This caused multi-argument aggregates like `covar_samp(data.y, data.x)` to fail with nil reference errors. This commit addresses this by evaluating extra args per-item inside the iterate loop using the item environment so all arguments resolve correctly. We also add few common aggregates: - `covar_pop`, `covar_samp`, `corr`: population/sample covariance and correlation coefficient using online co-moment algorithm. - `quantile(value, q, method)`: general quantile with interpolation methods: lower, higher, nearest, midpoint and default linear. - `percentile_cont(value, q)`: continuous percentile (linear) - `percentile_disc(value, q)`: discrete percentile (lower) Note: `percentile_cont` and `percentile_disc` share the `quantile` implementation through `ctx.name` at initialize time. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Update docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Make the ordering for quantile aggregates explicit Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Update docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Improve docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> --------- Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>
6.1 KiB
#maturity/experimental
The group by and having clauses of Space Lua/Lua Integrated Query support aggregate functions for grouped analysis, following SQL-style semantics.
After group by, each result row contains:
key: the group key (a single value or, for multi-key grouping, a table)group: a Lua table containing all items in that group
All aggregate functions (such as count, sum, min, max, avg, and custom aggregates) can be applied in select and having clauses. Aggregate expressions are available in both forms: with or without a variable binding in the from clause. The variable _ always refers to the current item.
Field names used in group by are exposed as locals in having, select, and order by. Use #group to obtain the item count per group.
Note
Note The
havingclause acts only on grouped output. For filtering individual items, usewhereprior to grouping.
Available aggregates
All registered aggregate functions — built-in, user-defined, and aliases — can be listed via index.aggregates():
${query[[
select
{
Name = '' .. name .. '',
Description = description,
Kind =
(builtin and 'builtin' or 'custom') ..
(target and ' alias for ' .. '' .. target .. '' or ''),
}
from
index.aggregates()
order by
builtin desc,
name
]]}
See Library/Std/APIs/Aggregate for how to define custom aggregates and aliases.
Examples
All example queries operate on tags.page, but will work with any query collection. As always, to see the underlying query, hover over the result table and click the Edit button to see the underlying query.
Counting with and without binding
Grouping pages by their first tag, and computing the count and aggregate statistics:
Without binding variable ${query [[ from tags.page group by tags[1] select { tag = key, total = count(name), min_size = min(size), max_size = max(size), avg_size = avg(size) } order by total desc ]]}
With binding variable ${query [[ from p = tags.page group by p.tags[1] select { tag = key, total = count(p.name), min_size = min(p.size), max_size = max(p.size), avg_size = avg(p.size) } order by total desc ]]}
Multi-key grouping and aggregate
${query[[ from p = tags.page group by p.tags[1], p.tags[2] select { first = key[1], second = key[2], count = count(p.name) } ]]}
Group filtering with having and aggregates
Only groups with more than two items and at least one tag set: ${query[[ from p = tags.page group by p.tags[1] having count(p.name) > 2 and key select { tag = key, total = count(p.name) } ]]}
Per-aggregate filtering with filter(where ...)
Individual aggregate expressions can include a filter(where <condition>) clause to restrict which rows contribute to that specific aggregate.
Unlike where (which filters rows before grouping) and having (which filters entire groups after aggregation), filter(where ...) applies per-aggregate, per-row within each group. Multiple aggregates in the same select can each have different filters.
${query [[ from p = index.tag 'page' group by p.tags[1] select { tag = key, total = count(p.name), big = count(p.name) filter(where p.size > 10), big_sz = sum(p.size) filter(where p.size > 10) } order by tag ]]}
The filter clause works with all aggregate functions: count, sum, min, max, avg, array_agg, and custom aggregates. When no rows match the filter condition, aggregates return their empty-group value: 0 for count, nil for sum, min, max, and avg, and an empty table {} for array_agg.
Intra-aggregate order by
Aggregate functions can include an order by clause inside the function call to control the order in which values are processed.
For commutative aggregates like sum, count, min, max, and avg, the intra-aggregate order by has no effect on the result because the value is the same regardless of iteration order. It is only meaningful for order-dependent aggregates like array_agg.
Ordered-set aggregates such as quantile, percentile_cont, and percentile_disc require an intra-aggregate order by clause to produce correct results, as they depend on the iteration order of input values. Without order by, results are undefined.
Basic example
Collect page names sorted alphabetically within each group:
${query [[ from p = index.tag 'page' group by p.tags[1] select { tag = key, names_asc = array_agg(p.name order by p.name asc), names_desc = array_agg(p.name order by p.name desc) } order by tag limit 5 ]]}
Combined with filter(where ...)
The order by and filter clauses can be used together. The filter is applied first (excluding rows), then the remaining rows are sorted before iteration:
${query [[ from p = index.tag 'page' group by p.tags[1] select { tag = key, big_by_size = array_agg(p.name order by p.size desc) filter(where p.size > 5) } order by tag limit 5 ]]}
Null handling
The nulls first and nulls last modifiers work inside intra-aggregate order by the same way they do in the query-level order by:
query [[
from
p = data
group
by p.category
select {
cat = key,
items = array_agg(p.name
order by
p.priority asc nulls last
)
}
]]
Field access after grouping
Non-aggregated field references, such as name in select, refer to the first item in the group, matching common SQL and MySQL semantics.
${query [[ from p = tags.page group by p.tags[1] select { tag = key, first_page = p.name, n = count(p.name) } ]]}
Custom aggregators
Custom aggregator functions may be defined by the user using Library/Std/APIs/Aggregate.
See also
- Space Lua/Lua Integrated Query/Grouping — grouping queries without aggregation
- Space Lua/Lua Integrated Query — full LIQ language reference and listing available aggregates