* [LIQ] Add new aggregate functions, aliases, and queryable aggregate registry * Extend with 13 new built-in aggregates: `product`, `string_agg`, `yaml_agg`, `json_agg`, `bit_and`, `bit_or`, `bit_xor`, `bool_and`, `bool_or`, `stddev_pop`, `stddev_samp`, `var_pop` and `var_samp`. * Introduce `aggregate.alias` API allowing users to define custom aliases for any aggregate. Standard aliases (`every`, `std`, `stddev` and `variance`) are now defined via this API rather than hardcoded. * Add `index.aggregates` queryable collection so users can discover all available aggregates directly from LIQ queries. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix config pass through query path so custom aggregates work Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Preserve `LuaTable`/`LuaFunction` values in aggregate config storage `config.set` uses `LuaNativeJSFunction` which calls `luaValueToJS` on all arguments. This converted the aggregate `LuaTable` to a plain JS object and wrapped `LuaFunction` callbacks in JS functions that also converted their returned values via `luaValueToJS`. The result was that state returned by initialize (a `LuaTable`) got converted to a plain JS object before being passed to `iterate`. Therefor Lua operations like `table.insert` on that were failing because they expected a `LuaTable` and not a plain JS array. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix formatting Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Improve aggregate functions descriptions, fix `sum` divergence Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Align `product` with `sum` Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: extract `alias` from `LuaTable` via `rawGet` in `aggregates()` registry Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Rename `alias` in `aggregates()` to `target` for clarity Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add a null guard at the top of `jsToLuaValue` This preserves `null`/`undefined` as-is (both map to Lua nil) and prevents them from falling through to the `typeof` "object" branch. For this PR it means that null `target` in our `aggregates` entries will correctly show as empty/`nil` in query results rather than `{}`. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Documentation reflects recent changes Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: make `sum`/`product` return null on empty input; stop `LIQ_NULL` leaks * `sum(`) and` product(`) now return null when no rows match (matching Postgres semantics) instead of returning 0 and 1 respectively. * Query result columns that hold null are internally preserved using a `LIQ_NULL` sentinel so that column keys survive in `LuaTable` storage. This sentinel was leaking into Lua code as "userdata" through three read paths: * `luaIndexValue`: `rawGet` returned the sentinel directly to Lua when accessing table fields, * `rawget` (stdlib): the builtin `rawget` function exposed the sentinel without converting it back to `nil`, * `createAugmentedEnv`: string interpolation unpacked table values via `rawGet` into local variables, making the sentinel visible in template expressions like `${var}`. All three now convert `LIQ_NULL` to `nil` at the read boundary, keeping the sentinel internal to table storage where it belongs. * Update affected test expectations accordingly. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Remove duplicated LIQ_NULL hazard, add guard for all builtin aggregate `iterate`s Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: `array_agg` preserves NULL positions Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add symbol guard to `json_agg` `JSON.stringify(Symbol(...))` in an array produces null by accident. That is a JS implementation detail we **MUST NOT** rely on. Explicit null push makes intent clear and avoids surprises if the `Symbol` representation ever changes. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add symbol guard to `yaml_agg` (ditto) `js-yaml` has no knowledge of the `LIQ_NULL` symbol. Passing null makes it emit YAML null (or `~`), which is the correct YAML representation of a missing value and matches standard `json_agg`/`yaml_agg` NULL-inclusion semantics. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Add intra-aggregate ordering null guards Without this, `LIQ_NULL` sort keys would fall through to `valA < valB` which is always false for `Symbol`s which is breaking the `nulls first`/`nulls last` contract... Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Ditto, but for `order by` null comparisons Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Guard `luaTypeName`, `luaTypeOf` and `luaToString` against `LIQ_NULL` sentinel Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Guard presentation layer against `LIQ_NULL` sentinel leaking as visible text Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Fix: Evaluate extra args per-item in `executeAggregate`; add new aggregates Extra arguments (2nd, 3rd, etc.) to aggregate functions were evaluated against the outer query environment where the object variable is not bound. This caused multi-argument aggregates like `covar_samp(data.y, data.x)` to fail with nil reference errors. This commit addresses this by evaluating extra args per-item inside the iterate loop using the item environment so all arguments resolve correctly. We also add few common aggregates: - `covar_pop`, `covar_samp`, `corr`: population/sample covariance and correlation coefficient using online co-moment algorithm. - `quantile(value, q, method)`: general quantile with interpolation methods: lower, higher, nearest, midpoint and default linear. - `percentile_cont(value, q)`: continuous percentile (linear) - `percentile_disc(value, q)`: discrete percentile (lower) Note: `percentile_cont` and `percentile_disc` share the `quantile` implementation through `ctx.name` at initialize time. Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Update docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Make the ordering for quantile aggregates explicit Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Update docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> * Improve docs Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz> --------- Signed-off-by: Matouš Jan Fialka <mjf@mjf.cz>
213 lines
6.1 KiB
Markdown
213 lines
6.1 KiB
Markdown
#maturity/experimental
|
|
|
|
The `group by` and `having` clauses of [[Space Lua/Lua Integrated Query]] support aggregate functions for grouped analysis, following SQL-style semantics.
|
|
|
|
After `group by`, each result row contains:
|
|
|
|
- `key`: the group key (a single value or, for multi-key grouping, a table)
|
|
- `group`: a Lua table containing all items in that group
|
|
|
|
All aggregate functions (such as `count`, `sum`, `min`, `max`, `avg`, and custom aggregates) can be applied in `select` and `having` clauses. Aggregate expressions are available in both forms: with or without a variable binding in the `from` clause. The variable `_` always refers to the current item.
|
|
|
|
Field names used in `group by` are exposed as locals in `having`, `select`, and `order by`. Use `#group` to obtain the item count per group.
|
|
|
|
> **note** Note
|
|
> The `having` clause acts only on grouped output. For filtering individual items, use `where` prior to grouping.
|
|
|
|
# Available aggregates
|
|
|
|
All registered aggregate functions — built-in, user-defined, and aliases — can be listed via `index.aggregates()`:
|
|
|
|
${query[[
|
|
select
|
|
{
|
|
Name = '`' .. name .. '`',
|
|
Description = description,
|
|
Kind =
|
|
(builtin and 'builtin' or 'custom') ..
|
|
(target and ' alias for ' .. '`' .. target .. '`' or ''),
|
|
}
|
|
from
|
|
index.aggregates()
|
|
order by
|
|
builtin desc,
|
|
name
|
|
]]}
|
|
|
|
See [[Library/Std/APIs/Aggregate|Aggregate API]] for how to define custom aggregates and aliases.
|
|
|
|
# Examples
|
|
All example queries operate on `tags.page`, but will work with any query collection. As always, to see the underlying query, hover over the result table and click the _Edit_ button to see the underlying query.
|
|
|
|
## Counting with and without binding
|
|
Grouping pages by their first tag, and computing the count and aggregate statistics:
|
|
|
|
**Without binding variable**
|
|
${query [[
|
|
from
|
|
tags.page
|
|
group by
|
|
tags[1]
|
|
select {
|
|
tag = key,
|
|
total = count(name),
|
|
min_size = min(size),
|
|
max_size = max(size),
|
|
avg_size = avg(size)
|
|
}
|
|
order by total desc
|
|
]]}
|
|
|
|
**With binding variable**
|
|
${query [[
|
|
from
|
|
p = tags.page
|
|
group by
|
|
p.tags[1]
|
|
select {
|
|
tag = key,
|
|
total = count(p.name),
|
|
min_size = min(p.size),
|
|
max_size = max(p.size),
|
|
avg_size = avg(p.size)
|
|
}
|
|
order by total desc
|
|
]]}
|
|
|
|
## Multi-key grouping and aggregate
|
|
${query[[
|
|
from
|
|
p = tags.page
|
|
group by
|
|
p.tags[1],
|
|
p.tags[2]
|
|
select {
|
|
first = key[1],
|
|
second = key[2],
|
|
count = count(p.name)
|
|
}
|
|
]]}
|
|
|
|
## Group filtering with `having` and aggregates
|
|
Only groups with more than two items and at least one tag set:
|
|
${query[[
|
|
from
|
|
p = tags.page
|
|
group by
|
|
p.tags[1]
|
|
having
|
|
count(p.name) > 2 and key
|
|
select {
|
|
tag = key,
|
|
total = count(p.name)
|
|
}
|
|
]]}
|
|
|
|
## Per-aggregate filtering with `filter(where ...)`
|
|
Individual aggregate expressions can include a `filter(where <condition>)` clause to restrict which rows contribute to that specific aggregate.
|
|
|
|
Unlike `where` (which filters rows before grouping) and `having` (which filters entire groups after aggregation), `filter(where ...)` applies per-aggregate, per-row within each group. Multiple aggregates in the same `select` can each have different filters.
|
|
|
|
${query [[
|
|
from
|
|
p = index.tag 'page'
|
|
group by
|
|
p.tags[1]
|
|
select {
|
|
tag = key,
|
|
total = count(p.name),
|
|
big = count(p.name) filter(where p.size > 10),
|
|
big_sz = sum(p.size) filter(where p.size > 10)
|
|
}
|
|
order by
|
|
tag
|
|
]]}
|
|
|
|
The filter clause works with all aggregate functions: `count`, `sum`, `min`, `max`, `avg`, `array_agg`, and custom aggregates. When no rows match the filter condition, aggregates return their empty-group value: `0` for `count`, `nil` for `sum`, `min`, `max`, and `avg`, and an empty table `{}` for `array_agg`.
|
|
|
|
## Intra-aggregate `order by`
|
|
Aggregate functions can include an `order by` clause **inside** the function call to control the order in which values are processed.
|
|
|
|
For commutative aggregates like `sum`, `count`, `min`, `max`, and `avg`, the intra-aggregate `order by` has no effect on the result because the value is the same regardless of iteration order. It is only meaningful for order-dependent aggregates like `array_agg`.
|
|
|
|
Ordered-set aggregates such as `quantile`, `percentile_cont`, and `percentile_disc` require an intra-aggregate `order by` clause to produce correct results, as they depend on the iteration order of input values. Without `order by`, results are undefined.
|
|
|
|
### Basic example
|
|
|
|
Collect page names sorted alphabetically within each group:
|
|
|
|
${query [[
|
|
from
|
|
p = index.tag 'page'
|
|
group by
|
|
p.tags[1]
|
|
select {
|
|
tag = key,
|
|
names_asc = array_agg(p.name order by p.name asc),
|
|
names_desc = array_agg(p.name order by p.name desc)
|
|
}
|
|
order by
|
|
tag
|
|
limit
|
|
5
|
|
]]}
|
|
|
|
### Combined with `filter(where ...)`
|
|
The `order by` and `filter` clauses can be used together. The filter is applied first (excluding rows), then the remaining rows are sorted before iteration:
|
|
|
|
${query [[
|
|
from
|
|
p = index.tag 'page'
|
|
group by
|
|
p.tags[1]
|
|
select {
|
|
tag = key,
|
|
big_by_size = array_agg(p.name order by p.size desc) filter(where p.size > 5)
|
|
}
|
|
order by
|
|
tag
|
|
limit
|
|
5
|
|
]]}
|
|
|
|
### Null handling
|
|
The `nulls first` and `nulls last` modifiers work inside intra-aggregate `order by` the same way they do in the query-level `order by`:
|
|
|
|
```lua
|
|
query [[
|
|
from
|
|
p = data
|
|
group
|
|
by p.category
|
|
select {
|
|
cat = key,
|
|
items = array_agg(p.name
|
|
order by
|
|
p.priority asc nulls last
|
|
)
|
|
}
|
|
]]
|
|
```
|
|
|
|
## Field access after grouping
|
|
Non-aggregated field references, such as `name` in `select`, refer to the first item in the group, matching common SQL and MySQL semantics.
|
|
|
|
${query [[
|
|
from
|
|
p = tags.page
|
|
group by
|
|
p.tags[1]
|
|
select {
|
|
tag = key,
|
|
first_page = p.name,
|
|
n = count(p.name)
|
|
}
|
|
]]}
|
|
|
|
## Custom aggregators
|
|
Custom aggregator functions may be defined by the user using [[Library/Std/APIs/Aggregate|dedicated API]].
|
|
|
|
# See also
|
|
* [[Space Lua/Lua Integrated Query/Grouping]] — grouping queries without aggregation
|
|
* [[Space Lua/Lua Integrated Query]] — full LIQ language reference and listing available aggregates
|