TTFT (first_token_ms) is only recorded for streaming requests, but the
ops dashboard weighted merged TTFT percentiles by success_count (all
successful requests, streaming + non-streaming). When non-streaming
traffic was present this diluted/skewed the merged TTFT figures shown
for longer (pre-aggregated) time ranges; the realtime path was exact.
Add a per-bucket ttft_sample_count (rows that actually recorded
first_token_ms) to ops_metrics_hourly / ops_metrics_daily and weight all
TTFT percentile merges by it instead of success_count:
- hourly/daily pre-agg upserts populate and propagate ttft_sample_count;
daily TTFT p50/p90/avg now weighted by ttft_sample_count.
- dashboard hourly-row merge and cross-segment combine weight TTFT by
the streaming sample count; queryUsageLatency returns it for raw
head/tail fragments.
duration stays weighted by success_count (recorded for every request);
p95/p99/max keep the conservative MAX merge (weight-independent).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>