- Add prompt_audit_events.full_prompt (migration 182) so admins can review
the exact unredacted prompt that triggered a finding; blocking mode writes
it from the snapshot, async mode reconstructs it from the Redis scan
payload so jobs rows stay redaction-only
- Event detail API returns full_prompt (list endpoint stays lean); text is
NUL-stripped and capped at 65536 runes
- Detail dialog shows the full prompt in a scrollable pane with fallback to
the legacy redacted preview; page copy updated to match the new behavior
- Rework filter deletion into a dedicated dialog with time-range presets and
criteria-change preview invalidation; localize decision/risk/category
labels across the events workspace
- Fix pre-existing i18n message-compile spec by declaring the
@intlify/message-compiler dev dependency
Admins often recreate groups with the same pricing, routing, and account membership. A server-side duplicate creates an inactive copy for review, preserves eligible account priorities, and recovers ambiguous retries without creating extra groups.
Constraint: Group has no neutral JSON metadata field for durable operation recovery
Constraint: Model routing references account IDs, so copied configuration requires matching bindings
Rejected: Rebuild from the list response | it omits configuration and account priority details
Rejected: Store operation identity in business configuration | it would pollute real group settings
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep duplicated groups inactive until an administrator reviews the copied configuration
Tested: Go unit and full tests, go vet, integration-tag compile, frontend Vitest, lint, typecheck, production build, and Playwright duplicate flow
Not-tested: PostgreSQL container integration locally because Docker is unavailable; CI will execute the database-backed suite
Follow-up fixes for the #3775 audit findings:
- Bill Grok video generation per second of output, matching the xAI rate
card: parse the request duration (1-15s, upstream default 8s) and compute
cost as per-second price x duration x count. The built-in rate card values
were already xAI per-second prices but were previously charged per video,
undercharging up to 15x with a user-controlled duration.
- Group video_price_* fields are now documented and surfaced as per-second
rates (USD/s); admin UI labels, placeholders and hints updated accordingly.
- Persist video_count/video_resolution/video_duration_seconds on usage_logs
(migration 172) so video billing is auditable, and exempt any row with
video_count > 0 from the image_size check constraint: a video billed via a
token-mode channel price produces billing_mode='token' with image_count=1
and no image_size, which the previous constraint rejected, dropping the
whole billing transaction.
- Only refetch the group in apiKeyWithFreshGroupMediaPricing when the group
object actually looks like it is missing media pricing fields (both media
multipliers zero and all prices nil, impossible for a normally loaded
group), removing a per-usage DB query for groups without overrides.
- Frontend: drop the unused admin.groups.mediaPricing locale block, map
cleared price inputs to null (create) / -1 (update, cleared via backend
normalizePrice) instead of sending "" that failed *float64 unmarshalling,
and align video price placeholders with the text-to-video default model
(grok-imagine-video 0.05/0.07, 1080p only on 1.5 at 0.25).
The risk control center's moderation log records had no field for the
keyword that triggered a keyword block, so the admin UI couldn't show
which keyword was hit (only the application slog logged it).
- migration 156: add matched_keyword column to content_moderation_logs
- ContentModerationLog gains MatchedKeyword; set it on keyword block
- repo CreateLog/ListLogs persist and read the column
- frontend: show "命中关键词" inline in the log table and detail modal
- i18n: add matchedKeyword (zh/en)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
TTFT (first_token_ms) is only recorded for streaming requests, but the
ops dashboard weighted merged TTFT percentiles by success_count (all
successful requests, streaming + non-streaming). When non-streaming
traffic was present this diluted/skewed the merged TTFT figures shown
for longer (pre-aggregated) time ranges; the realtime path was exact.
Add a per-bucket ttft_sample_count (rows that actually recorded
first_token_ms) to ops_metrics_hourly / ops_metrics_daily and weight all
TTFT percentile merges by it instead of success_count:
- hourly/daily pre-agg upserts populate and propagate ttft_sample_count;
daily TTFT p50/p90/avg now weighted by ttft_sample_count.
- dashboard hourly-row merge and cross-segment combine weight TTFT by
the streaming sample count; queryUsageLatency returns it for raw
head/tail fragments.
duration stays weighted by success_count (recorded for every request);
p95/p99/max keep the conservative MAX merge (weight-independent).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>