Files
hapi/ios/Packages/HapiKit/Sources/HapiProtocol/Models/VoiceApi.swift
T
weishu 05d063cc33 feat(ios): composer attachments + dictation (A-M3f)
Attachments (Android B-M3f semantics ported verbatim):
- HapiClient/Attachments/AttachmentPolicy — pure plan matrix (>4 MB
  recompressible image -> 2048 px JPEG q85 with .jpg rename, 50 MB hard
  reject, 192 MB image read cap, 512 px q80 previewUrl data-URL thumbs,
  data-URL parse/round-trip).
- HapiClient/Attachments/ComposerAttachments — upload-on-pick tray over an
  AttachmentUploading seam (APIClient conforms): uploading/ready/failed
  chips, retained payload for retry, remove -> best-effort delete,
  mid-upload removal deletes the orphan on completion, consume() ->
  AttachmentMetadata with JPEG data-URL previewUrl, discardAllDetached +
  deinit orphan cleanup (Android onCleared analogue).
- ChatInteractor: tray ownership, unsettled chips refuse the send with a
  notice, attachments-only sends post empty text, optimistic rows carry
  the metadata, appendDictatedText/postNotice/discardAttachments.
- App: AttachmentPreparer (capped security-scoped reads, ImageIO
  downscale/encode with EXIF transform, HEIC-undecodable fallback),
  PhotosPicker multi (videos via FileRepresentation temp files),
  UIImagePickerController camera capture, fileImporter; composer chip row
  (thumb/spinner/tap-to-retry/remove) with attachment-aware send gating;
  user bubbles upgrade chips to off-main-decoded previewUrl thumbnails
  (web-sent attachments included).

Dictation (Android B-M3ce port):
- HapiProtocol/Models/VoiceApi — TranscriptionResponse,
  TranscriptionProvidersResponse, TranscriptionProviderInfo.
- HapiClient/Endpoints/VoiceEndpoints — GET
  /api/voice/transcription/providers + multipart POST
  /api/voice/transcription (file/provider/mode/language, Android part
  order) over MultipartFormData; DictationTranscribing conformance.
- HapiClient/Voice/DictationController — idle/starting/recording/
  transcribing, transcribed/noProvider/error events, provider memoized
  (first standard-capable entry), appendTranscript port.
- App: AVAudioRecorderDictation (m4a/AAC mono 44.1 kHz 96 kbps, session
  activate/deactivate), mic button + recording chip (elapsed + cancel),
  record-permission request via AVAudioApplication.

Info.plist gains NSMicrophoneUsageDescription; the camera string now
covers attachment capture (modern PhotosPicker needs no photo-library
permission). Tests: policy matrix, tray over the real client with exact
base64 upload bodies + gated in-flight scenarios, dictation controller
suite with fake recorder/transport, voice endpoint request shapes, and
the interactor attachment-send matrix transcribed from the Android VM
tests (wire bodies byte-for-byte).
2026-08-18 11:27:28 +08:00

45 lines
1.6 KiB
Swift

import Foundation
// Wire types of the hub's voice surface (`hub/src/web/routes/voice.ts`),
// added for A-M3f dictation in lockstep with the Android port
// (`app.hapi.protocol.wire` — ApiResponses.kt). Only the standard
// (uploaded-file) transcription path is modeled; realtime tokens are out of
// scope for native v1.
/// Body of `POST /api/voice/transcription` (the one multipart endpoint).
public struct TranscriptionResponse: Codable, Equatable, Sendable {
public var text: String
public var language: String?
public init(text: String, language: String? = nil) {
self.text = text
self.language = language
}
}
/// Body of `GET /api/voice/transcription/providers` — only providers whose
/// keys are configured on the hub (`listConfiguredTranscriptionProviders`,
/// `shared/src/voice.ts`). Empty ⇒ dictation is unavailable.
public struct TranscriptionProvidersResponse: Codable, Equatable, Sendable {
public var providers: [TranscriptionProviderInfo]
public init(providers: [TranscriptionProviderInfo]) {
self.providers = providers
}
}
/// `TranscriptionProviderInfo` (`shared/src/voice.ts`).
public struct TranscriptionProviderInfo: Codable, Equatable, Sendable {
/// `openai | elevenlabs | deepgram | groq | openai-compatible | browser-local`.
public var id: String
public var label: String
/// Subset of `standard` / `realtime`; native dictation uses `standard`.
public var modes: [String]
public init(id: String, label: String, modes: [String]) {
self.id = id
self.label = label
self.modes = modes
}
}