Files
hapi/ios/Packages/HapiKit/Sources/HapiClient/Endpoints/VoiceEndpoints.swift
T
weishu 05d063cc33 feat(ios): composer attachments + dictation (A-M3f)
Attachments (Android B-M3f semantics ported verbatim):
- HapiClient/Attachments/AttachmentPolicy — pure plan matrix (>4 MB
  recompressible image -> 2048 px JPEG q85 with .jpg rename, 50 MB hard
  reject, 192 MB image read cap, 512 px q80 previewUrl data-URL thumbs,
  data-URL parse/round-trip).
- HapiClient/Attachments/ComposerAttachments — upload-on-pick tray over an
  AttachmentUploading seam (APIClient conforms): uploading/ready/failed
  chips, retained payload for retry, remove -> best-effort delete,
  mid-upload removal deletes the orphan on completion, consume() ->
  AttachmentMetadata with JPEG data-URL previewUrl, discardAllDetached +
  deinit orphan cleanup (Android onCleared analogue).
- ChatInteractor: tray ownership, unsettled chips refuse the send with a
  notice, attachments-only sends post empty text, optimistic rows carry
  the metadata, appendDictatedText/postNotice/discardAttachments.
- App: AttachmentPreparer (capped security-scoped reads, ImageIO
  downscale/encode with EXIF transform, HEIC-undecodable fallback),
  PhotosPicker multi (videos via FileRepresentation temp files),
  UIImagePickerController camera capture, fileImporter; composer chip row
  (thumb/spinner/tap-to-retry/remove) with attachment-aware send gating;
  user bubbles upgrade chips to off-main-decoded previewUrl thumbnails
  (web-sent attachments included).

Dictation (Android B-M3ce port):
- HapiProtocol/Models/VoiceApi — TranscriptionResponse,
  TranscriptionProvidersResponse, TranscriptionProviderInfo.
- HapiClient/Endpoints/VoiceEndpoints — GET
  /api/voice/transcription/providers + multipart POST
  /api/voice/transcription (file/provider/mode/language, Android part
  order) over MultipartFormData; DictationTranscribing conformance.
- HapiClient/Voice/DictationController — idle/starting/recording/
  transcribing, transcribed/noProvider/error events, provider memoized
  (first standard-capable entry), appendTranscript port.
- App: AVAudioRecorderDictation (m4a/AAC mono 44.1 kHz 96 kbps, session
  activate/deactivate), mic button + recording chip (elapsed + cancel),
  record-permission request via AVAudioApplication.

Info.plist gains NSMicrophoneUsageDescription; the camera string now
covers attachment capture (modern PhotosPicker needs no photo-library
permission). Tests: policy matrix, tray over the real client with exact
base64 upload bodies + gated in-flight scenarios, dictation controller
suite with fake recorder/transport, voice endpoint request shapes, and
the interactor attachment-send matrix transcribed from the Android VM
tests (wire bodies byte-for-byte).
2026-08-18 11:27:28 +08:00

63 lines
2.2 KiB
Swift

import Foundation
import HapiProtocol
/// Voice transcription (`docs/api/client-contract/rest.md`), added for A-M3f
/// dictation — mirrors the Android `HapiApi` voice methods.
extension APIClient {
/// `GET /api/voice/transcription/providers` — providers whose keys are
/// configured on the hub. Dictation picks the first entry supporting
/// `standard`; an empty list means no transcription provider is
/// configured.
public func transcriptionProviders() async throws -> TranscriptionProvidersResponse {
try await request(.get, "/api/voice/transcription/providers")
}
/// `POST /api/voice/transcription` — the one `multipart/form-data`
/// endpoint: `file` (≤ 25 MB audio), `provider`, `mode`, optional
/// `language` (BCP-47-ish). Field order mirrors the Android client
/// (file, provider, mode, language).
public func transcribeVoice(
audio: Data,
filename: String,
mimeType: String,
provider: String,
mode: String = "standard",
language: String? = nil
) async throws -> TranscriptionResponse {
var form = MultipartFormData()
form.appendFile(fieldName: "file", filename: filename, mimeType: mimeType, data: audio)
form.appendField(name: "provider", value: provider)
form.appendField(name: "mode", value: mode)
if let language {
form.appendField(name: "language", value: language)
}
return try await request(
.post,
"/api/voice/transcription",
rawBody: form.encodedBody(),
contentType: form.contentType
)
}
}
/// `DictationController` transport seam (the Android `HapiDictationApi`
/// adapter collapses to a conformance here).
extension APIClient: DictationTranscribing {
public func transcribe(
audio: Data,
filename: String,
mimeType: String,
provider: String,
language: String?
) async throws -> TranscriptionResponse {
try await transcribeVoice(
audio: audio,
filename: filename,
mimeType: mimeType,
provider: provider,
mode: "standard",
language: language
)
}
}