Files
chanora/crates/chanora_audio/Cargo.toml
T
EdisonJwa 99584fbc1a fix(audio,ios): decouple AudioHandler from VPIO render callback via ring buffer (rc.8+73)
User confirmed at +72 the symptom is 'voice + constant clicks +
choppy fragments'. The diagnostic data conclusively pointed to
iOS VPIO render-callback timing as the cause:

  * frames_changes=60+ per 100 callbacks at cb>=1800
    iOS keeps switching num_frames between 960 and 1104
    on roughly 60% of callbacks
  * peak_out_i16 is sensible (2500-16870, never clipping)
    when AudioHandler returns content
  * input_was_zero=true on most callbacks during active speech
    AudioHandler keeps entering buffering_samples state

The cause: previous render callback called fill_buffer
synchronously every iOS audio thread invocation. With iOS
calling at irregular rates with irregular sizes, AudioHandler's
jitter buffer (sized around 20 ms Opus frames) cannot satisfy
arbitrary-sized requests and falls back to returning silence
(&[] empty slice) on misaligned reads. The silent gaps in the
middle of the output buffer create discontinuities = audible
clicks; the missing-tail content produces choppy fragments.

Fix (architectural): decouple the AudioHandler decoder from the
VPIO render callback via a lock-free SPSC ring buffer.

  Producer (tokio task, 50 Hz):
    every 20 ms:
      fill_buffer(scratch_stereo_f32, 1920 = 20 ms stereo)
      downmix L+R -> mono i16 (960 samples)
      ring_buffer.push_slice(mono_i16)

  Consumer (VPIO render callback, iOS audio thread):
    every callback:
      pop num_frames samples from ring buffer into out
      zero-fill tail on underrun

Why it works:
  * Producer always asks AudioHandler for a stable 20 ms chunk
    (perfectly aligned with internal Opus frame size). No more
    buffering_samples false-triggers.
  * Consumer pulls whatever iOS asks for whenever iOS schedules
    it; ring buffer's 200 ms depth absorbs the callback jitter.
  * This is the standard pattern every production VoIP audio
    engine uses (WebRTC, Discord, FaceTime) to bridge bursty
    Opus decoders to bursty platform audio callbacks.

Implementation:
  * New dep: rtrb 0.3.4 (RustAudio realtime-safe SPSC ring
    buffer, 6.8M downloads, lock-free push/pop with no
    allocation on the audio thread).
  * RING_BUFFER_SAMPLES = 9600 (200 ms mono i16 at 48 kHz).
    Sized for 10x producer ticks of headroom.
  * PRODUCER_TICK_MS = 20 (matches Opus 50 Hz packet rate).
    set_missed_tick_behavior(Skip) to avoid burst catch-up on
    runtime stalls.
  * Producer task spawned in IosVoiceUnit::start, shutdown
    via tokio::oneshot when IosVoiceUnit drops.
  * Render callback is now just: pop into out, zero-fill tail,
    apply mute then gain.
  * Gain applied CONSUMER-side so user volume changes take
    effect within one callback (<= 200 ms latency).
  * Underrun diagnostics: count underrun callbacks + total
    zero-filled samples, log every 100 callbacks.

Threading + safety:
  * rtrb is lock-free SPSC. Audio thread never blocks.
  * Producer can block briefly on Arc<Mutex<AudioHandler>>
    contention with the inbound forwarder (handle_packet), but
    not with the audio thread.
  * Producer task is owned by tokio runtime; explicit shutdown
    channel ensures it exits when the engine stops.

Build verify:
  * Linux host: cargo check clean in 4.06s (downloads rtrb 0.3.4).
  * iOS Mac:   cargo check clean in 2.14s.

Build counter 72 -> 73.
2026-05-17 02:39:42 +08:00

145 lines
7.1 KiB
TOML

[package]
name = "chanora_audio"
description = "Chanora audio subsystem — cpal-based capture/playback, audiopus encode, tsclientlib AudioHandler for decode + jitter buffer + mix. DEC-011, DEC-011.1."
version.workspace = true
edition.workspace = true
rust-version.workspace = true
authors.workspace = true
license.workspace = true
repository.workspace = true
publish.workspace = true
[dependencies]
chanora_protocol = { path = "../chanora_protocol" }
thiserror.workspace = true
tracing.workspace = true
# Cross-platform audio I/O (DEC-011.1).
cpal = "0.16"
# Opus encoder. tsclientlib already pulls this; we depend explicitly so
# this crate can compile against it without going through tsclientlib.
audiopus = "0.3.0-rc.0"
# AudioHandler lives in the tsclientlib crate behind the `audio`
# feature. We import the crate just for the AudioHandler type; the
# Connection type stays inside chanora_protocol.
tsclientlib = { git = "https://github.com/ReSpeak/tsclientlib.git", rev = "04aa2491", default-features = false, features = ["audio"] }
# Force the reqwest TLS backend to native-tls (Security.framework on
# Apple, SChannel on Windows, system OpenSSL on Linux/BSD) instead of
# the rustls + aws-lc-rs combination tsclientlib's `default-tls`
# feature would otherwise pick. aws-lc-sys does not cross-compile
# cleanly to aarch64-apple-ios, and native-tls is the standard
# answer for "TLS that just works on every desktop+mobile platform
# without a vendored C dependency". Cargo's feature unification means
# this single direct dep applies to the transitive
# tsclientlib -> reqwest chain too.
reqwest = { version = "0.13", default-features = false, features = ["charset", "http2", "native-tls"] }
tokio = { version = "1", features = ["sync", "rt", "macros", "time"] }
[target.'cfg(target_os = "ios")'.dependencies]
# Direct CoreAudio AudioUnit access on iOS (DEC-011.x follow-up).
# cpal's iOS backend is unsuitable for VoIP: it opens
# kAudioUnitSubType_RemoteIO with a mono-only output element and no
# control over buffer size / sample rate, AND its AudioUnit stays
# bound to the route present at construction time so user-driven
# `overrideOutputAudioPort` flips do not actually move audio to the
# new transducer. Every production iOS VoIP client (Linphone, Mumble
# iOS, Signal, Jitsi, WebRTC reference) instead drives
# `kAudioUnitSubType_VoiceProcessingIO` (a.k.a. VPIO) directly. VPIO
# is Apple's recommended voice unit; it ships hardware AEC + AGC + NS
# and honours route changes natively because it IS the canonical
# voice unit on iOS. `coreaudio-rs` (RustAudio org, same maintainers
# as `cpal`, 8.6M downloads) gives us a safe wrapper around the
# AudioUnit C API. We use it on iOS only; cpal stays on macOS where
# its CoreAudio backend works well against HAL units.
#
# Default features keep `audio_toolbox` + `core_audio`, both required
# for AudioUnit construction + property access.
coreaudio-rs = "0.14"
# Lock-free SPSC ring buffer to decouple the VPIO render callback
# (consumer, runs on iOS's audio thread with strict realtime
# constraints) from the AudioHandler decoder (producer, runs on a
# tokio worker). Without this, the render callback calls
# AudioHandler::fill_buffer directly under a Mutex, and any time
# the decoder is mid-burst the callback either blocks or gets a
# partial-fill that produces clicks at the partial-fill boundary
# plus choppy fragments from the missing tail. iOS's VPIO calls
# the render callback at irregular intervals (we logged
# `frames_changes=60+ per 100 callbacks` = the buffer size flips
# on ~60% of callbacks); decoupling the two via a stable-rate
# ring buffer is the standard fix used by every production VoIP
# audio engine. `rtrb` 0.3.4 (mgeier, 6.8M downloads) is the
# realtime-safe SPSC ring buffer the Rust audio community
# converged on \u2014 lock-free push / pop with no allocation on
# the audio thread.
rtrb = "0.3"
[target.'cfg(target_os = "android")'.dependencies]
# JNI bindings to flip Android's AudioManager into MODE_IN_COMMUNICATION
# when the voice-comm preset is requested. ndk_context is initialised
# by the bridge crate's android_init shim.
jni = { version = "0.21", default-features = false }
ndk-context = "0.1"
[target.'cfg(target_os = "windows")'.dependencies]
# Real Windows global PTT (SDD-083 / SDD-084): RegisterRawInputDevices
# + WM_INPUT translation backed by a hidden message-only window, and
# SetWindowsHookExW(WH_KEYBOARD_LL / WH_MOUSE_LL) fallback. Both
# require a per-backend OS thread that owns a message pump.
windows = { version = "0.54", features = [
"Win32_Foundation",
"Win32_Graphics_Gdi",
"Win32_System_LibraryLoader",
"Win32_System_Threading",
"Win32_UI_Input",
"Win32_UI_Input_KeyboardAndMouse",
"Win32_UI_WindowsAndMessaging",
] }
[dev-dependencies]
# `test-util` enables `start_paused` / virtual-clock tests used by
# the missed-key-up watchdog unit tests.
tokio = { version = "1", features = ["sync", "rt", "macros", "time", "test-util"] }
# Cross-platform recording Layer for the SDD-090 / DEC-027 privacy
# invariant integration test (`tests/ptt_privacy.rs`).
tracing-subscriber = { version = "0.3", features = ["registry"] }
[target.'cfg(target_os = "linux")'.dependencies]
# GNOME-on-Wayland Global Push-to-Talk uses the freedesktop
# `org.freedesktop.portal.GlobalShortcuts` interface over D-Bus.
# `zbus` is the standard async D-Bus crate; the `tokio` runtime
# selector is mandatory in zbus 5; we share the tokio runtime
# the rest of the audio + core crates already depend on. The
# `blocking-api` feature is retained so the audio-engine
# probe path can do a synchronous portal-version read without
# starting an async runtime; the live session flow uses the
# async surface.
zbus = { version = "5", default-features = false, features = ["tokio", "blocking-api"] }
# Stream / sink utilities for consuming portal signals on the
# async path.
futures-util = { version = "0.3", default-features = false, features = ["std"] }
# Random token bytes for the portal handle_token / session_handle_token
# options. The portal recommends fresh tokens to scope its own
# object paths per call.
rand = "0.8"
# SDL2 audio for Linux. Replaces the cpal capture / playback paths
# on Linux only; cpal stays in use on Windows/macOS. Rationale: the
# cpal Linux backend opens raw ALSA `default`, which on most Arch
# / Fedora / Debian installs routes through `dmix` + `plug` with
# nearest-neighbour resampling and very small period sizes — the
# combination produces audible crackling/popping. SDL2 on the same
# systems routes through PipeWire's PulseAudio compat bridge (or
# real PulseAudio), both of which carry a high-quality resampler
# and a sensible default period. The upstream tsclientlib audio
# example (`tsclientlib/examples/audio_utils/ts_to_audio.rs`) and
# the official Qint client both use SDL2 in exactly this shape;
# this dep brings Chanora in line with that pattern.
#
# `bundled` is OFF deliberately — we link against the system
# libSDL2.so. Arch ships `sdl2-compat`; Debian/Ubuntu ship
# `libsdl2-2.0-0`; Fedora ships `SDL2`. The chanora-flutter Linux
# build documentation lists this as a runtime dependency.
sdl2 = { version = "0.37", default-features = false }