User confirmed at +72 the symptom is 'voice + constant clicks +
choppy fragments'. The diagnostic data conclusively pointed to
iOS VPIO render-callback timing as the cause:
* frames_changes=60+ per 100 callbacks at cb>=1800
iOS keeps switching num_frames between 960 and 1104
on roughly 60% of callbacks
* peak_out_i16 is sensible (2500-16870, never clipping)
when AudioHandler returns content
* input_was_zero=true on most callbacks during active speech
AudioHandler keeps entering buffering_samples state
The cause: previous render callback called fill_buffer
synchronously every iOS audio thread invocation. With iOS
calling at irregular rates with irregular sizes, AudioHandler's
jitter buffer (sized around 20 ms Opus frames) cannot satisfy
arbitrary-sized requests and falls back to returning silence
(&[] empty slice) on misaligned reads. The silent gaps in the
middle of the output buffer create discontinuities = audible
clicks; the missing-tail content produces choppy fragments.
Fix (architectural): decouple the AudioHandler decoder from the
VPIO render callback via a lock-free SPSC ring buffer.
Producer (tokio task, 50 Hz):
every 20 ms:
fill_buffer(scratch_stereo_f32, 1920 = 20 ms stereo)
downmix L+R -> mono i16 (960 samples)
ring_buffer.push_slice(mono_i16)
Consumer (VPIO render callback, iOS audio thread):
every callback:
pop num_frames samples from ring buffer into out
zero-fill tail on underrun
Why it works:
* Producer always asks AudioHandler for a stable 20 ms chunk
(perfectly aligned with internal Opus frame size). No more
buffering_samples false-triggers.
* Consumer pulls whatever iOS asks for whenever iOS schedules
it; ring buffer's 200 ms depth absorbs the callback jitter.
* This is the standard pattern every production VoIP audio
engine uses (WebRTC, Discord, FaceTime) to bridge bursty
Opus decoders to bursty platform audio callbacks.
Implementation:
* New dep: rtrb 0.3.4 (RustAudio realtime-safe SPSC ring
buffer, 6.8M downloads, lock-free push/pop with no
allocation on the audio thread).
* RING_BUFFER_SAMPLES = 9600 (200 ms mono i16 at 48 kHz).
Sized for 10x producer ticks of headroom.
* PRODUCER_TICK_MS = 20 (matches Opus 50 Hz packet rate).
set_missed_tick_behavior(Skip) to avoid burst catch-up on
runtime stalls.
* Producer task spawned in IosVoiceUnit::start, shutdown
via tokio::oneshot when IosVoiceUnit drops.
* Render callback is now just: pop into out, zero-fill tail,
apply mute then gain.
* Gain applied CONSUMER-side so user volume changes take
effect within one callback (<= 200 ms latency).
* Underrun diagnostics: count underrun callbacks + total
zero-filled samples, log every 100 callbacks.
Threading + safety:
* rtrb is lock-free SPSC. Audio thread never blocks.
* Producer can block briefly on Arc<Mutex<AudioHandler>>
contention with the inbound forwarder (handle_packet), but
not with the audio thread.
* Producer task is owned by tokio runtime; explicit shutdown
channel ensures it exits when the engine stops.
Build verify:
* Linux host: cargo check clean in 4.06s (downloads rtrb 0.3.4).
* iOS Mac: cargo check clean in 2.14s.
Build counter 72 -> 73.
145 lines
7.1 KiB
TOML
145 lines
7.1 KiB
TOML
[package]
|
|
name = "chanora_audio"
|
|
description = "Chanora audio subsystem — cpal-based capture/playback, audiopus encode, tsclientlib AudioHandler for decode + jitter buffer + mix. DEC-011, DEC-011.1."
|
|
version.workspace = true
|
|
edition.workspace = true
|
|
rust-version.workspace = true
|
|
authors.workspace = true
|
|
license.workspace = true
|
|
repository.workspace = true
|
|
publish.workspace = true
|
|
|
|
[dependencies]
|
|
chanora_protocol = { path = "../chanora_protocol" }
|
|
thiserror.workspace = true
|
|
tracing.workspace = true
|
|
|
|
# Cross-platform audio I/O (DEC-011.1).
|
|
cpal = "0.16"
|
|
# Opus encoder. tsclientlib already pulls this; we depend explicitly so
|
|
# this crate can compile against it without going through tsclientlib.
|
|
audiopus = "0.3.0-rc.0"
|
|
|
|
# AudioHandler lives in the tsclientlib crate behind the `audio`
|
|
# feature. We import the crate just for the AudioHandler type; the
|
|
# Connection type stays inside chanora_protocol.
|
|
tsclientlib = { git = "https://github.com/ReSpeak/tsclientlib.git", rev = "04aa2491", default-features = false, features = ["audio"] }
|
|
# Force the reqwest TLS backend to native-tls (Security.framework on
|
|
# Apple, SChannel on Windows, system OpenSSL on Linux/BSD) instead of
|
|
# the rustls + aws-lc-rs combination tsclientlib's `default-tls`
|
|
# feature would otherwise pick. aws-lc-sys does not cross-compile
|
|
# cleanly to aarch64-apple-ios, and native-tls is the standard
|
|
# answer for "TLS that just works on every desktop+mobile platform
|
|
# without a vendored C dependency". Cargo's feature unification means
|
|
# this single direct dep applies to the transitive
|
|
# tsclientlib -> reqwest chain too.
|
|
reqwest = { version = "0.13", default-features = false, features = ["charset", "http2", "native-tls"] }
|
|
|
|
tokio = { version = "1", features = ["sync", "rt", "macros", "time"] }
|
|
|
|
[target.'cfg(target_os = "ios")'.dependencies]
|
|
# Direct CoreAudio AudioUnit access on iOS (DEC-011.x follow-up).
|
|
# cpal's iOS backend is unsuitable for VoIP: it opens
|
|
# kAudioUnitSubType_RemoteIO with a mono-only output element and no
|
|
# control over buffer size / sample rate, AND its AudioUnit stays
|
|
# bound to the route present at construction time so user-driven
|
|
# `overrideOutputAudioPort` flips do not actually move audio to the
|
|
# new transducer. Every production iOS VoIP client (Linphone, Mumble
|
|
# iOS, Signal, Jitsi, WebRTC reference) instead drives
|
|
# `kAudioUnitSubType_VoiceProcessingIO` (a.k.a. VPIO) directly. VPIO
|
|
# is Apple's recommended voice unit; it ships hardware AEC + AGC + NS
|
|
# and honours route changes natively because it IS the canonical
|
|
# voice unit on iOS. `coreaudio-rs` (RustAudio org, same maintainers
|
|
# as `cpal`, 8.6M downloads) gives us a safe wrapper around the
|
|
# AudioUnit C API. We use it on iOS only; cpal stays on macOS where
|
|
# its CoreAudio backend works well against HAL units.
|
|
#
|
|
# Default features keep `audio_toolbox` + `core_audio`, both required
|
|
# for AudioUnit construction + property access.
|
|
coreaudio-rs = "0.14"
|
|
|
|
# Lock-free SPSC ring buffer to decouple the VPIO render callback
|
|
# (consumer, runs on iOS's audio thread with strict realtime
|
|
# constraints) from the AudioHandler decoder (producer, runs on a
|
|
# tokio worker). Without this, the render callback calls
|
|
# AudioHandler::fill_buffer directly under a Mutex, and any time
|
|
# the decoder is mid-burst the callback either blocks or gets a
|
|
# partial-fill that produces clicks at the partial-fill boundary
|
|
# plus choppy fragments from the missing tail. iOS's VPIO calls
|
|
# the render callback at irregular intervals (we logged
|
|
# `frames_changes=60+ per 100 callbacks` = the buffer size flips
|
|
# on ~60% of callbacks); decoupling the two via a stable-rate
|
|
# ring buffer is the standard fix used by every production VoIP
|
|
# audio engine. `rtrb` 0.3.4 (mgeier, 6.8M downloads) is the
|
|
# realtime-safe SPSC ring buffer the Rust audio community
|
|
# converged on \u2014 lock-free push / pop with no allocation on
|
|
# the audio thread.
|
|
rtrb = "0.3"
|
|
|
|
[target.'cfg(target_os = "android")'.dependencies]
|
|
# JNI bindings to flip Android's AudioManager into MODE_IN_COMMUNICATION
|
|
# when the voice-comm preset is requested. ndk_context is initialised
|
|
# by the bridge crate's android_init shim.
|
|
jni = { version = "0.21", default-features = false }
|
|
ndk-context = "0.1"
|
|
|
|
[target.'cfg(target_os = "windows")'.dependencies]
|
|
# Real Windows global PTT (SDD-083 / SDD-084): RegisterRawInputDevices
|
|
# + WM_INPUT translation backed by a hidden message-only window, and
|
|
# SetWindowsHookExW(WH_KEYBOARD_LL / WH_MOUSE_LL) fallback. Both
|
|
# require a per-backend OS thread that owns a message pump.
|
|
windows = { version = "0.54", features = [
|
|
"Win32_Foundation",
|
|
"Win32_Graphics_Gdi",
|
|
"Win32_System_LibraryLoader",
|
|
"Win32_System_Threading",
|
|
"Win32_UI_Input",
|
|
"Win32_UI_Input_KeyboardAndMouse",
|
|
"Win32_UI_WindowsAndMessaging",
|
|
] }
|
|
|
|
[dev-dependencies]
|
|
# `test-util` enables `start_paused` / virtual-clock tests used by
|
|
# the missed-key-up watchdog unit tests.
|
|
tokio = { version = "1", features = ["sync", "rt", "macros", "time", "test-util"] }
|
|
# Cross-platform recording Layer for the SDD-090 / DEC-027 privacy
|
|
# invariant integration test (`tests/ptt_privacy.rs`).
|
|
tracing-subscriber = { version = "0.3", features = ["registry"] }
|
|
|
|
[target.'cfg(target_os = "linux")'.dependencies]
|
|
# GNOME-on-Wayland Global Push-to-Talk uses the freedesktop
|
|
# `org.freedesktop.portal.GlobalShortcuts` interface over D-Bus.
|
|
# `zbus` is the standard async D-Bus crate; the `tokio` runtime
|
|
# selector is mandatory in zbus 5; we share the tokio runtime
|
|
# the rest of the audio + core crates already depend on. The
|
|
# `blocking-api` feature is retained so the audio-engine
|
|
# probe path can do a synchronous portal-version read without
|
|
# starting an async runtime; the live session flow uses the
|
|
# async surface.
|
|
zbus = { version = "5", default-features = false, features = ["tokio", "blocking-api"] }
|
|
# Stream / sink utilities for consuming portal signals on the
|
|
# async path.
|
|
futures-util = { version = "0.3", default-features = false, features = ["std"] }
|
|
# Random token bytes for the portal handle_token / session_handle_token
|
|
# options. The portal recommends fresh tokens to scope its own
|
|
# object paths per call.
|
|
rand = "0.8"
|
|
# SDL2 audio for Linux. Replaces the cpal capture / playback paths
|
|
# on Linux only; cpal stays in use on Windows/macOS. Rationale: the
|
|
# cpal Linux backend opens raw ALSA `default`, which on most Arch
|
|
# / Fedora / Debian installs routes through `dmix` + `plug` with
|
|
# nearest-neighbour resampling and very small period sizes — the
|
|
# combination produces audible crackling/popping. SDL2 on the same
|
|
# systems routes through PipeWire's PulseAudio compat bridge (or
|
|
# real PulseAudio), both of which carry a high-quality resampler
|
|
# and a sensible default period. The upstream tsclientlib audio
|
|
# example (`tsclientlib/examples/audio_utils/ts_to_audio.rs`) and
|
|
# the official Qint client both use SDL2 in exactly this shape;
|
|
# this dep brings Chanora in line with that pattern.
|
|
#
|
|
# `bundled` is OFF deliberately — we link against the system
|
|
# libSDL2.so. Arch ships `sdl2-compat`; Debian/Ubuntu ship
|
|
# `libsdl2-2.0-0`; Fedora ships `SDL2`. The chanora-flutter Linux
|
|
# build documentation lists this as a runtime dependency.
|
|
sdl2 = { version = "0.37", default-features = false }
|