PR #37 inlined VadWorkerPolicy into the desktop capture path but missed
the iOS files. Instead of restoring the deleted types, inline the same
direct if-let-Some pattern into ios_voice_unit.rs (the only remaining
iOS backend) and permanently remove VadWorkerPolicy and
callback_vad_worker_policy from vad/mod.rs.
iOS uses Option<AppleCoreMlVadWorker> with &mut self (no Mutex), so the
inline is simpler than the desktop's try_lock pattern.
Adopt a call-scoped VoIP audio session lifecycle so other apps' audio is
not stopped while Chanora is idle and the in-call session does not get
clobbered by media-server resets unrelated to voice.
AppDelegate.swift
- Set .ambient + .mixWithOthers as the idle baseline so the app does not
hold a VoiceChat session when no call is active.
- Switch to .playAndRecord + .voiceChat + .mixWithOthers + .duckOthers
on demand via the new chanora/ios_audio_session MethodChannel, and
revert to .ambient on deactivate.
- Gate the media-services-reset rebuild on voiceSessionActive so a
stray reset during idle no longer reactivates VoiceChat.
ios_audio_session_controller.dart (new)
- Thin Dart wrapper around chanora/ios_audio_session with activate/
deactivate; no-op on non-iOS; swallows PlatformException to keep
audio start/stop resilient to platform-side races.
main.dart
- Activate the iOS audio session on BridgeEvent_AudioStarted, deactivate
on BridgeEvent_AudioStopped, fire-and-forget via unawaited().
ios_voice_unit.rs
- Add the 10 local bindings required by the render-callback closure
preamble (wav_recorder_for_render, render_recorder_active,
render_ref_len, render_ref_accum, cb_count, last_num_frames,
num_frames_changes, callbacks_with_audio, callbacks_with_silence)
so the iOS target compiles cleanly with the new lifecycle wiring.
Tests
- 5 unit tests in test/services/ios_audio_session_controller_test.dart
cover activate/deactivate on iOS, no-op on non-iOS, and graceful
PlatformException handling.
Docs
- SRS SRS-110 expanded to cover the call-scoped lifecycle invariant.
- SysDes mobile-voice row updated to reflect the MethodChannel and
.ambient idle baseline.
- implementation-status-2026-05-28 voiceChat row flipped to done.
Verification
- flutter analyze: No issues found (2.8s)
- flutter test: 195 passed / 2 skipped / 0 failed
- cargo build -p chanora_audio --target aarch64-apple-ios: clean
- cargo build -p chanora_audio (macOS host): clean
Device QA matrix (Spotify-keeps-playing-while-idle, mix-during-call,
revert-on-call-end, media-services-reset-during-idle) remains pending
on physical hardware.
- ios_voice_unit.rs: add producer_shutdown AtomicBool flag (macOS only).
The macOS start path spawns a tokio producer task that holds clones
of Arc<Mutex<AudioHandler>>, Arc<ArrayQueue<f32>>, and the output
gain/muted atomics, then loops on a 20 ms tokio interval. Without
a shutdown signal the task runs forever on engine stop/restart and
leaks all four Arcs every cycle. Drop now stores 'true' on the
flag; the producer checks it at the top of each tick and exits,
releasing its clones within at most one 20 ms tick.
- macos/Runner/Release.entitlements: strengthen the existing
justification comment for com.apple.security.network.server.
Document the specific failure mode (tokio::net::UdpSocket::bind
-> sandbox 'network-outbound deny' -> EPERM) and explain why
network.client alone does not cover bind()-then-sendto. The
entitlement is required, not over-broad.
cargo check (host + aarch64-apple-darwin): clean
cargo test -p chanora_audio --lib: 133 passed
flutter test: 186 passed, 2 skipped
dart analyze: clean
The PR #27 refactor moved the iOS render callback to a direct-fill
path with its own peak_i16 == 0 check, dropping the !muted gate
that PR #28 added to the pre-refactor callback. Without this
gate, every muted callback fires a false-positive
increment_output_underrun() because the downmix helper writes
silence (peak_i16 = 0) by design when muted.
Re-apply PR #28's gate to the refactored iOS path so this PR does
not silently reintroduce the bug PR #28 was opened to fix.
Verified:
- cargo test -p chanora_audio --lib: 133 passed, 0 failed
- cargo build -p chanora_audio --target aarch64-apple-ios: clean
voice_render.rs: new limit_peak_inplace helper (single-pass, allocation-free peak scaler) with 4 unit tests. Applied in both the macOS and iOS render callbacks before the i16 conversion to prevent hard clipping on multi-client mixes that sum past 0 dBFS. Threshold 0.99 keeps the limiter transparent for normal voice levels (sub-millisecond per-frame latency at 48 kHz; pumping risk negligible for speech).
ios_voice_unit.rs: limit_peak_inplace call sites added to the macOS producer/ring render callback (per-frame, 2-element stack array) and the iOS direct-fill_buffer render callback (per-callback, on the scratch_stereo buffer); the downmix helper's hard-clamp is kept as defense-in-depth and is not expected to engage in the integrated flow.
- ios_voice_unit.rs:856-868: Revert iOS render callback from blocking
lock() back to try_lock() with silence-on-contention (WouldBlock
branch increments callback_xrun stat and returns the pre-zeroed
scratch buffer). Blocking lock() inside the CoreAudio HAL render
callback can stall the realtime IO thread when the decode task on
engine.rs:1309 holds the same AudioHandler Mutex, re-introducing
the underrun pattern this codebase already fixed elsewhere.
- ios_voice_unit.rs:782: Preallocate scratch_stereo to Apple's VPIO
MaximumFramesPerSlice (4096 frames * 2 channels = 8192 f32) at
setup time, so the realtime render callback never grows the Vec
via resize(). The defensive 'len() < needed' branch is kept for
the (impossible) case that the audio unit later raises max frames.
- main.dart:52-56: Gate _showAudioDebugOverlay behind kDebugMode &&
_isMacOS so the internal audio stats panel does not ship in
release builds. kDebugMode is a Dart compile-time const, so the
overlay subtree is tree-shaken out of release/profile binaries.
- Rebuilt macOS chanora_bridge.framework binary (universal arm64 +
x86_64) from the fixed source with the new CARGO_PROFILE_RELEASE_*
env vars (DWARF preserved for dsymutil). install_name patched back
to @rpath/chanora_bridge.framework/Versions/A/chanora_bridge.
Verified:
cargo test -p chanora_audio --lib: 129 passed, 0 failed
cargo check -p chanora_audio --target aarch64-apple-ios: clean
dart analyze lib/main.dart: no issues
xcrun lipo -archs: x86_64 arm64
xcrun otool -D: @rpath install_name preserved
macOS realtime audio was suffering buffer underruns on CoreAudio's VPIO
output callback. Root cause was twofold: AudioHandler was decoded under
a Mutex held across the realtime callback, and the render path
hard-coded mono i16 output regardless of the channel count the
callback actually exposed (CoreAudio occasionally hands the callback
stereo or quad output buffers, in which case writing only every Nth
sample produced silence + clicks).
This change brings macOS in line with the lock-free Android audio
architecture introduced for output stutter elimination:
* chanora_audio: AudioPacket / AudioCommand / AudioEventQueue
(previously gated to `target_os = "android"`) are now compiled on
macOS too. The decode loop in AudioEngine pushes inbound packets
into the queue; the VPIO render callback owns AudioHandler outright
and drains the queue, so the realtime thread never blocks on a
cross-thread mutex. set_client_volume also routes through the
command queue on macOS instead of locking the handler.
* voice_render.rs: new downmix_stereo_f32_to_interleaved_i16 helper
downmixes stereo f32 from AudioHandler to mono i16 and replicates
that mono sample across every output channel the callback exposes.
The existing downmix_stereo_f32_to_mono_i16 helper is retained for
iOS, where VPIO is reliably configured for single-channel output
via the AudioUnit stream format we pin at unit-create time.
Compile-gated to ios + test so the macos build doesn't warn on
dead code.
* ios_voice_unit.rs: render callback reads data.channels from the
args struct and forwards it to the new interleaved helper, so the
macOS path tolerates whatever channel count CoreAudio assigns. A
level decimation counter avoids running sqrt+log10 on every
callback (~93 Hz) when the Flutter consumer only reads at 30 Hz;
same regression class as the capture-side fix already in engine.rs.
* mobile_voice_backend.rs: VoiceAudioParams now carries
event_producer on macOS, and the AudioHandler is no longer wrapped
in Arc<Mutex<…>> on macOS because ownership moves into the render
callback. iOS keeps Arc<Mutex<…>> because its callback design
shares the handler with the decode task.
* lib.rs: audio_event_queue module is now compiled on macOS in
addition to android.
apps/chanora_flutter/lib/main.dart wraps the home tree in a Stack and
overlays AudioDebugStatsPanel on macOS so the live engine counters
(callback rate, drift, queue depth) used to diagnose the underrun are
visible while iterating on this code. iOS and other platforms are
unaffected.
apps/chanora_flutter/macos/Frameworks/chanora_bridge.framework binary
is rebuilt with these changes so flutter run on macOS picks up the new
realtime path without requiring developers to rebuild the Rust crate
locally. cargo check -p chanora_audio passes on macOS host.
ios_voice_unit.rs:1033: change the condition from `mix_stats.peak_i16 == 0` to `mix_stats.peak_i16 == 0 && !muted`. When the user mutes the channel via output_muted, the downmix helper fills the output buffer with silence (peak = 0), which previously falsely incremented the output_underrun counter. The mute toggle is intentional silence, not a real underrun.
Note: this does not address the separate false positive where peak_i16 == 0 with output unmuted but no audio incoming (e.g., just joined a channel with no remote speaking). A complete fix would require tracking whether the audio handler actually produced data; deferred to a follow-up.
* fix(diag): record bridge events in diagnostics page
Connection lost/reconnecting/disconnected and iOS audio interruption
events now appear in the diagnostics dialog alongside existing error
snackbar entries.
* fix(diag): reduce audio callback sample verbosity to debug
The render callback diagnostic sample logged every 100 callbacks
(~2s) at INFO level, flooding the 256-entry release log buffer and
pushing out useful events. Changed to DEBUG so it only appears in
debug builds with the larger 4096-entry buffer.
* fix(diag): add timestamps to Rust diagnostic log entries
The ring-buffer architecture (rc.8+73..+74) was making playback
strictly worse. Diagnostic data at +74 conclusively showed:
* Producer task ran perfectly at 50 Hz (250 ticks per 5 s).
* AudioHandler returned silence on 65-84% of fill_buffer calls
even when window_peak_f32 reached 0.98 (full-scale audio).
* Ring buffer never accumulated beyond 30 ms because consumer
(VPIO render callback at 43.5 Hz, ~1440 samples per call)
drained samples faster than the 50 Hz producer could push
them, in net effect.
The producer drained AudioHandler at 50 Hz \u2014 slightly faster
than iOS VPIO actually consumes audio. Each fill_buffer call
asked for 20 ms but adjacent Opus packets hadn't arrived yet, so
fill_buffer returned mostly silence. Linux/SDL's same pattern
works because SDL calls fill_buffer at EXACTLY the device
callback rate (50 Hz = 20 ms per buffer); the rates match.
Fix: revert to direct fill_buffer call from the render callback
(the SDL pattern in tsclientlib's own reference example at
tsclientlib/examples/audio_utils/ts_to_audio.rs). The render
callback now:
1. Resizes scratch_stereo Vec to 2 * num_frames f32 if needed
2. Zeros the live slice (fill_buffer is additive, not clearing)
3. Locks AudioHandler, calls fill_buffer(scratch_stereo)
4. Downmixes L+R -> mono i16 with master gain into out[]
5. Applies output_muted bypass
6. Tracks peak_out + audio/silence ratios for diagnostic
The closure owns scratch_stereo across callbacks for stable
allocation. Same memory model as Linux/SDL.
Removed:
* tokio::spawn producer task
* rtrb dep + RingBuffer<i16> + Producer/Consumer split
* tokio::sync::oneshot shutdown channel
* producer_shutdown_tx field on IosVoiceUnit struct
* RING_BUFFER_SAMPLES / PRODUCER_TICK_MS constants
* Producer-side diagnostic counters
Diagnostic kept: cb / num_frames / frames_changes /
callbacks_with_audio / callbacks_with_silence / peak_out_i16 /
gain. Logged every 100 callbacks.
The choppy / clicks symptom is independent of the buffer
architecture \u2014 it's whatever AudioHandler is doing on iOS
that's different from Linux. Next investigation step is to
either (a) switch from VPIO to RemoteIO unit (lose Apple's
voice processing entirely), or (b) understand why AudioHandler
returns silence so often on iOS-arrival packet timing patterns.
Build counter 74 -> 75.
External reviewer correctly identified that the +73 ring-buffer
commit didn't fix the symptom but the architecture is still
right. We need to distinguish two possible causes:
(a) Producer task isn't running (or running too rarely) so
ring stays underfilled.
(b) Producer IS running but fill_buffer returns zeros most of
the time (AudioHandler stuck in buffering_samples state
or no packets reaching it).
The +73 render-side diagnostic was insufficient: we logged
underruns + peak_out_i16 but not what the producer was
actually pushing. This commit adds producer-side metrics
rolled up every 5 s (250 ticks at 20 ms):
Producer task:
producer_ticks : timer firings (= ~250 per 5 s window;
fewer = tokio scheduler stalled)
produced_chunks : pushes into ring (= ticks - drops)
fill_buffer_calls : AudioHandler queries
fill_buffer_zero_returns : ticks where scratch came back all
zeros (no decoded content to play)
ring_full_drops : ticks where ring was full and we
skipped the push
ring_min/max_samples : depth envelope across window
ring_min/max_ms : same in milliseconds
window_peak_f32 : max scratch sample across window
window_rms_f32 : RMS of all scratch samples across
window
Render callback (per-100-callback as before, plus new fields):
ring_avail_before : Consumer::slots() before this read
(= how many samples were sitting in
the ring at callback entry)
read_frames : samples successfully popped
zero_filled : samples zero-filled because ring
was empty (= num_frames - read_frames)
underruns / underrun_samples / peak_out_i16 / clip_count_i16
: as before
Reading the next iteration's log:
If producer_ticks << 250 per 5 s window:
tokio scheduler isn't running the task fast enough.
Move producer to its own dedicated runtime, or use
std::thread + std::sync::mpsc + std::thread::sleep
instead of tokio.
If producer_ticks ~= 250 AND fill_buffer_zero_returns is
high (most ticks return silence):
AudioHandler isn't decoding packets fast enough OR is
stuck buffering. Bug is upstream in protocol layer
packet delivery or AudioHandler's jitter state machine.
The ring buffer architecture cannot fix this.
If producer_ticks ~= 250 AND fill_buffer_zero_returns is
low AND ring_min_ms stays >100ms AND underruns are low BUT
consumer's peak_out_i16 is still 0:
Something is wrong between push and pop. Lock-free
ring corruption, or wrong stride.
Pure diagnostic. No behavioural change beyond the logging.
Producer scratch envelope scan is O(scratch.len()) = 1920
samples per 20 ms tick = ~96k iterations/sec on the audio
producer thread \u2014 negligible CPU.
Build counter 73 -> 74.
User confirmed at +72 the symptom is 'voice + constant clicks +
choppy fragments'. The diagnostic data conclusively pointed to
iOS VPIO render-callback timing as the cause:
* frames_changes=60+ per 100 callbacks at cb>=1800
iOS keeps switching num_frames between 960 and 1104
on roughly 60% of callbacks
* peak_out_i16 is sensible (2500-16870, never clipping)
when AudioHandler returns content
* input_was_zero=true on most callbacks during active speech
AudioHandler keeps entering buffering_samples state
The cause: previous render callback called fill_buffer
synchronously every iOS audio thread invocation. With iOS
calling at irregular rates with irregular sizes, AudioHandler's
jitter buffer (sized around 20 ms Opus frames) cannot satisfy
arbitrary-sized requests and falls back to returning silence
(&[] empty slice) on misaligned reads. The silent gaps in the
middle of the output buffer create discontinuities = audible
clicks; the missing-tail content produces choppy fragments.
Fix (architectural): decouple the AudioHandler decoder from the
VPIO render callback via a lock-free SPSC ring buffer.
Producer (tokio task, 50 Hz):
every 20 ms:
fill_buffer(scratch_stereo_f32, 1920 = 20 ms stereo)
downmix L+R -> mono i16 (960 samples)
ring_buffer.push_slice(mono_i16)
Consumer (VPIO render callback, iOS audio thread):
every callback:
pop num_frames samples from ring buffer into out
zero-fill tail on underrun
Why it works:
* Producer always asks AudioHandler for a stable 20 ms chunk
(perfectly aligned with internal Opus frame size). No more
buffering_samples false-triggers.
* Consumer pulls whatever iOS asks for whenever iOS schedules
it; ring buffer's 200 ms depth absorbs the callback jitter.
* This is the standard pattern every production VoIP audio
engine uses (WebRTC, Discord, FaceTime) to bridge bursty
Opus decoders to bursty platform audio callbacks.
Implementation:
* New dep: rtrb 0.3.4 (RustAudio realtime-safe SPSC ring
buffer, 6.8M downloads, lock-free push/pop with no
allocation on the audio thread).
* RING_BUFFER_SAMPLES = 9600 (200 ms mono i16 at 48 kHz).
Sized for 10x producer ticks of headroom.
* PRODUCER_TICK_MS = 20 (matches Opus 50 Hz packet rate).
set_missed_tick_behavior(Skip) to avoid burst catch-up on
runtime stalls.
* Producer task spawned in IosVoiceUnit::start, shutdown
via tokio::oneshot when IosVoiceUnit drops.
* Render callback is now just: pop into out, zero-fill tail,
apply mute then gain.
* Gain applied CONSUMER-side so user volume changes take
effect within one callback (<= 200 ms latency).
* Underrun diagnostics: count underrun callbacks + total
zero-filled samples, log every 100 callbacks.
Threading + safety:
* rtrb is lock-free SPSC. Audio thread never blocks.
* Producer can block briefly on Arc<Mutex<AudioHandler>>
contention with the inbound forwarder (handle_packet), but
not with the audio thread.
* Producer task is owned by tokio runtime; explicit shutdown
channel ensures it exits when the engine stops.
Build verify:
* Linux host: cargo check clean in 4.06s (downloads rtrb 0.3.4).
* iOS Mac: cargo check clean in 2.14s.
Build counter 72 -> 73.
External code review pushed back on the 'iPhone speaker hardware
distortion' hypothesis and pointed out we need more than just
peak measurements. The reviewer's checklist:
* peak_i16
* rms_i16
* num_clipped_samples (abs >= 32767)
* zero_fill_count / underrun_count
* callback_frame_count variability
* decoded_packet_duration_ms
* input_was_zero
* actual ASBD / actual sample rate
The format diagnostics at +71 already showed iOS honoured
48 kHz Int16 mono on both buses and that .default mode +
.defaultToSpeaker routed to the speaker correctly with
outputVolume=0.45. So format + route are confirmed correct.
The remaining mystery is WHY 'loud but distorted' \u2014 we need
sample-level metrics to isolate where in the pipeline the
breakage occurs.
This commit instruments the VPIO render callback with:
* num_frames + frames_changes : detects iOS re-negotiating
buffer size between callbacks
(which would imply jitter the
fixed scratch_stereo Vec can't
absorb cleanly).
* peak_stereo + rms_stereo : characterises AudioHandler's
output BEFORE our downmix.
Distinguishes 'real audio
arriving' from 'silence'.
* peak_out_i16 + clip_count : measures what we hand VPIO.
clip_count > 0 means we're
clipping at our boundary even
with gain=1.0 \u2014 indicates
upstream is over-driven.
* input_was_zero : explicit silence/no-talker
indicator separate from peak=0
which could mean tiny content
rounded to 0.
Reviewer's preferred diagnostic path is to dump PCM to file
and play with ffplay externally; that's iOS-impractical
without a shared filesystem path the user can extract via
Files.app. Instead we sample the same metrics in-callback at
~2 Hz which gives us the same information at run time.
Pure diagnostic. No behavioural change. Counters live in the
FnMut closure so the audio thread cost is one branch +
counter increment per callback, plus a one-pass RMS sum +
peak scan every 100 callbacks.
Build counter 71 -> 72.
Reviewer also recommended a headphone test in parallel \u2014
that will be done by the user (out-of-band) at the next
test cycle to determine whether the symptom changes when
audio leaves the speaker path.
Per external review (helpful checklist from ChatGPT-style analysis
pointing out we never verified that iOS actually accepted our
preferred sample rate / channels / format): preferredSampleRate
and preferredIOBufferDuration are HINTS, not guarantees. iOS may
substitute its own values if the hardware can't satisfy our
preference. If VPIO is running at 44.1 kHz Float32 stereo while
our render callback writes 48 kHz Int16 mono into the buffer,
the symptoms would match what user reports (broken playback,
pitch shifted, severe distortion) and our previous diagnostics
wouldn't catch it because they only sampled signal-level metrics.
This commit adds two diagnostic emissions to verify:
1. AppDelegate.swift::activateAudioSession: after setActive
succeeds, log the ACTUAL session state \u2014 category, mode,
sampleRate, ioBufferDuration, current route (inputs +
outputs), outputVolume. Lets us see whether iOS honoured our
.default + .defaultToSpeaker setup and which physical route
it picked at launch.
2. ios_voice_unit.rs::IosVoiceUnit::start: after unit.start()
succeeds, log the actual OUTPUT and INPUT stream formats
VPIO accepted (sample_rate, channels, sample_format, flags).
If these differ from our requested 48 kHz Int16 mono, we
have a format-substitution problem.
Three possible outcomes from the next test:
* Both diagnostics confirm 48 kHz Int16 mono on both buses and
the session sampleRate=48000 -> format is correct; the
playback breakage is somewhere else (e.g. AudioHandler
jitter buffer behaviour, route binding, or hardware mixer).
* Session sampleRate != 48000 -> we need to insert a sample
rate converter or pin AVAudioSession's
setPreferredSampleRate(48000) explicitly in Swift before
setActive.
* VPIO substituted Float32 for our Int16 request -> our render
callback is writing i16 magnitudes into a Float32 buffer
which would explain the distortion. Fix: write Float32
directly using data::Interleaved<f32> instead of i16.
Build counter 70 -> 71. Pure diagnostic; no behavioural
change.
User report after the 8x boost commit (e85a6d3): playback STILL
broken, but now the diagnostic clearly shows the actual problem.
Render-callback peak_out_i16 SATURATES at 32767 on speech peaks
(cb=600, 1100, 1200, 2400) because the 8x boost amplifies an
already-loud signal into hard clipping. Quiet content reaches
audible level but loud peaks are catastrophically distorted.
The 8x boost was treating the wrong cause.
Real root cause (researched online after user prompted: 'this is
iOS a popular platform, there must be solutions'): iOS has TWO
independent audio channels:
In-call channel (.voiceChat / .videoChat modes)
* Routes through the phone-call audio path.
* Aggressively ducks non-voice content to the earpiece.
* Volume controlled by a separate in-call hardware
register, not the side buttons when not actively on a
phone call.
Media channel (.default mode)
* Routes through the standard media playback path.
* No automatic ducking.
* Volume controlled by the side volume buttons normally.
With AVAudioSession mode .voiceChat, iOS sends our output
through the in-call channel which plays at 'earpiece-level'
loudness on the speaker too. Signal is technically present but
buried under the speaker's noise floor. With mode .default +
.defaultToSpeaker option, output routes via media channel and
plays at normal loudness.
Both Twilio (video-quickstart-ios) and Daily.co (patched WebRTC
module) document the same workaround and use VPIO for AEC while
keeping the session mode at .default for loud playback:
github.com/twilio/video-quickstart-ios/issues/522
stackoverflow.com/questions/79834998 (Daily.co)
The user also noticed 'tx/rx almost no changes even receiving
packages' \u2014 likely a misinterpretation of the frames counter
not advancing as fast as expected during quiet voice; AudioHandler
returns silence when its jitter buffer is in buffering_samples
state which doesn't fire 'decode failed' but also doesn't
increment frames_received. The real issue is still the playback
ducking; the counter behaviour is a downstream symptom.
Changes:
1. AppDelegate.swift: AVAudioSession mode .voiceChat -> .default
with options [.defaultToSpeaker, .allowBluetoothHFP,
.allowBluetoothA2DP]. VPIO continues to do its job (AEC, NS,
AGC on the mic side); only the playback routing changes.
The earlier 'speaker selector silent under .default' bug
does NOT apply because we no longer use cpal RemoteIO \u2014
VPIO honours overrideOutputAudioPort under any mode.
2. ios_voice_unit.rs: revert the 8x output boost from e85a6d3.
With media-channel routing, signal levels are correct and
no software amplification is needed. Render callback restored
to plain (l+r)*0.5*gain downmix.
3. ios_voice_unit.rs: revert the BypassVoiceProcessing toggle
from c16318c. The VPIO chain stays enabled so we keep
capture-side AEC/AGC/NS for free \u2014 the playback breakage
it was trying to fix was the wrong layer all along.
4. ios_voice_unit.rs: drop the diagnostic render-callback log
line. Production-clean code; can be re-enabled by reverting
the diff in the closure if future debugging needs it.
Build counter 69 -> 70.
User-pasted log at +66 (https://pb.hit.moe/q8heratf.txt) shows
conclusive data over a 110-second continuous talker session:
Average peak_stereo_f32: ~0.005-0.010
Loud peak (one moment): ~0.234
peak_out_i16: ~150-300 (out of 32767)
The signal arriving at our render callback from
AudioHandler::fill_buffer is consistently at -40 dB FS for
normal human speech. The Opus decode path in tsclientlib is
correct (Channels::Stereo decoder, no attenuation in fill_buffer,
queue.volume defaults to 1.0). The remote (official TS3 client)
is simply transmitting voice at the level desktop TS3 clients
typically do \u2014 well below speaker-ready amplitude.
On Linux/macOS/Windows our cpal+SDL output paths play that
signal through OS audio mixers that apply additional system-
volume amplification, reaching the user's ears at sensible
loudness. iOS's VPIO output is NOT amplified by the system
mixer \u2014 it goes nearly raw to the speaker, so the same -40
dB signal is barely audible. Musicbot (which encodes near full
scale at ~-6 dB) plays fine; human voice does not.
Fix: apply a fixed 8x (+18 dB) iOS output boost on top of the
existing user-controllable output_gain. A -40 dB signal becomes
-22 dB (normal speakerphone level). User's volume slider
continues to function in a useful 0-2x range on top.
effective_gain = user_gain * IOS_OUTPUT_BOOST
Hard-clip at \u00b11.0 in the mono downmix prevents loud signals
(musicbot at peak 0.5 -> 4.0 -> clamped to 1.0) from
overflowing i16 wrap-around. Musicbot may distort on extreme
sustained content but voice remains intelligible at all
levels. Distortion ceiling matches the cpal-side FromF32 for
i16 conversion in engine.rs.
This is the same pattern Discord / Zoom / FaceTime iOS clients
apply: an internal output normalization on top of the user-
facing volume slider, calibrated so received voice is audible
at default settings.
Build counter 68 -> 69.
User report at +67 (.voiceChat mode): playback still 'broken'.
Even with VPIO's Apple-documented session-mode pairing,
its output-side gating chain (echo subtraction + adaptive
noise suppression) chops quiet inter-phoneme content of human
voice. Musicbot signal (loud, ~continuous) survives because it
stays above the gating threshold; speech does not.
Fix: set kAUVoiceIOProperty_BypassVoiceProcessing = 1 on the
unit immediately after EnableIO (before stream format / callbacks
/ initialize). This disables ALL VPIO voice processing \u2014 the
unit becomes effectively a vanilla RemoteIO with mic + speaker
buses. Raw samples pass through both directions.
Trade-off:
* Lost: Apple's hardware AEC + AGC + NS on the mic path. User
reports current capture is clean already, suggesting their
test environment (headset? non-speakerphone?) doesn't need
AEC. If echo loops back when speakerphone is engaged, we'll
re-evaluate \u2014 either re-enable VPIO selectively for
echo-prone routes or ship software AEC (DEC-007).
* Gained: playback is no longer gated. Quiet inter-phoneme
speech content reaches the speaker.
Property setter:
* Constant: kAUVoiceIOProperty_BypassVoiceProcessing = 2100
* Scope: Global, Element: Input (1) per WebRTC's reference iOS
ADM (voice_processing_audio_unit.mm).
* Value: u32 = 1 (= bypass).
* Soft-fail with warn log if the property is rejected on an
exotic iOS version (the unit still works, just with VPIO
defaults).
Build counter 67 -> 68.
User reports capture-side audio (mic -> remote) is clean but
local playback (remote -> speaker via VPIO render callback) is
'broken and poor' at +65. The pipeline appears correct on paper:
fill_buffer -> downmix (L+R)*0.5 -> gain -> clamp -> i16 -> VPIO.
No errors logged. To stop guessing, add structured logging
inside the render callback so the next test cycle yields data
about what's actually flowing through.
Diagnostic emitted every 100th callback (~2 s at iOS's typical
20-50 Hz callback rate):
ios VPIO render callback diagnostic sample
cb=<counter>
num_frames=<N> VPIO buffer size in mono samples.
Expected ~960 (20ms) or ~1104 (23ms).
Outliers point at format mismatch.
peak_stereo_f32=<f32> Peak |sample| of AudioHandler's
output BEFORE gain + downmix.
0.0 = handler is producing silence
(jitter underrun, no audio).
~1.0 = full-scale content reaching
the callback as expected.
peak_out_i16=<i16> Peak |sample| of the downmixed mono
i16 we write to VPIO. Zero with
non-zero peak_stereo = downmix bug.
Near 32767 = clipping pressure.
gain=<f32> Current master output gain.
What we'll be able to diagnose from a 5-second talker session:
* peak_stereo_f32 = 0 throughout
-> AudioHandler isn't producing samples. Inbound forwarder
may not be feeding it, or jitter buffer is stuck in
buffering_samples state. NOT a render-callback bug.
* peak_stereo_f32 oscillating, peak_out_i16 = 0
-> Downmix or i16 cast is broken. Math bug in the loop.
* num_frames wildly different from ~960-1104
-> StreamFormat got rejected and VPIO is delivering a
different rate. Format-pinning fight with the session.
* peak_stereo_f32 normal AND peak_out_i16 normal AND user
still says 'broken'
-> The signal reaches the device cleanly but iOS's VPIO
output processing (AEC residual subtraction, AGC
compression, NS gate) is mangling it after our callback
returns. That's a VPIO-config problem, not a render-
callback problem; fix is to disable specific VPIO
voice-processing properties on the unit before
initialize().
Pure diagnostic commit. No behavioural change beyond a
warn-rate-limited info log line every ~2 seconds. Cost in the
audio thread is one branch + counter increment + (every 100th)
a tracing macro invocation.
Build counter 65 -> 66.
Replace the silence-emitting render callback from commit 1 with
real playback that drives AudioHandler::fill_buffer and downmixes
its 48 kHz stereo f32 output to the i16 mono buffer VPIO expects.
Pipeline per render callback (mirrors the cpal-output + sdl_output
contracts so the platform-neutral playback path is preserved):
1. Lock the shared Arc<Mutex<AudioHandler>>, ask fill_buffer to
populate a stereo-f32 scratch slice of length 2*num_frames.
AudioHandler runs Opus decode + per-client jitter buffer + mix
internally. Same primitive every other platform calls.
2. If output_muted is true, zero the i16 output buffer and return.
We still ran fill_buffer in step 1 so the jitter buffer drains
while muted — preventing unbounded growth — which matches the
cpal/SDL backend contract.
3. Downmix stereo -> mono with master gain:
mono_f32 = (l + r) * 0.5 * gain
i16_out = (mono_f32.clamp(-1.0, 1.0) * i16::MAX) as i16
The 0.5 average preserves total signal energy with 3 dB
headroom against sum-of-correlated-peaks clipping. Multiply by
gain after the downmix saves one mul per sample. Hard-clip on
the i16 cast is acceptable because the upstream stereo signal
is already in [-1.0, 1.0] from the f32 mix; only gain >1.0
creates clipping pressure and that path is identical to every
other backend's i16 conversion.
Closure ownership:
* scratch_stereo: Vec<f32> moved into the FnMut closure. First
callback grows it to 2*num_frames; subsequent callbacks reuse
the backing allocation. The audio thread never hits the
allocator on steady-state callbacks.
* handler_for_render / output_gain_for_render / output_muted_for_render
are Arc clones taken before the closure literal.
Public API change: AudioEngine -> IosVoiceUnit::start parameters
that were previously underscored (commit 1 placeholder) are now
all consumed by the wiring. Signature is unchanged, just the
binder names lose the leading underscore. engine.rs call-site
is unaffected.
Build verify on Mac (target aarch64-apple-ios): cargo check
clean in 0.33s, no errors, no warnings.
Build counter 63 -> 64 — About dialog shows v1.0.0-rc.8+64.
Replace the no-op input callback from commit 1 with a real
capture pipeline that mirrors the cpal-side CaptureState in
engine.rs but is type-specialised for the i16 mono samples VPIO
delivers natively.
New IosCaptureState struct (private to ios_voice_unit.rs) owns:
* OpusEncoder configured for VoIP at 48 kHz mono (32 kbps,
complexity 10, inband FEC, packet-loss-perc 5 — identical
tuning to try_open_capture in engine.rs).
* pcm_accum: Vec<i16> with capacity 2*FRAME_SAMPLES_MONO, growing
if a VPIO callback ever delivers more than ~40 ms.
* opus_out: [u8; MAX_OPUS_FRAME] scratch.
* Cloned Arc<AtomicBool> transmit gate + Arc<AtomicU32> frames-sent
counter shared with AudioEngine.
ingest_i16 flow:
1. If PTT gate is off -> clear accumulator + return (matches cpal
behaviour, no pop on PTT-release edge).
2. Apply mic_gain. Fast-path when gain==1.0 skips the multiply +
saturate loop entirely; otherwise saturating mul-then-cast
keeps the signal in the i16 envelope.
3. Drain complete 20 ms / 960-sample frames from the accumulator,
encode via encoder.encode (i16 path, no float conversion
needed since VPIO already gave us i16), build OutPacket with
AudioData::C2S { codec: OpusVoice }, try_send on voice_out_tx.
4. Frame buffer is stack-allocated [i16; FRAME_SAMPLES_MONO] —
no per-callback heap allocation on the realtime audio thread.
VPIO setup changes in IosVoiceUnit::start:
* NEW: explicit kAudioOutputUnitProperty_EnableIO (=2003) with
value 1 on (Scope::Input, Element::Input) BEFORE the stream
format setters. VPIO's input element is OFF by default; without
this toggle no audio flows in and the input callback never
fires. Commit 1's comment claiming set_input_callback handles
this was wrong; coreaudio-rs's set_input_callback only installs
the kAudioOutputUnitProperty_SetInputCallback property, not
the EnableIO toggle.
* Apple's documented sequence (now matched):
1. AudioComponentInstanceNew -> AudioUnit::new_uninitialized
2. EnableIO on element 1 -> set_property(2003, ...)
3. Stream format both elems -> set_stream_format x2
4. Install callbacks -> set_input_callback + set_render_callback
5. AudioUnitInitialize -> unit.initialize
6. AudioOutputUnitStart -> unit.start
* set_input_callback closure now moves the IosCaptureState in
by value and calls ingest_i16 with args.data.buffer (the
&mut [i16] coreaudio-rs delivers after running AudioUnitRender
internally to pull the mic samples into a pre-allocated
AudioBufferList).
What this commit does NOT do:
* Output render callback is still a silence-emitting stub.
Commit 4 lands the AudioHandler::fill_buffer + i16 downmix.
* Route-change handling — commit 5.
Build verify on Mac (target aarch64-apple-ios): cargo check
clean in 1.18s, no errors, no warnings.
Build counter 62 -> 63 — About dialog shows v1.0.0-rc.8+63.
iOS build of chanora_audio (target aarch64-apple-ios) failed with
four compilation errors after commit 2 landed. Root causes were
all simple symbol-path / cfg-gating mistakes from the skeleton
commit; the underlying design is unchanged.
1. ios_voice_unit.rs: wrong import path for OutPacket. The
chanora_protocol crate re-exports it at the crate root
(`pub use ...::OutPacket` in lib.rs line 52), not from a
`voice` submodule (which doesn't exist).
- use chanora_protocol::voice::OutPacket;
+ use chanora_protocol::OutPacket;
2. ios_voice_unit.rs: LinearPcmFlags lives in
`coreaudio::audio_unit::audio_format`, not in
`stream_format` (the doc page lists it under StreamFormat but
the actual module path is the upstream Apple naming).
- use coreaudio::audio_unit::stream_format::LinearPcmFlags;
+ use coreaudio::audio_unit::audio_format::LinearPcmFlags;
3. ios_voice_unit.rs: `Ordering` import unused (commit 1
skeleton callbacks don't load atomics yet — that comes in
commits 3 + 4). Remove from the std::sync::atomic import to
silence the unused_imports warning.
4. engine.rs::AudioEngine::stop(): the existing body unconditionally
touched self._input_stream and self._output_stream, but commit
2 cfg-gated those fields away on iOS (and added an iOS-only
_ios_voice_unit field in their place). Split the field drop
logic with the same target_os = "ios" cfg so each platform
only touches the fields it actually has.
Also clean up two pre-existing warnings exposed by the iOS cfg
gating:
5. engine.rs: `tracing::{error, warn}` were imported
unconditionally but are only used inside cpal log lines.
Cfg-gate the import to not(target_os = "ios").
6. engine.rs: `AudioData`, `CodecType`, `OutAudio` from
chanora_protocol are only referenced in the Opus encoder feed
inside CaptureState — cpal-side only. Cfg-gate to
not(target_os = "ios"); keep `InboundVoice` + `OutPacket`
on the unconditional path because the inbound forwarder + (in
commit 3) the iOS capture pipeline both reference them.
Also fix a stray duplicate `#[cfg(not(target_os = "ios"))]`
attribute that landed on line 18 in commit 2.
Build verify (Linux host): cargo check -p chanora_audio clean
in 0.53s. iOS-side check pending on Mac.
Build counter bumped 60 -> 61 — the About dialog will display
v1.0.0-rc.8+61 so the user can confirm the build under test
matches this commit.
Skeleton scaffolding for the iOS VoiceProcessingIO backend that
will replace cpal on iOS. This commit lands the dependency + the
module + a constructable AudioUnit that emits silence and drops
input; nothing in engine.rs is wired up yet (that is commit 2).
Compilation contract for this commit:
* Linux / Windows / macOS / Android builds unaffected (the new
module is target_os='ios' gated, the new dep is in
'[target."cfg(target_os = \"ios\")"]').
* iOS build pulls in coreaudio-rs 0.14, constructs a VPIO unit,
pins stream format to 48 kHz Int16 mono on both buses, installs
no-op input + silence-emitting render callbacks, initializes,
and starts. No audio is actually moved until commits 3/4.
Why VPIO and not RemoteIO via cpal: cpal's iOS backend opens
RemoteIO with no control over stream format / buffer size /
channels and produces a mono-only output element that stays bound
to the route present at construction time. End-user symptom on
iPhone 16 Pro iOS 18.7.8: tapping Speaker in the picker flips
AVAudioSession.currentRoute.outputs to Speaker (confirmed in our
diagnostic logs from commit da631a2) but audio keeps coming out
the receiver because the AudioUnit's output binding is stale.
Every production iOS VoIP client (Mumble iOS, Linphone /
mediastreamer2, Signal-iOS, Jitsi Meet iOS, the WebRTC reference
impl) avoids RemoteIO and uses VoiceProcessingIO instead. VPIO is
Apple's recommended voice unit; it ships hardware AEC + AGC + NS
and re-binds the physical transducer correctly on route changes
because it IS the canonical voice unit on iOS — FaceTime's audio
path runs through it.
coreaudio-rs 0.14 (RustAudio org, 8.6M downloads, same maintainers
as cpal) gives us a safe wrapper around the AudioUnit C API on
iOS. Uses objc2-* crates underneath so links cleanly into iOS
builds. ios_voice_unit.rs sits next to sdl_output.rs as the iOS
sibling of the Linux SDL2 output path.
Build verify (Linux host): `cargo check -p chanora_audio`
finished clean in 16.09s. iOS build verification happens in
commit 2 when the module is exercised; for this commit the module
compiles in isolation but is dead code on iOS too (no caller).