Files
chanora/apps
EdisonJwa 99584fbc1a fix(audio,ios): decouple AudioHandler from VPIO render callback via ring buffer (rc.8+73)
User confirmed at +72 the symptom is 'voice + constant clicks +
choppy fragments'. The diagnostic data conclusively pointed to
iOS VPIO render-callback timing as the cause:

  * frames_changes=60+ per 100 callbacks at cb>=1800
    iOS keeps switching num_frames between 960 and 1104
    on roughly 60% of callbacks
  * peak_out_i16 is sensible (2500-16870, never clipping)
    when AudioHandler returns content
  * input_was_zero=true on most callbacks during active speech
    AudioHandler keeps entering buffering_samples state

The cause: previous render callback called fill_buffer
synchronously every iOS audio thread invocation. With iOS
calling at irregular rates with irregular sizes, AudioHandler's
jitter buffer (sized around 20 ms Opus frames) cannot satisfy
arbitrary-sized requests and falls back to returning silence
(&[] empty slice) on misaligned reads. The silent gaps in the
middle of the output buffer create discontinuities = audible
clicks; the missing-tail content produces choppy fragments.

Fix (architectural): decouple the AudioHandler decoder from the
VPIO render callback via a lock-free SPSC ring buffer.

  Producer (tokio task, 50 Hz):
    every 20 ms:
      fill_buffer(scratch_stereo_f32, 1920 = 20 ms stereo)
      downmix L+R -> mono i16 (960 samples)
      ring_buffer.push_slice(mono_i16)

  Consumer (VPIO render callback, iOS audio thread):
    every callback:
      pop num_frames samples from ring buffer into out
      zero-fill tail on underrun

Why it works:
  * Producer always asks AudioHandler for a stable 20 ms chunk
    (perfectly aligned with internal Opus frame size). No more
    buffering_samples false-triggers.
  * Consumer pulls whatever iOS asks for whenever iOS schedules
    it; ring buffer's 200 ms depth absorbs the callback jitter.
  * This is the standard pattern every production VoIP audio
    engine uses (WebRTC, Discord, FaceTime) to bridge bursty
    Opus decoders to bursty platform audio callbacks.

Implementation:
  * New dep: rtrb 0.3.4 (RustAudio realtime-safe SPSC ring
    buffer, 6.8M downloads, lock-free push/pop with no
    allocation on the audio thread).
  * RING_BUFFER_SAMPLES = 9600 (200 ms mono i16 at 48 kHz).
    Sized for 10x producer ticks of headroom.
  * PRODUCER_TICK_MS = 20 (matches Opus 50 Hz packet rate).
    set_missed_tick_behavior(Skip) to avoid burst catch-up on
    runtime stalls.
  * Producer task spawned in IosVoiceUnit::start, shutdown
    via tokio::oneshot when IosVoiceUnit drops.
  * Render callback is now just: pop into out, zero-fill tail,
    apply mute then gain.
  * Gain applied CONSUMER-side so user volume changes take
    effect within one callback (<= 200 ms latency).
  * Underrun diagnostics: count underrun callbacks + total
    zero-filled samples, log every 100 callbacks.

Threading + safety:
  * rtrb is lock-free SPSC. Audio thread never blocks.
  * Producer can block briefly on Arc<Mutex<AudioHandler>>
    contention with the inbound forwarder (handle_packet), but
    not with the audio thread.
  * Producer task is owned by tokio runtime; explicit shutdown
    channel ensures it exits when the engine stops.

Build verify:
  * Linux host: cargo check clean in 4.06s (downloads rtrb 0.3.4).
  * iOS Mac:   cargo check clean in 2.14s.

Build counter 72 -> 73.
2026-05-17 02:39:42 +08:00
..