Deduplicate audio processing toggle UI from voice_compact.dart and
voice_settings.dart into shared AudioProcessingPanel widget in
voice_settings_controls.dart. Centralize platform detection getters
in voice_platform.dart.
* feat(audio): add Silero ONNX VAD with WebRTC fallback
Introduce SileroOnnxVad and SileroOnnxVadWorker for desktop targets. The worker runs Silero v6 ONNX inference on a dedicated thread, accumulating 10 ms frames into the 512-sample 16 kHz input the model expects. Add VadOutput, VoiceActivityDetector trait, and WebRtcFallbackVad to provide a uniform VAD interface with graceful fallback when the ONNX model is unavailable. Wire the new VadBackend variants through AudioProcessingConfig and the snapshot stats so the bridge can report which detector is active.
* feat(audio): integrate desktop VAD worker into capture engine
Wire SileroOnnxVadWorker into the desktop capture path so voice activity can open the transmit gate before encoding. The capture callback now processes all audio through resample, downmix, and VAD unconditionally; transmit_active still gates Opus encoding.
Add new_desktop_audio_processing_state() to construct the config/stats/worker triple, and apply_desktop_vad_backend() to synchronously load or clear the worker on config changes. Override processing_backend to Noop for desktop so bridge diagnostics report the correct backend rather than the iOS-oriented PlatformVoiceProcessing default.
Includes review-driven cleanups: StreamConfig clone to deref per clippy, and a comment explaining why two try_lock calls on silero_vad_worker are structurally necessary (borrow checker requires the policy probe and the fallback path to not share a lock guard because mark_vad_fallback_active takes &mut self).
* fix(audio): modernize Windows PTT to current windows-rs API
Port the Raw Input plus low-level keyboard hook PTT backend to the newer windows-rs patterns: OptionalHandle, Result-returning CreateWindowExW, and None for CallNextHookEx. Replaces the old HHOOK(0) pointer casts. Add deterministic tests for mouse button 4 and 5 press and release driving the gate.
* build(windows): force MSVC release CRT for audiopus cmake builds
audiopus_sys calls cmake::build(opus_path), so downstream Cargo env cannot use cmake-rs Config::define() to override CMake's MSVC Debug CRT defaults. Point cmake-rs at a small wrapper that injects the policy and cache variables during configure while passing cmake --build, --version, and -E through unchanged. Keeps Opus Debug builds on Rust's release dynamic CRT (/MD) instead of CMake's default debug CRT (/MDd), which otherwise pulls in unresolved __imp__CrtDbgReportW symbols at test link.
Document that the iOS deployment target is intentionally absent from this file. It is enforced by tools/build-ios.sh and the Xcode project; setting it globally here would make native macOS cargo check runs try to link iPhone objects against the macOS SDK.
* build(flutter): update pubspec.lock after plugin additions
Regenerated lockfile reflecting the local_notifications and connectivity_plus plugin additions from the poke-notifications feature.
* fix(audio): address PR #37 review findings
Six fixes from independent PR review:
1. BLOCKER: Replace Windows-only cmake .cmd wrapper with cross-platform
CMake env vars. Setting CMAKE=tools/cmake-msvc-release-crt.cmd
globally broke non-Windows hosts because cmake-rs would try to
execute a .cmd file on macOS/Linux. Instead, set
CMAKE_POLICY_DEFAULT_CMP0091=NEW and CMAKE_MSVC_RUNTIME_LIBRARY=
MultiThreadedDLL as env vars that CMake reads natively. MSVC-
specific vars are safely ignored by GCC/Clang toolchains. Delete
the now-unnecessary wrapper script.
2. IMPORTANT: Join the Silero worker thread in Drop instead of
detaching it. The old code dropped the JoinHandle which detaches
the thread; the new code calls handle.join() after closing the
channel, ensuring the ONNX session is cleaned up before the
worker is replaced during config changes.
3. IMPORTANT: Single-try_lock refactor of the capture VAD callback.
The double try_lock (policy probe + send) is replaced by a single
scoped try_lock that both probes availability and sends the frame.
The guard is dropped before the fallback path, which needs &mut
self for mark_vad_fallback_active. This also eliminates the
VadWorkerPolicy enum and callback_vad_worker_policy function,
whose behavior is now inlined into the callback.
4. IMPORTANT: Remove tracing from the realtime capture callback.
mark_vad_fallback_active and sync_vad_backend emitted info!/warn!
from the audio thread. Replace with silent atomic state
publishing via SharedAudioProcessingStats; the bridge stats
stream already exposes vad_fallback_active for diagnostics.
5. IMPORTANT: Defer ONNX model load outside the worker mutex.
apply_desktop_vad_backend_to_worker now constructs the new worker
before taking the lock, then swaps it in under a short hold.
This prevents the realtime callback from being blocked during
model I/O + thread spawn.
6. MINOR: Remove unused VadBackend import from vad/mod.rs after
deleting the policy code.
* fix(audio): address PR #37 second-pass review findings
5-agent review found 5 blocking issues. All addressed:
1. BLOCKER: CMake env vars don't reach CMake cache. Restored .cmd wrapper
but scoped to Windows MSVC targets only via [target.x86_64-pc-windows-msvc]
and [target.aarch64-pc-windows-msvc] in .cargo/config.toml. Non-Windows
hosts are unaffected.
2. BLOCKER: processing_backend normalized in set_audio_processing_config
on desktop (cfg-gated override to Noop), mirroring startup default.
3. BLOCKER: Model-path reload was already wired via reload_audio_processing_config.
Fixed misleading doc comment in core/lib.rs.
4. BLOCKER: DEC-030 updated to reflect desktop VoiceActivity enablement.
Traceability docs (SRS, SysDes, SAD, SDD, implementation-status) updated.
5. Silero ONNX cfg narrowed to desktop-only (excludes macOS/Android).
Cargo.toml ort dependency target cfg narrowed similarly.
6. Realtime callback debt documented as TODO at CaptureState::ingest.
* fix(audio): exclude ort dep on Android target
ort does not provide first-class Android prebuilts in our pin, mirror the
iOS/macOS exclusion so cargo metadata succeeds for android targets.
* test(audio): fix stale select_ptt_backend import in ptt_privacy
The helper moved out of the ptt_backends submodule onto the crate root;
update the integration test imports so the test compiles again.
* build(windows): scope MSVC release CRT cmake wrapper via Cargo [env]
Cargo's [target.<triple>] table only forwards a fixed allowlist
(linker, runner, rustflags, rustdocflags, ar), so setting CMAKE there
was silently dropped and audiopus_sys kept linking the debug CRT,
producing LNK4098 'MSVCRTD conflicts' and __imp__CrtDbgReportW errors
on x86_64-pc-windows-msvc test builds.
Move the override to Cargo's [env] table using cc/cmake-rs's
target-suffixed CMAKE_<triple> lookup (force=true, relative=true) so it
applies to MSVC targets only and not to host tooling. Add stdout
markers to the wrapper so its invocation is provable in cargo -vv logs.
Verified: cargo test -p chanora_audio --target x86_64-pc-windows-msvc
--lib --no-run now links cleanly; CMakeCache.txt records
CMAKE_MSVC_RUNTIME_LIBRARY=MultiThreadedDLL and CMP0091=NEW.
* fix(flutter): gate VoiceActivity transmit mode by platform support
VoiceActivity relies on the native VAD worker, which is only wired up
on Windows, Linux, and Android. Showing the option on iOS, macOS, or
web let users select a mode that silently never transmitted.
Add voiceActivityTransmitAvailable + transmitModeSegmentsFor() helpers
in voice_settings_controls.dart, hide the VAD row in voice_compact.dart
and drop the VAD segment from the settings dialog when unsupported.
Keep the legacy const transmitModeSegments for the existing widget test
and add two new tests covering the gated helper.
Clarify ONNX Runtime guidance with direct-open install hints, restore desktop WebRTC VAD visibility, map mouse side buttons through focused PTT capture/runtime paths, and wait for server acks before showing chat sends as successful.
Constraint: Linux release UX must stay functional when ONNX Runtime is optional and GNOME portal availability varies
Rejected: Keep desktop VAD locked to Silero only | misleads users when ONNX Runtime is skipped
Confidence: medium
Scope-risk: moderate
Directive: Preserve the protocol send-ack wait path for chat so UI success always tracks real server acceptance
Tested: flutter analyze lib/main.dart lib/widgets/chat_views.dart lib/widgets/input_dialogs.dart lib/widgets/startup_dependency_screen.dart; flutter test test/widgets/input_dialogs_test.dart test/widgets/chat_views_test.dart test/services/startup_dependency_check_test.dart test/widgets/startup_dependency_screen_test.dart test/widgets/voice_settings_controls_test.dart test/widgets/audio_processing_config_state_test.dart; cargo test -p chanora_protocol --lib; cargo test -p chanora_audio ptt_backends --lib
Not-tested: Live manual GNOME portal rebind/global PTT on a real desktop session; observer-bot chat against a live server after the sender-name fallback change
User report from sideloaded iPhone build, in order of priority:
#4 'Could not join channel: audio: audio backend:
build_output_stream: The requested stream configuration is
not supported by the device.'
Cause: we forced cpal::BufferSize::Fixed(2048) on the output
and input streams unconditionally on non-Linux. iOS CoreAudio
RemoteIO units reject arbitrary buffer-size requests with that
exact error. Windows WASAPI needs the pinning for shared-mode
jitter, but macOS / iOS do not.
Fix: cfg-gate Fixed(2048) to target_os = 'windows'; everywhere
else use BufferSize::Default and let the platform HAL pick.
crates/chanora_audio/src/engine.rs.
#5 'Could not join channel: invariant violated:
voice_in already taken'
Cause: start_audio tore down the old engine BEFORE attempting
to construct the new one, and consumed voice_in (an mpsc
Receiver that can only be taken once) early. When the new
engine failed mid-construction (e.g. because of #4 above) the
session was left with: no audio engine, voice_in consumed,
no way to retry without reconnect. The second voice_join
attempt surfaced the invariant message.
Fix: build the new engine BEFORE tearing down the old. Only
swap state.audio if construction succeeded. crates/chanora_
core/src/lib.rs::ChanoraSession::start_audio. Additionally
added a put_voice_in helper to the protocol adapter (
crates/chanora_protocol/src/adapter.rs) for a future
broadcast-channel migration; the helper is unused on the
immediate fix path but documents the intent.
#3 'permission request would better on first open'
Cause: AVAudioSession only triggers the mic-permission
prompt the first time it tries to record. We never recorded
until voice_join, so the prompt fired then.
Fix iOS: AVAudioSession.sharedInstance().requestRecordPermission
in AppDelegate.swift::application(_:didFinishLaunchingWithOptions:).
Fix macOS: AVCaptureDevice.requestAccess(for: .audio) in
macos/Runner/AppDelegate.swift::applicationDidFinishLaunching.
Both run non-blocking; user can deny without crashing app
launch, and voice_join then surfaces a clearer downstream
error when the engine fails to open the input device.
#1 + #2 'one-column upper takes too much space; Push to Talk
button at bottom would be better'
Layout rework for narrow-mode (single column, mobile shape):
- Flipped the stacking order in main.dart so Voice Bar moves
to the BOTTOM of the body and the channel tree (Expanded)
fills above. Wide-mode (Row, >= 840 dp) layout unchanged.
- Inside the Voice Bar on touch-only hosts, moved the
on-screen Push to Talk button to be the LAST element of
the Voice Bar (was Row 3). Order now: pill + mutes, mode
badge + settings, level meter, stats line, release-tail
caption, PTT button. The button is closest to the user's
thumb when the Voice Bar is pinned to the bottom of a
narrow-layout screen.
#6 'remove right top debug badge'
debugShowCheckedModeBanner: false on the MaterialApp.
Release builds never showed it anyway; this only affects
local dev / debug builds.
#7 'what does the refresh button use for? nothing happened'
Removed. The snapshot updates via BridgeEvent::SnapshotChanged
are pushed from the bridge — a manual rust.snapshot() call
was redundant. Now only the Diagnostics + Disconnect actions
remain in the AppBar trailing row when connected.
#8 'Bind Key related function should not be added to a mobile
platform'
widgets/voice_settings.dart: bind-key OutlinedButton is now
#cfg'd out when Platform.isIOS || Platform.isAndroid. The
release-tail slider stays because it still applies to the
on-screen PTT button. Capability badge in voice_bar.dart
also hidden on mobile (it would always show L0Focused which
is redundant with the visible on-screen button).
Tests + analyze: chanora_audio 34/0/0 on macOS, workspace 78/0/1
on Linux; flutter analyze clean (6 pre-existing Radio.groupValue
infos). flutter build ios --release --no-codesign: 28.8 s clean
(Runner.app 29.9 MB).
1. Hard-mute now informs the server (setInputMuted) in addition to
clamping the local TransmitGate. Without the server-side flag,
other clients keep seeing us un-muted; without the local clamp
a beat of in-flight audio leaks through. Drive both together so
the mic icon and the actual silence land at the same time.
2. Split the badge's Configure affordance from the Voice Bar's
'Voice settings' gear. The gear opens the mode + release-tail
dialog (onConfigure); the badge's configure opens the bind-key
capture flow directly (new onBindKey). Previously both routed
to the settings dialog, so 'Voice settings' and the badge's
'Configure' were the same screen — useless duplication.
3. Bind-key label is now PTT-only. The mode-badge row no longer
prints 'PTT: Space' when Continuous / Voice Activity is
selected. A new PTT-only secondary line carries the bound key
plus the release-tail value together, hidden entirely for
non-PTT modes.
4. Release-tail row is now PTT-only in BOTH the Voice Bar and the
Voice settings dialog. The dialog previously kept the slider
visible across all modes; switching to Continuous left the
user staring at a control that did nothing.
5. PTT capability badge is now PTT-only. In Continuous and Voice
Activity modes there is no key binding to surface a capability
for, so the 'L0Focused (focused)' line + its info sheet and
the Configure button disappear from the Voice Bar when the
user isn't in PTT mode.
All five fixes are pure UI; no Rust changes needed. flutter analyze
remains clean (6 pre-existing Radio.groupValue deprecation infos).
Implement the SDD-094 / SDD-095 / SDD-096 / SDD-097 detailed designs
committed in dfa84ee.
Rust side
- chanora_audio::TransmitMode enum (Ptt/Continuous/VoiceActivity) with
serde-friendly u8 repr (SDD-095).
- chanora_audio::TransmitModeSelector: lock-free Atomic-backed selector
that is the sole writer of transmit_active (per SAD-083), applying
hard_mute as a final clamp. VoiceActivity falls through to Continuous
for v1 (DEC-030 placeholder).
- chanora_audio::ReleaseTailTimer: tokio-task-owning struct driving the
selector's ptt_held input; default 200 ms tail, configurable 0–500 ms
with AtomicU32 hot read; pending JoinHandle held in a std::sync::Mutex
touched only on PTT edge transitions (SDD-096).
- chanora_storage: AudioMeta persisted as audio_meta.json next to
identity.dek; get/set_transmit_mode + get/set_release_tail_ms with
0..=500 clamp on write.
- chanora_core::ChanoraSession: voice_join(channel, password) and
voice_leave() are the new lifecycle entry points; ensure_audio_running
and shutdown_audio_if_idle are private helpers around the existing
Option<AudioEngine> field. SessionEvent::VoiceState carries the
in_channel / transmit_mode / mute / release_tail_ms tuple. Selector
state survives reconnect; supervisor rewires it to each fresh engine
gate.
- chanora_bridge: drop start_audio; add voice_join, voice_leave,
set/get_transmit_mode, set/get_release_tail_ms, set_hard_mute.
BridgeEvent::VoiceState mirrors the core event. AudioStarted/Stopped
kept for backwards compat but Flutter ignores them in the new UI.
Flutter side
- New apps/chanora_flutter/lib/widgets/voice_bar.dart replaces the
legacy _AudioControls widget. Renders channel pill, mode badge,
mute toggle, level meter, PttCapabilityBadge, leave button. No
manual Start affordance anywhere.
- New apps/chanora_flutter/lib/widgets/voice_settings.dart dialog with
TransmitMode radio group (VoiceActivity disabled with 'Coming soon'
trailing label per DEC-030), bind-key button, release-tail slider
0–500 ms step 25.
- main.dart: state fields _inChannel, _transmitMode, _hardMute,
_releaseTailMs driven by BridgeEvent_VoiceState. Channel-tap now
calls voiceJoin instead of moveToChannel. Removed _onStartAudio,
_audioStarted-gated branch, and the FilledButton.
- l10n: 11 new strings in app_en.arb + app_zh.arb.
Verification
- cargo check --workspace: clean.
- cargo test --workspace --lib: 72 passed / 0 failed / 1 ignored
(chanora_audio: +12 new tests for TransmitMode/Selector/ReleaseTail;
chanora_storage: +2 new tests for audio_meta round-trip).
- flutter analyze: 0 errors, 0 warnings; 6 infos are the Flutter 3.32
Radio.groupValue deprecation (pre-existing API usage).
- FRB Dart/Rust bindings regenerated via flutter_rust_bridge_codegen.
Follow-up (intentionally deferred)
- PttController and per-platform PTT backends still drive AudioTransmitGate
directly via the legacy set_ptt path; routing those key edges through
ChanoraSession::release_tail_timer().{key_down,key_up} so the tail
applies to native PTT input is a contained wiring change in a follow-up.
- Real audio-level RMS in BridgeAudioStats (current meter is binary).
- VoiceActivity backend (DEC-030).