ascii-chat 0.11.33
Video chat in your terminal
Loading...
Searching...
No Matches
client_pipeline.h File Reference

Unified client-side audio processing pipeline. More...

Go to the source code of this file.

Data Structures

struct  client_audio_pipeline_flags_t
 Component enable/disable flags. More...
 
struct  client_audio_pipeline_config_t
 Pipeline configuration parameters. More...
 
struct  client_audio_pipeline_t
 Client audio pipeline state. More...
 

Macros

#define CLIENT_AUDIO_PIPELINE_FLAGS_ALL
 Default flags with all processing enabled.
 
#define CLIENT_AUDIO_PIPELINE_FLAGS_MINIMAL
 Minimal flags for testing (only codec, no processing)
 
Audio Pipeline Constants
#define CLIENT_AUDIO_PIPELINE_SAMPLE_RATE   48000
 
#define CLIENT_AUDIO_PIPELINE_FRAME_MS   20
 
#define CLIENT_AUDIO_PIPELINE_FRAME_SIZE   (CLIENT_AUDIO_PIPELINE_SAMPLE_RATE * CLIENT_AUDIO_PIPELINE_FRAME_MS / 1000)
 
#define CLIENT_AUDIO_PIPELINE_ECHO_REF_SIZE   (CLIENT_AUDIO_PIPELINE_SAMPLE_RATE / 2)
 
#define CLIENT_AUDIO_PIPELINE_MAX_OPUS_PACKET   4000
 

Typedefs

typedef struct SpeexPreprocessState_ SpeexPreprocessState
 
typedef struct JitterBuffer_ JitterBuffer
 
typedef struct webrtc_EchoCanceller3 webrtc_EchoCanceller3
 
typedef struct OpusEncoder OpusEncoder
 
typedef struct OpusDecoder OpusDecoder
 

Functions

client_audio_pipeline_config_t client_audio_pipeline_default_config (void)
 Get default configuration.
 
Lifecycle Functions
client_audio_pipeline_t * client_audio_pipeline_create (const client_audio_pipeline_config_t *config)
 Create a new client audio pipeline.
 
void client_audio_pipeline_destroy (client_audio_pipeline_t *pipeline)
 Destroy a client audio pipeline.
 
Flag Management
void client_audio_pipeline_set_flags (client_audio_pipeline_t *pipeline, client_audio_pipeline_flags_t flags)
 Set component enable flags.
 
client_audio_pipeline_flags_t client_audio_pipeline_get_flags (client_audio_pipeline_t *pipeline)
 Get current component enable flags.
 
Capture Path (Microphone → Network)
int client_audio_pipeline_capture (client_audio_pipeline_t *pipeline, const float *input, int num_samples, uint8_t *opus_out, int opus_out_size)
 Process captured audio and encode to Opus.
 
Playback Path (Network → Speakers)
int client_audio_pipeline_playback (client_audio_pipeline_t *pipeline, const uint8_t *opus_in, int opus_len, float *output, int max_samples)
 Decode Opus packet and process for playback.
 
int client_audio_pipeline_get_playback_frame (client_audio_pipeline_t *pipeline, float *output, int num_samples)
 Get audio frame from jitter buffer for playback callback.
 
Full-Duplex AEC3 Processing
void client_audio_pipeline_process_duplex (client_audio_pipeline_t *pipeline, const float *render_samples, int render_count, const float *capture_samples, int capture_count, float *processed_output)
 Process AEC3 inline in full-duplex callback.
 
Status and Diagnostics
int client_audio_pipeline_jitter_margin (client_audio_pipeline_t *pipeline)
 Get jitter buffer margin (buffered time in ms)
 
bool client_audio_pipeline_voice_detected (client_audio_pipeline_t *pipeline)
 Check if VAD detected voice activity in last capture.
 
void client_audio_pipeline_reset (client_audio_pipeline_t *pipeline)
 Reset pipeline state.
 

Detailed Description

Unified client-side audio processing pipeline.

This header provides a complete audio processing pipeline for ascii-chat clients, integrating WebRTC AEC3 (production-grade echo cancellation), SpeexDSP (noise suppression, AGC, VAD), Speex jitter buffer, Opus codec, and mixer components (compression, noise gate, filters).

PIPELINE ARCHITECTURE:

CAPTURE PATH (microphone → network): Mic Input (float32, 48kHz) ↓ Echo Cancellation (WebRTC AEC3 - automatic network delay estimation + adaptive filtering) ↓ Preprocessor (Noise/AGC/VAD) ↓ High-Pass Filter (remove rumble) ↓ Low-Pass Filter (remove hiss) ↓ Noise Gate (silence detection) ↓ Compressor (dynamic range control) ↓ Opus Encode → Network

PLAYBACK PATH (network → speakers): Network ↓ Speex Jitter Buffer ↓ Opus Decode ↓ Register with AEC (WebRTC AEC3 reference signal for echo learning) ↓ Soft Clipping → Speakers

COMPONENT FLAGS:

Each processing stage can be individually enabled/disabled via flags. This allows testing and debugging specific components.

THREAD SAFETY:

  • Pipeline has internal mutex for state protection
  • Capture and playback paths can run concurrently
  • Echo reference feeding is thread-safe
Author
Zachary Fogg me@zf.nosp@m.o.gg
Date
December 2025

Definition in file client_pipeline.h.

Macro Definition Documentation

◆ CLIENT_AUDIO_PIPELINE_ECHO_REF_SIZE

#define CLIENT_AUDIO_PIPELINE_ECHO_REF_SIZE   (CLIENT_AUDIO_PIPELINE_SAMPLE_RATE / 2)

Echo reference ring buffer size (500ms at 48kHz)

Definition at line 96 of file client_pipeline.h.

◆ CLIENT_AUDIO_PIPELINE_FLAGS_ALL

#define CLIENT_AUDIO_PIPELINE_FLAGS_ALL
Value:
.echo_cancel = true, \
.noise_suppress = true, \
.agc = true, \
.vad = true, \
.jitter_buffer = true, \
.compressor = true, \
.noise_gate = true, \
.highpass = true, \
.lowpass = true, \
})
Component enable/disable flags.

Default flags with all processing enabled.

Definition at line 133 of file client_pipeline.h.

134 { \
135 .echo_cancel = true, \
136 .noise_suppress = true, \
137 .agc = true, \
138 .vad = true, \
139 .jitter_buffer = true, \
140 .compressor = true, \
141 .noise_gate = true, \
142 .highpass = true, \
143 .lowpass = true, \
144 })

◆ CLIENT_AUDIO_PIPELINE_FLAGS_MINIMAL

#define CLIENT_AUDIO_PIPELINE_FLAGS_MINIMAL
Value:
.echo_cancel = false, \
.noise_suppress = false, \
.agc = false, \
.vad = false, \
.jitter_buffer = false, \
.compressor = false, \
.noise_gate = false, \
.highpass = false, \
.lowpass = false, \
})

Minimal flags for testing (only codec, no processing)

Definition at line 149 of file client_pipeline.h.

150 { \
151 .echo_cancel = false, \
152 .noise_suppress = false, \
153 .agc = false, \
154 .vad = false, \
155 .jitter_buffer = false, \
156 .compressor = false, \
157 .noise_gate = false, \
158 .highpass = false, \
159 .lowpass = false, \
160 })

◆ CLIENT_AUDIO_PIPELINE_FRAME_MS

#define CLIENT_AUDIO_PIPELINE_FRAME_MS   20

Default frame size in milliseconds

Definition at line 90 of file client_pipeline.h.

◆ CLIENT_AUDIO_PIPELINE_FRAME_SIZE

#define CLIENT_AUDIO_PIPELINE_FRAME_SIZE   (CLIENT_AUDIO_PIPELINE_SAMPLE_RATE * CLIENT_AUDIO_PIPELINE_FRAME_MS / 1000)

Default frame size in samples at 48kHz

Definition at line 93 of file client_pipeline.h.

◆ CLIENT_AUDIO_PIPELINE_MAX_OPUS_PACKET

#define CLIENT_AUDIO_PIPELINE_MAX_OPUS_PACKET   4000

Maximum Opus packet size

Definition at line 99 of file client_pipeline.h.

◆ CLIENT_AUDIO_PIPELINE_SAMPLE_RATE

#define CLIENT_AUDIO_PIPELINE_SAMPLE_RATE   48000

Default sample rate (48kHz, native for Opus)

Definition at line 87 of file client_pipeline.h.

Typedef Documentation

◆ JitterBuffer

typedef struct JitterBuffer_ JitterBuffer

Definition at line 68 of file client_pipeline.h.

◆ OpusDecoder

typedef struct OpusDecoder OpusDecoder

Definition at line 75 of file client_pipeline.h.

◆ OpusEncoder

typedef struct OpusEncoder OpusEncoder

Definition at line 74 of file client_pipeline.h.

◆ SpeexPreprocessState

typedef struct SpeexPreprocessState_ SpeexPreprocessState

Definition at line 67 of file client_pipeline.h.

◆ webrtc_EchoCanceller3

Definition at line 71 of file client_pipeline.h.

Function Documentation

◆ client_audio_pipeline_capture()

int client_audio_pipeline_capture ( client_audio_pipeline_t *  pipeline,
const float *  input,
int  num_samples,
uint8_t *  opus_out,
int  max_opus_len 
)

Process captured audio and encode to Opus.

Parameters
pipelinePipeline instance
inputInput samples (float32, -1.0 to 1.0)
num_samplesNumber of input samples (should match frame_size)
opus_outOutput buffer for Opus packet
opus_out_sizeSize of opus_out buffer (should be >= MAX_OPUS_PACKET)
Returns
Number of bytes written to opus_out, or negative on error

Processing pipeline (when flags enabled):

  1. Convert float → int16
  2. Echo cancellation (subtract speaker output from mic input)
  3. Speex preprocessor (noise suppression, AGC, VAD)
  4. Convert int16 → float
  5. High-pass filter
  6. Low-pass filter
  7. Noise gate
  8. Compressor
  9. Opus encode

Encode already-processed audio to Opus.

In full-duplex mode, AEC3 and DSP processing are done in process_duplex(). This function just does Opus encoding.

Definition at line 446 of file client_pipeline.cpp.

447 {
448 if (!pipeline || !input || !opus_out || num_samples != pipeline->frame_size) {
449 return -1;
450 }
451
452 // Input is already processed by process_duplex() in full-duplex mode.
453 // Just encode with Opus.
454 int opus_len = opus_encode_float(pipeline->encoder, input, num_samples, opus_out, max_opus_len);
455
456 if (opus_len < 0) {
457 log_error("Opus encoding failed: %d", opus_len);
458 return -1;
459 }
460
461 return opus_len;
462}
#define log_error(...)
Log an ERROR message.
Definition log/log.h:587

References client_audio_pipeline_t::encoder, client_audio_pipeline_t::frame_size, and log_error.

◆ client_audio_pipeline_create()

client_audio_pipeline_t * client_audio_pipeline_create ( const client_audio_pipeline_config_t *  config)

Create a new client audio pipeline.

Parameters
configPipeline configuration (NULL for defaults)
Returns
New pipeline instance, or NULL on failure

Allocates and initializes all audio processing components. Must be destroyed with client_audio_pipeline_destroy().

Create a new client audio pipeline.

This function:

  • Allocates the pipeline structure
  • Initializes Opus encoder/decoder
  • Sets up WebRTC AEC3 echo cancellation
  • Configures all audio processing parameters

Definition at line 157 of file client_pipeline.cpp.

157 {
159 if (!p) {
160 log_error("Failed to allocate client audio pipeline");
161 return NULL;
162 }
163
164 // Use default config if none provided
165 if (config) {
166 p->config = *config;
167 } else {
169 }
170
171 p->flags = p->config.flags;
173 p->agc_current_gain = 1.0f;
174
175 // No mutex needed - full-duplex means single callback thread handles all AEC3
176
177 // Initialize Opus encoder/decoder first (no exceptions)
178 int opus_error = 0;
179 p->encoder = opus_encoder_create(p->config.sample_rate, 1, OPUS_APPLICATION_VOIP, &opus_error);
180 if (!p->encoder || opus_error != OPUS_OK) {
181 log_error("Failed to create Opus encoder: %d", opus_error);
182 goto error;
183 }
184 opus_encoder_ctl(p->encoder, OPUS_SET_BITRATE(p->config.opus_bitrate));
185
186 // TODO: investigate this claim from Claude: DTX stops sending frames during silence, perhaps causing audible
187 // clicks/beeps when audio resumes. For now leave it enabled.
188 opus_encoder_ctl(p->encoder, OPUS_SET_DTX(0));
189
190 // Create Opus decoder
191 p->decoder = opus_decoder_create(p->config.sample_rate, 1, &opus_error);
192 if (!p->decoder || opus_error != OPUS_OK) {
193 log_error("Failed to create Opus decoder: %d", opus_error);
194 goto error;
195 }
196
197 // Create WebRTC AEC3 Echo Cancellation
198 // AEC3 provides production-grade acoustic echo cancellation with:
199 // - Automatic network delay estimation (0-500ms)
200 // - Adaptive filtering to actual echo path
201 // - Residual echo suppression via spectral subtraction
202 // - Jitter buffer handling via side information
203 if (p->flags.echo_cancel) {
204 // Configure AEC3 for better low-frequency (bass) echo cancellation
205 webrtc::EchoCanceller3Config aec3_config;
206
207 // Increase filter length for bass frequencies (default 13 blocks = ~17ms)
208 // Bass at 80Hz has 12.5ms period, so we need at least 50+ blocks (~67ms)
209 // to properly model the echo path for low frequencies
210 aec3_config.filter.main.length_blocks = 50; // ~67ms (was 13)
211 aec3_config.filter.shadow.length_blocks = 50; // ~67ms (was 13)
212 aec3_config.filter.main_initial.length_blocks = 25; // ~33ms (was 12)
213 aec3_config.filter.shadow_initial.length_blocks = 25; // ~33ms (was 12)
214
215 // More aggressive low-frequency suppression thresholds
216 // Lower values = more aggressive echo suppression
217 aec3_config.echo_audibility.audibility_threshold_lf = 5; // (was 10)
218
219 // Create AEC3 using the factory
220 auto factory = webrtc::EchoCanceller3Factory(aec3_config);
221
222 std::unique_ptr<webrtc::EchoControl> echo_control = factory.Create(static_cast<int>(p->config.sample_rate), // 48kHz
223 1, // num_render_channels (speaker output)
224 1 // num_capture_channels (microphone input)
225 );
226
227 if (!echo_control) {
228 log_warn("Failed to create WebRTC AEC3 instance - echo cancellation unavailable");
229 p->echo_canceller = NULL;
230 } else {
231 // Successfully created AEC3 - wrap in our C++ wrapper for C compatibility
232 auto wrapper = new WebRTCAec3Wrapper();
233 wrapper->aec3 = std::move(echo_control);
234 wrapper->config = aec3_config;
235 p->echo_canceller = wrapper;
236
237 log_info("✓ WebRTC AEC3 initialized (67ms filter for bass, adaptive delay)");
238
239 // Create persistent AudioBuffer instances for AEC3
240 p->aec3_render_buffer = new webrtc::AudioBuffer(48000, 1, 48000, 1, 48000, 1);
241 p->aec3_capture_buffer = new webrtc::AudioBuffer(48000, 1, 48000, 1, 48000, 1);
242
243 auto *render_buf = static_cast<webrtc::AudioBuffer *>(p->aec3_render_buffer);
244 auto *capture_buf = static_cast<webrtc::AudioBuffer *>(p->aec3_capture_buffer);
245
246 // Zero-initialize channel data
247 float *const *render_ch = render_buf->channels();
248 float *const *capture_ch = capture_buf->channels();
249 if (render_ch && render_ch[0]) {
250 memset(render_ch[0], 0, 480 * sizeof(float)); // 10ms at 48kHz
251 }
252 if (capture_ch && capture_ch[0]) {
253 memset(capture_ch[0], 0, 480 * sizeof(float));
254 }
255
256 // Prime filterbank state with dummy processing cycle
257 render_buf->SplitIntoFrequencyBands();
258 render_buf->MergeFrequencyBands();
259 capture_buf->SplitIntoFrequencyBands();
260 capture_buf->MergeFrequencyBands();
261
262 log_info(" - AudioBuffer filterbank state initialized");
263
264 // Warm up AEC3 with 10 silent frames to initialize internal state
265 for (int warmup = 0; warmup < 10; warmup++) {
266 memset(render_ch[0], 0, 480 * sizeof(float));
267 memset(capture_ch[0], 0, 480 * sizeof(float));
268
269 render_buf->SplitIntoFrequencyBands();
270 wrapper->aec3->AnalyzeRender(render_buf);
271 render_buf->MergeFrequencyBands();
272
273 wrapper->aec3->AnalyzeCapture(capture_buf);
274 capture_buf->SplitIntoFrequencyBands();
275 wrapper->aec3->SetAudioBufferDelay(0);
276 wrapper->aec3->ProcessCapture(capture_buf, false);
277 capture_buf->MergeFrequencyBands();
278 }
279 log_info(" - AEC3 warmed up with 10 silent frames");
280 log_info(" - Persistent AudioBuffer instances created");
281 }
282 }
283
284 // Initialize debug WAV writers for AEC3 analysis (if echo_cancel enabled)
285 p->debug_wav_aec3_in = NULL;
286 p->debug_wav_aec3_out = NULL;
287 if (p->flags.echo_cancel) {
288 // Open WAV files to capture AEC3 input and output
289 char tmp[PLATFORM_MAX_PATH_LENGTH];
290 if (platform_get_temp_dir(tmp, sizeof(tmp))) {
291 char path[PLATFORM_MAX_PATH_LENGTH];
292 safe_snprintf(path, sizeof(path), "%s/aec3_input.wav", tmp);
293 p->debug_wav_aec3_in = wav_writer_open(path, 48000, 1);
294 safe_snprintf(path, sizeof(path), "%s/aec3_output.wav", tmp);
295 p->debug_wav_aec3_out = wav_writer_open(path, 48000, 1);
296 }
297 if (p->debug_wav_aec3_in) {
298 log_info("Debug: Recording AEC3 input WAV to temp dir");
299 }
300 if (p->debug_wav_aec3_out) {
301 log_info("Debug: Recording AEC3 output WAV to temp dir");
302 }
303
304 log_info("✓ AEC3 echo cancellation enabled (full-duplex mode, no ring buffer delay)");
305 }
306
307 // Initialize audio processing components (compressor, noise gate, filters)
308 // These are applied in the capture path after AEC3 and before Opus encoding
309 {
310 float sample_rate = (float)p->config.sample_rate;
311
312 // Initialize compressor with config values
313 compressor_init(&p->compressor, sample_rate);
316 log_info("✓ Capture compressor: threshold=%.1fdB, ratio=%.1f:1, makeup=+%.1fdB", p->config.comp_threshold_db,
318
319 // Initialize noise gate with config values
320 noise_gate_init(&p->noise_gate, sample_rate);
323 log_info("✓ Capture noise gate: threshold=%.4f (%.1fdB)", p->config.gate_threshold,
324 20.0f * log10f(p->config.gate_threshold + 1e-10f));
325
326 // Initialize PLAYBACK noise gate - cuts quiet received audio before speakers
327 // Very low threshold - only cut actual silence, not quiet voice audio
328 // The server sends audio with RMS=0.01-0.02, so threshold must be below that
329 noise_gate_init(&p->playback_noise_gate, sample_rate);
331 0.002f, // -54dB threshold - only cut near-silence
332 1.0f, // 1ms attack - fast open
333 50.0f, // 50ms release - smooth close
334 0.4f); // Hysteresis
335 log_info("✓ Playback noise gate: threshold=0.002 (-54dB)");
336
337 // Initialize highpass filter (removes low-frequency rumble)
338 highpass_filter_init(&p->highpass, p->config.highpass_hz, sample_rate);
339 log_info("✓ Capture highpass filter: %.1f Hz", p->config.highpass_hz);
340
341 // Initialize lowpass filter (removes high-frequency hiss)
342 lowpass_filter_init(&p->lowpass, p->config.lowpass_hz, sample_rate);
343 log_info("✓ Capture lowpass filter: %.1f Hz", p->config.lowpass_hz);
344 }
345
346 p->initialized = true;
347
348 // Initialize startup fade-in to prevent initial microphone click
349 // 200ms at 48kHz = 9600 samples - gradual ramp from silence to full volume
350 // Longer fade-in (200ms vs 50ms) gives much smoother transition without audible pop
351 p->capture_fadein_remaining = (p->config.sample_rate * 200) / 1000; // 200ms worth of samples
352 log_info("✓ Capture fade-in: %d samples (200ms)", p->capture_fadein_remaining);
353
354 log_info("Audio pipeline created: %dHz, %dms frames, %dkbps Opus", p->config.sample_rate, p->config.frame_size_ns,
355 p->config.opus_bitrate / 1000);
356
357 (void)NAMED_REGISTER_CLIENT_AUDIO_PIPELINE(p, "audio_pipeline", NULL);
358
359 return p;
360
361error:
362 if (p->encoder)
363 opus_encoder_destroy(p->encoder);
364 if (p->decoder)
365 opus_decoder_destroy(p->decoder);
366 if (p->echo_canceller) {
367 delete static_cast<WebRTCAec3Wrapper *>(p->echo_canceller);
368 }
369 SAFE_FREE(p);
370 return NULL;
371}
client_audio_pipeline_config_t client_audio_pipeline_default_config(void)
Get default configuration.
void noise_gate_init(noise_gate_t *gate, float sample_rate)
Initialize a noise gate.
Definition mixer.c:925
void compressor_init(compressor_t *comp, float sample_rate)
Initialize a compressor.
Definition mixer.c:86
void compressor_set_params(compressor_t *comp, float threshold_dB, float ratio, uint64_t attack_ns, uint64_t release_ns, float makeup_dB)
Set compressor parameters.
Definition mixer.c:97
void highpass_filter_init(highpass_filter_t *filter, float cutoff_hz, float sample_rate)
Initialize a high-pass filter.
Definition mixer.c:1010
void lowpass_filter_init(lowpass_filter_t *filter, float cutoff_hz, float sample_rate)
Initialize a low-pass filter.
Definition mixer.c:1060
void noise_gate_set_params(noise_gate_t *gate, float threshold, uint64_t attack_ns, uint64_t release_ns, float hysteresis)
Set noise gate parameters.
Definition mixer.c:939
@ OPUS_APPLICATION_VOIP
Voice over IP (optimized for speech)
Definition opus.h:77
#define SAFE_FREE(ptr)
Definition common.h:376
#define SAFE_CALLOC(count, size, cast)
Definition common.h:274
#define NAMED_REGISTER_CLIENT_AUDIO_PIPELINE(pipeline, name, parent_ptr)
Register a client audio pipeline with automatic format specifier.
#define log_warn(...)
Log a WARN message.
Definition log/log.h:574
#define log_info(...)
Log an INFO message.
Definition log/log.h:561
int safe_snprintf(char *buffer, size_t buffer_size, const char *format,...)
Safe formatted string printing to buffer.
Definition system.c:148
bool platform_get_temp_dir(char *temp_dir, size_t path_size)
Get the system temporary directory path.
Definition util.c:74
C++ wrapper for WebRTC AEC3 (opaque to C code)
client_audio_pipeline_flags_t flags
Client audio pipeline state.
client_audio_pipeline_config_t config
client_audio_pipeline_flags_t flags
highpass_filter_t highpass
#define PLATFORM_MAX_PATH_LENGTH
Definition system.c:69
wav_writer_t * wav_writer_open(const char *filepath, int sample_rate, int channels)
Open WAV file for writing.
Definition wav_writer.c:48

References client_audio_pipeline_t::aec3_capture_buffer, client_audio_pipeline_t::aec3_render_buffer, client_audio_pipeline_t::agc_current_gain, client_audio_pipeline_t::capture_fadein_remaining, client_audio_pipeline_default_config(), client_audio_pipeline_config_t::comp_attack_ns, client_audio_pipeline_config_t::comp_makeup_db, client_audio_pipeline_config_t::comp_ratio, client_audio_pipeline_config_t::comp_release_ns, client_audio_pipeline_config_t::comp_threshold_db, client_audio_pipeline_t::compressor, compressor_init(), compressor_set_params(), client_audio_pipeline_t::config, client_audio_pipeline_t::debug_wav_aec3_in, client_audio_pipeline_t::debug_wav_aec3_out, client_audio_pipeline_t::decoder, client_audio_pipeline_flags_t::echo_cancel, client_audio_pipeline_t::echo_canceller, client_audio_pipeline_t::encoder, client_audio_pipeline_config_t::flags, client_audio_pipeline_t::flags, client_audio_pipeline_t::frame_size, client_audio_pipeline_config_t::frame_size_ns, client_audio_pipeline_config_t::gate_attack_ns, client_audio_pipeline_config_t::gate_hysteresis, client_audio_pipeline_config_t::gate_release_ns, client_audio_pipeline_config_t::gate_threshold, client_audio_pipeline_t::highpass, highpass_filter_init(), client_audio_pipeline_config_t::highpass_hz, client_audio_pipeline_t::initialized, log_error, log_info, log_warn, client_audio_pipeline_t::lowpass, lowpass_filter_init(), client_audio_pipeline_config_t::lowpass_hz, NAMED_REGISTER_CLIENT_AUDIO_PIPELINE, client_audio_pipeline_t::noise_gate, noise_gate_init(), noise_gate_set_params(), OPUS_APPLICATION_VOIP, client_audio_pipeline_config_t::opus_bitrate, platform_get_temp_dir(), PLATFORM_MAX_PATH_LENGTH, client_audio_pipeline_t::playback_noise_gate, SAFE_CALLOC, SAFE_FREE, safe_snprintf(), client_audio_pipeline_config_t::sample_rate, and wav_writer_open().

Referenced by audio_client_init().

◆ client_audio_pipeline_default_config()

client_audio_pipeline_config_t client_audio_pipeline_default_config ( void  )

Get default configuration.

Returns
Configuration with sensible defaults for voice chat

Definition at line 104 of file client_pipeline.cpp.

104 {
107 .frame_size_ns = CLIENT_AUDIO_PIPELINE_FRAME_MS,
108 .opus_bitrate = 24000,
109
110 .echo_filter_ns = 250,
111
112 .noise_suppress_db = -25,
113 .agc_level = 16000, // Increased from 8000 for louder output
114 .agc_max_gain = 35, // Increased from 30 dB to handle very quiet mics (35 dB = ~56x gain)
115
116 // Jitter margin: wait this long before starting playback
117 // Lower = less latency but more risk of underruns
118 // Must match AUDIO_JITTER_BUFFER_THRESHOLD in ringbuffer.h.
119 .jitter_margin_ns = 20, // 20ms = 1 Opus packet (optimized for LAN)
120
121 // Higher cutoff to cut low-frequency rumble and feedback
122 .highpass_hz = 150.0f, // Was 80Hz, increased to break rumble feedback loop
123 .lowpass_hz = 8000.0f,
124
125 // Compressor: only compress loud peaks, moderate makeup for volume
126 // User reported clipping with +6dB makeup gain
127 .comp_threshold_db = -12.0f, // Compress above -12dB (was -6dB)
128 .comp_ratio = 3.0f, // Gentler 3:1 ratio
129 .comp_attack_ns = 5 * NS_PER_MS_INT, // Fast attack for peaks
130 .comp_release_ns = 150 * NS_PER_MS_INT, // Slower release
131 .comp_makeup_db = 6.0f, // Increased from 2dB for more output volume
132
133 // Noise gate: VERY aggressive to cut quiet background audio completely
134 // User feedback: "don't amplify or play quiet background audio at all"
135 .gate_threshold = 0.08f, // -22dB threshold (was 0.02/-34dB) - cuts quiet audio hard
136 .gate_attack_ns = 500 * NS_PER_US_INT, // Very fast attack
137 .gate_release_ns = 30 * NS_PER_MS_INT, // Fast release (was 50ms)
138 .gate_hysteresis = 0.3f, // Tighter hysteresis = stays closed longer
139
141 };
142}
#define CLIENT_AUDIO_PIPELINE_FRAME_MS
#define CLIENT_AUDIO_PIPELINE_FLAGS_ALL
Default flags with all processing enabled.
#define CLIENT_AUDIO_PIPELINE_SAMPLE_RATE
#define NS_PER_MS_INT
Definition time.h:156
#define NS_PER_US_INT
Definition time.h:155
Pipeline configuration parameters.

References CLIENT_AUDIO_PIPELINE_FLAGS_ALL, CLIENT_AUDIO_PIPELINE_FRAME_MS, CLIENT_AUDIO_PIPELINE_SAMPLE_RATE, NS_PER_MS_INT, NS_PER_US_INT, and client_audio_pipeline_config_t::sample_rate.

Referenced by audio_client_init(), and client_audio_pipeline_create().

◆ client_audio_pipeline_destroy()

void client_audio_pipeline_destroy ( client_audio_pipeline_t *  pipeline)

Destroy a client audio pipeline.

Parameters
pipelinePipeline to destroy (can be NULL)

Frees all resources including SpeexDSP states, Opus codec, and all work buffers.

Definition at line 373 of file client_pipeline.cpp.

373 {
374 if (!pipeline)
375 return;
376
377 // Clean up WebRTC AEC3 AudioBuffer instances
378 if (pipeline->aec3_render_buffer) {
379 delete static_cast<webrtc::AudioBuffer *>(pipeline->aec3_render_buffer);
380 pipeline->aec3_render_buffer = NULL;
381 }
382 if (pipeline->aec3_capture_buffer) {
383 delete static_cast<webrtc::AudioBuffer *>(pipeline->aec3_capture_buffer);
384 pipeline->aec3_capture_buffer = NULL;
385 }
386
387 // Clean up WebRTC AEC3
388 if (pipeline->echo_canceller) {
389 delete static_cast<WebRTCAec3Wrapper *>(pipeline->echo_canceller);
390 pipeline->echo_canceller = NULL;
391 }
392
393 // Clean up Opus
394 if (pipeline->encoder) {
395 opus_encoder_destroy(pipeline->encoder);
396 pipeline->encoder = NULL;
397 }
398 if (pipeline->decoder) {
399 opus_decoder_destroy(pipeline->decoder);
400 pipeline->decoder = NULL;
401 }
402
403 // Clean up debug WAV writers
404 if (pipeline->debug_wav_aec3_in) {
406 pipeline->debug_wav_aec3_in = NULL;
407 }
408 if (pipeline->debug_wav_aec3_out) {
410 pipeline->debug_wav_aec3_out = NULL;
411 }
412
413 NAMED_UNREGISTER(pipeline);
414
415 SAFE_FREE(pipeline);
416}
#define NAMED_UNREGISTER(ptr)
Unregister a pointer.
WAV file writer context.
Definition wav_writer.h:23
void wav_writer_close(wav_writer_t *writer)
Close WAV file and finalize header.
Definition wav_writer.c:112

References client_audio_pipeline_t::aec3_capture_buffer, client_audio_pipeline_t::aec3_render_buffer, client_audio_pipeline_t::debug_wav_aec3_in, client_audio_pipeline_t::debug_wav_aec3_out, client_audio_pipeline_t::decoder, client_audio_pipeline_t::echo_canceller, client_audio_pipeline_t::encoder, NAMED_UNREGISTER, SAFE_FREE, and wav_writer_close().

Referenced by audio_cleanup().

◆ client_audio_pipeline_get_flags()

client_audio_pipeline_flags_t client_audio_pipeline_get_flags ( client_audio_pipeline_t *  pipeline)

Get current component enable flags.

Parameters
pipelinePipeline instance
Returns
Current flags

Definition at line 429 of file client_pipeline.cpp.

429 {
430 if (!pipeline)
432 // No mutex needed - flags are only written from main thread during setup
433 return pipeline->flags;
434}
#define CLIENT_AUDIO_PIPELINE_FLAGS_MINIMAL
Minimal flags for testing (only codec, no processing)

References CLIENT_AUDIO_PIPELINE_FLAGS_MINIMAL, and client_audio_pipeline_t::flags.

◆ client_audio_pipeline_get_playback_frame()

int client_audio_pipeline_get_playback_frame ( client_audio_pipeline_t *  pipeline,
float *  output,
int  num_samples 
)

Get audio frame from jitter buffer for playback callback.

Parameters
pipelinePipeline instance
outputOutput buffer for decoded samples (float32)
num_samplesNumber of samples to retrieve
Returns
Number of samples written, or negative on error

This is called from the audio output callback to get the next frame of audio. It pulls from the jitter buffer and decodes.

Get a processed playback frame (currently just returns decoded frame)

Definition at line 497 of file client_pipeline.cpp.

497 {
498 if (!pipeline || !output) {
499 return -1;
500 }
501
502 // No mutex needed - this is a placeholder
503 memset(output, 0, num_samples * sizeof(float));
504 return num_samples;
505}

◆ client_audio_pipeline_jitter_margin()

int client_audio_pipeline_jitter_margin ( client_audio_pipeline_t *  pipeline)

Get jitter buffer margin (buffered time in ms)

Parameters
pipelinePipeline instance
Returns
Current jitter buffer margin in milliseconds

Get jitter buffer margin

Definition at line 682 of file client_pipeline.cpp.

682 {
683 if (!pipeline)
684 return 0;
685 return pipeline->config.jitter_margin_ns;
686}

References client_audio_pipeline_t::config, and client_audio_pipeline_config_t::jitter_margin_ns.

◆ client_audio_pipeline_playback()

int client_audio_pipeline_playback ( client_audio_pipeline_t *  pipeline,
const uint8_t *  opus_in,
int  opus_len,
float *  output,
int  num_samples 
)

Decode Opus packet and process for playback.

Parameters
pipelinePipeline instance
opus_inOpus packet data (can be NULL for packet loss concealment)
opus_lenOpus packet length (0 if opus_in is NULL)
outputOutput buffer for decoded samples (float32)
max_samplesMaximum samples to write to output
Returns
Number of samples written to output, or negative on error

Processing pipeline (when flags enabled):

  1. Put Opus packet into jitter buffer
  2. Get packet from jitter buffer (handles reordering, loss)
  3. Opus decode
  4. Soft clipping

Note: Output should be fed to speakers AND to echo reference via client_audio_pipeline_feed_echo_ref().

Process network playback (decode and register with echo canceller as reference)

Definition at line 467 of file client_pipeline.cpp.

468 {
469 if (!pipeline || !opus_in || !output) {
470 return -1;
471 }
472
473 // No mutex needed - Opus decoder is only used from this thread
474
475 // Decode Opus
476 int decoded_samples = opus_decode_float(pipeline->decoder, opus_in, opus_len, output, num_samples, 0);
477
478 if (decoded_samples < 0) {
479 log_error("Opus decoding failed: %d", decoded_samples);
480 return -1;
481 }
482
483 // Apply playback noise gate - cut quiet background audio before it reaches speakers
484 if (decoded_samples > 0) {
485 noise_gate_process_buffer(&pipeline->playback_noise_gate, output, decoded_samples);
486 }
487
488 // NOTE: Render signal is queued to AEC3 in output_callback() when audio plays,
489 // not here. The capture thread drains the queue and processes AEC3.
490
491 return decoded_samples;
492}
void noise_gate_process_buffer(noise_gate_t *gate, float *buffer, int num_samples)
Process a buffer of samples through noise gate.
Definition mixer.c:982

References client_audio_pipeline_t::decoder, log_error, noise_gate_process_buffer(), and client_audio_pipeline_t::playback_noise_gate.

Referenced by audio_decode_opus().

◆ client_audio_pipeline_process_duplex()

void client_audio_pipeline_process_duplex ( client_audio_pipeline_t *  pipeline,
const float *  render_samples,
int  render_count,
const float *  capture_samples,
int  capture_count,
float *  processed_output 
)

Process AEC3 inline in full-duplex callback.

Parameters
pipelinePipeline instance
render_samplesAudio samples being played to speakers RIGHT NOW
render_countNumber of render samples
capture_samplesAudio samples from microphone RIGHT NOW
capture_countNumber of capture samples
processed_outputOutput buffer for processed capture (must be capture_count size)

Called from PortAudio's single full-duplex callback where render and capture happen at the EXACT same instant. No ring buffers, no timing mismatch.

Processing:

  1. AEC3 AnalyzeRender on render_samples
  2. AEC3 AnalyzeCapture + ProcessCapture on capture_samples
  3. Apply highpass, lowpass, noise gate, compressor

REAL-TIME SAFE: No mutexes, no allocations, no blocking.

Process AEC3 inline in full-duplex callback (REAL-TIME SAFE).

This is the PROFESSIONAL approach to AEC3 timing:

  • Called from a single PortAudio full-duplex callback
  • render_samples = what is being played to speakers RIGHT NOW
  • capture_samples = what microphone captured RIGHT NOW
  • Perfect synchronization - no timing mismatch possible

This function does ALL AEC3 processing inline:

  1. AnalyzeRender on render samples (speaker output)
  2. AnalyzeCapture + ProcessCapture on capture samples
  3. Apply filters, noise gate, compressor

Returns processed capture samples in processed_output. Opus encoding is done separately by the encoding thread.

Definition at line 524 of file client_pipeline.cpp.

526 {
527 if (!pipeline || !processed_output)
528 return;
529
530 // Copy capture samples to output buffer for processing
531 if (capture_samples && capture_count > 0) {
532 memcpy(processed_output, capture_samples, capture_count * sizeof(float));
533 } else {
534 memset(processed_output, 0, capture_count * sizeof(float));
535 return;
536 }
537
538 // Check for AEC3 bypass
539 static int bypass_aec3 = -1;
540 if (bypass_aec3 == -1) {
541 const char *env = platform_getenv("BYPASS_AEC3");
542 bypass_aec3 = (env && (strcmp(env, "1") == 0 || strcmp(env, "true") == 0)) ? 1 : 0;
543 if (bypass_aec3) {
544 log_warn("AEC3 BYPASSED (full-duplex mode) via BYPASS_AEC3=1");
545 }
546 }
547
548 // Debug WAV recording
549 if (pipeline->debug_wav_aec3_in) {
550 wav_writer_write((wav_writer_t *)pipeline->debug_wav_aec3_in, capture_samples, capture_count);
551 }
552
553 // Apply startup fade-in using smoothstep curve
554 if (pipeline->capture_fadein_remaining > 0) {
555 const int total_fadein_samples = (pipeline->config.sample_rate * 200) / 1000;
556 for (int i = 0; i < capture_count && pipeline->capture_fadein_remaining > 0; i++) {
557 float progress = 1.0f - ((float)pipeline->capture_fadein_remaining / (float)total_fadein_samples);
558 float gain = smoothstep(progress);
559 processed_output[i] *= gain;
560 pipeline->capture_fadein_remaining--;
561 }
562 }
563
564 // WebRTC AEC3 processing - INLINE, no ring buffer, no mutex
565 if (!bypass_aec3 && pipeline->flags.echo_cancel && pipeline->echo_canceller) {
566 auto wrapper = static_cast<WebRTCAec3Wrapper *>(pipeline->echo_canceller);
567 if (wrapper && wrapper->aec3) {
568 const int webrtc_frame_size = 480; // 10ms at 48kHz
569
570 auto *render_buf = static_cast<webrtc::AudioBuffer *>(pipeline->aec3_render_buffer);
571 auto *capture_buf = static_cast<webrtc::AudioBuffer *>(pipeline->aec3_capture_buffer);
572
573 if (render_buf && capture_buf) {
574 float *const *render_channels = render_buf->channels();
575 float *const *capture_channels = capture_buf->channels();
576
577 if (render_channels && render_channels[0] && capture_channels && capture_channels[0]) {
578 // Verify render_samples is valid before accessing
579 if (!render_samples && render_count > 0) {
580 log_warn_every(1000000, "AEC3: render_samples is NULL but render_count=%d", render_count);
581 return;
582 }
583
584 // Process in 10ms chunks (AEC3 requirement)
585 int render_offset = 0;
586 int capture_offset = 0;
587
588 // Process only complete 480-sample frames. The worker thread guarantees
589 // render_count == capture_count and both are multiples of 480, so no
590 // partial frames should occur. If they do (defensive), the tail samples
591 // pass through unprocessed rather than corrupting AEC3 with padded data.
592 while (render_offset + webrtc_frame_size <= render_count &&
593 capture_offset + webrtc_frame_size <= capture_count) {
594
595 // STEP 1: Feed render signal (what's playing to speakers)
596 copy_buffer_with_gain(&render_samples[render_offset], render_channels[0], webrtc_frame_size, 32768.0f);
597 render_buf->SplitIntoFrequencyBands();
598 wrapper->aec3->AnalyzeRender(render_buf);
599 render_buf->MergeFrequencyBands();
600 g_render_frames_fed.fetch_add(1, std::memory_order_relaxed);
601 render_offset += webrtc_frame_size;
602
603 // STEP 2: Process capture (microphone input)
604 copy_buffer_with_gain(&processed_output[capture_offset], capture_channels[0], webrtc_frame_size, 32768.0f);
605 wrapper->aec3->AnalyzeCapture(capture_buf);
606 capture_buf->SplitIntoFrequencyBands();
607 wrapper->aec3->ProcessCapture(capture_buf, false);
608 capture_buf->MergeFrequencyBands();
609
610 // Scale back to float range and soft clip
611 for (int j = 0; j < webrtc_frame_size; j++) {
612 float sample = capture_channels[0][j] / 32768.0f;
613 processed_output[capture_offset + j] = soft_clip(sample, 0.6f, 2.5f);
614 }
615
616 // Log AEC3 metrics periodically
617 static int duplex_log_count = 0;
618 if (++duplex_log_count % 100 == 1) {
619 webrtc::EchoControl::Metrics metrics = wrapper->aec3->GetMetrics();
620 log_info("AEC3 DUPLEX: ERL=%.1f ERLE=%.1f delay=%dms", metrics.echo_return_loss,
621 metrics.echo_return_loss_enhancement, metrics.delay_ms);
622 audio_analysis_set_aec3_metrics(metrics.echo_return_loss, metrics.echo_return_loss_enhancement,
623 metrics.delay_ms);
624 }
625 capture_offset += webrtc_frame_size;
626 }
627 }
628 }
629 }
630 }
631
632 // Debug WAV recording (after AEC3)
633 if (pipeline->debug_wav_aec3_out) {
634 wav_writer_write((wav_writer_t *)pipeline->debug_wav_aec3_out, processed_output, capture_count);
635 }
636
637 // Adjust gain from this frame's level. Applying agc_max_gain as a constant
638 // pre-gain amplifies silence and background noise by the maximum amount.
639 if (pipeline->flags.agc && capture_count > 0) {
640 double energy = 0.0;
641 for (int i = 0; i < capture_count; i++) {
642 const double sample = processed_output[i];
643 energy += sample * sample;
644 }
645 const float rms = static_cast<float>(std::sqrt(energy / capture_count));
646 const float target_rms = std::clamp(pipeline->config.agc_level / 32768.0f, 0.05f, 0.5f);
647 const float max_gain_db = std::clamp(static_cast<float>(pipeline->config.agc_max_gain), 0.0f, 36.0f);
648 const float max_gain = std::pow(10.0f, max_gain_db / 20.0f);
649 const float desired_gain = rms > 1e-5f ? std::clamp(target_rms / rms, 1.0f, max_gain) : 1.0f;
650
651 // Attenuate a loud frame quickly and raise quiet input more gradually.
652 const float smoothing = desired_gain < pipeline->agc_current_gain ? 0.25f : 0.05f;
653 pipeline->agc_current_gain += smoothing * (desired_gain - pipeline->agc_current_gain);
654 for (int i = 0; i < capture_count; i++) {
655 processed_output[i] *= pipeline->agc_current_gain;
656 }
657 }
658
659 // Apply capture processing chain: filters, noise gate, compressor
660 if (pipeline->flags.highpass) {
661 highpass_filter_process_buffer(&pipeline->highpass, processed_output, capture_count);
662 }
663 if (pipeline->flags.lowpass) {
664 lowpass_filter_process_buffer(&pipeline->lowpass, processed_output, capture_count);
665 }
666 if (pipeline->flags.noise_gate) {
667 noise_gate_process_buffer(&pipeline->noise_gate, processed_output, capture_count);
668 }
669 if (pipeline->flags.compressor) {
670 for (int i = 0; i < capture_count; i++) {
671 float gain = compressor_process_sample(&pipeline->compressor, processed_output[i]);
672 processed_output[i] *= gain;
673 }
674 // Apply soft clipping after compressor - threshold=0.7 gives 3dB headroom
675 soft_clip_buffer(processed_output, capture_count, 0.7f, 3.0f);
676 }
677}
void audio_analysis_set_aec3_metrics(double echo_return_loss, double echo_return_loss_enhancement, uint64_t delay_ns)
Set AEC3 echo cancellation metrics.
Definition analysis.c:510
void copy_buffer_with_gain(const float *src, float *dst, int count, float gain)
Copy buffer with gain scaling.
Definition mixer.c:1218
void soft_clip_buffer(float *buffer, int num_samples, float threshold, float steepness)
Apply soft clipping to a buffer.
Definition mixer.c:1122
float soft_clip(float sample, float threshold, float steepness)
Apply soft clipping to a sample.
Definition mixer.c:1109
float smoothstep(float t)
Compute smoothstep interpolation.
Definition mixer.c:1136
void lowpass_filter_process_buffer(lowpass_filter_t *filter, float *buffer, int num_samples)
Process a buffer of samples through low-pass filter.
Definition mixer.c:1095
void highpass_filter_process_buffer(highpass_filter_t *filter, float *buffer, int num_samples)
Process a buffer of samples through high-pass filter.
Definition mixer.c:1046
float compressor_process_sample(compressor_t *comp, float sidechain)
Process a single sample through compressor.
Definition mixer.c:131
const char * platform_getenv(const char *name)
Get an environment variable value.
Definition wasm/system.c:39
#define log_warn_every(interval_us, fmt,...)
Rate-limited WARN logging.
Definition log/log.h:708
int wav_writer_write(wav_writer_t *writer, const float *samples, int num_samples)
Write audio samples to WAV file.
Definition wav_writer.c:94

References client_audio_pipeline_t::aec3_capture_buffer, client_audio_pipeline_t::aec3_render_buffer, client_audio_pipeline_flags_t::agc, client_audio_pipeline_t::agc_current_gain, client_audio_pipeline_config_t::agc_level, client_audio_pipeline_config_t::agc_max_gain, audio_analysis_set_aec3_metrics(), client_audio_pipeline_t::capture_fadein_remaining, client_audio_pipeline_flags_t::compressor, client_audio_pipeline_t::compressor, compressor_process_sample(), client_audio_pipeline_t::config, copy_buffer_with_gain(), client_audio_pipeline_t::debug_wav_aec3_in, client_audio_pipeline_t::debug_wav_aec3_out, client_audio_pipeline_flags_t::echo_cancel, client_audio_pipeline_t::echo_canceller, client_audio_pipeline_t::flags, client_audio_pipeline_flags_t::highpass, client_audio_pipeline_t::highpass, highpass_filter_process_buffer(), log_info, log_warn, log_warn_every, client_audio_pipeline_flags_t::lowpass, client_audio_pipeline_t::lowpass, lowpass_filter_process_buffer(), client_audio_pipeline_flags_t::noise_gate, client_audio_pipeline_t::noise_gate, noise_gate_process_buffer(), platform_getenv(), client_audio_pipeline_config_t::sample_rate, smoothstep(), soft_clip(), soft_clip_buffer(), and wav_writer_write().

◆ client_audio_pipeline_reset()

void client_audio_pipeline_reset ( client_audio_pipeline_t *  pipeline)

Reset pipeline state.

Parameters
pipelinePipeline instance

Resets echo canceller, jitter buffer, and filter states. Call when starting a new audio session.

Reset pipeline state

Definition at line 691 of file client_pipeline.cpp.

691 {
692 if (!pipeline)
693 return;
694
695 // Reset global counters
696 g_render_frames_fed.store(0, std::memory_order_relaxed);
697 g_max_render_rms.store(0.0f, std::memory_order_relaxed);
698 pipeline->agc_current_gain = 1.0f;
699
700 log_info("Pipeline state reset");
701}

References client_audio_pipeline_t::agc_current_gain, and log_info.

◆ client_audio_pipeline_set_flags()

void client_audio_pipeline_set_flags ( client_audio_pipeline_t *  pipeline,
client_audio_pipeline_flags_t  flags 
)

Set component enable flags.

Parameters
pipelinePipeline instance
flagsNew flags to set

Thread-safe. Changes take effect on next capture/playback call.

Definition at line 422 of file client_pipeline.cpp.

422 {
423 if (!pipeline)
424 return;
425 // No mutex needed - flags are only read by capture thread
426 pipeline->flags = flags;
427}

References client_audio_pipeline_t::flags.

◆ client_audio_pipeline_voice_detected()

bool client_audio_pipeline_voice_detected ( client_audio_pipeline_t *  pipeline)

Check if VAD detected voice activity in last capture.

Parameters
pipelinePipeline instance
Returns
true if voice was detected, false otherwise