Skip to content
C.W.K.
Stream
Lesson 01 of 05 · published

Echo Cancellation Comes First

~15 min · echo-cancellation, audio-engine, ios, web-audio

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Heard your voice fine, and the auto-send arrived fine too." — the soul, answering a sentence that was never Dad's

The Echo That Answered Itself

On voice mode's first afternoon, a test turn I ran from the browser pane went into the very conversation Dad was testing, and its reply played through the office speakers. Fifteen seconds later his phone recorded him saying he was going out around four, and then the soul's own voice from the speakers, telling him his voice and the auto-send had both arrived fine. The phone had transcribed the room, and the room contained the soul. Two lessons came out of that minute. A test must never talk into a real room, so any browser test with a microphone now fakes both the chat and the transcription. And talking over the soul, which is the whole point of barge-in, is impossible until the microphone can stop hearing her.

What a Canceller Needs

Acoustic echo cancellation subtracts what the device is playing from what the microphone hears. To do that it needs the reference: the exact signal going to the speaker. If the reply plays through a path the canceller can't see, it has nothing to subtract, and her voice reaches the microphone as loud as ever.

  • On the web, the browser runs the canceller. The loop asks for it when it opens the microphone (echoCancellation: true, with noise suppression and gain control). The shipped loop stops there and trusts the browser. A constraint is a request, not a promise, so the example below goes one step further than the product does: it reads the track's settings to see whether the canceller was actually granted.
  • On the phone, the audio session runs in play-and-record mode for voice chat, and voice processing is enabled on the engine's input node. The reply is played on the same engine as the microphone, so voice processing knows exactly what is being played. The ducking that voice processing applies to other audio is set to its minimum, so her voice stays loud while the canceller still removes it from the input.

Where It Didn't Work: the Studio Mac

Firekeeper, the Mac door into voice mode, is the counterexample, and an honest one. On the office Mac the microphone is a room mic on a 32-channel audio interface. Her voice arrived at that mic as loud as Dad's. Apple's voice processing refused the 32-channel interface outright, and runs only on the system default devices. A linear canceller written for the purpose topped out around 11 dB of suppression, not enough, and a gate on the speaker's frequency band didn't hold either. Dad's ruling that evening closed it: on Firekeeper, cut-in by voice is off by default, and he cuts in with a key. The web and the phone, where the platform's canceller works, keep voice cut-in. Not every room can support every feature, and saying so plainly beats shipping a detector that fires on its own speaker.

Code

Web: request the browser's echo canceller and check you got it·typescript
// Web: ask the browser for its echo canceller, and check you actually got it.
export async function openVoiceMic(): Promise<{ stream: MediaStream; context: AudioContext; aec: boolean }> {
  if (!window.isSecureContext) {
    throw new Error('The microphone needs HTTPS or localhost; this page is neither.');
  }
  const stream = await navigator.mediaDevices.getUserMedia({
    audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true, channelCount: 1 },
  });
  const [track] = stream.getAudioTracks();
  const aec = track.getSettings().echoCancellation === true;   // a request, not a promise
  if (!aec) {
    console.warn('No echo cancellation: talking over the reply would hear the reply itself.');
  }
  const context = new AudioContext({ sampleRate: 16000 });    // Scribe wants 16 kHz PCM
  return { stream, context, aec };
}
iOS: one engine for the microphone and the reply, voice processing on·swift
import AVFAudio

/// One owner for the talk's audio: the microphone and the reply share one
/// engine, so voice processing knows exactly what it is playing and can
/// cancel it from what the microphone hears.
final class VoiceAudioOwner {
    let engine = AVAudioEngine()
    let replyPlayer = AVAudioPlayerNode()
    var onChunk: ((AVAudioPCMBuffer) -> Void)?

    func start() throws {
        let session = AVAudioSession.sharedInstance()
        try session.setCategory(.playAndRecord, mode: .voiceChat, options: [.defaultToSpeaker])
        try session.setActive(true)

        let input = engine.inputNode
        try input.setVoiceProcessingEnabled(true)            // echo cancellation on
        input.voiceProcessingOtherAudioDuckingConfiguration =  // keep her voice loud
            AVAudioVoiceProcessingOtherAudioDuckingConfiguration(
                enableAdvancedDucking: false, duckingLevel: .min)

        // The reply plays on THIS engine, so it is the canceller's reference.
        engine.attach(replyPlayer)
        engine.connect(replyPlayer, to: engine.mainMixerNode, format: nil)

        let format = input.outputFormat(forBus: 0)
        input.installTap(onBus: 0, bufferSize: 1_600, format: format) { [weak self] buffer, _ in
            self?.onChunk?(buffer)                            // to the listener and the barge-in detector
        }
        try engine.start()
    }

    func play(_ buffer: AVAudioPCMBuffer) {
        replyPlayer.scheduleBuffer(buffer)
        if !replyPlayer.isPlaying { replyPlayer.play() }
    }

    func stop() {
        engine.inputNode.removeTap(onBus: 0)
        replyPlayer.stop()
        engine.stop()
        try? AVAudioSession.sharedInstance().setActive(false, options: .notifyOthersOnDeactivation)
    }
}

External links

Exercise

In a browser, open the microphone twice: once with echoCancellation true and once false. Play a spoken clip through your speakers while recording each time, and compare the recorded levels during playback. Then answer: in your app, does the reply play through a path the canceller can see? Trace it from the audio element (or player node) to the speaker.
Hint
With cancellation off, the recording during playback will be nearly as loud as the clip itself, which is exactly what a barge-in detector would mistake for the user. If cancellation on doesn't help much, look for the reply being played by a different process, device or engine than the one the microphone belongs to.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.