Reconstruction
Light Commands crosses physical modalities: the attack carries an audio waveform in the intensity of a light beam rather than through air as sound. The attacker amplitude-modulates light with a chosen voice command and aims it at the aperture of a target MEMS microphone.
The researchers found that microphone hardware can respond to incident light and produce an electrical output corresponding to the modulation. Downstream voice-recognition software then receives what looks like microphone audio even though no ordinary acoustic command was spoken. In their tests, commands were injected into products using Alexa, Siri, Portal and Google Assistant, with demonstrations at distances up to 110 meters under favorable line-of-sight conditions.
The optical injection is only the first boundary. Consequence depends on what the voice interface is authorized to do without additional confirmation. The research therefore combines a sensor-transduction flaw with an authorization problem; it is not evidence that every MEMS microphone or current assistant is vulnerable at the published maximum range.
Mechanism & boundary
- 01
Prepare a voice command waveform
The attacker selects audio that the target voice-control stack would normally interpret as an instruction.
Boundary: attacker intent / audio command
- 02
Modulate the intensity of a light source
The audio waveform is encoded into rapid changes in optical intensity.
Boundary: audio signal / optical carrier
- 03
Aim the light at the microphone
A line-of-sight beam reaches the microphone aperture from outside the device, potentially through a window.
Boundary: external optical path / sensor hardware
- 04
Convert optical modulation into microphone output
The MEMS microphone and associated electronics respond to the modulated light and produce a signal resembling injected audio.
Boundary: light / electrical microphone signal
- 05
Execute whatever voice authority permits
The assistant processes the injected signal as a command; impact depends on authentication, confirmation and connected-device privileges.
Boundary: sensor input / authorized action
Claims & evidence
reported findingsupported
The researchers reported command injection against multiple voice-controlled products at distances up to 110 meters in tested conditions, including from another building.
reported findingsupported
Light Commands demonstrated that amplitude-modulated light can induce microphone output corresponding to attacker-chosen audio and thereby inject voice commands.
Implications
Light Commands shows that a sensor's accepted physical modality can be wider than its intended one. The demonstrated risk is remote command injection under specific optical and hardware conditions; downstream harm depends on weak voice-action authorization. A resilient design must therefore protect both the transducer and the authority granted to its output.
Controls & mitigations
- Block or attenuate direct light paths to microphone elements where hardware design permits it.
- Use multi-sensor consistency checks or other hardware defenses so a signal present on one optically exposed microphone is not automatically trusted.
- Require authentication or explicit confirmation for consequential voice actions, limiting damage even if a command reaches the speech recognizer.
What remains unknown
- Success depends on line of sight, aiming, optical power, microphone construction, device enclosure and the voice system's command behavior.
- The published 110-meter result is a tested maximum under specific conditions, not a universal operating distance.
- The public research reported no evidence of malicious exploitation in the wild at the time and does not establish current susceptibility of every tested product line.