MIT developed software that can recover sounds from video of vibrating objects
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
5 sources for · 0 against
Peer-reviewed literature and technology reporting confirm that researchers and engineers at MIT developed a technique and software to recover intelligible speech and sound from videos by measuring microscopic object vibrations.
The visual microphone recovers sound from object vibrations in a video by detecting the small vibrations of the objects caused by sound and converting the vibrations back into sound. Such vibrations are often very small and unnoticeable to humans. This difficulty is compounded by the fact that the target objects for sound recovery may not occupy a sizeable area of the video frame to allow for more pixels to be used in sound recovery. Changing the video resolution by up scaling and down scaling the video changes the number of pixels available for the visual microphone algorithm to work with and these effects on the quality of sound recovery are yet to be explored. This paper investigates the gap in the literature by up scaling videos with neural networks, down scaling videos, and a combination of both, on the recovered sound quality. It was found that changing the video resolution by any means gives insignificant improvements to the recovered sound quality at best and degrades the recovered sound quality in most of the studied cases. As a result, it is suggested not to apply any change in video resolution prior to sound recovery using the visual microphone.
By using a high-speed camera, researchers at MIT in 2014 where able to recover human speech from videos of minute vibrations of objects in a room. For example, in one experiment a 2,200fps camera was positioned outside a room behind sound-proof glass, videoing an empty crisp packet on the floor inside the room, while a researcher shouted “Mary had a little lamb” at the crisp packet. By detecting minute oscillations of the crisp packet of 1 μm (0.001 mm), and using hours of computer processing, a ten second audio clip could be produced that was recognisably “Mary had a little lamb” in an American accent.
The purpose of this study group was to investigate whether this tech- nique could be used in practice, with emphasis on the recovery of intel- ligible speech from a video feed of a room. During the week, the group investigated several aspects of the problem, including:
• how much an object vibrates due to sound;
• what can be done to maximize the vibration;
• how the MIT technique detects minute vibrations in videos; • what affects the quality of the resulting recording; and
• how good a recording is needed for intelligible speech.
It was discovered the MIT experiments would not have recovered intel- ligible speech from an ordinary conversation; their success depended on loud sounds and prior knowledge of “Mary had a little lamb”. Camera vibrations were also ignored by MIT; these are expected to be signifi- cant, but the technique could be adapted to be resilient to them. Other possibilities for enhancing their technique, by exploiting resonances or reflections, are discussed in the report. A high-speed low-noise cam- era is essential, and any existing video footage (such as from CCTV) is unlikely to be of sufficient quality. Further experiments with high-end high-speed cameras are needed to assess the feasibility of the technique in practice.
Since sound causes minute vibrations in objects surrounding the sound source, Visual Vibrometer allows us to recover the sound from remote point by optically measuring the vibration on the object's surface. For that system, Laser Doppler Velocimeters (LDVs) and high-speed cameras are commonly used, but there are issues in terms of equipment cost and data efficiency. Therefore, a technique has recently emerged to use event-based cameras in place of such equipment. Event-based cameras record only changes in brightness, independently, at each pixel, and its advantages such as high temporal resolution, high data efficiency, and simple device structure make it easier to measure. However, the technique is highly dependent on the object's surface characteristics, so the measurement conditions can prove difficult to realize. In this paper, we propose a new measurement system that observes changes in laser speckles caused by vibration using an event-based camera to achieve more robust measurement conditions, and a method for recovering sound from the event signals. The proposed sound recovery algorithm recovers the audio signal by noise reduction, sign assignment, and integration, assuming that the number of events produced by speckle pattern shift and detected at each time is closely related to the absolute value of the vibration speed. This method is very simple yet effective, and its performance is demonstrated by experiments in real environments.
The concept of visual microphones has become a compelling tool in non-contact sensing in recent years. The method involves recovering sound from silent videos by analyzing subtle object vibrations and it relies on multi-scale, orientation motion decomposition. Conventional approaches predominantly utilize the Complex Steerable Pyramid for phase-based motion extraction, which, despite its accuracy, is computationally overcomplete and resource-intensive. In this work, we propose a visual microphone framework that replaces the complex steerable pyramid with the more compact and efficient Riesz pyramid. Our method provides a phase-based motion extraction algorithm with the use of quaternioic representation of Riesz pyramid. Preliminary experiments demonstrate that the proposed approach achieves comparable quality in sound recovery with reduced computational overhead. This work opens up the potential for lightweight and scalable visual microphone implementations.
Caught on tape: cameras turn video into sound | New Scientist
Video: Invisible vibrations reveal hidden soundtrack
Now even the cameras have ears. Engineers at MIT have discovered a way to listen in on conversations simply by filming objects near someone talking loudly and measuring the tiny vibrations that sound causes in those objects. The technique uses high frame rate cameras to film an object. By tracking the position of an object down to 1/1000th of a pixel over time, the researchers were able to recover the vibration pattern in the material, and the sound that caused it.
Spies do already have laser microphones that allow them to listen in on far away conversations. But those are generally limited to listening to conversations behind clear panes of glass. The new technique can record sound from any object the camera can see, although some materials deliver better sound signatures than others.
## Seeing is hearing
“We were able to recover intelligible speech from maybe 15 feet away, from a bag of chips behind soundproof glass,” Abe Davis, who led the research. A recitation of “Mary Had A Little Lamb” is clearly audible in the visual recording (see video, above).
The same
Everything we examined (5)
This check searched the claim as stated. It did not run a separate search for evidence against it.