Hany Farid (UC Berkeley) says ten seconds of your voice is now enough to beat "voice as password"

Reddit r/artificial News

Summary

UC Berkeley researcher Hany Farid states that a ten-second voice clip is enough to compromise voice-based password systems, raising concerns about authentication security and fraud.

TL;DR: A UC Berkeley digital-forensics researcher just told an interviewer the exact number that quietly retires "voice as password." Ten seconds, per Hany Farid — that's the entire audio budget a bank's voice-biometric layer implicitly assumes an attacker can't clear. No breach, no malware, no privileged access: a short clip and consumer-grade cloning software beat an authentication layer marketed as secure a year ago. The hiring pipeline is failing on the identical premise. Fortune's reporting this month puts North Korean-linked fraud flags at 44% of U.S. remote-IT applications, up from 11% twelve months prior — the interview process assumed a face and a voice on a call meant something, and that assumption is now what's being priced against. Neither failure required new capability from the attacker. Both required the defender to keep trusting a signal that stopped being scarce. https://preview.redd.it/1eq5u9691bph1.jpg?width=1024&format=pjpg&auto=webp&s=d97a336f2dd5885afd8d7c093cd67577a1e52e4c Oo… Scams. All such techniques have one thing in common: quickly gain your trust (or stoke your fear), in order to swindle you for profit. Things were much more simpler decades ago. On another occasion, this call was a Malaysian Chinese guy – thick with Cantonese dialect. He was trying to pose as my long-lost friend, starting straight with a playful question, "哈罗,你点啊?认得我唧声冇?" (Translation: Hello! How are you? Recognize my voice?)… … As if I would be gullible enough to take the bait, and dig through my memory for any of my childhood pals that might sound like this imposter, and give him the name. Then he'll impersonate my pal from then on, and I'd be goner, for sure. So I just replied, "你打错电话。" (You've got the wrong number.) Even now, we still occasionally receiving SMS from scammers, trying to convince us that someone has wrongfully swiped our credit card of bank XYZ (which we don't own). And we should call the police at this number 03-ABCDEFG. Dismissed. Now I just treat them as noise - nominal entertainment for light conversation with my wife, at best. ___________________________ Every one of these calls tests the same blind spot: you'd catch it in a second if it happened in someone else's story. In your own, it just felt like a weird call you handled fine that one time. Hany Farid's whole episode is really about how much of our own trust is still running on that one-time save, untested since. Genuinely curious where you'd draw your own line here: is a family or team safe-word overkill, or already overdue? Drop your take below. 👇 Clip credit: Info-Tech Research Group (Digital Disruption w/ Geoff Nielson) — full episode on their channel. DM for credit or removal requests.
Original Article

Similar Articles

My voice-agent test now includes the 600-second cliff

Reddit r/AI_Agents

The author describes a voice agent call cut off at 600 seconds without warning, and proposes a testing approach to handle max duration gracefully, including pre-cutoff warnings and state preservation.

Voice AI Systems Are Vulnerable to Hidden Audio Attacks

Hacker News Top

New research shows that imperceptible audio signals can hijack large audio-language models (LALMs) with 79-96% success, forcing them to execute unauthorized commands like web searches or sending emails. The technique, dubbed AudioHijack, targets generative models and works regardless of user input, posing a serious security risk to voice AI systems.