On September 16, 2026, startup Deveillance released Kalypta — the first consumer-grade anti-transcription app. A local on-device model adds adversarial noise to the outgoing audio stream in real time, trained against Whisper and NVIDIA Canary architecture transcribers, causing AI notetakers like Granola, Wispr Flow, and Cluely to transcribe conversations less accurately. According to the company, the recognition error rate in live calls via Google Meet, Zoom, and Teams AI rises from 4.4% to 50.5–59.7%, while humans continue to understand the speaker, although the noise itself is audible to them. This is currently a beta version for macOS on Apple Silicon only, and all key figures were published by the vendor itself.
What happened
The Deveillance team opened beta access to Kalypta on September 16, 2026 — a utility that works as a local privacy filter for outgoing audio. The model runs entirely on the device: it analyzes the user's speech and mixes adversarial noise into the stream, trained against Whisper and NVIDIA Canary architecture transcription models, which underpin popular notetakers Granola, Wispr Flow, and Cluely. In the company's benchmark on 100 clips from the Mozilla Common Voice corpus, the word error rate (WER) in Google Meet, Zoom, and Teams AI increased from 4.4% on clean audio to 50.5–59.7% in live mode; on pre-recorded clips the effect is stronger — around 64–70% WER. In a pilot study where 7 participants read 14 sentences through the M2 model, the machine's recognition error increased by 51 percentage points, and the human error by 24.5 percentage points. The authors themselves acknowledge that in the current beta the noise is audible to interlocutors, and the model is updated almost weekly.
Context
The backstory is simple: AI notetakers have transformed from a novelty into a work habit over the past year, and a call participant increasingly finds out that their words have ended up in someone else's database only after the conversation — Granola, Wispr Flow, and Cluely transcribe meetings in the background, usually without the consent of all parties. Attacks on speech recognition through adversarial distortions have long been described in security research, but were considered a laboratory exotic: such perturbations are fragile and do not transfer well to real time. Kalypta is the first attempt to package this technique into a consumer product, i.e., a claim to a new category of 'local privacy filter for outgoing audio.' The technology is inherently dual-use: noise that protects against unauthorized recording also interferes with honest meeting documentation that everyone agreed to. Notably, the vendor published the evaluation methodology — WER, ESTOI, and PESQ metrics on 100 Common Voice clips under the CC0 license; this is a verifiable set, not a bare claim, although the measurements themselves currently come from an interested party.
Why this matters for the industry
For the industry, this is the start of the 'adversarial noise vs. ASR' race: transcription providers, from teams behind Whisper-like models to NVIDIA Canary, will have to regularly retrain against such distortions, otherwise their accuracy on protected calls will drop by about half. Meeting platforms — Google Meet, Zoom, and Teams — will have to formalize policy: determine whose side privacy is on and how to treat software that intentionally distorts the audio path. For startups, the launch is early confirmation of demand for voice privacy: the first two to three quarters are a window when the category can be claimed by brand and beta list before copies appear, but there is currently nothing to embed the solution into, because this is a signaling beta, not a building block. A separate value for ASR teams is the published methodology: the combination of WER with ESTOI and PESQ sets a template for honest evaluation of dual-use products, where machine error is measured together with speech quality for humans. The main lesson is that Kalypta solves one problem at the expense of another: protection from machines is achieved at the expense of intelligibility and comfort for live interlocutors, and it is this trade-off that will determine the fate of the category.
Why this matters for users
If you suspect that colleagues are silently recording you through Granola, Wispr Flow, or Cluely, Kalypta is the first chance to become 'inaudible' to AI while remaining understandable to humans. The real path now is one: sign up for the beta list and test the effect on your own scenarios, since the product is distributed through early access. Keep three practical limitations in mind: the app only works on macOS with Apple Silicon chips; the noise will be heard by your interlocutors, so the question 'what is that background in your mic' is inevitable; your own speech becomes less intelligible, which is critical for important negotiations. The model is updated almost weekly, so the effect may change noticeably from week to week — it is worth recording your own measurements rather than relying on feelings.
What is still unknown / limitations
All key figures are from the vendor itself: the benchmark of 100 clips and the pilot with 7 participants and 14 sentences are statistically a demonstration, not proof. The attack is trained against specific Whisper and NVIDIA Canary architectures — from available data, it does not follow that the effect transfers to other ASR models, and providers are able to retrain against such distortions. The real-time effect is weaker than on pre-recorded clips: about 50–60% WER vs. 64–70%. The assessment 'the model does not recognize 2 out of 3 words' should be taken with caution: at 50–60% WER, approximately half of the words are correctly recognized. Production metrics are not disclosed — latency, device load, price, and release timelines beyond macOS; the legal status of intentional audio distortion in corporate calls is also undetermined.
Sources
Author
Look at AI, editorial team
