Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs and support seven languages. Its company-published tests report an 11.1 ms time to first token for a 10-second clip on an Apple M4 Pro, but the results are not an independent evaluation.

Cactus Compute released Whistle, a speech-recognition model packaged as a 16.9 MB file and designed to transcribe audio locally on CPUs. The company says it supports English, German, French, Spanish, Italian, Dutch and Polish, and can return word-level timestamps or speech embeddings as well as transcripts.

Cactus Compute says Whistle processes 16 kHz mono audio clips of up to 30 seconds in one pass. It can detect the spoken language automatically or use a language specified by the user. The company also says audio stays on the device in its browser demonstration, where the initial use downloads the model file.

The release describes Whistle as running in the same C++ CPU engine as Needle, another Cactus Compute model. The company says both use the same container and quantisation, and that some model blocks use Needle’s code. Whistle has an encoder and a speech decoder; its audio-specific component is a gated cross-attention layer connecting them. Users can select decoder depth at load time, though the encoder always runs all eight of its blocks.

In benchmarks published with the announcement, Cactus Compute reports a 11.1 ms time to first token for a 10-second audio clip on an Apple M4 Pro CPU. It reports 1,319 decoded tokens per second for Whistle, compared with 266 for OpenAI’s Whisper base and 262 for Moonshine tiny v2 in the company’s test. The figures use each model’s official runtime at its defaults, according to Cactus Compute. The company also reports word-error-rate advantages for Whistle on several listed datasets, while Whisper base leads on TED-LIUM, AMI and the MLS average.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that runs on-device and shares a CPU engine with its Needle model.

A Smaller Model for Local Speech

A model that fits in 16.9 MB could make speech recognition more practical on devices with limited memory or unreliable connectivity, including phones, wearables, vehicles and embedded hardware. Local processing can also avoid sending audio to a remote service, although the release’s privacy statement applies to its own on-device setup and does not establish how every third-party integration will handle recordings.

The reported latency and size offer a point of comparison for developers weighing local recognition against larger models or cloud services. But the benchmark numbers are vendor-published results, not a third-party test. Performance can vary with hardware, runtime settings, audio quality and language, so the figures do not by themselves establish how Whistle will perform across consumer devices or real-world use.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Fits With Needle

Cactus Compute presents Whistle as a speech model built on infrastructure shared with Needle, rather than as a separate CPU engine. Its technical description says the encoder converts audio into frames, while the decoder generates text; cross-attention lets the decoder use the encoded clip. The model also supports keyword biasing, which the company describes as raising the likelihood of supplied phrases during decoding.

The release compares Whistle with Whisper base and Moonshine tiny v2 across file size, decoding speed and word error rate. Cactus Compute notes that some benchmark results are unavailable because model authors did not publish them. It also cautions that Whisper’s AMI figure uses a different subset from the AMI results for the other models. These caveats limit direct comparisons across every dataset.

““Speak, and Whistle transcribes it on your device.””

— Cactus Compute, in its October 2, 2026 release

Independent Tests Still Needed

The release does not provide an independent evaluation of the model’s speed or transcription accuracy. Its benchmark summary does not establish how Whistle performs on a broad range of processors, noisy recordings, accents or all seven supported languages. The available report also does not give a complete, directly comparable set of results for every model and dataset.

It is also not clear from the supplied release what licensing terms apply, where developers can obtain the model beyond the browser demonstration, or whether the stated 16.9 MB size includes all files needed for deployment. The company describes intended uses including microcontrollers and automotive systems, but the report does not provide device-specific measurements or deployment certifications.

Availability and Device Testing

The immediate next step for developers is to test Whistle on their own hardware and audio, checking transcription accuracy, memory use and response time under their intended settings. Cactus Compute has not, in the supplied announcement, set out a separate schedule for additional releases or independent benchmark work.

Readers should look for further documentation on model access, licensing and supported deployment environments. Independent comparisons using the same devices, audio clips and runtime conditions would help show whether the company’s reported size and speed advantages translate beyond its published tests.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says it runs locally on CPUs and can produce transcripts, word timestamps and speech embeddings.

Which languages does it support?

Cactus Compute lists English, German, French, Spanish, Italian, Dutch and Polish. It says the model can detect the language or use one specified by the user.

How large is the model?

The company gives the model file size as 16.9 MB. The announcement does not clarify whether that figure includes every file needed for all deployment setups.

Are the speed and accuracy results independently verified?

No independent verification is provided in the announcement. The speed and word-error-rate comparisons are reported by Cactus Compute, and the company notes missing results and differences among some benchmark datasets.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I’m A USB-C Maximalist

A tech enthusiast proclaims themselves a ‘USB-C Maximalist,’ emphasizing exclusive support for USB-C devices and accessories. The stance sparks debate.

Induction Cooking Basics That Make Switching Easier

mastering induction cooking basics makes switching seamless, but understanding the key factors unlocks its full potential—keep reading to discover how.

Essential Questions Before Your Big Move

Inquire about crucial factors before relocating to ensure a seamless transition, but are you ready to discover what you might be overlooking?

Farmhouse Apartments – Modern Living With Rustic Touches

Uncover the secrets to transforming your space into a cozy farmhouse apartment that beautifully blends modern living with rustic charm.