THE #1 AV NEWS PUBLICATION. PERIOD.

Deepgram Launches Flux, Conversational Speech Recognition Model

Flux

Deepgram has introduced Flux, a conversational speech recognition (CSR) model designed to support real-time voice agents. Unlike traditional automatic speech recognition (ASR), which was developed for transcription use cases like captions or meeting notes, Flux is built to understand the flow of dialogue, including when a speaker has finished and when to respond.

Traditional STT systems weren’t designed to participate in live dialogue. To recreate conversational flow, developers have historically pieced together transcription, voice activity detection and turn-taking logic.
Flux eliminates this by transforming speech recognition from simply transcribing words to modeling the flow of dialogue. This provides developers with the tools to build responsive voice agents.

Key features include:

  • Embedded turn-taking with context-aware detection and barge-in handling.
  • Low-latency performance, with ~260ms end-of-turn detection.
  • Simplified development, providing structured conversational cues to replace client-side logic.
  • Enterprise scalability, offering Nova-3 level accuracy and support for 100+ concurrent streams per GPU.

Flux redefines what speech recognition can do for real-time AI,” said Scott Stephenson, CEO and co-founder of Deepgram. “For decades, ASR was built to listen and record. Flux is different — it listens, understands, and guides conversations with human-like timing.”

Flux is generally available starting today. To mark the launch, Deepgram announced OktoberFLUX, allowing developers to use Flux free during October with up to 50 concurrent connections.

More information is available at deepgram.com/flux

Top