Transformers
Pretrained-model toolkit spanning text, vision, audio, and multimodal tasks.
Hey there, friend using a screen reader. I've carefully labeled every button and image on this site. Hope you have a smooth experience. If anything is inconvenient, my email is in the footer — just let me know and I'll fix it. — Store Owner
Loading…
Browse by topic
14 items
Pretrained-model toolkit spanning text, vision, audio, and multimodal tasks.
This issue is for keeping track of the recurrent Whisper asks as well as the linked on-going efforts to support that feature, if any. When a feature request has no linked PR, feel free to claim the work here if you want to help! - Related issues: https://github.com/vllm-project/vllm/issues/19556, https://github.com/vl…
Using GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1.
The official source provides no summary. Open the item to read the original.
Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.
Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.
Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ). To this end, several arena-based leaderboards have established themselves as useful reference points for the community:
Arabic WER: 20.92% · Parameters: 1.6B · Emirati WER (TII evaluation): 22.73%
Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:
One reason is that traditional benchmarks overlook many of the conditions and qualities that make voice systems reliable, natural, contextually appropriate, and effective in practice. That's why we recently introduced held-out sets in Real World VoiceEQ , the Open-ASR Leaderboard , and the Far-field ASR Leaderboard :…
Voice is rapidly becoming AI's primary interface. From customer support and healthcare to education, entertainment, and personal assistants, speech is increasingly replacing text as the way people interact with AI.
By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retrieving context, and generating a response. Text-to-speech (TTS) is the final step — and the one users notice most. If speech generation is slow, the whole experience f…