vLLM
An LLM inference engine — the place to understand KV caching and continuous batching.
Hey there, friend using a screen reader. I've carefully labeled every button and image on this site. Hope you have a smooth experience. If anything is inconvenient, my email is in the footer — just let me know and I'll fix it. — Store Owner
Loading…
Browse by topic
18 items
An LLM inference engine — the place to understand KV caching and continuous batching.
RankLLMRerank cannot be imported, unless vllm is installed. This is true even when using a backend that has nothing to do with vLLM. Reason: llama_index.postprocessor.rankllm_rerank.base imports PromptMode and Reranker from rank_llm.rerank at module level: In the current release 0.25.7, importing anything from rank_ll…
3.0.3 After upgrading Airflow from version 3.0.2 to 3.0.3 sensitive data in extra fields and extra fields JSON is visible in edit connection page. Sensitive data should be masked with Enter Edit Connection window in Airflow UI. Ubuntu 24.04 Other Docker-based deployment Airflow deployed with Docker Swarm. Settings in…
This issue is for keeping track of the recurrent Whisper asks as well as the linked on-going efforts to support that feature, if any. When a feature request has no linked PR, feel free to claim the work here if you want to help! - Related issues: https://github.com/vllm-project/vllm/issues/19556, https://github.com/vl…
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared…
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research .
Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.
GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
Kaggle’s AI Agents Intensive with Google brought learners together in a no-cost course to build and deploy the next frontier of AI.
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
How does DSpark work The decode phase in LLM inference is traditionally memory-bound. Most latency comes from streaming weights from DRAM into SRAM, not from intense computation. Speculative decoding addresses this by using a lightweight draft model to produce candidate tokens, then having the target model verify them…
By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retrieving context, and generating a response. Text-to-speech (TTS) is the final step — and the one users notice most. If speech generation is slow, the whole experience f…
You have an agent that can already record a demonstration and push it to the Hugging Face Hub . Now you want to run that loop continuously: collect episodes through the day, train a policy on the growing dataset, deploy it, and pull the next batch back to improve it. Run that loop once and every piece works. Run it ev…
gr.Workflow , built right into Gradio, makes the pipeline the interface . You describe your steps as a graph of typed nodes, and Gradio serves a drag-and-drop canvas where every node is runnable and every intermediate result is visible. The same graph is also a REST API and a one-command deploy to HF Mirror Spaces.
We're thrilled to share that Baseten is now a supported Inference Provider on the HF Mirror Hub!
Most of these backends are weight-only . This means that they store the weights in low precision and dequantize them back to high precision at compute time. This reduces memory usage significantly, but it usually does not make inference faster, and can even add a small latency overhead.
Of course, making AI research accessible requires a powerful search engine, so that humans and agents can quickly find relevant and related work, either through the website or the pwc search CLI command , which agents can use via the Skill .