Code
basic.py
Usage
1
Set up your virtual environment
2
Install dependencies
3
Start vLLM server
4
Set your API key
5
Run Agent
Save the code above as
basic.py, then run:Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Stream a vLLM agent’s response asynchronously with aprint_response().
import asyncio
from agno.agent import Agent
from agno.models.vllm import VLLM
agent = Agent(
model=VLLM(id="Qwen/Qwen2.5-7B-Instruct", top_k=20, enable_thinking=False),
markdown=True,
)
asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activate
uv venv --python 3.12
.venv\Scripts\activate
Install dependencies
uv pip install -U agno openai vllm
Start vLLM server
vllm serve Qwen/Qwen2.5-7B-Instruct \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--dtype float16 \
--max-model-len 8192 \
--gpu-memory-utilization 0.9
Set your API key
export VLLM_API_KEY=xxx
Run Agent
basic.py, then run:python basic.py
Was this page helpful?