Qwen3.8 系列旗舰模型, 面向智能体场景, 推理/编程/长周期任务能力强, 支持视觉理解. 1M context.
Model ID: qwen3.8-max · Type: chat · Provider: Alibaba
Endpoints: /v1/chat/completions · /v1/messages
| Input (per 1M tokens) | $1.4 USD |
| Output (per 1M tokens) | $4.2 USD |
| Cache read (per 1M tokens) | $0.119 USD |
| Cache write 5m (per 1M tokens) | $1.75 USD |
from openai import OpenAI
client = OpenAI(api_key="sk-...", base_url="https://api.router.ai/v1")
resp = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)Qwen3.8-Max (`qwen3.8-max`) is billed per usage at $1.4/1M in · $4.2/1M out, in USD. Current pricing is always listed at https://370.ai/models/qwen3.8-max.
Send a request to https://api.router.ai/v1/v1/chat/completions with the header `Authorization: Bearer <your API key>` and `"model": "qwen3.8-max"`. The API is OpenAI-compatible, so any OpenAI SDK works by changing base_url to https://api.router.ai/v1 — no other code change.
Qwen3.8-Max can be called on: /v1/chat/completions; /v1/messages.
Qwen3.8-Max accepts up to 991,808 input tokens and can return up to 131,072 output tokens. Requests exceeding the input limit are rejected before reaching the model.
Qwen3.8-Max supports: vision, function_calling, prompt_caching, reasoning, thinking, streaming.
Qwen3.8-Max is a chat model from Alibaba, available through the 370.AI gateway with the same API key as every other model.
Call it through the 370.AI OpenAI-compatible endpoint (API base: https://api.router.ai/v1). AI agents can discover and call every model on this gateway through MCP (https://mcp.router.ai/mcp) with no manual integration.