Qwen
Frontier language model for chat, reasoning and tool use.
- Modality
- Text · Vision
- Capacity
- Long context window
- Output
- Text · JSON · tools
- Latency
- Streaming, low latency
- Category
- LLM
- Provider
- Alibaba
- Model id
- qwen
- Billing
- Pay-as-you-go
- 99.9% SLA
- Official discount
- Pay-as-you-go
- Fast delivery
Try it
Type a prompt and preview a sample response. Real runs use your API key — pay only on success.
Est. cost: $0.0002 / run
curl https://api.modelplex.ai/v1/chat/completions \
-H "Authorization: Bearer $MODELPLEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen","messages":[{"role":"user","content":"Explain how a CDN works, in two sentences."}],"temperature":0.7,"max_tokens":512}'Overview
Qwen is Alibaba's language model, available through Modelplex. Frontier language model for chat, reasoning and tool use. Call it with the OpenAI-compatible API — one key, one balance, transparent pricing, and you only pay for what succeeds.
- 50K+
- active developers
- 99.9%
- uptime
- 2×
- faster routing
- 70%
- avg savings
Pricing details
- Unit price
- $0.30/1M
- Original price
- $0.38/1M
- Discount
- −21%
- Billing
- Pay-as-you-go
per 1M tokens
Prices include the official discount and you're charged only on success — failed generations are free.
Why it stands out
Deep reasoning
Multi-step problem solving, planning and reliable tool use.
Long context
Feed large documents and codebases without chunking gymnastics.
Drop-in compatible
OpenAI-compatible — point the SDK at APIMart and go.
Streaming
Token-by-token responses for snappy, interactive UX.
Transparent pricing
Per-token cost up front, billed only on success.
One key, one bill
Call it through APIMart alongside every other model.
What you can build
Up and running in 3 steps
- 1
Create an account
Sign up and grab free credits — no card required.
- 2
Point the SDK
Set the base URL and your APIMart key.
- 3
Make a call
Use the model id and pay only on success.
FAQ
Point the OpenAI SDK at Modelplex, set your key, and use the model id "qwen". The base URL and key are the only change to your code.
$0.30/1M, billed per call with our standard discount applied up front — and you're only charged on success.
Yes — new accounts start with free credits, so you can try Qwen before adding a balance.
Served on a 99.9% uptime target with automatic same-kind fallback if an upstream has a hiccup.
Yes. It's drop-in OpenAI-compatible — no code changes beyond the base URL and key.
Yes — set stream: true and tokens arrive server-sent as they're generated, exactly like the OpenAI API.
Qwen handles a large context for long documents and codebases — see the docs for the exact token limit.
Yes, where the model supports it: pass tools in the request and handle the tool calls returned in the response.
Per-key rate limits you set from the console, with sensible defaults on every new key. Need more headroom? Just ask support.
Calls are served from regional infrastructure on a 99.9% uptime target, and your inputs and outputs are never used to train models.
Related models
GPT
Frontier language model for chat, reasoning and tool use.
$1.25/1M$1.56/1M20% offClaude
Frontier language model for chat, reasoning and tool use.
$5/1M$6.25/1M20% offGemini
Frontier language model for chat, reasoning and tool use.
$1.25/1M$1.56/1M20% offDeepSeek
Frontier language model for chat, reasoning and tool use.
$0.28/1M$0.35/1M20% off