Dedicated model · Available as managed deployment
Alibaba's Qwen3 family — dense models from 0.6B to 32B and mixture-of-experts up to 235B — with a switchable thinking mode, 119 languages and Apache-2.0 weights. Validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.
Why AxForge
| One family, every size | 0.6B to 32B dense plus 30B-A3B and 235B-A22B MoE: pick the size for the job and keep the same behaviour and prompt style across them. |
|---|---|
| Thinking on demand | Qwen3 reasons step by step when you ask and answers directly when you do not — one model for both modes, switched per request. |
| Apache-2.0 | Permissive weights with no usage conditions: deploy, fine-tune and ship commercially without a licence conversation. |
Specifications
| Model | Qwen3 — Qwen |
|---|---|
| Modalities | Text |
| Sizes | 752M, 1.7B, 2.0B, 4.0B, 8.2B, 14.8B |
| Context window | 40,960 tokens |
| Licence | Open weights — apache-2.0; commercial use permitted |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the Qwen3 page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys Qwen3 on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with the model name you receive. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
8B and 14B for fast general use on a dedicated DGX Spark, 32B or 30B-A3B when quality matters; the 235B MoE is a multi-GPU deployment.
Qwen3.8 27B is the newer generation AxForge serves per token. Qwen3 is the broader family with more sizes — the choice when you want a specific size on your own machine.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
Hardware by the hour, week, month or year; the managed service is quoted per deployment — both confirmed in writing before anything is billed.