Alibaba Announced Aug 3, 2026

Qwen3.8-2.4T-A95B

The open-weight release of Qwen3.8-Max: a 2.4-trillion-parameter sparse MoE with 95B active parameters, positioned around autonomous coding, workplace tasks, and long-horizon execution.

Architecture
2.4T MoE · 95B active
512 experts with hybrid Gated DeltaNet and gated attention
Context
262K native
Extensible toward 1M; the hosted API documents a 1M window
Weights
Released
Qwen3.8-Max License: source-available with revenue-gated terms
Reasoning
Always on
The open checkpoint requires thinking mode for all interactions

Publisher pricing

Pricing below comes from the release announcement and does not represent a MicroRouter route or price.

Publisher-listed API price phases in US dollars per million tokens.
PhaseInput / 1MOutput / 1MBlended / 1MEffectiveSource
Model Studio API$1.65$4.95$2.48Aug 3, 2026open-endedAlibaba Cloud

Publisher-reported benchmarks

Results are grouped by unit. Percentage, Elo, and task-count results never share a scale.

BenchmarkReported resultReported byAs ofSource
Terminal Bench 2.186.6Alibaba CloudAug 3, 2026Announcement ↗
SWE-bench Pro67.7Alibaba CloudAug 3, 2026Announcement ↗
DeepSWE 1.156.6Alibaba CloudAug 3, 2026Announcement ↗
GPQA Diamond92.6Alibaba CloudAug 3, 2026Announcement ↗
HLE43.6Alibaba CloudAug 3, 2026Announcement ↗
  • Alibaba's hosted qwen3.8-max API is documented as multimodal, while the open-weight A95B checkpoint is text-only with mandatory thinking.
  • Listed pricing is the Model Studio rate for most regions; Singapore bills $2 / $6 per 1M tokens.
  • Alibaba reports the benchmark figures; MicroRouter has not reproduced them.

Independent measurements

Artificial Analysis runs its own evaluations and timing against each model's first-party API. These are the only numbers on this page not reported by the publisher.

Qwen3.8 2.4T A95B ranks #10 of 448 model families on their Intelligence Index.

Intelligence Index
57.7
Composite of their evaluations
Coding Index
71.9
Coding evaluations only
Blended price
$3
USD per 1M tokens, 3:1 input to output, first-party list
Output speed
39
Median tokens per second
First answer token
53s
Median seconds, including reasoning

Qwen3.8 2.4T A95B on Artificial Analysis ↗

Source: Artificial Analysis, read Sep 2, 2026. Measured on the model's first-party API, not a MicroRouter route.

What is not independently confirmed

  • Every figure on this page is self-reported by the cited publisher. MicroRouter did not run these benchmarks.
  • Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
  • Publisher pricing can change, and introductory phases have stated end dates.
  • Qwen3.8-2.4T-A95B is routable through MicroRouter today; live pricing on its model page is authoritative over the publisher figures here.

Related coverage

Other release dossiers and the current MicroRouter catalog.

Sources

  1. Alibaba Cloud announcement

    Published Aug 3, 2026

    Read at Alibaba Cloud
  2. Artificial Analysis model page

    Published Sep 2, 2026

    Read at Artificial Analysis

Editorial reference revised 2026-09-02. Catalog status comes from the generated routing snapshot.