QwenAPI.net API Access Pricing Benchmarks vs Kimi K3 About

Qwen3.8 Max Benchmarks: Every Real Score That Exists

Spoiler: almost none do. This page tracks verified results only — vendor claims are labeled as claims, and every number carries its date and conditions.

As of 2026-07-21, there is no official benchmark table for qwen3.8-max-preview. No model card, no scores, no license. Alibaba's "second only to Fable 5" line is currently an unverified claim.

The only independent, verifiable result so far: Trilogy AI's single blind StackPerf run — Qwen3.8 Max 80 vs Kimi K3 83 (one repository task, one run: a data point, not a verdict).

The verified-results table

ResultSource & conditionsTypeDate
StackPerf blind head-to-head: Qwen3.8-max-preview 80 vs Kimi K3 83. Kimi finished faster with fewer tokens; Qwen produced cleaner system boundariesTrilogy AI — one matched 269-file repository task, single runIndependent2026-07 (verified 2026-07-21)
"One of the most powerful models available today… second only to Fable 5"Alibaba announcement — no supporting table publishedVendor claim2026-07-19
0 of 321 tracked benchmarks have published scores for this modelBenchLM tracker (placeholder page)Absence-of-data, itself informative2026-07-21

Why the vacuum exists

Our own evaluation (in progress)

We are running a real-workload coding-agent evaluation of qwen3.8-max-preview against GPT-5.6 Sol, Claude Fable 5, and Kimi K3 — six repository-level tasks (multi-file bug fix, cross-module feature, constrained refactor, screenshot-to-code, tool-failure recovery, 1M-context codebase navigation), two independent runs each, hidden acceptance tests, blind review, and per-success cost accounting (time, tokens, credits, human interventions). It also includes a Qwen3.7-Max upgrade comparison and weekly anchor re-runs to detect preview drift.

Results publish here with full conditions and dates as tracks complete. What this page will never do: merge tracks into a single leaderboard score, or report a number without its test conditions.

How to read benchmark claims this month

  1. If a score has no date, discard it — preview drift makes undated numbers meaningless.
  2. If a "Qwen3.8 benchmark" cites MMLU-style knowledge scores, check the model ID: several circulating results actually belong to the unrelated old Qwen3-8B small model.
  3. Single-run agent results (including Trilogy's, and early X demos) are directional at best; agent variance across two runs is large.
Choosing a production model while the data settles? EvoLink lets you switch between 50+ models behind one API, so a benchmark surprise doesn't mean a rewrite.

Last verified 2026-07-21. This page updates when official scores, the model card, or our own evaluation results land.

Sources: Trilogy AI StackPerf run · BenchLM tracker · Announcement coverage