Grok 4.6, Qwen3.8-Max and Muse Spark 1.2 on our frontend benchmark
Three August releases ran through the Startrise LLM benchmark: twelve single-shot frontend briefs, gated in a browser, scored by a blind three-judge panel and a blind human. Grok 4.6 (72.1), Qwen3.8-Max (71.3) and Muse Spark 1.2 (70.3) land in a cluster, take three July task wins, and each breaks one engineering brief.
Three new entries land within 1.8 points of each other
Same twelve briefs, same contract, same three-judge panel and weights as the July flagship run, single-shot. Every gap between these three is inside run-to-run noise: read them as a cluster, not a ranking.
All three slot into the upper-middle of the 15-model table
Overall score, 0–100, N=1 per cell. Yellow: August additions. Source: STUDY.md leaderboard + Addenda A–C.
Three July task wins changed hands
Landing page
Qwen3.8-Max 87.3, previously Opus 5 at 85.8.
SVG icon system
Qwen3.8-Max 81.6, previously Fable 5 at 80.5.
HTML email
Grok 4.6 87.6, previously Opus 5 at 84.3.
| Brief | Grok 4.6 | Qwen 3.8 | Muse 1.2 |
|---|---|---|---|
| 01 Three.js scroll | 81.6 | 29.3 | 71.1 |
| 02 WebGL shader | 75.6 | 84.6 | 69.3 |
| 03 Landing page | 84.3 | 87.3 | 75.3 |
| 05 3D game | 26.2 | 47.7 | 67.6 |
| 06 Open creative | 77.0 | 75.4 | 70.7 |
| 07 Sell yourself | 71.0 | 77.8 | 66.5 |
| 08 Accessible UI | 51.7* | 81.3* | 69.3* |
| 09 Brownfield | 79.6* | 76.4* | 74.9* |
| 10 SVG icons | 76.8 | 81.6 | 67.0 |
| 11 Stateful app | 75.1* | 62.8* | 57.1* |
| 12 Zero JS | 78.6 | 77.6 | 79.1* |
| 13 HTML email | 87.6 | 73.4 | 76.0 |
Final score per brief. Yellow cell = task win on the pooled table. * = no blind human rating yet; human weight redistributed. Source: results/<run>/scores.json.
Each one still breaks on behaviour, single-shot
Lowest final cell per model. Grok 4.6's game shipped with a console error and no animation frames; Qwen's Three.js page barely rendered. All three posted record-tier landing-page verdicts and under-responded to the interaction probes. The August frontier is converging on taste and still stumbling on behaviour.
Muse is the cheap, fast one; Qwen the slow, heavy one
| Model | Output tokens | Wall clock | Suite cost | Red flags |
|---|---|---|---|---|
| Muse Spark 1.2 | 138,562 | 647 s | $0.62 | 69 |
| Grok 4.6 | 262,932 | 3,241 s | $1.63 | 35 |
| Qwen3.8-Max | 477,264 | 12,107 s | $2.92 | 40 |
All twelve briefs, at OpenRouter list rates. Source: STUDY.md Addenda A–C; tokens and red flags cross-checked against scores.json.
How to read these numbers
N=1
One run per cell, no error bars. Ties and sub-2-point gaps are noise.
Unrated cells
Accessible UI, brownfield and stateful app have no human rating on any of the three (Muse's zero-JS too).
One reviewer
The human column is one blind rater, same as every run on record.
Not yet audited
None of the three is in the judge-bias audit; they join at the next full re-run.
Contamination window
The briefs were public before these models shipped. No evidence they saw them; a holdout set is still required.
Does the ranking survive different weights? Re-weight gates, judges and human review, drop briefs, compare any two models and see cost vs score in the interactive Benchmark Explorer, built on the same per-cell data.
Queued: 13 newer models, added to the roster but untested
GPT-6 Astra · GPT-6.1 Sol · GPT-6 Luna · Claude Fable 5.1 · Claude Opus 5.5 · Claude Sonnet 5.5 · Grok 4.7 · GLM 5.3 · Muse Spark 1.3 · Gemini 3.8 Flash · DeepSeek V4.1 Flash · DeepSeek V4 Pro 0813 · Qwen 3.8 Max 0902. No scores exist for any of them yet. Cursor Composer 2.5 also remains unrun: it has no API and needs a manual run inside Cursor.
Sources
- sr-llm-benchmark STUDY.md, Addenda A (run 2026-08-09-2106), B (2026-08-12-0114), C (2026-08-12-2359). Working copy as of 2026-10-05, including uncommitted addenda edits.
- results/2026-08-09-2106, 2026-08-12-0114, 2026-08-12-2359: scores.json + human.json (per-cell finals, tokens, red flags re-verified)
- config/models.json (roster; entries marked "ADDED 2026-10-05, untested")
- config/scoring.json (weights 0.25 / 0.45 / 0.30)