Startrise Labs · Study
Benchmark addenda · Aug 2026

Grok 4.6, Qwen3.8-Max and Muse Spark 1.2 on our frontend benchmark

Three August releases ran through the Startrise LLM benchmark: twelve single-shot frontend briefs, gated in a browser, scored by a blind three-judge panel and a blind human. Grok 4.6 (72.1), Qwen3.8-Max (71.3) and Muse Spark 1.2 (70.3) land in a cluster, take three July task wins, and each breaks one engineering brief.

Three new entries land within 1.8 points of each other

72.1Grok 4.6 · ties GLM 5.2 · 9 of 12 cells human-rated
71.3Qwen3.8-Max · +7.7 over Qwen 3.7 Max · 9 of 12 rated
70.3Muse Spark 1.2 · 8 of 12 rated

Same twelve briefs, same contract, same three-judge panel and weights as the July flagship run, single-shot. Every gap between these three is inside run-to-run noise: read them as a cluster, not a ranking.

All three slot into the upper-middle of the 15-model table

Overall score, 0–100, N=1 per cell. Yellow: August additions. Source: STUDY.md leaderboard + Addenda A–C.

Three July task wins changed hands

Landing page

Qwen3.8-Max 87.3, previously Opus 5 at 85.8.

SVG icon system

Qwen3.8-Max 81.6, previously Fable 5 at 80.5.

HTML email

Grok 4.6 87.6, previously Opus 5 at 84.3.

BriefGrok 4.6Qwen 3.8Muse 1.2
01 Three.js scroll81.629.371.1
02 WebGL shader75.684.669.3
03 Landing page84.387.375.3
05 3D game26.247.767.6
06 Open creative77.075.470.7
07 Sell yourself71.077.866.5
08 Accessible UI51.7*81.3*69.3*
09 Brownfield79.6*76.4*74.9*
10 SVG icons76.881.667.0
11 Stateful app75.1*62.8*57.1*
12 Zero JS78.677.679.1*
13 HTML email87.673.476.0

Final score per brief. Yellow cell = task win on the pooled table. * = no blind human rating yet; human weight redistributed. Source: results/<run>/scores.json.

Each one still breaks on behaviour, single-shot

Lowest final cell per model. Grok 4.6's game shipped with a console error and no animation frames; Qwen's Three.js page barely rendered. All three posted record-tier landing-page verdicts and under-responded to the interaction probes. The August frontier is converging on taste and still stumbling on behaviour.

Muse is the cheap, fast one; Qwen the slow, heavy one

ModelOutput tokensWall clockSuite costRed flags
Muse Spark 1.2138,562647 s$0.6269
Grok 4.6262,9323,241 s$1.6335
Qwen3.8-Max477,26412,107 s$2.9240

All twelve briefs, at OpenRouter list rates. Source: STUDY.md Addenda A–C; tokens and red flags cross-checked against scores.json.

How to read these numbers

N=1

One run per cell, no error bars. Ties and sub-2-point gaps are noise.

Unrated cells

Accessible UI, brownfield and stateful app have no human rating on any of the three (Muse's zero-JS too).

One reviewer

The human column is one blind rater, same as every run on record.

Not yet audited

None of the three is in the judge-bias audit; they join at the next full re-run.

Contamination window

The briefs were public before these models shipped. No evidence they saw them; a holdout set is still required.

Does the ranking survive different weights? Re-weight gates, judges and human review, drop briefs, compare any two models and see cost vs score in the interactive Benchmark Explorer, built on the same per-cell data.

Results pending

Queued: 13 newer models, added to the roster but untested

GPT-6 Astra · GPT-6.1 Sol · GPT-6 Luna · Claude Fable 5.1 · Claude Opus 5.5 · Claude Sonnet 5.5 · Grok 4.7 · GLM 5.3 · Muse Spark 1.3 · Gemini 3.8 Flash · DeepSeek V4.1 Flash · DeepSeek V4 Pro 0813 · Qwen 3.8 Max 0902. No scores exist for any of them yet. Cursor Composer 2.5 also remains unrun: it has no API and needs a manual run inside Cursor.

Sources