Union Alpha Benchmark Results

Frontier-class coding scores at a fraction of frontier pricing — with one real tradeoff that the leaderboards do not show you.

Benchmark results

Every figure below is attributed to its source. None of it is vendor-confirmed, because no vendor has claimed Union Alpha as of September 17, 2026.

Benchmark Score Context Note
SWE-bench Verified 74.2% Real GitHub issue resolution Matches GPT and Opus-class models that cost roughly 50× more per task
DeepSWE 74.0% Software engineering resolution rate Audited figure; the number that first drew attention to the model
Terminal-Bench v4.0 ~52% Terminal agent execution, ~$1.60 per task Sits on the Pareto line between GLM-5.3-Flash (~33%, ~$0.30) and GPT-6 Astra high effort (~55%, ~$4)
SMF Official A 136/157 (86.6%) Community run of 157 tests by Mike Gannotti's team Zero errors, $0 cost. Reasoning and tool use rated perfect; code writing was the weak spot
Reliability sweep 8/8 at 100% benchable.ai success-rate measurement Returned consistent, usable responses across every benchmark attempted
Speed ranking 6th percentile Latency versus other catalogued models The honest caveat — noticeably slower than most alternatives

Cost per task

The benchmark scores only tell half the story. Union Alpha's real differentiator is how little each resolved task costs relative to models that score about the same.

Model Per task Note
Union Alpha ~$0.65 Free during preview; anticipated $0.50/$1.50 per 1M
GPT-class frontier $6.50 – $11.80 Comparable SWE-bench resolution rate
Opus-class frontier $6.50 – $11.80 Comparable SWE-bench resolution rate
GLM-5.3-Flash ~$0.30 Cheaper, but ~33% on Terminal-Bench v4 vs Union Alpha's ~52%
Speed caveat: Union Alpha ranks in the 6th percentile for speed against other catalogued models. The benchmark wins are real, but so is the wait — this is a throughput tradeoff, not a free lunch.

Benchmark questions

Are these benchmark numbers verified?
The DeepSWE figure of 74.0% is described as audited, and SWE-bench Verified at 74.2% is widely reported. Terminal-Bench v4 positioning comes from an Artificial Analysis chart that plots score against cost without printing exact values. The SMF result is a community run and should be read as indicative.
What is DeepSWE?
A software engineering benchmark that measures how reliably a model resolves real coding tasks. its 74.0% resolution rate is what first put the model on people's radar, because it matched models costing far more.
Why does the speed ranking matter?
Because benchmarks measure quality, not throughput. A 6th-percentile speed ranking means that for interactive use — chat, inline completion, anything a human waits on — the model will feel sluggish even though its output quality is high.
How does it compare on cost-performance?
On the Terminal-Bench v4 cost-versus-score plot it sits on the Pareto frontier, between GLM-5.3-Flash (cheaper but weaker) and GPT-6 Astra at high effort (stronger but several times the cost).

Keep reading