Union Alpha Benchmark Results
Frontier-class coding scores at a fraction of frontier pricing — with one real tradeoff that the leaderboards do not show you.
Benchmark results
Every figure below is attributed to its source. None of it is vendor-confirmed, because no vendor has claimed Union Alpha as of September 17, 2026.
| Benchmark | Score | Context | Note |
|---|---|---|---|
| SWE-bench Verified | 74.2% | Real GitHub issue resolution | Matches GPT and Opus-class models that cost roughly 50× more per task |
| DeepSWE | 74.0% | Software engineering resolution rate | Audited figure; the number that first drew attention to the model |
| Terminal-Bench v4.0 | ~52% | Terminal agent execution, ~$1.60 per task | Sits on the Pareto line between GLM-5.3-Flash (~33%, ~$0.30) and GPT-6 Astra high effort (~55%, ~$4) |
| SMF Official A | 136/157 (86.6%) | Community run of 157 tests by Mike Gannotti's team | Zero errors, $0 cost. Reasoning and tool use rated perfect; code writing was the weak spot |
| Reliability sweep | 8/8 at 100% | benchable.ai success-rate measurement | Returned consistent, usable responses across every benchmark attempted |
| Speed ranking | 6th percentile | Latency versus other catalogued models | The honest caveat — noticeably slower than most alternatives |
Cost per task
The benchmark scores only tell half the story. Union Alpha's real differentiator is how little each resolved task costs relative to models that score about the same.
| Model | Per task | Note |
|---|---|---|
| Union Alpha | ~$0.65 | Free during preview; anticipated $0.50/$1.50 per 1M |
| GPT-class frontier | $6.50 – $11.80 | Comparable SWE-bench resolution rate |
| Opus-class frontier | $6.50 – $11.80 | Comparable SWE-bench resolution rate |
| GLM-5.3-Flash | ~$0.30 | Cheaper, but ~33% on Terminal-Bench v4 vs Union Alpha's ~52% |
Benchmark questions
Are these benchmark numbers verified?
What is DeepSWE?
Why does the speed ranking matter?
How does it compare on cost-performance?
Keep reading
Model Comparison
See how these same benchmark numbers compare directly against GPT, Claude Opus and GLM.
Pricing Breakdown
What the preview costs today and what the anticipated per-token rate would mean for your workload.
Setup Guide
Set up OpenRouter or OpenCode and run these benchmarks yourself in about five minutes.