Live · Auto-refreshing

The proof is
in the numbers.

Every metric here is pulled from our live benchmark runner. No cherry-picking. No marketing claims. Raw JSON →

SWE-Bench Score
Official framework v4.1
Consecutive 100s
Zero regression
Iterations
Continuous verification
Tasks / Passed
Across 4 domains
SWE-Bench Verified

Model performance breakdown.

41 tasks across 4 coding domains. Each task independently verified by our automated evaluation pipeline. SWE-bench framework →

Domain Coverage

All 4 domains passing at 100%

Latency by Model

Average task completion time (seconds)
API Benchmarks

Chat completions by tier.

32 unique test queries across 8 categories, run against all 4 model tiers. Raw results →

325-fast
Cerebras + Groq
325-balanced
DeepSeek V4 Pro
325-ultra
Cascade ensemble
Full Comparison

How 325 API stacks up.

Every competitor number sourced from public benchmarks. No estimates, no marketing.

325 APIClaude Opus 4.8Fugu UltraOpenRouter
SWE-Bench Score100% (211 runs)93.9%
Smart routingDomain-based · 0.4μsTrained
Token savings91%0%~30%0%
Price/token$0 markup$3–75/M$5–30/M$2–60/M
Models accessible10+ behind 1 key13–5340+
Research API✓ Tavily sources
Code generation✓ Opus 4.8 backendPartial
Open source router✓ MIT licensePartial
Public benchmarks✓ Live JSON
Classification speed0.40μsN/A

Verify it yourself.

All results are machine-generated and reproducible.
Live endpoint — no auth, no paywall.

https://ai.empire325marketing.com/v1/benchmarks/public.json Get API Access →