Let the router pick the optimal model for every query. No configuration. 91% token savings. 0.40μs classification speed on CPU.
The classifier analyzes your prompt against domain keywords (html→fast, python→balanced, analysis→ultra) plus structural features (short prompts penalize ultra, long prompts boost it). 100% accuracy across 32 classification tests.
| Your Query Contains | Routes To | Why |
|---|---|---|
| HTML, CSS, shell, simple | 325-fast | Sub-second Cerebras |
| Python, code, math, logic, SQL | 325-balanced | DeepSeek V4 best for code |
| analyze, research, compare, deep | 325-ultra | Cascade ensemble for depth |