BENCHMARK DATA
A reproducible benchmark for Chinese AI model APIs
Use published measurements, explicit limitations, and a fixed evaluation protocol instead of unsupported model rankings.
Updated 2026-09-11
The first TokenAAS benchmark dataset publishes ten verified deepseek-v4-flash streaming requests: five direct and five through an authenticated regional relay. Both routes achieved 100% success with no timeouts; the dataset is a routing baseline, not yet a cross-model quality ranking.
Published benchmark results
Verified measurements are published with their date, route, sample count, and latency percentiles. Download the source data for independent review.
| Model | Route | Samples | First token P50 | First token P95 | Total P95 | Success |
|---|---|---|---|---|---|---|
| deepseek-v4-flash | Direct | 5 | 1240.62 ms | 1448.30 ms | 1492.09 ms | 100% |
| deepseek-v4-flash | Regional relay | 5 | 1240.14 ms | 1328.71 ms | 1395.36 ms | 100% |
Results are observations for the disclosed date, routes, and workload. They are not a continuing performance or availability guarantee.
Published baseline
The August 31, 2026 run used the same model and streaming scenario for both routes. Direct first-token latency was 1240.62 ms P50 and 1448.30 ms P95. Regional relay latency was 1240.14 ms P50 and 1328.71 ms P95.
- Model: deepseek-v4-flash
- Samples: 5 direct and 5 authenticated regional-relay requests
- Direct total latency: 1268.71 ms P50 and 1492.09 ms P95
- Relay total latency: 1300.00 ms P50 and 1395.36 ms P95
- Success rate: 100% and timeout rate: 0% on both routes
What the result does and does not show
This baseline verifies that both routing paths completed the selected workload and records their observed latency. It does not prove that one route is universally faster, and it does not compare answer quality between model families.
- Small initial sample: five requests per route
- One model and one streaming workload
- No universal latency or availability guarantee
- Repeat before making a production routing decision
Cross-model evaluation protocol
Each candidate should receive the same representative prompt set in the same time windows. Results should be reported by task category rather than collapsed into one unsupported overall score.
- Quality: fixed rubric and expected-output checks
- Performance: first-token latency, total latency, throughput, and error rate
- Economics: input tokens, output tokens, configured price, and cost per accepted result
- Operations: live capacity, retries, fallback behavior, and regional route
- Repetition: peak and off-peak windows with a disclosed sample size
Planned model matrix
Only measured results are published as results. The following IDs are benchmark candidates and remain pending until their runs pass data validation.
- DeepSeek: deepseek-v4-flash and deepseek-v4-pro
- Qwen: qwen3.7-flash
- GLM: glm-5.2 and glm-5.3
- Kimi: kimi-k2.6
Machine-readable dataset
The current dataset, methodology fields, limitations, and pending matrix are available as JSON at https://tokenaas.ai/data/chinese-ai-model-api-benchmark.json. The version date changes when verified results are added.
- Format: JSON
- License: CC BY 4.0
- Publisher: TokenAAS
- Current version: 2026-09-11
Frequently asked questions
Does this benchmark prove that the regional relay is always faster?
No. It reports one small production sample. Region, time, prompt, model version, network conditions, and upstream capacity can change the result.
Does TokenAAS publish model quality scores yet?
Not yet on this page. Cross-model quality results will only be published after the same prompt set and scoring rubric have been applied to every candidate.
Can I download the benchmark data?
Yes. The machine-readable JSON dataset is available at https://tokenaas.ai/data/chinese-ai-model-api-benchmark.json.
Which model families are planned for the next benchmark runs?
The current matrix includes selected DeepSeek, Qwen, GLM, and Kimi public model IDs. Live catalog availability determines which routes can be tested.