LATENCY BENCHMARK

Direct and regional-relay AI API latency measured from production infrastructure

A small controlled test shows why regional routing should be measured with first-token, total latency, success rate, and repeated time windows.

Updated 2026-09-11

In brief

On August 31, 2026, TokenAAS ran five streaming requests per route using deepseek-v4-flash. Both direct and authenticated regional-relay routes succeeded 100%. Relay first-token P95 was 1328.71 ms versus 1448.30 ms direct, an 8.26% lower value in this sample.

Test method

The same model and streaming scenario were tested from production infrastructure. Five requests used direct upstream access and five used an authenticated regional relay. The relay was protected by client authentication and source-network controls.

  • Model: deepseek-v4-flash
  • Scenario: streaming chat
  • Samples: 5 direct and 5 relay
  • Test date: 2026-08-31

Measured results

Direct first-token latency measured 1240.62 ms P50 and 1448.30 ms P95. Relay measured 1240.14 ms P50 and 1328.71 ms P95. Total latency P95 was 1492.09 ms direct and 1395.36 ms relay.

  • Success rate: 100% on both routes
  • Timeout rate: 0% on both routes
  • Relay first-token P95 difference: -8.26%
  • Relay total-latency P95 difference: -6.48%

How to interpret the result

This result does not prove that a relay is always faster. The sample is intentionally small and represents one model, one route configuration, and one time window.

  • Repeat during local morning, afternoon, evening, and peak periods
  • Measure application region and end-user region separately
  • Track errors and throttling with latency
  • Select routes using sustained observations rather than one run

Production routing guidance

A regional relay is useful when it improves reachability, policy enforcement, or operational control. Keep direct and relay routes observable and use a controlled switch so each account can be tested or rolled back independently.

Frequently asked questions

Does the regional relay always reduce latency?

No. Network conditions vary. This test observed lower P95 values for the relay, but route decisions require repeated measurements across time windows and workloads.

Why publish a small sample?

It documents the method and a reproducible baseline. It is not presented as a universal performance guarantee.