Skip to main content
6G AI Traffic Characterization Testbed

Metrics

At a glance
  • Four categories: A, traffic characteristics; B, response time and errors; C, service metrics for agents and computer use; D, task-related QoE.
  • Two vantage points: application-layer records logged per request, and packet captures parsed per packet and per flow.
  • What the study groups need: TR 26.870 maps each group of traffic characteristics it lists for other working groups (volume, temporal structure, latency, reliability, adaptivity) to these metrics.
  • Open point: whether and how TTFT, TTLT and tool-call latency can be used for AI traffic evaluation is for further study.

The categories​

CategoryCovers
ATraffic characteristics: byte volumes, UL/DL ratio, token throughput, streaming rate, burstiness
BResponse time and HTTP errors: end-to-end latency, TTFT, TTLT, stall events, success rate, error taxonomy
CService metrics: agent tool calling, loop factor, tool latency, computer-use actions
DTask-related QoE: video understanding accuracy. Task completion, answer quality, satisfaction and safety are for further study.

Open a category for its metrics.

A. Traffic characteristics

These describe the volume and temporal shape of AI service traffic.

  • Volume and tokens: mean uplink, downlink and total bytes; the UL/DL ratio, computed as the ratio of the means; mean input and output tokens, from the provider's usage report or estimated with a tokenizer; and the token output rate, computed per record and then averaged.
  • Streaming throughput: for chunked HTTP streaming, the byte rate between the first and the last token, and the mean number of content chunks.
  • Burstiness, application layer: the peak-to-mean ratio and the coefficient of variation of the bytes per request; and, with a configurable gap threshold (1 s by default) separating ON from OFF periods, the mean and count of ON and OFF gaps between requests.
  • Burstiness, packet layer: the instantaneous bit rate over short windows (1 ms to 100 ms recommended), with its mean, peak, peak-to-mean, coefficient of variation, active fraction and 95th and 99th percentiles; the packet arrival rate per window; and the inter-packet gap, with its mean, coefficient of variation and 99th percentile.
  • Communication frequency: application-layer interactions (requests, responses, coordination or signalling messages) per second, and the interval between them.
  • Small packets: the share of packets below a size threshold, 128 bytes by default.
B. Response time and errors
  • Response latency: the wall-clock time from sending the request to receiving the last byte, reported as mean, median, 95th and 99th percentile, minimum, maximum and standard deviation.
  • Time to first token (TTFT): from sending the request to receiving the first content token, or first byte where the token is not in the first bytes. It reflects the model's initial inference latency plus the network round trip, and is the primary indicator of perceived responsiveness for streaming.
  • Time to last token (TTLT): from sending the request to the final content token. For non-streaming responses it equals the full latency.
  • Streaming stalls: a stall is a gap between chunks above a threshold, 1 s by default. Reported as stall rate, stall count and stall duration.
  • Success and errors: success rate, then failures classified as timeouts, rate limiting (HTTP 429), server errors (5xx), tool failures and other errors.
C. Service metrics
  • Agents (MCP tool calling): tool calls per session; the loop factor, the mean number of model turns per session, above 1 for multi-step reasoning; and tool latency, the execution time inside the tool without model inference.
  • Computer use: actions and steps per session, where one step is one reasoning cycle of the model; action latency, the full screenshot, action and result cycle; screenshot bytes; and the action error rate.
D. Task-related QoE

An AI media service delivers an inference result, often text, rather than a reconstructed media signal. There is no reference signal to compare against, so reference-based media quality metrics, including those of TR 26.944, do not apply. A task-related metric is used instead, specific to each scenario because it depends on that scenario's ground truth.

For video understanding, the closed-form accuracy is the share of questions the model answers correctly in a test session. An open-ended accuracy, scored by a second model acting as judge on completeness and reliability, is for further study.

What each traffic characteristic is measured with

TR 26.870 lists groups of traffic characteristics as relevant to other working groups, and maps each to the analysis in annex C and to these metrics.

GroupExamples from the TRAnalysisMetrics
A. Volume and directionalityUplink and downlink volume, UL/DL asymmetry, session and flow durationClause C.4Category A
B. Temporal structureBurstiness, inter-arrival times, periodic or event-driven, streaming or request/responseClause C.5Category C (clause D.4.5)
C. Latency sensitivityOne-way and round-trip sensitivity, time to first media element, completion latencyClauses C.5 and C.6Category B
D. Reliability and continuitySensitivity to loss, tolerance to reordering or jitter, stalls, retransmissionsClause C.7Categories B and D
E. Adaptivity and dynamicsRate adaptation, dependency on inference timing, coupling across flows or modalities, changes within a sessionClauses C.7 and C.8Categories A and D

The TR lists as follow-up a careful review and detailed definition of these metrics, and the definition of how the characteristics are measured. Extending the mapping to media services other than AI is for further study.

Sources for this page
  • At a glance and the categories: TR 26.870 V0.6.1, clause D.4.1, table D.4.1-1; annex B.1 and its Editor's Note.
  • A. Traffic characteristics: TR 26.870 V0.6.1, clauses D.4.3.1 to D.4.3.6, tables D.4.3.2-1, D.4.3.3-1, D.4.3.4-1 and D.4.3.4-2.
  • B. Response time and errors: TR 26.870 V0.6.1, clauses D.4.4.1 to D.4.4.5 and their tables.
  • C. Service metrics: TR 26.870 V0.6.1, clauses D.4.5.2 and D.4.5.3, tables D.4.5.2-1 and D.4.5.3-1.
  • D. Task-related QoE: TR 26.870 V0.6.1, clauses D.4.6.0, D.4.6.1.0 to D.4.6.1.2.
  • What each characteristic is measured with: TR 26.870 V0.6.1, clause 6.3.3.3, table 6.3.3.3-1 and its Editor's Note.