AI Proxy · usage
2026-08-05 12:06 · everything the proxy has recorded
These models read 2.1B tokens and wrote 9.7M — 219 read for every one written.
That is the shape of agentic traffic: long context, short answers. It is why one tokens-per-second figure across every request reads low, and why decode rate is broken out by prompt depth further down.
Decode rate
39.8 tok/s
median, measured from the first token on
Reply length
226 tok
mean across every request
Throughput
96.6 /h
requests per hour over the period
Errors
2.5%
1,090 of 42,987
Ranked by request volume. Latency here is a mean over whole requests, so a few very long ones drag it upward — read it as a rough scale and use decode rate by depth for how fast the model actually generates.
| Model | Requests | Prompt | Completion | Mean latency | Errors | |
|---|---|---|---|---|---|---|
qwen3-coder-next | 18,118 | 1.7B | 4.4M | 10,499 ms | 835 | |
qwen/qwen3-coder-next | 5,114 | 376.6M | 776.4k | 70,721 ms | 43 | |
DeepSeek-V4-Flash-0731-UD-IQ2_XXS | 2,425 | 568.6k | 450.4k | 13,221 ms | 0 | |
(none) | 2,296 | 0 | 0 | 38 ms | 16 | |
ornith-nvfp4 | 1,276 | 28.6M | 371.8k | 9,245 ms | 60 | |
codellama:70b | 928 | 109.0k | 262.8k | 63,607 ms | 4 | |
qwen3.6:27b | 845 | 6.0M | 222.7k | 49,391 ms | 4 | |
ornith-1.0-35b | 725 | 6.9M | 228.8k | 371,528 ms | 10 | |
gemma3:27b | 689 | 68.1k | 162.3k | 23,271 ms | 0 | |
gemma4:26b | 689 | 70.2k | 162.9k | 5,636 ms | 0 | |
gemma4:latest | 689 | 67.5k | 293.0k | 9,589 ms | 0 | |
gpt-oss:120b | 689 | 101.4k | 355.2k | 19,212 ms | 0 |
Which applications the traffic came from, by their own fingerprint.
| Application | Requests | Conversations | Tokens | Mean latency | |
|---|---|---|---|---|---|
| hermes | 21,863 | 224 | 2.1B | 23,032 ms | |
| ai-proxy-bench | 17,644 | 107 | 17.1M | 32,019 ms | |
| vscode-copilot | 1,507 | 0 | 0 | 46 ms | |
| requests | 1,072 | 0 | 0 | 33 ms | |
| httpx | 357 | 0 | 0 | 59 ms | |
| openai-sdk | 277 | 210 | 4.5M | 257,693 ms | |
| hermes-safety | 227 | 15 | 106.5k | 26,221 ms | |
| curl | 16 | 8 | 953 | 517 ms | |
| browser-chrome | 10 | 0 | 0 | 46 ms | |
| ai-proxy-healthcheck | 4 | 2 | 60 | 318 ms | |
| ai-proxy-chat | 4 | 1 | 2.3k | 3,390 ms | |
| alias-test | 2 | 2 | 630 | 23,048 ms |
Decode slows as the KV cache grows, so a single median across every prompt size describes no real request. Compare a quoted tok/s figure against the bucket that matches its prompt size.
| Prompt size | Samples | p50 tok/s | p90 | max |
|---|---|---|---|---|
| < 4K | 17,043 | 39.6 | 75.3 | 532.8 |
| 4–16K | 248 | 30.9 | 59.3 | 71.3 |
| 16–64K | 2,286 | 36.2 | 59.3 | 534.3 |
| > 64K | 13,253 | 40.3 | 48.8 | 373.0 |
How much context a request carries as a conversation goes on, and what it costs. Latency rising with depth means every turn re-reads the conversation; latency staying flat, or falling, means the prefix cache is holding the shared history and only the new turn is being read.
| Depth | Requests | Prompt tokens | TTFT | Total |
|---|---|---|---|---|
| 1–4 turns | 7,328 | 161 | 2,080 ms | 12,835 ms |
Upstream time split between reading the prompt and writing the reply. With prompts this much larger than replies, prefill should dominate — that it does not is the prompt cache doing its job, so a jump in the prefill share is how a caching regression would first show up.
Prefill
16%
reading the prompt
Decode
84%
writing the reply
Upstream time
26.1 h
across 7,328 requests
| Upstream | Requests | Mean latency | |
|---|---|---|---|
| vllm | 19,394 | 10,417 ms | |
| ollama | 15,331 | 18,323 ms | |
| lmstudio | 5,837 | 108,108 ms | |
| llamacpp | 2,425 | 13,221 ms |
| Status | Requests |
|---|---|
| 200 | 41,904 |
| 404 | 588 |
| no response (aborted or still in flight) | 213 |
| 400 | 176 |
| 499 | 100 |
| 500 | 6 |
Tools the models actually invoked — useful for spotting definitions that are sent every turn and never used.
| Tool | Calls | |
|---|---|---|
terminal | 386 | |
run | 328 | |
read_file | 96 | |
skill_manage | 68 | |
skill_view | 49 | |
patch | 30 | |
skills_list | 21 | |
memory | 13 | |
computer_use | 12 | |
write_file | 11 | |
web_search | 5 | |
execute_code | 4 |
What this traffic would have cost bought elsewhere, against the electricity spent making it here. Nothing on this page was billed — the models run on your own hardware.
Tokens
2.2B
98% of them a cache read, 9.7M written
Cost to produce
$0.37
electricity, over 143 GPU hours and the idle time around them
Hosted open-weights, 30B class
$83
what the same tokens would have cost
Claude Sonnet 4.5
$920
what the same tokens would have cost
| Tokens | Produced here | Bought elsewhere | |||||
|---|---|---|---|---|---|---|---|
| Date | Input | Cached | Output | GPU hrs | Electricity | Claude Sonnet 4.5 | Hosted open-weights, 30B class |
| Jul 17 | 627.3k | 18.9M | 58.7k | 1.2 | $0.01 | $8.45 | $0.79 |
| Jul 18 | 4.9M | 143.6M | 250.5k | 18.2 | $0.03 | $61.44 | $5.92 |
| Jul 19 | 5.9M | 219.0M | 477.6k | 21.5 | $0.04 | $90.46 | $8.62 |
| Jul 20 | 2.2M | 47.7M | 530.5k | 12.5 | $0.03 | $28.79 | $2.40 |
| Jul 21 | 2.2M | 101.8M | 338.4k | 2.5 | $0.01 | $42.15 | $3.91 |
| Jul 22 | 5.0M | 177.5M | 879.9k | 5.2 | $0.02 | $81.40 | $7.35 |
| Jul 23 | 741.2k | 20.9M | 103.6k | 0.7 | $0.01 | $10.04 | $0.91 |
| Jul 24 | 2.8M | 120.6M | 498.4k | 2.7 | $0.01 | $52.17 | $4.77 |
| Jul 25 | 4.2M | 251.4M | 517.9k | 4.1 | $0.02 | $95.63 | $9.10 |
| Jul 26 | 7.8M | 468.9M | 907.6k | 8.3 | $0.02 | $177.67 | $16.95 |
| Jul 27 | 6.0M | 408.4M | 774.3k | 6.4 | $0.02 | $152.03 | $14.51 |
| Jul 28 | 502.6k | 14.2M | 90.3k | 0.6 | $0.01 | $7.12 | $0.63 |
| Jul 29 | 1.3M | 143.2M | 169.4k | 1.7 | $0.01 | $49.45 | $4.79 |
| Jul 31 | 0 | 0 | 0 | 0.0 | $0.01 | $0.00 | $0.00 |
| Aug 01 | 193.8k | 217.2k | 137.6k | 3.7 | $0.02 | $2.71 | $0.15 |
| Aug 02 | 21.4k | 287.1k | 628.5k | 12.0 | $0.02 | $9.58 | $0.39 |
| Aug 03 | 48.0k | 611.7k | 883.5k | 15.5 | $0.03 | $13.58 | $0.56 |
| Aug 04 | 49.0k | 1.2M | 1.7M | 22.0 | $0.04 | $25.60 | $1.05 |
| Aug 05 | 298 | 223.5k | 789.8k | 3.9 | $0.01 | $11.92 | $0.48 |
| Total | 44.3M | 2.1B | 9.7M | 142.6 | $0.37 | $920.18 | $83.28 |
watts and watts_idle from that.pricing in the rules config.