AI Proxy · usage

Usage report

Building the report — reading every request the proxy has recorded. A few seconds.

2026-08-05 12:06 · everything the proxy has recorded

Period 18d 12h
First seen 2026-07-17
Requests 42,987
Tokens 2.1B
GPU NVIDIA GB10
Proxy 0.2.0

These models read 2.1B tokens and wrote 9.7M219 read for every one written.

That is the shape of agentic traffic: long context, short answers. It is why one tokens-per-second figure across every request reads low, and why decode rate is broken out by prompt depth further down.

Decode rate

39.8 tok/s

median, measured from the first token on

Reply length

226 tok

mean across every request

Throughput

96.6 /h

requests per hour over the period

Errors

2.5%

1,090 of 42,987

Models

Ranked by request volume. Latency here is a mean over whole requests, so a few very long ones drag it upward — read it as a rough scale and use decode rate by depth for how fast the model actually generates.

ModelRequestsPromptCompletionMean latencyErrors
qwen3-coder-next18,1181.7B4.4M10,499 ms835
qwen/qwen3-coder-next5,114376.6M776.4k70,721 ms43
DeepSeek-V4-Flash-0731-UD-IQ2_XXS2,425568.6k450.4k13,221 ms0
(none)2,2960038 ms16
ornith-nvfp41,27628.6M371.8k9,245 ms60
codellama:70b928109.0k262.8k63,607 ms4
qwen3.6:27b8456.0M222.7k49,391 ms4
ornith-1.0-35b7256.9M228.8k371,528 ms10
gemma3:27b68968.1k162.3k23,271 ms0
gemma4:26b68970.2k162.9k5,636 ms0
gemma4:latest68967.5k293.0k9,589 ms0
gpt-oss:120b689101.4k355.2k19,212 ms0

Clients

Which applications the traffic came from, by their own fingerprint.

ApplicationRequestsConversationsTokensMean latency
hermes21,8632242.1B23,032 ms
ai-proxy-bench17,64410717.1M32,019 ms
vscode-copilot1,5070046 ms
requests1,0720033 ms
httpx3570059 ms
openai-sdk2772104.5M257,693 ms
hermes-safety22715106.5k26,221 ms
curl168953517 ms
browser-chrome100046 ms
ai-proxy-healthcheck4260318 ms
ai-proxy-chat412.3k3,390 ms
alias-test2263023,048 ms

Decode rate by prompt depth

Decode slows as the KV cache grows, so a single median across every prompt size describes no real request. Compare a quoted tok/s figure against the bucket that matches its prompt size.

Prompt sizeSamplesp50 tok/sp90max
< 4K17,04339.675.3532.8
4–16K24830.959.371.3
16–64K2,28636.259.3534.3
> 64K13,25340.348.8373.0

Conversation depth

How much context a request carries as a conversation goes on, and what it costs. Latency rising with depth means every turn re-reads the conversation; latency staying flat, or falling, means the prefix cache is holding the shared history and only the new turn is being read.

DepthRequestsPrompt tokensTTFTTotal
1–4 turns7,3281612,080 ms12,835 ms

Where the time goes

Upstream time split between reading the prompt and writing the reply. With prompts this much larger than replies, prefill should dominate — that it does not is the prompt cache doing its job, so a jump in the prefill share is how a caching regression would first show up.

Prefill

16%

reading the prompt

Decode

84%

writing the reply

Upstream time

26.1 h

across 7,328 requests

Upstreams

UpstreamRequestsMean latency
vllm19,39410,417 ms
ollama15,33118,323 ms
lmstudio5,837108,108 ms
llamacpp2,42513,221 ms

Response status

StatusRequests
20041,904
404588
no response (aborted or still in flight)213
400176
499100
5006

Tool calls

Tools the models actually invoked — useful for spotting definitions that are sent every turn and never used.

ToolCalls
terminal386
run328
read_file96
skill_manage68
skill_view49
patch30
skills_list21
memory13
computer_use12
write_file11
web_search5
execute_code4

Day by day

What this traffic would have cost bought elsewhere, against the electricity spent making it here. Nothing on this page was billed — the models run on your own hardware.

Tokens

2.2B

98% of them a cache read, 9.7M written

Cost to produce

$0.37

electricity, over 143 GPU hours and the idle time around them

Hosted open-weights, 30B class

$83

what the same tokens would have cost

Claude Sonnet 4.5

$920

what the same tokens would have cost

TokensProduced hereBought elsewhere
DateInputCachedOutputGPU hrsElectricityClaude Sonnet 4.5Hosted open-weights, 30B class
Jul 17627.3k18.9M58.7k1.2$0.01$8.45$0.79
Jul 184.9M143.6M250.5k18.2$0.03$61.44$5.92
Jul 195.9M219.0M477.6k21.5$0.04$90.46$8.62
Jul 202.2M47.7M530.5k12.5$0.03$28.79$2.40
Jul 212.2M101.8M338.4k2.5$0.01$42.15$3.91
Jul 225.0M177.5M879.9k5.2$0.02$81.40$7.35
Jul 23741.2k20.9M103.6k0.7$0.01$10.04$0.91
Jul 242.8M120.6M498.4k2.7$0.01$52.17$4.77
Jul 254.2M251.4M517.9k4.1$0.02$95.63$9.10
Jul 267.8M468.9M907.6k8.3$0.02$177.67$16.95
Jul 276.0M408.4M774.3k6.4$0.02$152.03$14.51
Jul 28502.6k14.2M90.3k0.6$0.01$7.12$0.63
Jul 291.3M143.2M169.4k1.7$0.01$49.45$4.79
Jul 310000.0$0.01$0.00$0.00
Aug 01193.8k217.2k137.6k3.7$0.02$2.71$0.15
Aug 0221.4k287.1k628.5k12.0$0.02$9.58$0.39
Aug 0348.0k611.7k883.5k15.5$0.03$13.58$0.56
Aug 0449.0k1.2M1.7M22.0$0.04$25.60$1.05
Aug 05298223.5k789.8k3.9$0.01$11.92$0.48
Total44.3M2.1B9.7M142.6$0.37$920.18$83.28