AI Proxy · benchmark

Benchmark — 38 configurations

2026-08-02 11:52 · proxy 0.2.0

GPU NVIDIA GB10 · 122 GB
Configurations 38
Requests 1,296
Graded suite coding-v1

gemma4:26b · ollama · cached leads: 100% fully correct at 63.5 tok/s out.

9 other configurations scored the same and were slower. Ranked by correctness first, then output rate. 38 configurations measured across 3 backends.

Run this — 70/30 weighted

gemma4:26b

100% correct · 63.5 tok/s · 16.8 GB · answers in 2.8s · score 76

Runner-up

qwen3-coder-next

100% correct · 62.3 tok/s · 48.2 GB · answers in 2.6s · score 76

Don't be fooled by

qwen3:0.6b

307 tok/s and only 22% correct

Not measured

2 cells

listed in the table as failures, not omitted

The trade-off

Every configuration, placed by correctness and output rate. A point below and to the left of another is beaten on both counts at once, so the dashed frontier is the shortlist — everything off it is dominated by something on it. Hover any point for its name.

0%20%40%60%80%100%1030100300output tokens/sec (log)better ↗TASKS FULLY CORRECT ↑gemma4:26b · ollama · cached — 100% correct at 63.5 tok/sgemma4:26b · 16.8 GB · ollama · cold — 100% correct at 63.5 tok/sqwen3-coder-next · NVFP4 · vllm · cold — 100% correct at 62.3 tok/sqwen3-coder-next · NVFP4 · vllm · cached — 100% correct at 62.1 tok/sqwen3-coder:tuned · 48.2 GB · ollama · cold — 100% correct at 60.6 tok/sqwen3-coder:tuned · ollama · cached — 100% correct at 60.5 tok/sqwen3-coder-next · ollama · cached — 100% correct at 51.6 tok/sqwen3-coder-next · 48.2 GB · ollama · cold — 100% correct at 51.6 tok/sllama4 · ollama · cached — 100% correct at 18.1 tok/sllama4 · 62.8 GB · ollama · cold — 100% correct at 18.1 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768 — 94% correct at 17.8 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144 — 94% correct at 17.8 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072 — 94% correct at 17.7 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768 — 94% correct at 17.7 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144 — 94% correct at 17.7 tok/sDeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072 — 94% correct at 17.6 tok/sqwen3.6:27b · ollama · cached — 94% correct at 12.1 tok/sqwen3.6:27b · 16.2 GB · ollama · cold — 94% correct at 12.1 tok/sgemma3:27b · 16.2 GB · ollama · cold — 94% correct at 11.5 tok/sgemma3:27b · ollama · cached — 94% correct at 11.5 tok/sllama3:70b-instruct · 37.2 GB · ollama · cold — 86% correct at 5.6 tok/sqwen3-coder:30b · ollama · cached — 83% correct at 83.9 tok/sqwen3-coder:30b · cachedqwen3-coder:30b · 17.3 GB · ollama · cold — 83% correct at 83.4 tok/sgemma4 · 8.9 GB · ollama · cold — 83% correct at 55.7 tok/sgemma4 · ollama · cached — 83% correct at 55.7 tok/sllama3:70b-instruct · ollama · cached — 83% correct at 5.6 tok/sgpt-oss:120b · ollama · cached — 81% correct at 38.5 tok/sgpt-oss:120b · 60.9 GB · ollama · cold — 81% correct at 37.8 tok/sminicpm-v4.5 · 5.7 GB · ollama · cold — 67% correct at 39.9 tok/sminicpm-v4.5 · ollama · cached — 64% correct at 39.9 tok/sqwen3:0.6b · ollama · cached — 22% correct at 307.1 tok/sqwen3:0.6b · 0.5 GB · ollama · cold — 17% correct at 305.5 tok/scodellama:70b · ollama · cached — 17% correct at 5.7 tok/scodellama:70b · 36.2 GB · ollama · cold — 17% correct at 5.7 tok/sqwen3:4b · ollama · cached — 11% correct at 72.7 tok/sqwen3:4b · 2.3 GB · ollama · cold — 11% correct at 72.1 tok/sgemma4:26b100% correct at 64 tok/son 16.8 GB of memoryqwen3:0.6b5× the winner's speed —and wrong 78% of the timellama4 — 63 GBbeaten on both axes by a model27% of its size

Held constant

Think

off

Temp

0.0

Parallel

1

Results

TTFT is the first token of any kind; TTFC the first content token — the gap between them is time the model spent reasoning. Decode rate is measured from the first token onward, so reasoning tokens count as generated work. Best value in each column is highlighted.

ConfigurationTTFT p50Decode p50TokensTotal p50Fully correctCasesvs bestOK
gemma4:26b · ollama · cached475 ms63.51662,780 ms100%100%4.3x36/36
gemma4:26b · 16.8 GB · ollama · cold487 ms63.51662,778 ms100%100%4.3x36/36
qwen3-coder-next · NVFP4 · vllm · cold150 ms62.31632,582 ms100%100%4.0x36/36
qwen3-coder-next · NVFP4 · vllm · cached147 ms62.11652,468 ms100%100%3.8x36/36
qwen3-coder:tuned · 48.2 GB · ollama · cold318 ms60.61572,989 ms100%100%4.6x36/36
qwen3-coder:tuned · ollama · cached252 ms60.51572,928 ms100%100%4.5x36/36
qwen3-coder-next · ollama · cached286 ms51.61573,435 ms100%100%5.3x36/36
qwen3-coder-next · 48.2 GB · ollama · cold351 ms51.61573,490 ms100%100%5.3x36/36
llama4 · ollama · cached819 ms18.11236,837 ms100%100%10.5x36/36
llama4 · 62.8 GB · ollama · cold831 ms18.11237,214 ms100%100%11.1x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768140 ms17.81668,112 ms94%93%12.4x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144143 ms17.81668,103 ms94%93%12.4x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072141 ms17.71668,118 ms94%93%12.4x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768336 ms17.71668,390 ms94%93%12.9x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144294 ms17.71668,407 ms94%93%12.9x36/36
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072298 ms17.61668,412 ms94%93%12.9x36/36
qwen3.6:27b · ollama · cached473 ms12.116910,656 ms94%93%16.3x36/36
qwen3.6:27b · 16.2 GB · ollama · cold545 ms12.116910,750 ms94%93%16.5x36/36
gemma3:27b · 16.2 GB · ollama · cold638 ms11.524919,465 ms94%94%29.8x36/36
gemma3:27b · ollama · cached611 ms11.524819,139 ms94%94%29.3x36/36
llama3:70b-instruct · 37.2 GB · ollama · cold474 ms5.610517,153 ms86%96%26.3x36/36
qwen3-coder:30b · ollama · cached211 ms83.91612,144 ms83%88%3.3x36/36
qwen3-coder:30b · 17.3 GB · ollama · cold264 ms83.41602,233 ms83%88%3.4x36/36
gemma4 · 8.9 GB · ollama · cold466 ms55.73306,498 ms83%83%10.0x36/36
gemma4 · ollama · cached447 ms55.73306,496 ms83%83%10.0x36/36
llama3:70b-instruct · ollama · cached432 ms5.610517,071 ms83%94%26.1x36/36
gpt-oss:120b · ollama · cached533 ms38.53498,825 ms81%81%13.5x36/36
gpt-oss:120b · 60.9 GB · ollama · cold658 ms37.83589,746 ms81%81%14.9x36/36
minicpm-v4.5 · 5.7 GB · ollama · cold243 ms39.91222,713 ms67%82%4.2x36/36
minicpm-v4.5 · ollama · cached238 ms39.91222,720 ms64%80%4.2x36/36
qwen3:0.6b · ollama · cached193 ms307.1146686 ms22%52%1.1x36/36
qwen3:0.6b · 0.5 GB · ollama · cold208 ms305.5143653 ms17%48%1.0x36/36
codellama:70b · ollama · cached257 ms5.724736,147 ms17%24%55.4x36/36
codellama:70b · 36.2 GB · ollama · cold320 ms5.724336,130 ms17%25%55.3x36/36
qwen3:4b · ollama · cached230 ms72.74877,276 ms11%14%11.1x36/36
qwen3:4b · 2.3 GB · ollama · cold261 ms72.14877,354 ms11%14%11.3x36/36
ornith-nvfp4 · NVFP4 · vllm · coldNone/None
ornith-nvfp4 · vllm · cachedNone/None

Weighted standings

Correctness alone is not a ranking — one point of correctness is not worth half the speed. Each model's best cell scores 70% × correctness + 30% × relative speed, where relative speed is decode rate against the fastest model in this report (qwen3:0.6b, 307.1 tok/s = 1.0). The bar is the weighting made visible: correctness · speed. Drag to change what you value; the ranking recomputes.

all correctness all speed70 / 30
#Model Correcttok/sScore
1gemma4:26b100%63.5
76
2qwen3-coder-next100%62.3
76
3qwen3-coder:tuned100%60.6
76
4llama4100%18.1
72
5DeepSeek-V4-Flash-0731-UD-IQ2_XXS94%17.8
68
6qwen3.6:27b94%12.1
67
7gemma3:27b94%11.5
67
8qwen3-coder:30b83%83.9
67
9gemma483%55.7
64
10llama3:70b-instruct86%5.6
61
11gpt-oss:120b81%38.5
60
12minicpm-v4.567%39.9
51
13qwen3:0.6b22%307.1
46
14qwen3:4b11%72.7
15
15codellama:70b17%5.7
12

Time to a finished answer

SECONDS UNTIL THE FULL ANSWER HAS ARRIVED (median task) →qwen3:0.6b0.7s · 22%qwen3-coder:30b2.1s · 83%qwen3-coder-next2.6s · 100%minicpm-v4.52.7s · 67%gemma4:26b2.8s · 100%qwen3-coder:tuned3.0s · 100%gemma46.5s · 83%llama46.8s · 100%qwen3:4b7.3s · 11%DeepSeek-V4-Flash-0731-U8.1s · 94%gpt-oss:120b8.8s · 81%qwen3.6:27b10.7s · 94%top 12 of 15 models — the full field is in the table

What memory buys

Position is footprint against correctness; bubble area is output speed. Models whose size cannot be read — vLLM checkpoints live inside their containers — are absent, not zero.

0%25%50%75%100%10204060GIGABYTES THE MODEL OCCUPIES → bubble area = output speedTASKS FULLY CORRECT ↑llama4 — 62.8 GB, 100%, 18.1 tok/sgpt-oss:120b — 60.9 GB, 81%, 38.5 tok/sqwen3-coder-next — 48.2 GB, 100%, 62.3 tok/sqwen3-coder:tuned — 48.2 GB, 100%, 60.6 tok/sllama3:70b-instruct — 37.2 GB, 86%, 5.6 tok/scodellama:70b — 36.2 GB, 17%, 5.7 tok/sqwen3-coder:30b — 17.3 GB, 83%, 83.9 tok/sgemma4:26b — 16.8 GB, 100%, 63.5 tok/sqwen3.6:27b — 16.2 GB, 94%, 12.1 tok/sgemma3:27b — 16.2 GB, 94%, 11.5 tok/sgemma4 — 8.9 GB, 83%, 55.7 tok/sminicpm-v4.5 — 5.7 GB, 67%, 39.9 tok/sqwen3:4b — 2.3 GB, 11%, 72.7 tok/sqwen3:0.6b — 0.5 GB, 22%, 307.1 tok/sgemma4:26b — the buyqwen3-coder-next & qwen3-coder:tunedgpt-oss:120bllama4

Same weights, two engines

The only controlled engine comparison a run can contain: identical weights reachable through more than one backend, paired by cache state. When output rates tie, the wait for the first token is what an engine buys.

TIME TO FIRST TOKEN, MS →qwen3-coder-next · cachedvllm: 147 ms147ollama: 286 ms286qwen3-coder-next · coldvllm: 150 ms150ollama: 351 ms351ollamavllm

Prompt cache — cold vs cached cells

Cold sends a uniquely salted prompt every time so nothing can be reused; cached repeats one identical prompt after a priming request. A backend whose prefix caching is off or unsupported shows roughly the same first-token latency in both columns — which looks like ordinary slowness rather than a misconfiguration.

ModelPromptThinkCold TTFT Cached TTFTSpeed-up
gemma4:26b0off487 ms475 ms1.0xno measurable reuse
qwen3-coder-next0off150 ms147 ms1.0xno measurable reuse
qwen3-coder:tuned0off318 ms252 ms1.3xno measurable reuse
qwen3-coder-next:latest0off351 ms286 ms1.2xno measurable reuse
llama4:latest0off831 ms819 ms1.0xno measurable reuse
DeepSeek-V4-Flash-0731-UD-IQ2_XXS0off298 ms141 ms2.1xcache is working
qwen3.6:27b0off545 ms473 ms1.2xno measurable reuse
gemma3:27b0off638 ms611 ms1.0xno measurable reuse
llama3:70b-instruct0off474 ms432 ms1.1xno measurable reuse
qwen3-coder:30b0off264 ms211 ms1.3xno measurable reuse
gemma4:latest0off466 ms447 ms1.0xno measurable reuse
gpt-oss:120b0off658 ms533 ms1.2xno measurable reuse
minicpm-v4.5:latest0off243 ms238 ms1.0xno measurable reuse
qwen3:0.6b0off208 ms193 ms1.1xno measurable reuse
codellama:70b0off320 ms257 ms1.2xno measurable reuse
qwen3:4b0off261 ms230 ms1.1xno measurable reuse
ornith-nvfp40off

Prompt cache — warm-up vs measured

The warm-up sends the same prompt the measured runs use, so its first-token time is that prompt’s cold prefill; everything after it is served warm. A backend whose prefix caching is off or unsupported shows no gap between these two columns.

ConfigurationPromptCold TTFTCached TTFTFaster by
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,7680439 ms140 ms3x
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,1440528 ms143 ms4x
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,0720432 ms141 ms3x
qwen3:0.6b · ollama · cached02,287 ms193 ms12x

Correctness by tier

The core tier confirms a model is not broken; it saturates for anything capable, which is exactly why the hard tier exists. Compare two models on the hard row when both score 100% on core.

12 configurations cleared both tiers in full and are not listed.

ConfigurationCoreHard
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768100%83%
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144100%83%
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072100%83%
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768100%83%
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144100%83%
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072100%83%
qwen3.6:27b · ollama · cached100%83%
qwen3.6:27b · 16.2 GB · ollama · cold100%83%
gemma3:27b · 16.2 GB · ollama · cold100%83%
gemma3:27b · ollama · cached100%83%
llama3:70b-instruct · 37.2 GB · ollama · cold83%92%
qwen3-coder:30b · ollama · cached75%100%
qwen3-coder:30b · 17.3 GB · ollama · cold75%100%
gemma4 · 8.9 GB · ollama · cold92%67%
gemma4 · ollama · cached92%67%
llama3:70b-instruct · ollama · cached83%83%
gpt-oss:120b · ollama · cached79%83%
gpt-oss:120b · 60.9 GB · ollama · cold79%83%
minicpm-v4.5 · 5.7 GB · ollama · cold67%67%
minicpm-v4.5 · ollama · cached62%67%
qwen3:0.6b · ollama · cached25%17%
qwen3:0.6b · 0.5 GB · ollama · cold21%8%
codellama:70b · ollama · cached17%17%
codellama:70b · 36.2 GB · ollama · cold17%17%
qwen3:4b · ollama · cached17%0%
qwen3:4b · 2.3 GB · ollama · cold17%0%

Per-task correctness

Share of responses that passed every case for that task. A model strong everywhere except one task and a model mediocre throughout can share an overall average.

TaskPerfect inConfigurations that missed it
balancedBracket matching for () [] {}32 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached
binary_searchFind a value's index in a sorted list32 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached
calculatorArithmetic with precedence, no eval12 of 38DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768, codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma3:27b · 16.2 GB · ollama · cold, gemma3:27b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3.6:27b · 16.2 GB · ollama · cold, qwen3.6:27b · ollama · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
compare_versionsCompare dotted version strings28 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
edit_distanceLevenshtein distance30 of 38minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
flattenFlatten arbitrarily nested lists30 of 38ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
group_anagramsGroup words that are anagrams26 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
lis_lengthLongest strictly increasing subsequence28 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
lru_cache_simSimulate an LRU cache's get/put sequence30 of 38gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
merge_intervalsMerge overlapping intervals30 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
romanInteger to Roman numeral27 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
run_lengthRun-length encode a string30 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
simplify_pathCanonicalise a Unix path with . and ..30 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
spiralMatrix in clockwise spiral order26 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
topo_sortTopological sort of a dependency graph29 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
two_sumIndices of the pair summing to a target30 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
word_breakSegment a string into dictionary words30 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached
word_freqTop-n most common words with tie rules22 of 38codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached

What the failures actually looked like

The first failing case per configuration: the call that was made, what came back, and what should have — or the compile error or timeout that stopped it. This is what a percentage point of correctness is made of.

calculator — Arithmetic with precedence, no eval · 4 examples
word_freq — Top-n most common words with tie rules · 4 examples
group_anagrams — Group words that are anagrams · 4 examples
spiral — Matrix in clockwise spiral order · 4 examples
roman — Integer to Roman numeral · 4 examples
lis_length — Longest strictly increasing subsequence · 4 examples
compare_versions — Compare dotted version strings · 4 examples
topo_sort — Topological sort of a dependency graph · 4 examples
flatten — Flatten arbitrarily nested lists · 4 examples
lru_cache_sim — Simulate an LRU cache's get/put sequence · 4 examples
edit_distance — Levenshtein distance · 4 examples
merge_intervals — Merge overlapping intervals · 4 examples
two_sum — Indices of the pair summing to a target · 4 examples
run_length — Run-length encode a string · 4 examples
simplify_path — Canonicalise a Unix path with . and .. · 4 examples
word_break — Segment a string into dictionary words · 4 examples
binary_search — Find a value's index in a sorted list · 4 examples
balanced — Bracket matching for () [] {} · 4 examples

Hardware

GPU

NVIDIA GB10 · 122 GB

CPU

Cortex-X925

Cores

20

OS

Ubuntu 24.04.4 LTS (aarch64)

Kernel

6.17.0-1018-nvidia

Proxy

0.2.0

CPU, OS and kernel read from the host at report time — this run predates their capture, and they do not change between runs.

What was tested

Every task asks for one answer in the task's language — Python, JavaScript under node, C and C++ under gcc, Rust, C#, or PHP, each run in a separate process under a timeout with the return value compared against the expected one. HTML and CSS tasks are graded structurally: the answer is parsed and checked against required structure (bindings, attributes, declarations in the right context) — a claim about the markup, not about how a browser renders it. Tasks whose toolchain is absent on the machine are skipped and listed here, never scored as zero. Fully correct counts only responses where every case for that task passed; cases is the share of individual cases that passed, so a near-miss still scores there. A response with no extractable code block scores zero — that measures instruction-following, not coding.

Suite

coding-v1

Tasks

18

Cases

81

Repeats

2 per task

Languages

1

Core — 12 tasks, 49 cases

  • binary_search Find a value's index in a sorted list
  • merge_intervals Merge overlapping intervals
  • word_freq Top-n most common words with tie rules
  • roman Integer to Roman numeral
  • balanced Bracket matching for () [] {}
  • flatten Flatten arbitrarily nested lists
  • two_sum Indices of the pair summing to a target
  • lru_cache_sim Simulate an LRU cache's get/put sequence
  • group_anagrams Group words that are anagrams
  • run_length Run-length encode a string
  • compare_versions Compare dotted version strings
  • spiral Matrix in clockwise spiral order

Hard — 6 tasks, 32 cases

  • edit_distance Levenshtein distance
  • lis_length Longest strictly increasing subsequence
  • simplify_path Canonicalise a Unix path with . and ..
  • calculator Arithmetic with precedence, no eval
  • word_break Segment a string into dictionary words
  • topo_sort Topological sort of a dependency graph

Cold-start cost

Time for the discarded warm-up request — the price of making the model resident, excluded from every measurement above.

SECONDS TO LOAD BEFORE THE FIRST USEFUL TOKEN →llama4 · 62.8 GB · ollama · co192sgpt-oss:120b · 60.9 GB · ollam101sllama3:70b-instruct · ollama ·93scodellama:70b · ollama · cache84sgemma3:27b · ollama · cached45sqwen3.6:27b · ollama · cached43sqwen3-coder-next · 48.2 GB · o35sDeepSeek-V4-Flash-0731-UD-IQ2_30sDeepSeek-V4-Flash-0731-UD-IQ2_30sDeepSeek-V4-Flash-0731-UD-IQ2_30s
ConfigurationCold start
llama4 · 62.8 GB · ollama · cold191.7 s
gpt-oss:120b · 60.9 GB · ollama · cold100.7 s
llama3:70b-instruct · ollama · cached93.2 s
codellama:70b · ollama · cached84.4 s
gemma3:27b · ollama · cached45.3 s
qwen3.6:27b · ollama · cached43.0 s
qwen3-coder-next · 48.2 GB · ollama · cold35.1 s
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,14429.9 s
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,07229.9 s
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,76829.8 s

+ 26 more under 29.8 s.

The trade-off — explore

Screen-only companion to the annotated chart above: hover a dot for its numbers, drag a box to zoom into the crowded band, double-click to reset. Hovering a model anywhere on this page highlights it everywhere. The printed report keeps the annotated version.