AI Proxy · benchmark
2026-08-02 11:52 · proxy 0.2.0
gemma4:26b · ollama · cached leads: 100% fully correct at 63.5 tok/s out.
9 other configurations scored the same and were slower. Ranked by correctness first, then output rate. 38 configurations measured across 3 backends.
Run this — 70/30 weighted
gemma4:26b
100% correct · 63.5 tok/s · 16.8 GB · answers in 2.8s · score 76
Runner-up
qwen3-coder-next
100% correct · 62.3 tok/s · 48.2 GB · answers in 2.6s · score 76
Don't be fooled by
qwen3:0.6b
307 tok/s and only 22% correct
Not measured
2 cells
listed in the table as failures, not omitted
Every configuration, placed by correctness and output rate. A point below and to the left of another is beaten on both counts at once, so the dashed frontier is the shortlist — everything off it is dominated by something on it. Hover any point for its name.
Think
off
Temp
0.0
Parallel
1
TTFT is the first token of any kind; TTFC the first content token — the gap between them is time the model spent reasoning. Decode rate is measured from the first token onward, so reasoning tokens count as generated work. Best value in each column is highlighted.
| Configuration | TTFT p50 | Decode p50 | Tokens | Total p50 | Fully correct | Cases | vs best | OK |
|---|---|---|---|---|---|---|---|---|
| gemma4:26b · ollama · cached | 475 ms | 63.5 | 166 | 2,780 ms | 100% | 100% | 4.3x | 36/36 |
| gemma4:26b · 16.8 GB · ollama · cold | 487 ms | 63.5 | 166 | 2,778 ms | 100% | 100% | 4.3x | 36/36 |
| qwen3-coder-next · NVFP4 · vllm · cold | 150 ms | 62.3 | 163 | 2,582 ms | 100% | 100% | 4.0x | 36/36 |
| qwen3-coder-next · NVFP4 · vllm · cached | 147 ms | 62.1 | 165 | 2,468 ms | 100% | 100% | 3.8x | 36/36 |
| qwen3-coder:tuned · 48.2 GB · ollama · cold | 318 ms | 60.6 | 157 | 2,989 ms | 100% | 100% | 4.6x | 36/36 |
| qwen3-coder:tuned · ollama · cached | 252 ms | 60.5 | 157 | 2,928 ms | 100% | 100% | 4.5x | 36/36 |
| qwen3-coder-next · ollama · cached | 286 ms | 51.6 | 157 | 3,435 ms | 100% | 100% | 5.3x | 36/36 |
| qwen3-coder-next · 48.2 GB · ollama · cold | 351 ms | 51.6 | 157 | 3,490 ms | 100% | 100% | 5.3x | 36/36 |
| llama4 · ollama · cached | 819 ms | 18.1 | 123 | 6,837 ms | 100% | 100% | 10.5x | 36/36 |
| llama4 · 62.8 GB · ollama · cold | 831 ms | 18.1 | 123 | 7,214 ms | 100% | 100% | 11.1x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768 | 140 ms | 17.8 | 166 | 8,112 ms | 94% | 93% | 12.4x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144 | 143 ms | 17.8 | 166 | 8,103 ms | 94% | 93% | 12.4x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072 | 141 ms | 17.7 | 166 | 8,118 ms | 94% | 93% | 12.4x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768 | 336 ms | 17.7 | 166 | 8,390 ms | 94% | 93% | 12.9x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144 | 294 ms | 17.7 | 166 | 8,407 ms | 94% | 93% | 12.9x | 36/36 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072 | 298 ms | 17.6 | 166 | 8,412 ms | 94% | 93% | 12.9x | 36/36 |
| qwen3.6:27b · ollama · cached | 473 ms | 12.1 | 169 | 10,656 ms | 94% | 93% | 16.3x | 36/36 |
| qwen3.6:27b · 16.2 GB · ollama · cold | 545 ms | 12.1 | 169 | 10,750 ms | 94% | 93% | 16.5x | 36/36 |
| gemma3:27b · 16.2 GB · ollama · cold | 638 ms | 11.5 | 249 | 19,465 ms | 94% | 94% | 29.8x | 36/36 |
| gemma3:27b · ollama · cached | 611 ms | 11.5 | 248 | 19,139 ms | 94% | 94% | 29.3x | 36/36 |
| llama3:70b-instruct · 37.2 GB · ollama · cold | 474 ms | 5.6 | 105 | 17,153 ms | 86% | 96% | 26.3x | 36/36 |
| qwen3-coder:30b · ollama · cached | 211 ms | 83.9 | 161 | 2,144 ms | 83% | 88% | 3.3x | 36/36 |
| qwen3-coder:30b · 17.3 GB · ollama · cold | 264 ms | 83.4 | 160 | 2,233 ms | 83% | 88% | 3.4x | 36/36 |
| gemma4 · 8.9 GB · ollama · cold | 466 ms | 55.7 | 330 | 6,498 ms | 83% | 83% | 10.0x | 36/36 |
| gemma4 · ollama · cached | 447 ms | 55.7 | 330 | 6,496 ms | 83% | 83% | 10.0x | 36/36 |
| llama3:70b-instruct · ollama · cached | 432 ms | 5.6 | 105 | 17,071 ms | 83% | 94% | 26.1x | 36/36 |
| gpt-oss:120b · ollama · cached | 533 ms | 38.5 | 349 | 8,825 ms | 81% | 81% | 13.5x | 36/36 |
| gpt-oss:120b · 60.9 GB · ollama · cold | 658 ms | 37.8 | 358 | 9,746 ms | 81% | 81% | 14.9x | 36/36 |
| minicpm-v4.5 · 5.7 GB · ollama · cold | 243 ms | 39.9 | 122 | 2,713 ms | 67% | 82% | 4.2x | 36/36 |
| minicpm-v4.5 · ollama · cached | 238 ms | 39.9 | 122 | 2,720 ms | 64% | 80% | 4.2x | 36/36 |
| qwen3:0.6b · ollama · cached | 193 ms | 307.1 | 146 | 686 ms | 22% | 52% | 1.1x | 36/36 |
| qwen3:0.6b · 0.5 GB · ollama · cold | 208 ms | 305.5 | 143 | 653 ms | 17% | 48% | 1.0x | 36/36 |
| codellama:70b · ollama · cached | 257 ms | 5.7 | 247 | 36,147 ms | 17% | 24% | 55.4x | 36/36 |
| codellama:70b · 36.2 GB · ollama · cold | 320 ms | 5.7 | 243 | 36,130 ms | 17% | 25% | 55.3x | 36/36 |
| qwen3:4b · ollama · cached | 230 ms | 72.7 | 487 | 7,276 ms | 11% | 14% | 11.1x | 36/36 |
| qwen3:4b · 2.3 GB · ollama · cold | 261 ms | 72.1 | 487 | 7,354 ms | 11% | 14% | 11.3x | 36/36 |
| ornith-nvfp4 · NVFP4 · vllm · cold | — | — | — | — | — | — | — | None/None |
| ornith-nvfp4 · vllm · cached | — | — | — | — | — | — | — | None/None |
Correctness alone is not a ranking — one point of correctness is not worth half the speed. Each model's best cell scores 70% × correctness + 30% × relative speed, where relative speed is decode rate against the fastest model in this report (qwen3:0.6b, 307.1 tok/s = 1.0). The bar is the weighting made visible: correctness · speed. Drag to change what you value; the ranking recomputes.
| # | Model | Correct | tok/s | Score |
|---|---|---|---|---|
| 1 | gemma4:26b | 100% | 63.5 | 76 |
| 2 | qwen3-coder-next | 100% | 62.3 | 76 |
| 3 | qwen3-coder:tuned | 100% | 60.6 | 76 |
| 4 | llama4 | 100% | 18.1 | 72 |
| 5 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS | 94% | 17.8 | 68 |
| 6 | qwen3.6:27b | 94% | 12.1 | 67 |
| 7 | gemma3:27b | 94% | 11.5 | 67 |
| 8 | qwen3-coder:30b | 83% | 83.9 | 67 |
| 9 | gemma4 | 83% | 55.7 | 64 |
| 10 | llama3:70b-instruct | 86% | 5.6 | 61 |
| 11 | gpt-oss:120b | 81% | 38.5 | 60 |
| 12 | minicpm-v4.5 | 67% | 39.9 | 51 |
| 13 | qwen3:0.6b | 22% | 307.1 | 46 |
| 14 | qwen3:4b | 11% | 72.7 | 15 |
| 15 | codellama:70b | 17% | 5.7 | 12 |
Position is footprint against correctness; bubble area is output speed. Models whose size cannot be read — vLLM checkpoints live inside their containers — are absent, not zero.
The only controlled engine comparison a run can contain: identical weights reachable through more than one backend, paired by cache state. When output rates tie, the wait for the first token is what an engine buys.
Cold sends a uniquely salted prompt every time so nothing can be reused; cached repeats one identical prompt after a priming request. A backend whose prefix caching is off or unsupported shows roughly the same first-token latency in both columns — which looks like ordinary slowness rather than a misconfiguration.
| Model | Prompt | Think | Cold TTFT | Cached TTFT | Speed-up | |
|---|---|---|---|---|---|---|
gemma4:26b | 0 | off | 487 ms | 475 ms | 1.0x | no measurable reuse |
qwen3-coder-next | 0 | off | 150 ms | 147 ms | 1.0x | no measurable reuse |
qwen3-coder:tuned | 0 | off | 318 ms | 252 ms | 1.3x | no measurable reuse |
qwen3-coder-next:latest | 0 | off | 351 ms | 286 ms | 1.2x | no measurable reuse |
llama4:latest | 0 | off | 831 ms | 819 ms | 1.0x | no measurable reuse |
DeepSeek-V4-Flash-0731-UD-IQ2_XXS | 0 | off | 298 ms | 141 ms | 2.1x | cache is working |
qwen3.6:27b | 0 | off | 545 ms | 473 ms | 1.2x | no measurable reuse |
gemma3:27b | 0 | off | 638 ms | 611 ms | 1.0x | no measurable reuse |
llama3:70b-instruct | 0 | off | 474 ms | 432 ms | 1.1x | no measurable reuse |
qwen3-coder:30b | 0 | off | 264 ms | 211 ms | 1.3x | no measurable reuse |
gemma4:latest | 0 | off | 466 ms | 447 ms | 1.0x | no measurable reuse |
gpt-oss:120b | 0 | off | 658 ms | 533 ms | 1.2x | no measurable reuse |
minicpm-v4.5:latest | 0 | off | 243 ms | 238 ms | 1.0x | no measurable reuse |
qwen3:0.6b | 0 | off | 208 ms | 193 ms | 1.1x | no measurable reuse |
codellama:70b | 0 | off | 320 ms | 257 ms | 1.2x | no measurable reuse |
qwen3:4b | 0 | off | 261 ms | 230 ms | 1.1x | no measurable reuse |
ornith-nvfp4 | 0 | off | — | — | — | — |
The warm-up sends the same prompt the measured runs use, so its first-token time is that prompt’s cold prefill; everything after it is served warm. A backend whose prefix caching is off or unsupported shows no gap between these two columns.
| Configuration | Prompt | Cold TTFT | Cached TTFT | Faster by |
|---|---|---|---|---|
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768 | 0 | 439 ms | 140 ms | 3x |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144 | 0 | 528 ms | 143 ms | 4x |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072 | 0 | 432 ms | 141 ms | 3x |
| qwen3:0.6b · ollama · cached | 0 | 2,287 ms | 193 ms | 12x |
The core tier confirms a model is not broken; it saturates for anything capable, which is exactly why the hard tier exists. Compare two models on the hard row when both score 100% on core.
12 configurations cleared both tiers in full and are not listed.
| Configuration | Core | Hard |
|---|---|---|
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768 | 100% | 83% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144 | 100% | 83% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072 | 100% | 83% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768 | 100% | 83% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144 | 100% | 83% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072 | 100% | 83% |
| qwen3.6:27b · ollama · cached | 100% | 83% |
| qwen3.6:27b · 16.2 GB · ollama · cold | 100% | 83% |
| gemma3:27b · 16.2 GB · ollama · cold | 100% | 83% |
| gemma3:27b · ollama · cached | 100% | 83% |
| llama3:70b-instruct · 37.2 GB · ollama · cold | 83% | 92% |
| qwen3-coder:30b · ollama · cached | 75% | 100% |
| qwen3-coder:30b · 17.3 GB · ollama · cold | 75% | 100% |
| gemma4 · 8.9 GB · ollama · cold | 92% | 67% |
| gemma4 · ollama · cached | 92% | 67% |
| llama3:70b-instruct · ollama · cached | 83% | 83% |
| gpt-oss:120b · ollama · cached | 79% | 83% |
| gpt-oss:120b · 60.9 GB · ollama · cold | 79% | 83% |
| minicpm-v4.5 · 5.7 GB · ollama · cold | 67% | 67% |
| minicpm-v4.5 · ollama · cached | 62% | 67% |
| qwen3:0.6b · ollama · cached | 25% | 17% |
| qwen3:0.6b · 0.5 GB · ollama · cold | 21% | 8% |
| codellama:70b · ollama · cached | 17% | 17% |
| codellama:70b · 36.2 GB · ollama · cold | 17% | 17% |
| qwen3:4b · ollama · cached | 17% | 0% |
| qwen3:4b · 2.3 GB · ollama · cold | 17% | 0% |
Share of responses that passed every case for that task. A model strong everywhere except one task and a model mediocre throughout can share an overall average.
| Task | Perfect in | Configurations that missed it |
|---|---|---|
balancedBracket matching for () [] {} | 32 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached |
binary_searchFind a value's index in a sorted list | 32 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached |
calculatorArithmetic with precedence, no eval | 12 of 38 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 131,072, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 262,144, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cold · 32,768, codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma3:27b · 16.2 GB · ollama · cold, gemma3:27b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3.6:27b · 16.2 GB · ollama · cold, qwen3.6:27b · ollama · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
compare_versionsCompare dotted version strings | 28 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
edit_distanceLevenshtein distance | 30 of 38 | minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
flattenFlatten arbitrarily nested lists | 30 of 38 | ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
group_anagramsGroup words that are anagrams | 26 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
lis_lengthLongest strictly increasing subsequence | 28 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
lru_cache_simSimulate an LRU cache's get/put sequence | 30 of 38 | gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
merge_intervalsMerge overlapping intervals | 30 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
romanInteger to Roman numeral | 27 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
run_lengthRun-length encode a string | 30 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
simplify_pathCanonicalise a Unix path with . and .. | 30 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
spiralMatrix in clockwise spiral order | 26 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
topo_sortTopological sort of a dependency graph | 29 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
two_sumIndices of the pair summing to a target | 30 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
word_breakSegment a string into dictionary words | 30 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3:0.6b · 0.5 GB · ollama · cold, qwen3:0.6b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
word_freqTop-n most common words with tie rules | 22 of 38 | codellama:70b · 36.2 GB · ollama · cold, codellama:70b · ollama · cached, gemma4 · 8.9 GB · ollama · cold, gemma4 · ollama · cached, gpt-oss:120b · 60.9 GB · ollama · cold, gpt-oss:120b · ollama · cached, llama3:70b-instruct · 37.2 GB · ollama · cold, llama3:70b-instruct · ollama · cached, minicpm-v4.5 · 5.7 GB · ollama · cold, minicpm-v4.5 · ollama · cached, ornith-nvfp4 · NVFP4 · vllm · cold, ornith-nvfp4 · vllm · cached, qwen3-coder:30b · 17.3 GB · ollama · cold, qwen3-coder:30b · ollama · cached, qwen3:4b · 2.3 GB · ollama · cold, qwen3:4b · ollama · cached |
The first failing case per configuration: the call that was made, what came back, and what should have — or the compile error or timeout that stopped it. This is what a percentage point of correctness is made of.
calculator — Arithmetic with precedence, no eval · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def calculate(expr: str) -> int:
"""
Evaluate an arithmetic expression with +, -, *, / and non-negative integers.
Normal precedence: * and / before + and -, left-to-right.
Division truncates toward zero.
"""
# Tokenize the expression: numbers and operators
tokens = []
i = 0
n = len(expr)
while i < n:
if expr[i].isspace():
i += 1
continue
if expr[i].isdigit():
num = 0
while i < n and expr[i].isdigit():
num = num * 10 + int(expr[i])
i += 1
tokens.append(num)
else:
tokens.append(expr[i])
i += 1
# First pass: handle * and /
stack = []
i = 0
while i < len(tokens):
tok = tokens[i]
if isinstance(tok, int):
stack.append(tok)
elif tok == '*' or tok == '/':
# Pop left operand
left = stack.pop()
i += 1
right = tokens[i] # must be int
if tok == '*':
stack.append(left * right)
else:
# Integer division truncating toward zero
if left * right < 0:
# Truncate toward zero: use abs and negate
res = -((-left) // right) if right > 0 else -(left // (-right))
# Simpler: use int(left / right) which truncates toward zero in Python
res = int(left / right)
else:
res = left // right
stack.append(res)
i += 1
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def calculate(expr: str) -> int:
"""
Evaluate an arithmetic expression with +, -, *, / and non-negative integers.
Normal precedence: * and / before + and -, left-to-right.
Division truncates toward zero.
"""
# Tokenize the expression: numbers and operators
tokens = []
i = 0
n = len(expr)
while i < n:
if expr[i].isspace():
i += 1
continue
if expr[i].isdigit():
num = 0
while i < n and expr[i].isdigit():
num = num * 10 + int(expr[i])
i += 1
tokens.append(num)
else:
tokens.append(expr[i])
i += 1
# First pass: handle * and /
stack = []
i = 0
while i < len(tokens):
tok = tokens[i]
if isinstance(tok, int):
stack.append(tok)
elif tok == '*' or tok == '/':
# Pop left operand
left = stack.pop()
i += 1
right = tokens[i] # must be int
if tok == '*':
stack.append(left * right)
else:
# Integer division truncating toward zero
if left * right < 0:
# Truncate toward zero: use abs and negate
res = -((-left) // right) if right > 0 else -(left // (-right))
# Simpler: use int(left / right) which truncates toward zero in Python
res = int(left / right)
else:
res = left // right
stack.append(res)
i += 1
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def calculate(expr: str) -> int:
"""
Evaluate an arithmetic expression with +, -, *, / and non-negative integers.
Normal precedence: * and / before + and -, left-to-right.
Division truncates toward zero.
"""
# Tokenize the expression: numbers and operators
tokens = []
i = 0
n = len(expr)
while i < n:
if expr[i].isspace():
i += 1
continue
if expr[i].isdigit():
num = 0
while i < n and expr[i].isdigit():
num = num * 10 + int(expr[i])
i += 1
tokens.append(num)
else:
tokens.append(expr[i])
i += 1
# First pass: handle * and /
stack = []
i = 0
while i < len(tokens):
tok = tokens[i]
if isinstance(tok, int):
stack.append(tok)
elif tok == '*' or tok == '/':
# Pop left operand
left = stack.pop()
i += 1
right = tokens[i] # must be int
if tok == '*':
stack.append(left * right)
else:
# Integer division truncating toward zero
if left * right < 0:
# Truncate toward zero: use abs and negate
res = -((-left) // right) if right > 0 else -(left // (-right))
# Simpler: use int(left / right) which truncates toward zero in Python
res = int(left / right)
else:
res = left // right
stack.append(res)
i += 1
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def calculate(expr: str) -> int:
"""
Evaluate an arithmetic expression with +, -, *, / and non-negative integers.
Normal precedence: * and / before + and -, left-to-right.
Division truncates toward zero.
"""
# Tokenize the expression: numbers and operators
tokens = []
i = 0
n = len(expr)
while i < n:
if expr[i].isspace():
i += 1
continue
if expr[i].isdigit():
num = 0
while i < n and expr[i].isdigit():
num = num * 10 + int(expr[i])
i += 1
tokens.append(num)
else:
tokens.append(expr[i])
i += 1
# First pass: handle * and /
stack = []
i = 0
while i < len(tokens):
tok = tokens[i]
if isinstance(tok, int):
stack.append(tok)
elif tok == '*' or tok == '/':
# Pop left operand
left = stack.pop()
i += 1
right = tokens[i] # must be int
if tok == '*':
stack.append(left * right)
else:
# Integer division truncating toward zero
if left * right < 0:
# Truncate toward zero: use abs and negate
res = -((-left) // right) if right > 0 else -(left // (-right))
# Simpler: use int(left / right) which truncates toward zero in Python
res = int(left / right)
else:
res = left // right
stack.append(res)
i += 1
… (truncated)word_freq — Top-n most common words with tie rules · 4 examplesword_freq("the cat the dog THE bird", 2) → [["the", 3], ["cat", 1]], expected [["the", 3], ["bird", 1]]import re
from collections import Counter
def word_freq(text: str, n: int) -> list[tuple[str, int]]:
words = re.findall(r'\b\w+\b', text.lower())
freq = Counter(words).most_common(n)
return [(word, count) for word, count in sorted(freq, key=lambda x: (-x[1], x[0]))]
word_freq("the cat the dog THE bird", 2) → [["the", 3], ["cat", 1]], expected [["the", 3], ["bird", 1]]import re
from collections import Counter
def word_freq(text: str, n: int) -> list[tuple[str, int]]:
# Extract words (runs of letters) and convert to lowercase
words = re.findall(r'[a-zA-Z]+', text.lower())
# Count frequency of each word
word_counts = Counter(words)
# Get n most common words, sorted by count descending then word ascending
return word_counts.most_common(n)
word_freq("the cat the dog THE bird", 2) → [["the", 3], ["cat", 1]], expected [["the", 3], ["bird", 1]]import re
from collections import Counter
def word_freq(text: str, n: int) -> list[tuple[str, int]]:
# Extract words (runs of letters) and convert to lowercase
words = re.findall(r'[a-zA-Z]+', text.lower())
# Count frequency of each word
word_counts = Counter(words)
# Get n most common words, sorted by count descending then word ascending
return sorted(word_counts.most_common(n), key=lambda x: (-x[1], x[0]))
SyntaxError: invalid syntax (<model>, line 1)import re
from collections import Counter
def word_freq(text: str, n: int) -> list[tuple[str, int]]:
"""
Returns the n most common lowercase words in text, as (word, count) tuples
sorted by count descending then word ascending. Words are runs of letters;
ignore case and punctuation.
Args:
text: The input string.
n: The number of top frequent words to return.
Returns:
A list of (word, count) tuples.
"""
# Convert to lowercase and find all sequences of letters (words)
words = re.findall(r'[a-zA-Z]+', text.lower())
# Count the frequency of each word
word_counts = Counter(words)
# Get all unique words and their counts as a list of tuples
items = list(word_counts.items())
# Sort the items:
# 1. By count (descending): -count
# 2. By word (ascending): word
sorted_items = sorted(items, key=lambda item: (-item[1], item[0]))
# Return the top n results
return sorted_items[:n]
if __name__ == '__main__':
# Example 1: Basic test
text1 = "The quick brown fox jumps over the lazy dog. Fox, fox!"
n1 = 3
print(f"Text: '{text1}'\nTop {n1} words:")
result1 = word_freq(text1, n1)
print(result1) # Expected: [('fox', 3), ('the', 2), ('quick', 1)] or similar based on tie-breaking
# Example 2: Case and punctuation test
text2 = "Apple banana! apple. BANANA, orange? Apple."
n2 = 4
print(f"\nText: '{text2}'\nTop {n2} words:")
result2 = word_freq(text2, n2)
# Counts: apple: 3, banana: 1, orange: 1. Wait, banana appears twice if we
… (truncated)group_anagrams — Group words that are anagrams · 4 examplesgroup_anagrams(["eat", "tea", "tan", "ate", "nat", "bat"]) → [["ate", "eat", "tea"], ["nat", "tan"], ["bat"]], expected [["ate", "eat", "tea"], ["bat"], ["nat", "tan"]]def group_anagrams(words: list[str]) -> list[list[str]]:
anagrams = {}
for word in words:
sorted_word = "".join(sorted(word))
if sorted_word not in anagrams:
anagrams[sorted_word] = [word]
else:
anagrams[sorted_word].append(word)
return [sorted(group) for group in anagrams.values()]
group_anagrams(["eat", "tea", "tan", "ate", "nat", "bat"]) → [["ate", "eat", "tea"], ["nat", "tan"], ["bat"]], expected [["ate", "eat", "tea"], ["bat"], ["nat", "tan"]]def group_anagrams(words: list[str]) -> list[list[str]]:
anagrams = {}
for word in words:
sorted_word = "".join(sorted(word))
if sorted_word not in anagrams:
anagrams[sorted_word] = [word]
else:
anagrams[sorted_word].append(word)
return [sorted(group) for group in anagrams.values()]
group_anagrams(["eat", "tea", "tan", "ate", "nat", "bat"]) → [["eat", "tea", "ate"], ["bat"], ["tan", "nat"]], expected [["ate", "eat", "tea"], ["bat"], ["nat", "tan"]]def group_anagrams(words: list[str]) -> list[list[str]]:
from collections import defaultdict
anagram_groups = defaultdict(list)
for word in words:
sorted_word = ''.join(sorted(word))
anagram_groups[sorted_word].append(word)
result = [group for group in sorted(anagram_groups.values(), key=lambda x: min(x))]
return result
group_anagrams(["eat", "tea", "tan", "ate", "nat", "bat"]) → [["eat", "tea", "ate"], ["bat"], ["tan", "nat"]], expected [["ate", "eat", "tea"], ["bat"], ["nat", "tan"]]def group_anagrams(words: list[str]) -> list[list[str]]:
from collections import defaultdict
anagram_groups = defaultdict(list)
for word in words:
sorted_word = ''.join(sorted(word))
anagram_groups[sorted_word].append(word)
result = [group for group in sorted(anagram_groups.values(), key=lambda x: min(x))]
return result
spiral — Matrix in clockwise spiral order · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def spiral(matrix: list[list[int]]) -> list[int]:
"""Return the elements of *matrix* in clockwise spiral order."""
if not matrix or not matrix[0]:
return []
res: list[int] = []
top, bottom = 0, len(matrix) - 1
left, right = 0, len(matrix[0]) - 1
while top <= bottom and left <= right:
# Traverse the top row from left to right.
for col in range(left, right + 1):
res.append(matrix[top][col])
top += 1
# Traverse the rightmost column from top to bottom.
for row in range(top, bottom + 1):
res.append(matrix[row][right])
right -= 1
if top <= bottom:
# Traverse the bottom row from right to left.
for col in range(right, left - 1, -1):
res.append(matrix[bottom][col])
bottom -= 1
if left <= right:
SyntaxError: invalid syntax (<model>, line 1)def spiral(matrix: list[list[int]]) -> list[int]:
"""Return the elements of *matrix* in clockwise spiral order."""
if not matrix or not matrix[0]:
return []
res: list[int] = []
top, bottom = 0, len(matrix) - 1
left, right = 0, len(matrix[0]) - 1
while top <= bottom and left <= right:
# Traverse the top row from left to right.
for col in range(left, right + 1):
res.append(matrix[top][col])
top += 1
# Traverse the rightmost column from top to bottom.
for row in range(top, bottom + 1):
res.append(matrix[row][right])
right -= 1
if top <= bottom:
# Traverse the bottom row from right to left.
for col in range(right, left - 1, -1):
res.append(matrix[bottom][col])
bottom -= 1
if left <= right:
spiral([[1, 2, 3], [4, 5, 6], [7, 8, 9]]) → [1, 2, 3, 6, 5, 4, 4, 5, 6, 8, 8, 7], expected [1, 2, 3, 6, 9, 8, 7, 4, 5]def spiral(matrix: list[list[int]]) -> list[int]:
result = []
while matrix:
# Top row
result += matrix.pop(0)
if len(matrix) == 0: break
# Right column
for col in reversed(matrix[0]):
result.append(col)
if len(matrix) == 0: break
matrix = [row[-1::-1] for row in matrix]
# Bottom row (reversed due to previous pop and append operations)
result += matrix.pop(0)[::-1]
if len(matrix) == 0: break
# Left column (reversed again because of the way we're popping rows from left side after reversing them once before appending top row in reversed order, so now they need another reversal to get correct direction for bottom-left to be rightmost)
matrix = [row[1:] for row in matrix]
if len(matrix) == 0: break
result += [matrix[-1][0]]
return result
spiral([[1, 2, 3], [4, 5, 6], [7, 8, 9]]) → [1, 2, 3, 6, 5, 4, 4, 5, 6, 7, 8, 9, 9, 8, 7], expected [1, 2, 3, 6, 9, 8, 7, 4, 5]def spiral(matrix: list[list[int]]) -> list[int]:
result = []
while matrix:
# Top row
result += matrix.pop(0)
if not matrix: break
# Right column
for col in reversed(matrix[0]):
result.append(col)
if not matrix: break
matrix = [row[-1::-1] for row in matrix]
# Bottom row (reversed due to previous pop and reverse steps)
result += matrix.pop(0)[::-1]
if not matrix: break
# Left column (reversed again because of the way we're popping rows from bottom up after reversing each row initially when removing right columns)
for col in reversed(matrix[0]):
result.append(col)
return result
roman — Integer to Roman numeral · 4 examplesno code in responseno code in responseto_roman(1) → "", expected "I"def to_roman(n: int) -> str:
if not (1 <= n <= 3999):
raise ValueError("Input must be between 1 and 3999")
roman_numerals = [
("M", 1000),
("CM", 900),
("D", 500),
("CD", 400),
("C", 100),
("XC", 90),
("L", 50),
("XL", 40),
("X", 10),
("IX", 9),
("V", 5),
("IV", 4)
]
result = []
for symbol, value in roman_numerals:
while n >= value:
result.append(symbol)
n -= value
return ''.join(result)
to_roman(1) → "", expected "I"def to_roman(n: int) -> str:
if not (1 <= n <= 3999):
raise ValueError("Input must be between 1 and 3999")
roman_numerals = [
("M", 1000),
("CM", 900),
("D", 500),
("CD", 400),
("C", 100),
("XC", 90),
("L", 50),
("XL", 40),
("X", 10),
("IX", 9),
("V", 5),
("IV", 4)
]
result = []
for symbol, value in roman_numerals:
while n >= value:
result.append(symbol)
n -= value
return ''.join(result)
lis_length — Longest strictly increasing subsequence · 4 examplesSyntaxError: invalid syntax (<model>, line 1)import bisect
def lis_length(nums: list[int]) -> int:
"""
Returns the length of the longest strictly increasing subsequence (LIS)
of the given list of numbers.
"""
if not nums:
return 0
# tails[i] stores the smallest tail of all increasing subsequences
# with length i+1. This array will always be sorted.
tails = []
for num in nums:
# Find the index 'i' where 'num' can replace an element
# to potentially form a shorter or equal-length subsequence
# ending with a smaller tail, but here we are looking for
# the smallest element in tails that is >= num.
# bisect_left returns an insertion point which comes before (to the left of)
# any existing entries of 'num'. Since we need strictly increasing,
# if tails[i] == num, we don't want to replace it because it doesn't help
# us find a *strictly* longer sequence ending with a smaller tail.
# We are looking for the smallest element in 'tails' that is >= num.
# If we find such an element at index 'i', replacing it with 'num'
# maintains the property of tails[i] being the smallest tail for length i+1,
# and since num <= tails[i], this replacement is valid or beneficial.
# For LIS, we want to find the first element in 'tails' that is >= num.
# If all elements are < num, then num extends the longest sequence found so far.
# bisect_left finds the insertion point for 'num' to maintain sorted order.
i = bisect.bisect_left(tails, num)
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import bisect
def lis_length(nums: list[int]) -> int:
"""
Returns the length of the longest strictly increasing subsequence (LIS)
of the given list of numbers.
"""
if not nums:
return 0
# tails[i] stores the smallest tail of all increasing subsequences
# with length i+1. This array will always be sorted.
tails = []
for num in nums:
# Find the index 'i' where 'num' can replace an element
# to potentially form a shorter or equal-length subsequence
# ending with a smaller tail, but here we are looking for
# the smallest element in tails that is >= num.
# bisect_left returns an insertion point which comes before (to the left of)
# any existing entries of 'num'. Since we need strictly increasing,
# if tails[i] == num, we don't want to replace it because it doesn't help
# us find a *strictly* longer sequence ending with a smaller tail.
# We are looking for the smallest element in 'tails' that is >= num.
# If we find such an element at index 'i', replacing it with 'num'
# maintains the property of tails[i] being the smallest tail for length i+1,
# and since num <= tails[i], this replacement is valid or beneficial.
# For LIS, we want to find the first element in 'tails' that is >= num.
# If all elements are < num, then num extends the longest sequence found so far.
# bisect_left finds the insertion point for 'num' to maintain sorted order.
i = bisect.bisect_left(tails, num)
… (truncated)lis_length([]) → null, expected 0def lis_length(nums: list[int]) -> int:
n = len(nums)
dp = [1] * n
for i in range(1, n):
for j in range(i):
if nums[i] > nums[j]:
dp[i] = max(dp[i], dp[j] + 1)
return dp[-1]
lis_length([]) → null, expected 0def lis_length(nums: list[int]) -> int:
n = len(nums)
dp = [1] * n
for i in range(1, n):
for j in range(i):
if nums[i] > nums[j]:
dp[i] = max(dp[i], dp[j] + 1)
return dp[-1]
compare_versions — Compare dotted version strings · 4 examplescompare_versions("1.0.0", "1") → null, expected 0def compare_versions(a: str, b: str) -> int:
def to_list(s):
return [int(x.lstrip('0')) for x in s.split('.')]
list_a = to_list(a)
list_b = to_list(b)
max_len = max(len(list_a), len(list_b))
padded_a = list_a + [0] * (max_len - len(list_a))
padded_b = list_b + [0] * (max_len - len(list_b))
for i in range(max_len):
if padded_a[i] < padded_b[i]:
return -1
elif padded_a[i] > padded_b[i]:
return 1
return 0
compare_versions("1.0.0", "1") → null, expected 0def compare_versions(a: str, b: str) -> int:
def to_list(s):
return [int(x.lstrip('0')) for x in s.split('.')]
list_a = to_list(a)
list_b = to_list(b)
max_len = max(len(list_a), len(list_b))
padded_a = list_a + [0] * (max_len - len(list_a))
padded_b = list_b + [0] * (max_len - len(list_b))
for i in range(max_len):
if padded_a[i] < padded_b[i]:
return -1
elif padded_a[i] > padded_b[i]:
return 1
return 0
compare_versions("1.2", "1.10") → -98, expected -1def compare_versions(a: str, b: str) -> int:
a = a.replace('.', '')
b = b.replace('.', '')
# Convert both to integers with leading zeros ignored
a = int(a, 10)
b = int(b, 10)
# Compare as integers
return a - b
compare_versions("1.2", "1.10") → -98, expected -1def compare_versions(a: str, b: str) -> int:
a = a.replace('.', '')
b = b.replace('.', '')
# Convert both to integers with leading zeros ignored
a = int(a, 10)
b = int(b, 10)
# Compare as integers
return a - b
topo_sort — Topological sort of a dependency graph · 4 examplesSyntaxError: invalid syntax (<model>, line 1)import heapq
from collections import defaultdict
def topo_sort(n: int, edges: list[list[int]]) -> list[int]:
"""
Computes a topological ordering of nodes 0..n-1 given directed edges.
If multiple orderings are valid, returns the lexicographically smallest one.
Returns an empty list if a cycle exists.
Args:
n: The number of nodes (0 to n-1).
edges: A list of edges, where each edge [a, b] means a must come before b.
Returns:
A list representing the lexicographically smallest topological ordering,
or an empty list if a cycle is detected.
"""
adj = defaultdict(list)
in_degree = [0] * n
for u, v in edges:
# Edge from u to v (u -> v)
adj[u].append(v)
in_degree[v] += 1
# Min-heap to store nodes with an in-degree of 0.
# Using a min-heap ensures that we always pick the smallest available node,
# guaranteeing the lexicographically smallest result.
min_heap = []
for i in range(n):
if in_degree[i] == 0:
heapq.heappush(min_heap, i)
topological_order = []
nodes_processed = 0
while min_heap:
# Get the smallest available node (lexicographically smallest choice)
u = heapq.heappop(min_heap)
topological_order.append(u)
nodes_processed += 1
# Process neighbors of u
# Sorting neighbors before adding them to the heap is not strictly necessary
# for correctness here because the heap handles the selection, but it's good practice.
for v in adj[u]:
in_degree[v] -= 1
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import heapq
from collections import defaultdict
def topo_sort(n: int, edges: list[list[int]]) -> list[int]:
"""
Computes a topological ordering of nodes 0..n-1 given directed edges.
If multiple orderings are valid, returns the lexicographically smallest one.
Returns an empty list if a cycle exists.
Args:
n: The number of nodes (0 to n-1).
edges: A list of edges, where each edge [a, b] means a must come before b.
Returns:
A list representing the lexicographically smallest topological ordering,
or an empty list if a cycle is detected.
"""
adj = defaultdict(list)
in_degree = [0] * n
for u, v in edges:
# Edge from u to v (u -> v)
adj[u].append(v)
in_degree[v] += 1
# Min-heap to store nodes with an in-degree of 0.
# Using a min-heap ensures that we always pick the smallest available node,
# guaranteeing the lexicographically smallest result.
min_heap = []
for i in range(n):
if in_degree[i] == 0:
heapq.heappush(min_heap, i)
topological_order = []
nodes_processed = 0
while min_heap:
# Get the smallest available node (lexicographically smallest choice)
u = heapq.heappop(min_heap)
topological_order.append(u)
nodes_processed += 1
# Process neighbors of u
# Sorting neighbors before adding them to the heap is not strictly necessary
# for correctness here because the heap handles the selection, but it's good practice.
for v in adj[u]:
in_degree[v] -= 1
… (truncated)topo_sort(4, [[0, 1], [0, 2], [1, 3], [2, 3]]) → [], expected [0, 1, 2, 3]def topo_sort(n: int, edges: list[list[int]]) -> list[int]:
# Build adjacency list
adj = [[] for _ in range(n)]
in_degree = [0] * n
# Build adjacency list and in-degree list
for u, v in edges:
adj[u].append(v)
in_degree[v] += 1
# Use a priority queue (min-heap) to find a topological order
import heapq
# Create a priority queue with nodes with in-degree 0
heap = []
for i in range(n):
if in_degree[i] == 0:
heapq.heappush(heap, i)
# Initialize result list
result = []
# Use a visited set to avoid cycles
visited = set()
while heap:
u = heapq.heappop(heap)
result.append(u)
for v in adj[u]:
if v not in visited:
in_degree[v] -= 1
if in_degree[v] == 0:
heapq.heappush(heap, v)
visited.add(v)
return result if len(result) == n else []
topo_sort(4, [[0, 1], [0, 2], [1, 3], [2, 3]]) → [0, 2, 1, 3], expected [0, 1, 2, 3]def topo_sort(n, edges):
graph = [[] for _ in range(n)]
visited = [False] * n
stack = []
# Create the adjacency list representation of the graph.
for edge in edges:
a, b = edge[0], edge[1]
graph[a].append(b)
def dfs(node):
visited[node] = True
for neighbor in graph[node]:
if not visited[neighbor]:
dfs(neighbor)
stack.append(node)
# Perform DFS on each unvisited node.
for i in range(n):
if not visited[i]:
dfs(i)
order = stack[::-1] # Reverse the stack to get the topological ordering.
# Check if there is a cycle by checking if all nodes were visited.
for i in range(n):
if not visited[i]:
return [] # Cycle exists, so we cannot produce a valid ordering.
return order
flatten — Flatten arbitrarily nested lists · 4 examplesSyntaxError: invalid syntax (<model>, line 1)```python code
def flatten(nested: list) -> list:
result = []
for item in nested:
if isinstance(item, list):
result.extend(flatten(item))
else:
result.append(item)
return result
```SyntaxError: invalid syntax (<model>, line 1)```python code
def flatten(nested: list) -> list:
result = []
for item in nested:
if isinstance(item, list):
result.extend(flatten(item))
else:
result.append(item)
return result
```flatten([1, [2, [3, [4]]], 5]) → [1, 5], expected [1, 2, 3, 4, 5]def flatten(nested: list) -> list:
result = []
stack = []
for item in nested:
if isinstance(item, list):
stack.extend(flatten(item))
else:
result.append(item)
return result
flatten([1, [2, [3, [4]]], 5]) → [1, 5], expected [1, 2, 3, 4, 5]def flatten(nested: list) -> list:
result = []
stack = []
for item in nested:
if isinstance(item, list):
stack.extend(flatten(item))
else:
result.append(item)
return result
lru_cache_sim — Simulate an LRU cache's get/put sequence · 4 examplesSyntaxError: invalid syntax (<model>, line 1)from collections import OrderedDict
def lru(capacity: int, ops: list) -> list:
"""
Simulate an LRU cache.
Parameters
----------
capacity : int
Maximum number of items the cache can hold.
ops : list
A sequence of operations. Each operation is a list where the first element
is either 'put' or 'get'. For 'put' the format is ['put', key, value];
for 'get' it is ['get', key].
Returns
-------
list
Results of all 'get' operations in order; -1 denotes a cache miss.
"""
if capacity <= 0:
# Cache cannot store anything; every get is a miss and puts are ignored.
return [-1 for op in ops if op[0] == 'get']
cache = OrderedDict() # key -> value, ordered by recent use (most recent at end)
results = []
for op in ops:
if not op: # safety guard
continue
typ = op[0].lower()
if typ == 'put':
_, key, val = op
if key in cache:
# Update value and mark as most recently used.
cache.move_to_end(key)
cache[key] = val
else:
# Evict least-recently-used item if needed.
if len(cache) >= capacity:
cache.popitemSyntaxError: invalid syntax (<model>, line 1)from collections import OrderedDict
def lru(capacity: int, ops: list) -> list:
"""
Simulate an LRU cache.
Parameters
----------
capacity : int
Maximum number of items the cache can hold.
ops : list
A sequence of operations. Each operation is a list where the first element
is either 'put' or 'get'. For 'put' the format is ['put', key, value];
for 'get' it is ['get', key].
Returns
-------
list
Results of all 'get' operations in order; -1 denotes a cache miss.
"""
if capacity <= 0:
# Cache cannot store anything; every get is a miss and puts are ignored.
return [-1 for op in ops if op[0] == 'get']
cache = OrderedDict() # key -> value, ordered by recent use (most recent at end)
results = []
for op in ops:
if not op: # safety guard
continue
typ = op[0].lower()
if typ == 'put':
_, key, val = op
if key in cache:
# Update value and mark as most recently used.
cache.move_to_end(key)
cache[key] = val
else:
# Evict least-recently-used item if needed.
if len(cache) >= capacity:
cache.popitemlru(2, [["put", 1, 1], ["put", 2, 2], ["get", 1], ["put", 3, 3], ["get", 2], ["get", 3]]) → null, expected [1, -1, 3]def lru(capacity: int, ops: list) -> list:
from collections import deque
cache = deque()
result = []
for op in ops:
if op == 'put':
key, value = op[1], op[2]
if len(cache) < capacity:
cache.append((key, value))
result.append(1)
else:
# Key exists
result.append(1)
cache.popleft()
cache.append((key, value))
else: # op == 'get'
key = op[1]
if key in cache:
result.append(-1)
result.append(1)
else:
result.append(-1)
cache.append((key, value))
return result
lru(2, [["put", 1, 1], ["put", 2, 2], ["get", 1], ["put", 3, 3], ["get", 2], ["get", 3]]) → [], expected [1, -1, 3]def lru(capacity: int, ops: list) -> list:
from collections import deque
cache = deque()
result = []
for op in ops:
if op == 'put':
key, value = op[1], op[2]
if len(cache) < capacity:
cache.append((key, value))
result.append(1)
else:
# Key exists
result.append(1)
cache.popleft()
cache.append((key, value))
elif op == 'get':
key = op[1]
if key in cache:
result.append(-1)
result.append(1)
else:
result.append(-1)
result.append(1)
return result
edit_distance — Levenshtein distance · 4 examplesedit_distance("kitten", "sitting") → 0, expected 3def edit_distance(a: str, b: str) -> int:
m = len(a)
n = len(b)
dp = [[0] * (n + 1) for _ in range(m + 1)]
for i in range(m + 1):
dp[i][0] = i
for j in range(n + 1):
dp[0][j] = j
for i in range(1, m + 1):
for j in range(1, n + 1):
if a[i - 1] == b[j - 1]:
dp[i][j] = dp[i - 1][j - 1]
else:
dp[i][j] = min(dp[i - 1][j], # Deletion
dp[i][j - 1], # Insertion
dp[i - 1][j - 1]) # Substitution
return dp[m][n]
edit_distance("kitten", "sitting") → 0, expected 3def edit_distance(a: str, b: str) -> int:
m = len(a)
n = len(b)
dp = [[0] * (n + 1) for _ in range(m + 1)]
for i in range(m + 1):
dp[i][0] = i
for j in range(n + 1):
dp[0][j] = j
for i in range(1, m + 1):
for j in range(1, n + 1):
if a[i - 1] == b[j - 1]:
dp[i][j] = dp[i - 1][j - 1]
else:
dp[i][j] = min(dp[i - 1][j], # Deletion
dp[i][j - 1], # Insertion
dp[i - 1][j - 1]) # Substitution
return dp[m][n]
edit_distance("kitten", "sitting") → null, expected 3def edit_distance(a: str, b: str) -> int:
# Base case: if one string is empty, the other is the distance
if a == b:
return 0
if len(a) == 0:
return len(b)
if len(b) == 0:
return len(a)
# If one is empty, the other is the distance
if len(a) == 1:
return edit_distance(a[0], b) + 1
if len(b) == 1:
return edit_distance(a, b[0]) + 1
# Recursive case
return min(
edit_distance(a[:i], b[:i]) + 1,
edit_distance(a[:i], b[i:]) + 1,
edit_distance(a[i:], b[:i]) + 1,
edit_distance(a[i:], b[i:]) + 1
)
edit_distance("kitten", "sitting") → 11, expected 3def edit_distance(a: str, b: str) -> int:
# Base case: if one string is empty, the other is the distance
if a == b:
return 0
if len(a) == 0:
return len(b)
if len(b) == 0:
return len(a)
# If one is empty, the other is the distance
if len(a) == 1 and len(b) == 1:
return 1
elif len(a) == 1:
return len(b) - 1
elif len(b) == 1:
return len(a) - 1
# Recursive case
return 1 + min(edit_distance(a[1:], b[1:]) + 1, edit_distance(a[1:], b) + 1, edit_distance(a, b[1:]))
merge_intervals — Merge overlapping intervals · 4 examplesmerge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) → [[1, 3], [2, 6], [8, 10], [15, 18]], expected [[1, 6], [8, 10], [15, 18]]def merge_intervals(intervals: list[list[int]]) -> list[list[int]]:
# Sort intervals by their start values
intervals.sort(key=lambda x: x[0])
result = []
for start, end in intervals:
if not result or result[-1][1] < end:
result.append([start, end])
else:
# Merge overlapping intervals
result[-1] = [min(result[-1][0], start), max(result[-1][1], end)]
return result
merge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) → [[1, 3], [2, 6], [8, 10], [15, 18]], expected [[1, 6], [8, 10], [15, 18]]def merge_intervals(intervals: list[list[int]]) -> list[list[int]]:
# Sort intervals by their start values
intervals.sort(key=lambda x: x[0])
result = []
for start, end in intervals:
if not result or result[-1][1] < end:
result.append([start, end])
else:
# Merge overlapping intervals
result[-1] = [min(result[-1][0], start), max(result[-1][1], end)]
return result
IndentationError: unexpected indent (<model>, line 1) def merge_intervals(intervals):
# Sort the intervals by their start time
sorted_intervals = sorted(intervals, key=lambda x: x[0])
merged_intervals = []
current_interval = None
for interval in sorted_intervals:
if not merged_intervals or merged_intervals[-1][1] < interval[0]:
# If the current interval is empty, add the new interval to it.
merged_intervals.append(interval)
current_interval = interval
else:
# If there's an overlap between the current and the new intervals, merge them.
if current_interval[1] < interval[1]:
current_interval[1] = interval[1]
return merged_intervals IndentationError: unexpected indent (<model>, line 1) def merge_intervals(intervals):
# Sort the intervals by their start time
sorted_intervals = sorted(intervals, key=lambda x: x[0])
merged_intervals = []
current_interval = None
for interval in sorted_intervals:
if not merged_intervals or merged_intervals[-1][1] < interval[0]:
# If the current interval is empty, add the new interval to it.
merged_intervals.append(interval)
current_interval = interval
else:
# If there's an overlap between the current and new intervals, merge them together.
if current_interval[1] < interval[1]:
current_interval[1] = interval[1]
return merged_intervals two_sum — Indices of the pair summing to a target · 4 examplestwo_sum([3, 2, 4], 6) → [0, 1, 1, 2], expected [1, 2]def two_sum(nums: list[int], target: int) -> list[int]:
indices = []
for i, num in enumerate(nums):
if (target - num) in nums[i:] and nums[i] != nums[i - 1]:
indices.append(i)
indices.append(i + 1)
return indices
two_sum([3, 2, 4], 6) → [0, 1, 1, 2], expected [1, 2]def two_sum(nums: list[int], target: int) -> list[int]:
indices = []
for i, num in enumerate(nums):
if (target - num) in nums[i:] and nums[i] != nums[i - 1]:
indices.append(i)
indices.append(i + 1)
return indices
IndentationError: unexpected indent (<model>, line 1) def two_sum(nums: list[int], target: int) -> list[int]:
num_indices = {}
for i, num in enumerate(nums):
if target - num in num_indices:
return [num_indices[target - num], i]
num_indices[num] = i
return [] IndentationError: unexpected indent (<model>, line 1) def two_sum(nums: list[int], target: int) -> list[int]:
num_indices = {}
for i, num in enumerate(nums):
complement = target - num
if complement not in num_indices:
num_indices[num] = i
else:
return [num_indices[complement], i]
raise ValueError("No solution found") run_length — Run-length encode a string · 4 examplestimeout after 10.0sdef rle(s: str) -> str:
result = []
i = 0
while i < len(s):
char = s[i]
count = 1
while i + count < len(s) and s[i + count] == char:
count += 1
i += count
result.append(char + str(count))
return ''.join(result)
timeout after 10.0sdef rle(s: str) -> str:
result = []
i = 0
while i < len(s):
char = s[i]
count = 1
while i + count < len(s) and s[i + count] == char:
count += 1
i += count
result.append(f"{char}{count}")
return "\n".join(result)
rle("aab") → "a2", expected "a2b1"def rle(s):
encoded_str = ""
count = 1
for i in range(len(s)):
if (i < len(s) - 1 and s[i] == s[i+1]) or i == len(s) - 1:
count += 1
continue
encoded_str += s[i] + str(count)
count = 1
return encoded_str
rle("aab") → "a2", expected "a2b1"def rle(s):
encoded_str = ""
count = 1
for i in range(len(s)):
if (i < len(s) - 1 and s[i] == s[i+1]) or (count == 1):
count += 1
else:
encoded_str += s[i] + str(count)
count = 1
return encoded_str
simplify_path — Canonicalise a Unix path with . and .. · 4 examplessimplify_path("/home/") → "home", expected "/home"def simplify_path(path: str) -> str:
# Split the path into parts
parts = path.split('/')
# Initialize the canonical path
canonical_path = []
# Process each part
for part in parts:
if part == '.':
canonical_path.append('/')
elif part == '..':
if canonical_path:
canonical_path.pop()
else:
canonical_path.append(part)
# If there's a trailing slash, remove it (except for the root)
if len(canonical_path) > 0:
canonical_path[-1] = '/' if len(canonical_path) == 1 else ''
return ''.join(canonical_path)
simplify_path("/home/") → "/home/", expected "/home"def simplify_path(path: str) -> str:
# Split the path into parts
parts = path.split('/')
# Process each part
i = 0
while i < len(parts):
part = parts[i]
if part == '.' or part == '..':
# Handle the '.' and '..' cases
if i == len(parts) - 1:
return '/' # Only return '/' if it's the root
else:
i += 1
else:
# Handle the other parts
i += 1
# Join the parts with '/' and return
return '/'.join(parts)
IndentationError: unexpected indent (<model>, line 1) def simplify_path(path):
stack = []
for token in path.split("/"):
if token == "..":
if stack:
stack.pop()
elif token != "." and token != "":
stack.append(token)
return "/" + "/".join(stack) IndentationError: unexpected indent (<model>, line 1) def simplify_path(path):
stack = []
for token in path.split("/"):
if token == "..":
if stack:
stack.pop()
elif token != "." and token != "":
stack.append(token)
return "/" + "/".join(stack) word_break — Segment a string into dictionary words · 4 examplesword_break("", ["a"]) → false, expected truedef word_break(s: str, words: list[str]) -> bool:
# Check if the words list is empty
if not words:
return True
# Check if the words list has at least one word
if len(words) == 0:
return False
# Check if the words list has at least one word
if len(words) == 1:
return s in words
# Try to break the string into words from the list
for word in words:
if s == word:
return True
if s.startswith(word):
return word_break(s[len(word):], words)
return False
word_break("leetcode", ["leet", "code"]) → false, expected truedef word_break(s: str, words: list[str]) -> bool:
# Check if the words list is empty
if not words:
return True
# Check if the words list has at least one word
if len(words) == 0:
return False
# Check if the words list has at least one word
if len(words) == 1:
return s in words
# Try to break the string into words from the list
for word in words:
if s == word:
return True
if s == word + s[1:]:
return True
return False
IndentationError: unexpected indent (<model>, line 1) def word_break(s, words):
n = len(s)
dp = [False] * (n + 1)
dp[0] = True
for i in range(1, n + 1):
for j in range(i):
if dp[j] and s[j:i] in words:
dp[i] = True
break
return dp[-1] IndentationError: unexpected indent (<model>, line 1) def word_break(s, words):
n = len(s)
dp = [False] * (n + 1)
dp[0] = True
for i in range(1, n + 1):
for j in range(i):
if dp[j] and s[j:i] in words:
dp[i] = True
break
return dp[-1] binary_search — Find a value's index in a sorted list · 4 examplesSyntaxError: invalid syntax (<model>, line 1)```python code
def binary_search(items: list[int], target: int) -> int:
left, right = 0, len(items) - 1
while left <= right:
mid = (left + right) // 2
if items[mid] == target:
return mid
elif items[mid] < target:
left = mid + 1
else:
right = mid - 1
return -1
```SyntaxError: invalid syntax (<model>, line 1)```python code
def binary_search(items: list[int], target: int) -> int:
left, right = 0, len(items) - 1
while left <= right:
mid = (left + right) // 2
if items[mid] == target:
return mid
elif items[mid] < target:
left = mid + 1
else:
right = mid - 1
return -1
```IndentationError: unexpected indent (<model>, line 1) def binary_search(items: list[int], target: int) -> int:
left = 0
right = len(items) - 1
while left <= right:
mid = (left + right) // 2
if items[mid] == target:
return mid
elif items[mid] < target:
left = mid + 1
else:
right = mid - 1
return -1 IndentationError: unexpected indent (<model>, line 1) def binary_search(items: list[int], target: int) -> int:
left = 0
right = len(items) - 1
while left <= right:
mid = (left + right) // 2
if items[mid] == target:
return mid
elif items[mid] < target:
left = mid + 1
else:
right = mid - 1
return -1 balanced — Bracket matching for () [] {} · 4 examplesis_balanced("({[]})") → false, expected truedef is_balanced(s: str) -> bool:
stack = []
i = 0
while i < len(s):
if s[i] in '([{':
stack.append(s[i])
i += 1
elif s[i] in ')]}':
if not stack:
return False
top = stack.pop()
if s[i] == ')' and top == '(':
continue
elif s[i] == ']' and top == '[':
continue
elif s[i] == '}' and top == '{':
continue
else:
return False
else:
i += 1
return True
is_balanced("({[]})") → false, expected truedef is_balanced(s: str) -> bool:
stack = []
i = 0
while i < len(s):
if s[i] in '([{':
stack.append(s[i])
i += 1
elif s[i] in ')]}':
if not stack:
return False
top = stack.pop()
if (s[i] == ')' and top == '(') or (s[i] == ']' and top == '[') or (s[i] == '}' and top == '{'):
continue
else:
return False
else:
i += 1
return True
IndentationError: unexpected indent (<model>, line 1) def is_balanced(s):
stack = []
brackets = {'(': ')', '[': ']', '{': '}'}
for char in s:
if char in brackets:
stack.append(char)
elif char in brackets.values():
if not stack or brackets[stack[-1]] != char:
return False
stack.pop()
return len(stack) == 0
The function iterates through the string character by character, pushing opening brackets onto a stack and checking for matching closing brackets when encountered. If a closing bracket is found without a corresponding opening bracket on top of the stack or if there are unmatched brackets left on the stack after iteration, it returns False; otherwise, it returns True. IndentationError: unexpected indent (<model>, line 1) def is_balanced(s):
stack = []
brackets = {")": "(", "]": "[", "}": "{"}
for char in s:
if char in brackets.values():
stack.append(char)
elif char in brackets:
if not stack or stack[-1] != brackets[char]:
return False
stack.pop()
return len(stack) == 0
run_unit_tests(tests_is_balanced) GPU
NVIDIA GB10 · 122 GB
CPU
Cortex-X925
Cores
20
OS
Ubuntu 24.04.4 LTS (aarch64)
Kernel
6.17.0-1018-nvidia
Proxy
0.2.0
CPU, OS and kernel read from the host at report time — this run predates their capture, and they do not change between runs.
Every task asks for one answer in the task's language — Python, JavaScript under node, C and C++ under gcc, Rust, C#, or PHP, each run in a separate process under a timeout with the return value compared against the expected one. HTML and CSS tasks are graded structurally: the answer is parsed and checked against required structure (bindings, attributes, declarations in the right context) — a claim about the markup, not about how a browser renders it. Tasks whose toolchain is absent on the machine are skipped and listed here, never scored as zero. Fully correct counts only responses where every case for that task passed; cases is the share of individual cases that passed, so a near-miss still scores there. A response with no extractable code block scores zero — that measures instruction-following, not coding.
Suite
coding-v1
Tasks
18
Cases
81
Repeats
2 per task
Languages
1
Core — 12 tasks, 49 cases
binary_search Find a value's index in a sorted listmerge_intervals Merge overlapping intervalsword_freq Top-n most common words with tie rulesroman Integer to Roman numeralbalanced Bracket matching for () [] {}flatten Flatten arbitrarily nested liststwo_sum Indices of the pair summing to a targetlru_cache_sim Simulate an LRU cache's get/put sequencegroup_anagrams Group words that are anagramsrun_length Run-length encode a stringcompare_versions Compare dotted version stringsspiral Matrix in clockwise spiral orderHard — 6 tasks, 32 cases
edit_distance Levenshtein distancelis_length Longest strictly increasing subsequencesimplify_path Canonicalise a Unix path with . and ..calculator Arithmetic with precedence, no evalword_break Segment a string into dictionary wordstopo_sort Topological sort of a dependency graphx-client-name: ai-proxy-bench.Time for the discarded warm-up request — the price of making the model resident, excluded from every measurement above.
| Configuration | Cold start |
|---|---|
| llama4 · 62.8 GB · ollama · cold | 191.7 s |
| gpt-oss:120b · 60.9 GB · ollama · cold | 100.7 s |
| llama3:70b-instruct · ollama · cached | 93.2 s |
| codellama:70b · ollama · cached | 84.4 s |
| gemma3:27b · ollama · cached | 45.3 s |
| qwen3.6:27b · ollama · cached | 43.0 s |
| qwen3-coder-next · 48.2 GB · ollama · cold | 35.1 s |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 262,144 | 29.9 s |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 131,072 | 29.9 s |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · cached · 32,768 | 29.8 s |
+ 26 more under 29.8 s.
Screen-only companion to the annotated chart above: hover a dot for its numbers, drag a box to zoom into the crowded band, double-click to reset. Hovering a model anywhere on this page highlights it everywhere. The printed report keeps the annotated version.