AI Proxy · benchmark
2026-08-03 12:55 · proxy 0.2.0
DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1 leads: 86% fully correct at 17.7 tok/s out, 31 s to load.
Ranked by correctness first, then output rate. 40 configurations measured across 3 backends.
Run this — 70/30 weighted
gemma4:26b
85% correct · 36.1 tok/s · 16.8 GB · answers in 7.1s · score 63
Runner-up
DeepSeek-V4-Flash-0731-UD-IQ2_XXS
86% correct · 17.7 tok/s · 84.6 GB · answers in 8.9s · score 62
Don't be fooled by
qwen3:0.6b
301 tok/s and only 5% correct
Every configuration, placed by correctness and output rate. A point below and to the left of another is beaten on both counts at once, so the dashed frontier is the shortlist — everything off it is dominated by something on it. Hover any point for its name.
Think
off
Cache
cached
Temp
0.0
TTFT is the first token of any kind; TTFC the first content token — the gap between them is time the model spent reasoning. Decode rate is measured from the first token onward, so reasoning tokens count as generated work. Best value in each column is highlighted.
| Configuration | Load | Resident | TTFT p50 | Decode p50 | Tokens | Total p50 | Fully correct | Cases | vs best | OK |
|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1 | 31 s | 0.4 GB | 143 ms | 17.7 | 190 | 8,906 ms | 86% | 85% | 17.8x | 87/87 |
| gemma4:26b · ollama · 4 | 9 s | 17.9 GB | 843 ms | 36.1 | 237 | 7,085 ms | 85% | 90% | 14.1x | 87/87 |
| gemma4:26b · 16.8 GB · ollama · 1 | 24 s | 17.9 GB | 496 ms | 63.5 | 247 | 3,920 ms | 83% | 88% | 7.8x | 87/87 |
| qwen3-coder-next · NVFP4 · vllm · 4 | 8 s | — | 385 ms | 41.4 | 208 | 4,388 ms | 82% | 85% | 8.7x | 87/87 |
| qwen3-coder-next · NVFP4 · vllm · 1 | 9 s | 0.1 GB | 160 ms | 62.1 | 211 | 2,943 ms | 79% | 84% | 5.9x | 87/87 |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4 | 29 s | — | 1,233 ms | 12.4 | 182 | 12,541 ms | 79% | 80% | 25.0x | 87/87 |
| ornith-nvfp4 · NVFP4 · vllm · 4 | 8 s | — | 326 ms | 48.6 | 236 | 3,887 ms | 78% | 83% | 7.7x | 87/87 |
| qwen3.6:35b-a3b · 22.3 GB · ollama · 1 | 31 s | 22.0 GB | 385 ms | 75.0 | 278 | 4,117 ms | 76% | 82% | 8.2x | 87/87 |
| qwen3.6:35b-a3b · ollama · 4 | 7 s | 22.0 GB | 4,781 ms | 75.0 | 278 | 7,460 ms | 76% | 82% | 14.9x | 87/87 |
| ornith-nvfp4 · NVFP4 · vllm · 1 | 8 s | 0.1 GB | 120 ms | 65.4 | 229 | 3,184 ms | 76% | 83% | 6.3x | 87/87 |
| qwen3-coder-next · ollama · 4 | 10 s | 48.7 GB | 5,599 ms | 51.3 | 219 | 8,839 ms | 76% | 81% | 17.6x | 87/87 |
| qwen3-coder-next · 48.2 GB · ollama · 1 | 73 s | 48.7 GB | 290 ms | 51.2 | 219 | 3,801 ms | 76% | 81% | 7.6x | 87/87 |
| qwen3-coder:tuned · ollama · 4 | 10 s | 50.6 GB | 5,716 ms | 50.8 | 215 | 8,603 ms | 76% | 82% | 17.2x | 87/87 |
| qwen3-coder:tuned · 48.2 GB · ollama · 1 | 66 s | 50.6 GB | 267 ms | 50.7 | 215 | 3,671 ms | 76% | 82% | 7.3x | 87/87 |
| gemma3:27b · 16.2 GB · ollama · 1 | 58 s | 17.9 GB | 643 ms | 11.5 | 238 | 18,462 ms | 76% | 76% | 36.8x | 87/87 |
| devstral-2:123b · 69.8 GB · ollama · 1 | 490 s | 92.9 GB | 661 ms | 2.7 | 189 | 57,047 ms | 76% | 85% | 113.7x | 87/87 |
| qwen3-coder:30b · ollama · 4 | 6 s | 24.0 GB | 440 ms | 41.0 | 190 | 4,290 ms | 75% | 87% | 8.6x | 87/87 |
| devstral-2:123b · 69.8 GB · ollama · 4 | 497 s | 92.9 GB | 2,655 ms | 2.4 | 189 | 64,218 ms | 75% | 85% | 128.0x | 87/87 |
| qwen3.6:27b · ollama · 4 | 43 s | 16.5 GB | 24,219 ms | 12.2 | 250 | 39,924 ms | 72% | 76% | 79.6x | 87/87 |
| qwen3.6:27b · 16.2 GB · ollama · 1 | 64 s | 16.5 GB | 491 ms | 12.2 | 250 | 16,566 ms | 72% | 76% | 33.0x | 87/87 |
| gemma3:27b · ollama · 4 | 45 s | 17.9 GB | 1,101 ms | 10.4 | 238 | 23,780 ms | 72% | 76% | 47.4x | 87/87 |
| qwen3-coder:30b · 17.3 GB · ollama · 1 | 21 s | 24.0 GB | 224 ms | 84.5 | 188 | 1,994 ms | 71% | 84% | 4.0x | 87/87 |
| devstral-small-2:24b · ollama · 4 | 38 s | 24.7 GB | 830 ms | 12.6 | 175 | 12,413 ms | 71% | 85% | 24.7x | 87/87 |
| devstral-small-2:24b · 14.1 GB · ollama · 1 | 49 s | 24.7 GB | 349 ms | 13.8 | 174 | 10,677 ms | 68% | 83% | 21.3x | 87/87 |
| llama4 · 62.8 GB · ollama · 1 | 194 s | 64.0 GB | 541 ms | 18.0 | 152 | 8,701 ms | 62% | 75% | 17.3x | 87/87 |
| llama4 · ollama · 4 | 30 s | 64.0 GB | 3,133 ms | 11.0 | 149 | 14,611 ms | 62% | 76% | 29.1x | 87/87 |
| gemma4 · 8.9 GB · ollama · 1 | 21 s | 4.0 GB | 471 ms | 56.1 | 367 | 8,217 ms | 54% | 54% | 16.4x | 87/87 |
| llama3:70b-instruct · 37.2 GB · ollama · 1 | 145 s | 42.4 GB | 433 ms | 5.5 | 130 | 21,611 ms | 52% | 72% | 43.1x | 87/87 |
| llama3:70b-instruct · ollama · 4 | 94 s | 42.4 GB | 1,124 ms | 5.3 | 127 | 22,715 ms | 51% | 72% | 45.3x | 87/87 |
| gemma4 · ollama · 4 | 10 s | 4.0 GB | 753 ms | 42.6 | 369 | 11,462 ms | 48% | 48% | 22.8x | 87/87 |
| gpt-oss:120b · 60.9 GB · ollama · 1 | 146 s | 60.4 GB | 508 ms | 37.9 | 435 | 13,924 ms | 43% | 37% | 27.8x | 87/87 |
| gpt-oss:120b · ollama · 4 | 14 s | 60.4 GB | 1,279 ms | 20.6 | 432 | 23,443 ms | 43% | 38% | 46.7x | 87/87 |
| codellama:70b · 36.2 GB · ollama · 1 | 107 s | 37.8 GB | 279 ms | 5.7 | 293 | 45,484 ms | 26% | 42% | 90.7x | 87/87 |
| codellama:70b · ollama · 4 | 64 s | 37.8 GB | 960 ms | 5.5 | 291 | 47,884 ms | 23% | 41% | 95.5x | 87/87 |
| minicpm-v4.5 · 5.7 GB · ollama · 1 | 22 s | 14.7 GB | 240 ms | 40.5 | 180 | 3,362 ms | 22% | 51% | 6.7x | 87/87 |
| minicpm-v4.5 · ollama · 4 | 13 s | 14.7 GB | 412 ms | 35.1 | 179 | 4,017 ms | 20% | 48% | 8.0x | 87/87 |
| qwen3:4b · 2.3 GB · ollama · 1 | 12 s | 12.6 GB | 229 ms | 73.1 | 503 | 7,220 ms | 7% | 8% | 14.4x | 87/87 |
| qwen3:0.6b · 0.5 GB · ollama · 1 | 6 s | 8.5 GB | 203 ms | 301.1 | 122 | 502 ms | 5% | 38% | 1.0x | 87/87 |
| qwen3:0.6b · ollama · 4 | 2 s | 8.5 GB | 323 ms | 190.5 | 124 | 732 ms | 5% | 38% | 1.5x | 87/87 |
| qwen3:4b · ollama · 4 | 7 s | 12.6 GB | 446 ms | 61.8 | 505 | 8,767 ms | 5% | 8% | 17.5x | 87/87 |
Correctness alone is not a ranking — one point of correctness is not worth half the speed. Each model's best cell scores 70% × correctness + 30% × relative speed, where relative speed is decode rate against the fastest model in this report (qwen3:0.6b, 301.1 tok/s = 1.0). The bar is the weighting made visible: correctness · speed. Drag to change what you value; the ranking recomputes.
| # | Model | Correct | tok/s | Score |
|---|---|---|---|---|
| 1 | gemma4:26b | 85% | 36.1 | 63 |
| 2 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS | 86% | 17.7 | 62 |
| 3 | qwen3-coder-next | 82% | 41.4 | 61 |
| 4 | qwen3.6:35b-a3b | 76% | 75.0 | 61 |
| 5 | ornith-nvfp4 | 78% | 48.6 | 60 |
| 6 | qwen3-coder:tuned | 76% | 50.8 | 58 |
| 7 | qwen3-coder:30b | 75% | 41.0 | 56 |
| 8 | gemma3:27b | 76% | 11.5 | 54 |
| 9 | devstral-2:123b | 76% | 2.7 | 53 |
| 10 | qwen3.6:27b | 72% | 12.2 | 52 |
| 11 | devstral-small-2:24b | 71% | 12.6 | 51 |
| 12 | llama4 | 62% | 18.0 | 45 |
| 13 | gemma4 | 54% | 56.1 | 43 |
| 14 | llama3:70b-instruct | 52% | 5.5 | 37 |
| 15 | gpt-oss:120b | 43% | 37.9 | 34 |
| 16 | qwen3:0.6b | 5% | 301.1 | 33 |
| 17 | minicpm-v4.5 | 22% | 40.5 | 19 |
| 18 | codellama:70b | 26% | 5.7 | 19 |
| 19 | qwen3:4b | 7% | 73.1 | 12 |
Position is footprint against correctness; bubble area is output speed. Models whose size cannot be read — vLLM checkpoints live inside their containers — are absent, not zero.
The only controlled engine comparison a run can contain: identical weights reachable through more than one backend, paired by cache state. When output rates tie, the wait for the first token is what an engine buys.
The warm-up sends the same prompt the measured runs use, so its first-token time is that prompt’s cold prefill; everything after it is served warm. A backend whose prefix caching is off or unsupported shows no gap between these two columns.
| Configuration | Prompt | Cold TTFT | Cached TTFT | Faster by |
|---|---|---|---|---|
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1 | 0 | 1,253 ms | 143 ms | 9x |
| gemma4:26b · 16.8 GB · ollama · 1 | 0 | 16,138 ms | 496 ms | 33x |
| qwen3-coder-next · NVFP4 · vllm · 1 | 0 | 437 ms | 160 ms | 3x |
| qwen3.6:35b-a3b · 22.3 GB · ollama · 1 | 0 | 24,671 ms | 385 ms | 64x |
| ornith-nvfp4 · NVFP4 · vllm · 1 | 0 | 223 ms | 120 ms | 2x |
| qwen3-coder-next · 48.2 GB · ollama · 1 | 0 | 62,884 ms | 290 ms | 217x |
| qwen3-coder:tuned · 48.2 GB · ollama · 1 | 0 | 56,133 ms | 267 ms | 210x |
| gemma3:27b · 16.2 GB · ollama · 1 | 0 | 13,444 ms | 643 ms | 21x |
| devstral-2:123b · 69.8 GB · ollama · 1 | 0 | 295,064 ms | 661 ms | 446x |
| devstral-2:123b · 69.8 GB · ollama · 4 | 0 | 300,909 ms | 2,655 ms | 113x |
| qwen3.6:27b · 16.2 GB · ollama · 1 | 0 | 22,023 ms | 491 ms | 45x |
| qwen3-coder:30b · 17.3 GB · ollama · 1 | 0 | 14,950 ms | 224 ms | 67x |
| devstral-small-2:24b · 14.1 GB · ollama · 1 | 0 | 11,850 ms | 349 ms | 34x |
| llama4 · 62.8 GB · ollama · 1 | 0 | 164,683 ms | 541 ms | 304x |
| gemma4 · 8.9 GB · ollama · 1 | 0 | 11,385 ms | 471 ms | 24x |
| llama3:70b-instruct · 37.2 GB · ollama · 1 | 0 | 51,159 ms | 433 ms | 118x |
| gpt-oss:120b · 60.9 GB · ollama · 1 | 0 | 131,971 ms | 508 ms | 260x |
| codellama:70b · 36.2 GB · ollama · 1 | 0 | 23,614 ms | 279 ms | 85x |
| minicpm-v4.5 · 5.7 GB · ollama · 1 | 0 | 9,129 ms | 240 ms | 38x |
| qwen3:4b · 2.3 GB · ollama · 1 | 0 | 5,013 ms | 229 ms | 22x |
| qwen3:0.6b · 0.5 GB · ollama · 1 | 0 | 4,179 ms | 203 ms | 21x |
The core tier confirms a model is not broken; it saturates for anything capable, which is exactly why the hard tier exists. Compare two models on the hard row when both score 100% on core.
| Configuration | Core | Hard |
|---|---|---|
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1 | 87% | 86% |
| gemma4:26b · ollama · 4 | 78% | 93% |
| gemma4:26b · 16.8 GB · ollama · 1 | 76% | 90% |
| qwen3-coder-next · NVFP4 · vllm · 4 | 84% | 79% |
| qwen3-coder-next · NVFP4 · vllm · 1 | 80% | 79% |
| DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4 | 80% | 79% |
| ornith-nvfp4 · NVFP4 · vllm · 4 | 76% | 81% |
| qwen3.6:35b-a3b · 22.3 GB · ollama · 1 | 80% | 71% |
| qwen3.6:35b-a3b · ollama · 4 | 80% | 71% |
| ornith-nvfp4 · NVFP4 · vllm · 1 | 69% | 83% |
| qwen3-coder-next · ollama · 4 | 73% | 79% |
| qwen3-coder-next · 48.2 GB · ollama · 1 | 73% | 79% |
| qwen3-coder:tuned · ollama · 4 | 73% | 79% |
| qwen3-coder:tuned · 48.2 GB · ollama · 1 | 73% | 79% |
| gemma3:27b · 16.2 GB · ollama · 1 | 80% | 71% |
| devstral-2:123b · 69.8 GB · ollama · 1 | 80% | 71% |
| qwen3-coder:30b · ollama · 4 | 69% | 81% |
| devstral-2:123b · 69.8 GB · ollama · 4 | 80% | 69% |
| qwen3.6:27b · ollama · 4 | 67% | 79% |
| qwen3.6:27b · 16.2 GB · ollama · 1 | 67% | 79% |
| gemma3:27b · ollama · 4 | 73% | 71% |
| qwen3-coder:30b · 17.3 GB · ollama · 1 | 64% | 79% |
| devstral-small-2:24b · ollama · 4 | 73% | 69% |
| devstral-small-2:24b · 14.1 GB · ollama · 1 | 71% | 64% |
| llama4 · 62.8 GB · ollama · 1 | 60% | 64% |
| llama4 · ollama · 4 | 64% | 60% |
| gemma4 · 8.9 GB · ollama · 1 | 60% | 48% |
| llama3:70b-instruct · 37.2 GB · ollama · 1 | 47% | 57% |
| llama3:70b-instruct · ollama · 4 | 49% | 52% |
| gemma4 · ollama · 4 | 53% | 43% |
| gpt-oss:120b · 60.9 GB · ollama · 1 | 51% | 33% |
| gpt-oss:120b · ollama · 4 | 49% | 36% |
| codellama:70b · 36.2 GB · ollama · 1 | 29% | 24% |
| codellama:70b · ollama · 4 | 27% | 19% |
| minicpm-v4.5 · 5.7 GB · ollama · 1 | 20% | 24% |
| minicpm-v4.5 · ollama · 4 | 18% | 21% |
| qwen3:4b · 2.3 GB · ollama · 1 | 7% | 7% |
| qwen3:0.6b · 0.5 GB · ollama · 1 | 0% | 10% |
| qwen3:0.6b · ollama · 4 | 2% | 7% |
| qwen3:4b · ollama · 4 | 7% | 2% |
Share of responses that passed every case for that task. A model strong everywhere except one task and a model mediocre throughout can share an overall average.
| Task | Perfect in | Configurations that missed it |
|---|---|---|
balanced_depthMax bracket nesting depth, -1 if unbalanced | 36 of 40 | qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
base_convertInteger between bases 2-36 with validation | 17 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
card_gridResponsive auto-fill card grid | 38 of 40 | qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4 |
clamp_addSaturating int addition — overflow trap | 30 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
clamp_mulSaturating int multiplication | 29 of 40 | llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3-coder-next · NVFP4 · vllm · 1, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
count_wordsCount words split on spaces and tabs | 33 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
csv_escapeRFC 4180 CSV field quoting | 31 of 40 | gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
csv_lineSplit one CSV record honouring quotes | 6 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
data_tableRevenue table with caption and scoped headers | 36 of 40 | minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
glob_matchGlob matching with ? and * | 32 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
group_rangesCollapse consecutive integers into range strings | 24 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
interval_intersectIntersect two interval lists | 31 of 40 | gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
json_pointerResolve an RFC 6901 JSON Pointer | 27 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
justifyFull text justification | 21 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · NVFP4 · vllm · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
login_formLogin form with labels bound to their inputs | 8 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
lru_opsLRU cache with eviction order | 26 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma3:27b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
mid_floorFloor midpoint of two i64s — overflow and negatives | 3 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
ordinalEnglish ordinal suffix — the 11th/12th/13th trap | 11 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3-coder-next · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
parse_queryParse a URL query string into an object | 22 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
path_normNormalise a POSIX path with . and .. | 27 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
pluckColumn from associative rows — null vs missing key | 25 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
roman_strictRoman to int, rejecting non-canonical forms | 2 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
round_toRound to nearest multiple, halves away from zero | 17 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
semver_cmpSemantic versions incl. pre-release precedence | 0 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · llamacpp · 1, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 1, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma3:27b · 16.2 GB · ollama · 1, gemma3:27b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:27b · 16.2 GB · ollama · 1, qwen3.6:27b · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
slugifyURL slug: symbol runs become one hyphen | 26 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
snake_to_camelsnake_case to camelCase — digits stop capitalisation | 11 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · 62.8 GB · ollama · 1, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, ornith-nvfp4 · NVFP4 · vllm · 1, ornith-nvfp4 · NVFP4 · vllm · 4, qwen3-coder-next · 48.2 GB · ollama · 1, qwen3-coder-next · ollama · 4, qwen3-coder:30b · 17.3 GB · ollama · 1, qwen3-coder:30b · ollama · 4, qwen3-coder:tuned · 48.2 GB · ollama · 1, qwen3-coder:tuned · ollama · 4, qwen3.6:35b-a3b · 22.3 GB · ollama · 1, qwen3.6:35b-a3b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
theme_varsDark-mode token inside a media query | 35 of 40 | codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · ollama · 4 |
tokenize_exprTokenise arithmetic, None on invalid input | 20 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, gemma4 · 8.9 GB · ollama · 1, gemma4 · ollama · 4, gemma4:26b · 16.8 GB · ollama · 1, gemma4:26b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, llama3:70b-instruct · 37.2 GB · ollama · 1, llama3:70b-instruct · ollama · 4, llama4 · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3-coder-next · NVFP4 · vllm · 1, qwen3-coder-next · NVFP4 · vllm · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
topo_lexSmallest topological order, None on cycle | 25 of 40 | DeepSeek-V4-Flash-0731-UD-IQ2_XXS · IQ2_XXS · 84.6 GB · llamacpp · 4, codellama:70b · 36.2 GB · ollama · 1, codellama:70b · ollama · 4, devstral-2:123b · 69.8 GB · ollama · 4, devstral-small-2:24b · 14.1 GB · ollama · 1, devstral-small-2:24b · ollama · 4, gpt-oss:120b · 60.9 GB · ollama · 1, gpt-oss:120b · ollama · 4, minicpm-v4.5 · 5.7 GB · ollama · 1, minicpm-v4.5 · ollama · 4, qwen3-coder-next · NVFP4 · vllm · 4, qwen3:0.6b · 0.5 GB · ollama · 1, qwen3:0.6b · ollama · 4, qwen3:4b · 2.3 GB · ollama · 1, qwen3:4b · ollama · 4 |
The first failing case per configuration: the call that was made, what came back, and what should have — or the compile error or timeout that stopped it. This is what a percentage point of correctness is made of.
semver_cmp — Semantic versions incl. pre-release precedence · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def semver_cmp(a: str, b: str) -> int:
# Split off build metadata
a_main, _, a_meta = a.partition('+')
b_main, _, b_meta = b.partition('+')
# Split main version into core and pre-release
a_core, _, a_pre = a_main.partition('-')
b_core, _, b_pre = b_main.partition('-')
# Compare core version numbers
a_parts = list(map(int, a_core.split('.')))
b_parts = list(map(int, b_core.split('.')))
# Pad with zeros to ensure 3 parts
while len(a_parts) < 3:
a_parts.append(0)
while len(b_parts) < 3:
b_parts.append(0)
for i in range(3):
if a_parts[i] < b_parts[i]:
return -1
elif a_parts[i] > b_parts[i]:
return 1
# Core versions equal, compare pre-release
if a_pre == '' and b_pre == '':
return 0
elif a_pre == '':
return 1 # no pre-release > with pre-release
elif b_pre == '':
return -1
# Both have pre-release, compare identifiers
a_ids = a_pre.split('.')
b_ids = b_pre.split('.')
for i in range(min(len(a_ids), len(b_ids))):
a_id = a_ids[i]
b_id = b_ids[i]
# Check if both are numeric
a_numeric = a_id.isdigit()
b_numeric = b_id.isdigit()
if a_numeric and b_numeric:
a_num = int(a_id)
b_num = int(b_id)
if a_num < b_num:
return -1
elif a_num > b_num:
return 1
elif a_numeric and not b_numeric:
return -1 # numeric < non-numeric
elif
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import re
def semver_cmp(a: str, b: str) -> int:
def parse_version(v: str):
# Remove build metadata (everything after '+')
base = v.split('+', 1)[0]
# Split version into core and pre-release parts
if '-' in base:
core_part, pre_part = base.split('-', 1)
pre_release = pre_part.split('.')
else:
core_part = base
pre_release = None
# Parse major, minor, patch
core = [int(x) for x in core_part.split('.')]
# Process pre-release identifiers
processed_pre = []
if pre_release is not None:
for part in pre_release:
if part.isdigit():
processed_pre.append((0, int(part))) # Numeric type 0
else:
processed_pre.append((1, part)) # String type 1
return core, processed_pre
core_a, pre_a = parse_version(a)
core_b, pre_b = parse_version(b)
# 1. Compare Major, Minor, Patch
for i in range(3):
if core_a[i] > core_b[i]: return 1
if core_a[i] < core_b[i]: return -1
# 2. Handle Pre-release precedence rule:
# A version WITH a pre-release has LOWER precedence than one WITHOUT.
if pre_a and not pre_b:
return -1
if not pre_a and pre_b:
return 1
if not pre_a and not pre_b:
return 0
# 3. Compare Pre-release identifiers
for part_a, part_b in zip(pre_a, pre_b):
# part is (type_priority, value)
# type_priority 0 for numeric, 1 for string ens
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import re
def semver_cmp(a: str, b: str) -> int:
def parse_version(v: str):
# Remove build metadata (everything after '+')
base = v.split('+', 1)[0]
# Split version into core and pre-release
if '-' in base:
core_part, pre_part = base.split('-', 1)
pre_release = pre_part.split('.')
else:
core_part = base
pre_release = None
# Parse major, minor, patch
major, minor, patch = map(int, core_part.split('.'))
# Process pre-release identifiers
processed_pre = []
if pre_release is not None:
for part in pre_release:
if part.isdigit():
processed_pre.append((0, int(part))) # Numeric type 0
else:
processed_pre.append((1, part)) # String type 1
return (major, minor, patch), processed_pre
core_a, pre_a = parse_version(a)
core_b, pre_b = parse_version(b)
# 1. Compare Major, Minor, Patch
if core_a < core_b:
return -1
if core_a > core_b:
return 1
# 2. Handle Pre-release precedence rules
# Rule: Version WITH pre-release has LOWER precedence than version WITHOUT
if pre_a and not pre_b:
return -1
if not pre_a and pre_b:
return 1
if not pre_a and not pre_b:
return 0
# 3. Compare Pre-release identifiers
for part_a, part_b in zip(pre_a, pre_b):
# part is (type_priority, value)
# type_priority 0 (numeric) < 1 (string)
if
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def semver_cmp(a: str, b: str) -> int:
def parse_version(version_str):
# Remove build metadata (everything after '+')
version_str = version_str.split('+')[0]
# Split into core version and pre-release
if '-' in version_str:
core, prerelease = version_str.split('-', 1)
else:
core = version_str
prerelease = None
# Parse core version (major.minor.patch)
parts = core.split('.')
major = int(parts[0])
minor = int(parts[1]) if len(parts) > 1 else 0
patch = int(parts[2]) if len(parts) > 2 else 0
# Parse pre-release identifiers if present
prerelease_ids = None
if prerelease is not None:
prerelease_ids = prerelease.split('.')
return (major, minor, patch, prerelease_ids)
def compare_prerelease_ids(ids_a, ids_b):
if ids_a is None and ids_b is None:
return 0
if ids_a is None: # No prerelease > has prerelease
return 1
if ids_b is None:
return -1
# Compare identifier by identifier
for i in range(min(len(ids_a), len(ids_b))):
id_a = ids_a[i]
id_b = ids_b[i]
# Check if both are numeric
a_is_num = id_a.isdigit()
b_is_num = id_b.isdigit()
if a_is_num and b_is_num:
# Both numeric: compare numerically
num_a = int(id_a)
num_b = int(id_b)
if num_a <
… (truncated)roman_strict — Roman to int, rejecting non-canonical forms · 4 examplesSyntaxError: unterminated string literal (detected at line 1) (<model>, line 1)def roman_strict(s: str) -> int | None:
# Mapping of Roman numeral symbols to their integer values
values = {
'I': 1, 'V': 5, 'X': 10, 'L': 50,
'C': 100, 'D': 500, 'M': 1000
}
# Define subtractive pairs and their values
subtractive_pairs = {
'IV': 4, 'IX': 9, 'XL': 40, 'XC': 90,
'CD': 400, 'CM': 900
}
# Check for invalid characters
for char in s:
if char not in values:
return None
# Check for invalid repetitions
# I, X, C can repeat at most 3 times
for char in 'IXC':
count = s.count(char)
if count > 3:
return None
# V, L, D can never repeat
for char in 'VLD':
if s.count(char) > 1:
return None
# Check for subtractive pairs appearing more than once
for pair in subtractive_pairs:
if s.count(pair) > 1:
return None
# Check for invalid subtractive combinations
# Only allowed subtractive pairs are IV, IX, XL, XC, CD, CM
# Any other subtractive pattern is invalid
for i in range(len(s) - 1):
if values[s[i]] < values[s[i+1]]:
pair = s[i:i+2]
if pair not in subtractive_pairs:
return None
# Check for non-increasing order (except for allowed subtractive pairs)
# Process the string to check order
i = 0
while i < len(s):
# Check if current pair is a valid subtractive pair
if i + 1 < len(s) and s[i:i+2] in subtractive_pairs:
i += 2
else:
# Check if curr
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def roman_strict(s: str) -> int | None:
# Define the value mapping
values = {'I': 1, 'V': 5, 'X': 10, 'L': 50, 'C': 100, 'D': 500, 'M': 1000}
# Check for invalid characters
if not all(c in values for c in s):
return None
# Check for invalid repetitions: V, L, D never repeat
for char in ['V', 'L', 'D']:
if s.count(char) > 1:
return None
# Check for invalid repetitions: I, X, C repeat at most three times
for char in ['I', 'X', 'C']:
if s.count(char) > 3:
return None
# Check for invalid subtractive pairs: only IV, IX, XL, XC, CD, CM allowed
valid_subtractive = {'IV', 'IX', 'XL', 'XC', 'CD', 'CM'}
for i in range(len(s) - 1):
if values[s[i]] < values[s[i+1]]:
pair = s[i:i+2]
if pair not in valid_subtractive:
return None
# Check that each subtractive pair appears at most once
for pair in valid_subtractive:
if s.count(pair) > 1:
return None
# Check non-increasing order (except for valid subtractive pairs)
i = 0
while i < len(s):
if i < len(s) - 1 and values[s[i]] < values[s[i+1]]:
# This should be a subtractive pair
pair = s[i:i+2]
if pair not in valid_subtractive:
return None
i += 2
else:
# Check non-increasing order
if i > 0 and values[s[i]] > values[s[i-1]]:
return None
i += 1
# Convert to integer
total = 0
i = 0
while i
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def roman_strict(s: str) -> int | None:
# Define Roman numeral values
values = {'I': 1, 'V': 5, 'X': 10, 'L': 50, 'C': 100, 'D': 500, 'M': 1000}
# Check for invalid characters
if not all(c in values for c in s):
return None
# Check for invalid repetitions
# V, L, D never repeat
for c in ['V', 'L', 'D']:
if s.count(c) > 1:
return None
# I, X, C repeat at most three times
for c in ['I', 'X', 'C']:
if s.count(c) > 3:
return None
# Check for invalid subtractive pairs
# Only IV, IX, XL, XC, CD, CM are allowed
invalid_pairs = ['IC', 'IL', 'ID', 'IM', 'VX', 'VL', 'VD', 'VM',
'XD', 'XM', 'LC', 'LD', 'LM', 'DM']
for pair in invalid_pairs:
if pair in s:
return None
# Check that subtractive pairs appear at most once
for pair in ['IV', 'IX', 'XL', 'XC', 'CD', 'CM']:
if s.count(pair) > 1:
return None
# Check non-increasing order (with exceptions for valid subtractive pairs)
i = 0
while i < len(s):
# Get current value
curr_val = values[s[i]]
# Check if this is the start of a subtractive pair
if i + 1 < len(s):
next_val = values[s[i + 1]]
if curr_val < next_val:
# Must be a valid subtractive pair
pair = s[i:i+2]
if pair not in ['IV', 'IX', 'XL', 'XC', 'CD', 'CM']:
return None
# Skip the pair
i += 2
conti
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def roman_strict(s: str) -> int | None:
# Define the values of Roman numeral symbols
values = {
'I': 1, 'V': 5, 'X': 10, 'L': 50,
'C': 100, 'D': 500, 'M': 1000
}
# Define the allowed subtractive pairs
subtractive_pairs = {'IV', 'IX', 'XL', 'XC', 'CD', 'CM'}
# Check for invalid characters
for ch in s:
if ch not in values:
return None
# Check for repeated V, L, D (never repeat)
for ch in 'VLD':
if s.count(ch) > 1:
return None
# Check for I, X, C repeat at most three times
for ch in 'IXC':
if s.count(ch) > 3:
return None
# Check for subtractive pairs appearing more than once
for pair in subtractive_pairs:
if s.count(pair) > 1:
return None
# Check for invalid subtractive combinations (like IC, IM, etc.)
# Only allowed subtractive pairs are the ones listed
# Also check that subtractive pairs are used correctly (i.e., not like IXI)
# We'll parse the string to check canonical order and subtractive usage
i = 0
n = len(s)
total = 0
prev_value = float('inf') # start with a large value
while i < n:
# Check for subtractive pair
if i + 1 < n and s[i:i+2] in subtractive_pairs:
pair = s[i:i+2]
# Check that this pair is allowed and not repeated (already checked count)
# Check that the pair is in correct position (non-increasing order)
# The pair value is the larger minus smaller
val = values[pa
… (truncated)mid_floor — Floor midpoint of two i64s — overflow and negatives · 4 examplescompile error: ns an unclosed delimiter
--> /tmp/bench_rust_li15gkm9/task.rs:42:3
|
3 | fn mid_floor(a: i64, b: i64) -> i64 {
| - unclofn mid_floor(a: i64, b: i64) -> i64 {
// To avoid overflow of (a + b), we use the property that
// floor((a + b) / 2) can be calculated using a and b individually.
// If both have the same sign, or one is zero, it's simple division.
// The tricky part is when signs differ, but even then:
// (a + b) / 2 == (a/2) + (b/2) + remainder_adjustment.
// A robust way to handle overflow and floor rounding:
// If a + b is even, result is (a + b) / 2.
// If a + b is odd, result is (a + b - 1) / 2 if we want floor.
// To avoid overflow, we use the identity:
// mid_floor(a, b) = a/2 + b/2 + adjustment
// where adjustment depends on the remainders of a and b.
let (q1, r1) = a.div_rem(2);
let (q2, r2) = b.div_rem(2);
// The sum of quotients is part of the answer.
// We just need to handle the remainders (-1, 0, or 1).
// Since we want floor((a+b)/2), we can use:
// (a + b) / 2 = a/2 + b/2 + (rem_a + rem_b)/2 (using floor division logic)
// However, the simplest overflow-free way to implement floor(a+b/2):
// If we use i128, we bypass all i64 overflow and truncation issues.
let sum = a as i128 + b as i128;
((sum / 2).floor() as i64) // Note: integer division of i128 truncates toward zero.
// To ensure floor (rounding towards negative infinity):
if sum >= 0 {
(sum / 2) as i64
} else {
// If sum is negative and oddcompile error: bench_rust_zmv43rmb/task.rs:36:3
|
3 | fn mid_floor(a: i64, b: i64) -> i64 {
| - unclosed delimiter
...
13 | } else iffn mid_floor(a: i64, b: i64) -> i64 {
// To avoid overflow of (a + b), we use the property that
// floor((a + b) / 2) can be calculated using a/2 + b/2 and handling the remainder.
// However, a simpler way to handle both overflow and the floor logic is:
// If a and b have the same sign, (a + b) / 2 is safe if we use a/2 + b/2 + adjustment.
// But even better: Use the property that mid_floor(a, b) = floor(a/2 + b/2 + (a%2 + b%2)/2).
// To avoid all overflow and handle negative rounding:
if a >= 0 && b >= 0 {
// Standard case for positive numbers
a / 2 + b / 2 + (a % 2 + b % 2) / 2
} else if a < 0 && b < 0 {
// Standard case for negative numbers
// In Rust, -3 / 2 = -1. We need floor(-1.5) = -2.
let res = a / 2 + b / 2;
let rem_sum = (a % 2).abs() + (b % 2).abs();
// If there is a remainder and the sum is negative, we need to adjust for floor
if (a % 2 != 0 || b % 2 != 0) && (a + b < 0) {
// This part is tricky with overflow. Let's use a more robust approach:
// mid_floor(a, b) = floor((a+b)/2).
// If we can't add them, we use the identity:
// floor((a+b)/2) = a/2 + b/2 + (1 if a%2 != 0 and b%2 != 0 and signs are same else 0) ... no.
}
// Let's use the most robust method:
// mid_floor(a, b) is simply (a/2 + b/2) plus an adjustment if both are odd.
mid_floor(-3, 0) → "-1", expected -2fn mid_floor(a: i64, b: i64) -> i64 {
if a <= 0 && b <= 0 {
(a.wrapping_add(b)) / 2
} else if a >= 0 && b >= 0 {
a + (b - a) / 2
} else {
// Handle mixed signs: use the formula (a + b - 1) / 2 for negative results
// to adjust for truncation toward zero
let sum = a.wrapping_add(b);
if sum >= 0 {
sum / 2
} else {
(sum - 1) / 2
}
}
}
compile error: 28`
|
help: you can convert an `i128` to an `i64` and panic if the converted value doesn't fit
|
13 | (sum / 2).try_into().unwrap()
| + fn mid_floor(a: i64, b: i64) -> i64 {
// Compute floor((a + b) / 2) without overflowing.
// We can use i128 to safely compute the sum.
let sum = a as i128 + b as i128;
// Now we need floor(sum / 2) where sum is an i128.
// In Rust, integer division of i128 truncates toward zero.
// For negative odd numbers, truncation toward zero gives -1, but floor gives -2.
// So if sum is negative and odd, we need to subtract 1 before dividing by 2.
if sum < 0 && sum % 2 != 0 {
(sum - 1) / 2
} else {
sum / 2
}
}
csv_line — Split one CSV record honouring quotes · 4 examplescsv_line("a,\"b,c\",d") → ["a", "bc,"], expected ["a", "b,c", "d"]def csv_line(line: str) -> list[str]:
if not line:
return ['']
result = []
current_field = []
in_quotes = False
i = 0
n = len(line)
while i < n:
char = line[i]
if not in_quotes:
if char == '"' and (not current_field or (len(current_field) == 1 and current_field[0] == '')):
# Only a quote at the very start of a field triggers quoted mode
in_quotes = True
i += 1
continue
elif char == ',':
result.append("".join(current_field))
current_field = []
i += 1
continue
else:
current_field.append(char)
else:
# Inside quotes
if i + 1 < n and char == '"' and line[i+1] == '"':
# Handle escaped double quotes ("")
current_field.append('"')
i += 2
continue
elif char == '"':
# Closing quote
in_quotes = False
i += 1
continue
else:
current_field.append(char)
i += 1
i += 1
# Append the last field processed
result.append("".join(current_field))
return result
csv_line("a,b\"c,d") → ["a", "bc,d"], expected ["a", "b\"c", "d"]def csv_line(line: str) -> list[str]:
if not line:
return ['']
fields = []
current_field = []
in_quotes = False
i = 0
while i < len(line):
char = line[i]
if not in_quotes:
if char == '"':
# Start of a quoted field
in_quotes = True
elif char == ',':
# End of a field
fields.append(''.join(current_field))
current_field = []
else:
# Regular character
current_field.append(char)
else:
# Inside quotes
if char == '"':
# Check if this is a doubled quote
if i + 1 < len(line) and line[i + 1] == '"':
# This is an escaped quote, add one quote
current_field.append('"')
i += 1 # Skip the next quote
else:
# End of quoted section
in_quotes = False
else:
current_field.append(char)
i += 1
# Add the last field
fields.append(''.join(current_field))
return fields
csv_line("a,b\"c,d") → ["a", "bc,d"], expected ["a", "b\"c", "d"]def csv_line(line: str) -> list[str]:
if not line:
return ['']
fields = []
current_field = []
in_quotes = False
i = 0
while i < len(line):
char = line[i]
if not in_quotes:
if char == '"':
# Start of a quoted field - only meaningful at the beginning of a field
in_quotes = True
i += 1
elif char == ',':
# End of a field
fields.append(''.join(current_field))
current_field = []
i += 1
else:
# Regular character
current_field.append(char)
i += 1
else: # in_quotes is True
if char == '"':
# Check if this is a doubled quote
if i + 1 < len(line) and line[i + 1] == '"':
# This is an escaped quote - add one quote and skip the next one
current_field.append('"')
i += 2
else:
# End of quoted section
in_quotes = False
i += 1
else:
# Regular character inside quotes
current_field.append(char)
i += 1
# Add the last field
fields.append(''.join(current_field))
return fields
csv_line("a,\"b,c\",d") → ["a", "b,c"], expected ["a", "b,c", "d"]def csv_line(line: str) -> list[str]:
"""Split a CSV line into fields, handling quoted fields."""
if line == '':
return ['']
fields = []
i = 0
n = len(line)
while i <= n:
if i == n:
# End of line, add empty field if needed
if not fields:
fields.append('')
break
if line[i] == '"':
# Start of a quoted field
i += 1
field = ''
while i < n:
if line[i] == '"':
# Check for doubled quote
if i + 1 < n and line[i + 1] == '"':
field += '"'
i += 2
else:
# End of quoted field
i += 1
# After closing quote, expect comma or end of line
if i < n and line[i] == ',':
fields.append(field)
field = ''
i += 1
continue
else:
fields.append(field)
break
else:
field += line[i]
i += 1
continue
else:
# Unquoted field
field = ''
while i < n and line[i] != ',':
field += line[i]
i += 1
if i < n:
fields.append(field)
i += 1 # skip comma
else
… (truncated)login_form — Login form with labels bound to their inputs · 4 examplesmethod of "form" → "POST", expected "post"<form method="POST" action="/login"> <label for="email">Email</label> <input type="email" id="email" name="email" required> <label for="password">Password</label> <input type="password" id="password" name="password" required> <button type="submit">Sign in</button> </form>
method of "form" → "POST", expected "post"<form action="/login" method="POST">
<div>
<label for="email">Email</label>
<input type="email" id="email" name="email" required>
</div>
<div>
<label for="password">Password</label>
<input type="password" id="password" name="password" required>
</div>
<button type="submit">Sign in</button>
</form>
method of "form" → "POST", expected "post"<form action="/login" method="POST">
<div>
<label for="email">Email</label>
<input type="email" id="email" name="email" required>
</div>
<div>
<label for="password">Password</label>
<input type="password" id="password" name="password" required>
</div>
<button type="submit">Sign in</button>
</form>
method of "form" → "POST", expected "post"<form action="/login" method="POST">
<div>
<label for="email">Email</label>
<input type="email" id="email" name="email" required>
</div>
<div>
<label for="password">Password</label>
<input type="password" id="password" name="password" required>
</div>
<button type="submit">Sign in</button>
</form>
ordinal — English ordinal suffix — the 11th/12th/13th trap · 4 examplescompile error: n/.cache/ai_proxy_cs/Program.cs(6,23): error CS0246: The type or namespace name 'ArgumentOutOfRangeException' could not be found (are you missing a using directive opublic static class Sol
{
public static string Ordinal(int n)
{
if (n < 0)
throw new ArgumentOutOfRangeException(nameof(n), "n must be non-negative");
// Special case for numbers ending in 11, 12, or 13
int lastTwoDigits = n % 100;
if (lastTwoDigits == 11 || lastTwoDigits == 12 || lastTwoDigits == 13)
return n + "th";
// Determine suffix based on last digit
int lastDigit = n % 10;
string suffix = lastDigit switch
{
1 => "st",
2 => "nd",
3 => "rd",
_ => "th"
};
return n + suffix;
}
}
compile error: ome/crimson/.cache/ai_proxy_cs/Program.cs(6,23): error CS0246: The type or namespace name 'ArgumentException' could not be found (are you missing a using directive opublic static class Sol
{
public static string Ordinal(int n)
{
if (n < 0)
throw new ArgumentException("Input must be non-negative", nameof(n));
int lastTwoDigits = n % 100;
string suffix;
if (lastTwoDigits >= 11 && lastTwoDigits <= 13)
suffix = "th";
else
{
switch (n % 10)
{
case 1: suffix = "st"; break;
case 2: suffix = "nd"; break;
case 3: suffix = "rd"; break;
default: suffix = "th"; break;
}
}
return $"{n}{suffix}";
}
}
compile error: n/.cache/ai_proxy_cs/Program.cs(6,23): error CS0246: The type or namespace name 'ArgumentOutOfRangeException' could not be found (are you missing a using directive opublic static class Sol
{
public static string Ordinal(int n)
{
if (n < 0)
throw new ArgumentOutOfRangeException(nameof(n), "n must be non-negative");
int lastTwoDigits = n % 100;
if (lastTwoDigits >= 11 && lastTwoDigits <= 13)
return $"{n}th";
int lastDigit = n % 10;
return lastDigit switch
{
1 => $"{n}st",
2 => $"{n}nd",
3 => $"{n}rd",
_ => $"{n}th"
};
}
}
compile error: n/.cache/ai_proxy_cs/Program.cs(6,23): error CS0246: The type or namespace name 'ArgumentOutOfRangeException' could not be found (are you missing a using directive opublic static class Sol
{
public static string Ordinal(int n)
{
if (n < 0)
throw new ArgumentOutOfRangeException(nameof(n), "n must be non-negative");
// Special case for numbers ending in 11, 12, or 13
int lastTwoDigits = n % 100;
if (lastTwoDigits == 11 || lastTwoDigits == 12 || lastTwoDigits == 13)
return n + "th";
// Determine suffix based on last digit
int lastDigit = n % 10;
string suffix = lastDigit switch
{
1 => "st",
2 => "nd",
3 => "rd",
_ => "th"
};
return n + suffix;
}
}
snake_to_camel — snake_case to camelCase — digits stop capitalisation · 4 examplessnake_to_camel("_leading") → "leading", expected "Leading"fn snake_to_camel(s: &str) -> String {
let mut result = String::with_capacity(s.len());
let mut capitalize_next = false;
let mut first_char_found = false;
// We iterate through characters, but we need to handle the "leading underscores"
// by skipping them until we find a non-underscore character.
let chars: Vec<char> = s.chars().collect();
let mut i = 0;
// Skip leading underscores
while i < chars.len() && chars[i] == '_' {
i += 1;
}
// Determine the end index to handle trailing underscores
let mut last_valid_idx = chars.len();
while last_valid_idx > i && chars[last_valid_idx - 1] == '_' {
last_valid_idx -= 1;
}
for j in i..last_valid_idx {
let c = chars[j];
if c == '_' {
capitalize_next = true;
} else {
if capitalize_next {
result.push(c.to_uppercase().next().unwrap_or(c));
capitalize_next = false;
} else {
result.push(c);
}
first_char_found = true;
}
}
// If the string was all underscores or empty, result is already empty.
// The logic above handles the "no underscores" case by simply pushing chars as they are.
result
}
snake_to_camel("_leading") → "leading", expected "Leading"fn snake_to_camel(s: &str) -> String {
let mut result = String::with_capacity(s.len());
let mut capitalize_next = false;
let mut first_char_found = false;
// We iterate through characters, but we need to handle the "leading underscores"
// by skipping them until we find a non-underscore character.
let chars: Vec<char> = s.chars().collect();
let mut i = 0;
// Skip leading underscores
while i < chars.len() && chars[i] == '_' {
i += 1;
}
// Determine the end index to handle trailing underscores
let mut last_valid_idx = chars.len();
while last_valid_idx > i && chars[last_valid_idx - 1] == '_' {
last_valid_idx -= 1;
}
for j in i..last_valid_idx {
let c = chars[j];
if c == '_' {
capitalize_next = true;
} else {
if capitalize_next {
result.push(c.to_ascii_uppercase());
capitalize_next = false;
} else {
result.push(c);
}
first_char_found = true;
}
}
result
}
snake_to_camel("hello_world") → "HelloWorld", expected "helloWorld"fn snake_to_camel(s: &str) -> String {
let mut result = String::with_capacity(s.len());
let mut chars = s.chars().peekable();
let mut skip_underscores = true;
while let Some(c) = chars.next() {
if c == '_' {
skip_underscores = true;
} else if skip_underscores {
if c.is_alphabetic() {
result.push(c.to_uppercase().next().unwrap());
} else {
result.push(c);
}
skip_underscores = false;
} else {
result.push(c);
}
}
result
}
snake_to_camel("_leading") → "leading", expected "Leading"fn snake_to_camel(s: &str) -> String {
let mut result = String::new();
let chars: Vec<char> = s.chars().collect();
let len = chars.len();
// Find the first non-underscore character to skip leading underscores
let mut i = 0;
while i < len && chars[i] == '_' {
i += 1;
}
// Process from the first non-underscore character
while i < len {
if chars[i] == '_' {
// Skip all consecutive underscores
while i < len && chars[i] == '_' {
i += 1;
}
// If we haven't reached the end, uppercase the next alphabetic character
if i < len {
let c = chars[i];
if c.is_alphabetic() {
result.push(c.to_ascii_uppercase());
} else {
result.push(c);
}
i += 1;
}
} else {
result.push(chars[i]);
i += 1;
}
}
// Remove trailing underscores (they would have been skipped in the loop above, but let's ensure)
// Actually, trailing underscores are already handled because after skipping them, i == len and nothing is added.
// But we need to make sure no trailing underscore was added. Since we only add when we see a non-underscore or after skipping underscores,
// trailing underscores are naturally omitted.
result
}
base_convert — Integer between bases 2-36 with validation · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def base_convert(s: str, frm: int, to: int) -> str | None:
if not (2 <= frm <= 36 and 2 <= to <= 36):
return None
if not s:
return None
# Handle sign
sign = ""
start_idx = 0
if s[0] == '-':
sign = "-"
start_idx = 1
if len(s) == 1:
return None
elif s[0] == '+':
start_idx = 1
if len(s) == 1:
return None
# Extract digits and validate
digits_str = s[start_idx:]
chars = "0123456789abcdefghijklmnopqrstuvwxyz"
char_to_val = {c: i for i, c in enumerate(chars)}
decimal_value = 0
for char in digits_str.lower():
if char not in char_to_val or char_to_val[char] >= frm:
return None
decimal_value = decimal_value * frm + char_to_val[char]
# Handle zero case (no sign for zero)
if decimal_value == 0:
return "0"
# Convert to target base
res = []
temp_val = decimal_value
while temp_val > 0:
res.append(chars[temperal_val % to])
temp_val //= to
result_str = "".join(reversed(res))
return sign + result_str
# Correcting the logic for a clean implementation
def base_convert(s: str, frm: int, to: int) -> str | None:
if not (2 <= frm <= 36 and 2 <= to <= 36):
return None
if not s:
return None
# Parse sign
is_negative = False
idx = 0
if s[0] == '-':
is_negative = True
idx += 1
elif s[0] == '+':
idx += 1
if idx == lenSyntaxError: invalid syntax (<model>, line 1)def base_convert(s: str, frm: int, to: int) -> str | None:
if not (2 <= frm <= 36 and 2 <= to <= 36):
return None
if not s or s in ('-', '+'):
return None
# Handle sign
is_negative = False
start_idx = 0
if s[0] == '-':
is_negative = True
start_idx = 1
elif s[0] == '+':
start_idx = 1
# Check if there's actual content after sign
digits_part = s[start_idx:]
if not digits_part:
return None
chars = "0123456789abcdefghijklmnopqrstuvwxyz"
char_to_val = {c: i for i, c in enumerate(chars)}
# Convert input string to decimal integer
decimal_value = 0
for char in digits_part.lower():
if char not in char_to_val or char_to_val[char] >= frm:
return None
decimal_value = decimal_value * frm + char_to_val[char]
# Handle zero case explicitly (no sign, no leading zeros)
if decimal_value == 0:
return "0"
# Convert decimal to target base
res = []
temp_val = decimal_value
while temp_val > 0:
res.append(chars[temperal_val % to])
temp_val //= to
result_str = "".join(reversed(res))
return ("-" if is_negative else "") + result_str
# Re-implementing logic cleanly for the final block
def base_convert(s: str, frm: int, to: int) -> str | None:
if not (2 <= frm <= 36 and 2 <= to <= 36):
return None
if not s or s in ('-', '+'):
return None
chars = "0123456789abcdefghijklmnopqrstuvwxyz"
char_to_val = {c: i for i, c in enumerate(no callable named 'base_convert'def base_conversion(s, frm, to):
if frm < 2 or frm > 36 or to < 2 or to > 36:
return None
if not s:
return None
sign = ''
if s[0] == '-':
sign = '-'
s = s[1:]
elif s[0] == '+':
s = s[1:]
if not s:
return None
digits = '0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ'
val = 0
for ch in s:
if ch not in digits[:frm]:
return None
val = val * frm + digits.index(ch)
if val == 0:
return '0'
res = ''
while val > 0:
res = digits[val % to] + res
val //= to
return sign + res
base_convert("-1A", 16, 10) → "26-", expected "-26"def base_convert(s: str, frm: int, to: int) -> str | None:
# Validate bases
if not (2 <= frm <= 36) or not (2 <= to <= 36):
return None
# Check for empty string
if not s:
return None
# Check if string is only a sign
if s in ('+', '-'):
return None
# Parse the sign
negative = False
if s[0] == '-':
negative = True
s = s[1:]
elif s[0] == '+':
s = s[1:]
# Check if after removing sign, string is empty
if not s:
return None
# Convert from base frm to integer
try:
num = int(s, frm)
except ValueError:
return None
# If the original string had a negative sign, negate
if negative:
num = -num
# Handle zero
if num == 0:
return '0'
# Convert integer to base 'to'
digits = '0123456789abcdefghijklmnopqrstuvwxyz'
if num < 0:
result = '-'
num = -num
else:
result = ''
while num > 0:
result = digits[num % to] + result
num //= to
return result
round_to — Round to nearest multiple, halves away from zero · 4 examplesround_to(-25, 10) → "-20", expected -30int round_to(int n, int m) {
if (m <= 0) return n; // Handle invalid m, though problem states m > 0
int quotient = n / m;
int remainder = n % m;
// For positive n: if remainder >= m/2, round up; else round down
// For negative n: if remainder <= -m/2, round down (more negative); else round up
// Since C truncates toward zero, remainder has same sign as n
if (n >= 0) {
if (remainder >= m - remainder) {
return (quotient + 1) * m;
} else {
return quotient * m;
}
} else {
// n < 0 case
if (remainder <= -(m - remainder)) {
return (quotient - 1) * m;
} else {
return quotient * m;
}
}
}
compile error: cted ';' before 'int'
6 | ```c
| ^
| ;
7 | int round_to(int n, int m) {
| ~~~
/tmp/bench_c_l2kd8s7i/task.c: In function 'round_toint round_to(int n, int m) {
// Round n to the nearest multiple of m, with exact halves rounding away from zero.
// For positive n: (n + m/2) / m * m works, but we need to handle negatives correctly.
// Use floor division: floor((n + m/2) / m) * m for positive, but for negative it fails.
// Better approach: compute the remainder and adjust.
// We can use the fact that C integer division truncates toward zero.
// To round to nearest multiple of m (m > 0), with ties away from zero:
// If n >= 0: result = ((n + m/2) / m) * m
// If n < 0: result = -((-n + m/2) / m) * m -- but this doesn't handle ties correctly.
// Let's think about it differently.
// The "round half away from zero" means:
// - For positive numbers, 25 rounds to 30 (since 25 is exactly halfway between 20 and 30, we go to 30)
// - For negative numbers, -25 rounds to -30
// A clean way: use floor division.
// floor_div(a, b) = (a >= 0 || a % b == 0) ? a / b : (a - b + 1) / b ... actually let's just compute it directly.
// Let's use: q = (n + m/2) / m for positive, but for negative we need different logic.
// Actually, let's just do it with a conditional or a clever formula.
// One approach:
// sign = n >= 0 ? 1 : -1
// abs_n = n >= 0 ? n : -n
// result = sign * ((abs_n + m/2) / m) * m
// But this fails for exact halves? Let's check: n=25, m=10: abs_n=25, (25+5)/10 = 30, *1 = 30. Good.
// n=-25, m=10: abs_nround_to(-25, 10) → "-35", expected -30int round_to(int n, int m) {
if (m <= 0) return n;
// Handle positive and negative numbers separately to avoid truncation issues
if (n >= 0) {
return ((n + m / 2) / m) * m;
} else {
// For negative numbers, we need to round away from zero on exact halves
// n is negative, so let's work with absolute value conceptually
// We want: if n = -k*m - r where 0 <= r < m, then:
// if r == 0: result is n (already a multiple)
// if r > 0 and r < m/2: round toward zero -> n + r = n - (-r) ... let's think differently
// Alternative approach: use division that rounds toward negative infinity for the quotient
// For negative n, C division truncates toward zero, so n/m is like ceil(|n|/m) in magnitude but positive direction
// Let's compute the remainder properly
int abs_n = -n; // since n < 0
int r = abs_n % m;
if (r == 0) {
return n;
}
// Now we have n = -(q*m + r) where q >= 0, 0 < r < m
// We want to round away from zero on exact halves
// So if r > m/2 or (r == m/2 and m is even, i.e., exact half), we round further from zero
// If r < m/2, we round toward zero
// For exact half: r == m/2 requires m to be even
int threshold = m / 2;
if (r > threshold || (m % 2 == 0 && r == threshold)) {
// Round away from zero: make the magnitude larger
// n is negative, so we subtract more
return n - m;
} else
… (truncated)round_to(-25, 10) → "-35", expected -30int round_to(int n, int m) {
if (m <= 0) return n;
// Handle positive and negative numbers separately to avoid truncation issues
if (n >= 0) {
return ((n + m / 2) / m) * m;
} else {
// For negative numbers, we need to round away from zero on exact halves
// n is negative, so let's work with absolute value conceptually
// We want: if n = -k*m - r where 0 <= r < m, then:
// if r == 0: result is n (already a multiple)
// if r > 0 and r < m/2: round toward zero -> n + r = n - (-r) ... let's think differently
// Alternative approach: use division that rounds toward negative infinity for the quotient
// For negative n, C division truncates toward zero, so n/m is like ceil(|n|/m) in magnitude but positive direction
// Let's compute the remainder properly
int abs_n = -n; // since n < 0
int r = abs_n % m;
if (r == 0) {
return n;
}
// Now we have n = -(q*m + r) where q >= 0, 0 < r < m
// We want to round away from zero on exact halves
// So if r > m/2 or (r == m/2 and m is even, i.e., exact half), we round further from zero
// If r < m/2, we round toward zero
// For exact half: r == m/2 requires m to be even
int threshold = m / 2;
if (r > threshold || (m % 2 == 0 && r == threshold)) {
// Round away from zero: make the magnitude larger
// n is negative, so we subtract more
return n - m;
} else
… (truncated)tokenize_expr — Tokenise arithmetic, None on invalid input · 4 examplesSyntaxError: invalid syntax (<model>, line 1)import re
def tokenize_expr(s: str) -> list | None:
# Regex to find all potential tokens: numbers, operators, or parentheses
# We use a pattern that captures sequences of digits/dots OR single characters
pattern = r'(\d+\.\d+|\d+|[+\-*/()])'
# First, check for invalid characters (anything not digit, dot, operator, paren, or whitespace)
allowed_chars = set("0123456789.+*/()- \t\n\r")
if not all(c in allowed_chars for c in s):
return None
tokens = []
# We use finditer to scan the string and manually validate structure
# However, a simpler way is to split by whitespace and then refine tokens.
# But since operators don't require whitespace (e.g., "1+2"), we use regex.
# Find all matches of our pattern
matches = list(re.finditer(pattern, s))
# Check if the sum of lengths of matched tokens and ignored whitespace equals string length
# This ensures no "garbage" characters were skipped by the regex
covered_indices = []
for m in matches:
tokens.append(m.group())
covered_indices.extend(range(m.start(), m.end()))
# Check if there are any non-whitespace characters not covered by tokens
for i, char in enumerate(s):
if char not in " \t\n\r" and i not in covered_indices:
return None
processed_tokens = []
for t in tokens:
# Check for invalid number format (multiple dots) - though regex handles most
if '.' in t:
if t.count('.') > 1:
return None
# Convert to float
try:
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import re
def tokenize_expr(s: str) -> list | None:
# Regex to find all potential tokens: numbers, operators, or parentheses
# We use a pattern that captures sequences of digits/dots OR single characters
pattern = r'(\d+\.\d+|\d+|[+\-*/()])'
# First, check for invalid characters (anything not digit, dot, operator, paren, or whitespace)
allowed_chars = set("0123456789.+*/()- \t\n\r")
if not all(c in allowed_chars for c in s):
return None
tokens = []
# We use finditer to scan the string and manually validate structure
# However, a simpler way is to split by whitespace and then refine tokens.
# But since operators don't require whitespace (e.g., "1+2"), we use regex.
# Find all matches of our pattern
matches = list(re.finditer(pattern, s))
# Check if the sum of lengths of matched tokens and ignored whitespace equals string length
# This ensures no "garbage" characters were skipped by the regex
covered_indices = []
for m in matches:
tokens.append(m.group())
covered_indices.extend(range(m.start(), m.end()))
# Check if there are any non-whitespace characters not covered by tokens
for i, char in enumerate(s):
if char not in " \t\n\r" and i not in covered_indices:
return None
processed_tokens = []
for t in tokens:
# Check for invalid number format (multiple dots) - though regex handles most
if '.' in t:
if t.count('.') > 1:
return None
# Convert to float
try:
… (truncated)tokenize_expr("1.2.3") → [1.2, 0.3], expected nulldef tokenize_expr(s: str) -> list | None:
tokens = []
i = 0
n = len(s)
while i < n:
if s[i].isspace():
i += 1
continue
if s[i] in '+-*/()':
tokens.append(s[i])
i += 1
continue
if s[i].isdigit() or s[i] == '.':
# Start of a number
start = i
has_dot = False
# Check for leading dot (e.g., ".5")
if s[i] == '.':
has_dot = True
i += 1
# Must have at least one digit after the dot
if i >= n or not s[i].isdigit():
return None
while i < n and s[i].isdigit():
i += 1
else:
# Digits before dot
while i < n and s[i].isdigit():
i += 1
# Check for decimal point
if i < n and s[i] == '.':
if has_dot: # Already found a dot
return None
has_dot = True
i += 1
# Must have at least one digit after the dot
if i >= n or not s[i].isdigit():
return None
while i < n and s[i].isdigit():
i += 1
# Extract the number string
num_str = s[start:i]
# Validate: no more than one dot (already handled above)
# Check for invalid characters in the num
… (truncated)tokenize_expr("1.2.3") → [1.2, 0.3], expected nulldef tokenize_expr(s: str) -> list | None:
tokens = []
i = 0
n = len(s)
while i < n:
if s[i].isspace():
i += 1
continue
if s[i] in '+-*/()':
tokens.append(s[i])
i += 1
continue
if s[i].isdigit() or s[i] == '.':
# Start of a number
start = i
has_dot = False
# Check for leading dot (like ".5")
if s[i] == '.':
has_dot = True
i += 1
# Must have at least one digit after the dot
if i >= n or not s[i].isdigit():
return None
else:
# Consume digits before any dot
while i < n and s[i].isdigit():
i += 1
# Check for decimal point
if i < n and s[i] == '.':
if has_dot: # Already had a dot
return None
has_dot = True
i += 1
# Must have at least one digit after the dot
if i >= n or not s[i].isdigit():
return None
# Consume remaining digits after the dot
while i < n and s[i].isdigit():
i += 1
num_str = s[start:i]
# Validate number format: must have at most one dot
if num_str.count('.') > 1:
return None
… (truncated)justify — Full text justification · 4 examplesjustify(["aa", "bb"], 5) → ["aa ", "bb "], expected ["aa bb"]def justify(words: list, width: int) -> list:
result = []
i = 0
n = len(words)
while i < n:
# Find how many words can fit in current line
line_words = []
line_length = 0
while i < n and line_length + len(words[i]) + (1 if line_words else 0) <= width:
line_words.append(words[i])
line_length += len(words[i]) + (1 if line_words else 0)
i += 1
# Build the line
if i == n or len(line_words) == 1:
# Last line or single word line: left justified
line = ' '.join(line_words)
line += ' ' * (width - len(line))
else:
# Distribute spaces evenly
total_spaces = width - sum(len(w) for w in line_words)
gaps = len(line_words) - 1
base_spaces = total_spaces // gaps
extra_spaces = total_spaces % gaps
# Build line with extra spaces on left gaps
line_parts = []
for j, word in enumerate(line_words):
line_parts.append(word)
if j < gaps: # Not the last word
spaces = base_spaces + (1 if j < extra_spaces else 0)
line_parts.append(' ' * spaces)
line = ''.join(line_parts)
result.append(line)
return result
SyntaxError: invalid syntax (<model>, line 1)def justify(words: list, width: int) -> list:
"""
Perform full text justification.
Args:
words: List of words to justify.
width: The target width for each line.
Returns:
List of strings, each of length 'width', with proper justification.
"""
if not words:
return []
# Build lines: each line is a list of words
lines = []
current_line = []
current_length = 0
for word in words:
# If adding this word would exceed width, start a new line
if current_line and current_length + 1 + len(word) > width:
lines.append(current_line)
current_line = [word]
current_length = len(word)
else:
current_line.append(word)
current_length += 1 + len(word) # +1 for the space
# Add the last line
if current_line:
lines.append(current_line)
# Now justify each line
result = []
num_lines = len(lines)
for i, line in enumerate(lines):
if i == num_lines - 1:
# Last line: left-justified
line_str = ' '.join(line)
line_str += ' ' * (width - len(line_str))
result.append(line_str)
elif len(line) == 1:
# Single word line: left-justified
line_str = line[0]
line_str += ' ' * (width - len(line_str))
result.append(line_str)
else:
# Full justification
total_words_length = sum(len(w) for w in line)
total_spaces = width - total_words_length
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def justify(words: list, width: int) -> list:
"""
Perform full text justification.
Pack as many words per line as fit with at least one space between them.
Pad each line to exactly width by distributing spaces between words as evenly
as possible, putting the extra spaces on the LEFT gaps first.
The last line, and any line holding a single word, is left justified with
the remaining space on the right.
"""
if not words:
return []
# Group words into lines
lines = []
current_line = []
current_length = 0
for word in words:
# Calculate the length needed if we add this word
if current_line:
needed = current_length + 1 + len(word)
else:
needed = len(word)
if needed <= width:
current_line.append(word)
current_length = needed if current_line else 0
# Recalculate current_length properly
current_length = sum(len(w) for w in current_line) + (len(current_line) - 1)
else:
lines.append(current_line)
current_line = [word]
current_length = len(word)
if current_line:
lines.append(current_line)
# Justify each line
result = []
for i, line in enumerate(lines):
if i == len(lines) - 1 or len(line) == 1:
# Last line or single word: left justified
line_str = ' '.join(line)
line_str = line_str + ' ' * (width - len(line_str))
result.append(line_str)
else:
# Full
… (truncated)justify(["This", "is", "an", "example", "of", "text", "justification."], 16) → null, expected ["This is an", "example of text", "justification. "]def justify(words: list, width: int) -> list:
def pack_line(start):
line_words = []
total_chars = 0
for i in range(start, len(words)):
word = words[i]
if total_chars + len(word) > width and line_words:
break
line_words.append(word)
total_chars += len(word) + 1 # +1 for space
return start + len(line_words), line_words
lines = []
i = 0
while i < len(words):
num_words, line_words = pack_line(i)
if not line_words:
break
lines.append((num_words, line_words))
i = num_words
justified_lines = []
for num_words, line_words in lines[:-1]:
total_spaces = width - sum(len(word) for word in line_words)
gaps = num_words - 1
if gaps == 0:
justified_lines.append(line_words[0].ljust(width))
continue
base_spaces = total_spaces // gaps
extra_spaces = total_spaces % gaps
line = []
for j in range(gaps):
spaces = base_spaces + (1 if j < extra_spaces else 0)
line.append(line_words[j])
line.append(' ' * spaces)
line.append(line_words[-1])
justified_lines.append(''.join(line))
# Handle last line
if lines:
num_words, line_words = lines[-1]
last_line = ' '.join(line_words).ljust(width)
justified_lines.append(last_line)
return justified_lines
parse_query — Parse a URL query string into an object · 4 examplesReferenceError: parseQuery is not defined/**
* Parses a URL query string into a plain object.
*
* @param {string} qs The query string to parse.
* @returns {object} A plain object representing the parsed query string.
*/
const parseQuery = (qs) => {
if (!qs) {
return {};
}
if (qs.startsWith('?')) {
qs = qs.slice(1);
}
const result = {};
const pairs = qs.split('&');
for (const pair of pairs) {
if (!pair) {
continue;
}
const [key, value] = pair.split('=', 1);
let decodedKey = key ? decodeURIComponent(key.replace(/\+/g, ' ')) : '';
let decodedValue = value !== undefined ? decodeURIComponent(value.replace(/\+/g, ' ')) : '';
if (result.hasOwnProperty(decodedKey)) {
if (!Array.isArray(result[decodedKey])) {
result[decodedKey] = [result[decodedKey]];
}
result[decodedKey].push(decodedValue);
} else {
result[decodedKey] = decodedValue;
}
}
return result;
};
ReferenceError: parseQuery is not defined/**
* Parses a URL query string into a plain object.
*
* @param {string} qs The query string to parse.
* @returns {object} A plain object representing the parsed query string.
*/
const parseQuery = (qs) => {
if (!qs) {
return {};
}
if (qs.startsWith('?')) {
qs = qs.slice(1);
}
const result = {};
const pairs = qs.split('&');
for (const pair of pairs) {
if (!pair) {
continue;
}
const [key, value] = pair.split('=', 1);
let decodedKey = key ? decodeURIComponent(key.replace(/\+/g, ' ')) : '';
let decodedValue = value !== undefined ? decodeURIComponent(value.replace(/\+/g, ' ')) : '';
if (result.hasOwnProperty(decodedKey)) {
if (!Array.isArray(result[decodedKey])) {
result[decodedKey] = [result[decodedKey]];
}
result[decodedKey].push(decodedValue);
} else {
result[decodedKey] = decodedValue;
}
}
return result;
};
parseQuery("a=%3D=&b") → {"a": "=", "b": ""}, expected {"a": "==", "b": ""}function parseQuery(qs) {
if (!qs || qs[0] === '?') qs = qs.substring(1);
const result = {};
if (!qs) return result;
qs.split('&').forEach(pair => {
if (!pair) return;
let [key, value] = pair.split('=', 2);
key = decodeURIComponent(key.replace(/\+/g, ' '));
value = value === undefined ? '' : decodeURIComponent(value.replace(/\+/g, ' '));
if (result[key]) {
if (!Array.isArray(result[key])) result[key] = [result[key]];
result[key].push(value);
} else {
result[key] = value;
}
});
return result;
}
parseQuery("a=%3D=&b") → {"a": "=", "b": ""}, expected {"a": "==", "b": ""}function parseQuery(qs) {
if (!qs || qs[0] === '?') qs = qs.substring(1);
const result = {};
if (!qs) return result;
qs.split('&').forEach(pair => {
if (!pair) return;
let [key, value] = pair.split('=', 2);
key = decodeURIComponent(key.replace(/\+/g, ' '));
value = value === undefined ? '' : decodeURIComponent(value.replace(/\+/g, ' '));
if (result[key]) {
if (!Array.isArray(result[key])) result[key] = [result[key]];
result[key].push(value);
} else {
result[key] = value;
}
});
return result;
}
group_ranges — Collapse consecutive integers into range strings · 4 examplesReferenceError: groupRanges is not defined/**
* Formats a strictly increasing array of integers as a compact range string.
* Consecutive runs collapse to 'start-end'; isolated values stand alone; parts join with ','.
* Two adjacent values are already a range ('1-2'). Negative numbers keep their sign, so -3..-1 renders as '-3--1'.
* An empty array returns ''.
* @param {number[]} nums The input array of integers.
* @returns {string} The formatted range string.
*/
const groupRanges = (nums) => {
if (!nums || nums.length === 0) {
return '';
}
const result = [];
let start = nums[0];
let end = nums[0];
for (let i = 1; i < nums.length; i++) {
if (nums[i] === end + 1) {
end = nums[i];
} else {
if (start === end) {
result.push(String(start));
} else {
result.push(`${start}-${end}`);
}
start = nums[i];
end = nums[i];
}
}
if (start === end) {
result.push(String(start));
} else {
result.push(`${start}-${end}`);
}
return result.join(',');
};
groupRanges([1, 2, 3, 7, 9, 10, 11]) → "1,3-2,7-3,9-7,10-9,11-10", expected "1-3,7,9-11"function groupRanges(nums) {
if (nums.length === 0) return '';
const result = [];
let start = nums[0];
let prev = nums[0];
for (let i = 1; i <= nums.length; i++) {
// When we reach the end or find a non-consecutive number
if (i === nums.length || nums[i] !== prev + 1) {
if (start === prev) {
result.push(String(start));
} else {
result.push(`${start}-${prev}`);
}
// Reset start for the next range
if (i < nums.length) {
start = nums[i];
}
}
prev = nums[i - 1];
}
return result.join(',');
}
groupRanges([1, 2, 3, 7, 9, 10, 11]) → "1,3-2,7-3,9-7,10-9,11-10", expected "1-3,7,9-11"function groupRanges(nums) {
if (nums.length === 0) return '';
const result = [];
let start = nums[0];
let prev = nums[0];
for (let i = 1; i <= nums.length; i++) {
// When we reach the end or find a non-consecutive number
if (i === nums.length || nums[i] !== prev + 1) {
if (start === prev) {
result.push(String(start));
} else {
result.push(`${start}-${prev}`);
}
// Reset start for the next range
if (i < nums.length) {
start = nums[i];
}
}
prev = nums[i - 1];
}
return result.join(',');
}
ReferenceError: groupRanges is not defined/**
* Formats a strictly increasing array of integers as a compact range string.
* Consecutive runs collapse to 'start-end'; isolated values stand alone; parts join with ','.
* Two adjacent values are already a range ('1-2'). Negative numbers keep their sign, so -3..-1 renders as '-3--1'.
* An empty array returns ''.
* @param {number[]} nums The input array of integers.
* @returns {string} The formatted range string.
*/
const groupRanges = (nums) => {
if (!nums || nums.length === 0) {
return '';
}
const result = [];
let start = nums[0];
let end = nums[0];
for (let i = 1; i < nums.length; i++) {
if (nums[i] === end + 1) {
end = nums[i];
} else {
if (start === end) {
result.push(String(start));
} else {
result.push(`${start}-${end}`);
}
start = nums[i];
end = nums[i];
}
}
if (start === end) {
result.push(String(start));
} else {
result.push(`${start}-${end}`);
}
return result.join(',');
};
topo_lex — Smallest topological order, None on cycle · 4 examplestopo_lex(4, [[0, 1], [1, 2], [2, 3]]) → null, expected [0, 1, 2, 3]from collections import defaultdict, deque
def topo_lex(n: int, edges: list) -> list | None:
# Build adjacency list and in-degree count
graph = defaultdict(list)
in_degree = [0] * n
for u, v in edges:
if 0 <= u < n and 0 <= v < n:
graph[u].append(v)
in_degree[v] += 1
# Use a min-heap (priority queue) to always get the smallest available node
# Since heapq is a min-heap, we can use it directly
heap = []
for i in range(n):
if in_degree[i] == 0:
heapq.heappush(heap, i)
result = []
while heap:
node = heapq.heappop(heap)
result.append(node)
for neighbor in graph[node]:
in_degree[neighbor] -= 1
if in_degree[neighbor] == 0:
heapq.heappush(heap, neighbor)
# If we couldn't process all nodes, there's a cycle
if len(result) != n:
return None
return result
topo_lex(2, [[0, 1], [1, 0]]) → [], expected nulldef topo_lex(n, edges):
from collections import defaultdict, deque
adj = defaultdict(list)
indeg = [0]*n
for u, v in edges:
adj[u].append(v)
indeg[v] += 1
q = deque([i for i in range(n) if indeg[i]==0])
res = []
while q:
# pop smallest available node
u = min(q)
q.remove(u)
res.append(u)
for v in adj[u]:
indeg[v] -= 1
if indeg[v]==0:
q.append(v)
if len(res) != n:
return [] # cycle
return res
topo_lex(3, [[1, 0]]) → [1, 2, 0], expected [1, 0, 2]def topo_lex(n: int, edges: list) -> list | None:
from collections import deque
# Build adjacency list and in-degree count
adj = [[] for _ in range(n)]
in_degree = [0] * n
for u, v in edges:
if u != v: # Ignore self-loops (they create cycles)
adj[u].append(v)
in_degree[v] += 1
# Initialize queue with nodes having zero in-degree
q = deque()
for i in range(n):
if in_degree[i] == 0:
q.append(i)
topo_order = []
while q:
# To get lexicographically smallest, process nodes in order
u = q.popleft()
topo_order.append(u)
for v in adj[u]:
in_degree[v] -= 1
if in_degree[v] == 0:
q.append(v)
# Check if topological sort is possible (no cycle)
if len(topo_order) != n:
return None
return topo_order
topo_lex(3, [[1, 0]]) → [1, 2, 0], expected [1, 0, 2]def topo_lex(n: int, edges: list) -> list | None:
from collections import defaultdict, deque
# Build adjacency list and in-degree count
adj = defaultdict(list)
in_degree = [0] * n
for u, v in edges:
if u == v or u < 0 or u >= n or v < 0 or v >= n:
continue # Skip invalid edges (though problem says nodes are 0..n-1)
adj[u].append(v)
in_degree[v] += 1
# Initialize queue with nodes having zero in-degree, sorted for lex order
q = deque()
for i in range(n):
if in_degree[i] == 0:
q.append(i)
topo_order = []
while q:
# Process nodes in lexicographical order
u = q.popleft()
topo_order.append(u)
for v in sorted(adj[u]):
in_degree[v] -= 1
if in_degree[v] == 0:
q.append(v)
if len(topo_order) != n:
return None # Cycle detected
return topo_order
pluck — Column from associative rows — null vs missing key · 4 examplespluck([{"a": 1}, {"b": 9}, {"a": 3}], "a") → {"0": 1, "2": 3}, expected [1, 3]function pluck(array $rows, string $key): array {
return array_map(
function($row) use ($key) {
return $row[$key];
},
array_filter($rows, function($row) use ($key) {
return array_key_exists($key, $row);
})
);
}
pluck([{"a": 1}, {"b": 9}, {"a": 3}], "a") → [1, null, 3], expected [1, 3]function pluck(array $rows, string $key): array {
return array_map(
fn($row) => array_key_exists($key, $row) ? $row[$key] : null,
$rows
);
}
pluck([{"a": 1}, {"b": 9}, {"a": 3}], "a") → [1, null, 3], expected [1, 3]function pluck(array $rows, string $key): array {
return array_map(
fn($row) => array_key_exists($key, $row) ? $row[$key] : null,
$rows
);
}
pluck([{"a": 1}, {"b": 9}, {"a": 3}], "a") → [1, null, 3], expected [1, 3]function pluck(array $rows, string $key): array {
return array_map(
fn($row) => array_key_exists($key, $row) ? $row[$key] : null,
$rows
);
}
lru_ops — LRU cache with eviction order · 4 exampleslru_ops(2, [["put", 1, 1], ["put", 2, 2], ["put", 1, 10], ["put", 3, 3], ["get", 2], ["get", 1]]) → null, expected [-1, 10]def lru_ops(capacity: int, ops: list) -> list:
"""Simulates an LRU cache and returns results of 'get' operations.
Args:
capacity: The maximum capacity of the cache.
ops: A list of operations to perform on the cache. Each operation is a list
of the form ['put', key, value] or ['get', key].
Returns:
A list of results of the 'get' operations only, using -1 for a miss.
"""
cache = {}
usage = [] # List to track usage order (most recent last)
results = []
for op in ops:
if op[0] == 'put':
key = op[1]
value = op[2]
if key in cache:
cache[key] = value
usage.remove(key) # Move to end (most recently used)
else:
if len(cache) >= capacity:
lru_key = usage.pop(0) # Evict least recently used
del cache[lru_key]
cache[key] = value
usage.append(key) # Add to end (most recently used)
elif op[0] == 'get':
key = op[1]
if key in cache:
usage.remove(key) # Move to end (most recently used)
usage.append(key)
results.append(cache[key])
else:
results.append(-1)
return results
SyntaxError: invalid syntax (<model>, line 1)```python code
def lru_ops(capacity: int, ops: list) -> list:
cache = {}
usage_order = []
def update_usage(key):
if key in usage_order:
usage_order.remove(key)
usage_order.append(key)
results = []
for op in ops:
if op[0] == 'put':
key, value = op[1], op[2]
if key in cache:
cache[key] = value
update_usage(key)
else:
if len(cache) >= capacity:
lru_key = usage_order.pop(0)
del cache[lru_key]
cache[key] = value
update_usage(key)
elif op[0] == 'get':
key = op[1]
if key in cache:
update_usage(key)
results.append(cache[key])
else:
results.append(-1)
return results
```lru_ops(2, [["put", 1, 1], ["put", 2, 2], ["get", 1], ["put", 3, 3], ["get", 2], ["get", 3]]) → [1, 2, 3], expected [1, -1, 3]def lru_ops(capacity: int, ops: list) -> list:
cache = {}
result = []
for op in ops:
if op[0] == 'put':
if len(cache) >= capacity and op[1] not in cache:
lru_key = min(cache, key=cache.get)
del cache[lru_key]
cache[op[1]] = (op[2], len(cache))
elif op[0] == 'get':
if op[1] in cache:
value, _ = cache[op[1]]
cache[op[1]] = (value, len(cache))
result.append(value)
else:
result.append(-1)
return result
lru_ops(2, [["put", 1, 1], ["put", 2, 2], ["get", 1], ["put", 3, 3], ["get", 2], ["get", 3]]) → null, expected [1, -1, 3]def lru_ops(capacity: int, ops: list) -> list:
cache = {}
result = []
for op in ops:
if op[0] == 'put':
if op[1] in cache:
del cache[op[1]]
elif len(cache) >= capacity:
lru_key = min(cache, key=cache.get)
del cache[lru_key]
cache[op[1]] = op[2]
elif op[0] == 'get':
if op[1] in cache:
result.append(cache[op[1]])
del cache[op[1]]
cache[op[1]] = cache.get(op[1], None)
else:
result.append(-1)
return result
slugify — URL slug: symbol runs become one hyphen · 4 examplesslugify("Hello, World!") → "Warning: preg_replace(): Unknown modifier '+' in /tmp/bench_\u2026", expected "hello-world"function slugify(string $s): string {
$slug = preg_replace('[^[:alnum:]]+', '-', strtolower($s));
return trim($slug, '-');
}
slugify("Hello, World!") → "Warning: preg_replace(): Unknown modifier '+' in /tmp/bench_\u2026", expected "hello-world"function slugify(string $s): string {
$slug = preg_replace('[^[:alnum:]]+', '-', strtolower($s));
return trim($slug, '-');
}
slugify("a--b__c") → "a-b__c", expected "a-b-c"function slugify(string $s): string {
// Convert to lowercase
$slug = strtolower($s);
// Replace non-alphanumeric characters with a hyphen, but only if they are not already hyphens.
// This handles cases like "a---b" -> "a-b" and "a!@#b" -> "a-b".
$slug = preg_replace('/[^\w\d]+/', '-', $slug);
// Trim leading and trailing hyphens
$slug = trim($slug, '-');
// If the original string contained nothing alphanumeric, the result after trimming might be empty.
// The regex replacement above ensures that if only special characters were present, it becomes a sequence of hyphens,
// which then gets trimmed to an empty string, satisfying the requirement.
return $slug;
}
slugify("Hello, World!") → "hello--world", expected "hello-world"function slugify(string $s): string {
$s = preg_replace('/[^a-z0-9]/i', '-', $s);
$s = trim($s, '-');
return strtolower($s);
}
json_pointer — Resolve an RFC 6901 JSON Pointer · 4 examplesno callable named 'json_pointer'
def resolve_pointer(doc, pointer):
if pointer == "":
return doc
tokens = pointer.split("/")[1:] # skip the leading '/'
current = doc
for token in tokens:
# decode ~1 -> / and ~0 -> ~ (order matters)
decoded = token.replace("~1", "/").replace("~0", "~")
if isinstance(current, list):
if decoded == "-":
# indicates the "next" element, but we don't support that here
return None
try:
idx = int(decoded)
except ValueError:
return None
if idx < 0 or idx >= len(current):
return None
current = current[idx]
elif isinstance(current, dict):
if decoded not in current:
return None
current = current[decoded]
else:
return None
return current
SyntaxError: invalid syntax (<model>, line 1)import re
def json_pointer(doc, pointer: str):
"""
Resolves an RFC 6901 JSON Pointer against doc, returning the referenced value or None if it does not resolve.
Args:
doc: The document (Python object) to query.
pointer: The JSON Pointer string.
Returns:
The referenced value, or None if the pointer is invalid or out of bounds.
"""
if not pointer:
return doc
# 1. Decode the pointer according to RFC 6901 rules for escaped characters
def decode_token(token):
# '~1' decodes to '/' and '~0' decodes to '~'
# The rule states '~1' must be decoded BEFORE '~0'.
decoded = token.replace('~1', '/').replace('~0', '~')
return decoded
# Split the pointer by '/', but we need a custom split/decode process because
# the components themselves might contain escaped slashes or tildes.
# A simpler approach is to first decode the whole string, then split by '/'
# However, standard JSON Pointer decoding requires careful handling of escapes *within* tokens.
# Let's use a regex-based iterative decoding/splitting process that respects escaping:
tokens = []
current_token = ""
i = 0
while i < len(pointer):
char = pointer[i]
if char == '/':
if current_token:
# Decode the accumulated token before appending it
decoded_token = decode_token(current_token)
tokens.append(decoded_token)
current_token = ""
else:
# Handle leading/consecutive slashes if th
… (truncated)SyntaxError: invalid syntax (<model>, line 1)import json
def json_pointer(doc, pointer: str):
"""
Resolves an RFC 6901 JSON Pointer against doc, returning the referenced value or None if it does not resolve.
Args:
doc: The document (Python object) to query.
pointer: The JSON Pointer string.
Returns:
The referenced value, or None if the pointer is invalid or out of bounds.
"""
if not pointer:
return doc
# 1. Decode the pointer according to RFC 6901 rules for escaped characters
def decode_token(token):
# '~1' decodes to '/' and '~0' decodes to '~'
# The rule states '~1' must be decoded BEFORE '~0'.
decoded = token.replace('~1', '/').replace('~0', '~')
return decoded
# Split the pointer by '/', but we need a custom split that handles escaped slashes correctly
# Since the input 'pointer' is assumed to be already processed for escaping,
# and the standard library approach of splitting on '/' will break if '/' was encoded as '\/',
# we rely on the fact that the tokens themselves are separated by unescaped '/'.
# The problem description implies a pre-processed pointer where only '~1' and '~0' need decoding,
# and the structure is token/token/...
tokens = []
current_token = ""
i = 0
while i < len(pointer):
char = pointer[i]
if char == '/':
tokens.append(decode_token(current_token))
current_token = ""
i += 1
else:
current_token += char
i += 1
tokens.append(decode_token(current_token))
# 2. Trave
… (truncated)no code in responsepath_norm — Normalise a POSIX path with . and .. · 4 examplespath_norm("/../") → "/..", expected "/"def path_norm(path: str) -> str:
if not path:
return '/'
# Split path into components
components = []
current = ''
for char in path:
if char == '/':
if current:
components.append(current)
current = ''
else:
# Handle multiple consecutive slashes
pass
else:
current += char
# Add the last component if it exists
if current:
components.append(current)
# Process components
stack = []
for component in components:
if component == '.' or component == '':
# Current directory or empty component - do nothing
continue
elif component == '..':
# Parent directory - go up one level if possible
if stack and stack[-1] != '..':
stack.pop()
elif not stack or stack[-1] == '..':
# At root or already at root, stay at root
stack.append('..')
else:
stack.append(component)
# Handle special case where we're left with only '..' components
# This represents the root directory
if not stack:
return '/'
# Reconstruct path
result = '/'
for i, component in enumerate(stack):
if i > 0:
result += '/'
result += component
return result
path_norm("/../") → "/..", expected "/"def path_norm(path: str) -> str:
if not path:
return '/'
# Split path into components
components = []
current = ''
for char in path:
if char == '/':
if current:
components.append(current)
current = ''
else:
current += char
# Add the last component if it exists
if current:
components.append(current)
# Process components
stack = []
for component in components:
if component == '.' or component == '':
# Skip current directory references and empty components
continue
elif component == '..':
# Go up one level if possible
if stack and stack[-1] != '..':
stack.pop()
else:
# Stay at root if already at root
stack.append(component)
else:
stack.append(component)
# Handle the case where we're at the root
# If we have a .. at the beginning, we should keep it
result = '/' + '/'.join(stack)
# If result is just '/', return it
if result == '/':
return '/'
# Remove trailing slash
if result.endswith('/'):
result = result[:-1]
return result
path_norm("/../") → "/..", expected "/"def path_norm(path: str) -> str:
stack = []
for component in filter(None, path.split('/')):
if component == '..':
if stack and stack[-1] != '..':
stack.pop()
elif not stack:
stack.append('..')
elif component != '.':
stack.append(component)
return '/' + '/'.join(stack) if stack else '/'
path_norm("/../") → "/..", expected "/"def path_norm(path: str) -> str:
stack = []
for component in filter(None, path.split('/')):
if component == '..':
if stack and stack[-1] != '..':
stack.pop()
elif not stack:
stack.append('..')
elif component != '.':
stack.append(component)
return '/' + '/'.join(stack) if stack else '/'
clamp_mul — Saturating int multiplication · 4 examplesSol.ClampMul(-3, -4) → "2147483647", expected 12public static class Sol
{
public static int ClampMul(int a, int b)
{
if (a == 0 || b == 0)
return 0;
// Check for overflow using division
if (a > 0)
{
if (b > 0)
{
if (a > int.MaxValue / b)
return int.MaxValue;
}
else // b < 0
{
if (b < int.MinValue / a)
return int.MinValue;
}
}
else // a < 0
{
if (b > 0)
{
if (a < int.MinValue / b)
return int.MinValue;
}
else // b < 0
{
if (a != 0 && b > int.MaxValue / a)
return int.MaxValue;
}
}
return a * b;
}
}
compile error: the current context [/home/crimson/.cache/ai_proxy_cs/csrunner.csproj]
/home/crimson/.cache/ai_proxy_cs/Program.cs(9,19): error CS0103: The name 'Math' does not exipublic static class Sol
{
public static int ClampMul(int a, int b)
{
if (a == 0 || b == 0) return 0;
bool negative = (a < 0) ^ (b < 0);
long la = Math.Abs((long)a);
long lb = Math.Abs((long)b);
long product = la * lb;
if (negative)
{
if (product > (long)(uint.MaxValue >> 1) + 1) // overflow below int.MinValue
return int.MinValue;
return -(int)product;
}
else
{
if (product > int.MaxValue)
return int.MaxValue;
return (int)product;
}
}
}
compile error: the current context [/home/crimson/.cache/ai_proxy_cs/csrunner.csproj]
/home/crimson/.cache/ai_proxy_cs/Program.cs(9,19): error CS0103: The name 'Math' does not exipublic static class Sol
{
public static int ClampMul(int a, int b)
{
if (a == 0 || b == 0) return 0;
bool negative = (a < 0) ^ (b < 0);
long la = Math.Abs((long)a);
long lb = Math.Abs((long)b);
long product = la * lb;
if (negative)
{
if (product > (long)(uint.MaxValue >> 1) + 1) // overflow below int.MinValue
return int.MinValue;
return -(int)product;
}
else
{
if (product > int.MaxValue)
return int.MaxValue;
return (int)product;
}
}
}
compile error: text [/home/crimson/.cache/ai_proxy_cs/csrunner.csproj]
Build FAILED.
/home/crimson/.cache/ai_proxy_cs/Program.cs(6,16): error CS0103: The name 'Math' does not exipublic static class Sol
{
public static int ClampMul(int a, int b)
{
long res = (long)a * b;
return Math.Clamp((int)res, int.MinValue, int.MaxValue);
}
}
clamp_add — Saturating int addition — overflow trap · 4 examplesclamp_add(100, 200) → "-2147483648", expected 300int clamp_add(int a, int b) {
if (a > INT_MAX - b) return INT_MAX;
if (a < INT_MIN - b) return INT_MIN;
return a + b;
}
clamp_add(100, 200) → "-2147483648", expected 300int clamp_add(int a, int b) {
if (a > INT_MAX - b) return INT_MAX;
if (a < INT_MIN - b) return INT_MIN;
return a + b;
}
no code in response Here is an example of a C function `clamp_add` that adds two integers while handling overflows correctly:
```cpp
#include <limits.h> // For INT_MAX and INT_MIN
int clamp_add(int a, int b) {
if (a > 0 && b > 0 && a + b < 0) {
return INT_MAX;
} else if (a < 0 && b < 0 && a + b > 0) {
return INT_MIN;
} else {
return a + b;
}
}
```
This function checks for overflow conditions and returns the appropriate values. If `a` and `b` are both positive and their sum is negative, it means that an overflow has occurred and `INT_MAX` is returned. Similarly, if `a` andno code in response Here is an example of a C function `clamp_add` that adds two integers while handling overflows correctly:
```cpp
#include <limits.h> // For INT_MAX and INT_MIN
int clamp_add(int a, int b) {
if (a > 0 && b > 0 && a + b < 0) {
return INT_MAX;
} else if (a < 0 && b < 0 && a + b > 0) {
return INT_MIN;
} else {
return a + b;
}
}
```
This function checks for overflow conditions and returns the appropriate values. If `a` and `b` are both positive and their sum is negative, it means that an overflow has occurred and `INT_MAX` is returned. Similarly, if `a` andcsv_escape — RFC 4180 CSV field quoting · 4 examplescsv_escape("say \"hi\"") → "\"say \"hi\"\"", expected "\"say \"\"hi\"\"\""std::string csv_escape(std::string field) {
bool needs_quoting = false;
for (char c : field) {
if (c == ',' || c == '"') {
needs_quoting = true;
break;
}
}
if (!needs_quoting) {
return field;
}
std::string result;
result.reserve(field.size() + 2);
result += '"';
for (char c : field) {
if (c == '"') {
result += "\"";
} else {
result += c;
}
}
result += '"';
return result;
}
compile error: pp
| ^
/tmp/bench_cpp_d_6lbomd/task.cpp:8:2: error: stray '`' in program
8 | ```cpp
| ^
/tmp/bench_cpp_d_6lbomd/task.cpp:8:3: error: stray '`' in pr#include <string>
std::string csv_escape(std::string field) {
// Determine whether quoting is required.
bool needQuote = false;
for (char c : field) {
if (c == ',' || c == '"') { // RFC 4180: quote if comma or double‑quote present
needQuote = true;
break;
}
}
// No quoting needed – return the original string unchanged.
if (!needQuote)
return field;
// Build quoted version, doubling interior quotes.
std::string result;
result.reserve(field.size() * 2 + 2); // worst case: every char is a quote
result.push_back('"');
for (char c : field) {
if (c == '"')
result.append("\"\""); // double the quote
else
result.push_back(c);
}
result.push_back('"');
return resultcompile error: pp
| ^
/tmp/bench_cpp_iblwew9h/task.cpp:8:2: error: stray '`' in program
8 | ```cpp
| ^
/tmp/bench_cpp_iblwew9h/task.cpp:8:3: error: stray '`' in pr#include <string>
std::string csv_escape(std::string field) {
// Determine whether quoting is required.
bool need_quotes = false;
for (char c : field) {
if (c == ',' || c == '"') { // comma or double‑quote triggers quoting
need_quotes = true;
break;
}
}
// If no quoting needed, return the original field unchanged.
if (!need_quotes)
compile error: = char; _Traits = std::char_traits<char>; _Alloc = std::allocator<char>; size_type = long unsigned int]'
2713 | find(_CharT __c, size_type __pos = 0) const _#include <string>
std::string csv_escape(const std::string& field) {
if (field.empty()) return "";
size_t comma_pos = field.find(',');
size_t quote_pos = field.find('"');
bool has_comma = (comma_pos != std::string::npos && comma_pos <= field.size() - 1);
bool has_quote = (quote_pos != std::string::npos && quote_pos <= field.size() - 2);
if (has_comma || has_quote) {
size_t start = 0;
std::string escaped;
while (start < field.size()) {
size_t end = field.find(',', start, 1);
if (end == std::string::npos) end = field.size();
size_t quote_end = field.find('"', start + 1);
if (quote_end != std::string::npos && quote_end <= end - 2) {
escaped += '"';
for (size_t i = start + 1; i < quote_end; ++i) {
char c = field[i];
if (c == '"') escaped += '"';
else escaped += c;
}
escaped += '"';
start = quote_end + 1;
} else {
size_t sub_start = end > start ? end : std::string::npos;
for (size_t i = start; i < field.size() && i <= sub_start; ++i) {
char c = field[i];
if (c == '"') escaped += '"';
else escaped += c;
}
start = sub_start + 1;
}
}
return '"' + escaped + '"';
} else {
return field;
}
}
interval_intersect — Intersect two interval lists · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def interval_intersect(a: list, b: list) -> list:
"""
Returns the intersection of two lists of closed intervals [start, end].
Each input list is sorted and its own intervals are disjoint.
Intervals that merely touch (one ends where the next begins) DO intersect
and the shared point is a valid result interval. The results are returned in order.
Args:
a: List of closed intervals for set A, e.g., [[s1, e1], [s2, e2], ...].
b: List of closed intervals for set B, e.g., [[s'1, e'1], [s'2, e'2], ...].
Returns:
List of intersecting closed intervals, sorted by start point.
"""
intersections = []
i = 0 # Pointer for list a
j = 0 # Pointer for list b
while i < len(a) and j < len(b):
interval_a = a[i]
interval_b = b[j]
start_a, end_a = interval_a
start_b, end_b = interval_b
# Calculate the intersection of [start_a, end_a] and [start_b, end_b]
# The intersection starts at the maximum of the two starts.
overlap_start = max(start_a, start_b)
# The intersection ends at the minimum of the two ends.
overlap_end = min(end_a, end_b)
# An intersection exists if the calculated start is less than or equal to the calculated end.
# Since intervals are closed, touching counts as an intersection (e.g., [1, 2] and [2, 3] intersect at [2, 2]).
if overlap_start <= overlap_end:
intersections.append([overlap_start, overlap_end])
# Advance the pointer associated with the interval that finishes earliest.
… (truncated)SyntaxError: invalid syntax (<model>, line 1)def interval_intersect(a: list[list[int]], b: list[list[int]]) -> list[list[int]]:
"""
Return the intersection of two sorted, disjoint lists of closed intervals.
Intervals that merely touch (e.g., [1, 3] and [3, 5]) are considered intersecting,
producing a single‑point interval [3, 3].
Parameters
----------
a, b : list of [start, end]
Each list is sorted by start coordinate and contains non‑overlapping intervals.
Returns
-------
list of [start, end]
Intersections in ascending order.
"""
i = j = 0
result: list[list[int]] = []
while i < len(a) and j < len(b):
# Compute overlap between a[i] and b[j]
start = max(a[i][0], b[j][0])
end = min(a[i][1], b[j][1])
if start <= end: SyntaxError: invalid syntax (<model>, line 1)def interval_intersect(a: list, b: list) -> list:
"""
Return the intersection of two sorted, internally disjoint lists of closed intervals.
Each interval is a two‑element list or tuple [start, end] (inclusive).
Intervals that only touch at an endpoint are considered intersecting,
and the shared point forms a valid result interval.
Parameters
----------
a : list of [int|float, int|float]
First list of intervals, sorted by start coordinate.
b : list of [int|float, int|float]
Second list of intervals, sorted by start coordinate.
Returns
-------
list of [int|float, int|float]
List of intersecting intervals in ascending order.
"""
i = j = 0
result = []
while i < len(a) and j < len(b):
a_start, a_end = a[i]
b_start, b_end = b[j]
# Overlap (including touching endpoints)
lo = max(a_start, b_start)
hi = min(a_end, b_end)
if lo <= hi: interval_intersect([[0, 2], [5, 10]], [[1, 5], [8, 12]]) → [[1, 2]], expected [[1, 2], [5, 5], [8, 10]]def interval_intersect(a: list, b: list) -> list:
i = j = 0
res = []
while i < len(a) and j < len(b):
start = max(a[i][0], b[j][0])
end = min(a[i][1], b[j][1])
if start <= end:
res.append([start, end])
if a[i][1] > b[j][1]:
i += 1
else:
j += 1
return res
glob_match — Glob matching with ? and * · 4 examplesSyntaxError: invalid syntax (<model>, line 1)def glob_match(pattern: str, text: str) -> bool:
"""
Return True if the whole `text` matches the wildcard `pattern`.
'?' – matches exactly one character.
'*' – matches any sequence of characters (including empty).
All other characters match themselves.
"""
m, n = len(pattern), len(text)
# dp[i][j] == pattern[:i] matches text[:j]
dp = [[False] * (n + 1) for _ in range(m + 1)]
dp[0][0] = True
# Patterns that consist only of '*' can match an empty string.
for i in range(1, m + 1):
if pattern[i - 1SyntaxError: invalid syntax (<model>, line 1)def glob_match(pattern: str, text: str) -> bool:
"""
Return True if the whole `text` matches the glob `pattern`.
'?' – matches exactly one character.
'*' – matches any sequence of characters (including empty).
All other characters match themselves.
"""
p_len, t_len = len(pattern), len(text)
# dp[j] == pattern[:i] matches text[:j]
dp = [False] * (t_len + 1)
dp[0] = True # empty pattern matches empty text
for i in range(1, p_len + 1):
pcglob_match("a?c", "ac") → null, expected falsedef glob_match(pattern: str, text: str) -> bool:
def match(text_index, pattern_index):
if text_index == len(text) and pattern_index == len(pattern):
return True # Both strings are fully matched.
elif text_index != len(text) and pattern_index == len(pattern):
return False # Pattern is exhausted but there's still text left.
elif text_index == len(text) and pattern[pattern_index] == '*':
return match(text_index, pattern_index + 1) # Text is exhausted; skip asterisks.
else:
if (pattern[pattern_index] == '?' and text_index < len(text)) or (pattern[pattern_index] == text[text_index]):
return match(text_index + 1, pattern_index + 1) # Matched a character or question mark.
elif pattern[pattern_index] == '*':
return match(text_index, pattern_index + 1) or match(text_index + 1, pattern_index) # Skip the asterisk or skip both.
else:
return False # Mismatch without wildcards.
return match(0, 0)
glob_match("a?c", "ac") → null, expected falsedef glob_match(pattern: str, text: str) -> bool:
def match(text_index, pattern_index):
if text_index == len(text) and pattern_index == len(pattern):
return True # Both strings are fully matched.
elif text_index != len(text) and pattern_index == len(pattern):
return False # Pattern is exhausted but there's still text left.
elif text_index == len(text) and pattern[pattern_index] == '*':
return match(text_index, pattern_index + 1) # Text is exhausted; skip asterisks.
else:
if (pattern[pattern_index] == '?' and text_index < len(text)) or (pattern[pattern_index] == text[text_index]):
return match(text_index + 1, pattern_index + 1) # Matched a character or question mark.
elif pattern[pattern_index] == '*':
return match(text_index, pattern_index + 1) or match(text_index + 1, pattern_index) # Skip the asterisk or skip both.
else:
return False # Mismatched characters.
return match(0, 0)
count_words — Count words split on spaces and tabs · 4 examplesno code in response Here is the C function to count words in a given string:
```cpp
#include <stdio.h>
#include <ctype.h>
int count_words(const char *s) {
int wordCount = 0;
while (*s != '\0') {
// Skip leading separators (spaces and tabs)
while (*s == ' ' || *s == '\t') {
s++;
}
if (*s != '\0') {
wordCount++;
// Skip the current word
while (*s != ' ' && *s != '\t' && *s != '\0') {
s++;
}
}
}
return wordCount;
}
```
This function iterates through theno code in response Here is the C function to count words in a given string:
```cpp
#include <stdio.h>
#include <ctype.h>
int count_words(const char *s) {
int wordCount = 0;
while (*s != '\0') {
// Skip leading separators (spaces and tabs)
while (*s == ' ' || *s == '\t') {
s++;
}
if (*s != '\0') {
wordCount++;
// Skip the current word
while (*s != ' ' && *s != '\t' && *s != '\0') {
s++;
}
}
}
return wordCount;
}
```
This function iterates through thecount_words("hello world") → "12", expected 2int count_words(const char *s) {
if (s == NULL || s[0] == '\0') return 0;
int count = 0, i = 0;
while (s[i]) {
if (!isspace(s[i])) {
if (i > 0 && isspace(s[i-1])) count++;
count++;
}
i++;
}
return count + (s[0] != '\0');
}
no code in responseWe are counting words in a string s. A word is a maximal run of non-space and non-tab characters. We note: - Leading, trailing, and repeated separators (spaces or tabs) are legal and do not count as words. - We return 0 for an empty string or a string that is all separators. Approach: We traverse the string and skip any leading separators (spaces or tabs). Then, we look for a run of non-separator characters (which is a word). Each such run is one word. We then skip the trailing separators. However, note: the problem says "maximal run of characters that are not spaces or ta
theme_vars — Dark-mode token inside a media query · 4 examplesno code in response1. Write CSS implementing a two-theme token: :root defines --ink as #111111; inside an @media (prefers-color-scheme: dark) block, :root redefines --ink as #eeeeee; and body sets color to var(--ink). The trap is scope — the dark value must live inside the media query, not beside it. Return only the CSS in a single ```css code block.
no code in response1. Write CSS implementing a two-theme token: :root defines --ink as #111111; inside an @media (prefers-color-scheme: dark) block, :root redefines --ink as #eeeeee; and body sets color to var(--ink). The trap is scope — the dark value must live inside the media query, not beside it. Return only the CSS in a single ```css code block.
decl → null, expected "var(--ink)":root {
--ink: #111111;
}
@media (prefers-color-scheme: dark) {
:root {
--ink: #eeeeee;
}
}
body {
--ink: #eeeeee;
}
decl → null, expected "var(--ink)":root {
--ink: #111111;
}
@media (prefers-color-scheme: dark) {
:root {
--ink: #eeeeee;
}
}
body {
--ink: #eeeeee;
}
data_table — Revenue table with caption and scoped headers · 4 examplescount of "thead th" → 2, expected 3<table>
<caption>Quarterly revenue</caption>
<thead>
<tr scope="col">Q1</scope>
<th scope="col">Q2</th>
<th scope="col">Q3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Revenue: $50,000</td>
<td>Revenue: $65,000</td>
<td>Revenue: $72,000</td>
</tr>
<tr>
<td>Growth: 10%</td>
<td>Growth: 8%</td>
<td>Growth: 4%</td>
</tr>
</tbody>
</table>
count of "thead th" → 2, expected 3<table>
<caption>Quarterly revenue</caption>
<thead>
<tr scope="col">Q1</scope>
<th scope="col">Q2</th>
<th scope="col">Q3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Revenue: $50,000</td>
<td>Revenue: $65,000</td>
<td>Revenue: $72,000</td>
</tr>
<tr>
<td>Growth: 10%</td>
<td>Growth: 8%</td>
<td>Growth: 4%</td>
</tr>
</tbody>
</table>
count of "thead th" → 4, expected 3We are to create an HTML fragment for a table with:
- A caption: exactly 'Quarterly revenue'
- A thead with a single row containing three th cells, each with scope="col"
- A tbody with exactly two rows, each row having three td cells.
We'll choose some example data for the cells.
Steps:
1. Start with the table tag: <table>
2. Add the caption: <caption>Quarterly revenue</caption>
3. Add the thead: <thead><tr><th scope="col">...</th><th scope="col">...</th><th scope="col">...</th></tr></thead>
4. Add the tbody: <tbody><tr><td>...</td><td>...</td><td>...</td></tr><tr><td>...</td><td>...</td><td>...</td></tr></tbody>
We have to choose three column headers and two rows of data.
Example:
Columns: Q1, Q2, Q3, Q4? But note: the problem says three columns. So let's do three quarters?
However, the problem says "quarterly revenue", so we can have three quarters? But typically quarterly revenue would be four quarters.
But the problem says: three <th> cells. So we have three columns.
Let's choose:
Column 1: "Q1"
Column 2: "Q2"
Column 3: "Q3"
And two rows of data (for two years? or two quarters? but we have three columns so we'll do two rows of three numbers)
Example data:
Row 1: $100,000, $150,000, $200,000
Row 2: $120,000, $160,000, $220,000
However, the problem says "cell contents are yours to choose". So we can pick any.
Let's write:
<caption>Quarterly revenue</caption>
<thead>
<tr>
<th scope="col">Q1</th>
<th scope="col">Q2</th>
<th scope="col">Q3</th>
</tr>
</thead>
<tbody>
<tr>
… (truncated)count of "table" → 2, expected 1 <table>
<caption>Quarterly revenue</caption>
<thead>
<tr>
<th scope="col">Q1</th>
<th scope="col">Q2</th>
<th scope="col">Q3</th>
</tr>
</thead>
<tbody>
<tr>
<td>$100balanced_depth — Max bracket nesting depth, -1 if unbalanced · 4 examplesno code in responseWe are to compute the maximum nesting depth of round brackets in the string s.
Steps:
1. We traverse the string and ignore non-bracket characters.
2. We maintain a counter for the current depth (starting at 0).
3. When we see an opening bracket '(', we increase the counter by 1.
4. When we see a closing bracket ')', we decrease the counter by 1.
5. However, if at any point the counter becomes negative, that means there's a closing bracket without an opener -> not balanced -> return -1.
6. Also, at the end, if the counter is not zero, then there are unclosed openers -> not balanced -> rebalanced_depth("(a(b)c)") → "-1", expected 2#include <string>
#include <cctype>
int balanced_depth(const std::string& s) {
int depth = 0;
for (char c : s) {
if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
} else {
continue;
}
}
return depth == 0 ? -1 : depth;
}
balanced_depth("(a(b)c)") → "-1", expected 2#include <string>
#include <cctype>
int balanced_depth(const std::string& s) {
int depth = 0;
for (char c : s) {
if (c == '(') {
depth++;
} else if (c == ')') {
depth--;
} else {
continue;
}
}
return depth == 0 ? -1 : depth;
}
no code in responseWe are to compute the maximum nesting depth of round brackets in the string s.
Steps:
1. We traverse the string.
2. We ignore non-bracket characters.
3. We maintain a counter for the current depth (starting at 0).
4. When we see an opening bracket '(', we increase the counter by 1.
5. When we see a closing bracket ')', we decrease the counter by 1.
6. However, if at any point the counter becomes negative, that means there's a closing bracket without an opener -> not balanced -> return -1.
7. Also, at the end, if the counter is not 0, then there are unclosed openers -> return -1.
But card_grid — Responsive auto-fill card grid · 2 examplesdecl → null, expected "translateY(-2px)".cards {
display: grid;
grid-template-columns: repeat(auto-fill, minmax(240px, 1fr));
gap: 16px;
}
.card {
border-radius: 8px;
transition: transform 0.3s;
}
decl → null, expected "translateY(-2px)".cards {
display: grid;
grid-template-columns: repeat(auto-fill, minmax(240px, 1fr));
gap: 16px;
}
.card {
border-radius: 8px;
transition: transform 0.3s;
cursor: pointer;
}
GPU
NVIDIA GB10 · 122 GB
Unified memory
122 GB
CPU
Cortex-X925
Cores
20
OS
Ubuntu 24.04.4 LTS (aarch64)
Kernel
6.17.0-1018-nvidia
Proxy
0.2.0
CPU, OS and kernel read from the host at report time — this run predates their capture, and they do not change between runs.
Every task asks for one answer in the task's language — Python, JavaScript under node, C and C++ under gcc, Rust, C#, or PHP, each run in a separate process under a timeout with the return value compared against the expected one. HTML and CSS tasks are graded structurally: the answer is parsed and checked against required structure (bindings, attributes, declarations in the right context) — a claim about the markup, not about how a browser renders it. Tasks whose toolchain is absent on the machine are skipped and listed here, never scored as zero. Fully correct counts only responses where every case for that task passed; cases is the share of individual cases that passed, so a near-miss still scores there. A response with no extractable code block scores zero — that measures instruction-following, not coding.
Suite
coding-v2
Tasks
29
Cases
194
Repeats
3 per task
Languages
9
Core — 15 tasks, 97 cases
group_ranges Collapse consecutive integers into range strings · jsclamp_add Saturating int addition — overflow trap · ccount_words Count words split on spaces and tabs · ccsv_escape RFC 4180 CSV field quoting · cppsnake_to_camel snake_case to camelCase — digits stop capitalisation · rustclamp_mul Saturating int multiplication · csharpslugify URL slug: symbol runs become one hyphen · phplogin_form Login form with labels bound to their inputs · htmlcard_grid Responsive auto-fill card grid · csssemver_cmp Semantic versions incl. pre-release precedencecsv_line Split one CSV record honouring quoteslru_ops LRU cache with eviction orderpath_norm Normalise a POSIX path with . and ..base_convert Integer between bases 2-36 with validationinterval_intersect Intersect two interval listsHard — 14 tasks, 97 cases
parse_query Parse a URL query string into an object · jsround_to Round to nearest multiple, halves away from zero · cbalanced_depth Max bracket nesting depth, -1 if unbalanced · cppmid_floor Floor midpoint of two i64s — overflow and negatives · rustordinal English ordinal suffix — the 11th/12th/13th trap · csharppluck Column from associative rows — null vs missing key · phpdata_table Revenue table with caption and scoped headers · htmltheme_vars Dark-mode token inside a media query · cssglob_match Glob matching with ? and *roman_strict Roman to int, rejecting non-canonical formstopo_lex Smallest topological order, None on cyclejustify Full text justificationjson_pointer Resolve an RFC 6901 JSON Pointertokenize_expr Tokenise arithmetic, None on invalid inputx-client-name: ai-proxy-bench.Time for the discarded warm-up request — the price of making the model resident, excluded from every measurement above.
| Configuration | Cold start |
|---|---|
| devstral-2:123b · 69.8 GB · ollama · 4 | 496.7 s |
| devstral-2:123b · 69.8 GB · ollama · 1 | 489.5 s |
| llama4 · 62.8 GB · ollama · 1 | 194.0 s |
| gpt-oss:120b · 60.9 GB · ollama · 1 | 145.8 s |
| llama3:70b-instruct · 37.2 GB · ollama · 1 | 144.9 s |
| codellama:70b · 36.2 GB · ollama · 1 | 107.4 s |
| llama3:70b-instruct · ollama · 4 | 94.2 s |
| qwen3-coder-next · 48.2 GB · ollama · 1 | 73.2 s |
| qwen3-coder:tuned · 48.2 GB · ollama · 1 | 66.5 s |
| qwen3.6:27b · 16.2 GB · ollama · 1 | 64.0 s |
+ 30 more under 64.0 s.
Screen-only companion to the annotated chart above: hover a dot for its numbers, drag a box to zoom into the crowded band, double-click to reset. Hovering a model anywhere on this page highlights it everywhere. The printed report keeps the annotated version.