2026-07-final-lfm2-5-8b
Ollama: lfm2.5:8b (9cf756159fc2)
Caveat: Visible thinking text affected strict formats; three coding failures were likely extraction false negatives.
Open Model Lab
Three local open models completed the same frozen 25-task July suite, producing 75 validated scored outputs with zero grader errors.
| Model | Passed | Pass rate | Mean score | Avg latency |
|---|---|---|---|---|
| lfm2.5-8b | 12/25 | 48.0% | 0.452 | 1,639.48 ms |
| gpt-oss | 23/25 | 92.0% | 0.920 | 3,597.04 ms |
| qwen3.6 | 12/25 | 48.0% | 0.540 | 11,112.48 ms |
The shared configuration used temperature 0.0, a 1024-token output limit, seed 42, and a 300-second timeout. Results apply only to this diagnostic suite.
Ollama: lfm2.5:8b (9cf756159fc2)
Caveat: Visible thinking text affected strict formats; three coding failures were likely extraction false negatives.
Ollama: gpt-oss:latest (17052f91a42e)
Caveat: One strict-format miss and one likely literal-rubric false negative were recorded.
Ollama: qwen3.6:latest (07d35212591f)
Caveat: Ten failures reached the fixed 1024-token limit; thinking metadata affected final-output completeness.