Open Model Lab

Runs

Three local open models completed the same frozen 25-task July suite, producing 75 validated scored outputs with zero grader errors.

Status

July baseline published. These are public summaries of local run artifacts. Raw task-level outputs remain excluded from version control; run identities, configurations, aggregate results, reports, and baseline hashes are publicly recorded.

July 2026 final runs

Model Passed Pass rate Mean score Avg latency
lfm2.5-8b 12/25 48.0% 0.452 1,639.48 ms
gpt-oss 23/25 92.0% 0.920 3,597.04 ms
qwen3.6 12/25 48.0% 0.540 11,112.48 ms

The shared configuration used temperature 0.0, a 1024-token output limit, seed 42, and a 300-second timeout. Results apply only to this diagnostic suite.

Run identities and caveats

validated

2026-07-final-lfm2-5-8b

Ollama: lfm2.5:8b (9cf756159fc2)

Caveat: Visible thinking text affected strict formats; three coding failures were likely extraction false negatives.

validated

2026-07-final-gpt-oss

Ollama: gpt-oss:latest (17052f91a42e)

Caveat: One strict-format miss and one likely literal-rubric false negative were recorded.

validated

2026-07-final-qwen3-6

Ollama: qwen3.6:latest (07d35212591f)

Caveat: Ten failures reached the fixed 1024-token limit; thinking metadata affected final-output completeness.

Run record fields

  • run id
  • date
  • model
  • model variant
  • eval suite
  • dataset version
  • config hash
  • score
  • cost
  • latency
  • failure mode distribution
  • known caveats
  • report link

Evidence