Why I Froze an Imperfect Eval Baseline
The July 2026 Open Model Lab baseline is imperfect. That is why I froze it.
The first Open Model Research Harness gate produced the shape I needed: 25 versioned tasks, three comparable local model runs, 75 validated scored outputs, deterministic comparison reports, and a manual audit of all 28 recorded failures. It also exposed problems in the measurement protocol itself. Some failures were caused by strict answer formats. Some were likely grader false negatives. Some were tied to the fixed 1024-token limit and runtime output handling.
That is exactly the point at which it becomes tempting to “clean up” the results. Fix the extractor. Increase the token limit. Relax the JSON check. Patch the rubric. Re-run the models. Replace the awkward score table with the one that better matches the manual audit.
I am not doing that to the July baseline.
A baseline is a measurement contract
An eval baseline is not just a score table. It is a contract about the exact dataset, model identity, runtime settings, grader behavior, report code, and known caveats used for a comparison.
If I change one of those after seeing the results, the old score and the new score no longer share the same meaning. The change may be justified. It may even make the measurement better. But it also creates a new measurement identity.
That distinction matters more than cosmetic accuracy. A flawed baseline with an honest identity is useful. A tidied baseline with invisible edits is harder to trust.
The failure audit changed the interpretation
The July report recorded three local Ollama runs under one shared configuration:
| Model | Passed | Failed | Pass rate | Mean score |
|---|---|---|---|---|
lfm2.5-8b | 12 | 13 | 48.0% | 0.452 |
gpt-oss | 23 | 2 | 92.0% | 0.920 |
qwen3.6 | 12 | 13 | 48.0% | 0.540 |
Those numbers are real recorded outcomes for that suite. They are not a general model ranking.
The audit made the table less convenient and more useful. It classified the 28 failed rows like this:
| Audit label | Count | What it means |
|---|---|---|
| Likely format or instruction failure | 11 | The model output probably violated the requested interface or final-answer shape. |
| Likely grader false negative | 7 | The grader likely rejected an output that was usable or semantically correct. |
| Truncation or runtime related | 10 | The failure was materially affected by the fixed output limit or runtime output handling. |
| Likely substantive model failure | 0 | No failed row was labeled as clearly wrong model reasoning under this audit. |
That last row is not a capability claim. It does not mean the models made no mistakes. It means the first run mostly surfaced protocol and measurement issues before it could support stronger conclusions about model behavior.
Why not repair the July results?
Because the repairs are exactly the things August and later gates need to measure deliberately.
If I improve code extraction, that is a grader change. If I increase the token limit, that is a run-configuration change. If I normalize Markdown-fenced JSON, that is an interface-tolerance change. If I revise a textual rubric after seeing a borderline output, that is a grader-policy change.
Each of those may be the right engineering move. None should mutate the July baseline in place.
The rule I want for this project is simple:
- Recorded grader outcomes remain recorded.
- Manual audit labels explain outcomes without overwriting them.
- Any material task, grader, runtime, or model-identity change gets a new evaluation identity.
- Later reports compare against July only when the comparison contract still holds.
- If the contract no longer holds, the report says so directly.
This is slower than polishing the score table. It is also less misleading.
What this means for August
The next Open Model Lab gate is SFT Pipeline + Data Quality. The question is not “can I make the August numbers look better than July?” The question is whether a controlled supervised fine-tuning experiment produces measurable behavior changes and regressions against a known baseline.
The July baseline already tells me where the measuring system is weak:
- strict final-answer formats need clearer handling;
- code extraction needs better evidence preservation;
- thinking text and final-answer channels need explicit policy;
- output-token limits can dominate some results;
- grader false negatives need to be auditable instead of hidden.
August can improve the harness. But if those improvements change the comparison contract, the result should be labeled as a new eval identity rather than a clean continuation of July.
That is the practical value of freezing an imperfect baseline. It keeps the research record from becoming a moving target.
The completed July gate, published report, run summaries, dataset record, and failure audit are linked from Open Model Lab. The implementation and frozen baseline live in the public Open Model Research Harness repository.