Why I Froze an Imperfect Eval Baseline

The July 2026 Open Model Lab baseline is imperfect. That is why I froze it.

The first Open Model Research Harness gate produced the shape I needed: 25 versioned tasks, three comparable local model runs, 75 validated scored outputs, deterministic comparison reports, and a manual audit of all 28 recorded failures. It also exposed problems in the measurement protocol itself. Some failures were caused by strict answer formats. Some were likely grader false negatives. Some were tied to the fixed 1024-token limit and runtime output handling.

That is exactly the point at which it becomes tempting to “clean up” the results. Fix the extractor. Increase the token limit. Relax the JSON check. Patch the rubric. Re-run the models. Replace the awkward score table with the one that better matches the manual audit.

I am not doing that to the July baseline.

Baseline pieceJuly stateStatusRule for later work
Evaluation suite25 draft tasks across 5 categoriesFrozenTask fixes require a new eval identity.
Run configurationTemperature 0.0, seed 42, 1024 output tokensFrozenGeneration-limit changes cannot rewrite July scores.
Recorded scores75 scored outputs, 28 failuresFrozenManual audit can qualify scores, not silently replace them.
Failure audit28 failed rows reviewedDoneKeep audit labels separate from grader outcomes.
August comparisonNot startedPlannedCompare against the frozen baseline or declare a new baseline.

A baseline is a measurement contract

An eval baseline is not just a score table. It is a contract about the exact dataset, model identity, runtime settings, grader behavior, report code, and known caveats used for a comparison.

If I change one of those after seeing the results, the old score and the new score no longer share the same meaning. The change may be justified. It may even make the measurement better. But it also creates a new measurement identity.

That distinction matters more than cosmetic accuracy. A flawed baseline with an honest identity is useful. A tidied baseline with invisible edits is harder to trust.

The failure audit changed the interpretation

The July report recorded three local Ollama runs under one shared configuration:

ModelPassedFailedPass rateMean score
lfm2.5-8b121348.0%0.452
gpt-oss23292.0%0.920
qwen3.6121348.0%0.540

Those numbers are real recorded outcomes for that suite. They are not a general model ranking.

The audit made the table less convenient and more useful. It classified the 28 failed rows like this:

Audit labelCountWhat it means
Likely format or instruction failure11The model output probably violated the requested interface or final-answer shape.
Likely grader false negative7The grader likely rejected an output that was usable or semantically correct.
Truncation or runtime related10The failure was materially affected by the fixed output limit or runtime output handling.
Likely substantive model failure0No failed row was labeled as clearly wrong model reasoning under this audit.

That last row is not a capability claim. It does not mean the models made no mistakes. It means the first run mostly surfaced protocol and measurement issues before it could support stronger conclusions about model behavior.

Why not repair the July results?

Because the repairs are exactly the things August and later gates need to measure deliberately.

If I improve code extraction, that is a grader change. If I increase the token limit, that is a run-configuration change. If I normalize Markdown-fenced JSON, that is an interface-tolerance change. If I revise a textual rubric after seeing a borderline output, that is a grader-policy change.

Each of those may be the right engineering move. None should mutate the July baseline in place.

The rule I want for this project is simple:

  1. Recorded grader outcomes remain recorded.
  2. Manual audit labels explain outcomes without overwriting them.
  3. Any material task, grader, runtime, or model-identity change gets a new evaluation identity.
  4. Later reports compare against July only when the comparison contract still holds.
  5. If the contract no longer holds, the report says so directly.

This is slower than polishing the score table. It is also less misleading.

What this means for August

The next Open Model Lab gate is SFT Pipeline + Data Quality. The question is not “can I make the August numbers look better than July?” The question is whether a controlled supervised fine-tuning experiment produces measurable behavior changes and regressions against a known baseline.

The July baseline already tells me where the measuring system is weak:

August can improve the harness. But if those improvements change the comparison contract, the result should be labeled as a new eval identity rather than a clean continuation of July.

That is the practical value of freezing an imperfect baseline. It keeps the research record from becoming a moving target.

The completed July gate, published report, run summaries, dataset record, and failure audit are linked from Open Model Lab. The implementation and frozen baseline live in the public Open Model Research Harness repository.