Journal

#model-behavior

3 posts

Why I Froze an Imperfect Eval Baseline

Jul 31, 2026 · note

A follow-up note on why the July Open Model Lab baseline stays frozen even after the failure audit found grader artifacts, strict-format misses, and token-limit effects.

Open Model Research Harness completes its July evaluation gate

Jul 29, 2026 · note

The first Open Model Research Harness gate is complete: 25 tasks, three local open models, 75 scored outputs, a deterministic comparison, and a failure audit that shows why raw pass rates need context.

Starting Open Model Research Harness

Jul 4, 2026 · note

A public 12-month project to build an eval-first research-engineering harness for open LLMs.