Journal
#llm-evals
Why I Froze an Imperfect Eval Baseline
Jul 31, 2026 · note
A follow-up note on why the July Open Model Lab baseline stays frozen even after the failure audit found grader artifacts, strict-format misses, and token-limit effects.
Open Model Research Harness completes its July evaluation gate
Jul 29, 2026 · note
The first Open Model Research Harness gate is complete: 25 tasks, three local open models, 75 scored outputs, a deterministic comparison, and a failure audit that shows why raw pass rates need context.
Starting Open Model Research Harness
Jul 4, 2026 · note
A public 12-month project to build an eval-first research-engineering harness for open LLMs.