Foundation + Eval Harness
Build the basic research infrastructure that can measure model behavior reliably before changing it.
Open Model Lab
A public engineering notebook for measuring open-model behavior across evals, post-training, agents, safety, monitorability, and systems work.
Open Model Lab is a public engineering notebook for measuring open-model behavior across evals, post-training, agents, safety, and systems work. The July 2026 baseline is frozen; later experiments are compared against it rather than rewriting it.
The July Foundation + Eval Harness gate completed on July 29 with a frozen 25-task suite, three comparable local model runs, 75 validated scored outputs, deterministic reporting, and a manual audit of all 28 recorded failures. August SFT and data-quality work is planned but has not been marked active yet.
Build the basic research infrastructure that can measure model behavior reliably before changing it.
Measure the behavioral difference between a base model and an instruction-tuned model.
Use preference data to improve SFT behavior in a more controlled way, then measure behavioral side effects.
Move the model from passive question-answering into a tool-using agent setup.
Measure agent success and failure taxonomy on multi-step tasks.
Measure safety behavior by quality and balance, not just refusal rate.
Evaluate not only final answers, but where reasoning processes break down.
Start analyzing model behavior with internal signals in addition to external scores.
Enter multimodal model work through screenshot and UI-understanding tasks.
Add systems and profiling knowledge to make research experiments more efficient.
Measure whether better data mixtures produce better behavior with the same compute.
Turn the 12-month work into a showable, reproducible, publishable portfolio.
Reproducible task execution, grading, metrics, and failure labels.
Instruction data cleaning, supervised fine-tuning, and behavior comparison.
Chosen/rejected pairs, DPO training, and side-effect measurement.
Tool registry, file operations, test execution, and replayable traces.
Multi-step coding tasks and trace-level failure taxonomy.
Refusal quality, over-refusal, under-refusal, jailbreak robustness, and safe completion.
Outcome scoring compared with observable process and self-correction signals.
Internal-signal probes for failure prediction and representation drift.
Screenshot QA, OCR plus reasoning, UI grounding, and visual failure modes.
Latency, throughput, batching, KV cache, quantization, and profiling.
Data mixtures, filtering, small scaling ladders, and score/GPU-hour.
Reproducible scripts, reports, dashboard, README, and portfolio packaging.
The completed July milestone report, comparison, limitations, and August handoff.
A versioned 25-task diagnostic suite across five categories; all task quality labels remain draft.
A frozen identity for the dataset, configs, graders, reports, and three local runs.
A manual review of all 28 recorded failures without rewriting the original scores.
The full July 2026 through June 2027 research-engineering plan.
A compact month-by-month view of gates, themes, and report targets.
Detailed monthly gates, focus areas, expected outputs, and decisions.
Planned and published report pages with explicit claim boundaries.
A compact reading and writing template for Open Model Lab reports.
Task schema, initial categories, graders, metrics, and failure taxonomy.
Published summaries for the three frozen July diagnostic runs.
Recorded identities and caveats for models used in published evaluations.
The July diagnostic eval suite plus planned training and later evaluation datasets.
Planned dashboards. No placeholder charts or synthetic metrics.
Public decision log for naming, eval-first scope, and claim boundaries.
Concise definitions for terms used throughout the lab section.