Open Model Lab

Open Model Lab

A public engineering notebook for measuring open-model behavior across evals, post-training, agents, safety, monitorability, and systems work.

Status facts

Project
Open Model Research Harness
Repository
GitHub
Timeframe
July 2026 - June 2027
Current status
July gate completed; August planned
Latest completed gate
July 2026 eval harness completed
Next gate
August 2026 SFT + data quality (planned)
Claims
Diagnostic baseline only; no general model-quality claim

July gate completed; August is next

Open Model Lab is a public engineering notebook for measuring open-model behavior across evals, post-training, agents, safety, and systems work. The July 2026 baseline is frozen; later experiments are compared against it rather than rewriting it.

The July Foundation + Eval Harness gate completed on July 29 with a frozen 25-task suite, three comparable local model runs, 75 validated scored outputs, deterministic reporting, and a manual audit of all 28 recorded failures. August SFT and data-quality work is planned but has not been marked active yet.

12-month timeline

July 2026

Foundation + Eval Harness

Build the basic research infrastructure that can measure model behavior reliably before changing it.

completed
August 2026

SFT Pipeline + Data Quality

Measure the behavioral difference between a base model and an instruction-tuned model.

planned
September 2026

Preference Optimization / DPO

Use preference data to improve SFT behavior in a more controlled way, then measure behavioral side effects.

planned
October 2026

Agent Harness / Tool Use

Move the model from passive question-answering into a tool-using agent setup.

planned

Research modules

completed

Eval Harness

Reproducible task execution, grading, metrics, and failure labels.

planned

SFT Pipeline

Instruction data cleaning, supervised fine-tuning, and behavior comparison.

planned

Preference Optimization / DPO

Chosen/rejected pairs, DPO training, and side-effect measurement.

planned

Agent Harness

Tool registry, file operations, test execution, and replayable traces.

planned

Long-Horizon Agent Evals

Multi-step coding tasks and trace-level failure taxonomy.

planned

Safety / Red Teaming

Refusal quality, over-refusal, under-refusal, jailbreak robustness, and safe completion.

planned

Reasoning Process Evaluation

Outcome scoring compared with observable process and self-correction signals.

planned

Monitorability

Internal-signal probes for failure prediction and representation drift.

planned

Multimodal UI Understanding

Screenshot QA, OCR plus reasoning, UI grounding, and visual failure modes.

planned

Systems Efficiency

Latency, throughput, batching, KV cache, quantization, and profiling.

planned

Data Efficiency / Scaling

Data mixtures, filtering, small scaling ladders, and score/GPU-hour.

planned

Final Integration

Reproducible scripts, reports, dashboard, README, and portfolio packaging.

Latest public artifacts

Published

July evaluation suite

A versioned 25-task diagnostic suite across five categories; all task quality labels remain draft.

Published

July baseline

A frozen identity for the dataset, configs, graders, reports, and three local runs.

Published

Failure audit

A manual review of all 28 recorded failures without rewriting the original scores.

Claim boundary

Diagnostic baseline, not a leaderboard. The July results apply only to a small 25-task suite on one local runtime. They do not establish general model quality or statistical significance. Planned later reports and dashboards remain empty until their own evidence exists.

12-month plan

The full July 2026 through June 2027 research-engineering plan.

Timeline

A compact month-by-month view of gates, themes, and report targets.

Months

Detailed monthly gates, focus areas, expected outputs, and decisions.

Reports

Planned and published report pages with explicit claim boundaries.

Report template

A compact reading and writing template for Open Model Lab reports.

Evals

Task schema, initial categories, graders, metrics, and failure taxonomy.

Runs

Published summaries for the three frozen July diagnostic runs.

Models

Recorded identities and caveats for models used in published evaluations.

Datasets

The July diagnostic eval suite plus planned training and later evaluation datasets.

Dashboards

Planned dashboards. No placeholder charts or synthetic metrics.

Decisions

Public decision log for naming, eval-first scope, and claim boundaries.

Glossary

Concise definitions for terms used throughout the lab section.