Open Model Lab

Datasets

The frozen July baseline includes a public 25-task diagnostic suite; later instruction, preference, agent, and safety datasets remain planned.

Purpose

The project needs dataset cards because behavior changes cannot be interpreted without knowing which data was used for evals, SFT, DPO, agent tasks, or safety checks.

Eval leakage control is mandatory: training data must not contain eval tasks.

Published evaluation dataset

Name
july_eval_v1
Version
v1
Status
Published diagnostic baseline
Tasks
25
Categories
5, balanced at 5 tasks each
Quality boundary
All 25 task records are marked draft

The suite is a small diagnostic baseline, not a statistically representative benchmark. Its content hash is 38c646401b74f92deac04c5e225dfeabd822f72c3e30ce52dfaa24d817375760.

Planned later datasets

  • instruction_v1
  • instruction_v2_filtered
  • preference_v1
  • agent_tasks_v1
  • safety_eval_v1

Dataset card fields

  • name
  • purpose
  • source
  • license/usage constraints
  • size
  • categories
  • filtering steps
  • deduplication
  • contamination/leakage checks
  • risks
  • version