Open Model Lab
Datasets
The frozen July baseline includes a public 25-task diagnostic suite; later instruction, preference, agent, and safety datasets remain planned.
Purpose
The project needs dataset cards because behavior changes cannot be interpreted without knowing which data was used for evals, SFT, DPO, agent tasks, or safety checks.
Eval leakage control is mandatory: training data must not contain eval tasks.
Published evaluation dataset
- Name
- july_eval_v1
- Version
- v1
- Status
- Published diagnostic baseline
- Tasks
- 25
- Categories
- 5, balanced at 5 tasks each
- Quality boundary
- All 25 task records are marked draft
The suite is a small diagnostic baseline, not a statistically representative benchmark. Its content hash is 38c646401b74f92deac04c5e225dfeabd822f72c3e30ce52dfaa24d817375760.
Planned later datasets
- instruction_v1
- instruction_v2_filtered
- preference_v1
- agent_tasks_v1
- safety_eval_v1
Dataset card fields
- name
- purpose
- source
- license/usage constraints
- size
- categories
- filtering steps
- deduplication
- contamination/leakage checks
- risks
- version