Choose theme
AI Resources
Monitorability Evals
Monitorability Evals contains OpenAI’s public evaluation data and scoring materials for studying whether a monitor detects features of another model’s behavior.
The repository accompanies the Monitoring Monitorability paper. Its supplied Python scaffold uses mock model and monitor outputs; that run is a worked scoring example, rather than a measurement of your own model. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Evaluation data for model monitoring
A research repository with public evaluation splits, prompt templates, dataset-to-prompt mappings, and metric code for model-monitoring experiments.
Why it stands out
Model outcome and monitor trigger are separate labels
The registry defines Y as the model outcome and Z as the monitor trigger. Their dataset-specific rules make it possible to trace what a positive result means.
Availability
Public repo with omitted restricted splits noted
The README lists included and omitted splits. AIME rows provide source links instead of problem text, and OpenAI marks Anti-Scheming and Memory as no longer used internally pending revised versions.
Why it matters
What makes it useful
For intervention experiments, the README describes two rows sharing an id: one perturbed and one unperturbed. That pair lets a reader connect the added hint or perturbation with its control case, rather than treating each row as an unrelated question. The id is also the grouping key for metric calculation.
What to know
Where it fits
The supplied run_eval_scaffold.py generates mock model and monitor outputs for eligible datasets with embedded prompts. Its documentation excludes memory-retrieval and tool-based evaluations because their interfaces vary across providers, and excludes AIME because those rows lack embedded problem text. A researcher connecting a real model therefore has integration work beyond running this example. The public suite organizes intervention, process and outcome-property evaluations. Inspecting those designs does not certify general model or monitoring safety.
Notable points
What stands out
OpenAI says it no longer uses Anti-Scheming and Memory internally until revised versions address two identified issues: the mismatch between a generic chain-of-thought sketchiness monitor and a task-specific answer grader, and Memory’s answer grading. The README links the GPT-5.4 Thinking system card for the explanation. Their presence in the repository is distinct from OpenAI continuing to use them.
Before using
What to review
Read the selected dataset’s y_rule and z_rule in dataset_registry.yaml alongside its prompt files.
Y labels the model outcome; Z labels whether the monitor fired. A positive monitor result is not automatically a correct model answer.
For process evaluations, the registry defines Y by answer correctness and Z by whether at least one labeled theorem path was fully detected. These are two different things to count.
Reader fit
Who may find it relevant
It fits researchers and builders comparing monitor prompts, graders and evaluation designs. If your experiment needs AIME questions, the public rows contain url and source fields instead of problem text. OpenAI directs readers to those fields to retrieve the corresponding statements from their linked sources. A loader that expects every row to contain problem text needs a separate retrieval step for this data. Before comparing two scaffold summaries, compare their saved config in summary.json: samples_per_row, max_rows_per_dataset, seed, n_bootstrap, selection_frac, and group_keep_frac. The writer records these settings next to dataset counts and metrics, so you can tell whether the runs used the same sampling setup before attributing a score difference to the monitor.
Editorial note
Why LifeHubber lists it
We list Monitorability Evals for researchers checking whether an apparent monitoring change could come from the evaluation's selection method. OpenAI's release post describes selecting eligible intervention instances on one sample fold and estimating the monitoring metric on a separate fold. This keeps the noise that selected an instance from being reused to score it, adding a way to examine selection bias beyond paired rows and saved run settings. It is an evaluation-design lesson, not certification that a model or monitor is safe.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Datasets
Keep browsing this category
Explore more datasets.
ParseBench
run-llama/ParseBench
A document parsing benchmark for AI-agent workflows, focused on whether parsed PDFs preserve structure and meaning for downstream evaluation.
UltraData-SFT-2605
openbmb/UltraData-SFT-2605
An OpenBMB supervised fine-tuning dataset with 15,036,178 thinking and non-thinking samples across math, code, knowledge, Chinese, instruction-following, and multilingual configurations, used in MiniCPM5-1B-SFT post-training.
LARYBench
meituan-longcat/LARYBench
A benchmark for evaluating latent action representations, with pipelines for action semantics, robotic control regression, and broader vision-to-action alignment.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.