LIFEHUBBER
Choose theme

AI Resources

Monitorability Evals

GitHub stars: 102 GitHub forks: 17 Declared license: Apache-2.0: Apache-2.0 Last pushed August 17, 2026: Pushed 1mo ago
Stats from GitHub

Monitorability Evals contains OpenAI’s public evaluation data and scoring materials for studying whether a monitor detects features of another model’s behavior.

The repository accompanies the Monitoring Monitorability paper. Its supplied Python scaffold uses mock model and monitor outputs; that run is a worked scoring example, rather than a measurement of your own model. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Evaluation data for model monitoring

A research repository with public evaluation splits, prompt templates, dataset-to-prompt mappings, and metric code for model-monitoring experiments.

Why it stands out

Model outcome and monitor trigger are separate labels

The registry defines Y as the model outcome and Z as the monitor trigger. Their dataset-specific rules make it possible to trace what a positive result means.

Availability

Public repo with omitted restricted splits noted

The README lists included and omitted splits. AIME rows provide source links instead of problem text, and OpenAI marks Anti-Scheming and Memory as no longer used internally pending revised versions.

Why it matters

What makes it useful

For intervention experiments, the README describes two rows sharing an id: one perturbed and one unperturbed. That pair lets a reader connect the added hint or perturbation with its control case, rather than treating each row as an unrelated question. The id is also the grouping key for metric calculation.

Notable points

What stands out

OpenAI says it no longer uses Anti-Scheming and Memory internally until revised versions address two identified issues: the mismatch between a generic chain-of-thought sketchiness monitor and a task-specific answer grader, and Memory’s answer grading. The README links the GPT-5.4 Thinking system card for the explanation. Their presence in the repository is distinct from OpenAI continuing to use them.

Before using

What to review

Read the selected dataset’s y_rule and z_rule in dataset_registry.yaml alongside its prompt files.

Y labels the model outcome; Z labels whether the monitor fired. A positive monitor result is not automatically a correct model answer.

For process evaluations, the registry defines Y by answer correctness and Z by whether at least one labeled theorem path was fully detected. These are two different things to count.

Reader fit

Who may find it relevant

It fits researchers and builders comparing monitor prompts, graders and evaluation designs. If your experiment needs AIME questions, the public rows contain url and source fields instead of problem text. OpenAI directs readers to those fields to retrieve the corresponding statements from their linked sources. A loader that expects every row to contain problem text needs a separate retrieval step for this data. Before comparing two scaffold summaries, compare their saved config in summary.json: samples_per_row, max_rows_per_dataset, seed, n_bootstrap, selection_frac, and group_keep_frac. The writer records these settings next to dataset counts and metrics, so you can tell whether the runs used the same sampling setup before attributing a score difference to the monitor.

Editorial note

Why LifeHubber lists it

We list Monitorability Evals for researchers checking whether an apparent monitoring change could come from the evaluation's selection method. OpenAI's release post describes selecting eligible intervention instances on one sample fold and estimating the monitoring metric on a separate fold. This keeps the noise that selected an instance from being reused to score it, adding a way to examine selection bias beyond paired rows and saved run settings. It is an evaluation-design lesson, not certification that a model or monitor is safe.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving