Choose theme
AI Resources
OpenAI Privacy Filter
OpenAI Privacy Filter detects and masks personally identifiable information in text. Its repository includes a local command-line toolkit for redaction, evaluation and finetuning.
The model labels tokens and decodes spans rather than answering questions. OpenAI describes it as a data-minimization aid, with trained categories and detection limits rather than an anonymization guarantee. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
A text detection and masking model
The model identifies spans such as private names, email addresses, dates and secrets. The CLI can process an example, a file or piped text.
Why it stands out
Decoding controls and policy training
Runtime decoding controls adjust the balance between missed spans and excessive masking. Changing the trained category policy requires finetuning.
Availability
Repository, weights and local CLI
OpenAI publishes the code, model weights, output schemas and evaluation and finetuning examples. The repository documents installation from its checkout.
Why it matters
What makes it useful
A name inside a sentence needs different handling from the sentence around it. Privacy Filter returns detected span boundaries and labels, so a text-processing workflow can replace the identified portions while retaining the surrounding text. Missed spans remain possible.
What to know
Where it fits
The CLI runs against a local checkpoint, but a first run can download the model when the default checkpoint is missing. The README documents OPF_CHECKPOINT and an explicit checkpoint path; on-premises execution and obtaining the weights are separate setup steps.
Notable points
What stands out
The evaluation guide separates matching categories from matching spans. With gold labels in its own taxonomy, typed mode reports category-level metrics. With a different taxonomy, untyped mode ignores category identity during matching and reports ground_truth_label_recall for the original labels. The same span detection can therefore be assessed without pretending the label names agree.
Before using
What to review
The model is primarily English. Its documented limitations include non-English text, unfamiliar names, novel secret formats and shifted span boundaries.
The publisher also describes excess masking of public entities, common nouns and benign strings. A masked result can lose useful context as well as leave sensitive text behind.
OpenAI presents it as one minimization layer, not anonymization, compliance or a safety guarantee; its guidance calls for in-domain evaluation and human review paths in sensitive workflows.
Reader fit
Who may find it relevant
Builders changing what counts as a private category have a training task, beyond choosing a broader masking operating point. The finetuning guide supports a custom label-space file and a separate validation dataset. Its tiny toy demos illustrate a category-policy change; they are not evidence that the resulting checkpoint fits a real text workflow.
Editorial note
Why LifeHubber lists it
A redacted preview and its JSON payload contain different things. The documented output includes original text and detected-span text alongside redacted_text; the evaluation predictions export also retains original text. A downstream consumer choosing redacted_text is making a different transfer from forwarding the whole payload. Masking the preview does not establish that the exported record has removed the original.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Ecosystem
Keep browsing this category
Explore more AI ecosystem resources.
Laya
NandhaKishorM/laya
An early Apache-2.0 model family and Python runtime for bounded choice, score, and yes-or-no decisions, with English, multilingual, and task-specialized checkpoints plus a router that selects between them.
AnyJev
nokia-applied-research/AnyJev
An Apache-2.0 Python framework that turns existing Transformers or vLLM-served language models into bounded choice, yes-or-no, and score decisions, with levels for zero-label bias correction, calibration, and question-specific decision heads.
Manifest
mnfst/manifest
An open-source model gateway for AI agents and apps, positioned around cost-aware routing, fallbacks, provider control, and self-hosted workflows.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.