LIFEHUBBER
Choose theme

AI Resources

DataDesigner

GitHub stars: 2.3K GitHub forks: 215 Declared license: Apache-2.0: Apache-2.0 Last pushed September 30, 2026: Pushed 1d ago
Stats from GitHub

DataDesigner is a Python framework for generating synthetic datasets from declared columns and dependencies. You can sample one field, use its value in another field’s generation prompt, and preview records before creating a larger dataset.

Generation and validation are separate steps. A validation column records whether a row meets a chosen check; a plausible-looking generated row can still fail that rule. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Build a dataset from columns

Define sampled or model-generated columns, optionally start with seed data, and create synthetic records. This is a configurable generation pipeline rather than a dataset ready to download.

Why it stands out

Let one field shape another

The getting-started example samples a language, then inserts that field into a greeting prompt. This makes the language choice an input to generation instead of sampling a greeting independently of it.

Availability

Python library and configured model endpoints

Install the library, configure the model endpoint and credentials, then preview a small sample. The repository includes examples and plugins; model requests go to the provider you configure.

Why it matters

What makes it useful

To reuse a column workflow with a different seed file, the getting-started guide shows a local Python config accepting a seed-path argument. Pass workflow arguments after -- so they reach that config instead of hardcoding a file path. This forwarding works only for local .py configs; YAML, JSON and remote config URLs reject those arguments, so the config format affects how you can reuse the workflow.

Notable points

What stands out

Validators add is_valid and check-specific metadata. Python code validation uses Ruff linting; SQL validation checks parsing. These checks do not establish that a program produces the intended result. A local callable can express a domain rule, while cross-row validation sees only the current batch, not every row in a full dataset.

Before using

What to review

Configure the model provider before generation. The repository describes NVIDIA Build endpoints as evaluation and testing services and says not to send confidential or personal data to them.

NEMO_TELEMETRY_ENABLED=false disables the library’s optional telemetry. It does not change where your configured model endpoint receives prompts or records.

Remote HTTP validators are currently unauthenticated. The validator guide describes a local proxy for an endpoint that requires authentication; built-in endpoint authentication is not established by that example.

Check the release notes when reusing older configurations: v0.9.3 replaces a retired Nano Build model and updates notebook dependencies. A configuration that names an older model still needs its endpoint checked.

Reader fit

Who may find it relevant

Developers creating structured synthetic inputs for an application or testing workflow.

Data teams defining related fields and explicit validation rules rather than accepting generated rows by appearance alone.

Readers seeking a ready-made dataset or a chat interface will still need a different delivery format.

Editorial note

Why LifeHubber lists it

Try the validator guide’s local price-greater-than-zero rule on a positive price and on known bad examples such as zero or a negative price. Inspect is_valid and error_message for each result before applying the rule to generated rows. This checks whether your chosen validation rule catches a known bad record; a plausible-looking sample alone cannot show that.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving