Theme
AI Resources
UltraData-SFT-2605
UltraData-SFT-2605 is an OpenBMB supervised fine-tuning dataset used in the post-training of MiniCPM5-1B-SFT. It contains 15,036,178 samples across math, code, knowledge, Chinese-general, instruction-following, multilingual math, and multilingual knowledge configurations.
Most configurations separate Deep Thinking examples from direct, Non-thinking responses. The dataset card also documents its filtering, answer-quality checks, training-based validation, and benchmark-decontamination process. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Post-training data for several core capabilities
The release groups supervised examples into seven configurations so builders can inspect or select math, code, knowledge, Chinese, instruction-following, and multilingual slices instead of treating the collection as one undifferentiated file.
Why it stands out
Thinking mode is visible in the splits
Math, code, knowledge, Chinese-general, and instruction-following include both think and no_think splits. The two multilingual configurations are released as no_think only, making the intended response style easier to inspect before training.
Availability
Large, gated, and not free of upstream conditions
The public repository lists about 319 GB of files, but access requires a Hugging Face account and agreement to share contact information. The dataset card also sets dataset-specific conditions and points to upstream terms; review the current wording at the official source before use or redistribution.
Why it matters
What makes it useful
Training-data releases are more useful when readers can see both the capability mix and the intended answer style. UltraData-SFT-2605 exposes domain-level configurations and thinking-mode splits, giving researchers a more inspectable starting point for selecting or comparing post-training data.
What to know
Where it fits
UltraData-SFT-2605 fits teams preparing supervised fine-tuning mixtures, comparing thinking with direct-response data, or tracing the data behind MiniCPM5-1B-SFT. It is a dataset to inspect and select from, not a ready-to-run model or training tool.
Notable points
What stands out
The publisher describes six management stages: validation and filtering of public queries, internal query construction, filtering of pre-training-format material, answer-quality filtering, training-based validation, and benchmark decontamination. These are publisher-reported procedures, not an independent audit of every sample.
Before using
What to review
The gated-access form and what contact information will be shared before requesting the files.
The terms of every upstream dataset used in a chosen configuration, not only the Apache-2.0 label shown on the repository.
The dataset card's additional conditions, reviewed in their current official wording rather than inferred from the repository's Apache-2.0 label.
Whether roughly 319 GB of downloads plus local processing, storage, deduplication, and training costs fit the project.
Potential contamination, harmful content, personal data, copyright, language imbalance, and answer-quality issues that may remain despite the publisher's filtering process.
Which configurations and think or no_think splits match the intended behavior before mixing them into a training run.
Reader fit
Who may find it relevant
Researchers and engineers preparing supervised fine-tuning data for language models.
Teams comparing thinking and direct-response examples across math, code, knowledge, or instruction-following tasks.
Readers tracing the post-training data used for MiniCPM5-1B-SFT.
Less relevant for people seeking a ready-to-run model, a small sample dataset, or data with simple unrestricted access.
Editorial note
Why LifeHubber lists it
UltraData-SFT-2605 is useful because it makes two training-data decisions inspectable: which capability slice is being selected, and whether the examples teach extended reasoning or a direct response. Its size, gated access, and layered usage terms should be evaluated just as carefully as its coverage.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Connect the dataset to a model and training workflow.
UltraData-SFT-2605 supplies supervised examples rather than a runnable model or trainer. Continue with the model it supported or inspect a separate fine-tuning workflow before planning a training run.
More in Datasets
Keep browsing this category
Explore more datasets.
ParseBench
run-llama/ParseBench
A document parsing benchmark for AI-agent workflows, focused on whether parsed PDFs preserve structure and meaning for downstream evaluation.
Terminal-Bench 2.0
harbor-framework/terminal-bench-2
A terminal-agent benchmark for evaluating AI agents on hard containerized command-line tasks, with Harbor run commands, task-level registry pages, GitHub and Hugging Face materials, docs, and paper links.
Monitorability Evals
openai/monitorability-evals
An OpenAI evaluation-data release for studying monitorability, with public eval splits, prompt templates, dataset mappings, and metric code from the Monitoring Monitorability paper.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.